Video stream data dynamic processing method and system based on memory calculation
By dividing dynamic memory pools in memory, building a space-time graph and using policy network decision-making processing actions, the problem that traditional video stream processing technology cannot cope with complex scenarios and changes in real time is solved, and the flexibility and efficiency of video stream processing is achieved, and video quality and resource utilization are improved.
Patent Information
- Application Number
- CN202510085484.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-20
AI Technical Summary
Traditional video stream processing technology cannot make real-time decisions based on the actual processing status of the video stream, and cannot flexibly respond to complex scenarios and changes, resulting in improper resource allocation, affecting the stable transmission and high-quality playback of the video stream.
A dynamic processing method for video stream data based on memory computing is proposed. By dividing a dynamic memory pool in memory, obtaining video stream data in real time, extracting frame-level features and optical flow information of keyframes, building a spatio-temporal map, analyzing the spatio-temporal domain characteristics, and processing actions based on the current state through a policy network, including adjusting the compression ratio, allocating computing resources and dynamically adjusting the transmission protocol.
It realizes dynamic adjustment of memory allocation according to the real-time requirements of video streams, improves the flexibility and efficiency of video stream processing, can process video data more accurately, improves video quality, and reduces the risk of resource waste and insufficient resources.
Smart Images

Figure CN120014512A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video processing technology, and in particular to a method and system for dynamically processing video stream data based on memory computing. Background Art
[0002] In traditional video stream processing technology, video data processing mainly relies on fixed processing strategies and manual rules. Usually, video streams are encoded, compressed, transmitted, and played based on preset algorithms and parameters.
[0003] However, this traditional processing technology has many defects: it is unable to make real-time decisions based on the actual processing status of the video stream, and is unable to flexibly respond to various complex scenarios and changes. For example, when the network conditions are poor, it is unable to automatically adjust the transmission protocol or compression ratio to ensure stable transmission and high-quality playback of the video stream; it often adopts fixed processing parameters and strategies, and is unable to dynamically allocate computing resources and network resources according to actual needs, which either makes the resources not fully utilized, resulting in waste, or makes the resources insufficient, affecting the processing effect of the video stream; for video streams with different resolutions and bit rates, different processing algorithms and parameters may need to be developed separately, which increases the development cost and maintenance difficulty; unstable quality may occur when processing video streams. For example, when the network conditions fluctuate, the video stream may experience problems such as freezes, screen distortion, or audio and video asynchrony, affecting the user experience. Summary of the invention
[0004] Based on this, the purpose of the present invention is to propose a method and system for dynamic processing of video stream data based on memory computing to solve the above-mentioned problems.
[0005] According to a method for dynamically processing video stream data based on memory computing proposed by the present invention, the method comprises:
[0006] A dynamic memory pool is allocated in the memory for storing and processing video stream data;
[0007] Get video stream data in real time and split it into individual frames;
[0008] Extract frame-level features of key frames;
[0009] Calculate optical flow for key frames;
[0010] Combining frame-level features with optical flow information to construct a spatiotemporal graph, and analyzing the spatiotemporal graph to extract spatiotemporal features;
[0011] The processing state of the video stream is represented as a state vector, a policy network is trained, and the next processing action is decided by the policy network according to the current state, wherein the processing state includes processing speed, memory usage, number of abnormal frames and spatiotemporal domain characteristics, and the processing actions include adjusting compression ratio, allocating computing resources, adjusting cache strategy and dynamically adjusting transmission protocol.
[0012] Furthermore, calculating the optical flow of the key frame includes:
[0013] Select feature points in keyframes;
[0014] Calculate the optical flow of feature points between adjacent key frames by using an optical flow algorithm, wherein the optical flow is used to represent the displacement vector of the feature points between consecutive key frames;
[0015] An optical flow feature is extracted from the optical flow data, wherein the optical flow feature includes the magnitude and direction of the optical flow vector.
[0016] Furthermore, the frame-level features and optical flow information are combined to construct a spatiotemporal graph, including:
[0017] Treat the key frame as a node in the spatiotemporal graph, where each node contains the frame-level features of the key frame;
[0018] The edges in the spatiotemporal graph are constructed according to the optical flow information between each pair of adjacent key frames, where each edge connects the nodes of adjacent key frames, and the weight or attribute of the edge is set as the optical flow feature.
[0019] Furthermore, the analyzing the spatiotemporal graph to extract spatiotemporal features includes:
[0020] The spatiotemporal graph is represented and learned through the graph representation learning algorithm to obtain a graph representation learning model;
[0021] Spatiotemporal features are extracted from the graph representation learning model, and the spatiotemporal features are used to represent the main change information of the video stream in two dimensions of time and space.
[0022] Furthermore, the processing state of the video stream is represented as a state vector, a policy network is trained, and the next processing action is decided by the policy network according to the current state, including:
[0023] Define the state vector S according to the processing state of the video stream t , wherein the processing status includes processing speed, memory usage, number of abnormal frames and spatiotemporal characteristics, which are used to describe the environmental status of the video stream processing system at the current moment or the status of the video stream itself;
[0024] An action space is defined according to processing actions of the video stream, wherein the processing actions include adjusting compression ratio, allocating computing resources, adjusting cache strategy, and dynamically adjusting transmission protocol, and each processing action is represented as a;
[0025] The neural network structure is used as a policy network, which receives the state vector as input and outputs the probability distribution of the action to be taken in the current state. The policy network is represented by π θ (a∣S t ), where θ is the parameter of the policy network, a is the action, S t is the current state;
[0026] Defining a reward function R(s, a) according to a processing objective for evaluating the quality of taking action a in state s, wherein the processing objective includes optimizing video quality, improving processing speed, reducing memory usage, and / or reducing abnormal frames;
[0027] The policy network is trained using reinforcement learning. During the training process, the expected cumulative reward is maximized by optimizing the policy network parameters θ, and the policy network is guided to automatically adjust the processing actions according to the current processing status by combining the evaluation of the value network. Through continuous trial and error and optimization, the optimal processing strategy is learned to improve the overall performance of video stream processing.
[0028] Apply the trained strategy network to decide the next processing action according to the current state, and the processing action includes:
[0029] Furthermore, the reinforcement learning is used to train the policy network. During the training process, the expected cumulative reward is maximized by optimizing the policy network parameters θ, and the policy network is guided to automatically adjust the processing action according to the current processing state in combination with the evaluation of the value network. Through continuous trial and error and optimization, the optimal processing strategy is learned to improve the overall performance of video stream processing, including:
[0030] Randomly initialize the parameters θ of the policy network;
[0031] Collect a sample of states, actions, and rewards;
[0032] Calculate the gradient of the objective function J(θ) with respect to the policy network parameters The calculation formula is:
[0033]
[0034] Among them, Q is estimated approximately through the value network πθ (S t , a), Q πθ (S t , a) means that under the strategy πθ, from the state vector S tThe expected cumulative reward after taking action a, E πθ [] represents the expectation under the strategy πθ, represents the logarithmic gradient of the policy network output action probability, π θ (a∣S t ) represents the policy network in the state vector S t The probability of taking action a under
[0035] Update the parameters θ of the policy network by the gradient ascent method. The update formula is:
[0036]
[0037] Among them, α is the learning rate, which controls the step size of parameter update;
[0038] Repeat the training until the policy network converges to obtain a trained policy network, which is used to improve the overall performance of video stream processing.
[0039] Furthermore, the value network is used to approximate Q πθ (S t , a), including:
[0040] Construct a value network based on a neural network, represented by Q ω (s, a), where ω is the parameter of the value network, the input of the value network is state s and action a, and the output is the Q value estimate of taking action a in state s;
[0041] Train the value network to make the Q value output by the value network close to the actual expected cumulative reward;
[0042] During the training process, the state vector S is propagated forward. t And action a is input into the value network to obtain the Q value estimate Q ω (S t , a);
[0043] Estimate the Q value Q ω (S t , a) as Q πθ (S t , an approximate estimate of a).
[0044] Furthermore, the application of the trained policy network determines the next processing action according to the current state, including:
[0045] Get the processing speed, memory usage, number of abnormal frames, and spatiotemporal features of the current video stream and combine them into a state vector S t ;
[0046] The state vector S tInput into the trained policy network to obtain the probability distribution π of the action to be taken in the current state θ (a∣S t );
[0047] Choose an action based on a probability distribution.
[0048] The present invention also proposes a video stream data dynamic processing system based on memory computing, which is used to implement the above-mentioned video stream data dynamic processing method based on memory computing. The system includes:
[0049] Partitioning module: used to divide a dynamic memory pool in the memory for storing and processing video stream data;
[0050] Data acquisition module: used to acquire video stream data in real time and split it into separate frames;
[0051] Feature extraction module: used to extract frame-level features of key frames;
[0052] Optical flow calculation module: used to calculate the optical flow of key frames;
[0053] Spatiotemporal feature module: used to combine frame-level features with optical flow information to construct a spatiotemporal graph, and analyze the spatiotemporal graph to extract spatiotemporal features;
[0054] Action decision module: used to represent the processing state of the video stream as a state vector, train a policy network, and decide the next processing action based on the current state through the policy network, wherein the processing state includes processing speed, memory usage, number of abnormal frames and spatiotemporal domain characteristics, and the processing action includes adjusting the compression ratio, allocating computing resources, adjusting the cache strategy and dynamically adjusting the transmission protocol.
[0055] In summary, the video stream data dynamic processing method based on memory computing of the present invention divides a dynamic memory pool in the memory, which is specifically used to store and process video stream data. The divided dynamic memory pool can dynamically adjust memory allocation according to the real-time demand of video stream data, thereby avoiding the problem of memory waste or shortage caused by fixed memory allocation.
[0056] Combine the frame-level features of key frames with optical flow information to construct a spatiotemporal graph and extract spatiotemporal features. Optical flow information is used to accurately capture the key changes in the video content to enhance the spatiotemporal graph's ability to express the dynamic content of the video. The spatiotemporal graph combines information in the time domain and space domain, so that the extracted spatiotemporal features can more accurately reflect the key changes in the video.
[0057] By quantifying the processing state of the video stream into a state vector, the policy network can fully and accurately perceive the current processing environment and conditions as well as the state of the video itself. The policy network makes intelligent decisions on processing actions based on the state vector and can quickly output the optimal processing action, avoiding decision delays and inefficiencies caused by reliance on artificial rules or fixed strategies. This intelligent decision-making mechanism enables the memory system to more efficiently respond to various complex video stream processing scenarios and changes, improve the overall processing efficiency and performance of the memory system, and more accurately process video data, improving video quality.
[0058] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description or will be learned through embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0060] Figure 1 This is a flow chart of a method for dynamically processing video stream data based on memory computing according to the first embodiment of the present invention;
[0061] Figure 2 This is a system block diagram of a video stream data dynamic processing system based on memory computing according to Embodiment 2 of the present invention. DETAILED DESCRIPTION
[0062] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0063] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items.
[0065] Embodiment 1
[0066] See also Figure 1 The present invention proposes a method for dynamically processing video stream data based on memory computing, the method comprising steps S101 to S106:
[0067] S101, a dynamic memory pool is allocated in the memory for storing and processing video stream data.
[0068] It should be noted that a dynamic memory pool is divided in the memory to efficiently store and process video stream data. The dynamic memory pool can dynamically adjust memory allocation according to the real-time needs of video stream data. Compared with static memory, it is more flexible and can avoid memory waste or insufficient problems caused by allocation.
[0069] The dynamic memory pool can be implemented by managing a series of memory blocks, each of which can store a certain size of video stream data. When new data needs to be stored, the dynamic memory pool will allocate a suitable memory block; when the data is no longer needed, the memory block will be recycled and added back to the memory pool.
[0070] S102, acquiring video stream data in real time and dividing it into separate frames.
[0071] It should be noted that the video stream data may come from various cameras, network video streams or existing video files.
[0072] Video stream data is composed of a series of continuous image frames, each of which contains information that constitutes the video picture. The system obtains data from the video stream in real time and divides the data into individual frames, which usually involves decoding the video stream data and then dividing the decoded data into individual frames according to time intervals (such as 24 frames per second, 30 frames per second, etc.). The divided frames need to be stored in a memory pool for subsequent processing and analysis.
[0073] S103: extract frame-level features of key frames.
[0074] It should be noted that frame-level features refer to representative information extracted from a single video frame. This information can reflect the content, structure or motion characteristics of the frame. Frame-level features may include color features, texture features, shape features, etc.
[0075] In the key frame extraction process, the frame-level features of the video frames can be extracted and the feature similarity between frames can be calculated to identify frames with significant content changes. These frames can be selected as key frames because they can better represent the main content of the video clip. For example, the color histogram can be used to calculate the color similarity between frames, or the optical flow method can be used to calculate the motion similarity between frames.
[0076] S104, calculating the optical flow for the key frame.
[0077] It should be noted that optical flow information is used to accurately capture motion information in videos, which can enhance the ability of spatiotemporal graphs to express dynamic content in videos. Calculating the optical flow between key frames focuses more on the key changes in the video content, thereby reducing unnecessary computational overhead and extracting effective motion information. This is particularly useful when processing long videos or real-time video streams, and can significantly improve processing efficiency and accuracy.
[0078] Further optionally, calculating the optical flow for the key frame includes:
[0079] Select feature points in keyframes;
[0080] Calculate the optical flow of feature points between adjacent key frames by using an optical flow algorithm, wherein the optical flow is used to represent the displacement vector of the feature points between consecutive key frames;
[0081] An optical flow feature is extracted from the optical flow data, wherein the optical flow feature includes the magnitude and direction of the optical flow vector.
[0082] Understandably, key frames are identified in the video stream, which can be frames where the video content changes significantly, or frames selected according to specific strategies (such as content difference thresholds, etc.). In the selected key frames, representative or easy-to-track feature points are selected, which will be used for tracking between adjacent key frames.
[0083] Optical flow algorithms (such as Lucas-Kanade method, Farneback method, etc.) are applied to calculate the optical flow of feature points between adjacent key frames. Optical flow represents the displacement vector of feature points between consecutive key frames, reflecting the motion information of objects between key frames.
[0084] The calculated optical flow data is analyzed to extract optical flow features, including the size (indicating the speed of movement) and direction of the optical flow vector. These optical flow features will be used for subsequent spatiotemporal graph construction and analysis to extract information about changes in the video stream in both time and space.
[0085] S105, combining the frame-level features and the optical flow information to construct a spatiotemporal graph, and analyzing the spatiotemporal graph to extract spatiotemporal domain features;
[0086] It should be noted that after extracting the frame-level features of the key frames and calculating the optical flow of the key frames, the frame-level features and optical flow information are combined to construct a spatiotemporal graph of the key frames. By analyzing the spatiotemporal graph based on the key frames, it is possible to extract important change information of the video stream in both time and space dimensions, namely, spatiotemporal features. By analyzing the spatiotemporal features of the key frames, the system can identify anomalies or key events in the video stream, such as the sudden appearance of objects, scene switching, etc., which is of great significance for triggering corresponding processing actions (such as adjusting the compression ratio, reallocating computing resources, etc.).
[0087] Further optionally, combining the frame-level features and the optical flow information to construct a spatiotemporal graph includes:
[0088] Treat the key frame as a node in the spatiotemporal graph, where each node contains the frame-level features of the key frame;
[0089] The edges in the spatiotemporal graph are constructed according to the optical flow information between each pair of adjacent key frames, where each edge connects the nodes of adjacent key frames, and the weight or attribute of the edge is set as the optical flow feature.
[0090] It can be understood that key frames are identified in the video stream, and each key frame is used as a node in the spatiotemporal graph, and each node contains frame-level features of the corresponding key frame image, such as color histogram, edge features, texture features, etc.
[0091] For each pair of adjacent keyframes, the edges in the spatiotemporal graph are constructed based on the optical flow information between them. Each edge connects the nodes of adjacent keyframes, and the weight or attribute of the edge is set to the optical flow feature, such as the magnitude and direction of the optical flow vector. These optical flow features reflect the motion information and degree of change of the object between adjacent keyframes.
[0092] The spatiotemporal graph constructed in this embodiment focuses more on the key changes in the video content, reduces the impact of non-key frames on the complexity of the spatiotemporal graph, and retains important change information of the video stream in both time and space dimensions. It is particularly effective when processing videos with significant motion or content changes, and can improve the accuracy and efficiency of spatiotemporal graph analysis.
[0093] Further optionally, analyzing the space-time graph to extract space-time domain features includes:
[0094] The spatiotemporal graph is represented and learned through the graph representation learning algorithm to obtain a graph representation learning model;
[0095] Spatiotemporal features are extracted from the graph representation learning model, and the spatiotemporal features are used to represent the main change information of the video stream in two dimensions of time and space, focusing on key changes and anomaly detection.
[0096] It is understandable that the graph attention network can be used to learn the representation of the key frame's spatiotemporal graph to obtain a graph representation learning model. The graph attention network dynamically allocates the importance of different edges through the attention mechanism, flexibly captures the structural information of the graph, and can process the node features, edge features and structural information of the spatiotemporal graph, thereby generating a graph representation learning model that fully reflects the characteristics of the spatiotemporal graph. Specifically, it includes:
[0097] The feature vectors of all nodes are combined into a feature matrix as the input of the graph attention layer; for each node, the attention weight between it and its neighboring nodes is calculated; according to the calculated attention weight, the features of the neighboring nodes are weighted summed to aggregate the information of the neighboring nodes. Multiple independent attention heads can be used to calculate the attention weights and weighted feature aggregation in parallel; the output results of multiple attention heads are merged; by stacking multiple graph attention layers, nodes can fuse more levels of neighboring node information, thereby learning deeper graph structure features.
[0098] Based on the graph representation learning model, spatiotemporal features are extracted, which may include the dynamic change pattern of nodes, the connection strength and time evolution of edges, the structure and evolution of subgraphs, etc. Through these features, the spatiotemporal correlation and changes between key frames in the video stream can be captured.
[0099] S106, representing the processing state of the video stream as a state vector, training a policy network, and using the policy network to decide the next processing action based on the current state, wherein the processing state includes processing speed, memory usage, number of abnormal frames, and spatiotemporal domain features, etc., and the processing actions include adjusting compression ratio, allocating computing resources, adjusting cache strategy, and dynamically adjusting transmission protocol, etc.
[0100] It should be noted that various state information in the video stream processing process is quantified to form a state vector. The processing state includes processing speed (indicating the current processing frame rate or processing delay of the video stream), memory usage (reflecting the system memory occupancy rate or remaining memory), number of abnormal frames (the number of abnormal frames detected in the past period of time, such as screen freezes, screen noise, etc.) and spatiotemporal characteristics of key frames (the characteristics of key frames in time and space, such as color distribution, texture features, motion trajectory, optical flow, etc.).
[0101] The processing state of the video stream is represented as a state vector, the policy network is trained, and the policy network decides the best processing action for the next step based on the current state (for example, when the memory usage is high and the number of abnormal frames increases, choose to reduce the compression ratio to retain more details; when the processing speed is slow and the motion trajectory is complex, choose to allocate more computing resources to speed up processing). This intelligent decision-making processing method for video streams can dynamically adapt to various changes in the video stream processing process and improve the flexibility and efficiency of the system.
[0102] Among them, the policy network is a deep learning model, such as a convolutional neural network combined with a recurrent neural network or a long short-term memory network, which can adapt to the high dimensionality and temporal nature of the state vector.
[0103] In the decision-making process, the current state vector is input into the policy network, and the network outputs one or more possible processing actions based on the learned mapping relationship. The system executes one of the actions according to a certain strategy (such as selecting the action with the highest probability, considering the cost and benefit of the action, etc.).
[0104] Further optionally, the processing state of the video stream is represented as a state vector, a policy network is trained, and the next processing action is decided according to the current state by the policy network, including:
[0105] Define the state vector S according to the processing state of the video stream t , wherein the processing status includes processing speed, memory usage, number of abnormal frames and spatiotemporal characteristics, which are used to describe the environmental status of the video stream processing system at the current moment or the status of the video stream itself;
[0106] An action space is defined according to processing actions of the video stream, wherein the processing actions include adjusting compression ratio, allocating computing resources, adjusting cache strategy, and dynamically adjusting transmission protocol, and each processing action is represented as a;
[0107] The neural network structure is used as a policy network, which receives the state vector as input and outputs the probability distribution of the action to be taken in the current state. The policy network is represented by π θ (a∣S t ), where θ is the parameter of the policy network, a is the action, and S t is the current state;
[0108] Defining a reward function R(s, a) according to a processing objective for evaluating the quality of taking action a in state s, wherein the processing objective includes optimizing video quality, improving processing speed, reducing memory usage, and / or reducing abnormal frames;
[0109] The policy network is trained using reinforcement learning. During the training process, the expected cumulative reward is maximized by optimizing the policy network parameters θ, and the policy network is guided to automatically adjust the processing actions according to the current processing status by combining the evaluation of the value network. Through continuous trial and error and optimization, the optimal processing strategy is learned to improve the overall performance of video stream processing.
[0110] Apply the trained strategy network to decide the next processing action according to the current state, and the processing action includes:
[0111] It is understandable that the present invention uses a reinforcement learning method to train the policy network, and combines the evaluation guidance of the value network, through continuous trial and error and optimization, maximizes the cumulative reward expectation, so that the policy network can quickly and automatically adjust the processing action according to the current processing state, and learn the optimal processing strategy. Apply the trained policy network to decide the next processing action according to the current state. These processing actions include adjusting the compression ratio, allocating computing resources, adjusting the cache strategy, and dynamically adjusting the transmission protocol, so as to optimize the overall performance of video stream processing in real time.
[0112] Further optionally, the policy network is trained using reinforcement learning. During the training process, the expectation of cumulative reward is maximized by optimizing the parameter θ of the policy network, and the policy network is guided to automatically adjust the processing action according to the current processing state in combination with the evaluation of the value network. The optimal processing strategy is learned through continuous trial and error and optimization to improve the overall performance of video stream processing, including:
[0113] Randomly initialize the parameters θ of the policy network;
[0114] You can collect a series of samples of states, actions, and rewards by simulating or actually running the video stream processing system.
[0115] Calculate the gradient of the objective function J(θ) with respect to the policy network parameters The calculation formula is:
[0116]
[0117] Among them, Q is estimated approximately through the value network πθ (S t , a), Q πθ (S t , a) means that under the strategy πθ, from the state vector S t The expected cumulative reward after taking action a, E πθ [] represents the expectation under the strategy πθ, represents the logarithmic gradient of the policy network output action probability, π θ (a∣S t ) represents the policy network in the state vector St The probability of taking action a under
[0118] Update the parameters θ of the policy network by the gradient ascent method. The update formula is:
[0119]
[0120] Among them, α is the learning rate, which controls the step size of parameter update;
[0121] Repeat the training until the policy network converges to obtain a trained policy network, which is used to improve the overall performance of video stream processing.
[0122] Optionally, the value network is used to approximate Q πθ (S t , a), including:
[0123] Construct a value network based on a neural network, represented by Q ω (s, a), where ω is the parameter of the value network, the input of the value network is state s and action a, and the output is the Q value estimate of taking action a in state s;
[0124] Train the value network, with the goal of making the Q value output by the value network as close as possible to the true expected cumulative reward;
[0125] During the training process, the state vector S is propagated forward. t And action a is input into the value network to obtain the Q value estimate Q ω (S t , a);
[0126] Estimate the Q value Q ω (S t , a) as Q πθ (S t , an approximate estimate of a).
[0127] It is understandable that the value network of the present invention is constructed based on a neural network, has a strong nonlinear fitting ability, and can accurately estimate the Q value in a complex environment, that is, the expected cumulative reward of the state-action pair. It not only quickly calculates the Q value estimate through forward propagation, improves the training efficiency, but also handles the problem of high-dimensional state and action space well. At the same time, it provides a stable gradient estimate for the policy network, which helps to prevent oscillation and divergence during training. It also defines an optimization goal for the policy network, which promotes the policy network to gradually learn a better processing strategy by maximizing the Q value, thereby significantly improving the overall performance of video stream processing.
[0128] Further optionally, the application of the trained policy network to decide the next processing action according to the current state includes:
[0129] Get the processing speed, memory usage, number of abnormal frames, and spatiotemporal features of the current video stream and combine them into a state vector S t ;
[0130] The state vector S t Input into the trained policy network to obtain the probability distribution π of the action to be taken in the current state θ (a∣S t );
[0131] Choose an action based on a probability distribution.
[0132] Understandably, in the process of applying the trained policy network to decide the next processing action, first of all, the key processing indicators of the current video stream are obtained, including processing speed, memory usage, number of abnormal frames, and spatiotemporal characteristics. These indicators together constitute the state vector that describes the current video stream processing status, and fully reflects the real-time status of the video stream processing system.
[0133] Next, this state vector is input into the trained policy network, which processes the input state vector based on the strategy and parameters learned internally, and outputs the probability distribution of each action to be taken in the current state. This probability distribution reflects the probability of obtaining the expected reward under different actions.
[0134] Finally, based on this probability distribution, an action can be selected to execute. This selection process can be based on a greedy strategy, an ε-greedy strategy, or other exploration and utilization balance strategies to ensure that the system can not only utilize the currently known optimal strategy, but also explore new possible strategies to a certain extent, thereby continuously optimizing and improving the selection of processing actions.
[0135] The present invention provides an efficient and adaptive processing action selection mechanism for a video stream processing system by executing intelligent decision-making based on reinforcement learning.
[0136] In summary, the video stream data dynamic processing method based on memory computing of the present invention divides a dynamic memory pool in the memory, which is specifically used to store and process video stream data. The divided dynamic memory pool can dynamically adjust memory allocation according to the real-time demand of video stream data, thereby avoiding the problem of memory waste or shortage caused by fixed memory allocation.
[0137] Combine the frame-level features of key frames with optical flow information to construct a spatiotemporal graph and extract spatiotemporal features. Optical flow information is used to accurately capture the key changes in the video content to enhance the spatiotemporal graph's ability to express the dynamic content of the video. The spatiotemporal graph combines information in the time domain and space domain, so that the extracted spatiotemporal features can more accurately reflect the key changes in the video.
[0138] By quantifying the processing state of the video stream into a state vector, the policy network can fully and accurately perceive the current processing environment and conditions as well as the state of the video itself. The policy network makes intelligent decisions on processing actions based on the state vector and can quickly output the optimal processing action, avoiding decision delays and inefficiencies caused by reliance on artificial rules or fixed strategies. This intelligent decision-making mechanism enables the memory system to more efficiently respond to various complex video stream processing scenarios and changes, improve the overall processing efficiency and performance of the memory system, and more accurately process video data, improving video quality.
[0139] Embodiment 2
[0140] See also Figure 2 The present invention proposes a video stream data dynamic processing system based on memory computing, the system comprising:
[0141] Partitioning module: used to divide a dynamic memory pool in the memory for storing and processing video stream data;
[0142] Data acquisition module: used to acquire video stream data in real time and split it into separate frames;
[0143] Feature extraction module: used to extract frame-level features of key frames;
[0144] Optical flow calculation module: used to calculate the optical flow of key frames;
[0145] Spatiotemporal feature module: used to combine frame-level features with optical flow information to construct a spatiotemporal graph, and analyze the spatiotemporal graph to extract spatiotemporal features;
[0146] Action decision module: used to represent the processing state of the video stream as a state vector, train a policy network, and decide the next processing action based on the current state through the policy network, wherein the processing state includes processing speed, memory usage, number of abnormal frames and spatiotemporal domain characteristics, and the processing action includes adjusting the compression ratio, allocating computing resources, adjusting the cache strategy and dynamically adjusting the transmission protocol.
[0147] Further optionally, the optical flow calculation module is also used for:
[0148] Select feature points in keyframes;
[0149] Calculate the optical flow of feature points between adjacent key frames by using an optical flow algorithm, wherein the optical flow is used to represent the displacement vector of the feature points between consecutive key frames;
[0150] An optical flow feature is extracted from the optical flow data, wherein the optical flow feature includes the magnitude and direction of the optical flow vector.
[0151] Further optionally, the spatiotemporal feature module is also used for:
[0152] Treat the key frame as a node in the spatiotemporal graph, where each node contains the frame-level features of the key frame;
[0153] The edges in the spatiotemporal graph are constructed according to the optical flow information between each pair of adjacent key frames, where each edge connects the nodes of adjacent key frames, and the weight or attribute of the edge is set as the optical flow feature.
[0154] Further optionally, the spatiotemporal feature module is also used for:
[0155] The spatiotemporal graph is represented and learned through the graph representation learning algorithm to obtain a graph representation learning model;
[0156] Spatiotemporal features are extracted from the graph representation learning model, and the spatiotemporal features are used to represent the main change information of the video stream in two dimensions of time and space.
[0157] Further optionally, the action decision module is also used for:
[0158] Define the state vector S according to the processing state of the video stream t , wherein the processing status includes processing speed, memory usage, number of abnormal frames and spatiotemporal characteristics, which are used to describe the environmental status of the video stream processing system at the current moment or the status of the video stream itself;
[0159] An action space is defined according to processing actions of the video stream, wherein the processing actions include adjusting compression ratio, allocating computing resources, adjusting cache strategy, and dynamically adjusting transmission protocol, and each processing action is represented as a;
[0160] The neural network structure is used as a policy network, which receives the state vector as input and outputs the probability distribution of the action to be taken in the current state. The policy network is represented by π θ (a∣S t ), where θ is the parameter of the policy network, a is the action, S t is the current state;
[0161] Defining a reward function R(s, a) according to a processing objective for evaluating the quality of taking action a in state s, wherein the processing objective includes optimizing video quality, improving processing speed, reducing memory usage, and / or reducing abnormal frames;
[0162] The policy network is trained using reinforcement learning. During the training process, the expected cumulative reward is maximized by optimizing the policy network parameters θ, and the policy network is guided to automatically adjust the processing actions according to the current processing status by combining the evaluation of the value network. Through continuous trial and error and optimization, the optimal processing strategy is learned to improve the overall performance of video stream processing.
[0163] Apply the trained strategy network to decide the next processing action according to the current state, and the processing action includes:
[0164] Further optionally, the spatiotemporal feature module is also used for:
[0165] Randomly initialize the parameters θ of the policy network;
[0166] Collect a sample of states, actions, and rewards;
[0167] Calculate the gradient of the objective function J(θ) with respect to the policy network parameters The calculation formula is:
[0168]
[0169] Among them, Q is estimated approximately by the value network πθ (S t , a), Q πθ (S t , a) means that under the strategy πθ, from the state vector S t The expected cumulative reward after taking action a, E πθ [] represents the expectation under the strategy πθ, represents the logarithmic gradient of the policy network output action probability, π θ (a∣S t ) represents the policy network in the state vector S t The probability of taking action a under
[0170] Update the parameters θ of the policy network by the gradient ascent method. The update formula is:
[0171]
[0172] Among them, α is the learning rate, which controls the step size of parameter update;
[0173] Repeat the training until the policy network converges to obtain a trained policy network, which is used to improve the overall performance of video stream processing.
[0174] Further optionally, the spatiotemporal feature module is also used for:
[0175] Construct a value network based on a neural network, represented by Q ω(s, a), where ω is the parameter of the value network, the input of the value network is state s and action a, and the output is the Q value estimate of taking action a in state s;
[0176] Train the value network to make the Q value output by the value network close to the actual expected cumulative reward;
[0177] During the training process, the state vector S is propagated forward. t And action a is input into the value network to obtain the Q value estimate Q ω (S t , a);
[0178] Estimate the Q value Q ω (S t , a) as Q πθ (S t , an approximate estimate of a).
[0179] Further optionally, the spatiotemporal feature module is also used for:
[0180] Get the processing speed, memory usage, number of abnormal frames, and spatiotemporal features of the current video stream and combine them into a state vector S t ;
[0181] The state vector S t Input into the trained policy network to obtain the probability distribution π of the action to be taken in the current state θ (a∣S t );
[0182] Choose an action based on a probability distribution.
[0183] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A method for dynamic processing of video stream data based on memory computing, characterized in that: The method comprises: A dynamic memory pool is allocated in the memory for storing and processing video stream data; Get video stream data in real time and split it into individual frames; Extract frame-level features of key frames; Calculate optical flow for key frames; Combining frame-level features with optical flow information to construct a spatiotemporal graph, and analyzing the spatiotemporal graph to extract spatiotemporal features; The processing state of the video stream is represented as a state vector, a policy network is trained, and the next processing action is decided by the policy network according to the current state, wherein the processing state includes processing speed, memory usage, number of abnormal frames and spatiotemporal domain characteristics, and the processing actions include adjusting compression ratio, allocating computing resources, adjusting cache strategy and dynamically adjusting transmission protocol.
2. The method for dynamically processing video stream data based on memory computing according to claim 1, characterized in that: The calculating of optical flow for the key frame comprises: Select feature points in keyframes; Calculate the optical flow of feature points between adjacent key frames by using an optical flow algorithm, wherein the optical flow is used to represent the displacement vector of the feature points between consecutive key frames; An optical flow feature is extracted from the optical flow data, wherein the optical flow feature includes the magnitude and direction of the optical flow vector.
3. The method for dynamically processing video stream data based on memory computing according to claim 2 is characterized in that: The process of combining frame-level features with optical flow information to construct a spatiotemporal graph includes: Treat the key frame as a node in the spatiotemporal graph, where each node contains the frame-level features of the key frame; The edges in the spatiotemporal graph are constructed according to the optical flow information between each pair of adjacent key frames, where each edge connects the nodes of adjacent key frames, and the weight or attribute of the edge is set as the optical flow feature.
4. The method for dynamically processing video stream data based on memory computing according to claim 3 is characterized in that: The analyzing the space-time graph to extract space-time domain features includes: The spatiotemporal graph is represented and learned through the graph representation learning algorithm to obtain a graph representation learning model; Spatiotemporal features are extracted from the graph representation learning model, and the spatiotemporal features are used to represent the main change information of the video stream in two dimensions of time and space.
5. The method for dynamically processing video stream data based on memory computing according to claim 1, characterized in that: The processing state of the video stream is represented as a state vector, a policy network is trained, and the next processing action is decided by the policy network according to the current state, including: Define the state vector S according to the processing state of the video stream t , wherein the processing status includes processing speed, memory usage, number of abnormal frames and spatiotemporal characteristics, which are used to describe the environmental status of the video stream processing system at the current moment or the status of the video stream itself; An action space is defined according to processing actions of the video stream, wherein the processing actions include adjusting compression ratio, allocating computing resources, adjusting cache strategy, and dynamically adjusting transmission protocol, and each processing action is represented as a; The neural network structure is used as a policy network, which receives the state vector as input and outputs the probability distribution of the action to be taken in the current state. The policy network is represented by π θ (a∣S t ), where θ is the parameter of the policy network, a is the action, and S t is the current state; Defining a reward function R(s, a) according to a processing objective for evaluating the quality of taking action a in state s, wherein the processing objective includes optimizing video quality, improving processing speed, reducing memory usage, and / or reducing abnormal frames; The policy network is trained using reinforcement learning. During the training process, the expected cumulative reward is maximized by optimizing the policy network parameters θ, and the policy network is guided to automatically adjust the processing actions according to the current processing status by combining the evaluation of the value network. Through continuous trial and error and optimization, the optimal processing strategy is learned to improve the overall performance of video stream processing. Apply the trained strategy network to decide the next processing action according to the current state, and the processing action includes:
6. The method for dynamically processing video stream data based on memory computing according to claim 5, characterized in that: The policy network is trained by reinforcement learning. During the training process, the expected cumulative reward is maximized by optimizing the policy network parameters θ, and the policy network is guided to automatically adjust the processing action according to the current processing state in combination with the evaluation of the value network. Through continuous trial and error and optimization, the optimal processing strategy is learned to improve the overall performance of video stream processing, including: Randomly initialize the parameters θ of the policy network; Collect a sample of states, actions, and rewards; Calculate the gradient of the objective function J(θ) with respect to the policy network parameters The calculation formula is: Among them, Q is estimated approximately by the value network πθ (S t , a), Q πθ (S t , a) means that under the strategy πθ, from the state vector S t The expected cumulative reward after taking action a, E πθ [] represents the expectation under the strategy πθ, represents the logarithmic gradient of the policy network output action probability, π θ (a∣S t ) represents the policy network in the state vector S t The probability of taking action a under Update the parameters θ of the policy network by the gradient ascent method. The update formula is: Among them, α is the learning rate, which controls the step size of parameter update; Repeat the training until the policy network converges to obtain a trained policy network, which is used to improve the overall performance of video stream processing.
7. The method for dynamically processing video stream data based on memory computing according to claim 6, characterized in that: The value network approximates Q πθ (S t , a), including: Construct a value network based on a neural network, represented by Q ω (s, a), where ω is the parameter of the value network, the input of the value network is state s and action a, and the output is the Q value estimate of taking action a in state s; Train the value network to make the Q value output by the value network close to the actual expected cumulative reward; During the training process, the state vector S is propagated forward. t And action a is input into the value network to obtain the Q value estimate Q ω (S t , a); Estimate the Q value Q ω (S t , a) as Q πθ (S t , an approximate estimate of a).
8. The method for dynamically processing video stream data based on memory computing according to claim 5, characterized in that: The strategy network trained by the application determines the next processing action according to the current state, including: Get the processing speed, memory usage, number of abnormal frames, and spatiotemporal features of the current video stream and combine them into a state vector S t ; The state vector S t Input into the trained policy network to obtain the probability distribution π of the action to be taken in the current state θ (a∣S t ); Choose an action based on a probability distribution.
9. A video stream data dynamic processing system based on memory computing, used to implement the video stream data dynamic processing method based on memory computing according to any one of claims 1 to 8, characterized in that: The system comprises: Partitioning module: used to divide a dynamic memory pool in the memory for storing and processing video stream data; Data acquisition module: used to acquire video stream data in real time and split it into separate frames; Feature extraction module: used to extract frame-level features of key frames; Optical flow calculation module: used to calculate the optical flow of key frames; Spatiotemporal feature module: used to combine frame-level features with optical flow information to construct a spatiotemporal graph, and analyze the spatiotemporal graph to extract spatiotemporal features; Action decision module: used to represent the processing state of the video stream as a state vector, train a policy network, and decide the next processing action based on the current state through the policy network, wherein the processing state includes processing speed, memory usage, number of abnormal frames and spatiotemporal domain characteristics, and the processing action includes adjusting the compression ratio, allocating computing resources, adjusting the cache strategy and dynamically adjusting the transmission protocol.
Citation Information
Patent Citations
Key frame recognition model training method and device, recognition method and device
CN112446342A
Video feature extraction method and device, equipment and storage medium
CN115115991A
Video analysis method and system of monitoring terminal
CN118781526A
Air-ground network optimization method and system based on MEC and digital twinning
CN119255263A
Chip platform hardware abstraction layer construction method
CN119312219A
Cited By
Operation management system based on data interaction
CN120568102A