Method and Application for Constructing a Video Transmission Configuration Model
By building a video transmission configuration model, using reinforcement learning and singular value decomposition, the adaptive video transmission parameter configuration of multi-view camera clusters is realized, which solves the problems of resource waste and accuracy, and improves video transmission efficiency and analysis accuracy.
Patent Information
- Application Number
- CN202210898346.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-07-28
AI Technical Summary
The prior art cannot realize the adaptive configuration of video transmission parameters with a small resource overhead, resulting in redundancy and waste of video data in multi-view camera cluster systems, and cannot take into account the accuracy of video transmission performance and analysis tasks.
By building a video transmission configuration model, using reinforcement learning network and singular value decomposition, the clustered camera matches edge computing nodes with similarity, dynamically adjusts video transmission parameters, and realizes adaptive configuration.
Reduces resource overhead, improves the efficiency of video transmission and analysis accuracy, and adapts to different network environments and task requirements.
Smart Images

Figure CN115357379B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video analysis, and more specifically, relates to a method for constructing a video transmission configuration model and its application. Background Art
[0002] With the advent of the big data era and the continuous development of communication networks and embedded technologies, cameras play an increasingly important role in smart cities due to their ability to capture rich information. Cameras are widely deployed to obtain rich and intuitive environmental information. These cameras are used in scenarios such as traffic control, security supervision, and factory monitoring. Due to the requirements of large-scale automation and low economic costs, the purpose of video analysis is to use computer vision technology to analyze the information collected by cameras, separate the background in the scene, and perform target analysis, which is widely applied in fields such as road monitoring, autonomous driving, target detection, and intelligent industry. The video source in video analysis usually comes from a multi-view camera cluster system. There are often multiple cameras in the multi-view cluster system, and the video transmission parameters (such as frame rate, resolution, and video segment size) of each camera are different. Its specific configuration will affect the performance of subsequent video analysis, such as accuracy and transmission bandwidth requirements. Therefore, it is of great significance to study a video transmission configuration method and its application.
[0003] Edge computing nodes have broad application prospects in video transmission and analysis. By adding an edge node architecture, video streams can initially be processed at the edge node instead of delivering all content to the cloud. Compared with the latter, it has the advantage of a smaller network burden at the edge node. Processing queries helps reduce service latency while maintaining high accuracy. Under the edge node architecture, existing video transmission configuration methods usually transmit the video captured by cameras to random edge nodes, and randomly match cameras with edge nodes, resulting in a large amount of redundancy in the video data collected by the multi-view camera cluster system, leading to waste of configuration resources. Moreover, the configuration parameters for each camera to transmit video are relatively fixed and cannot be adaptively adjusted according to specific analysis tasks and actual network environments, resulting in large resource overheads and being unable to balance video transmission performance and the accuracy of video analysis tasks. Summary of the Invention
[0004] Aiming at the above-mentioned defects or improvement requirements of the prior art, the present invention provides a method for constructing a video transmission configuration model and its application, so as to solve the technical problem that the prior art cannot achieve self-adaptive configuration of video transmission parameters with a small resource overhead.
[0005] To achieve the above object, the present invention provides a method for constructing a video transmission configuration model, including the following steps:
[0006] S1. Build a video transmission configuration model. Among them, the video transmission configuration model includes: multiple video transmission configuration units. The number of video transmission configuration units is the same as the number of edge computing nodes, and one video transmission configuration unit corresponds to one edge computing node. The video transmission configuration unit includes a reinforcement learning network.
[0007] The video transmission configuration unit is used to input the environmental state information when the corresponding edge computing node receives the video segment collected by the corresponding camera into its reinforcement learning network to obtain video transmission parameters. After uniformly configuring the cameras corresponding to the corresponding edge computing node according to the obtained video transmission parameters, calculate the reward value after the corresponding configuration under the above environmental state information, which is recorded as the reward value corresponding to the environmental state information. Among them, the environmental state information includes: transmission delay, transmission bandwidth, accuracy and type of video analysis. The video transmission parameters include: video frame rate, video resolution, and the size of the video segment for video analysis.
[0008] S2. Input the training sample subsets corresponding to each edge computing node into the corresponding video transmission configuration unit, and maximize the cumulative reward values corresponding to all environmental state information in each training sample subset respectively to train each video transmission configuration unit, so as to obtain a trained video transmission configuration model.
[0009] Among them, the training sample subset corresponding to the edge computing node includes: the environmental state information when the edge computing node receives the video segment collected by the corresponding camera at several historical moments. Among them, the method for determining the corresponding relationship between the edge computing node and the camera includes: clustering each camera according to the similarity of the video content it collects to obtain each camera cluster; and respectively matching each camera cluster with the edge computing node closest to it.
[0010] Further preferably, the method for determining the corresponding relationship between the edge computing node and the camera includes the following steps:
[0011] S01. Perform singular value decomposition on each video segment collected by each camera frame by frame to obtain the singular values of each video frame collected by each camera.
[0012] S02. For any two cameras, calculate the similarity of the video content they collect according to the following formula;
[0013]
[0014] Among them, sim pq is the similarity of the video content collected by the p-th camera and the q-th camera; S p is the set of singular values of each video frame collected by the p-th camera; |S q | is the singular value set Sp The number of singular values in it; S q is the set of singular values of each video frame collected by the q-th camera; |S q | is the number of singular values in the singular value set S q ; is the singular value set S q the i-th singular value in it, is the singular value set S q the i-th singular value in it,
[0015] S03. Cluster each camera to obtain each camera cluster, so that the similarity of the video content collected by any two cameras in each camera cluster is less than the preset similarity;
[0016] S04. Calculate the geographical distances between each camera cluster and each edge computing node respectively, and match each camera cluster with the edge computing node closest to it respectively.
[0017] Further preferably, denote the moment when the video segment output by the camera is transmitted to the cloud server for video analysis after the corresponding cameras corresponding to the edge computing node are uniformly configured according to the video transmission parameters obtained based on the environmental state information as t. At this time, the reward value corresponding to the obtained environmental state information is:
[0018] R t =(1 - ε)accuracy t - εdelay t
[0019] where ε is the preset weight; accuracy t is the accuracy of video analysis of the video segment transmitted at time t; delay t is the transmission delay of transmitting the video segment output by the camera to the cloud server at time t.
[0020] Further preferably, the reinforcement learning network includes: an Actor network and a Critic network; the Actor network is used to obtain video transmission parameters based on the environmental state information when the corresponding edge computing node receives the video segment collected by the corresponding camera; the Critic network is used to evaluate the video transmission parameters obtained by the Actor network, and train the Actor network with the goal of maximizing the evaluation result.
[0021] Further preferably, during the training process, for each video transmission configuration unit, every time an environmental state information sample is input, and after uniformly configuring each camera corresponding to the corresponding edge computing node according to the video transmission parameters, the QoE metric value after video analysis of the transmitted video segment is collected at this time, and it is determined whether the QoE metric value is less than a preset threshold. If so, the corresponding reward value is calculated, and the parameter value of the reinforcement learning network in the video transmission configuration unit is updated based on the reward value; otherwise, the training of the current round is ended, and the next environmental state information sample is directly input to start the next round of training.
[0022] Further preferably, in the above step S2, the training processes of the respective video transmission configuration units are executed in parallel.
[0023] Further preferably, the method for constructing the above video transmission configuration model further includes: every preset time period, re-collecting the training sample subset corresponding to the edge computing node, and performing reinforcement learning training on the video transmission configuration model by using step S2, so as to update the video transmission configuration model.
[0024] Further preferably, the corresponding relationship between the edge computing node and the camera is updated regularly according to its determination method.
[0025] In a second aspect, the present invention provides a video transmission configuration method, including: inputting the environmental state information when the camera to be configured collects a video segment into the video transmission configuration unit corresponding to the edge computing node corresponding to the camera to be configured in the video transmission configuration model, obtaining video transmission parameters, and configuring the camera to be configured according to the obtained video transmission parameters;
[0026] Among them, the video transmission configuration model is constructed by using the method for constructing the video transmission configuration model provided in the first aspect of the present invention.
[0027] In a third aspect, the present invention provides a video transmission configuration system, including: a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it executes the video transmission configuration method provided in the second aspect of the present invention.
[0028] In a fourth aspect, the present invention further provides a computer-readable storage medium, the computer-readable storage medium includes a stored computer program, wherein when the computer program is run by a processor, it controls the device where the storage medium is located to execute the method for constructing the video transmission configuration model provided in the first aspect of the present invention and / or the video transmission configuration method provided in the second aspect of the present invention.
[0029] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0030] 1. The present invention provides a method for constructing a video transmission configuration model. Based on the historical videos collected by each camera in advance, each camera is clustered into multiple camera clusters according to the similarity of the video content it collects, and is matched with the nearest edge computing node. The subsets obtained by clustering share the same edge computing node, so that each camera with similar video sources in the same camera cluster shares the same video transmission configuration and selectively transmits data, greatly reducing resource overhead. At the same time, a video transmission configuration unit is trained for each edge computing node based on the reinforcement learning method, and a video transmission configuration model is jointly constructed, realizing the adaptive configuration selection of dynamic video segments and enabling the self-adaptive configuration of video transmission parameters with relatively small resource overhead.
[0031] 2. The method for constructing the video transmission configuration model provided by the present invention uses a custom class KL distance combined with singular value decomposition to calculate the similarity between multi-view camera clusters, clusters those with high similarity into one cluster, and completes the partition matching of the camera network topology. This method is based on singular value decomposition, is easy to implement, has low computational complexity, uses logarithmic ratios to measure distances, and can quickly obtain the similarity of camera clusters.
[0032] 3. The method for constructing the video transmission configuration model provided by the present invention, the collected environmental state information includes: transmission delay, transmission bandwidth, accuracy and type of video analysis. In the case where the bandwidth between the camera and the edge node is fluctuating and limited, aiming at low transmission delay and high video analysis accuracy, the configuration problem of the video transmission parameters of the camera is transformed into a sequential decision-making problem, and a deep reinforcement learning algorithm is used for training, which can take into account both the video transmission performance and the accuracy of the video analysis task, so as to realize the adaptive adjustment of the video transmission parameters under specific analysis tasks and actual network environments.
[0033] 4. Considering that the training process of the model is mainly carried out under a black-box model and the whole process is decided by the machine, in order to prevent incorrect decisions and risky decisions, the method for constructing the video transmission configuration model provided by the present invention, during the training process, in addition to considering the obtained reward value, further considers the QoE metric value after video analysis of the transmitted video segment, and the priority of the QoE metric value is higher than the reward value. By introducing prior information to control the update of network parameters during the training process and using rich domain knowledge to improve robustness, the possibility of incorrect decisions and risky decisions is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic diagram of a video analysis architecture provided in Embodiment 1 of the present invention;
[0035] Figure 2Flowchart of camera network topology partitioning provided in Embodiment 1 of the present invention;
[0036] Figure 3 Flowchart of the method for constructing a video transmission configuration model provided in Embodiment 1 of the present invention;
[0037] Figure 4 Training flowchart of any video transmission configuration unit provided in Embodiment 1 of the present invention. Detailed implementation manners
[0038] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0039] Embodiment 1
[0040] This embodiment provides a method for constructing a video transmission configuration model, and the constructed video transmission configuration model serves subsequent video analysis operations. It should be noted that in a real environment (such as a traffic intersection), multiple cameras (camera clusters) are deployed to monitor and collect video data in different areas. Considering the complex scenario, the cameras are wirelessly connected to edge computing nodes. The camera cluster continuously collects videos (such as pedestrian and vehicle information at the zebra crossing, which can be adjusted according to the specific network fluctuations every once in a while), sends video segments to different edge nodes for preprocessing, and then the edge computing nodes upload the preprocessing results to the cloud server for video analysis. This is an "edge-cloud" collaborative video analysis architecture. As Figure 1 shown, the video analysis system includes cameras, edge computing nodes, and a cloud server. When offline, the edge computing node saves the video segments received locally at the node as historical training data. When online, the edge computing node selects the most information-rich frame set by analyzing spatio-temporal redundancy.
[0041] Although there are multiple cameras in the camera cluster, they may not be related or redundant to each other, which means there is room for further partitioning subsets in the topology of this camera cluster. Taking traffic cameras as an example, in the offline analysis part, traffic cameras facing highway sections may be related. For example, there may be more spatial overlaps in the scenes captured by traffic intersection cameras because the same vehicle may appear in the video feeds of multiple cameras, so they tend to share similar configurations. The video analyzer deployed on the edge computing node can further screen and partition according to the video information data, so as to obtain different subsets of camera clusters.
[0042] Specifically, in this embodiment, the correspondence between the edge computing nodes and the cameras is determined in the following manner: for each camera, clustering is performed according to the similarity of the video content it captures to obtain each camera cluster; and each camera cluster is respectively matched with the edge computing node closest to it, thereby completing the partition matching of the camera network topology. The methods for calculating the similarity of video content include mean square error, feature point detection, optical flow method, etc. In an alternative embodiment, the following method is used to calculate the similarity of video content, and the method for determining the correspondence between the corresponding edge computing nodes and the cameras includes the following steps:
[0043] S01. Perform singular value decomposition on the video segments captured by each camera frame by frame to obtain the singular values of each video frame captured by each camera;
[0044] Specifically, the singular value decomposition formula is as follows:
[0045]
[0046] The column vectors of U and V are eigenvectors, Σ is the singular value matrix, and the values on its diagonal are the singular values.
[0047] S02. For any two cameras, calculate the similarity of the video content they capture according to the following formula;
[0048]
[0049] where, sim pq is the similarity of the video content captured by the p-th camera and the q-th camera, which is used to reflect the redundancy between the two video contents; S p is the set of singular values of each video frame captured by the p-th camera; |S q | is the number of singular values in the singular value set S p ; S q is the set of singular values of each video frame captured by the q-th camera; |S q | is the number of singular values in the singular value set S q ; is the i-th singular value in the singular value set S p , is the i-th singular value in the singular value set S q ,
[0050] S03. Cluster each camera to obtain each camera cluster, such that the similarity of the video content collected by any two cameras in each camera cluster is less than a preset similarity; specifically, the similarity of the video content collected by each camera can be represented by a heat map, and cameras with high similarity are clustered into one cluster to complete the partition matching of the camera cluster. In this embodiment, after normalizing the similarity of the video content collected by any two cameras in the camera cluster, the preset similarity is set to 0.7.
[0051] S04. Calculate the geographical distances between each camera cluster and each edge computing node respectively, and match each camera cluster with the edge computing node closest to it.
[0052] The same camera cluster will share the same edge computing node to achieve configuration sharing, or in the case of limited resources, only selectively upload the video data of one camera in the subset (such as a counting task). Further, the correspondence between the edge computing node and the camera can be updated regularly according to its determination method based on the historical data set (in this embodiment, it is updated every 24 hours), and the whole process is as Figure 2 shown.
[0053] After obtaining the correspondence between the edge computing node and the camera in advance, as Figure 3 shown, the method for constructing the above video transmission configuration model includes the following steps:
[0054] S1. Build a video transmission configuration model; wherein, the video transmission configuration model includes: multiple video transmission configuration units; the number of video transmission configuration units is the same as the number of edge computing nodes, and one video transmission configuration unit corresponds to one edge computing node; the video transmission configuration unit includes a reinforcement learning network.
[0055] The video transmission configuration unit is used to input the environmental state information when the corresponding edge computing node receives the video segment collected by the corresponding camera into its reinforcement learning network to obtain video transmission parameters; after uniformly configuring the cameras corresponding to the corresponding edge computing node according to the obtained video transmission parameters, calculate the reward value after the corresponding configuration in the above environmental state information, which is recorded as the reward value corresponding to the environmental state information; wherein, the environmental state information includes: transmission delay, transmission bandwidth, accuracy obtained after video analysis, and video analysis type; the video analysis type includes: target counting, target tracking, target detection, etc.; the video transmission parameters include: video frame rate, video resolution, and the size of the video segment used for video analysis.
[0056] S2. Input the training sample subsets corresponding to each edge computing node into the corresponding video transmission configuration unit, and maximize the cumulative reward values corresponding to all environmental state information in each training sample subset respectively to train each video transmission configuration unit, so as to obtain a trained video transmission configuration model. Wherein, the training sample subset corresponding to the edge computing node includes: the environmental state information when the edge computing node receives the video segments collected by the corresponding camera at several historical moments.
[0057] It should be noted that in addition to basic video transmission parameters such as video frame rate and video resolution, since the camera uploads video frames in the form of video segments to the edge computing node during the real-time capture of video frames by the camera, and buffers the video frames in a set time window before uploading the next segment each time. Each workload and resource requirement is affected by the video content. Small segment sizes are vulnerable to noise, while large segments require buffering more frames. Therefore, the window size, that is, the size of the video segment used for video analysis, also needs to be concerned. Therefore, the video transmission parameters concerned by the present invention include: video frame rate, video resolution, and the size of the video segment used for video analysis. That is, the action space is the video transmission parameter, and the reinforcement learning network represents the selection of actions for each query according to the current environmental state information, that is, which configuration is used to encode the segment (i.e., the video transmission parameter), and resets the state through the reward feedback from the environment.
[0058] It should be noted that the reward value corresponding to the environmental state information can be calculated by weighing mutually exclusive indicators such as accuracy, energy consumption, bandwidth requirement, and delay. Specifically, in an optional implementation manner, when the bandwidth between the camera and the edge node is fluctuating and limited, aiming at low transmission delay and high video analysis accuracy, the configuration problem of the video transmission parameters of the camera is transformed into a sequential decision-making problem, and a deep reinforcement learning algorithm is used for training. Denote the moment when the video segments output by the camera are transmitted to the cloud server for video analysis after the cameras corresponding to the corresponding edge computing nodes are uniformly configured according to the video transmission parameters obtained based on the environmental state information as t. Since there is a conflict between the two goals of low transmission delay and high video analysis accuracy, therefore, the reward value corresponding to the environmental state information at this time is:
[0059] R t =(1 - ε)accuracy t - εdelay t
[0060] Wherein, ε is a preset weight used to weigh the proportion of accuracy and delay, and the value in this embodiment is 0.4; accuracy t is the accuracy of performing corresponding type of video analysis on the video segment transmitted at time t; delay tThe transmission delay for transmitting the video segment output by the camera to the cloud server at time t.
[0061] In this implementation, the training sample subset is input into the model in batches for training. The weight value of the reinforcement learning network is updated by maximizing the cumulative reward value corresponding to all environmental state information in the training sample subset, so that the reinforcement learning network can output the optimal video transmission parameters. The corresponding objective function is: where T is the number of samples in a batch, corresponding to one training cycle. Thus, the cumulative reward is defined as the sum of the accuracy and delay functions of all queries for an event within the cycle. The present invention maximizes the average of low delay and accuracy through the cumulative reward. The present invention makes the most beneficial configuration decisions through the interaction between the agent and the external environment, makes a fine-grained adaptive configuration for the currently collected video segments, selects to encode them independently, and provides a reference selection for the configuration of subsequent segments.
[0062] Further, in an alternative implementation, as Figure 4 shown is the training flow chart of any video transmission configuration unit. Among them, the reinforcement learning network includes: an Actor network and a Critic network; the Actor network is used to obtain video transmission parameters based on the environmental state information when the corresponding edge computing node receives the video segment collected by the corresponding camera; the Critic network is used to evaluate the video transmission parameters obtained by the Actor network, and trains the Actor network with the goal of maximizing the evaluation result. Under this reinforcement learning network, the Actor network learns the maximum expected return, and the Critic provides suggestions; at this time, the Critic learns a central value function and can observe global information. Specifically, the A3C reinforcement learning algorithm is used for training. During the training process, the environmental state information in the training sample subset is first input into the Actor network to obtain the corresponding video transmission parameters; the environmental state information and the corresponding video transmission parameters are input into the Critic network to obtain its evaluation result, and the Actor network is trained with the goal of maximizing the evaluation result; at the same time, after the cameras corresponding to the corresponding edge computing nodes are uniformly configured according to the obtained video transmission parameters, the reward value after the corresponding configuration is calculated under the above environmental state information; the environmental state information, the corresponding video transmission parameters and the reward value are stored in the experience pool as a tuple information; when the experience pool is full of data, tuple information data is randomly sampled from the experience pool to train the Critic network.
[0063] Further, in order to accelerate the training speed and improve the training efficiency, in an alternative implementation, the training processes of the respective video transmission configuration units can be executed in parallel. That is, based on multi-agent task learning, agents in different threads explore different strategies, and the learning speed is increased through parallel agents.
[0064] Further, since the above training process is mainly carried out under a black-box model and the entire process is machine-decision-making, in order to prevent incorrect decisions and risky decisions, in an alternative implementation, when updating the parameters in the video transmission configuration model during the training process, in addition to considering the obtained reward value, the QoE (Quality of Experience) metric value after video analysis of the transmitted video segments is further considered, and the priority of the QoE metric value is higher than that of the reward value. Specifically, during the training process, for each video transmission configuration unit, every time an environmental state information sample is input, and after the respective cameras corresponding to the corresponding edge computing nodes are uniformly configured according to the video transmission parameters, the QoE metric value (set to have a value range of [0, 10]) after video analysis of the transmitted video segments at this time is collected, and it is determined whether the QoE metric value is less than a preset threshold (set to 8 in this embodiment). If so, the corresponding reward value is calculated, and the parameter values of the reinforcement learning network in the video transmission configuration unit are updated based on the reward value; otherwise, the training of the current round is ended, and the next environmental state information sample is directly input, thereby starting the next round of training. By introducing prior information to control the update of the network parameters during the training process, the possibility of incorrect decisions and risky decisions is reduced.
[0065] Further, in an alternative implementation, the method for constructing the above video transmission configuration model further includes: every preset time period (set to 4h in this embodiment), the training sample subset corresponding to the edge computing node is recollected, and the video transmission configuration model is trained by reinforcement learning using step S2, thereby updating the video transmission configuration model.
[0066] In summary, the present invention discloses a method for constructing a video transmission configuration model, which is mainly used to reduce resource consumption and meet the requirements of low latency and high accuracy. On the one hand, the present invention utilizes a historical data set, calculates the inter-frame redundant similarity in the offline stage, and divides and configures a multi-view camera cluster using the overall structure. The divided subsets share the same edge node, selectively transmit data, and reduce resource overhead. On the other hand, an improved deep reinforcement learning algorithm with prior knowledge is proposed to achieve an adaptive configuration selection for the dynamic video segments of the configuration knob, and by adjusting the video coding configuration, a balance between accuracy and latency requirements is achieved under the conditions of bandwidth fluctuations and uncertain query tasks. On the basis of combining edge computing nodes, a flexible and efficient solution is provided for the configuration selection analysis of multimedia video streams.
[0067] Embodiment 2
[0068] A video transmission configuration method includes: inputting the environmental state information when a camera to be configured captures a video segment into a video transmission configuration unit corresponding to an edge computing node corresponding to the camera to be configured in a video transmission configuration model, obtaining video transmission parameters, and configuring the camera to be configured according to the obtained video transmission parameters.
[0069] Among them, the video transmission configuration model is constructed by using the construction method of the video transmission configuration model provided in Embodiment 1 of the present invention.
[0070] The related technical solutions are the same as those in Embodiment 1 and will not be elaborated here.
[0071] Embodiment 3
[0072] A video transmission configuration system includes: a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it executes the video transmission configuration method provided in Embodiment 2 of the present invention.
[0073] The related technical solutions are the same as those in Embodiment 2 and will not be elaborated here.
[0074] Embodiment 4
[0075] A computer-readable storage medium includes a stored computer program. Among them, when the computer program is run by a processor, it controls the device where the storage medium is located to execute the construction method of the video transmission configuration model provided in Embodiment 1 of the present invention and / or the video transmission configuration method provided in Embodiment 2 of the present invention.
[0076] The related technical solutions are the same as those in Embodiment 1 and Embodiment 2 and will not be elaborated here.
[0077] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for constructing a video transmission configuration model, characterized in that Including the following steps: S1. Build a video transmission configuration model; the video transmission configuration model includes: a plurality of video transmission configuration units; the number of video transmission configuration units is the same as the number of edge computing nodes, and one video transmission configuration unit corresponds to one edge computing node; the video transmission configuration unit includes a reinforcement learning network; The video transmission configuration unit is used to input the environmental state information when the corresponding edge computing node receives the video segment collected by the corresponding camera into its reinforcement learning network to obtain video transmission parameters; after uniformly configuring the cameras corresponding to the corresponding edge computing node according to the obtained video transmission parameters, calculate the reward value after the corresponding configuration under the environmental state information, denoted as the reward value corresponding to the environmental state information; the environmental state information includes: transmission delay, transmission bandwidth, accuracy and type of video analysis; the video transmission parameters include: video frame rate, video resolution, and the size of the video segment for video analysis; S2. Input the training sample subsets corresponding to each edge computing node into the corresponding video transmission configuration unit, and maximize the cumulative reward values corresponding to all environmental state information in each training sample subset respectively to train each video transmission configuration unit, so as to obtain a trained video transmission configuration model; Among them, the training sample subset corresponding to the edge computing node includes: the environmental state information when the edge computing node receives the video segment collected by the corresponding camera at several historical moments; A method for determining the corresponding relationship between the edge computing node and the camera includes: clustering each camera according to the similarity of the video content it collects to obtain each camera cluster; and respectively matching each camera cluster with the edge computing node closest to it.
2. The method for constructing a video transmission configuration model according to claim 1, wherein The method for determining the corresponding relationship between the edge computing node and the camera includes the following steps: S01. Perform singular value decomposition on each video segment collected by each camera frame by frame to obtain the singular values of each video frame collected by each camera; S02. For any two cameras, calculate the similarity of the video content they collect according to the following formula; where, sim pq is the similarity of the video content collected by the p-th camera and the q-th camera; S p is the set of singular values of each video frame collected by the p-th camera; |S q | is the number of singular values in the singular value set S p ; S q is the set of singular values of each video frame collected by the q-th camera; |S q | is the number of singular values in the singular value set S q ; is the i-th singular value in the singular value set S p ; is the i-th singular value in the singular value set S q ; S03. Cluster each camera to obtain each camera cluster, so that the similarity of the video content collected by any two cameras in each camera cluster is less than the preset similarity; S04. Calculate the geographical distances between each camera cluster and each edge computing node respectively, and respectively match each camera cluster with the edge computing node closest to it.
3. The method for constructing a video transmission configuration model according to claim 1, wherein Denote the moment when the video segment output by the camera is transmitted to the cloud server for video analysis after uniformly configuring the cameras corresponding to the corresponding edge computing node according to the video transmission parameters obtained based on the environmental state information as t. At this time, the reward value corresponding to the obtained environmental state information is: R t = (1 - ε)accuracy t - εdelay t where ε is a preset weight; accuracy t is the accuracy of video analysis for the video segment transmitted at time t; delay t is the transmission delay of transmitting the video segment output by the camera to the cloud server at time t.
4. The method for constructing a video transmission configuration model according to claim 1, wherein The reinforcement learning network includes: an Actor network and a Critic network; the Actor network is used to obtain video transmission parameters based on the environmental state information when the corresponding edge computing node receives the video segments collected by the corresponding camera; the Critic network is used to evaluate the video transmission parameters obtained by the Actor network, and train the Actor network with the goal of maximizing the evaluation result.
5. The method for constructing a video transmission configuration model according to claim 1, wherein In step S2, the training processes of each video transmission configuration unit are executed in parallel.
6. The method for constructing a video transmission configuration model according to any one of claims 1-5, characterized in that, During the training process, for each video transmission configuration unit, every time an environmental state information sample is input, and after the corresponding cameras of the corresponding edge computing node are uniformly configured according to the video transmission parameters, the QoE metric value after video analysis of the transmitted video segment is collected at this time, and it is judged whether the QoE metric value is less than a preset threshold. If so, the corresponding reward value is calculated, and the parameter values of the reinforcement learning network in the video transmission configuration unit are updated based on the reward value; otherwise, the training of the current round is ended, and the next environmental state information sample is directly input to start the next round of training.
7. The method for constructing a video transmission configuration model according to any one of claims 1-5, characterized in that It further includes: Every preset time period, a training sample subset corresponding to the edge computing node is re-collected, and the video transmission configuration model is trained by reinforcement learning using step S2 to update the video transmission configuration model. The corresponding relationship between the edge computing node and the camera is updated regularly according to its determination method.
8. A video transmission configuration method, characterized in that, It includes: Input the environmental state information when the camera to be configured collects video segments into the video transmission configuration unit corresponding to the edge computing node corresponding to the camera to be configured in the video transmission configuration model, obtain video transmission parameters, and configure the camera to be configured according to the obtained video transmission parameters. Among them, the video transmission configuration model is constructed by using the construction method of the video transmission configuration model described in any one of claims 1-7.
9. A video transmission configuration system, comprising: A memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it executes the video transmission configuration method described in claim 8.
10. A computer-readable storage medium, the computer-readable storage medium including a stored computer program, wherein, When the computer program is run by the processor, it controls the device where the storage medium is located to execute the construction method of the video transmission configuration model described in any one of claims 1-7 and / or the video transmission configuration method described in claim 8.
Citation Information
Patent Citations
Method for scheduling interference workloads on edge network resources
EP4024212A1
Digital telepathology and virtual control of a microscope using edge computing
US20200372643A1