Training method, application method, device and equipment of video bit rate adaptive network
By adopting meta-training and meta-testing methods based on historical experience in video bit rate adaptive networks, the problem that traditional methods are difficult to adapt to complex network environments is solved, and efficient adaptation to diversified high-speed fluctuating network environments is achieved.
Patent Information
- Application Number
- CN202210762758.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-06-30
AI Technical Summary
Traditional video bit rate adaptive methods are difficult to adapt to complex and diverse network environments, resulting in the inability to work properly in wireless networks.
Through a meta-training method based on historical experience, the video bit rate adaptive network is initialized so that it can quickly learn and generate the optimal video bit rate that is adapted to the network state group. The network parameters are then fine-tuned based on metatests to optimize the adaptability to a diverse high-speed fluctuating network environment.
It realizes efficient adaptation of video code rate adaptive network in a diverse high-speed fluctuating network environment, and improves the stability and quality of video transmission.
Smart Images

Figure CN115499657B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a training method, application method, device and equipment for a video bit rate adaptive network. Background Art
[0002] With the rapid development of information technology, emerging applications such as the Metaverse are constantly emerging. Since emerging applications such as the Metaverse require the interactive delay (motion-to-photon latency) to be at least less than 20ms, the buffer size of Internet middleware and Internet endpoints is greatly limited, and the frequency and amplitude of end-to-end network fluctuations are significantly increased. Traditional bitrate adaptation methods usually rely on empirical network congestion signals to adjust encoding and decoding parameters and network card sending rates. They are difficult to adapt to wireless network environments with drastic channel fading fluctuations, resulting in the inability of traditional bitrate adaptation methods to work properly.
[0003] Therefore, when it is assumed that user mobility is significantly slower than the change in channel characteristics, the industry proposes a learning-based rate adaptation method that expects to automatically summarize the channel fading characteristics of the current user through a deep neural network to automatically summarize a better rate adjustment strategy. This method usually uses a fixed neural network mode and does not make further adjustments to network parameters during the application phase. However, this method is difficult to meet the current increasingly diverse network conditions (such as wifi, 4G, 5G, wired, etc.), complex communication environments (such as mobility, indoor / outdoor, population density, etc.) and other influencing factors (such as preferences, etc.), resulting in the rate adaptation method's ability to adapt to complex environments such as diverse network conditions. Summary of the invention
[0004] The embodiments of the present application provide a training method, application method, device and equipment for a video bit rate adaptive network, which are used to perform meta-training on an initialized video bit rate adaptive network based on historical experience, so that the initialized video bit rate adaptive network has the ability to quickly learn and quickly generate an optimal video bit rate that adapts to a network state group, and then, can further fine-tune the network parameters of the video bit rate adaptive network to be adjusted based on meta-testing, so as to optimize the adaptability of the target video bit rate adaptive network to a diverse, high-speed, fluctuating network environment.
[0005] On the one hand, an embodiment of the present application provides a method for training a video bit rate adaptive network, including:
[0006] Based on N meta-learning tasks, N first network state groups of different classes are sampled, and the N first network state groups of different classes are assigned to N initialized video bit rate adaptive networks to obtain N first networks, wherein each first network state group is used to simulate a first network environment of each first network, and the initialized video bit rate adaptive network is set with an initialized network parameter, and N is an integer greater than 1;
[0007] Based on the decision code rate output by each first network and the first network environment, a first reward value corresponding to each decision code rate is obtained to perform an inner loop update on the initialized network parameters of the first network;
[0008] Resampling N second network state groups of different classes corresponding to the N meta-learning tasks, and assigning the N second network state groups of different classes to the N first networks updated through the inner loop, wherein each second network state group is used to simulate the second network environment of each first network updated through the inner loop;
[0009] Based on each decision bit rate output by the first network after inner loop update and the second network environment, a second reward value corresponding to each decision bit rate is obtained to perform outer loop update on the initialized network parameters of the initialized video bit rate adaptive network;
[0010] Repeat the operations of sampling the first network state group, obtaining the first reward value, inner loop updating, sampling the second network state group, obtaining the second reward value, and outer loop updating until the reward values of the N meta-learning tasks after the inner loop updating meet the convergence condition, and initialize the video bitrate adaptive network update to obtain the video bitrate adaptive network to be adjusted;
[0011] Obtaining a current category of network status group corresponding to the target video transmission process, wherein the current category of network status group is used to simulate the current category of network environment of the video bit rate adaptive network to be adjusted;
[0012] Based on the decision bitrate output by the video bitrate adaptive network to be adjusted and the current category network environment, the reward value corresponding to each decision bitrate is obtained to perform an inner loop update on the network parameters of the video bitrate adaptive network to be adjusted until the reward value meets the convergence condition, thereby obtaining the target video bitrate adaptive network corresponding to the current category.
[0013] On the other hand, the present application provides an application method of a video bit rate adaptive network, including:
[0014] Obtain the current network status group, current observation quantity and target decision bit rate at the current moment corresponding to the target video transmission process;
[0015] If the category of the current network state group has not changed, the current observation amount and the target decision bit rate at the current moment are input into the target decision network in the target video bit rate adaptive network, and the target decision bit rate at the next moment is output through the target decision network;
[0016] Adjust the sending rate and the video editing bit rate based on the target decision bit rate at the next moment;
[0017] The target video is transmitted based on the adjusted sending rate and video editing bit rate.
[0018] On the other hand, the present application provides a training device for a video bit rate adaptive network, comprising:
[0019] A processing unit, configured to sample N first network state groups of different classes based on N meta-learning tasks, and assign the N first network state groups of different classes to N initialized video bit rate adaptive networks to obtain N first networks, wherein each first network state group is used to simulate a first network environment of each first network, and the initialized video bit rate adaptive network is provided with an initialized network parameter, and N is an integer greater than 1;
[0020] An acquisition unit, configured to acquire a first reward value corresponding to each decision code rate based on the decision code rate output by each first network and the first network environment, so as to perform an inner loop update on the initialized network parameters of the first network;
[0021] The processing unit is further used to resample N second network state groups of different classes corresponding to the N meta-learning tasks, and distribute the N second network state groups of different classes to the N first networks updated through the inner loop, wherein each second network state group is used to simulate the second network environment of each first network updated through the inner loop;
[0022] The acquisition unit is further used to acquire a second reward value corresponding to each decision bit rate based on the decision bit rate output by each first network after inner loop update and the second network environment, so as to perform outer loop update on the initialized network parameters of the initialized video bit rate adaptive network;
[0023] The processing unit is further used to repeatedly perform the operations of sampling the first network state group, obtaining the first reward value, inner loop updating, sampling the second network state group, obtaining the second reward value, and outer loop updating until the reward values of the N meta-learning tasks after the inner loop updating meet the convergence condition, and initialize the video bit rate adaptive network update to obtain the video bit rate adaptive network to be adjusted;
[0024] The acquisition unit is further used to acquire a network status group of the current category corresponding to the target video transmission process, wherein the network status group of the current category is used to simulate the current category network environment of the video bit rate adaptive network to be adjusted;
[0025] The acquisition unit is further used to obtain a reward value corresponding to each decision bit rate based on the decision bit rate output by the video bit rate adaptive network to be adjusted and the current category network environment, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted until the reward value meets the convergence condition, thereby obtaining a target video bit rate adaptive network corresponding to the current category.
[0026] In a possible design, in an implementation of another aspect of the embodiment of the present application,
[0027] The acquisition unit is further used to acquire a first historical network trajectory, wherein the first historical network trajectory includes a throughput, a packet loss rate, and a round-trip delay of the network in a first historical time period;
[0028] The processing unit is further used to perform bandwidth estimation and filtering processing on the first historical network trajectory based on the time dimension to obtain a mean, a standard deviation, and a fluctuation difference corresponding to each time window;
[0029] A determination unit, used to determine a mean interval threshold, a standard deviation interval threshold, and a fluctuation difference interval threshold based on the mean, standard deviation, and fluctuation difference corresponding to each time window;
[0030] The processing unit can be specifically used to: based on N meta-learning tasks, mean interval thresholds, standard deviation interval thresholds, and fluctuation difference interval thresholds, sample N different types of first network state groups, and assign the N different types of first network state groups to N initialized video bit rate adaptive networks to obtain N first networks.
[0031] In a possible design, in an implementation of another aspect of the embodiment of the present application, the processing unit may be specifically used for:
[0032] Based on the time dimension, the bandwidth of the second historical network trajectory is estimated and filtered to obtain the unit time bandwidth corresponding to each time window;
[0033] Based on the unit time bandwidth corresponding to each time window, the network state group distribution probability statistics are performed on the second historical network trajectory to obtain a first probability density function;
[0034] Sample N meta-learning tasks based on the first probability density function;
[0035] Based on N meta-learning tasks, a mean interval threshold, a standard deviation interval threshold, and a fluctuation difference interval threshold, a network trajectory generator is used to generate K bandwidth trajectories corresponding to each meta-learning task to obtain N first network state groups of different classes, wherein each first network state group includes K bandwidth trajectories, and K is an integer greater than 1.
[0036] In a possible design, in an implementation of another aspect of the embodiment of the present application,
[0037] The processing unit is further used to perform distribution probability statistics of the network state group on the network trajectory within the target time period during the target video transmission process to obtain a current probability density function corresponding to the target time period;
[0038] The processing unit is further used to calculate the cross entropy between the current probability density function and the first probability density function;
[0039] The acquisition unit is further used to re-update the initialization network parameters of the initialized video bit rate adaptive network based on the current probability density function if the cross entropy is greater than the cross threshold, so as to obtain a new video bit rate adaptive network to be adjusted.
[0040] In a possible design, in an implementation of another aspect of the embodiment of the present application, the processing unit may be specifically used for:
[0041] Based on the unit time bandwidth corresponding to each time window and the definition of the network state group, a network state scatter plot corresponding to the second historical network trajectory is mapped;
[0042] The density of the scattered points in the network state scatter diagram is calculated and linear interpolated to obtain a first probability density function.
[0043] In a possible design, in an implementation of another aspect of the embodiment of the present application, the acquisition unit may be specifically used to:
[0044] Simulating a first network environment based on the first network state group, wherein each first network environment includes K first bandwidth trajectories of different time lengths;
[0045] For each first network, K observations and K decision code rates at the previous moment are input into the first decision network of the first network, and K decision code rates at the current moment are output through the first decision network;
[0046] The K decision code rates at the current moment interact with the K first bandwidth trajectories at the current moment in the first network environment respectively to obtain K observation quantities at the current moment;
[0047] Inputting K observation quantities and K decision code rates at the current moment into the first evaluation network of the first network, and outputting the evaluation value at the next moment through the first evaluation network;
[0048] Based on the K observations at the current moment, K first reward values corresponding to the K decision code rates at the current moment are calculated;
[0049] Based on the K first reward values corresponding to the K decision code rates at the current moment, the initialized network parameters of the first decision network and the first evaluation network are updated to perform an inner loop update on the initialized network parameters of the first network.
[0050] In a possible design, in an implementation of another aspect of the embodiment of the present application, the processing unit may be specifically used for:
[0051] Resample N new meta-learning tasks based on the first probability density function;
[0052] Based on N new meta-learning tasks, a mean interval threshold, a standard deviation interval threshold, and a fluctuation difference interval threshold, a network trajectory generator is used to generate K new bandwidth trajectories corresponding to each new meta-learning task to obtain N second network state groups of different classes, where each second network state group includes K new bandwidth trajectories.
[0053] In a possible design, in an implementation of another aspect of the embodiment of the present application, the acquisition unit may be specifically used to:
[0054] Simulating a second network environment based on the second network state group, wherein each second network environment includes K second bandwidth trajectories of different time lengths;
[0055] For each first network updated through the inner loop, K observations and K decision code rates at the previous moment are input into the first decision network of the first network updated through the inner loop, and K new decision code rates at the current moment are output;
[0056] The K new decision code rates at the current moment interact with the K second bandwidth trajectories at the current moment in the second network environment respectively to obtain K new observation quantities at the current moment;
[0057] Based on the K new observations at the current moment, K second reward values corresponding to the K new decision code rates at the current moment are calculated;
[0058] The K second reward values corresponding to the K new decision code rates at the current moment based on the N meta-learning tasks are added and averaged to obtain the average reward value at the current moment;
[0059] The initialized network parameters of the initialized video bitrate adaptive network are updated based on the average reward value at the current moment, so as to perform an outer loop update on the initialized network parameters of the initialized video bitrate adaptive network.
[0060] In a possible design, in an implementation of another aspect of the embodiment of the present application,
[0061] The processing unit is further used to calculate the mean difference, standard deviation difference and fluctuation difference between the network status group of the current category and the network status group of the previous time during the target video transmission process;
[0062] The processing unit is further used to determine that the category of the network state group of the current category has changed if any one of the mean difference, the standard deviation difference or the fluctuation difference difference satisfies the change condition of the interval range, and execute the decision bit rate output based on the video bit rate adaptive network to be adjusted and the current category network environment, obtain the reward value corresponding to each decision bit rate, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted, until the reward value satisfies the convergence condition, and obtain the target video bit rate adaptive network corresponding to the current category;
[0063] The processing unit is also used to determine that the category of the network status group of the current category has not changed if the mean difference, the standard deviation difference and the fluctuation difference do not meet the change conditions of the interval range, and then use the target video bit rate adaptive network corresponding to the network status group at the previous time as the target video bit rate adaptive network corresponding to the current category.
[0064] In a possible design, in an implementation of another aspect of the embodiment of the present application, the acquisition unit may be specifically used to:
[0065] Simulating a current category network environment based on a current category network state group, wherein the current category network environment includes K current bandwidth trajectories of different time lengths;
[0066] Obtain the current observation value and current decision bit rate corresponding to the target video transmission process;
[0067] Inputting the current observation amount and the current decision bit rate into the decision network to be adjusted of the video bit rate adaptive network to be adjusted, and outputting the decision bit rate at the next moment through the decision network to be adjusted;
[0068] The decision code rate at the next moment is interacted with the K current bandwidth trajectories at the next moment in the current category network environment to obtain K observations at the next moment;
[0069] Inputting the K observation quantities at the next moment and the decision bit rate at the next moment into the evaluation network to be adjusted of the video bit rate adaptive network to be adjusted, and outputting the evaluation value through the evaluation network to be adjusted;
[0070] Based on the K observations at the next moment, the K reward values at the next moment are calculated;
[0071] Based on the K reward values at the next moment, the network parameters of the decision network to be adjusted and the evaluation network to be adjusted are updated, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted.
[0072] In a possible design, in an implementation of another aspect of the embodiment of the present application,
[0073] The processing unit is further used to compare the current packet loss rate and the current round-trip delay corresponding to the network status group of the current category with the packet loss threshold and the delay threshold respectively;
[0074] The processing unit is also used to start the protection fallback mechanism and suspend the use of the decision network of the video bit rate adaptive network to be adjusted for bit rate selection when the current packet loss rate is greater than the packet loss threshold or the current round-trip delay is greater than the delay threshold.
[0075] On the other hand, the present application provides an application device of a video bit rate adaptive network, including:
[0076] An acquisition unit, used to acquire the current network state group, the current observation amount and the target decision bit rate at the current moment corresponding to the target video transmission process;
[0077] A processing unit, configured to input the current observation amount and the target decision bit rate at the current moment into the target decision network in the target video bit rate adaptation network if the category of the current network state group has not changed, and output the target decision bit rate at the next moment through the target decision network;
[0078] A determination unit, configured to adjust a transmission rate and a video editing bit rate based on a target decision bit rate at a next moment;
[0079] The processing unit is further used to transmit the target video based on the adjusted sending rate and video editing bit rate.
[0080] In a possible design, in an implementation of another aspect of the embodiment of the present application,
[0081] The processing unit is further configured to simulate the current network environment of the video bit rate adaptive network to be adjusted based on the network state group of the current category if the category of the current network state group changes;
[0082] The acquisition unit is further used to acquire a reward value corresponding to each decision bit rate based on the decision bit rate output by the video bit rate adaptive network to be adjusted and the current network environment, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted until the reward value meets the convergence condition, thereby obtaining a new target video bit rate adaptive network corresponding to the current network state group;
[0083] The processing unit is further used to input the current observation amount and the target decision bit rate at the current moment into the target decision network in the new target video bit rate adaptive network, and output the target decision bit rate at the next moment through the target decision network;
[0084] The determination unit is further used to adjust the transmission rate and the video editing bit rate based on the target decision bit rate at the next moment;
[0085] The processing unit is further used to transmit the target video based on the adjusted sending rate and video editing bit rate.
[0086] On the other hand, the present application provides a computer device, including: a memory, a processor, and a bus system;
[0087] Wherein, the memory is used to store programs;
[0088] The processor is used to implement the above-mentioned methods when executing the program in the memory;
[0089] The bus system is used to connect the memory and the processor so that the memory and the processor can communicate with each other.
[0090] Another aspect of the present application provides a computer-readable storage medium, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer is enabled to execute the above-mentioned methods.
[0091] It can be seen from the above technical solutions that the embodiments of the present application have the following beneficial effects:
[0092] By sampling N different types of first network state groups and distributing them to N initialized video bit rate adaptive networks, N first networks are obtained to simulate the first network environment of each first network, so that the first reward value corresponding to each decision bit rate can be obtained based on the decision bit rate output by each first network and the first network environment, so as to perform an inner loop update on the initialized network parameters of the first network, and re-sample N different types of second network state groups to simulate the second network environment of each first network updated through the inner loop, so that the second reward value corresponding to each decision bit rate can be obtained based on the decision bit rate output by each first network updated through the inner loop and the second network environment, An outer loop is performed to update the initialized network parameters of the initialized video bit rate adaptive network, and then the above operation is repeated until the reward value after the inner loop update meets the convergence condition, so as to obtain the video bit rate adaptive network to be adjusted, and then, the network state group corresponding to the current category during the target video transmission process is obtained to simulate the current category network environment, and the reward value corresponding to each decision bit rate can be obtained based on the decision bit rate output by the video bit rate adaptive network to be adjusted and the current category network environment, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted, until the reward value meets the convergence condition, and the target video bit rate adaptive network corresponding to the current category is obtained. Through the above-mentioned method, the first reward value corresponding to each decision bit rate can be obtained in an offline state, and the initialized network parameters of the first network can be updated in an inner loop, and the second reward value corresponding to each decision bit rate can be obtained, and the initialized network parameters of the initialized video bit rate adaptive network can be updated in an outer loop, so as to realize meta-training of the initialized video bit rate adaptive network based on historical experience, so that the initialized video bit rate adaptive network has the ability to quickly learn and quickly generate the best video bit rate for the adaptive network state group, and then, the network state group of the current category can be detected in real time in an online state to obtain the reward value, and the network parameters of the video bit rate adaptive network to be adjusted can be updated in an inner loop to obtain the target video bit rate adaptive network, and the network parameters of the video bit rate adaptive network to be adjusted can be further fine-tuned based on the meta-test to optimize the adaptability of the target video bit rate adaptive network to a diversified high-speed fluctuating network environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] Figure 1 This is a schematic diagram of the architecture of a video data control system in an embodiment of the present application;
[0094] Figure 2 is a flow chart of an embodiment of a training method for a video bit rate adaptive network in an embodiment of the present application;
[0095] Figure 3 is another embodiment flow chart of the training method of the video bit rate adaptive network in the embodiment of the present application;
[0096] Figure 4 is another embodiment flow chart of the training method of the video bit rate adaptive network in the embodiment of the present application;
[0097] Figure 5 is another embodiment flow chart of the training method of the video bit rate adaptive network in the embodiment of the present application;
[0098] Figure 6 is another embodiment flow chart of the training method of the video bit rate adaptive network in the embodiment of the present application;
[0099] Figure 7 is another embodiment flow chart of the training method of the video bit rate adaptive network in the embodiment of the present application;
[0100] Figure 8 is another embodiment flow chart of the training method of the video bit rate adaptive network in the embodiment of the present application;
[0101] Fig. 9 is another embodiment flow chart of the training method of the video bit rate adaptive network in the embodiment of the present application;
[0102] Fig.10 is another embodiment flow chart of the training method of the video bit rate adaptive network in the embodiment of the present application;
[0103] Fig.11 is another embodiment flow chart of the training method of the video bit rate adaptive network in the embodiment of the present application;
[0104] Fig.12 is another embodiment flow chart of the training method of the video bit rate adaptive network in the embodiment of the present application;
[0105] Fig.13 This is a flow chart of an embodiment of the application method of the video bit rate adaptive network in the embodiment of the present application;
[0106] Fig.14 is another embodiment flow chart of the application method of the video bit rate adaptive network in the embodiment of the present application;
[0107] Fig.15 It is a schematic diagram of a principle flow of a training method for a video bit rate adaptive network in an embodiment of the present application;
[0108] Fig.16 It is a schematic diagram of a meta-training process of a training method for a video bit rate adaptive network in an embodiment of the present application;
[0109] Fig.17 It is a network update schematic diagram of the training method of the video bit rate adaptive network in the embodiment of the present application;
[0110] Fig.18 It is a schematic diagram of a meta-test process of a training method for a video bit rate adaptive network in an embodiment of the present application;
[0111] Fig.19 It is a schematic diagram of a preprocessing flow of a training method for a video bit rate adaptive network in an embodiment of the present application;
[0112] Fig. 20 It is a schematic diagram of an interactive video transmitted through a video bit rate in a training method of a video bit rate adaptive network in an embodiment of the present application;
[0113] FIG. 21( a ) is a schematic diagram of an interactive video interface of a method for training a video bitrate adaptive network in an embodiment of the present application;
[0114] FIG. 21( b ) is a schematic diagram of another interactive video interface of the video bit rate adaptive network training method in an embodiment of the present application;
[0115] FIG. 21( c ) is a schematic diagram of an interface of a video application of a training method for a video bit rate adaptive network in an embodiment of the present application;
[0116] Fig. 22 It is a schematic diagram of an embodiment of a training device for a video bit rate adaptive network in an embodiment of the present application;
[0117] Fig.23 This is a schematic diagram of an embodiment of an application device of a video bit rate adaptive network in an embodiment of the present application;
[0118] Fig.24 It is a schematic diagram of an embodiment of a computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0119] The embodiments of the present application provide a training method, application method, device and equipment for a video bit rate adaptive network, which are used to perform meta-training on an initialized video bit rate adaptive network based on historical experience, so that the initialized video bit rate adaptive network has the ability to quickly learn and quickly generate an optimal video bit rate that adapts to a network state group, and then, can further fine-tune the network parameters of the video bit rate adaptive network to be adjusted based on meta-testing, so as to optimize the adaptability of the target video bit rate adaptive network to a diverse, high-speed, fluctuating network environment.
[0120] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "corresponding to" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0121] It can be understood that in the specific implementation of the present application, related data such as the first historical network trajectory, the second historical status trajectory and the network status group are involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0122] It is understandable that the training method of the video bit rate adaptive network disclosed in the present application involves cloud technology. The cloud technology is further introduced below. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network in a wide area network or a local area network to realize the calculation, storage, processing, and sharing of data. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model application, which can form a resource pool, which is used on demand and is flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, each item may have its own identification mark in the future, and all need to be transmitted to the background system for logical processing. Data of different levels will be processed separately, and all kinds of industry data require strong system backing support, which can only be achieved through cloud computing.
[0123] Cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called a "cloud". From the user's perspective, the resources in the "cloud" are infinitely scalable and can be obtained at any time, used on demand, expanded at any time, and paid for by use.
[0124] As a provider of basic cloud computing capabilities, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service) platform) will be established, and various types of virtual resources will be deployed in the resource pool for external customers to choose to use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.
[0125] According to the logical function division, the PaaS (Platform as a Service) layer can be deployed on the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on the PaaS layer. SaaS can also be deployed directly on IaaS. PaaS is a platform for software operation, such as databases, web containers, etc. SaaS is a variety of transaction software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.
[0126] Secondly, cloud security refers to the general term for security software, hardware, users, organizations, and security cloud platforms based on cloud computing business model applications. Cloud security integrates emerging technologies and concepts such as parallel processing, grid computing, and unknown virus behavior judgment. Through a large number of networked clients, it monitors abnormal software behavior in the network, obtains the latest information on Trojans and malicious programs on the Internet, and sends it to the server for automatic analysis and processing, and then distributes virus and Trojan solutions to each client.
[0127] Secondly, cloud storage is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (storage devices are also called storage nodes) in the network through application software or application interfaces to work together and provide external data storage and transaction access functions.
[0128] At present, the storage method of the storage system is: create a logical volume, and when creating a logical volume, allocate physical storage space for each logical volume. The physical storage space may be composed of disks of a storage device or several storage devices. The client stores data on a logical volume, that is, stores the data on the file system. The file system divides the data into many parts, each of which is an object. The object contains not only data but also additional information such as data identification (ID, ID entity). The file system writes each object into the physical storage space of the logical volume, and the file system records the storage location information of each object, so that when the client requests to access the data, the file system can allow the client to access the data according to the storage location information of each object.
[0129] The process of the storage system allocating physical storage space to a logical volume is as follows: based on the estimated capacity of the objects stored in the logical volume (this estimate often has a large margin relative to the actual capacity of the objects to be stored) and the grouping of independent redundant disk arrays (RAID, Redundant Array of Independent Disks), the physical storage space is pre-divided into stripes. A logical volume can be understood as a stripe, thereby allocating physical storage space to the logical volume.
[0130] It is understandable that the training method of the video bitrate adaptive network disclosed in the present application also involves artificial intelligence (AI) technology, and artificial intelligence technology is further introduced below. Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making.
[0131] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0132] Secondly, natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0133] Secondly, Machine Learning (ML) is a multi-disciplinary interdisciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0134] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, drones, robots, smart medical care, smart customer service, etc. I believe that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0135] It should be understood that the training method of the video bitrate adaptive network provided in the present application can be applied to various scenarios, including but not limited to artificial intelligence, cloud technology, maps, smart transportation, etc., and is used to complete the transmission of the target video by obtaining the decision bitrate adapted to each network state, so as to be applied to scenarios such as multimedia live broadcast, cloud game applications, intelligent video interaction systems, and intelligent video calls.
[0136] In order to solve the above problems, this application proposes a video bit rate adaptive network training method, which is applied to Figure 1 The video data control system shown is shown in Figure 1 , Figure 1 FIG. 1 is a schematic diagram of the architecture of a video data control system in an embodiment of the present application. Figure 1As shown, the server distributes N different types of first network state groups provided by the sampling terminal device to N initialized video bit rate adaptive networks to obtain N first networks to simulate the first network environment of each first network, so that the first reward value corresponding to each decision bit rate can be obtained based on the decision bit rate output by each first network and the first network environment, so as to perform an inner loop update on the initialized network parameters of the first network, and resample N different types of second network state groups to simulate the second network environment of each first network updated by the inner loop, so that the decision bit rate output by each first network updated by the inner loop and the second network environment can be obtained. The second reward value is used to perform an outer loop update on the initialized network parameters of the initialized video bit rate adaptive network, and then the above operation is repeated until the reward value after the inner loop update meets the convergence condition, so as to obtain the video bit rate adaptive network to be adjusted, and then, obtain the network state group of the current category corresponding to the target video transmission process to simulate the current category network environment, and obtain the reward value corresponding to each decision bit rate based on the decision bit rate output by the video bit rate adaptive network to be adjusted and the current category network environment, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted, until the reward value meets the convergence condition, and obtain the target video bit rate adaptive network corresponding to the current category. Through the above-mentioned method, the first reward value corresponding to each decision bit rate can be obtained in an offline state, and the initialized network parameters of the first network can be updated in an inner loop, and the second reward value corresponding to each decision bit rate can be obtained, and the initialized network parameters of the initialized video bit rate adaptive network can be updated in an outer loop, so as to realize meta-training of the initialized video bit rate adaptive network based on historical experience, so that the initialized video bit rate adaptive network has the ability to quickly learn and quickly generate the best video bit rate for the adaptive network state group, and then, the network state group of the current category can be detected in real time in an online state to obtain the reward value, and the network parameters of the video bit rate adaptive network to be adjusted can be updated in an inner loop to obtain the target video bit rate adaptive network, and the network parameters of the video bit rate adaptive network to be adjusted can be further fine-tuned based on the meta-test to optimize the adaptability of the target video bit rate adaptive network to a diversified high-speed fluctuating network environment.
[0137] Understandably, Figure 1 Only one type of terminal device is shown in the figure. In actual scenarios, more types of terminal devices may participate in the data processing process. Terminal devices include but are not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, car terminals, etc. The specific number and type depend on the actual scenario and are not limited here. Figure 1One server is shown in the figure, but in actual scenarios, multiple servers may also be involved, especially in the scenario of multi-model training interaction. The number of servers depends on the actual scenario and is not limited here.
[0138] It should be noted that in this embodiment, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The terminal device and the server can be directly or indirectly connected through wired or wireless communication, and the terminal device and the server can be connected to form a blockchain network, which is not limited in this application.
[0139] Combined with the above introduction, the training method of the video bit rate adaptive network in this application will be introduced below. Figure 2 In one embodiment of the present application, a method for training a video bit rate adaptive network includes:
[0140] In step S101, based on N meta-learning tasks, N different types of first network state groups are sampled, and the N different types of first network state groups are assigned to N initialized video bit rate adaptive networks to obtain N first networks, wherein each first network state group is used to simulate a first network environment of each first network, and the initialized video bit rate adaptive network is set with an initialized network parameter, and N is an integer greater than 1;
[0141] In this embodiment, in order to better optimize the network parameters of the initialized video bit rate adaptive network to improve the adaptability of the video bit rate adaptive network, as shown in FIG. Fig.15 As shown in Figure 2, meta-learning includes a meta-training phase and a meta-testing phase. Fig.15 The semi-offline stage shown is the meta-training stage of meta-learning. Based on N meta-learning tasks, N first network state groups of different classes can be sampled, and the N first network state groups of different classes can be assigned to N initialized video bitrate adaptive networks to obtain N first networks, so that the initialization network parameters of the initialized video bitrate adaptive network can be better updated based on the first network state groups.
[0142] Among them, the initialized video rate adaptive network is a network constructed based on strong learning technology, the initialized video rate adaptive network includes a decision network and an evaluation network, and the initialized video rate adaptive network is set with initialized network parameters. Each first network state group is used to simulate the first network environment of each first network, N is an integer greater than 1, and a meta-learning task corresponds to a class of first network state groups. The meta-learning task is based on meta-learning technology, through offline learning of different tasks, and can adapt to different categories of network states through less sample training and rapid update of neural network parameters.
[0143] Specifically, based on N meta-learning tasks, N different types of first network state groups are sampled. Specifically, based on the time dimension, bandwidth estimation and filtering are performed on the historical network trajectories collected within a period of time (such as the second historical network trajectory, including the network throughput, packet loss rate, and round-trip delay within a period of time) to obtain the unit time bandwidth corresponding to each time window. Then, based on the unit time bandwidth corresponding to each time window and the definition of the network state group, the historical network trajectories collected within a period of time (such as the second historical network trajectory, including the network throughput, packet loss rate, and round-trip delay within a period of time) are subjected to network state group distribution probability statistics (such as Fig.16 The network state group distribution probability statistics shown in the figure) is used to obtain the first probability density function.
[0144] Furthermore, N meta-learning tasks can be sampled based on the first probability density function, and based on the definition of the N meta-learning tasks, a network trajectory generator (such as Fig.16 The network trajectory generator shown in FIG. 1 is used to simulate K bandwidth trajectories (such as 40ms bandwidth trajectory, 80ms bandwidth trajectory, etc.) corresponding to each meta-learning task to obtain N first network state groups of different classes. Then, the N first network state groups of different classes can be assigned to N initialized video bitrate adaptive networks, and the network environment of each initialized video bitrate adaptive network can be simulated to obtain N first networks.
[0145] In step S102, based on the decision code rate output by each first network and the first network environment, a first reward value corresponding to each decision code rate is obtained to perform an inner loop update on the initialized network parameters of the first network;
[0146] In this embodiment, after obtaining N different types of first network state groups, the first reward value corresponding to each decision code rate can be calculated based on the decision code rate output by each first network and the first network environment, and then the initialized network parameters of the first network can be updated based on the first reward value, thereby realizing an inner loop update of the initialized network parameters of the first network.
[0147] Specifically, Fig.16 The meta-training update initialization parameter shown in the figure can obtain the first reward value corresponding to each decision code rate based on the decision code rate output by each first network and the first network environment after obtaining the first network state groups of N different classes. Specifically, it can be as follows Fig.17 As shown, the first network environment is simulated based on the first network state group to simulate K first bandwidth trajectories of different time lengths, and then, for each first network, the first bandwidth trajectory based on Fig.17 The K decision code rates of the last moment output by a first decision network in the decision network shown in the figure, and the K observations of the last moment obtained by interacting the K decision code rates of the last moment with the first bandwidth trajectories of K different time lengths (such as Fig.17 The throughput, delay, delay jitter, and packet loss rate (as shown in the figure) are input to the first decision network of the first network, and K decision code rates at the current moment are output through the first decision network.
[0148] Further, the K decision code rates at the current moment are respectively interacted with the K first bandwidth trajectories at the current moment in the first network environment to obtain K observations at the current moment. Then, the K observations at the current moment and the K decision code rates can be input into Fig.17 The first evaluation network in the evaluation network shown in the figure outputs the evaluation value of the next moment through the first evaluation network. At the same time, the K first reward values corresponding to the K decision code rates at the current moment can be calculated based on the K observations at the current moment. Specifically, it can be based on the preset smoothness and each observation in the K observations, where each observation includes throughput, delay and packet loss rate, and the reward formula: reward = + throughput (Mbps) - μ delay (s) - ρ packet loss rate - σ smoothness (Mbps) is used to calculate the K first reward values. Further, the initialization network parameters of the first decision network and the first evaluation network can be updated based on the K first reward values corresponding to the K decision code rates at the current moment, so as to perform an inner loop update on the initialization network parameters of the first network to maximize the reward.
[0149] In step S103, N second network state groups of different categories corresponding to the N meta-learning tasks are resampled, and the N second network state groups of different categories are assigned to the N first networks that have been updated through the inner loop, wherein each second network state group is used to simulate the second network environment of each first network that has been updated through the inner loop;
[0150] In this embodiment, after the initialized network parameters of the first network are updated in an inner loop, N different types of second network state groups corresponding to N meta-learning tasks can be resampled based on the first probability density function, so that the initialized network parameters of the initialized video bitrate adaptive network can be better updated based on the second network state group.
[0151] Each second network state group is used to simulate the second network environment of each first network updated through the inner loop.
[0152] Specifically, resampling the N different types of second network state groups corresponding to the N meta-learning tasks may specifically be resampling the N new meta-learning tasks based on the first probability density function, and based on the definition of the N new meta-learning tasks, through a network trajectory generator (such as Fig.16 The network trajectory generator shown in the figure is used to simulate K new bandwidth trajectories (such as 60ms bandwidth trajectory, 90ms bandwidth trajectory, etc.) corresponding to each new meta-learning task to obtain N second network state groups of different categories. Further, after obtaining N second network state groups of different categories, the N second network state groups of different categories can be assigned to N first networks that have been updated through the inner loop to simulate the second network environment of each first network that has been updated through the inner loop.
[0153] In step S104, based on the decision bitrate output by each first network after inner loop update and the second network environment, a second reward value corresponding to each decision bitrate is obtained to perform outer loop update on the initialized network parameters of the initialized video bitrate adaptive network;
[0154] In this embodiment, after obtaining N different types of second network state groups corresponding to N meta-learning tasks, the second reward value corresponding to each decision bit rate can be calculated based on the decision bit rate output by each first network after the inner loop update and the second network environment. Then, the initialized network parameters of the initially constructed initialized video bit rate adaptive network can be updated based on the second reward values of the N meta-learning tasks, so as to perform an outer loop update on the initialized network parameters of the initialized video bit rate adaptive network.
[0155] Specifically, Fig.16 The meta-training update initialization parameters shown in the figure can obtain the second reward value corresponding to each decision code rate based on the decision code rate output by each first network after inner loop update and the second network environment after obtaining N different types of second network state groups. Specifically, it can be as follows Fig.17As shown, firstly, based on the second network state group, the second network environment of the first network after the inner loop update is simulated to simulate and obtain K second bandwidth trajectories of different time lengths. Then, for each first network after the inner loop update, the second bandwidth trajectories based on Fig.17 The K decision code rates of the last moment output by a first decision network after inner loop update in the decision network shown in the figure, and the K observations of the last moment obtained by interacting the K decision code rates of the last moment with the K second bandwidth trajectories of different time lengths (such as Fig.17 The throughput, delay, delay jitter and packet loss rate (as shown in the figure) are input to the first decision network of the decision network that has been updated through the inner loop at the current moment, and the first decision network that has been updated through the inner loop at the current moment outputs K new decision code rates at the current moment.
[0156] Further, the K new decision code rates at the current moment are respectively interacted with the K second bandwidth trajectories at the current moment in the second network environment to obtain K new observation quantities at the current moment. Then, the K new observation quantities and the K new decision code rates at the current moment can be input into Fig.17 The first evaluation network in the evaluation network after the inner loop update at the current moment shown in the diagram outputs the evaluation value of the next moment through the first evaluation network after the inner loop update at the current moment. At the same time, the K second reward values corresponding to the K new decision bit rates at the current moment can be calculated based on the K new observations at the current moment. Specifically, the K second reward values can be calculated based on the preset smoothness and the throughput, delay and packet loss rate included in each of the K new observations, using the above reward formula. Furthermore, the K second reward values corresponding to the K new decision bit rates at the current moment based on the N new meta-learning tasks can be added and averaged to obtain the reward average value at the current moment, and then, based on the reward average value at the current moment, the initialization network parameters of the initialized video bit rate adaptive network initially constructed are updated to perform an outer loop update on the initialization network parameters of the initialized video bit rate adaptive network.
[0157] In step S105, the operations of sampling the first network state group, obtaining the first reward value, inner loop updating, sampling the second network state group, obtaining the second reward value, and outer loop updating are repeatedly performed until the reward values of the N meta-learning tasks after the inner loop updating meet the convergence condition, and the video bit rate adaptive network is initialized to update and obtain the video bit rate adaptive network to be adjusted;
[0158] Specifically, the operations of sampling the first network state group, obtaining the first reward value, inner loop updating, sampling the second network state group, obtaining the second reward value, and outer loop updating are repeatedly performed until the reward values of the N meta-learning tasks after the inner loop updating meet the convergence conditions, such as until the mean value of the reward of each task no longer increases after the inner loop, that is, convergence is reached, then the video bitrate adaptive network is initialized to update and obtain the video bitrate adaptive network to be adjusted.
[0159] For example, Fig.17 As shown, the input of each first network includes the past packet loss rate sequence, delay sequence, delay jitter sequence, throughput sequence and bit rate selection sequence, and each first decision network of each first network outputs the target bit rate, and interacts with the video transmission environment (such as the first network environment) to obtain a reward value (such as the first reward value), and then the initialization network parameters of each first network are updated, that is, the inner loop update, and then, the input of each first network updated by the inner loop includes the past packet loss rate sequence, delay sequence, delay jitter sequence, throughput sequence and bit rate selection sequence, and each first decision network of each first network updated by the inner loop outputs a new target bit rate, and interacts with the video transmission environment (such as the second network environment) to obtain a reward value (such as the second reward value), and then the initialization network parameters of the initially constructed initialization video bit rate adaptive network are updated, that is, the outer loop update. Among them, the reward value (such as the first reward value and the second reward value, etc.) is affected by throughput, delay, packet loss rate and smoothness.
[0160] In step S106, a current category network state group corresponding to the target video transmission process is obtained, wherein the current category network state group is used to simulate the current category network environment of the video bit rate adaptive network to be adjusted;
[0161] In this embodiment, after the video bit rate adaptive network to be adjusted is updated, the online phase shown in 15 can be entered into the meta-test phase based on meta-learning, such as Fig.18 The online monitoring network status category shown can obtain the network status group of the current category corresponding to the target video transmission process, so that the video bit rate adaptive network to be adjusted can be fine-tuned based on the network status group of the current category in the future, so as to better optimize the adaptability of the video bit rate adaptive network to be adjusted.
[0162] Specifically, Fig.15 As shown, meta-learning includes a meta-training phase and a meta-testing phase. After completing the meta-training phase of step S101 to step S105, the meta-testing phase based on meta-learning in the online phase shown in 15 can be entered, such as Fig.18The online monitoring network status category shown can obtain the network status group of the current category corresponding to the target video transmission process, and judge whether the meta-test start condition is currently met based on the network status group of the current category. When the meta-test start condition is met, the video bit rate adaptive network is fine-tuned for the network status group based on the current category. Conversely, when the meta-test start condition is not met, the target video bit rate adaptive network obtained from the previous meta-test is used as the target video bit rate adaptive network adapted for the network status group of the current category.
[0163] In step S107, based on the decision bit rate output by the video bit rate adaptive network to be adjusted and the current category network environment, a reward value corresponding to each decision bit rate is obtained to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted until the reward value meets the convergence condition, thereby obtaining the target video bit rate adaptive network corresponding to the current category.
[0164] In this embodiment, after obtaining the network status group of the current category corresponding to the target video transmission process, if it is determined to start the meta test, then enter Fig.18 In the meta-test phase shown, the reward value corresponding to each decision bit rate can be obtained based on the decision bit rate output by the video bit rate adaptive network to be adjusted and the current category network environment. Then, based on the reward value corresponding to each decision bit rate, the network parameters of the video bit rate adaptive network to be adjusted are updated in an inner loop until the reward value meets the convergence condition, so as to obtain the target video bit rate adaptive network corresponding to the current category.
[0165] Specifically, after obtaining the network status group of the current category corresponding to the target video transmission process, such as Fig.18 The meta-test stimulus condition design shown in the figure can determine whether the current meta-test stimulus condition is met based on the network state group of the current category, that is, whether the meta-test start condition is met. Specifically, the mean difference, standard deviation difference and fluctuation difference between the network state group of the current category and the network state group of the previous time during the target video transmission process are calculated. If the mean difference, standard deviation difference and fluctuation difference difference all meet the interval range change condition, it is determined that the category of the network state group of the current category has changed, and the current category network environment can be simulated based on the network state group of the current category to simulate K bandwidth trajectories of different time lengths, and then the current observation quantity (such as Fig.17 The throughput, delay, delay jitter and packet loss rate (as shown) and the current decision bit rate are input to the decision network to be adjusted of the video bit rate adaptive network to be adjusted, and the decision bit rate at the next moment is output through the decision network to be adjusted.
[0166] Further, the decision bitrate at the next moment is interacted with the K current bandwidth trajectories at the next moment in the current category network environment respectively to obtain K observations at the next moment, and then the K observations at the next moment and the decision bitrate at the next moment can be input into the evaluation network to be adjusted of the video bitrate adaptive network to be adjusted, and the evaluation value is output through the evaluation network to be adjusted. At the same time, the K reward values at the next moment are calculated based on the K observations at the next moment and the decision bitrate at the next moment. Specifically, the K reward values can be calculated based on the preset smoothness and each observation in the K observations at the next moment, wherein each observation includes throughput, delay and packet loss rate, and the above reward formula is used to calculate the K reward values. Further, the network parameters of the decision network to be adjusted and the evaluation network to be adjusted can be updated based on the K reward values at the next moment, so as to perform an inner loop update on the network parameters of the video bitrate adaptive network to be adjusted, until the reward value meets the convergence condition, such as until the mean of the reward value no longer increases after the inner loop, that is, convergence is achieved, and the target video bitrate adaptive network corresponding to the current category is obtained.
[0167] On the contrary, if any of the mean difference, standard deviation difference or fluctuation difference difference does not meet the change condition of the interval range, it is determined that the category of the network status group of the current category has not changed, and the target video bit rate adaptive network corresponding to the network status group at the previous time can be used as the target video bit rate adaptive network corresponding to the current category.
[0168] In an embodiment of the present application, a training method for a video bit rate adaptive network is provided. Through the above-mentioned method, the first reward value corresponding to each decision bit rate can be obtained in an offline state, and the initialization network parameters of the first network can be updated in an inner loop, and the second reward value corresponding to each decision bit rate can be obtained. The initialization network parameters of the video bit rate adaptive network are updated in an outer loop to achieve meta-training of the initialized video bit rate adaptive network based on historical experience, so that the initialized video bit rate adaptive network has the ability to quickly learn and quickly generate the optimal video bit rate for the adaptive network state group. Then, the network state group of the current category can be detected in real time in an online state to obtain the reward value, and the network parameters of the video bit rate adaptive network to be adjusted are updated in an inner loop to obtain the target video bit rate adaptive network. The network parameters of the video bit rate adaptive network to be adjusted can be further fine-tuned based on the meta-test to optimize the adaptability of the target video bit rate adaptive network to a diversified high-speed fluctuating network environment.
[0169] Optionally, in the above Figure 2 Based on the corresponding embodiment, in another optional embodiment of the training method of the video bit rate adaptive network provided in the embodiment of the present application, as Figure 3As shown, step S101 samples N first network state groups of different classes based on N meta-learning tasks, and distributes the N first network state groups of different classes to N initialized video bit rate adaptive networks to obtain N first networks. The method further includes: steps S301 to S303; step S101 includes step S304;
[0170] In step S301, a first historical network trace is obtained, wherein the first historical network trace includes a network throughput, a packet loss rate, and a round-trip delay in a first historical time period;
[0171] In this embodiment, in order to better help the video bitrate adaptive network to quickly learn and quickly adapt to new tasks, before sampling N different types of first network state groups based on N meta-learning tasks, that is, Fig.15 As shown, before entering the meta-training stage based on meta-learning, a preprocessing stage of relevant data and tasks can be performed. The definition of the meta-learning task and the network state groups of different categories can be defined. At the same time, the first historical network trajectory can be obtained, and the bandwidth estimation and filtering processing can be performed on the first historical network trajectory based on the time dimension to obtain the mean, standard deviation and fluctuation difference corresponding to each time window. Then, based on the mean, standard deviation and fluctuation difference corresponding to each time window, the mean interval threshold, standard deviation interval threshold and fluctuation difference interval threshold can be designed, so that the mean interval, standard deviation interval and fluctuation difference interval of each network state group can be determined based on the mean interval threshold, standard deviation interval threshold and fluctuation difference interval threshold to obtain the category of each network state group, so that the subsequent design of the mean interval threshold, standard deviation interval threshold and fluctuation difference interval threshold, as well as the definition of the category of each network state group, can better and more accurately simulate the corresponding network environment and understand the distribution of the network state group category, so as to more quickly help the video bit rate adaptive network to quickly learn and quickly adapt to new tasks.
[0172] Specifically, Fig.19 As shown, the entire process of the preprocessing stage is performed offline and may include: Fig.19 The process shown in the figure includes the definition of meta-learning tasks, the definition of network state categories, the collection of network trajectories, the design of bandwidth estimation and filter schemes, and the design of network state group intervals. Since meta-learning is the process of learning different tasks offline and then quickly adapting to new tasks through a small number of sample training processes, Fig.19 In the definition stage of the meta-learning task shown in the figure, this embodiment can first define the meta-learning task Γ iIt is defined as the adaptability of video (such as interactive video) transmission rate adaptation technology to different categories of network status groups, so that the video rate adaptation network can continuously generate new video rate selection strategies that adapt to the current network status through less sample training and rapid update of network parameters, thereby achieving seamless adaptation to dynamic network status changes, which can improve the adaptability of the video rate adaptation network to a certain extent.
[0173] Furthermore, in the meta-learning task Γ i After definition, you can enter Fig.19 The network state category is defined in the stage shown in the figure to further clarify the different meta-learning tasks in the meta-learning, that is, one meta-learning task corresponds to a category of network state group. In this embodiment, the network state group i can be defined as <m i ,d i ,w i ,Δm i ,Δd i ,Δw i >, where m, d, w, Δm, Δd, and Δw represent the mean center point, standard deviation center point, fluctuation difference center point, mean interval, standard deviation interval, and fluctuation difference interval of the network status group, respectively. The fluctuation difference can be calculated using the following formula (1):
[0174]
[0175] Among them, b t represents the network bandwidth value at time t, W b is a preset sliding window. Among them, the design of Δm, Δd, and Δw is crucial, because the larger the interval, the greater the diversity of network states within the network state group. Although it makes the specificity of the video bitrate adaptive network worse, it will make the connection between classes better during real-time operation, thereby reducing the number of meta-test startups. It is understandable that the mean interval threshold, standard deviation interval threshold, and fluctuation difference interval threshold of the network state group need to be evaluated through the network track state change characteristics collected offline.
[0176] Furthermore, if Fig.19 As shown, while defining the network status group and the network status group category, in order to better obtain the mean interval threshold, standard deviation interval threshold and fluctuation difference interval threshold of the appropriate network status group, such as Fig.19 In the collection phase of the network trajectory shown, a network status trajectory for a period of time (i.e., the first historical network trajectory) is collected, wherein the network trajectory includes network throughput, packet loss rate Loss, and round-trip delay RTT, etc. It can be understood that the longer the collection time is, the more accurate and comprehensive the statistics of the current network status are.
[0177] In step S302, bandwidth estimation and filtering are performed on the first historical network trajectory based on the time dimension to obtain the mean, standard deviation, and fluctuation difference corresponding to each time window;
[0178] Specifically, Fig.19 As shown, after collecting the first historical network trajectory, it can be based on Fig.19 The bandwidth estimation and filter scheme shown in the figure is designed. Based on the bandwidth estimation and filter algorithm, the bandwidth of the first historical network trajectory is estimated and filtered to obtain the mean, standard deviation and fluctuation difference corresponding to each time window. Specifically, it can be considered that interactive video is more likely to fail to fill the actual bandwidth than VoD video transmission. Therefore, if the bandwidth is estimated based on the throughput measured at the receiving end (such as the terminal device used by the target object watching the video), it is easy to obtain a smaller bandwidth estimation value. This may be because the adjustment of the video application and the transmission layer bit rate under the WebRTC architecture (the current classic video transmission architecture based on web pages) is tightly coupled and synchronized, and the interactive video is usually shot and transmitted. Therefore, once the target bit rate is lower than the actual bandwidth, the underfill situation will occur. Therefore, this embodiment can collect the first historical network trajectory in the past period of time in advance, that is, the RTT (round trip time), Loss (packet loss rate) and throughput of each data packet of the network to estimate the Th network bandwidth.
[0179] Furthermore, it should be noted that this embodiment follows the principle that when there is packet loss or waiting delay caused by congestion, it can be understood that the bandwidth is full at this time, that is, the throughput can be equal to the network bandwidth; and when the above two conditions are not met, that is, there is no packet loss or waiting delay caused by congestion, it can be understood that the bandwidth is not full at this time, that is, the throughput is less than the bandwidth estimate. Therefore, this embodiment can design the bandwidth estimation algorithm and construct the bandwidth filter as follows: if the packet loss rate Loss within a unit time (which can be taken as 1s) is greater than 5%, or the average round-trip delay RTT is greater than the propagation delay (that is, there is no queuing delay) + 10, the bandwidth estimate for the unit time can be the throughput; and when the above conditions are not met, that is, if the packet loss rate Loss within a unit time (which can be taken as 1s) is less than or equal to 5%, and the average round-trip delay RTT is less than or equal to the propagation delay RTT prop(i.e. no queuing delay) + 10, the bandwidth estimation value can be increased by 1.25 times based on the throughput; and if the above conditions are not met in the previous unit time, the current bandwidth estimation value can be further increased by 1.25 times based on the bandwidth estimation value at the previous moment, and then the bandwidth estimation value at the previous moment can be further increased by 1.25 times and compared with 1.25 times of the current throughput, and the maximum value is taken as the bandwidth estimation value. It can be understood that this process imitates the upward sniffing mechanism of the congestion control mechanism, which is used to prevent the occurrence of long-term bandwidth underfilling, that is, the above bandwidth estimation value b t The estimation method can be shown in the following formula (2):
[0180]
[0181] Among them, up t =Loss t >0.05||RTT>RTT prop When +10 is 0, it means that the packet loss rate Loss within the unit time (which can be 1s) is not greater than 5%, or the average round-trip delay RTT is greater than the propagation delay RTT. prop (ie no queue delay) + 10 means that the sniffing mechanism is not enabled. Similarly, up t =Loss t >0.05||RTT>RTT prop When +10 is 1, the packet loss rate Loss within a unit time (which can be 1s) is greater than 5%, or the average round-trip delay RTT is greater than the propagation delay RTT prop (i.e. no queue delay) + 10 means that the sniffing mechanism is enabled.
[0182] The packet loss rate Loss=5% threshold represents non-congestion packet loss (such as wireless link loss, router port flip, etc.), and the RTT threshold RTT prop The calculation method is RTT prop =min(RTT t' ), t'∈[max(tW RTT ,0),t], indicating that the current target video (such as an interactive video call) is in the time window W RTT The round-trip delay RTT of each group of data packets t' The minimum value of W RTT It can be in the range of tens of seconds or minutes. However, since many target videos are short in duration during transmission (such as video calls), we can collect offline videos with a call length greater than W. RTT of videos and their respective RTT propPerform statistics, remove statistical outliers, and take data within a certain confidence interval (such as a 90% confidence interval) as the initial value of the propagation delay When the target video is being transmitted (such as in a real-time video call), the latest estimated RTT prop Less than the initial value Then update the propagation delay of the current link. That is, the above RTT prop The estimation method is shown in the following formula (3):
[0183] RTT prop =min(RTT t' ), t'∈[max(tW RTT ,0),t] (3);
[0184] Among them, the above RTT prop The estimation method and bandwidth estimation value b t The estimation method is more suitable for real-time measurement, that is, the future network status changes are unknown. However, in a non-real-time state, the future network status changes are known. For example, although the bandwidth is not full at time t, the bandwidth estimation value of the next nearest full bandwidth can be obtained. In this embodiment, the bandwidth estimation value b is further calculated. later Optimize and fine-tune as shown in the following formula (4):
[0185] up t =(Loss t >0.05||RTT>RTT prop )||(1.25*Th t >b later ) (4);
[0186] The above formula (4) shows that if the sniffing bandwidth is larger than the bandwidth in the subsequent full state, it can be understood that the current throughput is likely to be not much different from the actual network bandwidth, and the up t Change to 1 to indicate that the bandwidth is full at the current moment. t =0,up t-1 = 1, further estimate the bandwidth b later Fine-tune as shown in the following formula (5):
[0187]
[0188] Furthermore, the above-mentioned adjusted formula (4) and formula (5) can be recalculated based on the bandwidth estimation according to the up t The mean, standard deviation, and fluctuation of the bandwidth are calculated.
[0189] In step S303, based on the mean, standard deviation and fluctuation difference corresponding to each time window, a mean interval threshold, a standard deviation interval threshold and a fluctuation difference interval threshold are determined;
[0190] In step S304, based on N meta-learning tasks, mean interval thresholds, standard deviation interval thresholds, and fluctuation difference interval thresholds, N different types of first network state groups are sampled, and the N different types of first network state groups are assigned to N initialized video bit rate adaptive networks to obtain N first networks.
[0191] Specifically, Fig.16 As shown, after calculating the mean, standard deviation and fluctuation difference corresponding to each time window based on the above bandwidth estimation and filtering algorithm, you can enter the following Fig.19 The network status group interval design stage shown in the figure may specifically be: first, based on the collected network status trajectory (such as the first network trajectory), using the time length W b The sliding window of the network state group is used to estimate the bandwidth of each unit time in each time window according to the above bandwidth estimation and filtering algorithm, namely formula (2) to formula (5), and then the mean and standard deviation of the bandwidth are calculated; secondly, according to the time length T required for the meta-test of each category of network status group test , statistical interval The changes Δm, Δd, Δw of the mean m, standard deviation d, and volatility w in the sliding window of the sliding window can be removed. Then, the sliding window can be removed when the last moment is in protection fallback mode or the last two moments are up. t =0,up t-1 = 1, and then, by testing different W b and different thresholds Δm', Δd', Δw' of Δm, Δd, Δw, so as to obtain a certain number (such as 90% or more) of points satisfying the real-time Δm'>Δm, Δd'>Δd, Δw'>Δw under the condition of minimizing these thresholds, then the network status group categories smaller than these thresholds can be deleted, that is, the network status group categories are deleted. <m i ,d i ,w i ,Δm i <Δm',Δd i <Δd',Δw i <Δw'> corresponds to the network status group category i, so that the network status group classification only has fixed thresholds and a part of the classes exceeding the threshold on the Δm, Δd, and Δw attributes. This part of the classes is mainly composed of large-scale mutations in the network status or continuous sniffing when the network is not fully occupied.
[0192] It is understandable that in these cases, the interval attribute of the network state group should change with the calculated values of real-time Δm, Δd, and Δw. This is very necessary, especially when dealing with large network bandwidth fluctuations or upward sniffing, and it is necessary to quickly learn classes with larger intervals, while reducing the aggressiveness of the neural network (larger Δ intervals indicate that some original class characteristics are included) while maintaining the exploration of new network states to enhance adaptability. Secondly, the design of Δ interval thresholds such as mean interval threshold, standard deviation interval threshold, and fluctuation difference interval threshold also ensures that the video bitrate adaptive network can seamlessly adapt to network state changes through meta-testing in most time periods, thereby improving the adaptability of the video bitrate adaptive network to a certain extent.
[0193] Furthermore, after determining the mean interval threshold, the standard deviation interval threshold, and the fluctuation difference interval threshold, based on N meta-learning tasks and in combination with the definition of the network state group category i and the determined mean interval threshold, the standard deviation interval threshold, and the fluctuation difference interval threshold, N different types of first network state groups can be sampled and then the N different types of first network state groups can be assigned to N initialized video bitrate adaptive networks to simulate the network environment of each initialized video bitrate adaptive network to obtain N first networks.
[0194] Optionally, in the above Figure 3 Based on the corresponding embodiment, in another optional embodiment of the training method of the video bit rate adaptive network provided in the embodiment of the present application, as Figure 4 As shown, step S304 samples N first network state groups of different classes based on N meta-learning tasks, mean interval thresholds, standard deviation interval thresholds, and fluctuation difference interval thresholds, including:
[0195] In step S401, bandwidth estimation and filtering are performed on the second historical network trajectory based on the time dimension to obtain the unit time bandwidth corresponding to each time window;
[0196] In step S402, based on the unit time bandwidth corresponding to each time window, the network state group distribution probability statistics are performed on the second historical network trajectory to obtain a first probability density function;
[0197] Specifically, Fig.16 The network state group distribution probability statistics phase shown in the figure can be started from the end of the last meta-training, and the real-time utilization time is W b The sliding window W is used to estimate the bandwidth of each unit time in each window. Further, the sliding window W is used to estimate the bandwidth of each unit time in each window. b, and slide the sliding window at unit time intervals (such as 1s) to count the number of items in each window<m,d,w,Δm,Δd,Δw> , a scatter plot of the network state corresponding to the collected network trajectory within a period of time (such as the second historical network trajectory) is obtained, where Δ is the difference between adjacent windows. Then, the density of the scatter points can be calculated based on the obtained scatter plot, and linear interpolation is performed to obtain the probability density function of the network state group category distribution, that is, the first probability density function f<m,d,w,Δm,Δd,Δw> .
[0198] In step S403, N meta-learning tasks are sampled based on the first probability density function;
[0199] In step S404, based on the N meta-learning tasks, the mean interval threshold, the standard deviation interval threshold, and the fluctuation difference interval threshold, K bandwidth trajectories corresponding to each meta-learning task are generated by a network trajectory generator to obtain N first network state groups of different classes, wherein each first network state group includes K bandwidth trajectories, and K is an integer greater than 1.
[0200] Specifically, after obtaining the first probability density function, you can enter the following Fig.16 In the network state group category sampling stage shown in the figure, the meta-learning tasks required for the meta-training process can be sampled based on the probability density function of the network state group category distribution, that is, the first probability density function, and the number of tasks N required to be sampled during each outer loop update during meta-training can be used to randomly sample N different categories of network state groups. For example, the sampled network state group category is i→ <m i ,d i ,w i ,Δm i ,Δd i ,Δw i >,i=1,2,...,N.
[0201] Furthermore, after sampling N meta-learning tasks, we can proceed as follows Fig.16 The generation of network state group intervals shown in the figure may specifically generate the network state group interval of the category i according to the definition of the sampled network state group category i as the mean m∈[m i -Δm i ,m i +Δm i ], standard deviation d∈[d i -Δd i ,d i +Δd i ], fluctuation difference w∈[w i -Δw i ,w i +Δw i ].
[0202] Further, after obtaining the network status group interval, you can enter Fig.16 The network trajectory generator stage shown in the figure can be specifically that after a network state group category i is given, the mean and standard deviation can be randomly sampled in the interval, and L random values (i.e., the network trajectory length of the first historical network trajectory during meta-training) that meet the mean interval and the standard deviation interval and range between [0, MAX] are generated using Beta distribution, and then the sampled values are rearranged so that the trajectory formed by these random values is in each sliding window W R The time period complies with the fluctuation difference interval, mean interval and standard deviation interval. If not, the sliding window W is regenerated. R Random values are generated until they meet the interval standard. The above process is repeated until K bandwidth trajectories that meet the interval standard are generated, forming the category i network state group that needs to perform the learning task during meta-training, that is, the first network state group.
[0203] Optionally, in the above Figure 4 Based on the corresponding embodiment, in another optional embodiment of the training method of the video bit rate adaptive network provided in the embodiment of the present application, as Figure 5 As shown, after step S106 obtains the network status group of the current category corresponding to the target video transmission process, the method further includes:
[0204] In step S501, during the target video transmission process, the network trajectory within the target time period is subjected to distribution probability statistics of the network state group to obtain a current probability density function corresponding to the target time period;
[0205] In step S502, a cross entropy between the current probability density function and the first probability density function is calculated;
[0206] In step S503, if the cross entropy is greater than the cross threshold, the initialization network parameters of the initialized video bit rate adaptive network are updated again based on the current probability density function to obtain a new video bit rate adaptive network to be adjusted.
[0207] In this embodiment, during the target video transmission process, the distribution probability statistics of the network state group can be performed on the detected network trajectory within the target time period according to a preset time period to obtain the current probability density function corresponding to the target time period, and then the cross entropy between the current probability density function and the first probability density function can be calculated. If the cross entropy is greater than the cross threshold, the meta-training can be restarted, that is, the initialization network parameters of the initialized video bit rate adaptive network are re-updated based on the current probability density function to obtain a new video bit rate adaptive network to be adjusted.
[0208] The target time period is the time after the last meta-training ends and the meta-test begins. According to the preset time period, such as one hour, the duration of the preset time period experienced by the target video transmission process is collected, such as collecting the network trajectory of the target video transmission within one hour. The crossover threshold is set according to the actual application requirements and is not specifically limited here.
[0209] Specifically, Fig.15 As shown in FIG, after the last meta-training is completed, the meta-test begins. During the target video transmission process, the network trajectories detected within the target time period can be analyzed according to the preset time period, and the distribution probability statistics of the network state group can be performed on the detected network trajectories, that is, through the preset sliding window W b , continuously counting the network status groups in each time window<m,d,w,Δm,Δd,Δw> , to obtain a scatter plot of the network state corresponding to the network trajectory, and then calculate the probability density function of the network state group category distribution during this period (i.e., the target time period) to obtain the current probability density function.
[0210] Further, the cross entropy between the current probability density function and the first probability density function can be calculated. When the cross entropy between the current probability density function and the first probability density function during the last element training is greater than a certain threshold (i.e., the cross threshold), it can be understood that the probability density function at this time has undergone a significant change, that is, the network state distribution at this time has changed significantly from before, and the network parameters of the currently used video bit rate adaptive network to be adjusted are not suitable for the network state distribution at this time, that is, the adaptation effect of the network parameters of the currently used video bit rate adaptive network to be adjusted is not good. Therefore, enter the following example Fig.15 In the restart meta-training phase shown, the initialization network parameters of the initialized video bit rate adaptive network are updated again based on the current probability density function to obtain a new video bit rate adaptive network to be adjusted.
[0211] Optionally, in the above Figure 4 Based on the corresponding embodiment, in another optional embodiment of the training method of the video bit rate adaptive network provided in the embodiment of the present application, as Figure 6 As shown, step S402 performs network state group distribution probability statistics on the second historical network trajectory based on the unit time bandwidth corresponding to each time window to obtain a first probability density function, including:
[0212] In step S601, based on the unit time bandwidth corresponding to each time window, a network state scatter plot corresponding to the second historical network trajectory is mapped;
[0213] In step S602, density calculation and linear interpolation processing are performed on the network status scatter plot to obtain a first probability density function.
[0214] Specifically, Fig.16 The network state group distribution probability statistics phase shown in the figure can be started from the end of the last meta-training, and the real-time utilization time is W b The sliding window W is used to estimate the bandwidth of each unit time in each window. Further, the sliding window W is used to estimate the bandwidth of each unit time in each window. b , and slide the sliding window at unit time intervals (such as 1s) to count the number of items in each window<m,d,w,Δm,Δd,Δw> , a scatter plot of the network state corresponding to the collected network trajectory within a period of time (such as the second historical network trajectory) is obtained, where Δ is the difference between adjacent windows. Then, the density of the scatter points can be calculated based on the obtained scatter plot, and linear interpolation is performed to obtain the probability density function of the network state group category distribution, that is, the first probability density function f<m,d,w,Δm,Δd,Δw> .
[0215] Optionally, in the above Figure 2 Based on the corresponding embodiment, in another optional embodiment of the training method of the video bit rate adaptive network provided in the embodiment of the present application, as Figure 7 As shown, step S102 obtains a first reward value corresponding to each decision code rate based on the decision code rate output by each first network and the first network environment, so as to perform an inner loop update on the initialized network parameters of the first network, including:
[0216] In step S701, a first network environment is simulated based on a first network state group, wherein each first network environment includes K first bandwidth trajectories of different time lengths;
[0217] In step S702, for each first network, K observations and K decision code rates at the previous moment are input into the first decision network of the first network, and K decision code rates at the current moment are output through the first decision network;
[0218] In step S703, the K decision code rates at the current moment are respectively interacted with the K first bandwidth trajectories at the current moment in the first network environment to obtain K observation quantities at the current moment;
[0219] In step S704, the K observation values and K decision code rates at the current moment are input into the first evaluation network of the first network, and the evaluation value at the next moment is output through the first evaluation network;
[0220] In step S705, K first reward values corresponding to the K decision code rates at the current moment are calculated based on the K observation quantities at the current moment;
[0221] In step S706, the initialized network parameters of the first decision network and the first evaluation network are updated based on the K first reward values corresponding to the K decision code rates at the current moment, so as to perform an inner loop update on the initialized network parameters of the first network.
[0222] Specifically, after sampling N different types of first network state groups based on N meta-learning tasks, we can proceed as follows: Fig.16 The meta-training update initialization parameter stage shown in the figure. The entire meta-learning process is based on the MAML algorithm, and the PPO reinforcement learning algorithm is used to update the network parameters. Other optimization algorithms can also be used to update the network parameters, which are not specifically limited here.
[0223] Among them, meta-learning includes meta-training and meta-testing stages, and the optimal neural network initialization parameters can be obtained through offline meta-training. The specific steps include: based on the probability density function, that is, the first probability density function, sampling N meta-learning tasks; then, for each learning task, using the network trajectory generator to generate K corresponding bandwidth trajectories, and obtaining the first network state group corresponding to each category; then, simulating each first network state group on the simulator to obtain the first network environment corresponding to each first network state group, and obtaining the observation amount and reward value under the current bitrate adaptation strategy based on the first network environment. Specifically, it can be based on the initialized video bitrate adaptation network constructed based on reinforcement learning technology, and N meta-learning tasks are learned, that is, N first networks can be obtained, each first network includes a first decision network and a first evaluation network, and then, for each first network, the K observation amounts (such as packet loss rate, delay, delay jitter and throughput, etc. as shown in 17) and K decision bit rates of the previous moment can be input into the first decision network of the first network, and through The K decision bit rates at the current moment are output through the first decision network. Furthermore, since the bandwidth trajectory is a trajectory of a period of time, the K decision bit rates at the current moment can be interacted with the K first bandwidth trajectories at the current moment in the first network environment respectively to obtain K observations at the current moment. At the same time, the K observations at the current moment and the K decision bit rates can be input into the first evaluation network of the first network. The evaluation value at the next moment is output through the first evaluation network to evaluate the reward value at the next moment and to indicate the effect of the current video bit rate adaptation strategy. At the same time, the above-mentioned reward formula can be used to calculate the K first reward values corresponding to the K decision bit rates at the current moment based on the K observations at the current moment to update the initialization network parameters of the first decision network and the first evaluation network at the current moment. Repeating the above process can realize the inner loop update of the initialization network parameters of each first network. Further, for different network state group categories, the corresponding neural network updates (inner loops) are performed from the initialization network parameters to maximize the reward.
[0224] For example, assuming that there are 3 categories of first network state groups, which are assigned to 3 initialized video bit rate adaptive networks, there are 3 corresponding first networks. For one first network, K (such as 5) observation quantities (such as packet loss rate, delay, delay jitter and throughput, etc. as shown in 17) and K (such as 5) decision bit rates at the previous moment can be input into the first decision network of the first network, and the first decision network outputs K (such as 5) decision bit rates at the current moment, such as 0.1Mbps, 0.2Mbps, 0.4Mbps, 0.6Mbps, 1.1Mbps, etc., and the decision bit rates 0.1Mbps, 0.2Mbps, 0.4Mbps, 0.6Mbps, 1.1Mbps at the current moment interact with the K (such as 5) first bandwidth trajectories at the current moment in the first network environment, respectively, so as to K (e.g., 5) observation quantities at the current moment are obtained. At the same time, the K (e.g., 5) observation quantities at the current moment and K decision code rates of 0.1Mbps, 0.2Mbps, 0.4Mbps, 0.6Mbps, and 1.1Mbps can be respectively input into the first evaluation network of the first network, and the K (e.g., 5) evaluation values at the next moment are output through the first evaluation network. At the same time, the above-mentioned reward formula can be used to calculate the K first reward values corresponding to the K decision code rates of 0.1Mbps, 0.2Mbps, 0.4Mbps, 0.6Mbps, and 1.1Mbps at the current moment based on the K (e.g., 5) observation quantities at the current moment, such as 0.24, 0.28, 0.34, 0.48, and 0.65, to initialize the network parameters of the first decision network and the first evaluation network at the current moment.
[0225] Optionally, in the above Figure 4 Based on the corresponding embodiment, in another optional embodiment of the training method of the video bit rate adaptive network provided in the embodiment of the present application, as Figure 8 As shown, step S103 resamples N different types of second network state groups corresponding to N meta-learning tasks, including:
[0226] In step S801, N new meta-learning tasks are resampled based on the first probability density function;
[0227] In step S802, based on N new meta-learning tasks, the mean interval threshold, the standard deviation interval threshold and the fluctuation difference interval threshold, K new bandwidth trajectories corresponding to each new meta-learning task are generated by a network trajectory generator to obtain N second network state groups of different classes, wherein each second network state group includes K new bandwidth trajectories.
[0228] Specifically, after the initialized network parameters of the first network are updated in an inner loop, the probability density function of the category distribution of the network state group, i.e., the first probability density function, can be re-based on the number of tasks N required to be sampled during each outer loop update during meta-training, and N new meta-learning tasks can be resampled to randomly sample new network state groups of N different categories.
[0229] Furthermore, after sampling N new meta-learning tasks, the network state group interval of this category can be regenerated as the mean m∈[m i -Δm i ,m i +Δm i ], standard deviation d∈[d i -Δd i ,d i +Δd i ], fluctuation difference w∈[w i -Δw i ,w i +Δw i ].
[0230] Furthermore, after regenerating the network status group interval of this category, the mean and standard deviation can be randomly sampled in the interval according to the given network status group category i, and L random values that can meet the mean interval and standard deviation interval and range between [0, MAX] can be regenerated using Beta distribution, and then the re-obtained sampled values are rearranged so that the trajectory formed by these random values is in each sliding window W R The time period complies with the fluctuation difference interval, mean interval and standard deviation interval. If not, the sliding window W is regenerated. R Random values are generated until they meet the interval criteria. The above process is repeated until K new bandwidth trajectories that meet the interval criteria are generated, and the network state group of category i that needs to perform the learning task during meta-training is reorganized, that is, the second network state group.
[0231] Optionally, in the above Figure 2 Based on the corresponding embodiment, in another optional embodiment of the training method of the video bit rate adaptive network provided in the embodiment of the present application, as Fig. 9 As shown, step S104 obtains a second reward value corresponding to each decision bit rate based on the decision bit rate output by each first network after inner loop update and the second network environment, so as to perform an outer loop update on the initialized network parameters of the initialized video bit rate adaptive network, including:
[0232] In step S901, a second network environment is simulated based on a second network state group, wherein each second network environment includes K second bandwidth trajectories of different time lengths;
[0233] In step S902, for each first network that has been updated through the inner loop, K observations and K decision code rates at the previous moment are input into the first decision network of the first network that has been updated through the inner loop, and K new decision code rates are output at the current moment;
[0234] In step S903, the K new decision code rates at the current moment are respectively interacted with the K second bandwidth trajectories at the current moment in the second network environment to obtain K new observation quantities at the current moment;
[0235] In step S904, K second reward values corresponding to the K new decision code rates at the current moment are calculated based on the K new observation quantities at the current moment;
[0236] In step S905, K second reward values corresponding to K new decision code rates at the current moment based on the N meta-learning tasks are added and averaged to obtain the reward average value at the current moment;
[0237] In step S906, the initialized network parameters of the initialized video bit rate adaptive network are updated based on the average reward value at the current moment, so as to perform an outer loop update on the initialized network parameters of the initialized video bit rate adaptive network.
[0238] Specifically, after resampling N second network state groups of different categories corresponding to N meta-learning tasks, N new meta-learning tasks can be resampled based on the probability density function, that is, the first probability density function; then, for each new meta-learning task, K corresponding new bandwidth trajectories are generated using the network trajectory generator to obtain the second network state group corresponding to each category; then, each second network state group is simulated on the simulator to obtain the second network environment corresponding to each second network state group, and based on the second network environment, the observation amount and reward value under the current bit rate adaptation strategy are obtained, which can be specifically for each first network updated after the inner loop, the K observation amounts (packet loss rate, delay, etc. as shown in 17) at the previous moment are obtained. The K decision code rates are input to the first decision network based on the first network updated by the inner loop, that is, the first decision network updated by the K observations at the current moment and the K decision code rates, and the K new decision code rates at the current moment are output through the updated first decision network. Furthermore, since the new bandwidth trajectory is also a trajectory of a period of time, the K new decision code rates at the current moment can be interacted with the K second bandwidth trajectories at the current moment in the second network environment respectively to obtain the K new observations at the current moment. At the same time, the above-mentioned reward formula can be used to calculate the K second reward values corresponding to the K new decision code rates at the current moment based on the K new observations at the current moment.
[0239] Furthermore, the rewards of each task can be averaged, that is, the K second reward values corresponding to the K new decision bit rates at the current moment based on the N meta-learning tasks are added and averaged to obtain the average reward value, and the initialized network parameters of the initialized video bit rate adaptive network are updated based on the average reward value at the current moment. The above process is repeated to perform an outer loop update on the initialized network parameters of the initialized video bit rate adaptive network.
[0240] Optionally, in the above Figure 4 Based on the corresponding embodiment, in another optional embodiment of the training method of the video bit rate adaptive network provided in the embodiment of the present application, as Fig.10 As shown, after step S106 obtains the network status group of the current category corresponding to the target video transmission process, the method further includes:
[0241] In step S1001, the mean difference, standard deviation difference and fluctuation difference between the network status group of the current category and the network status group of the previous time during the target video transmission process are calculated;
[0242] In step S1002, if any one of the mean difference, the standard deviation difference or the fluctuation difference difference satisfies the change condition of the interval range, it is determined that the category of the network state group of the current category has changed, and the decision bit rate output based on the video bit rate adaptive network to be adjusted and the current category network environment are executed to obtain the reward value corresponding to each decision bit rate, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted until the reward value meets the convergence condition, and the target video bit rate adaptive network corresponding to the current category is obtained;
[0243] In step S1003, if the mean difference, the standard deviation difference and the fluctuation difference do not meet the change conditions of the interval range, it is determined that the category of the network status group of the current category has not changed, and the target video bit rate adaptive network corresponding to the network status group at the previous time is used as the target video bit rate adaptive network corresponding to the current category.
[0244] Specifically, after obtaining the network status group of the current category corresponding to the target video transmission process, you can enter the following Fig.18 In the meta-test stimulus condition design stage shown in the figure, the meta-test start condition can be set as follows: compared with the network status group m at the last meta-test last d last 、w last , we can calculate the network state group of the current category and the network state group m at the previous time during the target video transmission process last d last 、w lastThe difference in mean, standard deviation and fluctuation between them, when the newly detected network status group category meets |mm last |>Δm last / 2,|dd last |>Δd last / 2,|ww last |>Δw last / 2, that is, more than half of the interval range of the network status group tested in the previous unit, it can be understood that any one of the mean difference, standard deviation difference or fluctuation difference difference satisfies the change condition of the interval range, then it is determined that the category of the current category of the network status group has changed, and further, it can be entered as follows Fig.18 The meta-test phase shown is similar to the offline meta-training phase. Based on the current network group category, a network trajectory generator can be used to generate K corresponding bandwidth trajectories to obtain the network state group of this category; and the network state group is simulated on the simulator to obtain the observation quantity and reward value under the current bitrate adaptation strategy; for this network state group category, the network parameters of the video bitrate adaptation network to be adjusted are updated in an inner loop until the reward value meets the convergence condition, and the target video bitrate adaptation network corresponding to the current category is obtained.
[0245] On the contrary, if the mean difference, standard deviation difference and fluctuation difference do not meet the change conditions of the interval range, that is, the current network status group category does not meet the |mm last |>Δm last / 2,|dd last |>Δd last / 2 and |ww last |>Δw last / 2, it is determined that the category of the network status group of the current category has not changed, so the target video bit rate adaptive network corresponding to the network status group at the previous time can be used as the target video bit rate adaptive network corresponding to the current category.
[0246] Optionally, in the above Figure 2 Based on the corresponding embodiment, in another optional embodiment of the training method of the video bit rate adaptive network provided in the embodiment of the present application, as Fig.11 As shown, step S107 obtains a reward value corresponding to each decision bit rate based on the decision bit rate output by the video bit rate adaptive network to be adjusted and the current category network environment, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted, including:
[0247] In step S1101, a current category network environment is simulated based on a current category network state group, wherein the current category network environment includes K current bandwidth trajectories of different time lengths;
[0248] In step S1102, the current observation value and the current decision bit rate corresponding to the target video transmission process are obtained;
[0249] In step S1103, the current observation amount and the current decision bit rate are input to the decision network to be adjusted of the video bit rate adaptive network to be adjusted, and the decision bit rate at the next moment is output through the decision network to be adjusted;
[0250] In step S1104, the decision code rate at the next moment is interacted with the K current bandwidth trajectories at the next moment in the current category network environment to obtain K observation quantities at the next moment;
[0251] In step S1105, the K observation quantities at the next moment and the decision bit rate at the next moment are input into the evaluation network to be adjusted of the video bit rate adaptive network to be adjusted, and the evaluation value is outputted through the evaluation network to be adjusted;
[0252] In step S1106, K reward values at the next moment are calculated based on the K observation values at the next moment and the K decision code rates at the next moment;
[0253] In step S1107, the network parameters of the decision network to be adjusted and the evaluation network to be adjusted are updated based on the K reward values at the next moment, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted.
[0254] Specifically, after determining that the category of the network status group of the current category has changed, you can enter the following Fig.18The meta-test phase shown in the figure is similar to the offline meta-training phase. Based on the current network group category, a network trajectory generator can be used to generate K corresponding bandwidth trajectories to obtain the network state group of this category, and the network state group can be simulated on the simulator to obtain the current category network environment to obtain the observation value and reward value under the current bitrate adaptation strategy. Specifically, the current observation value and the current decision bitrate corresponding to the target video transmission process are obtained, and then the current observation value and the current decision bitrate are input into the decision network to be adjusted of the video bitrate adaptation network to be adjusted, and the decision bitrate at the next moment is output through the decision network to be adjusted, and the decision bitrate at the next moment is output. The rate interacts with the K current bandwidth trajectories at the next moment in the current category network environment respectively to obtain K observations at the next moment. Then, the above reward formula can be sampled, and the K reward values at the next moment are calculated based on the K observations at the next moment and the K decision bit rates at the next moment. Furthermore, for the network state group category, the network parameters of the decision network to be adjusted and the evaluation network to be adjusted are updated based on the K reward values at the next moment, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted until the reward value meets the convergence condition, so as to obtain the target video bit rate adaptive network corresponding to the current category.
[0255] Optionally, in the above Figure 2 Based on the corresponding embodiment, in another optional embodiment of the training method of the video bit rate adaptive network provided in the embodiment of the present application, as Fig.12 As shown, after step S106 obtains the network status group of the current category corresponding to the target video transmission process, the method further includes:
[0256] In step S1201, the current packet loss rate and the current round-trip delay corresponding to the network status group of the current category are compared with the packet loss threshold and the delay threshold respectively;
[0257] In step S1202, when the current packet loss rate is greater than the packet loss threshold or the current round-trip delay is greater than the delay threshold, the protection fallback mechanism is enabled, and the decision network of the video bit rate adaptive network to be adjusted is temporarily stopped for bit rate selection.
[0258] In this embodiment, after obtaining the network status group of the current category corresponding to the target video transmission process, the current packet loss rate and the current round-trip delay corresponding to the network status group of the current category can be compared with the packet loss threshold and the delay threshold, respectively. If the current packet loss rate is greater than the packet loss threshold or the current round-trip delay is greater than the delay threshold, it can be understood that the adaptation effect of the current decision network of the video bit rate adaptive network to be adjusted is not good. Then, the protection fallback mechanism can be enabled to suspend the use of the decision network of the video bit rate adaptive network to be adjusted for bit rate selection, so as to better protect the video bit rate adaptive network to be adjusted.
[0259] The packet loss threshold and the delay threshold are set according to actual application requirements and are not specifically limited here.
[0260] Specifically, after obtaining the network status group of the current category corresponding to the target video transmission process, the current packet loss rate and the current round-trip delay corresponding to the network status group of the current category can be compared with the packet loss threshold and the delay threshold, respectively. When the packet loss rate Loss is greater than a certain threshold (such as the packet loss threshold) or the round-trip delay RTT is greater than a certain threshold (such as the delay threshold), it means that the current video bit rate adaptive algorithm of the video bit rate adaptive network to be adjusted based on meta-reinforcement learning is poor, and the protection fallback mechanism (such as Fig.18 The protection fallback mechanism shown in the figure) is used, and other rate selection algorithms, such as the GCC algorithm, can be used again. Other rate selection algorithms, which are not specifically limited here, are used to re-select the rate, and the selected rate is used to replace the decision network of the video rate adaptive network to be adjusted to select the rate to obtain the decision rate.
[0261] The following is an introduction to the application method of the video bit rate adaptive network in this application. Fig.13 In one embodiment of the present application, a method for applying a video bit rate adaptive network includes:
[0262] In step S1301, the current network state group, the current observation amount and the target decision bit rate at the current moment corresponding to the target video transmission process are obtained;
[0263] In step S1302, if the category of the current network state group has not changed, the current observation amount and the target decision bit rate at the current moment are input into the target decision network in the target video bit rate adaptive network, and the target decision bit rate at the next moment is output through the target decision network;
[0264] In step S1303, the sending rate and the video editing bit rate are adjusted based on the target decision bit rate at the next moment;
[0265] In step S1304, the target video is transmitted based on the adjusted sending rate and the video editing bit rate.
[0266] In this embodiment, after the target video bit rate adaptation network is obtained in the meta-test phase based on meta-learning, the decision network in the target video bit rate adaptation network can be directly used to make a decision on the video bit rate of the target video being transmitted, that is, the current network state group, the current observation amount and the target decision bit rate at the current moment corresponding to the target video transmission process can be obtained. If the category of the current network state group has not changed, it can be understood that the decision network in the current target video bit rate adaptation network has a good adaptation effect to the current network state, and the current observation amount and the target decision bit rate at the current moment can be directly input into the target decision network in the target video bit rate adaptation network, and the target decision bit rate at the next moment adapted to the current network state group can be obtained through the target decision network, so that the sending rate and the video editing bit rate can be adjusted based on the target decision bit rate at the next moment, and then, the target video can be transmitted better and faster based on the adjusted sending rate and video editing bit rate.
[0267] Specifically, Fig. 20 As shown, assuming that the WebRTC platform is used, the GCC bit rate decision result is replaced by the target bit rate result obtained by the decision network in the target video bit rate adaptive network. Fig. 20 The encoder in the server (shown as an example) generates a video of corresponding quality according to the target decision bit rate, and the pacer mechanism adjusts the corresponding sending rate according to the target decision bit rate. The specific process is: at the sending end (such as Fig. 20 The server shown in the figure detects the network status in real time (the current network status group fed back by the receiving end, the current observation value and the target decision bit rate at the current moment), and uploads it to the server, inputs the current observation value and the target decision bit rate at the current moment into the target decision network in the target video bit rate adaptive network, obtains the corresponding target bit rate result, that is, the target decision bit rate at the next moment, and returns it to the sending end, so that the sending end can adjust to the corresponding video editing bit rate and sending rate through the encoder and pacer based on the target decision bit rate at the next moment, and then the sending end transmits the target video to the receiving end (such as Peer-A and Peer-B) based on the adjusted sending rate and video editing bit rate.
[0268] It can be understood that this embodiment can be applied to the intelligent video call process shown in FIG. 21(a), the multimedia live broadcast process shown in FIG. 21(b) Figure 1 ), VR live broadcast, cloud gaming applications as shown in Figure 21(c), remote surgery, and smart car networking.
[0269] Optionally, in the above Fig.13 Based on the corresponding embodiment, in another optional embodiment of the application method of the video bit rate adaptive network provided in the embodiment of the present application, as Fig.14 As shown, after step S1301 obtains the current network state group, the current observation amount and the target decision bit rate at the current moment corresponding to the target video transmission process, the method further includes:
[0270] In step S1401, if the category of the current network status group changes, the current network environment of the video bit rate adaptive network to be adjusted is simulated based on the network status group of the current category;
[0271] In step S1402, based on the decision bit rate output by the video bit rate adaptive network to be adjusted and the current network environment, a reward value corresponding to each decision bit rate is obtained to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted until the reward value meets the convergence condition, thereby obtaining a new target video bit rate adaptive network corresponding to the current network state group;
[0272] In step S1403, the current observation amount and the target decision bit rate at the current moment are input into the target decision network in the new target video bit rate adaptation network, and the target decision bit rate at the next moment is output through the target decision network;
[0273] In step S1404, the sending rate and the video editing bit rate are adjusted based on the target decision bit rate at the next moment;
[0274] In step S1405, the target video is transmitted based on the adjusted sending rate and the video editing bit rate.
[0275] Specifically, after determining that the category of the network status group of the current category has changed, the following can be re-entered: Fig.18The meta-test phase shown in the figure is similar to the offline meta-training phase. The current network environment of the video bitrate adaptive network to be adjusted can be simulated by the network state group of the current category. That is, based on the network state group of the current category, a network trajectory generator is used to generate K corresponding bandwidth trajectories to obtain the network state group of the category, and the network state group is simulated on the simulator to obtain the current network environment to obtain the observation amount and reward value under the current bitrate adaptive strategy. Specifically, the current observation amount corresponding to the target video transmission process and the decision bitrate at the current moment are obtained, and then the current observation amount and the decision bitrate at the current moment are input into the video bitrate adaptive network to be adjusted. The decision network to be adjusted of the network is outputted through the decision network to be adjusted, and the decision bit rate at the next moment is interacted with the K current bandwidth trajectories at the next moment in the current category network environment respectively to obtain K observations at the next moment. Furthermore, for the network state group category, the network parameters of the decision network to be adjusted and the evaluation network to be adjusted are updated based on the K reward values corresponding to the K decision bit rates at the next moment, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted until the reward value meets the convergence condition, so as to obtain a new target video bit rate adaptive network corresponding to the current network state group.
[0276] Furthermore, the current observation quantity and the target decision bit rate at the current moment are input into the decision network of the new target video bit rate adaptive network corresponding to the current network state group, and the corresponding target bit rate result is obtained, and it is returned to the sending end, so that the sending end can adjust to the corresponding video editing bit rate and sending rate through the encoder and the pacer based on the corresponding target bit rate result, and then, the sending end transmits the target video to the receiving end (such as Peer-A and Peer-B) based on the adjusted sending rate and video editing bit rate.
[0277] The following is a detailed description of the training device of the video bit rate adaptive network in this application. Fig. 22 , Fig. 22 This is a schematic diagram of an embodiment of a video bit rate adaptive network training device in an embodiment of the present application. The video bit rate adaptive network training device 20 includes:
[0278] The processing unit 201 is used to sample N first network state groups of different classes based on N meta-learning tasks, and assign the N first network state groups of different classes to N initialized video bit rate adaptive networks to obtain N first networks, wherein each first network state group is used to simulate a first network environment of each first network, and the initialized video bit rate adaptive network is set with an initialized network parameter, and N is an integer greater than 1;
[0279] An acquisition unit 202 is used to acquire a first reward value corresponding to each decision code rate based on the decision code rate output by each first network and the first network environment, so as to perform an inner loop update on the initialized network parameters of the first network;
[0280] The processing unit 201 is further used to resample N second network state groups of different classes corresponding to the N meta-learning tasks, and distribute the N second network state groups of different classes to the N first networks updated through the inner loop, wherein each second network state group is used to simulate the second network environment of each first network updated through the inner loop;
[0281] The acquisition unit 202 is further used to acquire a second reward value corresponding to each decision bit rate based on the decision bit rate output by each first network after inner loop update and the second network environment, so as to perform outer loop update on the initialized network parameters of the initialized video bit rate adaptive network;
[0282] The processing unit 201 is further used to repeatedly perform the operations of sampling the first network state group, obtaining the first reward value, inner loop updating, sampling the second network state group, obtaining the second reward value, and outer loop updating until the reward values of the N meta-learning tasks after the inner loop updating meet the convergence condition, and initialize the video bit rate adaptive network to update and obtain the video bit rate adaptive network to be adjusted;
[0283] The acquisition unit 202 is further used to acquire a network status group of the current category corresponding to the target video transmission process, wherein the network status group of the current category is used to simulate the current category network environment of the video bit rate adaptive network to be adjusted;
[0284] The acquisition unit 202 is further used to obtain a reward value corresponding to each decision bit rate based on the decision bit rate output by the video bit rate adaptive network to be adjusted and the current category network environment, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted until the reward value meets the convergence condition, thereby obtaining a target video bit rate adaptive network corresponding to the current category.
[0285] Optionally, in the above Fig. 22 Based on the corresponding embodiment, in another embodiment of the training device for the video bit rate adaptive network provided by the embodiment of the present application,
[0286] The acquisition unit 202 is further configured to acquire a first historical network trajectory, wherein the first historical network trajectory includes a throughput, a packet loss rate, and a round-trip delay of the network in a first historical time period;
[0287] The processing unit 201 is further used to perform bandwidth estimation and filtering processing on the first historical network trajectory based on the time dimension to obtain a mean, a standard deviation, and a fluctuation difference corresponding to each time window;
[0288] A determination unit 203 is used to determine a mean interval threshold, a standard deviation interval threshold, and a fluctuation difference interval threshold based on the mean, standard deviation, and fluctuation difference corresponding to each time window;
[0289] The processing unit 201 can be specifically used to: based on N meta-learning tasks, mean interval thresholds, standard deviation interval thresholds, and fluctuation difference interval thresholds, sample N different types of first network state groups, and assign the N different types of first network state groups to N initialized video bit rate adaptive networks to obtain N first networks.
[0290] Optionally, in the above Fig. 22 On the basis of the corresponding embodiment, in another embodiment of the training device for the video bit rate adaptive network provided by the embodiment of the present application, the processing unit 201 can be specifically used for:
[0291] Based on the time dimension, the bandwidth of the second historical network trajectory is estimated and filtered to obtain the unit time bandwidth corresponding to each time window;
[0292] Based on the unit time bandwidth corresponding to each time window, the network state group distribution probability statistics are performed on the second historical network trajectory to obtain a first probability density function;
[0293] Sample N meta-learning tasks based on the first probability density function;
[0294] Based on N meta-learning tasks, a mean interval threshold, a standard deviation interval threshold, and a fluctuation difference interval threshold, a network trajectory generator is used to generate K bandwidth trajectories corresponding to each meta-learning task to obtain N first network state groups of different classes, wherein each first network state group includes K bandwidth trajectories, and K is an integer greater than 1.
[0295] Optionally, in the above Fig. 22 Based on the corresponding embodiment, in another embodiment of the training device for the video bit rate adaptive network provided by the embodiment of the present application,
[0296] The processing unit 201 is further used to perform distribution probability statistics of the network state group on the network trajectory within the target time period during the target video transmission process to obtain a current probability density function corresponding to the target time period;
[0297] The processing unit 201 is further configured to calculate a cross entropy between the current probability density function and the first probability density function;
[0298] The acquisition unit 202 is further configured to re-update the initialization network parameters of the initialized video bit rate adaptive network based on the current probability density function if the cross entropy is greater than the cross threshold, so as to obtain a new video bit rate adaptive network to be adjusted.
[0299] Optionally, in the above Fig. 22 On the basis of the corresponding embodiment, in another embodiment of the training device for the video bit rate adaptive network provided by the embodiment of the present application, the processing unit 201 can be specifically used for:
[0300] Based on the unit time bandwidth corresponding to each time window and the definition of the network state group, a network state scatter plot corresponding to the second historical network trajectory is mapped;
[0301] The density of the scattered points in the network state scatter diagram is calculated and linear interpolated to obtain a first probability density function.
[0302] Optionally, in the above Fig. 22 On the basis of the corresponding embodiment, in another embodiment of the training device for the video bit rate adaptive network provided by the embodiment of the present application, the acquisition unit 202 can be specifically used for:
[0303] Simulating a first network environment based on the first network state group, wherein each first network environment includes K first bandwidth trajectories of different time lengths;
[0304] For each first network, K observations and K decision code rates at the previous moment are input into the first decision network of the first network, and K decision code rates at the current moment are output through the first decision network;
[0305] The K decision code rates at the current moment interact with the K first bandwidth trajectories at the current moment in the first network environment respectively to obtain K observation quantities at the current moment;
[0306] Inputting K observation quantities and K decision code rates at the current moment into the first evaluation network of the first network, and outputting the evaluation value at the next moment through the first evaluation network;
[0307] Based on the K observations at the current moment, K first reward values corresponding to the K decision code rates at the current moment are calculated;
[0308] Based on the K first reward values corresponding to the K decision code rates at the current moment, the initialized network parameters of the first decision network and the first evaluation network are updated to perform an inner loop update on the initialized network parameters of the first network.
[0309] Optionally, in the above Fig. 22 On the basis of the corresponding embodiment, in another embodiment of the training device for the video bit rate adaptive network provided by the embodiment of the present application, the processing unit 201 can be specifically used for:
[0310] Resample N new meta-learning tasks based on the first probability density function;
[0311] Based on N new meta-learning tasks, a mean interval threshold, a standard deviation interval threshold, and a fluctuation difference interval threshold, a network trajectory generator is used to generate K new bandwidth trajectories corresponding to each new meta-learning task to obtain N second network state groups of different classes, where each second network state group includes K new bandwidth trajectories.
[0312] Optionally, in the above Fig. 22 On the basis of the corresponding embodiment, in another embodiment of the training device for the video bit rate adaptive network provided by the embodiment of the present application, the acquisition unit 202 can be specifically used for:
[0313] Simulating a second network environment based on the second network state group, wherein each second network environment includes K second bandwidth trajectories of different time lengths;
[0314] For each first network updated through the inner loop, K observations and K decision code rates at the previous moment are input into the first decision network of the first network updated through the inner loop, and K new decision code rates at the current moment are output;
[0315] The K new decision code rates at the current moment interact with the K second bandwidth trajectories at the current moment in the second network environment respectively to obtain K new observation quantities at the current moment;
[0316] Based on the K new observations at the current moment, K second reward values corresponding to the K new decision code rates at the current moment are calculated;
[0317] The K second reward values corresponding to the K new decision code rates at the current moment based on the N meta-learning tasks are added and averaged to obtain the average reward value at the current moment;
[0318] The initialized network parameters of the initialized video bitrate adaptive network are updated based on the average reward value at the current moment, so as to perform an outer loop update on the initialized network parameters of the initialized video bitrate adaptive network.
[0319] Optionally, in the above Fig. 22 Based on the corresponding embodiment, in another embodiment of the training device for the video bit rate adaptive network provided by the embodiment of the present application,
[0320] The processing unit 201 is further used to calculate the mean difference, standard deviation difference and fluctuation difference between the network status group of the current category and the network status group of the previous time during the target video transmission process;
[0321] The processing unit 201 is further used to determine that the category of the network state group of the current category has changed if any one of the mean difference, the standard deviation difference or the fluctuation difference difference satisfies the change condition of the interval range, and execute the decision bit rate output based on the video bit rate adaptive network to be adjusted and the current category network environment, obtain the reward value corresponding to each decision bit rate, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted, until the reward value satisfies the convergence condition, and obtain the target video bit rate adaptive network corresponding to the current category;
[0322] The processing unit 201 is also used to determine that the category of the network status group of the current category has not changed if the mean difference, the standard deviation difference and the fluctuation difference do not meet the change conditions of the interval range, and then use the target video bit rate adaptive network corresponding to the network status group at the previous time as the target video bit rate adaptive network corresponding to the current category.
[0323] Optionally, in the above Fig. 22 On the basis of the corresponding embodiment, in another embodiment of the training device for the video bit rate adaptive network provided by the embodiment of the present application, the acquisition unit 202 can be specifically used for:
[0324] Simulating a current category network environment based on a current category network state group, wherein the current category network environment includes K current bandwidth trajectories of different time lengths;
[0325] Obtain the current observation value and current decision bit rate corresponding to the target video transmission process;
[0326] Inputting the current observation amount and the current decision bit rate into the decision network to be adjusted of the video bit rate adaptive network to be adjusted, and outputting the decision bit rate at the next moment through the decision network to be adjusted;
[0327] The decision code rate at the next moment is interacted with the K current bandwidth trajectories at the next moment in the current category network environment to obtain K observations at the next moment;
[0328] Inputting the K observation quantities at the next moment and the decision bit rate at the next moment into the evaluation network to be adjusted of the video bit rate adaptive network to be adjusted, and outputting the evaluation value through the evaluation network to be adjusted;
[0329] Based on the K observations at the next moment and the decision code rate at the next moment, the K reward values at the next moment are calculated;
[0330] Based on the K reward values at the next moment, the network parameters of the decision network to be adjusted and the evaluation network to be adjusted are updated, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted.
[0331] Optionally, in the above Fig. 22 Based on the corresponding embodiment, in another embodiment of the training device for the video bit rate adaptive network provided by the embodiment of the present application,
[0332] The processing unit 201 is further configured to compare the current packet loss rate and the current round trip delay corresponding to the network status group of the current category with the packet loss threshold and the delay threshold respectively;
[0333] The processing unit 201 is further configured to enable a protection fallback mechanism and suspend the use of the decision network of the video bit rate adaptive network to be adjusted for bit rate selection when the current packet loss rate is greater than a packet loss threshold or the current round trip delay is greater than a delay threshold.
[0334] The following is a detailed description of the application device of the video bit rate adaptive network in this application. Fig.23 , Fig.23 This is a schematic diagram of an embodiment of an application device of a video bit rate adaptive network in an embodiment of the present application. The application device 30 of the video bit rate adaptive network includes:
[0335] An acquisition unit 301 is used to acquire a current network state group, a current observation amount and a target decision bit rate at a current moment corresponding to a target video transmission process;
[0336] The processing unit 302 is used for inputting the current observation amount and the target decision bit rate at the current moment into the target decision network in the target video bit rate adaptation network if the category of the current network state group has not changed, and outputting the target decision bit rate at the next moment through the target decision network;
[0337] A determination unit 303, configured to adjust a transmission rate and a video editing bit rate based on a target decision bit rate at a next moment;
[0338] The processing unit 302 is further configured to transmit the target video based on the adjusted sending rate and video editing bit rate.
[0339] Optionally, in the above Fig.23 On the basis of the corresponding embodiment, in another embodiment of the application device of the video bit rate adaptive network provided by the embodiment of the present application,
[0340] The processing unit 302 is further configured to simulate the current network environment of the video bit rate adaptive network to be adjusted based on the network state group of the current category if the category of the current network state group changes;
[0341] The acquisition unit 301 is further used to acquire a reward value corresponding to each decision bit rate based on the decision bit rate output by the video bit rate adaptive network to be adjusted and the current network environment, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted until the reward value meets the convergence condition, thereby obtaining a new target video bit rate adaptive network corresponding to the current network state group;
[0342] The processing unit 302 is further configured to input the current observation amount and the target decision bit rate at the current moment into the target decision network in the new target video bit rate adaptation network, and output the target decision bit rate at the next moment through the target decision network;
[0343] The determination unit 303 is further configured to adjust the transmission rate and the video editing bit rate based on the target decision bit rate at the next moment;
[0344] The processing unit 302 is further configured to transmit the target video based on the adjusted sending rate and video editing bit rate.
[0345] On the other hand, the present application provides another schematic diagram of a computer device, such as Fig.24 As shown, Fig.24 3 is a schematic diagram of a computer device structure provided in an embodiment of the present application. The computer device 300 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 (for example, one or more mass storage devices) storing application programs 331 or data 332. Among them, the memory 320 and the storage medium 330 may be temporary storage or permanent storage. The program stored in the storage medium 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the computer device 300. Furthermore, the central processing unit 310 may be configured to communicate with the storage medium 330 to execute a series of instruction operations in the storage medium 330 on the computer device 300.
[0346] The computer device 300 may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating systems 333, such as Windows Server 2000, Windows Server 2003, Windows Server 2003E, Windows Server 2003R ... TM , Mac OS X TM , Unix TM ,Linux TM , FreeBSD TM etc.
[0347] The computer device 300 is also used to execute the following steps: Figures 2 to 12 The steps in the corresponding embodiments, and the execution of Figure 13 to Figure 14 The steps in the corresponding embodiments.
[0348] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following Figures 2 to 12 The steps in the method described in the illustrated embodiment, and the implementation of Figure 13 to Figure 14 The steps in the corresponding embodiments.
[0349] Another aspect of the present application provides a computer program product comprising a computer program, which, when executed by a processor, implements the following Figures 2 to 12 The steps in the method described in the illustrated embodiment, and the implementation of Figure 13 to Figure 14 The steps in the corresponding embodiments.
[0350] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0351] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0352] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0353] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0354] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
Claims
1. A video bit rate adaptive network training method, characterized in that: include: Based on N meta-learning tasks, N first network state groups of different classes are sampled, and the first network state groups of N different classes are assigned to N initialized video bit rate adaptive networks to obtain N first networks, wherein each of the first network state groups is used to simulate the first network environment of each of the first networks, and the initialized video bit rate adaptive network is provided with an initialized network parameter, and N is an integer greater than 1; Based on the decision code rate output by each of the first networks and the first network environment, a first reward value corresponding to each decision code rate is obtained to perform an inner loop update on the initialized network parameters of the first network, including: simulating the first network environment based on the first network state group, wherein each of the first network environments includes K first bandwidth trajectories of different time lengths, and K is an integer greater than 1; for each of the first networks, K observations and K decision code rates at the previous moment are input into the first decision network of the first network, and K decision code rates at the current moment are output through the first decision network; the K decision code rates at the current moment are input into the first decision network of the first network. Respectively interact with the K first bandwidth trajectories at the current moment in the first network environment to obtain K observation quantities at the current moment; input the K observation quantities and K decision code rates at the current moment into the first evaluation network of the first network, and output the evaluation value at the next moment through the first evaluation network; calculate and obtain K first reward values corresponding to the K decision code rates at the current moment based on the K observation quantities at the current moment; update the initialization network parameters of the first decision network and the first evaluation network based on the K first reward values corresponding to the K decision code rates at the current moment, so as to perform an inner loop update on the initialization network parameters of the first network; Resampling N second network state groups of different classes corresponding to the N meta-learning tasks, and assigning the N second network state groups of different classes to the N first networks updated by the inner loop, wherein each second network state group is used to simulate the second network environment of each first network updated by the inner loop; Based on the decision bitrate output by each first network after the inner loop update and the second network environment, a second reward value corresponding to each decision bitrate is obtained to perform an outer loop update on the initialized network parameters of the initialized video bitrate adaptive network, including: simulating the second network environment based on the second network state group, wherein each second network environment includes K second bandwidth trajectories of different time lengths; for each first network after the inner loop update, the K observations and K decision bitrates at the previous moment are input into the first decision network of the first network after the inner loop update, and K new decision bitrates at the current moment are output; the current moment is updated. The K new decision bit rates at the previous moment interact with the K second bandwidth trajectories at the current moment in the second network environment respectively to obtain K new observation quantities at the current moment; based on the K new observation quantities at the current moment, K second reward values corresponding to the K new decision bit rates at the current moment are calculated; based on the K second reward values corresponding to the K new decision bit rates at the current moment of the N meta-learning tasks, the K second reward values are added and averaged to obtain the reward average value at the current moment; based on the reward average value at the current moment, the initialized network parameters of the initialized video bit rate adaptive network are updated to perform an outer loop update on the initialized network parameters of the initialized video bit rate adaptive network; Repeating the operations of sampling the first network state group, obtaining the first reward value, the inner loop update, sampling the second network state group, obtaining the second reward value, and the outer loop update until the reward values of the N meta-learning tasks after the inner loop update meet the convergence condition, and the initialized video bitrate adaptive network is updated to obtain the video bitrate adaptive network to be adjusted; Acquire a current category network state group corresponding to the target video transmission process, wherein the current category network state group is used to simulate the current category network environment of the video bit rate adaptive network to be adjusted; Based on the decision bit rate output by the video bit rate adaptive network to be adjusted and the current category network environment, a reward value corresponding to each decision bit rate is obtained to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted until the reward value meets the convergence condition, thereby obtaining a target video bit rate adaptive network corresponding to the current category.
2. The method according to claim 1, characterized in that Before sampling N first network state groups of different classes based on N meta-learning tasks and allocating the N first network state groups of different classes to N initialized video bitrate adaptive networks to obtain N first networks, the method further includes: Acquire a first historical network trace, wherein the first historical network trace includes a network throughput, a packet loss rate, and a round-trip delay in a first historical time period; Based on the time dimension, bandwidth estimation and filtering are performed on the first historical network trajectory to obtain a mean, standard deviation, and fluctuation difference corresponding to each time window; Based on the mean, standard deviation and fluctuation difference corresponding to each time window, determine the mean interval threshold, standard deviation interval threshold and fluctuation difference interval threshold; Based on the N meta-learning tasks, N first network state groups of different classes are sampled, and the N first network state groups of different classes are distributed to N initialized video bit rate adaptive networks to obtain N first networks: Based on the N meta-learning tasks, the mean interval threshold, the standard deviation interval threshold and the fluctuation difference interval threshold, N different classes of the first network state groups are sampled, and the N different classes of the first network state groups are assigned to the N initialized video bit rate adaptive networks to obtain N first networks.
3. The method according to claim 2, characterized in that The sampling of N different types of the first network state groups based on the N meta-learning tasks, the mean interval threshold, the standard deviation interval threshold, and the fluctuation difference interval threshold includes: Based on the time dimension, the bandwidth estimation and filtering processing are performed on the second historical network trajectory to obtain the unit time bandwidth corresponding to each time window; Based on the unit time bandwidth corresponding to each of the time windows, performing network state group distribution probability statistics on the second historical network trajectory to obtain a first probability density function; Sampling N of the meta-learning tasks based on the first probability density function; Based on the N meta-learning tasks, the mean interval threshold, the standard deviation interval threshold, and the fluctuation difference interval threshold, K bandwidth trajectories corresponding to each of the meta-learning tasks are generated by a network trajectory generator to obtain N different classes of the first network state groups, wherein each of the first network state groups includes K bandwidth trajectories, and K is an integer greater than 1.
4. The method according to claim 3, characterized in that After obtaining the network status group of the current category corresponding to the target video transmission process, the method further includes: During the target video transmission process, performing distribution probability statistics of the network state group on the network trajectory within the target time period to obtain a current probability density function corresponding to the target time period; Calculating the cross entropy between the current probability density function and the first probability density function; If the cross entropy is greater than the cross threshold, the initialized network parameters of the initialized video bit rate adaptive network are updated again based on the current probability density function to obtain a new video bit rate adaptive network to be adjusted.
5. The method according to claim 3, characterized in that: The performing network state group distribution probability statistics on the second historical network trajectory based on the unit time bandwidth corresponding to each of the time windows to obtain the first probability density function includes: Based on the unit time bandwidth corresponding to each of the time windows, a network state scatter plot corresponding to the second historical network trajectory is obtained by mapping; Density calculation and linear interpolation processing are performed on the network status scatter plot to obtain the first probability density function.
6. The method according to claim 3, characterized in that The resampling of N different types of second network state groups corresponding to the N meta-learning tasks includes: Resampling N new meta-learning tasks based on the first probability density function; Based on the N new meta-learning tasks, the mean interval threshold, the standard deviation interval threshold and the fluctuation difference interval threshold, K new bandwidth trajectories corresponding to each of the new meta-learning tasks are generated by a network trajectory generator to obtain N second network state groups of different classes, wherein each of the second network state groups includes K new bandwidth trajectories.
7. The method according to claim 3, characterized in that After obtaining the network status group of the current category corresponding to the target video transmission process, the method further includes: Calculate the mean difference, standard deviation difference and fluctuation difference between the network status group of the current category and the network status group at the previous time during the target video transmission process; If any one of the mean difference, the standard deviation difference or the fluctuation difference difference satisfies the change condition of the interval range, it is determined that the category of the network state group of the current category has changed, and the decision bit rate output based on the video bit rate adaptive network to be adjusted and the current category network environment are executed to obtain the reward value corresponding to each decision bit rate, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted until the reward value satisfies the convergence condition, and the target video bit rate adaptive network corresponding to the current category is obtained; If the mean difference, the standard deviation difference and the fluctuation difference do not meet the change conditions of the interval range, it is determined that the category of the network status group of the current category has not changed, and the target video bit rate adaptive network corresponding to the network status group at the previous time is used as the target video bit rate adaptive network corresponding to the current category.
8. The method according to claim 1, characterized in that The step of obtaining a reward value corresponding to each decision bit rate based on the decision bit rate output by the to-be-adjusted video bit rate adaptive network and the current category network environment, so as to perform an inner loop update on the network parameters of the to-be-adjusted video bit rate adaptive network, includes: Simulating the current category network environment based on the current category network state group, wherein the current category network environment includes K current bandwidth trajectories of different time lengths; Obtain the current observation value and current decision bit rate corresponding to the target video transmission process; Inputting the current observation amount and the current decision bit rate into the decision network to be adjusted of the video bit rate adaptive network to be adjusted, and outputting the decision bit rate at the next moment through the decision network to be adjusted; The decision code rate at the next moment interacts with the K current bandwidth trajectories at the next moment in the current category network environment to obtain K observation quantities at the next moment; Inputting the K observation quantities at the next moment and the decision bit rate at the next moment into the evaluation network to be adjusted of the video bit rate adaptive network to be adjusted, and outputting an evaluation value through the evaluation network to be adjusted; Calculating K reward values at the next moment based on the K observation quantities at the next moment and the decision code rate at the next moment; The network parameters of the decision network to be adjusted and the evaluation network to be adjusted are updated based on the K reward values at the next moment, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted.
9. The method according to claim 1, characterized in that: After obtaining the network status group of the current category corresponding to the target video transmission process, the method further includes: Compare the current packet loss rate and the current round-trip delay corresponding to the network status group of the current category with the packet loss threshold and the delay threshold respectively; When the current packet loss rate is greater than the packet loss threshold or the current round-trip delay is greater than the delay threshold, a protection fallback mechanism is enabled, and the use of the decision network of the video bit rate adaptive network to be adjusted for bit rate selection is suspended.
10. An application method of a video bit rate adaptive network, characterized in that: include: Obtain the current network status group, current observation quantity and target decision bit rate at the current moment corresponding to the target video transmission process; If the category of the current network state group has not changed, the current observation amount and the target decision bit rate at the current moment are input into the target decision network in the target video bit rate adaptation network according to any one of claims 1 to 9, and the target decision bit rate at the next moment is output through the target decision network; Adjusting the sending rate and the video editing bit rate based on the target decision bit rate at the next moment; The target video is transmitted based on the adjusted sending rate and video editing bit rate.
11. The method according to claim 10, characterized in that After obtaining the current network status group, the current observation amount and the target decision bit rate at the current moment corresponding to the target video transmission process, the method further includes: If the category of the current network state group changes, simulating the current network environment of the video bit rate adaptive network to be adjusted based on the network state group of the current category; Based on the decision bit rate output by the video bit rate adaptive network to be adjusted and the current network environment, a reward value corresponding to each decision bit rate is obtained to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted until the reward value meets the convergence condition, thereby obtaining a new target video bit rate adaptive network corresponding to the current network state group; Inputting the current observation amount and the target decision bit rate at the current moment into the target decision network in the new target video bit rate adaptive network, and outputting the target decision bit rate at the next moment through the target decision network; Adjusting the sending rate and the video editing bit rate based on the target decision bit rate at the next moment; The target video is transmitted based on the adjusted sending rate and video editing bit rate.
12. A training device for a video bit rate adaptive network, characterized in that: include: A processing unit, configured to sample N first network state groups of different classes based on N meta-learning tasks, and assign the N first network state groups of different classes to N initialized video bit rate adaptive networks to obtain N first networks, wherein each of the first network state groups is used to simulate a first network environment of each of the first networks, and the initialized video bit rate adaptive network is provided with an initialized network parameter, and N is an integer greater than 1; An acquisition unit is used to acquire a first reward value corresponding to each decision code rate based on the decision code rate output by each first network and the first network environment, so as to perform an inner loop update on the initialized network parameters of the first network, including: simulating the first network environment based on the first network state group, wherein each first network environment includes K first bandwidth trajectories of different time lengths, and K is an integer greater than 1; for each first network, inputting K observations and K decision code rates at the previous moment into the first decision network of the first network, and outputting K decision code rates at the current moment through the first decision network; and The decision code rates interact with the K first bandwidth trajectories at the current moment in the first network environment respectively to obtain K observation quantities at the current moment; the K observation quantities and the K decision code rates at the current moment are input into the first evaluation network of the first network, and the evaluation value at the next moment is output through the first evaluation network; based on the K observation quantities at the current moment, the K first reward values corresponding to the K decision code rates at the current moment are calculated; based on the K first reward values corresponding to the K decision code rates at the current moment, the initialization network parameters of the first decision network and the first evaluation network are updated to perform an inner loop update on the initialization network parameters of the first network; The processing unit is further used to resample N second network state groups of different classes corresponding to the N meta-learning tasks, and distribute the N second network state groups of different classes to the N first networks updated by the inner loop, wherein each second network state group is used to simulate the second network environment of each first network updated by the inner loop; The acquisition unit is further used to acquire the second reward value corresponding to each decision bit rate based on the decision bit rate output by each first network after the inner loop update and the second network environment, so as to perform an outer loop update on the initialized network parameters of the initialized video bit rate adaptive network, including: simulating the second network environment based on the second network state group, wherein each second network environment includes K second bandwidth trajectories of different time lengths; for each first network after the inner loop update, inputting the K observations and K decision bit rates at the previous moment into the first decision network of the first network after the inner loop update, and outputting K new decision bit rates at the current moment. ; The K new decision bit rates at the current moment interact with the K second bandwidth trajectories at the current moment in the second network environment respectively to obtain K new observation quantities at the current moment; based on the K new observation quantities at the current moment, the K second reward values corresponding to the K new decision bit rates at the current moment are calculated; based on the K second reward values corresponding to the K new decision bit rates at the current moment of the N meta-learning tasks, the K second reward values are added and averaged to obtain the reward average value at the current moment; based on the reward average value at the current moment, the initialized network parameters of the initialized video bit rate adaptive network are updated to perform an outer loop update on the initialized network parameters of the initialized video bit rate adaptive network; The processing unit is further configured to repeatedly perform the operations of sampling the first network state group, obtaining the first reward value, the inner loop updating, sampling the second network state group, obtaining the second reward value, and the outer loop updating until the reward values of the N meta-learning tasks after the inner loop updating meet the convergence condition, and the initialized video bitrate adaptive network is updated to obtain the video bitrate adaptive network to be adjusted; The acquisition unit is further used to acquire a network status group of a current category corresponding to the target video transmission process, wherein the network status group of the current category is used to simulate the current category network environment of the video bit rate adaptive network to be adjusted; The acquisition unit is further used to acquire the reward value corresponding to each decision bit rate based on the decision bit rate output by the video bit rate adaptive network to be adjusted and the current category network environment, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted until the reward value meets the convergence condition, thereby obtaining the target video bit rate adaptive network corresponding to the current category.
13. The device according to claim 12, characterized in that The device also includes a determining unit; The acquisition unit is further configured to acquire a first historical network trajectory, wherein the first historical network trajectory includes a throughput, a packet loss rate, and a round-trip delay of the network in a first historical time period; The processing unit is further used to perform bandwidth estimation and filtering processing on the first historical network trajectory based on the time dimension to obtain a mean, a standard deviation, and a fluctuation difference corresponding to each time window; The determining unit is used to determine a mean interval threshold, a standard deviation interval threshold, and a fluctuation difference interval threshold based on the mean, standard deviation, and fluctuation difference corresponding to each time window; The processing unit is specifically used for: Based on the N meta-learning tasks, the mean interval threshold, the standard deviation interval threshold and the fluctuation difference interval threshold, N different classes of the first network state groups are sampled, and the N different classes of the first network state groups are assigned to the N initialized video bit rate adaptive networks to obtain N first networks.
14. The device according to claim 13, characterized in that The processing unit is specifically used for: Based on the time dimension, the bandwidth estimation and filtering processing are performed on the second historical network trajectory to obtain the unit time bandwidth corresponding to each time window; Based on the unit time bandwidth corresponding to each of the time windows, performing network state group distribution probability statistics on the second historical network trajectory to obtain a first probability density function; Sampling N of the meta-learning tasks based on the first probability density function; Based on the N meta-learning tasks, the mean interval threshold, the standard deviation interval threshold, and the fluctuation difference interval threshold, K bandwidth trajectories corresponding to each of the meta-learning tasks are generated by a network trajectory generator to obtain N different classes of the first network state groups, wherein each of the first network state groups includes K bandwidth trajectories, and K is an integer greater than 1.
15. The device according to claim 14, characterized in that The processing unit is further used to perform distribution probability statistics of the network state group on the network trajectory within the target time period during the target video transmission process to obtain a current probability density function corresponding to the target time period; and calculate the cross entropy between the current probability density function and the first probability density function; The acquisition unit is further configured to re-update the initialized network parameters of the initialized video bit rate adaptive network based on the current probability density function if the cross entropy is greater than a cross threshold, so as to obtain a new video bit rate adaptive network to be adjusted.
16. The device according to claim 14, characterized in that The processing unit is specifically used for: Based on the unit time bandwidth corresponding to each of the time windows, a network state scatter plot corresponding to the second historical network trajectory is obtained by mapping; Density calculation and linear interpolation processing are performed on the network status scatter plot to obtain the first probability density function.
17. The device according to claim 14, characterized in that The acquisition unit is specifically used for: Resampling N new meta-learning tasks based on the first probability density function; Based on the N new meta-learning tasks, the mean interval threshold, the standard deviation interval threshold and the fluctuation difference interval threshold, K new bandwidth trajectories corresponding to each of the new meta-learning tasks are generated by a network trajectory generator to obtain N second network state groups of different classes, wherein each of the second network state groups includes K new bandwidth trajectories.
18. The device according to claim 14, characterized in that The processing unit is further used for: Calculate the mean difference, standard deviation difference and fluctuation difference between the network status group of the current category and the network status group at the previous time during the target video transmission process; If any one of the mean difference, the standard deviation difference or the fluctuation difference difference satisfies the change condition of the interval range, it is determined that the category of the network state group of the current category has changed, and the decision bit rate output based on the video bit rate adaptive network to be adjusted and the current category network environment are executed to obtain the reward value corresponding to each decision bit rate, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted until the reward value satisfies the convergence condition, and the target video bit rate adaptive network corresponding to the current category is obtained; If the mean difference, the standard deviation difference and the fluctuation difference do not meet the change conditions of the interval range, it is determined that the category of the network status group of the current category has not changed, and the target video bit rate adaptive network corresponding to the network status group at the previous time is used as the target video bit rate adaptive network corresponding to the current category.
19. The device according to claim 12, characterized in that The acquisition unit is specifically used for: Simulating the current category network environment based on the current category network state group, wherein the current category network environment includes K current bandwidth trajectories of different time lengths; Obtain the current observation value and current decision bit rate corresponding to the target video transmission process; Inputting the current observation amount and the current decision bit rate into the decision network to be adjusted of the video bit rate adaptive network to be adjusted, and outputting the decision bit rate at the next moment through the decision network to be adjusted; The decision code rate at the next moment interacts with the K current bandwidth trajectories at the next moment in the current category network environment to obtain K observation quantities at the next moment; Inputting the K observation quantities at the next moment and the decision bit rate at the next moment into the evaluation network to be adjusted of the video bit rate adaptive network to be adjusted, and outputting an evaluation value through the evaluation network to be adjusted; Calculating K reward values at the next moment based on the K observation quantities at the next moment and the decision code rate at the next moment; The network parameters of the decision network to be adjusted and the evaluation network to be adjusted are updated based on the K reward values at the next moment, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted.
20. The device according to claim 12, characterized in that The processing unit is further used for: Compare the current packet loss rate and the current round-trip delay corresponding to the network status group of the current category with the packet loss threshold and the delay threshold respectively; When the current packet loss rate is greater than the packet loss threshold or the current round-trip delay is greater than the delay threshold, a protection fallback mechanism is enabled, and the use of the decision network of the video bit rate adaptive network to be adjusted for bit rate selection is suspended.
21. An application device of a video bit rate adaptive network, characterized in that: include: An acquisition unit, used to acquire the current network state group, the current observation amount and the target decision bit rate at the current moment corresponding to the target video transmission process; A processing unit, configured to input the current observation amount and the target decision bit rate at the current moment into a target decision network in the target video bit rate adaptation network according to any one of claims 1 to 9 if the category of the current network state group has not changed, and output the target decision bit rate at the next moment through the target decision network; A determination unit, configured to adjust a transmission rate and a video editing bit rate based on the target decision bit rate at the next moment; The processing unit is further used to transmit the target video based on the adjusted sending rate and video editing bit rate.
22. The device according to claim 21, characterized in that The processing unit is further configured to simulate the current network environment of the video bit rate adaptive network to be adjusted based on the network state group of the current category if the category of the current network state group changes; The acquisition unit is further used to acquire a reward value corresponding to each decision bit rate based on the decision bit rate output by the video bit rate adaptive network to be adjusted and the current network environment, so as to perform an inner loop update on the network parameters of the video bit rate adaptive network to be adjusted until the reward value meets the convergence condition, thereby obtaining a new target video bit rate adaptive network corresponding to the current network state group; The processing unit is further used to input the current observation amount and the target decision bit rate at the current moment into the target decision network in the new target video bit rate adaptive network, and output the target decision bit rate at the next moment through the target decision network; The determining unit is further configured to adjust the sending rate and the video editing bit rate based on the target decision bit rate at the next moment; The processing unit is further used to transmit the target video based on the adjusted sending rate and video editing bit rate.
23. A computer device comprising a memory, a processor and a bus system, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the processor implements the steps of the method according to any one of claims 1 to 9, and implements the steps of the method according to any one of claims 10 to 11; The bus system is used to connect the memory and the processor so that the memory and the processor can communicate with each other.
24. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 9 and the computer program implements the steps of the method according to any one of claims 10 to 11.
25. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 9 and the computer program implements the steps of the method according to any one of claims 10 to 11.
Citation Information
Patent Citations
Real-time video code rate self-adaptive regulation and control method and system based on reinforcement learning
CN111901642A
Subjective quality evaluation method for real-time video communication
CN114401364A