Network high-definition player streaming media decoding method, device and storage medium
By building a dual-branch neural network and a multi-level cache system, combined with the event triggering mechanism, the decoding instability problem in ultra-high-definition video streams is solved, efficient decoding control and playback stability is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202510273088.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-10
AI Technical Summary
When handling 4K/8K ultra-high-definition video streams, existing decoding technology cannot effectively deal with uncertain factors such as network bandwidth fluctuations, video code rate changes, and decoding chip heating, resulting in playback lag, mosaics and audio and video out-synchronization, affecting the user experience.
By building a dual-branch neural network structure, combining timing feature extraction and spatial feature extraction, a multi-level streaming media cache system is built, and an event triggering mechanism and adaptive decoding control strategy are adopted to monitor and adjust decoding parameters in real time, and optimize the coordinated work of the network layer, decoding layer and display layer.
It significantly improves the accuracy of decoding prediction, reduces the playback start delay and lag rate, improves system stability, and ensures continuous and stable playback of ultra-high-definition video streams.
Smart Images

Figure CN119815042B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of streaming media decoding, and in particular to a method, device and storage medium for decoding streaming media of a network high-definition player. Background Art
[0002] With the popularization of 5G networks and the continuous improvement of video content quality, 4K / 8K ultra-high-definition video streaming services have become an important application scenario for network HD players. However, in the actual decoding process, due to the influence of uncertain factors such as network bandwidth fluctuations, dynamic changes in video bit rates, and heating of decoding chips, it is easy to cause problems such as freezes, mosaics, and asynchrony between audio and video in the playback screen, which seriously affects the user viewing experience.
[0003] Existing decoding technologies mainly use fixed parameter decoding schemes and MMSE-based inter-frame prediction methods. When processing ultra-high-definition video streams of new generation coding standards such as H.265 / HEVC, the complex inter-frame dependencies and limited computing power of playback devices often lead to increased decoding delays. In addition, traditional fixed caching strategies are difficult to cope with network fluctuations and video bit rate changes, and cannot meet the needs of high-quality streaming media playback. Summary of the invention
[0004] The present invention provides a method, device and storage medium for decoding streaming media of a network high-definition player, which realizes the coordinated optimization of the network layer, decoding layer and display layer, reduces the playback startup delay and improves the overall stability of the network high-definition player.
[0005] In a first aspect, the present invention provides a method for decoding streaming media of a network high-definition player, the method comprising:
[0006] The real-time temperature data of the decoding chip in the network high-definition player, the network bandwidth jitter data, and the video bit rate fluctuation data are collected and stored in parallel through multi-threading to obtain the streaming media dynamic feature data set;
[0007] Classify and train the data in the streaming media dynamic feature data set according to the inter-frame prediction difficulty to obtain an inter-frame correlation prediction model;
[0008] Constructing a multi-level streaming media cache system based on the inter-frame correlation prediction model;
[0009] Real-time monitoring of the cache status, network status and chip status in the multi-level streaming media cache system to obtain a decoding event sequence;
[0010] The decoding parameters are adjusted in real time according to the decoding event sequence to obtain streaming media decoding control instructions.
[0011] In a second aspect, the present invention provides a streaming media decoding device for a network high-definition player, and the streaming media decoding device for the network high-definition player includes:
[0012] A parallel acquisition module, configured to perform multi-threaded parallel acquisition and storage on real-time temperature data, network bandwidth jitter data, and video bitrate fluctuation data of a decoding chip in the network high-definition player to obtain a streaming media dynamic feature data set;
[0013] A classification training module, configured to classify and train the data in the streaming media dynamic feature data set according to the inter-frame prediction difficulty to obtain an inter-frame correlation prediction model;
[0014] A construction module, configured to construct a multi-level streaming media cache system based on the inter-frame correlation prediction model;
[0015] A real-time monitoring module, configured to perform real-time monitoring on the cache state, network state, and chip state in the multi-level streaming media cache system to obtain a decoding event sequence;
[0016] A real-time adjustment module, configured to perform real-time adjustment on decoding parameters according to the decoding event sequence to obtain a streaming media decoding control instruction.
[0017] In a third aspect of the present invention, there is provided a computer-readable storage medium storing instructions, which when run on a computer, cause the computer to execute the above-mentioned streaming media decoding method for a network high-definition player.
[0018] In the technical solution provided by the present invention, by constructing a dual-branch neural network structure and combining temporal feature extraction and spatial feature extraction, the decoding prediction accuracy for different types of video frames is significantly improved; based on the hierarchical design of the multi-level streaming media cache system, the collaborative optimization of the network layer, decoding layer, and display layer is realized, and the playback startup delay is reduced; an adaptive decoding control strategy using an event trigger mechanism reduces the system resource occupancy and playback stuttering rate; an attention fusion module is introduced to improve the accuracy of inter-frame correlation prediction, making the decoding resource allocation more reasonable and effectively avoiding the problem of overheating of the decoding chip; through the optimization training of the multi-task loss function, the decoding control instruction can simultaneously take into account temperature control, performance optimization, and resource allocation, significantly improving the overall stability of the system; based on the dynamic scheduling mechanism of event priorities, the system can respond to various abnormal events in a timely manner, ensuring the continuous and stable playback of the ultra-high-definition video stream. Description of the Drawings
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required in the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a schematic diagram of the steps of the streaming media decoding method of the network high-definition player in the embodiments of the present invention;
[0021] Figure 2 It is a schematic diagram of the structure of the streaming media decoding device of the network high-definition player in the embodiments of the present invention. Specific embodiments
[0022] The embodiments of the present invention provide a streaming media decoding method, device and storage medium for a network high-definition player. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the term "comprising" or "having" and any deformation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0023] For ease of understanding, the following describes the specific process of the embodiments of the present invention. Please refer to Figure 1 , an embodiment of the streaming media decoding method of the network high-definition player in the embodiments of the present invention includes:
[0024] Step S1: Multithreadedly and parallelly collect and store the real-time temperature data, network bandwidth jitter data, and video bitrate fluctuation data of the decoding chip in the network high-definition player to obtain a streaming media dynamic feature data set;
[0025] It can be understood that the execution subject of the present invention can be a streaming media decoding device of a network high-definition player, or a terminal or a server. Specifically, it is not limited here. The embodiments of the present invention are described by taking the server as the execution subject as an example.
[0026] Specifically, a data acquisition mechanism is designed in the player system to ensure that the real-time temperature data of the decoding chip, network bandwidth jitter data, video bitrate fluctuation data, and system resource status data are collected and stored in parallel in a multi-threaded manner, thereby constructing a streaming media dynamic feature dataset. The temperature data of the decoding chip is obtained through periodic sampling, including the measurement of the core temperature and the surface temperature. The core temperature is directly read from the internal sensor of the decoding chip, while the surface temperature is supplemented and monitored by an external temperature sensor, reflecting the heat dissipation and thermal stability of the decoding chip. During the data acquisition process, each sample is attached with an accurate timestamp to ensure the accuracy of subsequent data alignment. At the same time, a sliding window detection is performed on the network bandwidth data to obtain the bandwidth jitter situation. Sliding window detection is a data processing method that continuously calculates the change trend of the bandwidth within a specified time window and can capture instantaneous fluctuations, thereby obtaining a network bandwidth jitter data stream. This data stream includes two key parameters, the bandwidth change rate and the packet loss rate. The former reflects the dynamic change of the network transmission rate, while the latter measures the severity of data loss. These two together determine the smoothness of streaming media playback. These parameters are calculated by regularly capturing network traffic data packets and trend prediction is performed in combination with historical data to take appropriate preventive measures when signs of network instability appear. At the same time, the bitrate of the input video stream is dynamically monitored. Since the bitrate fluctuation of the video stream directly affects the computational load of the decoding chip and the network transmission pressure, the video bitrate is analyzed in real time and its change trend is recorded. A method based on packet parsing is used to extract the bitrate information of each segment of video data from the transport stream, or the video bitrate data before decoding is directly obtained through the decoder module of the player to form a video bitrate fluctuation data stream. The CPU occupancy rate, memory usage rate, and GPU usage rate are monitored in real time to form a system resource status data stream. The monitoring of the CPU occupancy rate is obtained through the process management module of the operating system, while the memory usage rate is calculated by analyzing the memory allocation situation and cache utilization. At the same time, the GPU usage rate is queried through the API interface provided by the graphics card driver. These data together constitute the system resource status data stream to help judge the resource bottleneck problem during the decoding process. The real-time temperature data stream, network bandwidth jitter data stream, video bitrate fluctuation data stream, and system resource status data stream are timestamp-aligned to ensure that data from different sources are matched according to the same time axis, constructing a multi-dimensional state data matrix. The multi-dimensional state data matrix is stored in a circular buffer to obtain a streaming media dynamic feature dataset. A circular buffer is a storage structure that can maintain a high access efficiency in scenarios where the data volume is large and continuous updates are required, and avoid memory overflow problems. Through circular buffer storage, old data will be automatically overwritten by new data, thus ensuring that the dataset always maintains the latest state without occupying too much storage resources.
[0027] Step S2: Classify and train the data in the streaming media dynamic feature dataset according to the inter-frame prediction difficulty to obtain an inter-frame correlation prediction model;
[0028] Specifically, data preprocessing is performed on the dynamic feature dataset of the streaming media. The temperature features, network features, and video bitrate features of the decoding chip are standardized to form a temperature feature vector, a network feature vector, and a bitrate feature vector respectively. The temperature feature vector consists of the core temperature and surface temperature data of the decoding chip, and the mean normalization method is used to ensure that all data is within the same scale range. The network feature vector is composed of network bandwidth jitter data and packet loss rate. The sliding window mean filtering method is used to reduce data noise while maintaining the temporal characteristics of the data. The bitrate feature vector is composed of video bitrate fluctuation data, and normalization transformation is used to ensure that the bitrate features of different videos can be compared within the same numerical range. The data is classified according to the H.265 / HEVC coding standard to distinguish the different characteristics of I-frames, P-frames, and B-frames. An I-frame is a completely independent frame that contains complete image information and does not depend on other frames for decoding. Its encoding complexity is usually low. A P-frame is predicted based on the previous frame and contains motion vector information, and its encoding complexity is relatively high. A B-frame depends on bidirectional prediction of the previous and next frames, so it is more complex in motion vector calculation. To quantify this complexity, the entropy coding complexity and motion vector distribution of each frame type are calculated. The entropy coding complexity is measured by statistically analyzing the dispersion of each pixel block within the frame, and the motion vector distribution is obtained by analyzing the motion direction and amplitude of the prediction blocks within the frame. These information constitute the frame type feature matrix. Based on the frame type feature matrix, a dual-branch neural network structure is constructed, which consists of a temporal feature extraction branch and a spatial feature extraction branch. Among them, the temporal feature extraction branch uses a three-layer bidirectional LSTM network, with 256 LSTM units in each layer. This structure captures both forward and backward time dependencies, enabling the network to better understand the dynamic change patterns between video frames. The spatial feature extraction branch uses a ResNet structure, which contains four residual block groups. Each residual block group consists of three convolutional layers and skip connections. The convolutional layers can extract the spatial features within the frame, and the design of the skip connections can alleviate the problem of gradient disappearance and improve the training stability of the deep network. The combination of the dual-branch structure enables the model to learn both the temporal correlation between frames and extract the spatial features within the frame, thereby enhancing the accuracy of decoding prediction. An attention fusion module is added to the output end of the dual-branch network. This module consists of a channel attention sub-module and a spatial attention sub-module. The channel attention sub-module enhances the attention to key features by calculating the importance weights of each feature channel, while the spatial attention sub-module highlights the recognition ability of the target area by calculating the weighted distribution of spatial features. After being processed by the attention fusion module, the obtained fused feature map can more effectively express the dynamic change patterns between frames during the decoding process. The fused feature map is input into the decoding prediction head network, which consists of three parallel branches, respectively used to predict the inter-frame correlation, decoding resource requirements, and cache allocation strategy.Among them, the inter-frame correlation prediction branch estimates the similarity between the current frame and the previous and next frames, thereby guiding the adjustment of the weights of inter-frame prediction during the decoding process; the decoding resource requirement prediction branch is used to estimate the computing resources required during the decoding process, including CPU occupancy, memory requirements, and GPU load, in order to dynamically optimize the decoding parameters; the cache allocation strategy prediction branch calculates a reasonable cache management scheme to ensure that there will be no stuttering problems caused by insufficient or overloaded caches during the decoding process. All three branches use a two-layer fully connected network for prediction. The number of neurons in the first layer of the fully connected layer is 512, providing sufficient feature expression ability, while the number of neurons in the second layer of the fully connected layer matches the respective prediction target dimensions to ensure the accuracy of the prediction results. During the training process, a multi-task loss function is used to optimize the decoding prediction head network. Among them, the inter-frame correlation loss is used to measure the error of the predicted inter-frame similarity, the resource prediction loss is used to evaluate the accuracy of the decoding resource requirement prediction, and the cache optimization loss is used to optimize the rationality of the cache management strategy. These loss functions are combined in a weighted manner and the Adam optimizer is used for gradient update to ensure that the model can converge efficiently and has good generalization ability. After multiple rounds of training and optimization iterations, the model can accurately predict the inter-frame correlation, decoding resource requirements, and cache allocation strategy, thereby effectively guiding the decoding process of the streaming media player and improving the decoding efficiency and playback stability.
[0029] Step S3: Construct a multi-level streaming media cache system based on the inter-frame correlation prediction model;
[0030] Specifically, hierarchical mapping is performed on the inter-frame correlation, decoding resource requirements, and cache allocation strategy output by the inter-frame correlation prediction model to obtain multi-level cache configuration parameters. The multi-level cache configuration parameters include the network layer cache size, the decoding layer frame type cache ratio, and the display layer buffer capacity. These parameters determine the resource allocation method of the cache system at different levels, thus affecting the stability and efficiency of the entire streaming media decoding. Based on the multi-level cache configuration parameters, cache space allocation is performed on the network layer to form a network layer buffer structure. The size of the network layer buffer structure is dynamically adjusted according to the current video bitrate, and its capacity range is set between twice and four times the video bitrate, so as to provide sufficient buffer space during network fluctuations to reduce the probability of packet loss or stuttering. The specific adjustment method of the network layer cache depends on the real-time network bandwidth jitter situation. When the network is stable, the cache occupancy is reduced to reduce data latency. When the network jitter is large or the packet loss rate is high, the cache capacity is appropriately increased to enhance the anti-jitter ability. At the same time, the data storage strategy of the network layer cache needs to match the data acquisition mechanism of the decoding layer, and an intelligent cache replacement strategy is adopted, such as a discard strategy based on temporal priority, to ensure that important frames are preferentially retained without affecting the continuity of streaming media playback. After the network layer cache allocation is completed, cache partitioning is performed on the decoding layer to form a decoding layer cache structure. The cache structure of the decoding layer consists of an I-frame buffer, a P-frame buffer, and a B-frame buffer, and its capacity is dynamically configured according to the analysis results of the inter-frame correlation prediction model. Since the I-frame is a key frame for independent decoding, the priority of its cache area is the highest, so the capacity of the I-frame buffer is set to the first target value to ensure that the I-frame can be stored and read in time. The P-frame depends on the previous frame for decoding, and the capacity of its cache area is set to the second target value. The B-frame decodes by referring to both the previous and next frames, and its cache requirement is lower. However, to prevent decoding anomalies caused by inter-frame prediction errors, the capacity of the B-frame buffer is set to the third target value. The partitioning of the decoding layer cache not only needs to consider the cache requirements of different frame types but also needs to combine the decoding resource requirement prediction results to adjust the cache strategy when the chip resources are limited. For example, when the decoding pressure is high, the capacity of the B-frame buffer is reduced to reduce the decoding calculation burden and improve the decoding efficiency. Based on the multi-level cache configuration parameters, cache configuration is performed on the display layer to ensure that the video frame is smoothly transmitted to the display unit after decoding. The core of the display layer cache structure is two buffers, and their sizes are both set to 1.5 times the pixel data required by the current display resolution. This design provides sufficient buffer space during display frame switching to avoid screen tearing or frame loss. The strategy of the display layer cache works in coordination with the decoding layer cache, that is, when the cache space of the decoding layer is tight, part of the display layer cache is preferentially released to ensure the smooth data transmission of the decoding layer. At the same time, when the display frame is stable, the double-buffer mechanism is used for seamless switching of the display frame to reduce screen stuttering and improve the playback smoothness.After the configuration of each layer's cache structure is completed, optimize the data transfer channels between the network layer, decoding layer, and display layer to form an efficient inter-layer data transfer strategy. The design of the inter-layer data transfer strategy takes into account the priority of the data stream. Under normal circumstances, the transmission priority of I-frame data is the highest, followed by P-frames, and the transmission priority of B-frames is the lowest. Therefore, the data channel dynamically adjusts the transmission scheduling order according to the prediction result of the inter-frame correlation to ensure the priority transmission of key data during the decoding process. At the same time, the data transfer strategy combines the decoding resource status to avoid buffer overflow caused by insufficient processing capacity when the decoding layer has data accumulation. In implementation, an adaptive flow control mechanism is adopted to dynamically adjust the transmission rate of the data stream by monitoring the length of the decoding queue to ensure the stable operation of the decoder. At the same time, the data transfer of the display layer adopts an intelligent scheduling mechanism to reduce the phenomenon of frame loss. For example, when the display layer cache is about to be exhausted, fetch the next frame of data from the decoding layer cache in advance to maintain smooth playback. By integrating the network layer buffer structure, decoding layer cache structure, display layer cache structure, and inter-layer data transfer strategy, a multi-level streaming media cache system is constructed.
[0031] Parse the cache ratio of the decoded layer frame type and the encoding parameters in the multi-level cache configuration parameters to determine the cache requirements for each frame type. The cache ratio of the decoded layer frame type reflects the allocation ratio of I-frames, P-frames, and B-frames in the cache, while the encoding parameters include the size characteristics and decoding complexity of different frame types under the H.265 / HEVC encoding standard. During the parsing process, combine the analysis results of the inter-frame correlation prediction model to calculate the frame grouping weight values for different frame types. These weight values are used to measure the importance of various frames in the cache and guide the subsequent cache allocation strategy. Since I-frames are key frames, they have strong decoding independence and relatively large storage volumes, so a higher cache weight value is assigned. P-frames have moderate cache requirements due to their high temporal correlation, and B-frames rely on the previous and next frames for bidirectional prediction, so the storage space is appropriately reduced during cache allocation to optimize the overall cache efficiency. Calculate the cache capacity according to the frame grouping weight values and the total system memory capacity to generate a frame type cache allocation plan. The core of cache capacity calculation is to ensure that the system memory can be reasonably allocated to different types of frames, while ensuring that the decoding process will not cause data loss or decoding stuttering due to insufficient cache. During the calculation process, comprehensively consider the maximum available memory of the decoding chip, the dynamic memory requirements during system operation, and the bitrate characteristics of the current video stream, and make dynamic adjustments in combination with the historical data provided by the inter-frame correlation prediction model. For example, in the case of a high-bitrate video stream, appropriately increase the capacity of the I-frame cache, while in a low-bitrate scenario, reduce the I-frame cache and leave more storage space for P-frames and B-frames. Through the cache capacity calculation method, ensure that the cache space allocation for different frame types is reasonable and maximize the decoding efficiency within the range allowed by the system memory. Based on the frame type cache allocation plan, perform partition management on the system memory to generate a physical address table for the cache area. The partition management of the system memory considers data alignment and the efficiency of cache access, so fixed-size memory blocks are used for partitioning. Each partition block corresponds to a specific cache area and is indexed through the physical address table. For example, map the I-frame cache area to a high-priority memory area to ensure its access speed, while map the P-frame and B-frame cache areas to medium-priority and low-priority memory areas respectively to optimize memory utilization. Based on the physical address table of the cache area, construct a multi-level page table structure to form a cache access control table. The role of the multi-level page table structure is to provide a more flexible and efficient cache management mechanism, enabling the decoding chip to perform fast address conversion when reading cache data and reducing the memory fragmentation problem. The multi-level page table structure includes a global page table, a first-level page table, and a second-level page table. The global page table is used to manage the distribution of the entire cache space, the first-level page table is used to map the frame type cache area, and the second-level page table is used to index specific storage blocks. In this way, achieve efficient address lookup during the decoding process, reduce the latency of memory access, and at the same time ensure the dynamic management ability of the cache space, enabling the cache space to be reallocated according to the load situation during system operation.Design a cache scheduling scheme based on the cache access control table to achieve efficient reading and storage of the cache. The core of the cache scheduling scheme lies in optimizing the access order of data and reducing unnecessary cache replacement operations, thereby improving the continuity of decoding. The data access priority of I-frames is the highest. Therefore, the cache scheduling scheme needs to ensure that I-frame data can be loaded into the decoder as soon as possible and the cache can be released in a timely manner after decoding to reduce the risk of cache overflow. The cache access scheduling of P-frames needs to consider their temporal correlation with I-frames, and the data of the upcoming P-frames to be used is pre-loaded through a prediction algorithm. The cache access scheduling of B-frames adopts a delayed loading strategy to reduce the occupancy of system memory. To optimize the efficiency of cache scheduling, combined with the cache prefetching technology, the data to be accessed in the future is predicted by analyzing the playback mode of the video stream and pre-loaded into the cache in advance, thereby reducing the stuttering problem caused by the untimely arrival of data during the decoding process. Create a decoding layer cache structure based on the buffer physical address table, cache access control table, and cache scheduling scheme. The cache structure includes an I-frame buffer with a capacity of the first target value, a P-frame buffer with a capacity of the second target value, and a B-frame buffer with a capacity of the third target value.
[0032] Step S4: Monitor the cache status, network status, and chip status in the multi-level streaming media cache system in real time to obtain a decoding event sequence;
[0033] Specifically, perform distributed sampling on the cache layer status in the multi-level streaming media cache system to obtain multi-layer cache status data. Since the cache system involves multiple levels, including network layer cache, decoding layer cache, and display layer cache, a distributed sampling strategy is adopted during the sampling process to ensure data integrity and real-time performance. Sampling the network layer cache status requires monitoring the queuing length, data loss rate, and buffer utilization rate of the cache queue, while sampling the decoding layer cache status focuses on the cache occupancy of I-frames, P-frames, and B-frames, as well as the cache hit rate. Sampling the display layer cache status tracks the buffer switching rate of frames and the picture refresh delay. Through the distributed sampling method, high-precision multi-layer cache status data is obtained with low computational overhead. Conduct multi-dimensional threshold analysis on the multi-layer cache status data to identify potential anomalies and determine event trigger conditions. Since different video coding standards and hardware resource configurations will affect the stability of the decoding process, the threshold analysis needs to be dynamically adjusted in combination with video coding characteristics and hardware resource limitations. Establish a multi-dimensional feature space for the cache status data, and statistically calculate the optimal operating range of different parameters based on historical decoding data, so as to determine reasonable hierarchical trigger thresholds. For example, under the H.265 / HEVC coding standard, the cache hit rate of I-frames should be maintained within a certain high range, while the cache utilization rates of P-frames and B-frames are adaptively adjusted according to different bitrates. In terms of hardware resources, the temperature, CPU occupancy rate, and GPU load of the decoding chip also need to be considered in the threshold analysis to ensure that the system will not affect the decoding quality due to overheating or resource overload. In this way, dynamically determine the event trigger conditions based on the cache status data, and ensure that the threshold settings can adapt to different video streams and hardware environments. Build an adaptive priority scheduling mechanism based on the event trigger conditions, so as to dynamically assign priorities according to the impact degree of events on the decoding quality. Since there are various types of events during the decoding process, such as cache overflow, network jitter, frame loss, and excessive decoding load, establish a flexible priority scheduling mechanism to dynamically adjust the decoding strategy in different situations. Adopt a weighted scoring method to calculate the priority scores according to the impact range, duration, and historical impact degree of various events, and generate a multi-level event priority table. During this process, high-impact events (such as I-frame cache loss or decoder overload) are assigned higher priorities, while low-impact events (such as short-term network fluctuations or minor cache jitters) are arranged at lower priorities, so as to ensure that the system can give priority to processing critical events and avoid affecting the overall decoding performance due to secondary events occupying too much computing resources. Rearrange the multi-level event priority table according to time sequence, and sort it in combination with the timestamp information and priority weights of events to obtain a weighted event sequence. Adopt a time-weighted sorting algorithm, that is, consider the time factor when calculating the priority of events, so that newly occurred high-priority events can be responded to faster, while historical events are appropriately adjusted according to their remaining impact degree.Input the weighted event sequence into the policy mapping network to match response policies for various events by leveraging historical decoding control experience. The policy mapping network, based on deep learning or rule matching methods, learns the best event response solutions from past decoding logs and control policies and provides optimal processing strategies for the current event sequence. For example, when the system detects a sharp drop in network bandwidth, the policy mapping network automatically adjusts the bitrate or increases the cache capacity based on historical experience to reduce playback stuttering. When the system detects that the temperature of the decoding chip exceeds the safe range, it reduces the decoding frequency or adopts a hardware acceleration mode to lower the temperature. The policy mapping network also needs to have an adaptive learning ability, that is, it can dynamically optimize the policy matching scheme in a continuously updated decoding environment to improve the accuracy and efficiency of responses. Conduct spatio-temporal correlation analysis based on the weighted event sequence and the event processing policy group to identify the temporal dependence relationships and spatial coupling effects between different events, and obtain the decoded event sequence. The temporal dependence relationships are reflected in the occurrence order and the duration of influence of events. For example, a buffer overflow event often leads to the accumulation of the decoding queue and further causes frame loss. Therefore, establish causal relationships in the event sequence and prioritize the key events that trigger chain reactions. The spatial coupling effects are reflected in the mutual influence between different decoding resources. For example, a high GPU load causes CPU resource tension, thus affecting the execution efficiency of the decoding task. Therefore, during the construction of the event sequence, conduct cross-analysis of the influence ranges of different events to identify potential resource competition problems and adjust the decoding scheduling policy when necessary. By combining temporal analysis and spatial correlation analysis, ensure that the decoded event sequence can accurately reflect the dynamic evolution of various events during the decoding process.
[0034] Step S5: According to the decoded event sequence, adjust the decoding parameters in real time to obtain the streaming media decoding control instruction.
[0035] Specifically, the events in the decoding event sequence are classified and parsed so as to extract corresponding control parameters for different types of events. Decoding events are classified into temperature events, performance events, network events, cache events and bitrate events. Each type of event has different degrees of impact on the decoding process. Therefore, its key features are extracted during the parsing process, and the corresponding decoding control target matrix is established. Temperature events mainly involve the changing trends of the core temperature and surface temperature of the decoding chip. This information is used to adjust the frequency and voltage of the decoding chip to prevent overheating from causing performance degradation or system protection shutdown. Performance events focus on the usage of CPU, GPU and memory, which determines the execution capability and resource scheduling method of the decoding task. Network events involve parameters such as bandwidth jitter, packet loss rate and data delay. This information is crucial for adaptive bitrate adjustment and cache control, while cache events reflect the status of the multi-level cache system, including the cache occupancy rate of I frames, P frames and B frames and the cache overflow. These data are used to optimize the cache management strategy. Bitrate events are mainly related to the dynamic bitrate changes of the input video stream, which affects the decoding complexity and computing resource requirements. Therefore, the decoding strategy needs to be adjusted in a targeted manner. By parsing all these events and constructing a decoding control target matrix, a set of decoding control parameters is formed. Hardware resource control is performed based on the decoding control target matrix to optimize resource allocation and performance management during the decoding process. The frequency and voltage of the decoding chip are adjusted to reduce power consumption at low load and improve computing power at high load, thereby achieving a balance between performance and energy efficiency. The number of decoding threads also needs to be dynamically adjusted to adapt to different decoding task requirements. For example, when playing high-resolution videos, parallel decoding threads are added to increase the decoding speed, while when playing low-bitrate videos, threads are reduced to save computing resources. At the same time, the cache allocation ratio is dynamically adjusted according to the needs of the decoding task. For example, when the network fluctuates greatly, the capacity of the network layer cache is increased to reduce data loss, and when the decoder resources are tight, the occupancy of the B frame cache is reduced to prioritize the storage space of I frames and P frames. Through the above hardware resource control measures, a dynamically optimized hardware control parameter sequence is obtained, which guides the decoding chip to make adaptive adjustments under different conditions. The hardware control parameter sequence is applied to the decoding process of the H.265 / HEVC encoded stream, and the appropriate decoding strategy is matched to ensure that the decoding process can maintain smoothness and optimize the utilization of computing resources. The decoding strategy matching of the H.265 / HEVC encoded stream is mainly based on the decoding priority and frame skipping strategy of I frames, P frames and B frames. Among them, the I frame has the highest priority because it is a key frame and must be fully decoded to ensure the integrity of the basic picture information of the video, while the decoding priority of the P frame is second. When resources are limited, the B frame is skipped to reduce computing overhead.When the system resources are scarce or the network condition is poor, some B-frames are skipped through an intelligent frame skipping strategy, thereby reducing the decoding burden and improving the playback fluency. When the resources are sufficient, all frames are ensured to be decoded completely to enhance the picture quality. By combining the decoding priority and the frame skipping strategy, the playback effect is optimized under different decoding environments, thus ensuring the stability of streaming media playback. After the decoding execution plan is determined, its parallelism is analyzed to optimize the execution order of decoding tasks and reduce the processing delay. The core of the parallelism analysis is to calculate the dependency relationships and execution timings among decoding tasks, thereby constructing a task scheduling sequence, which includes the execution order and time window of decoding tasks. During the decoding process, the decoding of I-frames needs to be prioritized, while the decoding of P-frames and B-frames is processed in parallel. Therefore, it is necessary to reasonably arrange the decoding order during task scheduling to make full use of computing resources. For example, a double-buffer mechanism is adopted to preload the data of the next frame while decoding the current frame, thereby reducing the waiting time and improving the decoding efficiency. At the same time, the thread parallelism is optimized by means of task grouping. For example, on a multi-core processor, a high-priority thread is separately allocated for the decoding of I-frames, while the decoding of P-frames and B-frames is executed in parallel by multiple threads to maximize the utilization of hardware resources. By optimizing the dependency relationships and execution timings of decoding tasks, an efficient task scheduling sequence is generated, thereby increasing the overall decoding speed. Based on the task scheduling sequence, a resource allocation model is constructed, and by calculating the dynamic balance points of CPU occupancy, memory usage, and GPU usage, an optimal resource allocation scheme is obtained. During the calculation process, an adaptive load balancing algorithm is adopted to reasonably allocate decoding tasks among different computing units. For example, when the CPU load is high, some decoding tasks are transferred to the GPU for processing, while when the GPU load is too heavy, the CPU is used for supplementary calculation, thereby avoiding the overload of a single computing resource. The memory allocation also needs to be dynamically adjusted to ensure that the decoding cache will not cause data loss or decoding failure due to insufficient memory. By calculating the dynamic balance points of computing resources, the execution efficiency of decoding tasks is ensured to be maximized, while avoiding the degradation of decoding performance caused by insufficient or overloaded resources. After optimizing the hardware control parameter sequence, the decoding execution plan, and the resource allocation plan, these parameters are integrated and mapped to generate streaming media decoding control instructions. The streaming media decoding control instructions include temperature control instructions, decoding control instructions, and resource control instructions. Among them, the temperature control instructions are used to dynamically adjust the power consumption mode of the decoding chip to ensure that the chip operates within a reasonable temperature range. The decoding control instructions are used to adjust the decoding strategy, such as optimizing the frame priority or enabling the frame skipping mechanism, to adapt to different decoding environments. The resource control instructions are used to dynamically allocate computing resources to ensure that the decoding tasks can be executed efficiently. By sending these control instructions to the decoding system, precise control of streaming media playback is achieved, enabling it to maintain the best playback effect under different network environments and hardware conditions and effectively improving the decoding efficiency and user experience.
[0036] In the embodiments of the present invention, by constructing a dual-branch neural network structure and combining temporal feature extraction and spatial feature extraction, the decoding prediction accuracy for different types of video frames is significantly improved; based on the hierarchical design of the multi-level streaming media caching system, the collaborative optimization of the network layer, decoding layer, and display layer is realized, and the playback startup delay is reduced; the adaptive decoding control strategy adopting the event-triggered mechanism reduces the system resource occupancy and playback jitter rate; the attention fusion module is introduced to improve the accuracy of inter-frame correlation prediction, making the decoding resource allocation more reasonable and effectively avoiding the overheating problem of the decoding chip; through the optimization training of the multi-task loss function, the decoding control instruction can take into account temperature control, performance optimization, and resource allocation simultaneously, significantly improving the overall stability of the system; based on the dynamic scheduling mechanism of event priorities, the system can respond to various abnormal events in a timely manner, ensuring the continuous and stable playback of the ultra-high-definition video stream.
[0037] In a specific embodiment, the process of executing step S1 may specifically include the following steps:
[0038] Periodically sample the core temperature and surface temperature of the decoding chip in the network high-definition player to obtain a real-time temperature data stream;
[0039] Perform a sliding window detection on the network bandwidth data to obtain a network bandwidth jitter data stream, and the network bandwidth jitter data stream includes the bandwidth change rate and packet loss rate;
[0040] Dynamically monitor the bit rate of the input video stream to obtain a video bit rate fluctuation data stream;
[0041] Collect the status of the system resources, and obtain a system resource status data stream by monitoring the CPU occupancy rate, memory usage rate, and GPU usage rate;
[0042] Align the time stamps of the real-time temperature data stream, network bandwidth jitter data stream, video bit rate fluctuation data stream, and system resource status data stream to obtain a multi-dimensional status data matrix;
[0043] Store the multi-dimensional status data matrix in a circular buffer to obtain a streaming media dynamic feature data set.
[0044] Specifically, an efficient multi-threaded data acquisition mechanism is established to ensure that all data streams can be obtained in real time and form a complete multi-dimensional status data matrix through time stamp alignment. In terms of the temperature monitoring of the decoding chip, the core temperature represents the temperature inside the decoding chip, while the surface temperature represents the temperature on the surface of the decoding chip package. The change trends of the two are used to judge the heat dissipation status and workload of the decoding chip. Assume that the sampling interval is , then at any moment , the core temperature and surface temperature are expressed as:
[0045]
[0046]
[0047] Among them, is the power consumption of the decoding chip, is the heat dissipation capacity, represents the ambient temperature, respectively represent the influence coefficients of power consumption converted into temperature change, internal conduction of the chip, package heat transfer, and ambient heat dissipation. After these data are sampled periodically, a real-time temperature data stream is formed. At the same time, network bandwidth jitter detection is an important part to ensure the stability of streaming media playback. Set the sliding window length to , then the bandwidth change rate within the window is calculated as:
[0048]
[0049] Among them, represents the network bandwidth at the th sampling, and another important indicator of the bandwidth jitter data stream is the packet loss rate , and its calculation formula is:
[0050]
[0051] Among them, represents the number of lost data packets within the window , represents the total number of data packets within the window. These data reflect the changes in the network conditions and play an important role in adaptive bitrate adjustment and buffer management during the decoding process. For example, in practical applications, if the bandwidth change rate is lower than the threshold 0.05 for a long time, and at the same time the packet loss rate exceeds 0.02, it indicates that the network has significant fluctuations and the video bitrate needs to be appropriately reduced to prevent stuttering. For the bitrate fluctuation monitoring of the input video stream, it is set that within the time window , the bitrate is calculated as:
[0052]
[0053] Among them, represents the size of the video data at time , represents the data packet transmission time. This data stream is used to evaluate the video transmission stability in real time. For example, if A sudden increase within a short period indicates that the video stream enters the high bitrate segment. At this time, the decoding system needs to reserve more buffer space. If decreases rapidly, the buffer pre-storage amount is appropriately reduced to save memory resources. Meanwhile, the status of system resources is collected. The CPU occupancy rate , memory usage and GPU usage are obtained through the operating system interface and recorded periodically. Assume that the calculation of the CPU occupancy rate is based on the executed task time and idle time , then:
[0054]
[0055] Similarly, the memory usage and GPU usage are calculated through similar ratios. The real-time temperature data stream, network bandwidth jitter data stream, video bitrate fluctuation data stream, and system resource status data stream are timestamp-aligned to form a multi-dimensional state data matrix. Set the data sampling interval to be consistent as , then the data matrix at time is expressed as:
[0056]
[0057] This matrix contains the decoding chip temperature, network bandwidth jitter, video bitrate change, and system resource status information. To efficiently manage this data, a circular buffer storage mechanism is adopted to ensure that storage overflow does not occur due to the growth of data volume during long-term operation. Assume that the maximum capacity of the circular buffer is , then when new data enters the buffer, it is stored in the following manner:
[0058]
[0059] When the buffer is full, the new data will overwrite the oldest data, thus always maintaining the status information of the most recent time steps.
[0060] In a specific embodiment, the process of executing step S2 may specifically include the following steps:
[0061] Perform data preprocessing on the dynamic feature dataset of the streaming media to obtain a standardized feature dataset, and the standardized feature dataset includes a temperature feature vector, a network feature vector, and a bitrate feature vector;
[0062] Classify the standardized feature dataset according to the I-frame, P-frame, and B-frame types in the H.265 / HEVC coding standard, and calculate the entropy coding complexity and motion vector distribution of each frame type to obtain a frame type feature matrix;
[0063] Construct a dual-branch neural network structure based on the frame type feature matrix. The dual-branch neural network structure includes a temporal feature extraction branch and a spatial feature extraction branch. The temporal feature extraction branch consists of three layers of bidirectional LSTM, with 256 LSTM units in each layer. The spatial feature extraction branch adopts a ResNet structure, including four residual block groups, and each residual block group contains three convolutional layers and skip connections;
[0064] Add an attention fusion module at the output ends of the two branches of the dual-branch neural network structure. The attention fusion module includes a channel attention sub-module and a spatial attention sub-module, and obtains a fused feature map by calculating the channel weight and spatial weight of the feature map;
[0065] Input the fused feature map into the decoding prediction head network. The decoding prediction head network includes three parallel branches, which respectively predict the inter-frame correlation, decoding resource requirements, and cache allocation strategy. Each branch contains two fully connected layers, where the number of neurons in the first layer is 512, and the number of neurons in the second layer matches the dimension of the prediction target;
[0066] Use a multi-task loss function to optimize and train the decoding prediction head network. The multi-task loss function includes inter-frame correlation loss, resource prediction loss, and cache optimization loss. The network parameters are iteratively updated through the Adam optimizer to obtain an inter-frame correlation prediction model.
[0067] Specifically, standardize the dynamic feature dataset of the streaming media to eliminate the scale differences between different data sources so that it can be efficiently processed by the neural network. During the standardization process, the data needs to be converted into a temperature feature vector, a network feature vector, and a bitrate feature vector. The temperature feature vector consists of the core temperature and the surface temperature of the decoding chip. The network feature vector consists of the bandwidth jitter and the packet loss rate while the bitrate feature vector consists of the instantaneous bitrate of the video stream and the bitrate distribution of the frame types To make the numerical ranges of different features consistent, the mean normalization method is adopted:
[0068]
[0069] where, represents the original data, represents the mean, Represents the standard deviation. All features are normalized to a distribution with zero mean and unit variance to improve training stability. The standardized feature dataset is classified according to the H.265 / HEVC coding standard to distinguish the different characteristics of I-frames, P-frames, and B-frames, and the entropy coding complexity and motion vector distribution of each frame type are calculated to construct a frame type feature matrix. Entropy coding complexity Reflects the amount of encoded information of pixel blocks in a frame. The calculation formula is as follows:
[0070]
[0071] Where, is the probability distribution of the th coding unit in the frame, and represents the number of coding units. The entropy coding complexity of I-frames is usually higher as it needs to store the complete picture information, while the entropy coding complexity of P-frames and B-frames is relatively lower as they are predicted based on reference frames. Motion vector distribution is an important indicator to measure the complexity of inter-frame prediction, and its calculation method is:
[0072]
[0073] Where, and represent the horizontal and vertical components of the th motion vector respectively, and is the total number of motion vectors in the frame. For B-frames, since it depends on the previous and next frames for prediction, it has a more complex motion vector distribution, while the motion vectors of P-frames mainly depend on the previous frame and its motion vector distribution is relatively single. By calculating the entropy coding complexity and motion vector distribution of each frame type, a frame type feature matrix is formed. Based on the frame type feature matrix, a dual-branch neural network structure is constructed, which consists of a temporal feature extraction branch and a spatial feature extraction branch. The temporal feature extraction branch uses a three-layer bidirectional LSTM, with each layer containing 256 LSTM units. The bidirectional LSTM captures both the forward and backward dependencies between frames simultaneously, and the calculation formula is as follows:
[0074]
[0075] Where, is the hidden state at the current time step, is the input feature, and are the weight matrices, is the bias term. The design of the bidirectional LSTM makes full use of the temporal correlation between video frames to improve the model's perception ability of dynamic changes. The spatial feature extraction branch adopts the ResNet structure, which contains four residual block groups. Each residual block group consists of three convolutional layers and skip connections. The calculation formula is:
[0076]
[0077] Among them, represents the input feature, represents the weight matrix of the convolutional layer, represents the non-linear transformation. The residual connection can effectively alleviate the vanishing gradient problem of deep networks and improve the training stability. Through the combination of the dual-branch structure, temporal features and spatial features are extracted simultaneously, providing more comprehensive information for decoding optimization. At the output end of the dual-branch neural network structure, an attention fusion module is added to enhance the expression ability of features. The attention fusion module consists of a channel attention sub-module and a spatial attention sub-module. The calculation method of the channel attention sub-module is:
[0078]
[0079] Among them, represents the channel weight, is the channel attention weight matrix. The calculation method of the spatial attention sub-module is similar:
[0080]
[0081] Among them, represents the spatial position weight, is the spatial attention weight matrix. By calculating the channel weight and spatial weight of the feature map, the expression ability of key features is enhanced, and the accuracy of decoding prediction is improved. The fused feature map is input into the decoding prediction head network, which contains three parallel branches, respectively used to predict the inter-frame correlation, decoding resource requirements, and cache allocation strategy. Each branch consists of two fully connected layers. The first layer contains 512 neurons for feature transformation. The calculation formula is:
[0082]
[0083] Among them, is the weight matrix, is the bias term, It is a non-linear activation function. The number of neurons in the second layer matches the dimension of the prediction target to ensure that the output result meets the requirements of decoding optimization. The decoding prediction head network is optimized and trained using a multi-task loss function, which includes inter-frame correlation loss, resource prediction loss, and cache optimization loss. The calculation formulas are as follows:
[0084]
[0085] Among them, is the inter-frame correlation loss, is the resource prediction loss, is the cache optimization loss, is the loss weight coefficient. The optimization uses the Adam optimizer, and the update rule is as follows:
[0086]
[0087] Among them, represents the model parameters, represents the learning rate, and are the first-order and second-order moment estimates respectively, is the numerical stability term. After iterative training, an inter-frame correlation prediction model is finally obtained. This model is used to optimize the decoding process of the streaming media player and improve the smoothness and stability of video playback.
[0088] Among them, the decoding prediction head network is optimized and trained using a multi-task loss function, including: performing denoising diffusion probability modeling on the outputs of the three parallel branches of the decoding prediction head network, obtaining the diffusion probability transition matrix by constructing a Markov diffusion chain; gradually injecting noise into the prediction result based on the diffusion probability transition matrix, obtaining a noisy prediction sequence through a forward diffusion process with T time steps; constructing a reverse denoising process for the noisy prediction sequence, obtaining the noise estimation value at each time step by designing a noise prediction network with a U-Net structure; inputting the noise estimation value into the variational inference network, optimizing the parameters by minimizing the evidence lower bound, and obtaining the posterior distribution parameters; performing probability resampling on the output of the decoding prediction head network based on the posterior distribution parameters to generate diverse prediction results, obtaining a set of prediction results; performing ensemble learning on the set of prediction results, fusing multiple sampling results through weighted voting to obtain the final prediction output; calculating the multi-task loss function based on the prediction output and the true label, performing weighted combination of the inter-frame correlation loss, resource prediction loss, and cache optimization loss to obtain the overall loss value; optimizing and updating the network parameters based on the overall loss value, and obtaining the optimized decoding prediction head network through iterative training using the Adam optimizer.
[0089] In a specific embodiment, the process of executing step S3 may specifically include the following steps:
[0090] Perform hierarchical mapping on the inter-frame correlation, decoding resource requirements, and cache allocation strategy output by the inter-frame correlation prediction model to obtain multi-level cache configuration parameters. The multi-level cache configuration parameters include the network layer cache size, the cache ratio of frame types in the decoding layer, and the buffer capacity of the display layer;
[0091] Based on the multi-level cache configuration parameters, allocate cache space for the network layer to obtain a network layer buffer structure, and the size of the network layer buffer structure is dynamically adjusted between 2 times and 4 times the video bit rate;
[0092] Based on the multi-level cache configuration parameters, perform cache partitioning on the decoding layer to obtain a decoding layer cache structure. The decoding layer cache structure includes an I-frame buffer with a capacity of a first target value, a P-frame buffer with a capacity of a second target value, and a B-frame buffer with a capacity of a third target value;
[0093] Based on the multi-level cache configuration parameters, perform cache configuration on the display layer to obtain a display layer cache structure. The display layer cache structure includes two buffers with a size 1.5 times the pixel data of the current display resolution;
[0094] Configure data transfer channels for the network layer buffer structure, the decoding layer cache structure, and the display layer cache structure to obtain an inter-layer data transfer strategy;
[0095] Combine the network layer buffer structure, the decoding layer cache structure, the display layer cache structure, and the inter-layer data transfer strategy to obtain a multi-level streaming media cache system.
[0096] Specifically, parse the inter-frame correlation , decoding resource requirements , and cache allocation strategy . The inter-frame correlation reflects the redundancy degree between adjacent frames. When its value is relatively high, it indicates that the inter-frame information is relatively similar, and less computing resources are used for prediction compensation during the decoding process; while when its value is relatively low, a larger cache and more complex decoding calculations are required. The decoding resource requirements measure the load conditions of the CPU, GPU, and memory, and the cache allocation strategy is used to guide how the cache system allocates resources between different levels. Based on these input parameters, calculate multi-level cache configuration parameters, including the network layer cache size , the cache ratio of frame types in the decoding layer , and the buffer capacity of the display layer . In the network layer cache allocation, the size of the network layer buffer structure Perform dynamic adjustment to ensure the continuity of data transmission. Set the cache size within the range of 2 to 4 times the video bitrate, and its calculation expression is:
[0097]
[0098] Among them, is the dynamic adjustment factor, is the network buffer compensation time. When the network condition is relatively stable, takes a smaller value to reduce cache occupancy. When the network jitter is large, takes a larger value to improve data availability, thereby reducing the possibility of stuttering. In the cache division of the decoding layer, the cache structure of the decoding layer consists of the I-frame buffer , the P-frame buffer and the B-frame buffer , and its capacity allocation is based on the frame type cache ratio , and the calculation formula is as follows:
[0099]
[0100] Among them, is mainly used to ensure the complete storage of l-frame data, is used for the storage of P-frame prediction data, Since the B-frame depends on the front and back frames, its cache requirement is relatively low. The setting of the cache ratio is affected by the inter-frame correlation . When is relatively high, reduce the I-frame and P-frame caches and increase the E-frame cache. When is relatively low, increase the cache ratio of the I-frame and P-frame to enhance the stability of decoding. In the display layer cache configuration, the size of the display layer buffer depends on the current display resolution and the data storage requirement for each pixel . To prevent tearing or display delay of the picture, use two buffers, and the size of each buffer should be at least 1.5 times the current display resolution. The calculation formula is as follows:
[0101]
[0102] Among them, represents the number of storage bytes required for each pixel, ensuring that buffer overflow does not occur during data switching and affecting the display effect. After completing the cache allocation of each layer, configure the inter-layer data transmission channel to optimize the data transmission efficiency between different layers and obtain the inter-layer data transmission strategy. Set the transmission rate from the network layer to the decoding layer as , and the transmission rate from the decoding layer to the display layer as , then it should satisfy:
[0103]
[0104] Ensure that there is no transmission bottleneck during data transfer. If lower than , resulting in the decoding layer being unable to obtain sufficient data, thus causing playback stuttering. Therefore, in the data transmission strategy, it is necessary to prioritize ensuring stability. Integrate the network layer buffer structure , the decoding layer cache structure , the display layer cache structure and the inter-layer data transmission strategy to form a multi-level streaming media cache system. This system adaptively adjusts the cache allocation strategy under different network environments and decoding conditions to optimize the stability of streaming media playback. When the network bandwidth is low and the jitter is large, the system automatically increases the network layer cache size , reducing the probability of data loss. When the decoding resources are limited, the decoding layer cache division will prioritize ensuring the storage of I-frames and adjust the cache ratio of P-frames and B-frames to reduce the computational burden. At the display layer, the double-buffer mechanism ensures the smoothness of picture switching and improves the fluency of video playback. Through the optimization of the multi-level cache system, streaming media playback can maintain a high playback quality under different conditions, while reducing the consumption of decoding resources and improving playback stability.
[0105] Among them, perform resource optimization configuration on the multi-level streaming media cache system to obtain the optimized multi-level streaming media cache system, including: performing resource correlation modeling on the network layer buffer structure and the decoding layer cache structure in the multi-level streaming media cache system, and obtaining the resource utility matrix by calculating the mapping relationship between the cache space and the decoding performance; constructing a multi-objective optimization function based on the resource utility matrix, taking the network bandwidth utilization rate, the decoding layer cache utilization rate, and the display layer cache switching frequency as optimization objectives, and obtaining the resource configuration constraint conditions; performing Lagrangian decomposition on the resource configuration constraint conditions, and obtaining the resource allocation objective function by introducing the network cost coefficient and the cache cost coefficient; designing an iterative optimization algorithm based on the resource allocation objective function, and obtaining the frame cache allocation strategy by calculating the optimal allocation ratio of resources in each cache layer; dynamically configuring the I-frame cache area, P-frame cache area, and B-frame cache area according to the frame cache allocation strategy, and obtaining the cache optimization plan by real-time adjusting the proportion of the cache space occupied by various types of frames; performing performance evaluation on the cache optimization plan, and obtaining the optimization effect evaluation index by calculating the cache hit rate, access latency, and resource utilization rate; constructing a feedback adjustment mechanism based on the optimization effect evaluation index, and obtaining the optimization parameter update strategy by dynamically adjusting the resource allocation weight and the optimization parameter; applying the cache optimization plan and the optimization parameter update strategy to the multi-level streaming media cache system to obtain the optimized multi-level streaming media cache system.
[0106] In a specific embodiment, the process of performing cache partitioning on the decoding layer based on multi-level cache configuration parameters to obtain the decoding layer cache structure may specifically include the following steps:
[0107] Parse the frame type cache ratio and encoding parameters in the multi-level cache configuration parameters to obtain the frame grouping weight value;
[0108] Perform cache capacity calculation according to the frame grouping weight value and the total system memory capacity to obtain the frame type cache allocation scheme, and perform partition management on the system memory based on the frame type cache allocation scheme to obtain the physical address table of the cache area;
[0109] Establish a multi-level page table structure for the physical address table of the cache area to obtain the cache access control table;
[0110] Construct a cache scheduling scheme based on the cache access control table, and create a decoding layer cache structure based on the physical address table of the cache area, the cache access control table, and the cache scheduling scheme. The decoding layer cache structure includes an I-frame cache area with a capacity of a first target value, a P-frame cache area with a capacity of a second target value, and a B-frame cache area with a capacity of a third target value.
[0111] Specifically, analyze the encoding structure of the video stream, and combine the output of the inter-frame correlation prediction model to calculate the cache requirements for different frame types (I-frames, P-frames, and B-frames). Set the cache weight of the I-frame to , the cache weight of the P-frame to , and the cache weight of the B-frame to . Then, according to the H.265 / HEVC encoding standard, the frame grouping weight value is calculated as:
[0112]
[0113] Among them, respectively represent the data sizes of I-frames, P-frames, and B-frames, respectively represent the average decoding times of the corresponding frames. Since I-frames need to store complete picture information, their data sizes are relatively large, while P-frames and B-frames rely on other frames for prediction, and their data sizes are relatively small. However, since B-frames require two-way prediction from the previous and next frames, their cache requirements are dynamically adjusted according to specific decoding strategies. By calculating the frame grouping weight value, the priority of each frame type in the decoding layer cache is determined, thus providing a basis for subsequent cache allocation. After calculating the frame grouping weight value, perform cache capacity calculation according to the total system memory capacity and determine the specific allocation scheme for frame type caches. Set the total capacity of the decoding layer cache to , then the cache capacities of each frame type are calculated as:
[0114]
[0115]
[0116] Among them, respectively represent the buffer capacities of I-frames, P-frames, and B-frames. Through this allocation method, the buffer ratio is dynamically adjusted according to the decoding requirements of different frame types to ensure that high-priority frames (such as I-frames) obtain sufficient storage space, while lower-priority frames (such as B-frames) appropriately reduce buffer occupancy when resources are limited. After determining the frame type buffer allocation scheme, the system memory is partitioned and managed to generate a physical address table for the buffer. Assume that the minimum storage unit of the memory partition is , then the physical address table is mapped in the following way:
[0117]
[0118] Among them, represents the physical address index of the buffer area, and each address index corresponds to a memory block of a fixed size . To improve memory access efficiency, an allocation method based on continuous address assignment is adopted to align the buffer blocks of I-frames, P-frames, and B-frames in terms of physical addresses, ensuring that the required data can be quickly read during the decoding process. After establishing the physical address table for the buffer area, a multi-level page table structure is constructed to obtain a buffer access control table. It is set that the page table structure is divided into a global page table and a local page table . Among them, the global page table is used to manage the distribution of the entire buffer space, while the local page table is used to index the buffer areas of different frame types. The specific mapping method is as follows:
[0119]
[0120]
[0121] Among them, respectively represent the local page tables of I-frames, P-frames, and B-frames. Each page table stores the index pointing to the physical address, ensuring that the decoder can quickly locate the corresponding buffer area. Based on the buffer access control table, a buffer scheduling scheme is constructed to optimize the data reading and storage strategy. It is set that the buffer scheduling priority is affected by the inter-frame correlation and the buffer occupancy , then the scheduling priority is calculated as:
[0122]
[0123] Among them, and are adjustment coefficients, and the calculation method is:
[0124]
[0125] When the cache occupancy rate is too high, adjust the cache policy, such as increasing the I-frame cache space to improve decoding stability, or reducing the B-frame cache to release memory resources. Combine the cache access control table and the scheduling scheme to form a complete cache structure for the decoding layer. After completing the cache allocation and scheduling optimization, the cache structure for the decoding layer includes an I-frame buffer with a capacity of and a P-frame buffer with a capacity of and a B-frame buffer with a capacity of . This structure ensures that different types of frame data are stored and accessed according to the preset policy, improving decoding efficiency and stability.
[0126] In a specific embodiment, the process of executing step S4 may specifically include the following steps:
[0127] Perform distributed sampling on the cache layer status in the multi-level streaming media cache system to obtain multi-layer cache status data;
[0128] Perform multi-dimensional threshold analysis on the multi-layer cache status data, set hierarchical trigger thresholds according to video coding characteristics and hardware resource limitations, and obtain event trigger conditions;
[0129] Construct an adaptive priority scheduling mechanism based on the event trigger conditions, perform dynamic priority assignment according to the impact degree of the event on the decoding quality, and obtain a multi-level event priority table;
[0130] Rearrange the multi-level event priority table according to time sequence, sort according to the timestamp information and priority weight of the event, and obtain a weighted event sequence;
[0131] Input the weighted event sequence into the policy mapping network, perform response policy matching on various events according to historical decoding control experience, and obtain an event processing policy group;
[0132] Perform spatio-temporal correlation analysis based on the weighted event sequence and the event processing policy group, and obtain a decoding event sequence by considering the temporal dependence relationship and spatial coupling effect of the event.
[0133] Specifically, construct an efficient data acquisition system to monitor the status of the cache layer in real time. The cache system includes a network layer cache, a decoding layer cache, and a display layer cache. The status of each level needs to be sampled separately to obtain the complete cache usage. Set the sampling interval of the cache status to , then the cache status at any time is represented as:
[0134]
[0135] Among them, represents the multi-level cache status data, represents the cache utilization rate of the network layer, represents the cache occupancy rate of the decoding layer, represents the occupancy of the display layer buffer. Through the distributed sampling method, these data are recorded at different time points, and a cache status time series is established. After obtaining the cache status data, multi-dimensional threshold analysis is performed to determine whether there are abnormal situations in the system, and hierarchical trigger thresholds are set according to video coding characteristics and hardware resource limitations. The trigger thresholds for each cache layer are set respectively as , then when any cache status data exceeds the corresponding threshold, an event can be triggered:
[0136]
[0137] Among them, represents the event trigger signal. When the value is 1, it means that the current cache status exceeds the threshold, and the system needs to take corresponding control measures. The setting of the threshold depends on the video coding standard and hardware resources. For example, in the H.265 coding format, the cache occupancy of the decoding layer for high-bitrate videos is usually high. Therefore, needs to be appropriately increased, while in a low-bandwidth network environment, needs to be decreased to prevent cache data overload. Based on the event trigger conditions, an adaptive priority scheduling mechanism is constructed to dynamically adjust the priorities of different events. The priority value is affected by the event type and the decoding quality impact factor . Then the priority calculation formula is as follows:
[0138]
[0139] Among them, is the adjustment coefficient, represents the weight of the event type. For example, the weight of the cache overflow event is higher, while the weight of short-term network fluctuations is lower. And reflects the change in decoding quality. For example, when the frame loss rate is high, increases, making the priority of related events higher. In this way, it is ensured that key events are responded to in a timely manner and the decoding stability of the system is improved. After determining the event priorities, the events are rearranged according to the time sequence and sorted according to the timestamp information and priority weights of the events to obtain a weighted event sequence. Suppose each event in the event list has a timestamp and a priority , then the event sorting rule is as follows:
[0140]
[0141] Among them, the sorting is based on the timestamp in descending order. At the same time, events with higher priorities are in the front to ensure that critical events can be processed first. After completing the event sorting, the weighted event sequence is input into the policy mapping network to match response policies for various events using historical decoding control experience. Set the policy mapping function as , then the event processing policy group is calculated as follows:
[0142]
[0143] Among them, represents the event processing policy group, represents the historical decoding control experience. The policy mapping network adopts a deep learning model or a rule-based matching method. For example, it uses a neural network to learn the mapping relationship between historical events and decoding control policies, or sets predefined policies based on empirical rules. For example, when the network jitter exceeds a threshold, the video bitrate is reduced to ensure smooth playback. After obtaining the event processing policy group, spatio-temporal correlation analysis is performed to comprehensively consider the temporal dependence relationship and spatial coupling effect of events, and finally a decoded event sequence is generated. Set the spatio-temporal correlation function as , then the decoded event sequence is calculated as follows:
[0144]
[0145] Among them, represents the spatial coupling factor. For example, buffer overflow in the decoding layer causes frame delay in the display layer. Therefore, it is necessary to introduce spatio-temporal correlation analysis into the event sequence to optimize the event processing order. The decoded event sequence includes different types of events such as abnormal cache status, network jitter, and decoding resource overload, and is optimized and sorted according to priorities and time sequences. This event sequence is used to adjust the streaming media decoding parameters in real time, improve the playback stability, reduce the risk of stuttering, and ensure that the system can provide the best playback experience under different networks and decoding loads. For example, when the system detects that the decoding layer buffer is about to overflow, the decoding calculation amount is reduced by adjusting the frame dropping policy, and when the network jitter is severe, the bitrate is dynamically reduced to improve the smoothness.
[0146] In a specific embodiment, the process of executing step S5 may specifically include the following steps:
[0147] Classify and analyze the events in the decoded event sequence, extract control parameters according to the characteristics of temperature events, performance events, network events, cache events, and bitrate events, and obtain a decoded regulation target matrix;
[0148] Based on the decoding regulation target matrix, perform hardware resource control, dynamically configure the frequency and voltage of the decoding chip, the number of decoding threads, and the cache allocation ratio to obtain a sequence of hardware control parameters;
[0149] According to the sequence of hardware control parameters, perform decoding strategy matching on the H.265 / HEVC encoded stream, and based on the decoding priorities and frame skipping strategies of I-frames, P-frames, and B-frames, obtain a decoding execution plan;
[0150] Conduct a parallelism analysis on the decoding execution plan, and by calculating the dependency relationships and execution timings of each decoding task, obtain a task scheduling sequence, which includes the execution order and time window of the decoding tasks;
[0151] Based on the task scheduling sequence, construct a resource allocation model, and by calculating the dynamic balance points of CPU occupancy, memory usage, and GPU usage, obtain a resource allocation plan;
[0152] Integrate and map the sequence of hardware control parameters, the decoding execution plan, and the resource allocation plan to obtain a streaming media decoding control instruction, which includes a temperature control instruction, a decoding control instruction, and a resource control instruction.
[0153] Specifically, extract key features from the decoding event sequence and classify them according to temperature events, performance events, network events, cache events, and bitrate events. Set the decoding event sequence as , and then split it into five categories of sub-event sets:
[0154]
[0155] Among them, represents temperature events, involving the core temperature and surface temperature represents performance events, involving CPU occupancy , GPU load and memory usage represents network events, involving bandwidth jitter and packet loss rate represents cache events, involving cache occupancy and cache overflow rate represents bitrate events, involving the instantaneous bitrate and bitrate fluctuation . By analyzing these events and extracting their respective control parameters, construct a decoding regulation target matrix:
[0156]
[0157] Among them, each row represents a specific decoding control dimension. For example, the first row focuses on hardware temperature and performance, the second row focuses on network and cache status, and the third row focuses on bitrate fluctuations. Based on the decoding regulation target matrix, the operating parameters of the decoding chip are adjusted to ensure that the decoding task can be executed in an optimized hardware resource environment. Set the chip frequency to , the voltage to , the number of decoding threads to , and the cache allocation ratio to . Then the hardware control parameter sequence is calculated as:
[0158]
[0159]
[0160]
[0161]
[0162] Among them, represent the default chip frequency, voltage, number of decoding threads, and cache allocation ratio respectively, is the set reference threshold, is the adjustment coefficient. When the temperature exceeds the reference value , the chip frequency will be reduced to reduce heat generation. And when the CPU occupancy is higher than , the voltage will be appropriately adjusted to improve performance. At the same time, when the GPU load is too high, the number of decoding threads will also be reduced to reduce the GPU pressure, and the cache allocation ratio will be dynamically adjusted according to the cache occupancy to ensure that the cache does not overflow. After obtaining the hardware control parameter sequence, match the decoding strategy of the H.265 / HEVC encoded stream, and formulate a decoding execution plan according to the decoding priorities and frame skipping strategies of I-frames, P-frames, and B-frames. Set the I-frame priority to , the P-frame priority to , and the B-frame priority to . Then the decoding priority is calculated as follows:
[0163]
[0164] The frame skipping strategy is affected by the decoding resource status . Set the frame skipping ratio . The calculation formula is as follows:
[0165]
[0166] When exceeds the threshold the system will preferentially skip B-frames to reduce the computational load. After determining the decoding execution scheme, parallelism analysis is performed to optimize the execution order of the decoding tasks. Set the task scheduling sequence by the decoding task execution order and the time window consists of, then the parallelism calculation is as follows:
[0167]
[0168] When exceeds a certain value, the decoding tasks need to reallocate threads to ensure balanced computational load. Based on the task scheduling sequence, a resource allocation model is constructed, and the dynamic balance points of the CPU, GPU, and memory are calculated. Set the CPU utilization rate , GPU utilization rate , memory usage , then the resource allocation scheme calculation is as follows:
[0169]
[0170] If any one of the values exceeds 1, it is necessary to reduce the decoding load, such as adjusting the number of threads or the frame skipping ratio . Integrate and map the hardware control parameter sequence, decoding execution scheme, and resource allocation scheme to obtain the complete streaming media decoding control instructions. Through these control instructions, the system adaptively adjusts the decoding parameters, improves the decoding efficiency, reduces the stuttering phenomenon, and ensures the smoothness and stability of video playback under different network conditions and computational loads.
[0171] The above describes the streaming media decoding method of the network high-definition player in the embodiments of the present invention. Next, the streaming media decoding device of the network high-definition player in the embodiments of the present invention will be described. Please refer to Figure 2 , an embodiment of the streaming media decoding device of the network high-definition player in the embodiments of the present invention includes:
[0172] A parallel acquisition module for multi-threaded parallel acquisition and storage of real-time temperature data, network bandwidth jitter data, and video bitrate fluctuation data of the decoding chip in the network high-definition player to obtain a streaming media dynamic feature data set;
[0173] A classification training module for classifying and training the data in the streaming media dynamic feature data set according to the inter-frame prediction difficulty to obtain an inter-frame correlation prediction model;
[0174] A building block for constructing a multi-level streaming media caching system based on an inter-frame correlation prediction model;
[0175] A real-time monitoring module for real-time monitoring of the cache status, network status, and chip status in the multi-level streaming media caching system to obtain a decoding event sequence;
[0176] A real-time adjustment module for real-time adjustment of decoding parameters according to the decoding event sequence to obtain a streaming media decoding control instruction.
[0177] Through the collaborative cooperation of the above-mentioned various components, by constructing a dual-branch neural network structure, combining temporal feature extraction and spatial feature extraction, the decoding prediction accuracy for different types of video frames is significantly improved; based on the hierarchical design of the multi-level streaming media caching system, the collaborative optimization of the network layer, decoding layer, and display layer is achieved, reducing the playback startup delay; the adaptive decoding control strategy using an event-triggered mechanism reduces the system resource occupancy and playback jitter rate; the introduction of an attention fusion module improves the accuracy of inter-frame correlation prediction, makes the decoding resource allocation more reasonable, and effectively avoids the problem of overheating of the decoding chip; through the optimization training of the multi-task loss function, the decoding control instruction can take into account temperature control, performance optimization, and resource allocation at the same time, significantly improving the overall stability of the system; the dynamic scheduling mechanism based on event priorities enables the system to respond to various abnormal events in a timely manner, ensuring the continuous and stable playback of ultra-high-definition video streams.
[0178] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0179] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in the present invention and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.
[0180] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.
[0181] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc., which can store program codes.
[0182] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for streaming media decoding of a network high-definition player, characterized in that, The method includes: Performing multi-threaded parallel acquisition and storage on the real-time temperature data, network bandwidth jitter data, and video bitrate fluctuation data of the decoding chip in the network high-definition player to obtain a streaming media dynamic feature dataset; Classifying and training the data in the streaming media dynamic feature dataset according to the inter-frame prediction difficulty to obtain an inter-frame correlation prediction model; specifically including: performing data preprocessing on the streaming media dynamic feature dataset to obtain a standardized feature dataset, where the standardized feature dataset includes a temperature feature vector, a network feature vector, and a bitrate feature vector; classifying the standardized feature dataset according to the I-frame, P-frame, and B-frame types in the H.265 / HEVC encoding standard, calculating the entropy coding complexity and motion vector distribution of each frame type to obtain a frame type feature matrix; constructing a dual-branch neural network structure based on the frame type feature matrix, where the dual-branch neural network structure includes a temporal feature extraction branch and a spatial feature extraction branch, and the temporal feature extraction branch is composed of three layers of bidirectional LSTM, with 256 LSTM units in each layer, and the spatial feature extraction branch adopts a ResNet structure, including four residual block groups, and each residual block group contains three convolutional layers and skip connections; adding an attention fusion module at the output ends of the two branches of the dual-branch neural network structure, where the attention fusion module includes a channel attention sub-module and a spatial attention sub-module, and obtaining a fused feature map by calculating the channel weight and spatial weight of the feature map; inputting the fused feature map into a decoding prediction head network, where the decoding prediction head network includes three parallel branches, respectively predicting the inter-frame correlation, decoding resource requirements, and cache allocation strategy, and each branch includes two fully connected layers, where the number of neurons in the first layer is 512, and the number of neurons in the second layer matches the prediction target dimension; optimizing and training the decoding prediction head network using a multi-task loss function, where the multi-task loss function includes an inter-frame correlation loss, a resource prediction loss, and a cache optimization loss, and iteratively updating the network parameters through an Adam optimizer to obtain an inter-frame correlation prediction model; Constructing a multi-level streaming media cache system based on the inter-frame correlation prediction model; Performing real-time monitoring on the cache status, network status, and chip status in the multi-level streaming media cache system to obtain a decoding event sequence; Performing real-time adjustment on the decoding parameters according to the decoding event sequence to obtain a streaming media decoding control instruction.
2. The network high-definition player streaming media decoding method according to claim 1, wherein The performing multi-threaded parallel acquisition and storage on the real-time temperature data, network bandwidth jitter data, and video bitrate fluctuation data of the decoding chip in the network high-definition player to obtain a streaming media dynamic feature dataset includes: Periodically sampling the core temperature and surface temperature of the decoding chip in the network high-definition player to obtain a real-time temperature data stream; Performing a sliding window detection on the network bandwidth data to obtain a network bandwidth jitter data stream, where the network bandwidth jitter data stream includes a bandwidth change rate and a packet loss rate; Performing dynamic monitoring on the bitrate of the input video stream to obtain a video bitrate fluctuation data stream; Collect the status of system resources, and obtain the system resource status data stream by monitoring the CPU occupancy rate, memory usage rate, and GPU usage rate; Align the timestamps of the real-time temperature data stream, network bandwidth jitter data stream, video bitrate fluctuation data stream, and system resource status data stream to obtain a multi-dimensional status data matrix; Store the multi-dimensional status data matrix in a circular buffer to obtain a streaming media dynamic feature data set.
3. The method for streaming media decoding of the network high-definition player according to claim 1, characterized in that, The construction of a multi-level streaming media cache system based on the inter-frame correlation prediction model includes: Perform hierarchical mapping on the inter-frame correlation, decoding resource requirements, and cache allocation strategy output by the inter-frame correlation prediction model to obtain multi-level cache configuration parameters, where the multi-level cache configuration parameters include the cache size of the network layer, the cache ratio of frame types in the decoding layer, and the buffer capacity of the display layer; Allocate cache space for the network layer based on the multi-level cache configuration parameters to obtain a network layer buffer structure, and the size of the network layer buffer structure is dynamically adjusted between 2 times and 4 times the video bitrate; Perform cache partitioning on the decoding layer based on the multi-level cache configuration parameters to obtain a decoding layer cache structure, where the decoding layer cache structure includes an I-frame buffer with a capacity of a first target value, a P-frame buffer with a capacity of a second target value, and a B-frame buffer with a capacity of a third target value; Perform cache configuration on the display layer based on the multi-level cache configuration parameters to obtain a display layer cache structure, where the display layer cache structure includes two buffers with a size 1.5 times the pixel data of the current display resolution; Configure the data transmission channels for the network layer buffer structure, the decoding layer cache structure, and the display layer cache structure to obtain an inter-layer data transmission strategy; Combine the network layer buffer structure, the decoding layer cache structure, the display layer cache structure, and the inter-layer data transmission strategy to obtain a multi-level streaming media cache system.
4. The network high-definition player streaming media decoding method according to claim 3, characterized in that, The performing cache partitioning on the decoding layer based on the multi-level cache configuration parameters to obtain a decoding layer cache structure includes: Analyze the cache ratio of frame types in the decoding layer and the encoding parameters in the multi-level cache configuration parameters to obtain frame grouping weight values; Calculate the cache capacity according to the frame grouping weight values and the total system memory capacity to obtain a frame type cache allocation scheme, and perform partition management on the system memory based on the frame type cache allocation scheme to obtain a cache area physical address table; Establish a multi-level page table structure for the cache area physical address table to obtain a cache access control table; Construct a cache scheduling scheme based on the cache access control table, and create a decoding layer cache structure based on the cache area physical address table, the cache access control table, and the cache scheduling scheme. The decoding layer cache structure includes an I-frame buffer with a capacity of a first target value, a P-frame buffer with a capacity of a second target value, and a B-frame buffer with a capacity of a third target value.
5. The method for streaming media decoding of the network high-definition player according to claim 1, characterized in that The real-time monitoring of the cache status, network status, and chip status in the multi-level streaming media cache system to obtain a decoding event sequence includes: Perform distributed sampling on the cache layer status in the multi-level streaming media cache system to obtain multi-layer cache status data; Perform multi-dimensional threshold analysis on the multi-layer cache status data, set hierarchical trigger thresholds according to video coding characteristics and hardware resource limitations, and obtain event trigger conditions; Construct an adaptive priority scheduling mechanism based on the event trigger conditions, perform dynamic priority assignment according to the degree of influence of events on decoding quality, and obtain a multi-level event priority table; Rearrange the multi-level event priority table according to time sequence, and sort according to the timestamp information and priority weight of the events to obtain a weighted event sequence; Input the weighted event sequence into the policy mapping network, match response strategies for various events according to historical decoding control experience, and obtain an event processing strategy group; Perform spatio-temporal correlation analysis based on the weighted event sequence and the event processing strategy group, and obtain a decoding event sequence by considering the temporal dependence relationship and spatial coupling effect of the events; 6. The network high-definition player streaming media decoding method according to claim 1, characterized in that, The decoding parameters are adjusted in real time according to the decoding event sequence to obtain a streaming media decoding control instruction, including: Classify and analyze the events in the decoding event sequence, extract control parameters according to the characteristics of temperature events, performance events, network events, cache events, and bitrate events, and obtain a decoding regulation target matrix; Perform hardware resource control based on the decoding regulation target matrix, dynamically configure the frequency and voltage of the decoding chip, the number of decoding threads, and the cache allocation ratio, and obtain a hardware control parameter sequence; Match the decoding strategy for the H.265 / HEVC encoded stream according to the hardware control parameter sequence, and obtain a decoding execution plan based on the decoding priorities and frame skipping strategies of I-frames, P-frames, and B-frames; Perform parallelism analysis on the decoding execution plan, calculate the dependency relationship and execution time sequence of each decoding task, and obtain a task scheduling sequence, where the task scheduling sequence includes the execution order and time window of the decoding tasks; Construct a resource allocation model based on the task scheduling sequence, calculate the dynamic balance point of CPU occupancy rate, memory usage rate, and GPU usage rate, and obtain a resource allocation plan; Integrate and map the hardware control parameter sequence, the decoding execution plan, and the resource allocation plan to obtain a streaming media decoding control instruction, where the streaming media decoding control instruction includes a temperature control instruction, a decoding control instruction, and a resource control instruction.
7. A streaming media decoding device for a network high-definition player, characterized in that, For executing the streaming media decoding method of the network high-definition player as described in any one of claims 1-6, the streaming media decoding device of the network high-definition player includes: A parallel acquisition module for multi-threaded parallel acquisition and storage of real-time temperature data, network bandwidth jitter data, and video bitrate fluctuation data of the decoding chip in the network high-definition player to obtain a streaming media dynamic feature data set; A classification training module for classifying and training the data in the streaming media dynamic feature data set according to the inter-frame prediction difficulty to obtain an inter-frame correlation prediction model; A construction module for constructing a multi-level streaming media cache system based on the inter-frame correlation prediction model; A real-time monitoring module for real-time monitoring of the cache status, network status, and chip status in the multi-level streaming media cache system to obtain a decoding event sequence; A real-time adjustment module, configured to perform real-time adjustment on decoding parameters according to the decoded event sequence, so as to obtain a streaming media decoding control instruction.
8. A computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, the processor is caused to execute the streaming media decoding method of the network high-definition player according to any one of claims 1 to 6.
Citation Information
Patent Citations
Standards-compliant model-based video encoding and decoding
US20130230099A1
Intra-coded video frame caching for video telephony sessions
US20180063548A1