Video quality monitoring system
By using multiple machine learning models to generate presentation quality indicators in the edge device of the video network, the logic and resource problems of evaluating the cross-network presentation quality of video content in the prior art are solved, and effective monitoring and rating of the presentation quality of video content in the video network is realized.
Patent Information
- Application Number
- CN202411581446.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-10
- Filing Date
- 2024-11-07
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, when evaluating the cross-network presentation quality of video content, there are logic and resource problems, and it is difficult to effectively monitor and rate video quality in video networks.
By receiving network packets containing content encapsulated in multiple layers in an edge device, presentation quality metrics for extracted content are generated in a hierarchical order using multiple machine learning models and providing these metrics to the server over the network.
It realizes effective monitoring and rating of the quality of video content presentation in the video network, improves the consistency and quality of user experience, and reduces network bandwidth consumption.
Smart Images

Figure CN119996644A_ABST
Abstract
Description
Technical Field
[0001] The present description generally relates to a video network, including, for example, video quality analysis across the network. Background Art
[0002] Video networks can stream content to millions of viewers spread across diverse geographic locations. Maintaining the quality of streamed content across the network is crucial to providing a consistent and positive user experience. Traditionally, the image quality of video content has been assessed by comparing reference images from the video content with corresponding images from the video content after it has traversed the network to the user's location. While this approach may be feasible for limited environments where issues have been identified and flagged, comparing delivered video content to reference video content presents logistical and resource challenges. Summary of the Invention
[0003] In one aspect, the present disclosure relates to a device comprising: a computer-readable storage medium storing one or more instruction sequences; and a processing circuit system configured to execute the one or more instruction sequences so as to: receive a plurality of network packets containing content encapsulated in a plurality of layers; process the received plurality of network packets to extract the content for presentation; generate a predicted presentation quality indicator for the extracted content using a plurality of machine learning models in a hierarchical order, wherein data generated during the processing of the received plurality of network packets is used as input to the plurality of machine learning models; and provide the predicted presentation quality indicator for the extracted content to a server via a network, wherein the data generated during the processing of the received plurality of network packets is correlated across the plurality of layers to generate the predicted presentation quality indicator.
[0004] In another aspect, the present disclosure relates to a method comprising: receiving a plurality of network packets containing content encapsulated in a plurality of layers; processing the received plurality of network packets to extract the content for presentation; generating a predicted presentation quality indicator for the extracted content using a plurality of machine learning models in a hierarchical order, wherein data generated during the processing of the received plurality of network packets is used as input to the plurality of machine learning models; and providing the predicted presentation quality indicator for the extracted content to a server via a network, wherein the machine learning models in the plurality of machine learning models are associated with corresponding layers in the plurality of layers, and wherein the data generated during the processing of the received plurality of network packets is correlated across the plurality of layers to generate the predicted presentation quality indicator.
[0005] In another aspect, the present disclosure relates to a system comprising: a server; and a plurality of edge devices configured to communicate with the server via a network, wherein each edge device of the plurality of edge devices comprises: a computer-readable storage medium storing one or more instruction sequences; and a processing circuit system configured to execute the one or more instruction sequences to: receive a plurality of network packets containing content encapsulated in a plurality of layers; process the received plurality of network packets to extract the content for presentation; generate predicted presentation quality indicators for the extracted content using a plurality of machine learning models in a hierarchical order, wherein data generated during processing of the received plurality of network packets is used as input to the plurality of machine learning models; and provide the predicted presentation quality indicators for the extracted content to the server via the network, wherein the data generated during processing of the received plurality of network packets is correlated across the plurality of layers to generate the predicted presentation quality indicators, wherein the server is configured to correlate the predicted presentation quality indicators provided by the plurality of edge devices in order to evaluate the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Certain features of the technology are set forth in the appended claims. However, for illustrative purposes, several aspects of the technology are set forth in the following figures.
[0007] Figure 1 Illustrated is an example of a network environment in which aspects of the present technology may be implemented.
[0008] Figure 2 is a block diagram illustrating the general operation of a video quality monitoring system in accordance with aspects of the present technique.
[0009] Figure 3 is a block diagram illustrating components of customer premises equipment in accordance with aspects of the present technology.
[0010] Figure 4 is a block diagram illustrating an exemplary integration of multiple models.
[0011] Figure 5 is a block diagram illustrating aspects of a video presentation quality monitor in accordance with aspects of the present technology.
[0012] Figure 6 is a block diagram illustrating the operation of a video presentation quality monitoring system in accordance with aspects of the present technology.
[0013] Figures 7 to 10 is a block diagram illustrating the operation of a video presentation quality monitoring system through a sequence of inference stages in accordance with aspects of the present technique.
[0014] Figure 11is a flow chart depicting an exemplary process for monitoring video content quality in accordance with aspects of the present technology.
[0015] Figure 12 is a block diagram illustrating an electronic system in which aspects of the present technology may be implemented. DETAILED DESCRIPTION
[0016] The detailed description set forth below is intended to illustrate various configurations of the present technology and is not intended to represent the only configurations in which the present technology may be practiced. The accompanying drawings are incorporated herein and form a part of the detailed description. For the purpose of providing a thorough understanding of the present technology, the detailed description includes specific details. However, the present technology is not limited to the specific details set forth herein and may be practiced without one or more of these specific details. In some instances, structures and components are shown in block diagram form to avoid obscuring the concepts of the present technology.
[0017] Human perception of whether image quality is good or bad is relatively straightforward. However, attempting to replicate human perception using machine analysis of image data is less straightforward. Typically, machine analysis of image data involves comparing an instance of an image to a reference image and computing the difference between the two images using a reference metric such as mean squared error (MSE). The ability to mimic human perception of images to some extent without requiring comparison to a reference image is extremely attractive, especially to operators of large-scale video networks.
[0018] This technology provides a solution that facilitates monitoring and rating the presentation of content (e.g., video content) across a network by utilizing machine learning models to predict the quality of the content presented to users after traversing the network. According to various aspects of this technology, an end-to-end video analytics system can be provided that integrates intelligent detection capabilities within edge devices of a network. Edge devices include customer premises devices such as set-top boxes, smart TVs, modems, routers, etc. This technology does not limit the application of machine learning models to the analysis of low-level data, such as image artifacts in decoded image data. Instead, the solution provided by this technology extends to domains and protocols involved in content delivery and presentation. For example, machine learning models can be used to analyze IP (Internet Protocol), MPEG (Moving Picture Experts Group) transports, video and audio bitstreams, pixels, symbols, and metadata to track key quality and performance indicators. Data collected at the edge can be analyzed at the edge and / or sent to servers, such as cloud servers, for processing. The ability to analyze data at the edge and limit communication with the central office to results and / or critical portions of the data frees up network bandwidth for content delivery rather than analysis traffic. Examples and descriptions of this technology are provided in detail below.
[0019] Figure 1The diagram illustrates an example of a network environment 100 in which aspects of the present technology may be implemented. However, not all depicted components are required, and one or more implementations may include additional components not shown in the diagram. The arrangement and types of components may be varied without departing from the spirit or scope of the claims as set forth herein. Unless expressly stated otherwise, depicted or described connections and couplings (including electrical and communication connections and couplings) between components are not limited to direct connections or direct couplings, but may be implemented with one or more intermediate components.
[0020] Exemplary network environment 100 includes a headend video quality monitoring (HE-VQM) server 110, a content server 120, customer premises equipment (CPE) 130-160, and a network 170. CPE 130-160 include, but are not limited to, a set-top box (STB) 130, a smart television (TV) 140, a router 150, and a modem 160. HE-VQM server 110 and content server 120 may be configured to communicate with CPE 130-160 via network 170. Network 170 may include one or more public communication networks (e.g., the Internet, a cable distribution network, a cellular data network, etc.) and / or one or more private communication networks (e.g., a private local area network (LAN), a leased line, etc.). Network 170 may also include, but is not limited to, any one or more of the following network topologies, including, but not limited to, a bus network, a star network, a ring network, a mesh network, a star-bus network, a tree or hierarchical network, and the like. In one or more implementations, network 170 may include transmission lines, such as coaxial transmission lines, fiber optic transmission lines, or generally any transmission lines, that communicatively couple HE-VQM server 110 and content server 120 to CPEs 130 to 160. HE-VQM server 110 and content server 120 may communicate with CPEs 130 to 160 via the same network connection or via different respective network connections.
[0021] The HE-VQM server 110 and the content server 120 may be co-located at a video central office (e.g., a facility containing equipment configured to receive and process content from various sources for distribution to customer premises) of a cable operator or some other type of content distributor, or may be located at different respective locations. The HE-VQM server 110 and the content server 120 may be implemented together on a common server, or may be implemented in separate respective servers. In addition, the HE-VQM 110 and / or the content server 120 may be implemented using a single computing device, or may be implemented using multiple computing devices configured to work together to perform their respective functions (e.g., a cloud computing system, a distributed system, etc.).
[0022] In short, content server 120 may be configured to communicate with CPEs 130-160 to deliver content, such as video content, audio content, data, etc., as a network packet stream via network 170. As discussed in more detail below, CPEs 130-160 may include a CPE video quality monitor (VQM) configured to analyze or evaluate the content delivered by content server 120 and generate presentation quality metrics that estimate the quality of presentation of that content to consumers of the content. Reports containing the presentation quality metrics may be provided to HE-VQM 110 via network 170 for further analysis, either individually or in conjunction with reports received from other CPEs. Figure 2 is a block diagram illustrating an example of the general operation of a video quality monitoring system according to aspects of the present technology. Figure 2 , a video server 210 of a video central office 200 provides video in a network packet stream to a CPE 220. In addition to presenting video to users of the CPE 220, the CPE 220 also includes a CPE-VQM 230 that evaluates the video content delivered from the video server 210 and generates a video quality report 240 containing data estimating the quality of the video data presented to the user. The CPE-VQM 230 can provide the video quality report 240 to the HE-VQM 250 for system-wide evaluation.
[0023] Figure 3 is a block diagram illustrating components of a CPE, such as a Figure 1 The CPE 130 depicted in Figure 2 220 is depicted in FIG. However, not all of the depicted components are required, and one or more implementations may include additional components not shown in the figures. Changes may be made to the arrangement and types of components without departing from the spirit or scope of the claims as set forth herein. Unless expressly stated otherwise, depicted or described connections and couplings (including electrical and communicative connections and couplings) between components are not limited to direct connections or couplings, but may be implemented with one or more intermediate components.
[0024] exist Figure 3In the example depicted in FIG, CPE 300 includes a system on a chip (SOC) 305, external memory 310, and an interface 315. Interface 315 may include suitable circuitry, logic, and / or code to enable communication of network packets with CPE 300. The present technology is not limited to any particular network protocol and / or configuration. CPE 300 may include a single interface through which all communication of network packets is performed. Alternatively, CPE 300 may include multiple interfaces 315 of the same type or different corresponding types to facilitate communication with different entities (e.g., different servers and / or other network devices).
[0025] According to aspects of the present technology, the SOC 305 may include a central processing unit (CPU) 320, a security processor 325, a delivery engine 330, a stream processor 335, a video / audio codec 340, a machine learning core 345, on-chip memory 350, and registers 355. The SOC 305 and its components, individually or collectively as a group of two or more components, represent processing circuitry configured to perform the operations described herein. The SOC 305 or one or more of its components may be implemented in hardware using circuitry such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device (PLD), a controller, a state machine, gating logic, discrete hardware components, or any other suitable device. One or more components of the SOC 305 (e.g., the stream processor 335) may include or may be implemented using software / firmware (e.g., instructions, code, subroutines, etc.) that is executed by the processing circuitry (e.g., the CPU 320) to provide the operations described herein.
[0026] The CPU 320 may comprise suitable logic, circuitry, and / or code to enable processing of data and / or control of the operation of the CPE 300. In this regard, the CPU 320 may be configured to provide control signals to various other components of the CPE 300. The CPU 320 may also control the transfer of data between components within the CPE 300 and between the CPE 300 and other devices or systems external to the CPE 300.
[0027] The security processor 325 may include suitable logic, circuitry, and / or code to enable management of a secure content pipeline for protected content, such as premium video content. Management of the secure content pipeline may include encryption / decryption of protected content. The security processor 325 may work with other components of the SOC 305 to securely handle protected content. Figure 3Although not depicted, SOC 305 may include a secure CPU for performing operations involving protected content, as well as secure registers and a secure portion of on-chip memory used during processing of the protected content. Other components of SOC 305 may recognize the protected content being processed and use the secure registers and secure on-chip memory instead of publicly accessible storage locations within and outside of SOC 305.
[0028] The transport engine 330 may include suitable logic, circuitry, and / or code for managing and monitoring the communication of network packets sent and received by the CPE 300. Network packets may be sent and received using a number of different transport protocols, including, but not limited to, the Moving Picture Experts Group (MPEG) transport protocol and / or the Internet Protocol (IP) transport protocol. While managing the transport of network packets, the transport engine 330 may be configured to extract / capture and make available various delivery metrics that may be used in various aspects of the video quality monitoring system described herein. For example, MPEG content delivery loss may be detected by audio and video packet identifier (PID) counter discontinuities. Other MPEG delivery metrics may include video buffer errors (overflow, underflow), program clock reference (PCR) value out of range, PCR discontinuities, etc. An exemplary metric is a payload integrity failure count that tracks whether a packet payload was able to be correctly parsed. An integrity check failure indicates that the packet payload could not be correctly parsed. Consequently, the payload data will subsequently not be decoded by the decoder. Consequently, a black screen and service interruption may be expected. Payload integrity may fail for many reasons, such as data corruption during delivery or invalid security keys used to decrypt protected video content.
[0029] Similarly, for the purpose of monitoring presentation quality, various IP transport metrics can be extracted / captured and made available. For example, such metrics may include network jitter measured by inter-packet arrival times in one or both of the time and frequency domains, transmission patterns measured by the number of packets per unit time and the length of the packets, and flow characteristics measured by duration, size, and / or byte value distribution.
[0030] The video / audio codec 340 may include appropriate logic, circuitry, and / or code to enable decoding of content in the received stream. The present technology is not limited to any particular type of encoding / decoding standard and may be used with a variety of coding standards. For example, video data may be encoded / decoded using H.264, H.265, H.266, VP9, AV1, etc., and audio data may be encoded / decoded using AC3, AAC, He-AAC, MP3, WAV, etc. The present technology may be configured to monitor data generated during the decoding of video and audio data to serve as an indicator of possible problems with decoding and presenting the video / audio content. For example, video decodability may be tracked by counting the number of pictures that were not decoded. Note that some decodable pictures may be decoded incorrectly, while the remaining pictures are decoded without error. Error-decoded pictures may be tracked by frame type (e.g., I-frames, P-frames, and B-frames). The present technology may also track decoder performance. For example, performance metrics such as current frame decode time, average frame decode time, and maximum frame decode time can be tracked to identify decoder problems and delivery issues.
[0031] The stream processor 335 may include suitable logic, circuitry, and / or code to coordinate streaming of operations performed by components of the SOC 305 and to collect / extract data generated during processing of received content for use in various aspects of the present technology. For example, the stream processor 335 may be configured to format and / or store data generated during processing of received content in internal memory and / or external memory locations for use in monitoring presentation quality of the content.
[0032] The machine learning core 345 may include suitable logic, circuitry, and / or code to enable operation of machine learning models, such as neural networks, used during presentation quality monitoring. The machine learning core 345 may include a framework for implementing one or more models in accordance with various aspects of the present technology. In the case of a neural network model, the model may be the result of a machine learning architecture trained using one or more data sets and defined by a set of parameters, which may specify node operations and edge weights. The machine learning core 345 may also include a framework for other types of mathematical algorithms used as models. Such models may be used to process various types of data associated with the processing and presentation of content. For example, during processing by the delivery engine 330 and / or stream processor 335, high-level or semantic data, such as program and channel information associated with the content, may be extracted and used as input to the model. Additionally, more complex, lower-level signal data, such as picture pixels and / or audio symbols, may be processed using a trained neural network model. While the model implementation has been described using the machine learning core 345, the present technology may also implement one or more models using a CPU 320 executing one or more instruction sequences, rather than utilizing the machine learning core 345.
[0033] The on-chip memory 350 may include suitable logic, circuitry, and / or code that enables storage and access of various types of data by the components of the SOC 305 as described herein. The on-chip memory 350 may include, for example, random access memory (RAM), read-only memory (ROM), flash memory, etc. The on-chip memory 350 may include multiple types of memory, such as volatile memory that provides temporary workspace for the components of the SOC 305 and non-volatile memory that provides storage space that preserves data across power cycles. As described above, the on-chip memory 350 may include a portion of secure memory for use when protected content is being processed.
[0034] Registers 355 may include suitable logic and circuitry to provide storage for data that may be written to and read by components of the SOC 305. Registers 355 may provide faster access to smaller amounts of data than provided by on-chip memory 350. In addition, registers 355 may include secure registers for use when protected content is being processed.
[0035] The external memory 310 may include suitable logic, circuitry, and / or code that enables storage of various types of information, such as received data, generated data, code, and / or configuration information. The external memory 310 may include, for example, random access memory (RAM), read-only memory (ROM), flash memory, magnetic storage devices, optical storage devices, and the like. The external memory 310 may include various types of memory, such as volatile memory and non-volatile memory, and, similar to the on-chip memory 350, may include a portion of secure memory for use by the CPE 300 when processing and presenting protected content. Figure 3 As depicted, external memory 310 contains an operating system 360 and a VQM application 365 in accordance with aspects of the present technology.
[0036] According to aspects of the present technology, operating system 360 comprises a computer program having one or more sequences of instructions or codes and associated data and settings. Upon execution of the instructions or codes, for example by CPU 320, one or more processes are initiated to manage the resources and operations of CPE 300 to implement the processes described herein. In addition to operating system 360, external memory 310 may also include a trusted operating system (not shown). The trusted operating system may be executed by a secure CPU in SOC 305 to manage access to secure memory and register locations and to manage resources associated with executing trusted applications that may be used to process and present protected content.
[0037] According to aspects of the present technology, VQM app 365 comprises one or more computer programs having one or more sequences of instructions or code and associated data and settings. Upon execution of the instructions or code, one or more processes may be initiated to perform the quality monitoring operations described herein. VQM app 365 may be configured to reference and utilize data generated and / or extracted during content processing, content and metadata describing the content, and data used to select and configure a machine learning model for generating a video quality report, which may include one or more predicted presentation quality metrics including picture quality scores, processing statistics, processing errors, and the like. VQM app 365 may be configured to execute the instructions or code on one or more processors, including but not limited to CPU 320, machine learning core 345, and video / audio codec 340. The processor used by VQM app 365 may depend on factors such as execution speed, power consumption, memory resource constraints, and processor availability in SOC 305.
[0038] According to aspects of the present technology, the VQM app 365 can integrate multiple models to generate a predicted presentation quality metric. Two or more of the models used by the VQM app 365 can be executed in a hierarchical order. Additionally, the VQM app 365 can be configured to run the models in parallel and / or sequentially. According to aspects of the present technology, Figure 4 is a block diagram illustrating an exemplary integration of multiple models. Figure 4 As depicted in , the models can be connected to work in parallel and / or sequentially, with respective output results combined with prescribed weights to generate an output index (e.g., a predicted presentation quality indicator). Figure 4 The number and / or arrangement of the models depicted in the drawings are not limited, and other arrangements and numbers of models may be used for implementation.
[0039] refer to Figure 4 , Models 1, 1.1, and 1.2 are connected sequentially, with the selection of Model 1.1 or Model 1.2 determined by the output of Model 1. Similar to Model 1, Model 3 is connected sequentially to Models 3.1 and 3.2. However, both Models 3.1 and 3.2 receive the weighted output of Model 3, which is processed in parallel. Similarly, Model 3.1 is connected sequentially to Models 3.1.1 and 3.1.2, to process the weighted output of Model 3.1 in parallel. On the other hand, Model 3.2 is similar to Model 1, with Model 3.2 being connected sequentially, with the selection of Model 3.2.1 or 3.2.2 determined based on the output of Model 3.2. In this example, the weighted outputs of Models 1.1 or 1.2, Model 2, Model 3.1.1, Model 3.1.2, and Models 3.2.1 or 3.2.2 are combined to generate an output index comprising a predicted presentation quality indicator.
[0040] refer to Figure 4 , Model 1 can be configured to process video metadata for program and channel guide information. Model 1 can predict whether the type of content being processed is artificial content (e.g., content generated using a computer) or natural scene content (e.g., images / content captured and generated using a camera). Based on the results of Model 1, if the content is predicted to be artificial content, the system selects Model 1.1, and if the content is predicted to be natural scene content, the system selects Model 1.2. In this regard, artificial content tends to have signal level differences compared to natural scene content in terms of color, texture, shape, etc. In this arrangement, Model 1 processes high-level semantic data and can be a decision tree algorithm, while Models 1.1 and 1.2 can process signal level data, such as picture pixels, and can be implemented using neural networks to achieve content-aware video quality prediction.
[0041] When video transport packets are lost, the consequences may not be visible and may not affect the user experience of the rendered video content. For example, the lost packets may be null packets that have little or no effect on the rendering of the video content. The lost packets may be program service information (PSI) related packets that are repeatedly transmitted in the transport stream. The loss of these packets has little or no effect on the rendering of the video content. Models can be configured and trained to effectively detect the visibility of packet loss. According to various aspects of the present technology, Model 3.2 can be an MPEG transport engine model that monitors packet continuity counters (CCs) for packet loss and other attributes. If Model 3.2 detects packet loss and the packet is a video packet, then Video Model 3.2.1 can be selected for further processing. Model 3.2.1 can be configured to detect whether any frame errors occur in video decoding. Model 3.2.1 can be integrated with a video neural network model (Model 2) trained to detect picture pixel artifacts such as blockiness or jitter noise. If the packet is an audio packet, then Model 3.2.2 can be selected. Model 3.2.2 can be configured to detect audio frame errors in decoding. In this example, Model 3.2 can be a transport model, while Model 3.2.1 and Model 3.2.2 can be a video model and an audio model, respectively, and Model 2 can be a neural network model.
[0042] According to aspects of the present technology, Figure 5 is a block diagram illustrating aspects of a video presentation quality monitor. Figure 5 In the example shown in FIG, a hierarchical application of a group model is illustrated for the case of packet loss. Figure 5 , video content 500 may be received by a CPE. A delivery engine model 510 may identify and output a packet continuity counter failure indicating one or more packet losses observed by the delivery engine while processing network packets. A video codec model 520 may identify and output an indicator of video frame decoding errors that occurred while the video codec was processing a bitstream containing video content. This indicator of video frame decoding errors confirms that one or more of the lost packets are video packets. A neural network picture model 530 may detect a degradation in picture quality based on identifying blocky or noisy frames in pixel data decoded by the video codec, which further confirms that one or more of the lost packets are video packets. Figure 5 The video rendering quality monitor represented in generates a picture quality score 540 based on the outputs of the three models. The picture quality score 540 may be provided to the HE-VQM for further processing.
[0043] Figure 5The example depicted in illustrates content encapsulation in multiple layers across multiple domains. When evaluated by the delivery engine model 510, the video content 500 is encapsulated in network packets. When evaluated by the video codec model 520, the video content has been removed from the network packet layer and is presented in one or more bitstreams. When evaluated by the NN picture model 530, the video content in the bitstream layer has been decoded into pixel values. In addition to these layers, when evaluated by the delivery engine model 510 and the video codec model 520, the video content may exist in a compressed state, and when evaluated by the NN picture model 530, the video content may exist in a decompressed state. As illustrated in this brief example, the present technology is effective in monitoring presentation quality by detecting problems at different layers of encapsulation and at different stages of processing the received video content for presentation. Using information such as timestamps associated with the video content at different encapsulation layers or other information that can be used to identify the content portion encapsulated in each of the layers, data generated during different processing stages can be correlated across different layers. In this way, problems identified by the three models can be correlated to confirm that the problems occur with respect to the same portion of the content contained in the encapsulation layer.
[0044] According to aspects of the present technology, Figure 6 is a block diagram illustrating the operation of a video quality monitor. Figure 6 As depicted in FIG, the operation of CPE VQM 600 includes a sequence of inference stages (e.g., inference stage 1, inference stage 2, inference stage 3, ..., inference stage N), wherein when each inference stage is executed, the video content is in a specific format based on the encapsulation layer. Figures 7 to 10 An example of four inference stages is depicted in the block diagram illustrated in FIG. For explanation purposes, the blocks illustrating the inference stages may be described herein as occurring serially or linearly. However, two or more blocks of the illustrated inference stages may be executed in parallel. Additionally, Figures 7 to 10 The blocks depicted in the flowchart may be executed in an order different from that shown, and the inference phase may not execute the illustrated blocks and / or may include one or more additional blocks.
[0045] Reference below Figures 7 to 10In the examples described, video content is encapsulated in multiple layers that are removed at different stages of processing the video content. For example, the first layer of encapsulation could be IP encapsulated data, followed by a second layer of encapsulation of an encrypted MPEG2 transport, a third layer of encapsulation of an H.264 compressed bitstream, and finally, H.264 decoded pixels. The present technology is not limited to these standards / protocols and can be implemented for systems utilizing other standards / protocols. Although not detailed in the examples below, once the content's protection is stripped during processing, the protected video content should be stored in secure memory and register locations. For example, the IP encapsulated data and the encrypted MPEG2 transport data do not need to be stored in secure memory locations. However, the decrypted data found in the H.264 compressed bitstream and H.264 decoded pixels should be stored in secure memory locations.
[0046] According to various aspects of the present technology, a stream processor within a CPE's SOC can operate as a data aggregator in a VQM system. For example, the stream processor can internally collect data from delivery engines, video / audio codecs, and other sources and store the collected data in external memory for access by the VQM app and machine learning models. Alternatively, the CPU can replace the stream processor in operating as the data aggregator.
[0047] like Figure 7 As depicted in the example of , for inference phase 1, the encapsulation of the received video content is IP encapsulated data, which may contain encrypted MPEG-2 transport streams with H.264 and AC3 encoded video and audio compression formats. Again, the present technology is not limited to these or any particular standards / protocols and can be implemented for systems that use other standards / protocols to deliver and present video content and audio content. The measurements performed by transport model 1 are feature extraction, including but not limited to IP packet inter-arrival times (IPTs), assisted by stream processor 710, which outputs a series of time-domain IPT data to external memory 720. During the measurement, stream processor 710 may store intermediate measurements in internal on-chip memory for faster processing. When the measurement is complete, stream processor 710 stores the final IPT time series data to external memory 720.
[0048] The CPU 730 may transform the IPT time series data into the frequency domain, specifically, into the frequency of packet arrival periodicity. The CPU 730 may transfer the IPT time series data to internal memory for faster processing. After completing the transformation, the CPU 730 may store the transformed IPT data in the external memory 720. Transport Model 1, configured by the CPU 730 with Transport Model 1 data, may be configured to detect network jitter by correlating the transformed IPT data over a period of time. The index output from Transport Model 1 may be a binary classification indicating whether jitter is detected.
[0049] like Figure 8 As depicted in the example of FIGURE 1, transport engine 840 can be configured to process incoming encrypted MPEG2 transport data to perform measurements of various transport characteristics, including, but not limited to, packet continuity as reflected in continuity counter (CC) errors. A packet continuity counter (CC) circuit in transport engine 840 processes incoming data in its own internal (secure) memory and detects whether a counter discontinuity has occurred. If a counter discontinuity is detected, transport engine 840 sets associated (secure) on-chip registers 850. Stream processor 810 reads the registers set by transport engine 840 and stores the CC and other transport information in on-chip secure memory, along with associated time-domain information (e.g., timestamps). After any formatting or additional processing, stream processor 810 prepares the final transport engine data in a defined structure and stores the data in external secure memory 820.
[0050] The MPEG2-TS transport model can be executed on the CPU 830. The model can be configured to read data stored by the stream processor 810 from the external memory 820 and determine additional MPEG2 transport metrics. Intermediate model data, inputs, and outputs can be stored in on-chip secure memory on the CPU 830. The model process detects and classifies whether a detected packet loss is significant and outputs an index indicating whether the packet loss is significant to the external memory 820. In other examples, the model can be used with a codec model to predict the impact of packet loss, such as loss visibility.
[0051] like Figure 9As depicted in the example of , video codec 940 can be configured to process H.264 compressed bitstream video input and perform measurements such as decoding speed, latency, error statistics, etc. during processing. Based on the measurements, video codec 940 detects whether frame decoding errors have occurred and, if necessary, uses on-chip secure memory for data processing. Video codec 940 can set associated secure registers indicating detected frame errors and other video data. Stream processor 910 can read data related to detected frame decoding errors and store the data and other information in its on-chip secure memory for processing. After processing, stream processor 910 can prepare final video codec data in a data structure, including frame errors, timestamps, etc., and store the data structure in external memory 920.
[0052] The video codec model 960 can be executed on the CPU 830. The video codec model 960 can read data stored in external memory by the stream processor during the current inference phase and the previous inference phase and capture video metrics (e.g., H.264 video metrics). Intermediate model data, inputs, and outputs can be stored in the CPU 930's on-chip secure memory. The model process detects and classifies whether packet loss detected or measured by the delivery engine is incorrectly confirmed by decoding and may be visible. In this case, if packet loss is visible, the video codec module model 960 predicts that the output video codec model 960 is exponential.
[0053] like Figure 10 As depicted in the example of , the inference stage 4 is configured to process H.264 decoded pixels. For example, the machine learning core 1070 can be configured to process the H.264 decoded pixels to perform measurements, such as feature extraction of video attributes (e.g., noise, artificial edges, etc.). The machine learning core 1070 can be configured to use internal memory to store data related to feature extraction and model data for configuring a neural network (NN) video quality model 1080. The NN video quality model 1080 can be configured to predict a video quality score based on the extracted video attributes and store the predicted video quality score in the external memory 1020.
[0054] According to aspects of the present technique, the predicted video score may be an output index of the NN video quality model 1080. Alternatively, or in addition to the predicted video score, the video quality model 1090 may be executed on the CPU 1030 to perform a joint inference using as input data measured and stored during a previous inference stage, wherein the packet loss data from inference stage 2, the decoding error data from inference stage 3, and the predicted video score generated in inference stage 4 are used to jointly verify or confirm that the detected packet loss is visible to a user viewing the presentation of the content.
[0055] According to aspects of the present technology, the HE-VQM may include one or more of the same machine learning models as the machine learning models incorporated into the CPE-VQM of the CPE device, thereby allowing the HE-VQM to replicate the video content analysis performed by the CPE device. The HE-VQM may more easily access the original or reference content delivered by the video server across the network. In some embodiments, the HE-VQM may perform some or all of the monitoring and analysis performed by the CPE-VQM on the original or reference content to generate an expected presentation quality score. The HE-VQM may provide the expected presentation quality score to the video server for inclusion in the metadata transmitted to the CPE device along with the video content. Using the results generated by the HE-VQM, the CPE-VQM may compare its generated results with the results provided from the HE-VQM to assess whether any problems found in the content presentation are inherent in the original or reference content. If the problem is inherent, the CPE-VQM may not notify the HE-VQM of its results.
[0056] According to aspects of the present technology, a CPE may be configured to verify or confirm problems predicted for content presented directly to a user of the CPE. For example, when a CPE-VQM identifies a possible visible problem with the presentation of content, the CPE may provide a prompt to the user viewing the content requesting confirmation of the possible visible problem. The prompt may be a visual prompt placed on the screen where the content is being viewed, and / or may be an audio prompt with text-to-speech technology played through the speakers of the device used to view the content. The viewing user may respond to the prompt to confirm or deny the existence of the possible visible problem in the presentation of the content. The response may be made using any of several mechanisms including a user interface, including but not limited to user voice confirmation with automatic speech recognition on the CPE or the device where the content is being presented, gestures, and the like. If the user denies the existence of a possible visible problem with the presentation of the content, the system may postpone or cancel sending the video quality report to the HE-VQM.
[0057] Figure 11 is a flow chart depicting an exemplary process for monitoring video content quality according to aspects of the present technology. For explanation purposes, the blocks of the illustrated process may be described herein as occurring serially or linearly. However, two or more blocks of the illustrated process may be executed in parallel. Additionally, Figure 11 The blocks depicted in the drawings may be executed in an order different from that shown, and a process may not perform one or more of the illustrated blocks and / or may include one or more additional blocks.
[0058] According to aspects of the present technology, process 1100 includes receiving a plurality of network packets containing content encapsulated in multiple layers (block 1110). For example, the network packets may be received by a CPE device, such as a set-top box. The network packets may be processed to extract the encapsulated content for presentation (block 1120). The content may include video data, audio data, etc. Data generated during the processing of the network packets may be used to generate a predicted presentation quality metric for the extracted content using a machine learning model in a hierarchical order (block 1130). The predicted presentation quality metric may be provided to a server, such as an HE-VQM, for further processing (block 1140).
[0059] Figure 12 The electronic system 1200 is conceptually illustrated and can be used to implement one or more embodiments of the present technology. However, not all of the depicted components are required, and one or more embodiments may include additional components not shown in the figures. The arrangement and types of components may be varied without departing from the spirit or scope of the claims as set forth herein. Unless expressly stated otherwise, depicted or described connections and couplings (including electrical and communication connections and couplings) between components are not limited to direct connections or couplings, but may be implemented with one or more intermediate components.
[0060] For example, the electronic system 1200 may be an HE-VQM or video server as described above. Such an electronic system 1200 includes various types of computer-readable media and interfaces for various other types of computer-readable media. The electronic system 1200 includes a bus 1208, one or more processing units 1212, a system memory 1204, a read-only memory (ROM) 1210, a permanent storage device 1202, an input device interface 1214, an output device interface 1206, and a network interface 1216, or subsets and variations thereof.
[0061] The bus 1208 collectively represents all system, peripheral, and chipset buses that communicatively connect the various internal devices of the electronic system 1200. In one or more embodiments, the bus 1208 communicatively connects the one or more processing units 1212 with the ROM 1210, the system memory 1204, and the permanent storage device 1202. The one or more processing units 1212 retrieve instructions to execute and data to process from these various memory units in order to perform the processes of the present disclosure. In different embodiments, the one or more processing units 1212 can be a single processor or a multi-core processor.
[0062] ROM 1210 stores static data and instructions required by one or more processing units 1212 and other modules of the electronic system. Permanent storage 1202, on the other hand, is a read-write memory device. Permanent storage 1202 is a non-volatile memory unit that stores instructions and data even when the electronic system 1200 is turned off. One or more embodiments of the present disclosure use a mass storage device (e.g., a solid-state drive, or a magnetic or optical disk and its corresponding magnetic disk drive) as permanent storage 1202.
[0063] Other embodiments use removable storage devices (e.g., flash memory drives, optical disks and their corresponding disk drives, external magnetic hard disk drives, etc.) as permanent storage 1202. Like permanent storage 1202, system memory 1204 is a read-write memory device. However, unlike permanent storage 1202, system memory 1204 is a volatile read-write memory, such as random access memory. System memory 1204 stores any instructions and data required by one or more processing units 1212 during execution. In one or more embodiments, the processes of the present disclosure are stored in system memory 1204, permanent storage 1202, and / or ROM 1210. One or more processing units 1212 retrieve instructions to execute and data to process from these various memory units in order to perform the processes of one or more embodiments.
[0064] Bus 1208 is also connected to input device interface 1214 and output device interface 1206. Input device interface 1214 enables a user to communicate information and select commands to the electronic system. Input devices used with input device interface 1214 include, for example, an alphanumeric keyboard and a pointing device (also known as a "cursor control device"). Output device interface 1206 is capable of, for example, displaying images generated by electronic system 1200. Output devices used with output device interface 1206 include, for example, a printer and a display device, such as a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a flexible display, a flat panel display, a solid-state display, a projector, or any other device for outputting information. One or more embodiments include a device that functions as both an input device and an output device, such as a touch screen. In these embodiments, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound, voice, or tactile input.
[0065] Finally, if Figure 12As shown in FIG1 , bus 1208 also couples electronic system 1200 to one or more networks (not shown) via one or more network interfaces 1216. In this manner, the computer may be part of one or more computer networks, such as a local area network (LAN), a wide area network (WAN), or an intranet, or a network of networks, such as the Internet. Any or all components of electronic system 1200 may be used in conjunction with the present disclosure.
[0066] Embodiments within the scope of the present disclosure may be implemented in part or in whole using a tangible computer-readable storage medium (or multiple tangible computer-readable storage media of one or more types) encoding one or more instructions.The tangible computer-readable storage medium may also be non-transitory in nature.
[0067] Computer-readable storage media can be any storage medium that can be read, written, or otherwise accessed by a general-purpose or special-purpose computing device, including any processing electronics and / or processing circuitry capable of executing instructions. By way of example and not limitation, computer-readable media can include any volatile semiconductor memory, such as RAM, DRAM, SRAM, T-RAM, Z-RAM, and TTRAM. Computer-readable media can also include any non-volatile semiconductor memory, such as ROM, PROM, EPROM, EEPROM, NVRAM, Flash, nvSRAM, FeRAM, FeTRAM, MRAM, PRAM, CBRAM, SONOS, RRAM, NRAM, racetrack memory, FJG, and Millipede memory.
[0068] Furthermore, the computer-readable storage medium may include any non-semiconductor memory, such as optical disk storage devices, magnetic disk storage devices, magnetic tape, other magnetic storage devices, or any other medium capable of storing one or more instructions. In some embodiments, the tangible computer-readable storage medium may be directly coupled to the computing device, while in other embodiments, the tangible computer-readable storage medium may be indirectly coupled to the computing device, for example, via one or more wired connections, one or more wireless connections, or any combination thereof.
[0069] Instructions can be executed directly or can be used to develop executable instructions. For example, instructions can be implemented as executable or non-executable machine code, or as instructions in a high-level language that can be compiled to produce executable or non-executable machine code. In addition, instructions can also be implemented as data or can include data. Computer-executable instructions can also be organized in any format, including routines, subroutines, programs, data structures, objects, modules, applications, applets, functions, etc. As recognized by those skilled in the art, the details including but not limited to the number, structure, sequence, and organization of instructions can vary significantly without changing the underlying logic, functionality, processing, and output.
[0070] While the above discussion primarily refers to microprocessors or multi-core processors executing software, one or more implementations are performed by one or more integrated circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In one or more implementations, such integrated circuits execute instructions stored on the circuits themselves.
[0071] According to aspects of the present technology, a device is provided, comprising: a computer-readable storage medium storing one or more instruction sequences; and a processing circuit system configured to execute the one or more instruction sequences to: receive a plurality of network packets containing content encapsulated in a plurality of layers; process the received plurality of network packets to extract the content for presentation; generate a predicted presentation quality indicator for the extracted content using a plurality of machine learning models in a hierarchical order, wherein data generated during the processing of the received plurality of network packets is used as input to the plurality of machine learning models; and provide the predicted presentation quality indicator for the extracted content to a server via a network, wherein the data generated during the processing of the received plurality of network packets is correlated across the plurality of layers to generate the predicted presentation quality indicator.
[0072] The machine learning models of the plurality of machine learning models may be associated with corresponding layers of the plurality of layers. The data used as input to the plurality of machine learning models may come from a plurality of different domains, each corresponding to one or more layers of the plurality of layers. The different domains may include at least one of a packet-level domain, a bitstream-level domain, or a symbol-level domain. The output data generated by at least one of the plurality of machine learning models may be provided as input data to another of the plurality of machine learning models.
[0073] The content may include at least one of audio content or video content. The received plurality of network packets may further contain an expected presentation quality indicator, and providing the predicted presentation quality score to the server may be based on a comparison of the expected presentation quality score and the predicted presentation quality score. The processing circuitry may be further configured to: provide a prompt to confirm the predicted presentation quality indicator presented to a user; and receive a user response to the prompt, wherein providing the predicted presentation quality indicator to the server is based on the user's response to the prompt. The prompt may include at least one of an audio prompt or a video prompt. The processing circuitry may include at least one of a delivery engine, a stream processor, a codec, and a machine learning core.
[0074] According to aspects of the present technology, a method is provided, comprising: receiving a plurality of network packets containing content encapsulated in a plurality of layers; processing the received plurality of network packets to extract the content for presentation; generating a predicted presentation quality indicator for the extracted content using a plurality of machine learning models in a hierarchical order, wherein data generated during processing of the received plurality of network packets is used as input to the plurality of machine learning models; and providing the predicted presentation quality indicator for the extracted content to a server via a network, wherein the machine learning models in the plurality of machine learning models are associated with corresponding layers in the plurality of layers, and wherein the data generated during processing of the received plurality of network packets is correlated across the plurality of layers to generate the predicted presentation quality indicator.
[0075] The data used as input to the plurality of machine learning models may come from a plurality of different domains, each corresponding to one or more of the plurality of layers, and wherein the different domains may include at least one of a packet-level domain, a bitstream-level domain, or a symbol-level domain. The method may further include providing an output generated by at least one of the plurality of machine learning models as the input to another of the plurality of machine learning models.
[0076] The received plurality of network packets may further contain an expected presentation quality indicator, and the predicted presentation quality score of the server may be based on a comparison of the expected presentation quality score and the predicted presentation quality score. The method may further include: providing a prompt to confirm the predicted presentation quality indicator presented to a user; and receiving a user response to the prompt, wherein providing the predicted presentation quality indicator to the server is based on the user's response to the prompt.
[0077] According to aspects of the present technology, a system is provided that includes: a server; and a plurality of edge devices configured to communicate with the server over a network. Each of the plurality of edge devices includes: a computer-readable storage medium storing one or more sequences of instructions; and a processing circuit system configured to execute the one or more sequences of instructions to: receive a plurality of network packets containing content encapsulated in a plurality of layers; process the received plurality of network packets to extract the content for presentation; generate a predicted presentation quality metric for the extracted content using a plurality of machine learning models in a hierarchical order, wherein data generated during processing of the received plurality of network packets is used as input to the plurality of machine learning models; and provide the predicted presentation quality metric for the extracted content to the server over the network, wherein the data generated during processing of the received plurality of network packets is correlated across the plurality of layers to generate the predicted presentation quality metric, wherein the server is configured to correlate the predicted presentation quality metric provided by the plurality of edge devices in order to evaluate the system.
[0078] The machine learning models of the plurality of machine learning models may be associated with corresponding layers of the plurality of layers, and the data used as input to the plurality of machine learning models may come from a plurality of different domains, each corresponding to one or more layers of the plurality of layers. The output data generated by at least one of the plurality of machine learning models may be provided as input data to another of the plurality of machine learning models. The server may be configured to: generate an expected presentation quality metric for the content based on the original source of the content, wherein the plurality of network packets received by the plurality of edge devices further contain the expected presentation quality metric generated by the server, and wherein providing the predicted presentation quality score to the server is based on a comparison of the expected presentation quality score with the predicted presentation quality score. The processing circuitry of the plurality of edge devices may further be configured to: provide a prompt to confirm the predicted presentation quality metric presented to a user; and receive a user response to the prompt, wherein providing the predicted presentation quality metric to the server is based on the user's response to the prompt.
[0079] The preceding description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the various aspects shown herein, but should be given the full scope consistent with the language of the claims, wherein reference to an element in the singular is not intended to mean "one and only one" (unless specifically stated), but rather "one or more." Unless specifically stated otherwise, the term "some" refers to one or more. Masculine pronouns (e.g., his) include feminine and neuter genders (e.g., her and its), and vice versa. If any, titles and subtitles are used only for convenience and do not limit the present disclosure.
[0080] The predicates "configured to," "operable to," and "programmed to" do not imply any specific tangible or intangible modification of an object, but are intended to be used interchangeably. For example, a processor configured to monitor and control the operation of a component may also mean a processor programmed to monitor and control the operation or a processor operable to monitor and control the operation. Similarly, a processor configured to execute code may be configured as a processor programmed to execute code or operable to execute code.
[0081] Phrases such as "aspects" do not imply that such aspect is essential to the present technology or that such aspect applies to all configurations of the present technology. Disclosure related to an aspect may apply to all configurations or to one or more configurations. Phrases such as aspects may refer to one or more aspects, and vice versa. Phrases such as "configurations" do not imply that such configuration is essential to the present technology or that such configuration applies to all configurations of the present technology. Disclosure related to a configuration may apply to all configurations or to one or more configurations. Phrases such as configurations may refer to one or more configurations, and vice versa.
[0082] The word “exemplary” is used herein to mean “serving as an example or illustration.” Any aspect or design described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects or designs.
[0083] All structural and functional equivalents of the elements of the various aspects described throughout this disclosure that are known or later known to those skilled in the art are expressly incorporated herein by reference and are intended to be covered by the claims. In addition, nothing disclosed herein is intended to be committed to the public, regardless of whether this disclosure is explicitly stated in the claims. Any claim element will not be interpreted under the provisions of 35 U.S.C. § 112 (f) unless the element is explicitly stated using the phrase "means for..." or, in the case of a method claim, the element is stated using the phrase "step for..." In addition, with respect to the use of the terms "including," "having," and the like in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the meaning of the term "comprising" when "comprising" is used as a transition word in the claims.
[0084] Those skilled in the art will appreciate that the various illustrative blocks, modules, elements, components, methods, and algorithms described herein can be implemented as electronic hardware, computer software, or a combination of both. To illustrate this interchangeability of hardware and software, various illustrative blocks, modules, elements, components, methods, and algorithms have generally been described above in terms of their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. A skilled person may implement the described functionality in varying ways for each specific application. The various components and blocks may be arranged in different ways (e.g., arranged in a different order, or partitioned in different ways), all without departing from the scope of the present technology.
[0085] The predicates "configured to," "operable to," and "programmed to" do not imply any specific tangible or intangible modification of an object, but are intended to be used interchangeably. For example, a processor configured to monitor and control the operation of a component may also mean a processor programmed to monitor and control the operation or a processor operable to monitor and control the operation. Similarly, a processor configured to execute code may be configured as a processor programmed to execute code or operable to execute code.
Claims
1. A device comprising: A computer-readable storage medium storing one or more sequences of instructions; and processing circuitry configured to execute the one or more sequences of instructions to: receiving a plurality of network packets including content encapsulated in a plurality of layers; processing the received plurality of network packets to extract the content for presentation; generating a predicted presentation quality indicator for the extracted content using a plurality of machine learning models in a hierarchical order, wherein data generated during processing of the received plurality of network packets is used as input to the plurality of machine learning models; and The predicted presentation quality indicator for the extracted content is provided to a server via a network, wherein the data generated during processing of the received plurality of network packets is correlated across the plurality of layers to generate the predicted presentation quality indicator.
2. An apparatus according to claim 1, wherein the machine learning model among the multiple machine learning models is associated with a corresponding layer among the multiple layers.
3. A device according to claim 2, wherein the data used as input to the multiple machine learning models is from multiple different domains, each corresponding to one or more of the multiple layers.
4. The device of claim 3, wherein the different domains comprise at least one of a packet level domain, a bitstream level domain, or a symbol level domain.
5. An apparatus according to claim 3, wherein output data generated by at least one of the multiple machine learning models is provided as input data to another one of the multiple machine learning models. The device of claim 1 , wherein the content comprises at least one of audio content or video content.
7. The apparatus of claim 1, wherein the received plurality of network packets further contain an expected rendering quality indicator, and Wherein providing the predicted presentation quality score to the server is based on a comparison of an expected presentation quality score and the predicted presentation quality score.
8. The device of claim 1, wherein the processing circuitry is further configured to: providing a prompt to confirm the predicted presentation quality indicator presented to the user; and receiving a user response to the prompt, Wherein providing the predicted presentation quality indicator to the server is based on a response of the user to the prompt.
9. The device of claim 8, wherein the prompt comprises at least one of an audio prompt or a video prompt.
10. The device of claim 1, wherein the processing circuitry comprises at least one of a transport engine, a stream processor, a codec, and a machine learning core.
11. A method comprising: receiving a plurality of network packets including content encapsulated in a plurality of layers; processing the received plurality of network packets to extract the content for presentation; generating a predicted presentation quality indicator for the extracted content using a plurality of machine learning models in a hierarchical order, wherein data generated during processing of the received plurality of network packets is used as input to the plurality of machine learning models; and providing the predicted presentation quality indicator for the extracted content to a server via a network, wherein the machine learning model in the plurality of machine learning models is associated with a corresponding layer in the plurality of layers, and wherein the data generated during processing of the received plurality of network packets is correlated across the plurality of layers to generate the predicted presentation quality indicator.
12. The method of claim 11, wherein the data used as input to the plurality of machine learning models is from a plurality of different domains each corresponding to one or more of the plurality of layers, and Wherein the different domains include at least one of a packet level domain, a bitstream level domain, or a symbol level domain.
13. The method of claim 11, further comprising providing an output generated by at least one of the plurality of machine learning models as the input to another of the plurality of machine learning models.
14. The method of claim 11, wherein the received plurality of network packets further contain an expected presentation quality indicator, and Wherein providing the predicted presentation quality score to the server is based on a comparison of an expected presentation quality score and the predicted presentation quality score.
15. The method according to claim 11, further comprising: providing a prompt to confirm the predicted presentation quality indicator presented to the user; and receiving a user response to the prompt, Wherein providing the predicted presentation quality indicator to the server is based on a response of the user to the prompt.
16. A system comprising: server; as well as A plurality of edge devices configured to communicate with the server via a network, wherein each edge device of the plurality of edge devices comprises: A computer-readable storage medium storing one or more sequences of instructions; and processing circuitry configured to execute the one or more sequences of instructions to: receiving a plurality of network packets including content encapsulated in a plurality of layers; processing the received plurality of network packets to extract the content for presentation; generating a predicted presentation quality indicator for the extracted content using a plurality of machine learning models in a hierarchical order, wherein data generated during processing of the received plurality of network packets is used as input to the plurality of machine learning models; and providing the predicted presentation quality indicator for the extracted content to the server via the network, wherein the data generated during processing of the received plurality of network packets is correlated across the plurality of layers to generate the predicted rendering quality indicator, Wherein the server is configured to correlate the predicted presentation quality indicators provided by the plurality of edge devices in order to evaluate the system.
17. The system of claim 16, wherein the machine learning model of the plurality of machine learning models is associated with a corresponding layer of the plurality of layers, and Wherein the data used as input to the plurality of machine learning models is from a plurality of different domains each corresponding to one or more of the plurality of layers.
18. A system according to claim 17, wherein output data generated by at least one of the multiple machine learning models is provided as input data to another one of the multiple machine learning models.
19. The system of claim 18, wherein the server is configured to: generating an expected rendering quality indicator for the content based on an original source of the content, wherein the plurality of network packets received by the plurality of edge devices further contain the expected presentation quality indicator generated by the server, and Wherein providing the predicted presentation quality score to the server is based on a comparison of an expected presentation quality score and the predicted presentation quality score.
20. The system of claim 16, wherein the processing circuitry of the plurality of edge devices is further configured to: providing a prompt to confirm the predicted presentation quality indicator presented to the user; and receiving a user response to the prompt, Wherein providing the predicted presentation quality indicator to the server is based on a response of the user to the prompt.