Reference-free generative machine listener
A reference-free DNN initialized from a pre-trained model estimates subjective listening scores in adaptive streaming, addressing the need for variability estimation and real-time monitoring without reference signals, enhancing audio quality assessment efficiency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2026-03-12
AI Technical Summary
Existing audio quality assessment methods in adaptive streaming environments lack the ability to provide an estimate of variability in subjective listening scores and require access to reference signals, which is not feasible for streaming clients, leading to inefficiencies in real-time quality monitoring.
A reference-free deep neural network (DNN) is configured to estimate subjective listening scores by initializing weights from a pre-trained DNN, allowing it to predict individual listening scores, including probability distributions and confidence intervals, without requiring a reference signal, and is trained using Gammatone spectrograms to emulate human hearing.
The solution enables efficient, real-time audio quality monitoring on streaming clients, providing accurate estimates of subjective listening scores and variability, supporting broader Quality-of-Experience delivery efforts.
Smart Images

Figure EP2025074361_12032026_PF_FP_ABST
Abstract
Description
REFERENCE-FREE GENERATIVE MACHINE LISTENER Cross-Reference to Related Applications
[0001] This application claims the benefit of priority from U.S. Provisional Application No. 63 / 691,557, filed on 6 September 2024, which is incorporated by reference herein in its entirety. Technical Field
[0002] The present disclosure relates to techniques for evaluating playout performance in an adaptive streaming environment, and in relation thereto, to techniques for configuring a deep neural network (DNN) for estimating indications of a subjective listening score for a test audio signal. In particular, the present disclosure relates to techniques for configuring a reference-free DNN for estimating indications of a subjective listening score for a test audio signal. In particular, the present disclosure relates to techniques for evaluating playout performance directly on a client device. Background
[0003] Algorithms for objective quality assessment of coding schemes can be divided into two categories. There are so-called intrusive algorithms, which require that a reference signal is provided in addition to the test signal, and there are also non-intrusive algorithms, which only analyze the test signal. In general, intrusive algorithms work better than non- intrusive algorithms. For example, in the context of lossy coding it may be difficult to assess the performance of a codec without the reference, since an artistic intent is unknown in this situation (e.g., a bandlimited audio signal could be a result of a deliberate processing step applied in content production rather than an artifact from a codec). Furthermore, any metrics provided by a non-intrusive quality assessment tool may be hard to interpret, as such tools typically lack grounding in well-established subjective quality assessment methodologies. For example, a Multi-Stimulus Test with Hidden Reference and Anchor (MUSHRA) test would be commonly used to evaluate audio codecs, and it involves usage of the reference signal.
[0004] Objective quality assessment may be of particular relevance in adaptive streaming environments where a quality (e.g., bitrate) of delivered content can be dynamically adjusted. What is of interest here is the playout audio quality at respective streaming clients. Intrusive algorithms however are not applicable in this case, as the streaming clients typically do not have access to reference signals.
[0005] Further, listener scores achieved in listening tests such as MUSHRA tests could be predicted by a system which takes the signal under test and the reference signal as inputs. For a given pair of input signals, there will be a certain degree of variability in the listener scores and there is value in capturing that aspect of the data for automated estimation of quality of experience, for example in entertainment delivery systems. In an actual subjective MUSHRA test with multiple listeners, the mean and standard deviation of the listener scores can be computed. The standard deviation can then be converted to a confidence interval given the number of listeners and a statistical model.
[0006] However, current automated implementations of subjective listening test such as MUSHRA tests merely provide a mean value of the MUSHRA score and cannot provide an idea of variability of subjective listening scores that would be obtained from a population of test listeners.
[0007] Thus, there is a need for improvements in assessing audio quality in adaptive streaming environments. There is further need for improvements in automated predictions of subjective listening scores. There is particular need for techniques that can provide an estimate of a variability of subjective listening scores in automated solutions.
[0008] The usage of the reference signal however would be problematic in particular if the test were to be performed on a client (as it would require out-of-band placement of the reference signal). Based on deep learning, non-intrusive metrics have evolved to provide competitive accuracy. Considering a streaming application, it is necessary to monitor the audio quality of streaming in real-time, even though no reference may be available at the streaming client.
[0009] Accordingly, a reference-free objective audio quality metric has been demanded in plenty of circumstances where a clean and aligned reference is missing. Major reference-free quality models currently focus on mono-speech signals at low sample rates. State-of-the-art methodologies with the aid of neural network like SESQA(https: / / ieeexplore.ieee.org / document / 9414052) works on 48 kHz mono speech signals and SQUIM from Meta Inc. (https: / / pytorch.org / audio / main / tutorials / squim_tutorial.html) deals with 16 kHz speech signals. Despite their advance in algorithms, these two methods in practice have shown strong content-dependency, low accuracy with subjective scores, or reliance on semi-reference signals.
[0010] Thus, there is a need for improvements in assessing audio quality with non- intrusive schemes. There is further need for improvements in providing real-time predictions of subjective listening scores on a client side, which does not require access to the reference. Summary
[0011] In view of at least some of these needs, the present disclosure provides methods and apparatus for configuring a deep neural network (DNN) for estimating an indication of a subjective listening score, methods for estimating an indication of a subjective listening score using a DNN, DNNs, as well as corresponding apparatus, programs, and computer-readable storage media.
[0012] The present disclosure further provides methods and apparatus for evaluating playout performance in an adaptive streaming environment, methods of providing playout- related information, as well as corresponding apparatus, programs and computer-readable storage media.
[0013] The present disclosure further provides methods for configuring a reference- free DNN for estimating an indication of a subjective listening score for an audio signal, methods for estimating an indication of a subjective listening score for an audio signal using a reference-free DNN, a reference-free DNN for estimating an indication of a subjective listening score for an audio signal, as well as corresponding apparatus, programs, and computer-readable storage media.
[0014] The present disclosure further provides methods for an audio steaming device estimating an indication of a subjective listening score for an audio signal played out by the audio streaming device, as well as corresponding apparatus, programs, and computer-readable storage media.
[0015] An aspect of the present disclosure relates to a method of configuring a reference-free DNN for estimating an indication of a subjective listening score for an audiosignal (e.g., test audio signal). The method may be a method of training the reference-free DNN, for example. The listening score may be a score according to a listening test performed according to a predefined listening test methodology. The predefined listening test methodology may be a standardized listening test methodology. Further, the listening test may apply a predefined test metric and / or test scenario. The method may include providing an output stage of the reference-free DNN to generate the indication of the listening score. The method may include initializing weights for at least one of a plurality of layers of the reference-free DNN. The initialization may be performed for any one or more of the plurality of layers, for example, the shallowest layer(s) of the reference-free DNN. The method may further include, after initializing said weights, training the reference-free DNN by, in a training epoch among a plurality of training epochs, inputting one or more training data items, each indicative of a respective value of the listening score and further indicative of a representation of the audio signal, into an input stage of the reference-free DNN. Training the reference-free DNN may further include, in the training epoch, determining respective indications of the listening score based on the one or more training data items. Training the reference-free DNN may further include, in the training epoch, determining respective loss values for the one or more training data items by evaluating a loss function. Here, the loss function may depend on the indication of the listening score. Training the reference-free DNN may yet further include, in the training epoch, adjusting one or more internal parameters of the reference-free DNN based on the determined loss values. The internal parameters of the reference-free DNN may be model parameters, for example, such as coefficients (e.g., filter coefficients) of a plurality of layers of the DNN. The method may further include initializing the weights of the reference-free DNN based on a pre-trained DNN for estimating an indication of a subjective listening score for the audio signal based on the audio signal and a reference audio signal for the audio signal. The pre-trained DNN may be a full-reference DNN, e.g., a full-reference generative machine listener. According to the method, the pre- trained DNN may comprise an input layer for receiving a representation of the audio signal and a representation of the reference audio signal for the audio signal, a plurality of layers for performing processing based on the representation of the audio signal and the representation of the reference audio signal, and an output stage for generating the indication of the listening score. The reference-DNN and the pre-trained DNN may have similar configurations other than the configuration relating to the reference signal. The method may further comprise initializing the weights of the reference-free DNN based on first weights of the input layer of the pre-trained DNN, which are weights associated with the representation of the audio signal,and / or second weights of the input layer of the pre-trained DNN, which are weights associated with the representation of the reference audio signal.
[0016] Accordingly, the present disclosure provides a reference-free audio quality metric with simplified training, in particular by utilizing parameters from a pre-trained audio quality metric. This improves the training efficiency since using pre-trained parameters may shorten the training complexity and time needed for the reference-free audio quality metric.
[0017] Further, the reference-free DNN is trained not on mean values of the subjective listening score, but rather on individual listening scores. This allows to adapt the reference- free DNN for predicting parameters beyond the mean listening score, including, for example, a probability distribution, standard deviation, and / or confidence interval of the listening score.
[0018] In some embodiments, the initialized weights of the reference-free DNN may be weights of the input stage of the reference-free DNN. This provides a simple initialization of the weights of the reference-free DNN.
[0019] In some embodiments, the initialized weights of the reference-free DNN may be initialized based on the first weights of the pre-trained DNN. Since the reference signal is not necessary in the reference-free DNN, initializing the weights of the reference-free DNN using the first weights of the pre-trained DNN associated with the representation of the audio signal provides improved training efficiency.
[0020] In some embodiments, a value of each of the initialized weights of the reference-free DNN may be based on an average value of a respective first weight and a respective second weight. Using both the first weights of the pre-trained DNN associated with the representation of the audio signal and the second weights of the pre-trained DNN associated with the representation of the reference audio signal allows a more comprehensive utilization of the pre-trained DNN, which also provides improved training efficiency.
[0021] In some embodiments, the one or more training data items are each indicative of a respective representation of a binaural transformation of the audio signal, wherein the audio signal is a multi-channel audio signal or an object-based audio signal. This allows efficient playout performance also for spatial audio signals.
[0022] In some embodiments, the indication of the listening score may relate to a probability distribution of the listening score, with the output stage being adapted for generating the probability distribution of the listening score. The probability distribution mayemulate listening scores obtained by a plurality of listening tests for the audio signal. The plurality of listening tests emulated by the probability distribution may be independent listening tests. Further, the probability distribution may be parameterized by two or more parameters of the probability distribution. Then, determining respective indications of the listening score based on the one or more training data items may include determining respective parameters of the probability distribution based on the one or more training data items. Determining the parameters of the probability distribution may be based at least in part on the value of the subjective listening score. Further, determining the parameters of the probability distribution may be based on a current state of the reference-free DNN, for example the current values of internal parameters of the reference-free DNN. The loss function may depend on the parameters of the distribution.
[0023] In some embodiments, training the reference-free DNN may be based on a maximum likelihood principle.
[0024] In some embodiments, the loss function may relate to a negative log likelihood, NLL, loss.
[0025] In some embodiments, the probability distribution may relate to a Gaussian ^ distribution parameterized by a mean ^ and a variance ^ . Then, the loss function ^ may^^^^^be given by ^^ = log ^ +^^^^+ ^, where ^ is a constant and ^ is the subjectivelistening score. The constant ^ may be given by ^ = 2^, for
[0026] In some embodiments, the probability distribution may relate to a logistic distribution parameterized by a mean ^ and a scale ^. Then, the loss function ^ may be^^^^^^^^given by ^ = log ^ + 2 log sech + ^, where ^ is a constant and ^ is the subjective $listening score. The constant ^ may be given by ^ = log 4, for example.
[0027] Both these parameterizations of the probability distribution have been found to provide for efficient training at the training stage, and meaningful output at inference.
[0028] In some embodiments, the representation of the audio signal and the representation of the reference audio signal may relate to Gammatone spectrograms.
[0029] Gammatone spectrograms are auditory features specifically adapted to human hearing and perception and therefore allow for achieving meaningful results at reduced computational complexity.
[0030] In some embodiments, the predefined listening test may be a Multi-Stimulus Test with Hidden Reference and Anchor (MUSHRA) listening test, for example as standardized under ITU-R recommendation BS.1534. Additionally or alternately, the predefined listening test may be MUSHRA-like listening test, but without presenting (or exposing the known) reference to the subjects, which is basically the same as a MUSHRA test, but without an open reference and anchor.
[0031] In some embodiments, the DNN may implement a generative model.
[0032] In some embodiments, the pre-trained DNN may have been trained based onthe negative log likelihood loss that may be given by ^'(( = − log *+^^|-, / ^, where*+^^|-, / ^ is the probability distribution for test score ^ given a representation of the audiosignal / and a representation of a reference audio signal - for the audio signal / , and 0 indicates the internal parameters of the pre-trained DNN.
[0033] Using the maximum likelihood principle, or correspondingly, negative log likelihood loss allows to efficiently train the DNN based on individual listening scores to provide an indication of listening scores that could be expected in an actual subjective listening test with multiple listeners.
[0034] Another aspect of the disclosure relates to a method of estimating an indication of a subjective listening score for an audio signal using a reference-free DNN. The listening score may be a score according to a predefined listening test. The reference-free DNN may include an input stage for receiving test data items indicative of a representation of the audio signal. The reference-free DNN may further include a plurality of layers for performing processing based on the received test data items. Processing by the plurality of layers may be further based on a current state of the reference-free DNN, for example the current values of internal parameters of the reference-free DNN. The reference-free DNN may yet further include an output stage, connected to a last one of the plurality of layers, for generating the indication of the listening score. The method may include inputting the test data items to the input stage. The method may further include determining a representation of the indication of the listening score based on an output of the output stage.
[0035] In some embodiments, the test data items may each be indicative of a respective representation of a binaural transformation of the audio signal, wherein the audio signal is a multi-channel audio signal or an object-based audio signal.
[0036] In some embodiments, the method may be performed by an audio streaming device, wherein the audio signal may be played out by the audio streaming device. Accordingly, the audio streaming device performs reference-free audio quality analysis in an efficient manner as described above.
[0037] In some embodiments, the indication of the listening score may relate to a probability distribution of the listening score, with the output stage of the reference-free DNN being adapted for generating the probability distribution of the listening score. The probability distribution may emulate listening scores obtained by a plurality of (subjective) listening tests for the audio signal. The probability distribution may be parameterized by two or more parameters of the probability distribution.
[0038] In some embodiments, determining the representation of the probability distribution may include determining at least one of a mean, a standard deviation, and a confidence interval from the output of the output stage.
[0039] In some embodiments, the confidence interval may be determined based on the output of the output stage and a number (e.g., count) of listeners to the listening test to be emulated.
[0040] In some embodiments, the probability distribution may relate to a Gaussian distribution parameterized by a mean ^ and a variance ^^or to a logistic distribution parameterized by a mean ^ and a scale ^.
[0041] In some embodiments, the representation of the audio signal and the representation of the reference audio signal may relate to Gammatone spectrograms.
[0042] In some embodiments, the predefined listening test may be a MUSHRA listening test and / or MUSHRA-like listening test, but without presenting (or exposing the known) reference to the subjects.
[0043] Another aspect of the disclosure relates to a reference-free DNN for estimating an indication of a subjective listening score for an audio signal. The listening score may be a score according to a predefined listening test. The reference-free DNN may include an input stage for receiving test data items indicative of a representation of the audio signal. The DNN reference-free may further include a plurality of layers for performing processing based on the received test data items. Processing by the plurality of layers may be further based on a current state of the reference-free DNN, for example the current values of internal parametersof the reference-free DNN. The reference-free DNN may yet further include an output stage, coupled to a last one of the plurality of layers, for generating the indication of the listening score.
[0044] In some embodiments, the test data items may each be indicative of a respective representation of a binaural transformation of the audio signal, wherein the audio signal is a multi-channel audio signal or an object-based audio signal.
[0045] In some embodiments, the reference-free DNN may have been configured by initializing training weights for at least one of the plurality of layers and training the reference-free DNN by, in a training epoch among a plurality of training epochs, inputting one or more training data items, each indicative of a representation of a training audio signal and a respective value of the listening score for the training audio signal, into the input stage of the reference-free DNN. Training the reference-free DNN, in the training epoch, may further include determining respective indications of the listening score based on the one or more training data items. Training the reference-free DNN, in the training epoch, may further include determining respective loss values for the one or more training data items by evaluating a loss function. This loss function may depend on the indication of the listening score. Training the reference-free DNN, in the training epoch, may yet further include adjusting one or more internal parameters of the reference-free DNN based on the determined loss values. Configuring the reference-free DNN may involve or correspond to obtaining internal parameters of the reference-free DNN by training the reference-free DNN. The training weights of the reference-free DNN may be initialized based on a pre-trained DNN for estimating an indication of a subjective listening score for the training audio signal based on the training audio signal and a reference audio signal for the training audio signal. The pre- trained DNN may comprise an input layer for receiving a representation of the training audio signal and a representation of the reference audio signal for the training audio signal, a plurality of layers for performing processing based on the representation of the training audio signal and the representation of the reference audio signal, and an output stage for generating the indication of the listening score. The training weights of the reference-free DNN may be initialized based on first weights of the input layer of the pre-trained DNN, which are weights associated with the representation of the training audio signal, and / or second weights of the input layer of the pre-trained DNN, which are weights associated with the representation of the reference audio signal.
[0046] In some embodiments, the initialized training weights of the reference-free DNN may be weights of the input stage of the reference-free DNN.
[0047] In some embodiments, the initialized training weights of the reference-free DNN may be initialized based on the first weights of the pre-trained DNN.
[0048] In some embodiments, a value of each of the initialized training weights of the reference-free DNN may be based on an average value of a respective first weight and a respective second weight.
[0049] In some embodiments, the indication of the listening score may relate to a probability distribution of the listening score, with the output stage of the reference-free DNN being adapted for generating the probability distribution of the listening score. The probability distribution may emulate listening scores obtained by a plurality of listening tests for the audio signal. Further, the probability distribution may be parameterized by two or more parameters of the probability distribution.
[0050] In some embodiments, determining respective indications of the listening score based on the one or more training data items may include determining respective parameters of the probability distribution based on the one or more training data items. Determining the parameters of the probability distribution may be based at least in part on the value of the subjective listening score. Further, determining the parameters of the probability distribution may be based on a current state of the reference-free DNN, for example the current values of internal parameters of the reference-free DNN. The loss function may depend on the parameters of the distribution.
[0051] In some embodiments, the probability distribution may relate to a Gaussian distribution parameterized by a mean ^ and a variance ^^or to a logistic distribution parameterized by a mean ^ and a scale ^.
[0052] In some embodiments, the representation of the audio signal and the representation of the reference audio signal may relate to (one or more) Gammatone spectrograms.
[0053] In some embodiments, the predefined listening test may be a MUSHRA listening test and / or MUSHRA-like listening test, but without presenting (or exposing the known) reference to the subjects.
[0054] Another aspect of the disclosure relates to an apparatus comprising a processor and a memory coupled to the processor, and storing instructions for the processor. The processor may be adapted to carry out the methods according to the foregoing aspects and their embodiments.
[0055] Another aspect of the disclosure relates to an apparatus comprising a processor and a memory coupled to the processor, and storing instructions for the processor. The processor may be adapted to implement DNNs according to the foregoing aspects and their embodiments.
[0056] Another aspect of the disclosure relates to a program comprising instructions that, when executed by a processor, cause the processor to carry out the methods according to the foregoing aspects and their embodiments.
[0057] Another aspect of the disclosure relates to a program comprising instructions that, when executed by a processor, cause the processor to implement the DNN according to the foregoing aspects and their embodiments.
[0058] Another aspect of the disclosure relates to a computer-readable storage medium storing any of the aforementioned programs.
[0059] Another aspect of the disclosure relates to a method of an audio streaming device estimating an indication of a subjective listening score for an audio signal played out by the audio streaming device. The listening score may be a score according to a predefined listening test according to a predefined listening test methodology. The predefined listening test methodology may be a standardized listening test methodology. The method may comprise obtaining metadata relating to the played out audio signal. The method may comprise providing, based on the obtained metadata, test data for being input into a trained reference-free deep neural network, DNN, for estimating the indication of the listening score. The test data may be obtained in real-time during data streaming on the audio steaming device. The reference-free DNN may be trained and obtained in accordance with the previously described aspects and embodiments and may comprise an input stage for receiving the test data, a plurality of layers for performing processing based on the received test data, and an output stage for generating the indication of the listening score. The method may further comprise inputting the test data into the input stage. The method may further comprise determining a representation of the indication of the listening score based on an output of the output stage.
[0060] According to the method, an efficient way is provided to automate the monitoring of the quality of audio in a streaming scenario and also in real time. As a result, one can objectively monitor the quality of audio streams on the client side and can preemptively understand and adjust the delivery mechanism on the server side. Further, this method supports broader Quality-of-Experience delivery efforts.
[0061] In some embodiments, the method may further comprise selecting, based on the obtained metadata and from a plurality of trained algorithms, an algorithm to be applied in the trained reference-free DNN for estimating the indication of the listening score. This allows an appropriate selection of the trained model or algorithm to be applied for the obtained metadata, providing improved estimation of the listening score specifically for the currently obtained metadata, i.e., the currently played out audio signal.
[0062] In some embodiments, each of the plurality of trained algorithms may correspond to a respective playout mode for being used by the audio streaming device, wherein the metadata may comprise information relating to the playout mode currently used on the audio streaming device. This allows an appropriate selection of the trained model or algorithm to be applied for the currently used playout mode, providing improved estimation of the listening score specifically for the currently used playout mode.
[0063] In some embodiments, the playout modes may comprise a headphone mode, a discrete speaker mode and a soundbar mode.
[0064] In some embodiments, the metadata may comprise information relating to one or more of the following analyses: an analysis of the playout buffer provided in the audio streaming device, an analysis of a bitrate ladder used for providing the audio streaming device with streaming audio data, an analysis of the audio streaming device, and an analysis of playout conditions. Accordingly, specific playout-relevant information is obtained for providing improved estimation of the listening score.
[0065] In some embodiments, the test data may be indicative of a representation of a binaural transformation of the played out audio signal, wherein the played out audio signal may relate to a multi-channel signal or an object-based signal.
[0066] In some embodiments, the test data may be indicative of a representation of the audio signal and the representation of the audio signal may relate to Gammatone spectrograms.
[0067] In some embodiments, the representation of the audio signal may comprise segments of Gammatone spectrograms, each segment being of a respective pre-determined time length in the played out audio signal, wherein the test data may comprise information relating to respective segment names and / or corresponding bitrates used for the respective segments. This provides detailed analysis of the played out audio signal.
[0068] In some embodiments, the method may comprise estimating indications of respective listening scores for a plurality of audio streaming device clusters, each cluster corresponding to a respective content delivery infrastructure, and each cluster comprising one or more audio streaming devices. The method may further comprise estimating for each cluster an indication of a listening score based on metadata collected from the one or more audio streaming devices included in the each cluster, and comparing the estimated listening scores of the clusters.
[0069] Accordingly, an offline setting of the listening score estimation procedure on the audio streaming device is provided, wherein the listening scores for multiple client devices or multiple clusters of client devices are obtained offline after the respective test data is obtained from the multiple devices / clusters. The listening scores can therefore be provided offline in batch, for enabling further analysis with regard to the multiple devices / clusters.
[0070] In some embodiments, the method may further comprise providing the estimated indication of respective listening score to a content encoding and packaging infrastructure for adjusting a number of levels used in a bitrate ladder for providing the audio streaming device with streaming audio data and / or for adjusting bitrates applied in the content encoding and packaging infrastructure. Accordingly, it is provided a feedback loop using the estimated listening score as a feedback for adjusting the performance of the content encoding and packaging infrastructure.
[0071] In some embodiments, the reference-free DNN may have been configured by initializing training weights for at least one of the plurality of layers and training the reference-free DNN by, in a training epoch among a plurality of training epochs: inputting one or more training data items, each indicative of a representation of a training audio signal and further indicative of a respective value of a listening score for the training audio signal; determining respective indications of the listening score based on the one or more training data items; determining respective loss values for the one or more training data items by evaluating a loss function, wherein the loss function depends on the indication of the listeningscore; and adjusting one or more internal parameters of the reference-free DNN based on the determined loss values, wherein the training weights of the reference-free DNN are initialized based on a pre-trained DNN for estimating an indication of a listening score for the training audio signal based on the training audio signal and a reference audio signal for the training audio signal, wherein the pre-trained DNN comprises: an input layer for receiving a representation of the training audio signal and a representation of the reference audio signal for the training audio signal; a plurality of layers for performing processing based on the representation of the training audio signal and the representation of the reference audio signal; and an output stage for generating the indication of the listening score wherein the training weights of the reference-free DNN are initialized based on first weights of the input layer of the pre-trained DNN, which are weights associated with the representation of the training audio signal, and / or second weights of the input layer of the pre-trained DNN, which are weights associated with the representation of the reference audio signal.
[0072] Accordingly, the reference-free DNN may have been configured in accordance with the above aspects and embodiments, and with the corresponding technical advantages.
[0073] In some embodiments, the initialized training weights of the reference-free DNN may be weights of the input stage of the reference-free DNN.
[0074] In some embodiments, the initialized training weights of the reference-free DNN may be initialized based on the first weights of the pre-trained DNN.
[0075] In some embodiments, a value of each of the initialized training weights of the reference-free DNN may be based on an average value of a respective first weight and a respective second weight.
[0076] In some embodiments, the indication of the listening score may relate to a probability distribution of the listening score, with the output stage of the reference-free DNN being adapted for generating the probability distribution of the listening score, wherein the probability distribution emulates listening scores obtained by a plurality of listening tests for the audio signal, and wherein the probability distribution is parameterized by two or more parameters of the probability distribution.
[0077] In some embodiments, determining the representation of the probability distribution may comprise determining at least one of a mean, a standard deviation, and a confidence interval from the output of the output stage.
[0078] In some embodiments, the predefined listening test may be a Multi-Stimulus Test with Hidden Reference and Anchor, MUSHRA, listening test and / or MUSHRA-like listening test, but without presenting (or exposing the known) reference to the subjects.
[0079] Another aspect of the disclosure relates to an apparatus comprising a processor and a memory coupled to the processor, and storing instructions for the processor, wherein the processor may be adapted to carry out the method according to any one of the above aspects and their embodiments.
[0080] Another aspect of the disclosure relates to a program comprising instructions that, when executed by a processor, may cause the processor to implement the DNN according to any one of the above aspects and their embodiments.
[0081] Another aspect of the disclosure relates to a computer-readable storage medium storing the aforementioned programs.
[0082] Another aspect of the disclosure relates to a method of evaluating playout performance in an adaptive streaming environment. Playout performance may relate to (subjective) playout quality, for example. The method may include obtaining playout-related information from a streaming client. The playout-related information may include, correspond to, or be in the form of metadata, for example. The method may further include estimating a representation of a test audio signal based on the playout-related information. The test audio signal may be an audio signal played out by the streaming client. Estimating the representation of the test audio signal may involve or correspond to reconstructing the test audio signal or a representation thereof. The representation may relate to a set of features or spectrograms of the test audio signal, for example. The method may yet further include determining, using an audio quality assessment algorithm, an estimate of an audio quality of the test audio signal based on the estimated representation of the test audio signal. The audio quality assessment algorithm may be an objective audio quality assessment algorithm. Further, the audio quality assessment algorithm may emulate an intrusive audio quality test (e.g., listening test), such as a MUSHRA test, for example. It is understood that the estimated representation of the test audio signal is in a form suitable for input to the audio quality assessment algorithm. For example, the representation of the test audio signal may relate to a predefined number of segments to accommodate for requirements of the audio quality assessment algorithm relating to a time span covered by input audio signals or representations thereof.
[0083] Thereby, the proposed method allows to estimate results of an intrusive listening test without requiring knowledge of the reference signal at the streaming client. Configured as described above, the intrusive listening test can be performed at a network node removed from the streaming client. The streaming client is only required to provide lightweight metadata to the network node performing the test. Both, a version of the played out audio as well as a reference for the played out audio can be derived at said network node, using the metadata. As a result, the proposed method can yield a meaningful and readily interpretable estimate of an audio quality of audio content played out by the streaming client without significant additional signaling overhead to or from the streaming client.
[0084] In some embodiments, the method may further include generating a representation of a reference audio signal for the test audio signal. This may include obtaining (e.g., receiving) audio content or a representation thereof from a content repository. It is understood that the representation of the reference audio signal is in a form suitable for input to the audio quality assessment algorithm. For example, the representation of the reference audio signal may relate to a predefined number of segments to accommodate for requirements of the audio quality assessment algorithm relating to a time span covered by input audio signals or representations thereof.
[0085] In some embodiments, the method may further include obtaining, from the streaming client, an indication of audio content processed by the streaming client. The indication of the audio content may comprise an identifier, such as a file name, etc. of the audio content, a bitrate level of the audio content, and / or information on a segmentation of the audio content. The indication of audio content may be used to determine a sequence of audio segments received by the streaming client. The audio content processed by the streaming client may be audio content received by the streaming client, for example from a content delivery network.
[0086] In some embodiments, estimating the representation of the test audio signal may be further based on the indication of the audio content.
[0087] In some embodiments, generating the representation of the reference audio signal may be based on the indication of the audio content. This may comprise obtaining (e.g., receiving) audio content or a representation thereof from a content repository, based on the indication of the audio content that is played out by the streaming client.
[0088] In some embodiments, the playout-related information may include bitrate information indicating a bitrate of the audio signal played out by the streaming client. Then, estimating the representation of the test audio signal may be based on the bitrate information. The bitrate information may be provided for each of a plurality of segments of the audio signal played out by the streaming client.
[0089] In some embodiments, the audio quality assessment algorithm may use a set of pretrained models for audio quality assessment. Generating the estimate of the audio quality may include selecting a pretrained model among the set of pretrained models based on the playout-related information.
[0090] Thereby, it can be ensured that the optimal model is used for each relevant situation (e.g., a model particularly trained for the situation), thereby improving reliability of the estimation of playout performance.
[0091] In some embodiments, the playout related information may include information relating to a playout device associated with the streaming client. Then, the pretrained model may be selected based on the information relating to the playout device. The information relating to the playout device may include an indication of the playout device (e.g., headphones, soundbar, discrete speakers, etc.) and / or an indication of characteristics of the playout conditions (e.g., SNR, etc.).
[0092] Accordingly, an appropriate model for audio quality assessment may be used for each of a plurality of different playout device configurations, thereby improving reliability of the estimation of playout performance.
[0093] In some embodiments, the audio quality assessment algorithm may be implemented by a deep neural network, DNN, for estimating an indication of a subjective listening score for the representation of a test audio signal as the estimate of the audio quality. The listening score may be a score according to a predefined listening test. The DNN may include an input stage for receiving the representation of the test audio signal and a representation of a reference audio signal for the test audio signal. The DNN may further include a plurality of layers for performing processing based on the representation of the test audio signal and the representation of the reference audio signal. The DNN may yet further include an output stage for generating the indication of the listening score.
[0094] In some embodiments, the DNN may have been configured by training the DNN by, in a training epoch among a plurality of training epochs, inputting one or more training data items, each indicative of a respective value of the listening score. The DNN may have further been configured by, in the training epoch, determining respective indications of the listening score based on the one or more training data items. The DNN may have further been configured by, in the training epoch, determining respective loss values for the one or more training data items by evaluating a loss function, wherein the loss function depends on the indication of the listening score. The DNN may have yet further been configured by, in the training epoch, adjusting one or more internal parameters of the DNN based on the determined loss values.
[0095] In some embodiments, the method may be implemented at a different network node than the streaming client.
[0096] In some embodiments, the representation of the estimate of the test audio signal may relate to one or more Gammatone spectrograms.
[0097] In some embodiments, the representation of the reference audio signal may relate to one or more Gammatone spectrograms.
[0098] In some embodiments, the method may further include outputting the estimate of the audio quality of the test audio signal to a network node different from a network node associated with the streaming client. The network nodes may be network nodes in a cloud- based framework, for example.
[0099] In some embodiments, the estimate of the audio quality of the test audio signal may be output to a network node for performing encoding and / or packaging of the audio content. Then, the method may further include optimizing the encoding and / or packaging based on the estimate of the audio quality of the test audio signal.
[0100] In some embodiments, the method may further include determining an optimal number of quality levels in a bitrate ladder for distribution by a content delivery network, based on the estimate of the audio quality of the test audio signal. This may involve inputting the estimate of the audio quality of the test audio signal to a utility function used for determining the optimal number of quality levels, for example.
[0101] In some embodiments, the method may further include determining a configuration and / or set of coding tools based on the estimate of the audio quality of the test audio signal.
[0102] In some embodiments, the method may further include determining estimates of audio quality of test audio signal for streaming clients in each of a plurality of populations of streaming clients. The method may yet further include comparing the estimates of audio quality determined for the plurality of populations of streaming clients.
[0103] This may allow to compare different content delivery methods and / or playout methods for determining an optimal delivery method and / or playout method.
[0104] Another aspect of the disclosure relates to a method of providing playout- related information at a streaming client that processes audio content in an adaptive streaming environment. The method may include generating the playout-related information by one or more of: analyzing a playout buffer associated with the streaming client for determining bitrate information indicating a bitrate of segments of an audio signal played out by the streaming client; analyzing manifest information associated with the audio content; and analyzing characteristics of a playout device associated with the streaming clients. The method may further include outputting the playout-related information to a network node different from a network node associated with the streaming client.
[0105] Another aspect of the disclosure relates to an apparatus comprising a processor and a memory coupled to the processor, and storing instructions for the processor. The processor may be adapted to carry out the methods according to any one of the two preceding aspects and their embodiments.
[0106] Another aspect of the disclosure relates to a program comprising instructions that, when executed by a processor, cause the processor to carry out the methods according to the two aforementioned aspects and their embodiments.
[0107] Another aspect of the disclosure relates to a computer-readable storage medium storing the program of the preceding aspect.
[0108] It should be noted that the methods and systems including its preferred embodiments as outlined in the present disclosure may be used stand-alone or in combination with the other methods and systems disclosed in this document. Furthermore, all aspects of the methods, apparatus, and systems outlined in the present disclosure may be arbitrarilycombined. In particular, the features of the claims may be combined with one another in an arbitrary manner.
[0109] It will be appreciated that apparatus features and method steps may be interchanged in many ways. In particular, the details of the disclosed method(s) can be realized by the corresponding apparatus, and vice versa, as the skilled person will appreciate. Moreover, any of the above statements made with respect to the method(s) (and, e.g., their steps) are understood to likewise apply to the corresponding apparatus (and, e.g., their blocks, stages, units), and vice versa. Brief Description of the Drawings
[0110] The invention is explained below in an exemplary manner with reference to the accompanying drawings, wherein like numbers indicate like elements, and where
[0111] Fig.1A is a block diagram schematically illustrating an example of a framework with reference signal for evaluating playout performance in an adaptive streaming environment according to embodiments of the disclosure;
[0112] Fig.1B is a block diagram schematically illustrating an example of a reference- free framework for evaluating playout performance in an adaptive streaming environment according to embodiments of the disclosure;
[0113] Fig.2A is a flowchart illustrating an example of a method of evaluating playout performance in an adaptive streaming environment according to embodiments of the disclosure;
[0114] Fig.2B is a flowchart illustrating an example of a method of evaluating playout performance on an audio streaming device according to embodiments of the disclosure;
[0115] Fig.3A is a block diagram schematically illustrating playout analysis and generation of playout-related information at a streaming client according to embodiments of the disclosure;
[0116] Fig.3B is a block diagram schematically illustrating model selection for evaluating playout performance with an intrusive scheme in accordance with embodiments of the disclosure;
[0117] Fig.3C is a block diagram schematically illustrating model selection for evaluating playout performance with a non-intrusive scheme in accordance with embodiments of the disclosure;
[0118] Fig.4A and Fig.4B are flowcharts illustrating an example of a method of generating playout-related information in accordance with embodiments of the disclosure;
[0119] Fig.5A is a block diagram schematically illustrating an example of evaluation of playout performance using precomputed data with an intrusive scheme according to embodiments of the disclosure;
[0120] Fig.5B is a block diagram schematically illustrating an example of evaluation of playout performance using precomputed data with a non-intrusive scheme according to embodiments of the disclosure;
[0121] Fig.6 schematically illustrates an example of a graphical user interface for representing playout performance according to embodiments of the disclosure;
[0122] Fig.7 is a flowchart illustrating an example of using the evaluated playout performance for comparing populations of streaming clients according to embodiments of the disclosure;
[0123] Fig.8A is a block diagram schematically illustrating an example of a framework with reference signal for comparing different content delivery networks according to embodiments of the disclosure;
[0124] Fig.8B is a block diagram schematically illustrating an example of a reference-free framework for comparing different content delivery networks according to embodiments of the disclosure;
[0125] Fig.9A is a block diagram schematically illustrating an example of a framework with reference signal for comparing different content delivery methods or playout methods according to embodiments of the disclosure;
[0126] Fig.9B is a block diagram schematically illustrating an example of a reference-free framework for comparing different content delivery methods or playout methods according to embodiments of the disclosure;
[0127] Fig.10 is a block diagram schematically illustrating an example of a framework for optimizing encoding and / or packaging of content in an adaptive streaming environment according to embodiments of the disclosure;
[0128] Fig.11 is a flowchart illustrating an example of a method of optimizing encoding and / or packaging of content in an adaptive streaming environment according to embodiments of the disclosure;
[0129] Fig.12 is a block diagram schematically illustrating an example of configuring a DNN for estimating indications of subjective listening scores according to embodiments of the disclosure;
[0130] Fig.13 and Figs.14A, 14B are flowcharts illustrating examples of methods of configuring a DNN for estimating indications of subjective listening scores according to embodiments of the disclosure, with Fig.14A relating to an intrusive scheme and Fig.14B relating to a non-intrusive scheme;
[0131] Fig.15A is a flowchart illustrating an example of a method of estimating indications of subjective listening scores using a full-reference DNN according to embodiments of the disclosure;
[0132] Fig.15B is a flowchart illustrating an example of a method of estimating indications of subjective listening scores using a reference-free DNN according to embodiments of the disclosure;
[0133] Fig.16 to Fig.20 are diagrams illustrating examples of performance of trained or partially trained DNNs with reference according to embodiments of the disclosure;
[0134] Fig.21 shows an example training loss plot for training with different reference-free schemes and schemes with reference according to embodiments of the disclosure;
[0135] Fig.22 shows an example scaling of non-intrusive scheme scores with codec bitrates according to embodiments of the disclosure;
[0136] Fig.23 shows an example of predicted quality scores versus audio bandwidth according to embodiments of the disclosure;
[0137] Fig.24 is a block diagram showing an example conversion from intrusive schemes to non-intrusive schemes according to embodiments of the disclosure;
[0138] Fig.25 is a block diagram showing an example transfer of weights from intrusive schemes to non-intrusive schemes according to embodiments of the disclosure;
[0139] Fig.26 is a block diagram showing an example binaural transformation of multichannel audio signals according to embodiments of the disclosure; and
[0140] Fig.27 is a block diagram showing an example of an apparatus for performing methods or implementing DNNs according to embodiments of the disclosure. Detailed Description
[0141] The present disclosure relates to techniques for estimating audio quality of played-out content in an adaptive streaming environment and to techniques for configuring (e.g., training) and using DNNs for estimating audio quality, which will be described in turn.
[0142] The present disclosure further relates to techniques for configuring (e.g., training) and using reference-free DNNs for estimating audio quality and to techniques for providing estimated listening scores on audio streaming devices without any reference signals utilizing trained reference-free DNNs, which will be described in turn. MACHINE LISTENER-BASED AUDIO STREAMING QUALITY
[0143] Broadly speaking, part of the present disclosure relates to techniques (e.g., methods, apparatus, and systems) for performing objective quality testing of audio in the context of content streaming (e.g., adaptive streaming, in particular adaptive audio streaming), for example over the Internet. A system implementing such techniques may include a cloud- based service receiving playout-related information (e.g., playout-related parameters) from a client (streaming client) or a set of clients (i.e., population of clients) and then computing an objective quality score by emulating a subjective quality assessment test, for example according to the MUSHRA methodology (e.g., as standardized under ITU-R recommendation BS.1534).
[0144] The proposed techniques facilitate evaluation of audio quality from encoding to the actual client playout. They also allow to evaluate a population of clients in terms of the audio quality that is delivered to these clients.
[0145] Further, the proposed techniques can be used for monitoring of audio experience. The techniques can also be used to perform A / B testing (e.g., bucket testing orsplit-run testing) on real-world populations of streaming clients (e.g., for experiments comparing performance of bitrate ladders and / or codecs). The system can perform quality analysis in an online mode and an offline setting. In the online mode, the audio quality may be estimated in real-time according to the progress of streaming based on a feedback channel between a client and the service. In the offline setting, the service can first collect all parametric data from clients taking part in an experiment, and then perform the analysis according to that data.
[0146] In particular, the present disclosure proposes a method for performing objective quality testing of audio in the context of streaming over the Internet. With full- reference schemes, a system implementing a cloud-based service receiving playout-related parameters from a client (or a set of clients) and then computing an objective quality score by emulating subjective quality assessment tests according to the MUSHRA methodology was proposed. According to the present disclosure, emulation of the full-reference metric was done by recreating the test signal in the cloud-based service.
[0147] Notably, the proposed techniques may apply the so-called generative machine listener described later in the present disclosure. As will be explained in more detail later, the generative machine listener is a neural network (e.g., DNN) trained to evaluate audio by comparing it to a reference signal and providing the evaluation result for example as a probability distribution with a mean value corresponding to the MUSHRA score and a confidence value corresponding to a confidence interval expected in a listening test performed on such material. REFERENCE-FREE QUALITY ESTIMATION OF AUDIO
[0148] Further, part of the present disclosure relates to techniques (e.g., methods, apparatus, and systems) for performing objective quality testing of audio with non-intrusive or reference-schemes, and particularly relates to techniques for configuring reference-free DNNs based on pre-trained DNNs used in an intrusive scheme.
[0149] For example, the present disclosure provides an application of the reference- free generative machine listener. The reference-free generative listener is a neural network trained to evaluate audio without comparing it to the reference signal and providing the evaluation result as a probability distribution with a mean quality value (corresponding to, e.g., MUSHRA score) and a confidence value corresponding to a confidence interval typically expected in a listening test performed on such material.
[0150] The proposed methods and apparatuses further facilitate the evaluation of audio quality from encoding to the actual client playout. It allows for the evaluation of a population of clients in terms of the audio quality delivered to these clients. The method can be used for monitoring audio quality-of-experience (QoE).
[0151] The proposed methods and apparatuses may use a streaming client with an instrumented client that captures playout-related metadata. The metadata is used to recreate the test signal in real-time and fed into the “reference-free” machine listener.
[0152] The "reference-free" machine listener is preferably implemented as a service on the client that performs audio analysis on the fly. According to the present disclosure, it is as if a listener is placed at the streaming client and is instructed to evaluate the audio quality as streaming happens.
[0153] As playout happens, the instrumented player supplies playout-related metadata to the quality assessment service on the client. The service uses the machine listener to obtain a probability distribution describing the listening experience, and this information can be supplied wherever it is needed in the content distribution chain. The estimated quality score can be updated accordingly by analyzing the playout and the actual signal that gets to the client. Since the model used is probabilistic, a real-time update on the confidence of the assessment is obtained, which can capture more sophisticated aspects of the quality of experience.
[0154] Therefore, reference-free schemes according to the present disclosure enable (1) computing the quality of audio streaming, particularly with parameters from the player (client), and (2) predicting the quality with the "reference-free" generative machine listener and a confidence interval (CI); it also demonstrates how the quality score changes depending on events happening in the player and allows for performing quality prediction on the client device without the need to supply the reference signal to the client device.
[0155] Unlike in full-reference schemes, the proposed reference-free schemes are not based on a feedback channel between a client and the cloud service but on directly estimating the quality in the client and feeding back the estimated quality to the cloud service. Definitions
[0156] Intrusive quality assessment requires access to the reference signal and the test signal. Well-established subjective testing methodologies use this approach.
[0157] Non-intrusive quality methods require only access to the test signal.
[0158] Objective quality assessment algorithm facilitates estimation of quality of experience for human observers without using human observers. For example, a generative machine listener performs objective quality assessment by predicting the quality scores that would be achieved in subjective testing with human observers. In particular, the machine listeners for example facilitates estimation of mean performance scores along with the associated confidence intervals.
[0159] Adaptive streaming is a content delivery method where the content is available in multiple quality versions, which are associated with different bitrates. The higher the bitrate the higher the quality is. A content player includes a policy that attempts to determine the highest possible bitrate that results in delivery of segments in time before they are due to play out. In other words, the adaptive streaming policy attempts to maximize the Quality of Experience (e.g., in a setting where content segments are being downloaded, inserted into a playout buffer, and played out), while maintaining the probability of depletion of the playout buffer below some reasonably low threshold (i.e., the probability of rebuffering remains small).
[0160] Bitrate ladder is a set of versions of the content that are associated with different quality versions of the content. The bitrates of content in the bitrate ladder are designed to facilitate streaming over diverse throughput scenarios (e.g., very low to very high throughput). An adaptive streaming policy will select an appropriate quality level from the bitrate ladder on a per-segment basis. Information about the bitrate ladder is typically supplied to the content player in a so-called manifest. Description of Example Embodiments
[0161] Fig.1A depicts an example of a quality assessment service that incorporates a machine listener (e.g., as part of a quality assessment service) in a framework using intrusive schemes for adaptive streaming. The machine listener facilitates intrusive quality assessment of audio experience at the client, without a need to provide an uncoded reference to the client.
[0162] As noted above, using intrusive algorithms for quality assessment in (adaptive) streaming environments is hindered by the fact that streaming clients typically do not have access to reference signals. Providing the streaming clients with reference signals typicallywould require out-of-band placement of the reference signals, which is strongly disfavored by bandwidth limitations.
[0163] Embodiments of the present disclosure facilitate execution of an intrusive test for example within a cloud or network service, at a node (test node, network node) where the reference signal can be supplied. Instead of sending the client playout signal upstream to the test node, the test signal is reassembled at the test node based on lightweight playout-related metadata, which can be obtained by an instrumented client and then sent upstream to the service.
[0164] In the example of Fig.1A, the streaming client 10 receives coded audio content 5 (e.g., audio content or video content with associated audio content), for example via the Internet, from a Content Delivery Network (CDN) 105. The CDN 105 may provide different versions of given content, for example at different bitrates (e.g., using different settings within a predefined bitrate ladder), depending on streaming client configuration and / or network conditions, etc.. The streaming client 10 on the other hand may be configured to employ adaptive bitrate control to request content at different bitrates for maximizing playout quality and / or user experience.
[0165] After appropriate decoding, the streaming client 10 plays out the audio content, for example in a segment-by-segment manner, via a playout buffer. At the same time, the streaming client 10 performs playout analysis, for example via playout analysis block 120, to generate playout-related information 20 (e.g., playout-related metadata, or playout metadata). In this sense, the streaming client 10 acts as, implements, or comprises an instrumented client that collects and forwards playout-related information 20. The playout-related information 20 is provided to or is retrieved by a quality assessment service 150A (e.g., machine listener service). In general, the playout-related information 20 may be said to be provided to or retrieved by a test node. Further, an indication of audio content processed by the streaming client 10 is provided to or is retrieved by the quality assessment service 150A (or test node), to enable the quality assessment service 150A to generate a reference signal relating to the content processed by the streaming client 10, for intrusive quality assessment. Here, the indication of the audio content may comprise an identifier, such as a file name such as a file name (e.g., file name of an audio segment), etc. of the audio content, a bitrate level of the audio content, and / or information on a (current) segmentation of the audio content. The indication of audio content may be used to determine a sequence of audio segments received and played out by the streaming client 10. The audio content processed by the streamingclient 10 may be audio content received by the streaming client 10, for example from the CDN 105. The indication of the audio content may be obtained, for example, by intercepting a request of the streaming client 10 to the CDN 105.
[0166] The quality assessment service 150A may be in the form of a network service or cloud service. Further, the quality assessment service 150A may be configured for performing methods of evaluating playout performance in an adaptive streaming environment, such as method 200A described below, for example. To this end, the quality assessment service 150A may comprise a trained network 40A and a model selector 145A for selecting an appropriately trained model among a set of models based on the playout-related information 20. The quality assessment service 150A may further comprise a recreate test signal block 130 for estimating the test signal 30A, and a reference lookup block 160 for estimating a reference signal 60 for the test signal 30A.
[0167] An example of a method 200A of evaluating playout performance in an adaptive streaming environment (e.g., by the quality assessment service 150A of Fig.1A) is illustrated in the flowchart of Fig.2A. Playout performance may relate to (subjective) playout quality, for example. Method 200A comprises steps S210A through S230A as well as an optional step S240A. The method may be implemented at a different network node (e.g., test node) than the streaming client 10. Further, it may be implemented at a different network node than the CDN 105. Here, the network nodes may be network nodes in a cloud-based framework, for example.
[0168] At step S210A, playout-related information is obtained from the streaming client. The playout-related information may include, correspond to, or be in the form of metadata, for example.
[0169] At step S220A, a representation of the test audio signal is estimated based on the playout-related information. Here, the test audio signal is an audio signal played out by the streaming client. Estimating the representation of the test audio signal may involve or correspond to reconstructing the test audio signal or a representation thereof. The representation may relate to a set of features or spectrograms (e.g., Gammatone spectrograms) of the test audio signal, for example.
[0170] At step S230A, an estimate of an audio quality of the test audio signal is determined, using an audio quality assessment algorithm, based on the estimated representation of the test audio signal. The audio quality assessment algorithm may be anobjective audio quality assessment algorithm. Further, the audio quality assessment algorithm may emulate an intrusive audio quality test (e.g., listening test), such as a MUSHRA test, for example. It is understood that the estimated representation of the test audio signal is in a form suitable for input to the audio quality assessment algorithm. For example, the representation of the test audio signal may relate to a predefined number of segments to accommodate for requirements of the audio quality assessment algorithm relating to a time span covered by input audio signals or representations thereof.
[0171] At step S240A, which may be optional, the estimate of the audio quality of the test audio signal is output to a network node different from a network node associated with the streaming client. Non-limiting examples for using the estimate of the audio quality of the test audio signal will be described below with reference to Fig.6, Fig.7, Fig.8A, Fig.9A, Fig.10, and Fig.11.
[0172] Configured as described above, techniques according to the present disclosure use a quality assessment service (e.g., comprising a machine listener) as a cloud service for streaming quality evaluation. The quality assessment service (e.g., machine listener) is a component independent from the content delivery system and independent from the streaming client(s). The quality assessment service (e.g., machine listener) has the following properties: • It allows for performing intrusive quality testing of playout by the client device without need of supplying the reference signal to the client device. • It allows for performing objective quality evaluation that can emulate well established subjective testing methodology such as MUSHRA, for example. • By appropriate choice of the audio quality assessment algorithm, it can provide results as a probability distribution (e.g., MUSHRA score + confidence interval).
[0173] Potential applications and advantages of techniques according to embodiments of the present disclosure may include the following: • Techniques according to embodiments of the disclosure facilitate delivery with generic CDNs and allow to decouple the quality assessment system from the CDN infrastructure, which may be advantageous. Thus, CDN in general does not need to be involved in operating the machine listener service, and there is no need to store the reference signals in the CDN. The proposed techniques also facilitate multi-CDN delivery.• Techniques according to embodiments of the disclosure facilitate experimentation such as evaluation of scenarios where the quality estimates cannot be precomputed. One example of such scenarios is where the number of combinations of ABR ladder levels in a playout buffer and the number of distinct ways of playout (e.g., speaker in handheld device, headphone, discrete speakers) may be prohibitively large. • An example application of the system shown in the example of Fig.1A is AB testing of bitrate ladders or codecs in real-world content delivery scenarios (e.g., on real populations of clients over real content delivery situations).
[0174] Configured as described above, systems and methods according to embodiments of the disclosure comprise means / steps to allow an emulation of an intrusive listening test (e.g., MUSHRA test) by operating the (generative) machine listener in a cloud, where the test signal is reconstructed (or partially reconstructed) within the service by using a feedback channel from an instrumented client sending playout-related metadata. Further, these systems and methods entail selection of an appropriate model used by the (generative) machine listener from a collection of pretrained models, based on the playout-related metadata.
[0175] In addition to the above, method 200A may also comprise a step (not shown in Fig.2A) of generating a representation of a reference audio signal for the test audio signal, for use by the audio quality assessment algorithm. This may include obtaining (e.g., receiving) audio content or a representation thereof from a content repository (or content origin in general). It is understood that the representation of the reference audio signal should be in a form suitable for input to the audio quality assessment algorithm. For example, the representation of the reference audio signal may relate to a predefined number of segments to accommodate for requirements of the audio quality assessment algorithm relating to a time span covered by input audio signals or representations thereof.
[0176] In addition to the playout-related information 20, the quality assessment service 150A may further require an indication (e.g., identification) of the audio content processed by the streaming client 10. Thus, method 200A may further comprise a step (not shown in Fig.2A) of obtaining, from the streaming client 10, an indication of audio content processed by the streaming client 10. As noted above, the indication of the audio content may comprise an identifier, such as a file name, etc. of the audio content, a bitrate level of the audio content, and / or information on a segmentation of the audio content. The indication ofaudio content may be used to determine a sequence of audio segments received by the streaming client 10. The audio content processed by the streaming client 10 may be audio content received by the streaming client 10, for example from the CDN 105. Then, with the indication of the audio content available, estimating the representation of the test audio signal at step S220A may be further based on the indication of the audio content. Further, also the generation of the representation of the reference audio signal may be based on the indication of the audio content. This may comprise obtaining (e.g., receiving) audio content or a representation thereof from a content repository, based on the indication of the audio content that is played out by the streaming client.
[0177] If the playout-related information 20 comprises bitrate information indicating a bitrate of the audio signal played out by the streaming client 10, estimating the representation of the test audio signal at step S220A may be based on the bitrate information. This bitrate information may be provided for each of a plurality of segments of the audio signal played out by the streaming client 10.
[0178] An example process for estimating representations of the test signal and the reference signal will be described below with reference to Fig.5A.
[0179] Fig.1B depicts an example of a quality assessment service that incorporates a machine listener (e.g., as part of a quality assessment service) in a framework using non- intrusive schemes for adaptive streaming. For description relating to the same numerals in Fig.1B as in Fig.1A, reference is made to Fig.1A.
[0180] In particular, Fig.1B depicts the quality assessment service that incorporates the “reference-free” machine listener. The reference-free machine listener facilitates non- intrusive quality assessment of audio QoE at the client, without the need to provide an uncoded reference to the client.
[0181] The difference between the example in Fig.1B and that in Fig.1A is that the playout-related information 20 is provided to or is retrieved by a quality assessment service 150B (e.g., machine listener service) that does not require a reference signal. The reference- free machine listener service 150B is preferably provided on a client device playing out the audio signal being evaluated.
[0182] In the reference-free framework shown in Fig.1B, the test signal 30B generated by the recreate test signal block 130 is preferably configured to be input to a trainedreference-free network 40B. More preferably, the test signal 30B is indicative of a representation of the played out audio signal; in some cases, the test signal 30B may be indicative of a binaural transformation of the played out audio signal, wherein the played out audio signal relates to a multi-channel signal or an object-based signal.
[0183] Further, based on the played out metadata 20, or the metadata obtained based on the played out audio signal, a model selector 145B is configured to select an appropriately trained model or algorithm among a set of models or algorithms based on the playout-related information 20, so as to better reflect the result of the playout analysis, i.e., the information relating to the played out audio signal.
[0184] The reference-free trained network 40B is configured to output an audio quality score 50B preferably in the form of a probability distribution, based on the non- intrusive quality assessment.
[0185] The test signals 30A, 30B and the model selectors 145A, 145B in Figs.1A and 1B may be the same or different, depending on the specific configurations of the trained reference-free DNN.
[0186] A further example of a method 200B of evaluating playout performance in an adaptive streaming environment (e.g., by the non-intrusive quality assessment service 150B of Fig.1B) is illustrated in the flowchart of Fig.2B. Playout performance may relate to (subjective) playout quality, for example. Method 200B comprises steps S210B through S240B. The method may be implemented at a streaming client 10 that is playing out the audio signal being evaluated.
[0187] At step S210B, metadata relating to the played out audio signal is obtained. Playout-related information may include, correspond to, or be in the form of metadata, for example.
[0188] At step S220B, based on the obtained metadata, test data for being input into a trained reference-free deep neural network, DNN, is provided, for estimating the indication of the listening score. Here, the test data or the test audio signal relates to an audio signal played out by the streaming client, or is indicative of a representation of the audio signal played out by the streaming client. The representation may relate to a set of features or spectrograms (e.g., Gammatone spectrograms) of the audio signal, for example.
[0189] At step S230B, the test data is input into an input stage of the trained reference- free DNN. The trained reference-free DNN may include an audio quality assessment algorithm that may be an objective audio quality assessment algorithm. Further, the audio quality assessment algorithm may emulate a non-intrusive audio quality test (e.g., listening test), such as a MUSHRA test or MUSHRA-like test without reference signal, for example. It is understood that the estimated representation of the audio signal is in a form suitable for input to the audio quality assessment algorithm. For example, the representation of the audio signal may relate to a predefined number of segments to accommodate for requirements of the audio quality assessment algorithm relating to a time span covered by input audio signals or representations thereof.
[0190] At step S240B, a representation of the indication of the listening score is determined based on an output of an output stage of the trained reference-free DNN. Non- limiting examples for using the estimate of the audio quality of the test audio signal will be described below with reference to Fig.6, Fig.7, Fig.8B, Fig.9B, Fig.10, and Fig.11.
[0191] The above-proposed methods and apparatuses according to the present disclosure facilitate executing a non-intrusive test within a client. Unlike the methods for the intrusive schemes, playout-related metadata obtained by an instrumented client does not need to be sent upstream to recreate the test signal.
[0192] The present disclosure uses the reference-free machine listener directly as a client service for streaming quality evaluation as an independent component of a content delivery system. • It allows for performing non-intrusive testing on the client device without the need to supply the reference signal to the client device. • It allows for performing an objective quality evaluation that tries to emulate the well- established subjective testing methodology (e.g., MUSHRA or MUSHRA-like) without a reference signal. • It provides results as a probability distribution (e.g., MUSHRA score + confidence interval). Applications: • The present disclosure facilitates delivery with a generic CDN (the CDN does not need to be involved in operating the service, and there is no need to store the reference in theCDN). It also facilitates multi-CDN delivery. The main idea behind this invention is to decouple the quality assessment system from the CDN infrastructure, which may be advantageous. • The present disclosure facilitates experimentation, such as the evaluation of scenarios where the quality estimates cannot be precomputed. For example, where the number of combinations of ABR ladder levels in a playout buffer and the number of distinct ways of playout (e.g., speaker in a handheld device, headphone, discrete speakers) may be prohibitively large. • An example application of the system from Figure 1B is A / B testing of bitrate ladders or codecs in real-world content delivery scenarios (i.e., on real populations of clients over real content delivery situations).
[0193] The technical advantages of the system in Figure 1B comprises steps to allow an emulation of a MUSHRA (or MUSHRA-like) test by operating the reference-free generative machine listener in a client. According to another aspect of the present disclosure, the technical advantages comprise the selection of an appropriate model used by the reference-free generative machine listener from a collection of pre-trained models based on the playout-related metadata.
[0194] Fig.3A illustrates an example of a streaming client 10 that collects playout- related information (e.g., playout-related metadata), which is then aggregated for providing test signal and for model selection as explained above. Fig.3A thus relates to operation of the instrumented client implemented or comprised by the streaming client 10. The playout metadata comprises information on the composition of the player buffer (determined by the Playout Buffer Analysis), the composition of the currently used bitrate ladder (determined by Manifest Analysis), and the identification of the playout device (Playout Device Analysis).
[0195] The example illustrated in Fig.3A can be applied to intrusive schemes, wherein the metadata or the playout-related information is sent upstream to facilitate operation of the quality assessment service 150A (e.g., machine listener service) shown in Fig.1A.
[0196] The example illustrated in Fig.3A can also be applied to non-intrusive schemes in the quality assessment service 150B (e.g., machine listener service) shown in Fig. 1B. In this non-intrusive case, the metadata is not required to be sent upstream to facilitatecomputation of the quality assessment; instead, the assessment can be done directly in the client.
[0197] When being applied to non-intrusive schemes, the example illustrated in Fig. 3A comprises the operation of the instrumented client, wherein: • The scheme requires an instrumented client (which can be easily uploaded to the client device) but does not require any other operations (such as placing a reference on the client). • The playout analysis extracts playout-relevant information (e.g., the sequence of segments in the playout buffer, playout device, content of the manifest). • The metadata can be used to perform appropriate model selection and recreate the test signal, or features of the test signal (e.g., spectrogram). Given the test signals, upon selection of an appropriate trained model, a probability distribution representing the quality score may be computed in the client. Application of the example illustrated in Fig.3A applied to non-intrusive schemes: • The system may differentiate between different playout scenarios. For example, headphone playout may be in some cases more critical than speaker playout. This can be reflected in the quality score generated by a client. To achieve this, the client service performing the quality assessment may include a set of pretrained models. For example, there could be a model trained on listening test data from tests performed over headphones. There could be another model trained on listening tests performed over discrete speakers. Since the two playout scenarios generally differ in terms of how critical they are, it may be beneficial to use dedicated models for these scenarios, and then use an appropriate model to perform the evaluation of the quality.
[0198] An example of a corresponding method 400 of providing playout-related information at a streaming client that processes audio content in an adaptive streaming environment is shown in the flowchart of Fig.4A. Method 400, performed at the streaming client, comprises steps S410 and S420. Method 450 shown in the flowchart of Fig.4B relates to details of step S410. Method 450 comprises steps S460 through S480.
[0199] At step S410, the playout-related information is generated.
[0200] At step S420, the playout-related information is output to a network node (e.g., test node) different from a network node associated with the streaming client. For example, asexplained above, the playout related information may be output to the quality assessment service 150A or the quality assessment service 150B.
[0201] Method 400 may further comprise a step (not shown in the figure) of providing an indication of audio content processed by the streaming client, as described above in the context of Fig.1A and 1B.
[0202] Steps S460 through S480 of method 450 in Fig.4B relate to details and potential implementations of step S410 in method 400. It is understood that step S410 may comprise one or more, potentially all, of steps S460 through S480.
[0203] At step S460, a playout buffer associated with the streaming client is analyzed for determining bitrate information indicating a bitrate of segments of an audio signal played out by the streaming client. Accordingly, the bitrate information may indicate respective bitrates for each of a plurality of sequential segments of the played-out content. Analysis of the playout buffer may also yield further information on a composition of the playout buffer, in addition to the bitrate information.
[0204] Here, it is understood that the playout buffer typically contains a sequence of segments. A change of bitrate may occur on a per-segment basis, for example due to an action of the ABR policy operating on the streaming client.
[0205] At step S470, manifest information associated with the audio content is analyzed. Analyzing the manifest information may yield information on the currently used bitrate ladder.
[0206] At step S480, characteristics of a playout device associated with the streaming clients are analyzed. Characteristics of the playout device may relate to a device type and / or a type of reproduction system used, for example headphone playout or speaker playout.
[0207] Thus, returning to Fig.3A, playout analysis, for example by playout analysis block 120, may comprise one or more of playout buffer analysis (e.g., at Playout Buffer Analysis block 310), manifest analysis (e.g., at Manifest Analysis block 320), and playout device analysis (e.g., at Playout Device Analysis block 330).
[0208] Further, in accordance with the above, the playout-related information (e.g., metadata) may comprise information on a composition of a player buffer (e.g., determined by the Playout Buffer Analysis), composition of the currently used bitrate ladder (e.g.,determined by Manifest Analysis), and / or an identification of the playout device (e.g., determined by Playout Device Analysis).
[0209] Operation and properties of the instrumented client associated with the streaming client 10 can be briefly summarized as follows. • Techniques according to the present disclosure require an instrumented client (which can be easily uploaded to the client device), but do not require any other operations (such as placing a reference on the client, for example). • The playout analysis extracts playout-relevant information (e.g., the sequence of segments in the playout buffer, information on the playout device, information on the content of the manifest) and sends it upstream as metadata. • The metadata can be used to recreate the test signal, or features of the test signal (e.g., spectrogram, such as Gammatone spectrograms), and can be used to perform a look up for a relevant reference signal. Given the two signals (i.e., test signal and reference signal), upon selection of an appropriately trained model, an indication of the quality score (e.g., probability distribution representing the quality score) may be computed by the quality assessment service 150A or 150B.
[0210] Further, potential applications and advantages of the instrumented client and / or its operation may include the following: • The system may differentiate between different playout scenarios. For example, headphone playout may be in some cases more critical than speaker playout. This can be reflected in the quality score generated by a client. To achieve this, the cloud service performing the quality assessment may include a set of pretrained models. For example, there could be a model trained on listening test data from tests performed over headphones. There could be another model trained on listening tests performed over discrete speakers. Since the two playout scenarios generally differ in terms of how critical they are, it may be beneficial to use dedicated models for these scenarios, and then use an appropriate model to perform evaluation of playout quality.
[0211] Fig.3B is a block diagram schematically illustrating an example process of selection of a model (from a set of pretrained models) that can be used by the quality assessment service 150A (e.g., machine listener). The model selection is based on the playout-related information 20 sent upstream by the instrumented client 10.
[0212] For example, different models may have been trained based on different training data, relating to respective different use cases. In one embodiment, different models may have been trained for different device characteristics, for example for different device type and / or different reproduction systems (e.g., headphones or speakers).
[0213] Model selection may be performed at Model Selection block 145A. Selection may be made from a collection of models 340A, comprising individual models 350-1, 350-2, 350-3, … . Each of these models may relate to a machine listener or full-reference DNN trained for audio quality assessment under specific circumstances. For example, the models in the collection of models 340A may have been trained for different device characteristics, as described above.
[0214] Fig.3C is a block diagram schematically illustrating an example process of selection of a model (from a set of pretrained models) that can be used by the quality assessment service 150B (e.g., machine listener). The model selection is based on the playout- related information 20 obtained at / on / in the instrumented client 10.
[0215] Similar to Fig.3B, in Fig.3C, for example, different models may have been trained based on different training data, relating to respective different use cases. In one embodiment, different models may have been trained for different device characteristics, for example for different device type and / or different reproduction systems (e.g., headphones or speakers).
[0216] Model selection may be performed at Model Selection block 145B. Selection may be made from a collection of models 340B, comprising individual models 360-1, 360-2, 360-3, … . Each of these models may relate to a machine listener or reference-free DNN trained for audio quality assessment under specific circumstances. For example, the models in the collection of models 340B may have been trained for different device characteristics, as described above.
[0217] Returning to method 200A and 200B of Fig.2A and Fig.2B, the audio quality assessment algorithm employed by this method may use a set of pretrained models for audio quality assessment. Then, at step S230A and at step S240B, determining (e.g., generating) the estimate of the audio quality may comprise selecting a pretrained model among the set of pretrained models based on the playout-related information.
[0218] For instance, as explained above, the playout related information 20 may comprise information relating to a playout device associated with the streaming client. Then, the pretrained model may be selected based on the information relating to the playout device. The information relating to the playout device may include an indication of the playout device (e.g., headphones, soundbar, discrete speakers, etc.) and / or an indication of characteristics of the playout conditions (e.g., SNR, etc.).
[0219] In some embodiments, the quality assessment service may relate to a generative machine listener (e.g., stereo generative machine listener), implemented by a respective (intrusive or non-intrusive) DNN, using (e.g., as the aforementioned audio quality assessment algorithm) the algorithm described below under section MACHINE LISTENER. This algorithm may operate on spectrograms (e.g., Gammatone spectrograms) computed from or for the test and reference signals, instead of operating directly on the waveforms. This means that both the test and the reference signals (when present) can be assembled from precomputed blocks (e.g., precomputed blocks of spectrograms), which can reduce cloud storage requirements and computational load. This is of particular advantage from the point of view reducing the cost of running the machine listener in a cloud or directly in a client device.
[0220] The spectrograms (or other audio features) may be computed on per-segment basis according to the segmentation introduced by the transport mechanism that is used by the content delivery system (e.g., by CDN 105).
[0221] Thus, the (stereo) generative listener may operate on the Gammatone spectrograms of left, right, mid, and side signals of reference and coded stereo signals (e.g., as described in [5]). Gammatone filters are a popular approximation to the filtering performed by the ear. Gammatone-based spectrogram can thus be considered as a more perceptually motivated representation than the traditional spectrogram. The Gammatone spectrograms of the audio signal may be calculated for example with a window size of 80 ms, hop size of 20 ms, and 32 frequency bands ranging from 50 Hz up to 24 kHz. The resulting Gammatone spectrograms of short segments (e.g., 1 second of signal in ABR ladder) of reference and coded signals can be precomputed, paired and stacked along channel dimension, which results in an input size of 8×32×50 (channels×bands×time-frame) to the neural network, for example.
[0222] Thus, in some embodiments, the aforementioned representation of the estimate of the test audio signal and the aforementioned representation of the reference audio signalmay each relate to one or more Gammatone spectrograms (e.g., Left (L), Right (R), Mid (M), and Side (S) spectrograms).
[0223] However, it is understood that techniques (e.g., methods and apparatus) according to the present disclosure are not limited to using spectrograms (e.g., Gammatone spectrograms), but that these techniques may likewise operate on waveforms or other audio features. Also, spectrograms other than Gammatone spectrograms may be used for this purpose, such as other perceptually motivated spectrograms. Nevertheless, for conciseness of presentation, without intended limitation, reference will be made in the following to spectrograms, in particular, Gammatone spectrograms, instead of generic spectrograms or waveforms.
[0224] Fig.5A schematically illustrates the processes of assembling the reference signal and reconstructing the test signal based on playout-related information 20 (e.g., metadata) that is sent upstream to the quality assessment service 150A / 540A by the instrumented client 10, i.e., in the intrusive framework as shown in Fig.1A. In some embodiments, the data flow process used to feed the quality assessment model (e.g., generative listener model, such as trained network 40A in Fig.1A) may include several precomputed steps.
[0225] Precomputed spectrograms (e.g., Gammatone spectrograms) are held, on a per segment basis, in reference repository 580. Based on the indication of the (segments of the) audio content played out by the streaming client 10 (e.g., ID and segmentation of content item played out), the reference repository 580 is queried by Reference Lookup block 585 for assembly of the reference signal, segment by segment, at Assemble Reference block 590. The assembled reference signal 560 is provided to the audio quality assessment algorithm, such as the machine listener, for audio quality assessment at Machine Listener Analysis block 540A.
[0226] Further, the playout-related information 20, together with the indication of the (segments of the) audio content played out by the streaming client 10 (e.g., ID and segmentation of content item played out) is provided to Assemble Test Signal block 575. Based on the playout-related information 20 (e.g., bitrate, codec config), a content repository 570 is queried to assemble the test signal, again segment by segment. This yields the aforementioned representation of the test audio signal 530, for input to the audio quality assessment algorithm, such as the machine listener, for audio quality assessment at MachineListener Analysis block 540A. Assembly of the representation of the test audio signal 530 may correspond to step S220A of method 200A, for example.
[0227] Although not shown in Fig.5A, it is understood that the playout-related information 20 may also be provided to Machine Listener Analysis block 540A, for model selection as described above.
[0228] Based on the assembled representation of the test audio signal 530 and the assembled reference signal 560, the audio quality assessment algorithm can generate the estimate of the audio quality of the test audio signal, as explained above with reference to step S230A of method 200A.
[0229] Fig.5B schematically illustrates the processes of reconstructing the test signal based on playout-related information 20 (e.g., metadata) that is provided to the quality assessment service 150B / 540B on or at or in the instrumented client 10, i.e., in the non- intrusive framework as shown in Fig.1B. In some embodiments, the data flow process used to feed the quality assessment model (e.g., generative listener model, such as trained network 40B in Fig.1B) may include several precomputed steps.
[0230] The difference between the example in Fig.5B and that in Fig.5A is the application of a non-intrusive or an intrusive machine listener analysis. Therefore, for the same numerals in Fig.5B, reference is made to Fig.5A. Compared to Fig.5A, Fig.5B does not require assembling of the reference signal.
[0231] Fig.6 is a non-limiting example of a graphical user interface showing the results provided by techniques according to embodiments of the disclosure. The GUI comprises indicators / selectors of available levels of the bitrate ladder 610 through 650, as well as an indication 660 of the audio score for the test signal. This indication 660 may comprise, for example, a mean and a confidence interval for a subjective listening score, such as a MUSHRA or a MUSHRA-like quality or CI score, for example. Downstream Applications for Estimates of Audio Quality
[0232] Techniques according to the present disclosure can operate in an online and an offline setting, both in the intrusive framework and in the non-intrusive framework.
[0233] In the online setting, the quality analysis may be performed on the fly and the performance score (e.g., the estimate of audio quality) can be distributed wherever it is needed in the delivery system.
[0234] The offline setting comprises aggregation of the playout-related information (e.g., playout metadata) from multiple streaming clients (e.g., two sets of clients used for AB testing). The test signals can be constructed in an offline manner (e.g., after the experiment is finished) from the collected playout-related information. The quality assessment service (e.g., machine listener) can then perform quality analysis offline, providing the performance statistics for the clients (or sets of clients).
[0235] In general, regardless of whether the online setting or the offline setting applies, techniques according to the present disclosure can be used for comparing streaming clients in different populations of streaming clients, or for comparing different populations of streaming clients.
[0236] An example of a corresponding method 700 is schematically illustrated in the flowchart of Fig.7. Method 700 comprises steps S710 and S720 and may be performed subsequent to or in conjunction with methods 200A and 200B described above.
[0237] At step S710, estimates of audio quality of test audio signals for streaming clients in each of a plurality of populations of streaming clients are determined. This may be done as described above in the context of methods 200A and 200B. In particular, the step comprises estimating indications of respective listening scores for a plurality of audio streaming device clusters, each cluster corresponding to a respective content delivery infrastructure, and each cluster comprising one or more audio streaming devices.
[0238] At step S720, the estimates of audio quality determined for the plurality of populations of streaming clients are compared to each other. In general, the determined estimates are analyzed. In particular, the step comprises estimating for each cluster an indication of a listening score based on metadata collected from the one or more audio streaming devices included in the each cluster, and comparing the estimated listening scores of the clusters.
[0239] Examples, application, and use cases of such comparison are schematically illustrated in the block diagrams of Figs.8A, 8B and Figs.9A, 9B. In these Figures, the same reference numerals refer to the same configurations / operations.
[0240] In the example of Fig.8A, a content server 850 (content origin) provides respective (audio) content 852, 854 to first and second content delivery networks CDN1, 830, and CDN2, 840. The first CDN 830 provides content 835 to a first population 810 of streaming clients 815. The second CDN 840 provides content 845 to a second population 820 of streamlining clients 825. Streaming clients of both populations 810, 820 provide respective sets of playout-related information 870, 880 to quality assessment service 860A, which determines respective estimates of audio quality 865 (or estimates of playout performance in general) for the populations 810, 820, based on respective sets of playout-related information 870, 880, by techniques as set out above. Determining respective estimates of audio quality 865 may require receiving the downloaded segments 855 (provided to the streaming clients) or a representation thereof from the content server 850. Comparing the estimates of audio quality 865 for the two populations 810, 820 allows to infer information about different performances of the different CDNs 830, 840, for example. This information may be used to optimize content delivery to the populations of streaming clients.
[0241] Referring to the example in Fig.8B, only the differences to that in Fig.8A are explained. While Fig.8A relates to an intrusive scheme performed by intrusive quality assessment service 860A, Fig.8B relates to a non-intrusive scheme performed by non- intrusive quality assessment service 860B.
[0242] In the example of Fig.9A, a content server 950 (content origin) provides (audio) content 952 to a CDN 910 which provides content 932 to a first population 910 of streaming clients 915 and provides content 934 to a second population 920 of streamlining clients 925. Streaming clients of both populations 910, 920 provide respective sets of playout- related information 970, 980 to quality assessment service 960A, which determines respective estimates of audio quality 965 (or estimates of playout performance in general) for the populations 910, 920, based on respective sets of playout-related information 970, 980, by techniques as set out above. Determining respective estimates of audio quality 965 may require receiving the downloaded segments 955 (provided to the streaming clients) or a representation thereof from the content server 950. Comparing the estimates of audio quality 965 for the two populations 910, 920 of streaming clients allows to infer information about different performances of the different populations, for example in cases is which different delivery methods and / or playout methods are employed for / by the different populations. This information may be used to optimize content delivery to the populations of streaming clients and / or playout by the populations of streaming clients.
[0243] Referring to the example in Fig.9B, only the differences to that in Fig.9A are explained. While Fig.9A relates to an intrusive scheme performed by intrusive quality assessment service 960A, Fig.9B relates to a non-intrusive scheme performed by non- intrusive quality assessment service 960B.
[0244] As a further application or use case, the proposed techniques may be used to provide a feedback loop for edge processing of content. For example, if the content encoder is located at the edge of the network, the estimated audio quality (e.g., MUSHRA or MUSHRA- like score) could be used to fine-tune the bitrate ladder used to deliver audio content to a client.
[0245] An example of a framework involving such feedback loop is schematically illustrated in Fig.10. A content server 1050 (content origin) provides (audio) content to an Encode and Packaging Coordination Engine 1090 (encoding and packaging engine) that encodes and packages content in accordance with a set of one or more rules, and provides the encoded and packaged content to a Point of Presence (PoP) in CDN 1030 (or plural CDNs).
[0246] The rules employed by the encoding and packaging engine 1090 may relate to maximizing a value function and / or minimizing a cost function, for example. For example, the one or more rules may relate to minimizing a number of levels in the bitrate ladder while optimizing an average (worst case) performance, and / or determining optimal bitrates for maximizing average (worst case) performance.
[0247] The CDN 1030 provides (audio) content to a population / cluster 1010 of streaming clients 1015, which in turn provide playout-related information 1070 to a quality assessment service 1060. The quality assessment service 1060 determines a playout performance for the population 1010 of streaming clients 1015 (e.g., an average playout performance, a worst-case playout performance, etc.) by techniques as set out above, and provides an indication of the determined playout performance to the encoding and packaging engine 1090.
[0248] In line with techniques described above, the playout performance may relate to the aforementioned estimates of audio quality, or a quantity derived therefrom. The quality assessment service 1060 may be based on an intrusive scheme or a non-intrusive scheme.
[0249] Thus, in general, step S240A and Step 240B of method 200A and method 200B described above may comprise or relate to outputting the determined estimate of theaudio quality of the test audio signal (e.g., as a playout performance) to a network node for performing encoding and / or packaging of the audio content (e.g., the aforementioned encoding and packaging engine 1090).
[0250] Based on the playout performance 1065, the encoding and packaging engine 1090 in the example of Fig.11 can then optimize encoding and / or packaging of the content.
[0251] For example, in a framework as shown in the example of Fig.10, one or more of steps S1110 to S1130 of method 1100 shown in the flowchart of Fig.11 may be performed.
[0252] Step S1110 relates to or comprises optimizing the encoding and / or packaging based on the estimate of the audio quality of the test audio signal.
[0253] Step S1120 relates to or comprises determining an optimal number of quality levels in a bitrate ladder for distribution by a content delivery network, based on the estimate of the audio quality of the test audio signal. This may involve inputting the estimate of the audio quality of the test audio signal to a utility function (e.g., value function and / or cost function) used for determining the optimal number of quality levels, for example.
[0254] Step S1130 relates to or comprises determining a configuration and / or set of coding tools based on the estimate of the audio quality of the test audio signal. Apparatus for Implementing Methods According to the Disclosure
[0255] Finally, while reference above may be mainly made to methods according to the present disclosure, the present disclosure likewise relates to apparatus (e.g., computer- implemented apparatus) for performing methods and techniques described throughout the present disclosure. An example of such apparatus 2100, which is described in more detail below, is schematically illustrated in Fig.27. Such an apparatus may implement, for example, the quality assessment service 150A, 150B, or the (instrumented) streaming client 10. The apparatus (e.g., processor 2110 thereof) may receive, among others, suitable input data (e.g., playout-related information or an indication of audio content processed by a streaming client), depending on use cases and / or implementations. The apparatus 2100 (e.g., processor 2110 thereof) may be adapted to carry out the methods / techniques described throughout the present disclosure (e.g., method 200A of Fig.2A, method 200B of Fig.2B, method 400 of Fig.4A, method 450 of Fig.4B, method 700 of Fig.7, and / or method 1100 of Fig.11) and to generatecorresponding output data 1240 (e.g., an estimate of an audio quality), depending on use cases and / or implementations.
[0256] The present disclosure likewise relates to corresponding computer programs and computer-readable storage media. MACHINE LISTENER
[0257] One non-limiting example for implementing the quality assessment service described above, or the algorithm employed by the quality assessment service, is a so-called machine listener (e.g., generative machine listener) as described in the following. A machine listener (e.g., generative machine listener) may be an intrusive machine listener or a non- intrusive machine listener, wherein the latter is reference-free. In the following, when not specifying to an intrusive or a non-intrusive scheme, the description applies to both schemes.
[0258] Broadly speaking, an intrusive machine listener (e.g., generative machine listener) according to embodiments of this disclosure is a neural network (e.g., deep neural network) trained to evaluate audio by comparing it to the relevant reference signal and providing the evaluation result, for example as a probability distribution of predicted listener scores in the currency of subjective listening tests such as Multi-Stimulus Tests with Hidden Reference and Anchor (MUSHRA).
[0259] A non-intrusive or reference-free machine listener (e.g., reference-free generative machine listener) according to embodiments of this disclosure is a reference-free neural network (e.g., deep neural network) trained to evaluate audio without requiring a relevant reference signal and providing the evaluation result, for example as a probability distribution of predicted listener scores in the currency of subjective listening tests such as MUSHRA tests or MUSHRA-like tests which are similar to MUSHRA tests but without presenting (or exposing the known) reference to the subjects.
[0260] In general, listener scores achieved for example in MUSHRA tests can be predicted by an intrusive system that takes the signal under test and the reference signal as inputs. Examples of such systems including neural networks are given in [5], which is hereby included by reference in its entirety.
[0261] For a given pair of input signals, there typically will be a certain degree of variability in the listener scores perceived by different listeners. It has been found that there is value in capturing that aspect of the data for automated estimation of quality of experience inentertainment delivery systems. In an actual subjective MUSHRA test, the mean and standard deviation of the listener scores obtained from different listeners can be computed, and the standard deviation can then be converted to a confidence interval given the number of listeners and a statistical model.
[0262] It has been found by the inventors to be challenging to train a neural network to directly output both mean values and confidence intervals. One feasible alternative for quantifying prediction variability may be to use a bootstrapping approach by training multiple models with the mean scores as objectives on randomly sampled subsets of the data. Each of the trained models will then introduce a slightly different prediction. The variability of these predictions then quantifies the level of confidence. However, the disadvantages of the bootstrapping method are twofold. First, using the bootstrapping method implies high complexity: one needs as many models as the number of listeners to be simulated. Second, there is a risk of modeling prediction variability rather than listener score distribution.
[0263] Another feasible alternative for quantifying prediction variability may be to consider separate modeling of the bias (see [1], [2]) or bias and inconsistency (see [3], [4]) of individual listeners across the signals under test. The main application of the latter method is prescreening of listeners with outlier behavior. A common inconvenience for these methods is that one needs to keep track of the listener identity in the dataset.
[0264] The present disclosure seeks to provide improved techniques for quantifying prediction variability of subjective listening tests. At the application stage (i.e., at inference) the trained model (e.g., generative model) according to the present disclosure provides a distribution of scores from which it is easy to extract mean scores as well as standard deviations and / or confidence intervals (CI) for any number of listeners.
[0265] At the training stage, unlike existing approaches where mean subjective scores are used as a target for training, the model according to the present disclosure utilizes individual listener scores. This has been found to simplify the preprocessing and allows the dataset to influence the training in direct proportion to the human effort for listening test data with varying number of listeners. Further, the maximum likelihood principle may be used for parameter estimation.
[0266] Techniques according to the present invention have been found to have the following advantages. The trained model (e.g., generative model) reaches similar performance as typical non-generative models in predicting the mean, but is also capable of predicting theconfidence interval. Further, the trained model is more robust to conditions unseen in typical listening tests.
[0267] The present disclosure further seeks to provide improved techniques for evaluating audio quality when reference signals are not present or provided, for instance when a client device directly performs evaluation of the played out audio signal.
[0268] Generative Machine Listener (GML) has shown in experiments and practice its high accuracy and robustness in evaluating the decoded mono / stereo / binaural signals regardless of test conditions (codecs, bitrates, etc.) – but it requires corresponding reference signals. The structure of InSE-Net (on which GML is based on) is generic and could be easily modified as a reference-free scheme by simply removing the reference signals fed to the input layers. The learned weights from the GML with reference signals (i.e., intrusive GML model), could be wisely selected, and transferred to the reference-free case according to the present disclosure.
[0269] For example, according to the present disclosure, the reference-free Generative Machine Listener (rf-GML) is a robust tool to evaluate 48 kHz mono / stereo / binaural audio without requiring any form of reference signals. rf-GML according to the present disclosure may inherit: (1) the structure of InSE-Net and GML and (2) the training scheme of GML which required the reference signals (i.e., intrusive or full-reference GML). Non-intrusive schemes provide an objective audio quality tool to assess general audio at higher sample rates. The core idea of non-intrusive schemes e.g., rf-GML is to transfer learned weights from its counterpart with reference, which makes retraining and maintenance of rf-GML more easy and more efficient. The performance of rf-GML could be boosted by any new updates of the intrusive GML.
[0270] The rf-GML: (1) utilizes the pre-trained weights; (2) is re-trained with the same data without corresponding reference; and (3) is re-trained only for a few iterations (or epochs), to adapt to reference-less training signals. The best model has been selected according to its performance on validation sets and has later been tested on test datasets. So far rf-GML has shown very competitive results on speech-only signals as SESQA and significantly better results than SQUIM from Meta. On mono signals at low bitrates, the proposed model is around 7% worse than SESQA, however, on stereo signals at low bitrates the proposed model has a close tie with SESQA, and on stereo signals at high bitrates the proposed model has around 15% to 20% performance boost over SESQA. Note that SESQAis a mono model and was trained for mono speech-only signals; and SQUIM for speech-only signals as well. Whereas, rf-GML is trained on 48 kHz stereo / binaural general audio and has never seen mono signals in its training, which has already shown its superiority and versatility in the algorithm. For general audio, rf-GML on average achieves 0.7 to 0.8 Spearman (rank) correlation and 0.78 to 0.90 Pearson (linear) correlation.
[0271] Note also that the rf-GML could easily be extended to evaluate multichannel and spatial audio signals.
[0272] Fig.12 shows a comparison between a conventional model (e.g., DNN) for predicting a mean listening score (e.g., MUSHRA score) and a model (e.g., DNN implementing the generative machine listener) according to embodiments of the disclosure. The conventional model (non-generative approach) shown on the left-hand side, given ref- coded audio (or features thereof, such as Gammatone spectrograms) predicts a mean subjective listening score (e.g., MUSHRA score). On the other hand, the generative machine listener model shown on the right-hand side provides a distribution of listening scores (e.g., MUSHRA scores). At training, the present disclosure proposes to utilize individual listening scores. Without intended limitation, the architecture of the DNN implementing the generative machine listener may be the one described in [5], with the difference that the output stage is configured to provide for more than one output, for example suitable for representing the distribution of listening scores (e.g., MUSHRA scores). Description of Example Embodiments
[0273] Given an original signal - and a signal under test / the generative listener model (or the DNN implementing same) gives an indication of a probability distribution of (subjective) listening scores (e.g., MUSHRA scores) ^ for / , for example as a parametrized probability density *+^^|-, / ^(1) The parameters 0 of the model are trained by the maximum likelihood principle. One can then simulate a listening test with 1 listeners by sampling the model 1 times. The negative log likelihood (NLL) loss, used for training of the model parameters 0 for a listener score value ^in the dataset may be −log *+^^|-, / ^.
[0274] In general, at the training stage, the input is given by training data items, each indicative of a respective value of the listening score ^. The loss function depends on the indication of the listening score ^. The training data items may each be further indicative of a representation of the audio signal (signal under test / ) and a representation of the reference audio signal (original signal -) for the audio signal. Here, the representation of the audio signal / and the representation of the reference audio signal - may relate to Gammatone spectrograms, for example. Each training data item may be obtained by performing a standardized listening test (e.g., MUSHRA test) for a test signal / and a corresponding reference signal -, yielding the listening score ^. Performing such test multiple times, for example with different listeners, will yield multiple training data items, tentatively denoted as^^, / , -^. Test signals / and corresponding reference signals - may be obtained from suitableaudio libraries, for example. Here and in the remainder of the disclosure, it is understood that the reference signal corresponds to an uncoded signal. The test audio signal corresponds to a coded audio signal (e.g., a signal obtained after encoding, decoding, and if necessary, time alignment with the reference signal for compensating coding delays).
[0275] If the representation of the test signal / and the reference signal - relate to Gammatone spectrograms (typically, L, R, M, and S spectrograms), each actual sound signal may yield 4 training data items (e.g., one per L, R, M, and S spectrogram).
[0276] An example of a method 1300 of configuring (e.g., training) a DNN for estimating an indication of a subjective listening score for an audio signal is illustrated by the flowchart of Fig.13. Method 1300 comprises steps S1310 and S1320.
[0277] It is understood that the DNN implements the model (e.g., generative model) under consideration. For example, the DNN may implement the aforementioned generative machine listener, non-intrusive or intrusive.
[0278] Further, the listening score to be estimated or predicted by the DNN may be a score according to a predefined (e.g., standardized) listening test. The listening test may apply a predefined test metric and / or test scenario. One example of such listening test is a MUSHRA listening test.
[0279] At step S1310, an output stage of the DNN to generate the indication of the listening score is provided.
[0280] At step S1320, the DNN is trained, in (at least) a training epoch among a plurality of training epochs.
[0281] The DNN may be DNN utilizing reference (preferably full-reference) or a reference-free DNN. DNN utilizing reference
[0282] Method 1400A as illustrated by the flowchart of Fig.14A is an example of a possible implementation of training a full-reference DNN, in a training epoch among the plurality of training epochs. Method 1400A comprises steps S1410A, S1420, S1430 and S1440.
[0283] At step S1410A, one or more training data items are input. For example, a mini batch of training data items (e.g., 8 training data items) may be input. As described above, each training data item is indicative of a respective value of the listening score ^. As also described above, each training data item may be further indicative of a representation of the audio signal (signal under test / ) and a representation of the reference audio signal (original signal -) for the audio signal.
[0284] At step S1420, respective indications of the listening score are determined based on the one or more training data items. These indications may be determined, for example, based on the representation of the audio signal and the representation of the reference audio signal.
[0285] In some embodiments, the indication of the listening score may relate to a probability distribution (e.g., probability density function) of the listening score, with the output stage being adapted for generating the probability distribution of the listening score.This probability distribution (e.g., *+^^|-, / ^ as per Eq. (1)) may emulate listening scoresobtained by a plurality of (independent) listening tests for the audio signal. Further, the probability distribution may be parameterized by two or more parameters of the probability distribution. Examples for possible parameterizations of the probability distribution will be described below.
[0286] If the indication of the listening score relates to the probability distribution of the listening score, determining respective indications of the listening score based on the one or more training data items may comprise determining respective parameters of the probability distribution based on the one or more training data items. Here, determining theparameters of the probability distribution may be based at least in part on the value of the subjective listening score included in the training data item. Further, determining the parameters of the probability distribution may be based on a current state of the full-reference DNN, for example the current values of internal parameters of the full-reference DNN.
[0287] At step S1430, respective loss values for the one or more training data items are determined by evaluating a loss function. This loss function depends on the indication of the listening score.
[0288] In some embodiments, if the indication of the listening score relates to the probability distribution of the listening score, the loss function may depend on the parameters of the distribution. Examples of loss functions will be described below.
[0289] At step S1440, one or more internal parameters of the full-reference DNN are adjusted based on the determined loss values, for example by using well-known regression and back-propagation techniques. The internal parameters of the full-reference DNN may be model parameters, for example, such as coefficients (e.g., filter coefficients) of a plurality of layers of the full-reference DNN.
[0290] If multiple training data items are input per training epoch, adjusting the internal parameters may be based on an aggregate of the loss values for the training data items, such as a mean or average thereof, for example.
[0291] As noted above, training the full-reference DNN, for example via method 1400, may be based on the maximum likelihood principle. Accordingly, the loss function employed for the training (e.g., the loss function evaluated at step S1430 of method 1400) my relate to a negative log likelihood (NLL) loss. As also noted above, the negative log likelihood loss may be given by ^'(( = − log *+^^|-, / ^,(2)where *+^^|-, / ^ is the probability density function (as an example of a probabilitydistribution) for test score ^ given a representation of the audio signal / and a representation of a reference audio signal - for the audio signal / , and 0 indicates the internal parameters of the DNN.
[0292] A first non-limiting example of the probability density function *+^^|-, / ^ isthat of a Gaussian distribution parametrized by mean ^ and variance ^^. The NLL loss may then be given by, for example ^^ ^^ − ^^^23455 = ^ log 2^ + log ^ +2^^,(3)with a parametrization via ^ = ^+^-, / ^ and log ^ = 6+^-, / ^ for training.
[0293] Thus, the probability distribution at step S1420 of method 1400 may relate to a Gaussian distribution parameterized by a mean ^ and a variance ^^. Then, the loss function ^^ may be given by ^ = log ^ + ^^^^^^^^^^ ^^^^^ ^^^ + ^, where ^ is a constant and ^ is thesubjective ^ listening score. The constant ^ may be given by ^ = log 2^, for example, in line^with Eq. (3).
[0294] A second non-limiting example of the probability density function * ^^|-, / ^+is that of a logistic distribution parametrized by mean ^ and scale ^. The NLL loss in this case may be given by, for example ^= log 4 + log ^ + 2 log sechwith a parametrization via ^ = ^ ^-, / ^ and log ^ = 6 ^-, / ^ for training.+ +
[0295] Thus, the probability distribution at step S1420 of method 1400 may relate to a logistic distribution parameterized by a mean ^ and a scale ^. Then, the loss function ^^^^^^^^^may be given by ^ = log ^ + 2 log sech + ^, where ^ is a constant and ^ is the $subjective listening score. The constant ^ may be given by ^ = log 4, for example, in linewith Eq. (4).
[0296] Models with more than two parameters, such as mixtures of Gaussians or logistics, or even categorical distributions, may have the capability to model multimodal listener score distributions. On the other hand, there is a potential drawback of requiring more data for successful training.
[0297] At inference, an estimate of an indication of a subjective listening score for an audio signal can be determined using an appropriately trained full-reference DNN, forexample a full-reference DNN trained as described above. As above, the listening score is assumed to be a score according to a predefined listening test. Further, the full-reference DNN is in general assumed to comprise an input stage for receiving a representation of the audio signal and a representation of a reference audio signal for the audio signal, a plurality of layers for performing processing based on the representation of the audio signal and the representation of the reference audio signal, and an output stage, coupled to a last one of the plurality of layers, for generating the indication of the listening score. Here, processing by the plurality of layers may be further based on a current state of the full-reference DNN, for example the current values of internal parameters of the full-reference DNN.
[0298] An example of a corresponding method 1500A using this full-reference DNN is illustrated in the flowchart of Fig.15A. Method 1500A comprises steps S1510A and S1520.
[0299] At step S1510A, the representation of the audio signal and the representation of the reference audio signal are input to the input stage of the full-reference DNN.
[0300] At step S1520, a representation of the indication of the listening score is determined based on an output of the output stage of the full-reference DNN.
[0301] As above, the indication of the listening score may relate to a probability distribution of the listening score, with the output stage of the full-reference DNN being adapted for generating the probability distribution of the listening score. This probability distribution may be seen as emulating listening scores obtained by a plurality of listening tests for the audio signal. Further, the probability distribution may be parameterized by two or more parameters of the probability distribution. Thus, the representation of the indication of the listening score determined at step S1520 may relate to the parameters of the probability distribution, for example.
[0302] Also, having available the output of the output stage of the full-reference DNN, the representation of the probability distribution may be determined for example via determining at least one of a mean, a standard deviation, and a confidence interval from the output of the output stage.
[0303] For instance, the confidence interval may be determined based on the output of the output stage and a number of listeners to the listening test to be emulated. Referring to theabove example of a Gaussian parameterization of the probability distribution, once the parameter ^ has been determined, the 95% confidence interval (<=>?) can be computed as <=>? = 1.96^ . √1 (5) It is understood that an analogous determination can be applied to the case of a parametrization of the probability distribution as a logistic distribution, based on the scale ^.
[0304] Further to the methods described above, the present disclosure likewise relates to a DNN for estimating an indication of a subjective listening score for an audio signal. Again, the listening score may be a score according to a predefined listening test, such as a MUSHRA test, for example. Such a full-reference DNN may comprise an input stage for receiving a representation of the audio signal (e.g., one or more Gammatone spectrograms) and a representation of a reference audio signal for the audio signal (e.g., one or more Gammatone spectrograms), a plurality of layers for performing processing based on the representation of the audio signal and the representation of the reference audio signal, and an output stage for generating the indication of the listening score. It is understood that a first one of the plurality of layers is coupled to the input stage and that a last one of the plurality of layers is coupled to the output stage. Processing by the plurality of layers may be further based on a current state of the full-reference DNN, for example the current values of internal parameters of the full-reference DNN.
[0305] It is understood that the full-reference DNN may be implemented by any suitable computing system, such as the apparatus shown in Fig.27, for example.
[0306] Further, the full-reference DNN may have been configured (e.g., trained) by training the DNN in accordance with method 1400A described above.
[0307] In particular, the full-reference DNN may have been trained by, in a training epoch among a plurality of training epochs: inputting one or more training data items, each indicative of a respective value of the listening score, determining respective indications of the listening score based on the one or more training data items, determining respective loss values for the one or more training data items by evaluating a loss function, wherein the loss function depends on the indication of the listening score, and adjusting one or more internal parameters of the DNN based on the determined loss values.
[0308] As above, the indication of the listening score may relate to a probability distribution of the listening score, with the output stage of the full-reference DNN being adapted for generating the probability distribution of the listening score. This probability distribution may emulate listening scores obtained by a plurality of listening tests for the audio signal, and may be parameterized by two or more parameters of the probability distribution. For example, the probability distribution may relate to a Gaussian distribution parameterized by a mean ^ and a variance ^^or to a logistic distribution parameterized by a mean ^ and a scale ^, as described above.
[0309] When the indication of the listening score relates to a probability distribution of the listening score, determining respective indications of the listening score based on the one or more training data items may comprise determining respective parameters of the probability distribution based on the one or more training data items, for example based at least in part on the value of the subjective listening score. Further in this case, the loss function will depend on the parameters of the distribution. It is also understood that determining the parameters of the probability distribution may be based on a current state of the full-reference DNN, for example the current values of internal parameters of the full- reference DNN. Reference-free DNN
[0310] The above description relating to the full-reference DNN applies to Reference- free DNN, with the differences being illustrated in Figs.14B and 15B.
[0311] Method 1400B as illustrated by the flowchart of Fig.14B is an example of a possible implementation of training a reference-free DNN, in a training epoch among the plurality of training epochs. Method 1400B comprises steps S1410B, S1410C, S1420, S1430 and S1440.
[0312] Therein, steps S1420, S1430 and S1440 are the same as those in Fig.14A. The differences relating to the reference-free DNN as shown in Fig.14B are the steps S1410B, S1410C.
[0313] At step S1410B, weights for at least one of a plurality of layers of the reference-free DNN are initialized based on a pre-trained DNN. The pre-trained DNN is preferably a full-reference DNN as described above. The weights of the reference-free DNN may be initialized based on trained weights of one or more layers of the pre-trained DNN.Initialization of the weights of the reference-free DNN based on a pre-trained DNN shortens the training time of the reference-free DNN. For example, the initialized weights of the (input layer of the) reference-free DNN are initialized based on the weights of the input layer of the pre-trained DNN that are associated with the representation of the audio signal, and / or based on the weights of the input layer of the pre-trained DNN that are associated with the representation of the reference audio signal. A specific manner of utilizing the parameters of the pre-trained DNN is not limited by the above examples.
[0314] At step S1410C, one or more training data items, each indicative of a respective value of the listening score and further indicative of a representation of the audio signal, are input into an input stage of the reference-free DNN.
[0315] Based on a pre-trained DNN, a reference-free DNN may be configured and trained with minimum amount of costs and time as the reference-free DNN may have the same configuration as the pre-trained DNN other than in the parts relating to the reference audio signal.
[0316] An example of a corresponding method 1500B using this reference-free DNN is illustrated in the flowchart of Fig.15B. Method 1500B comprises steps S1510B and S1520.
[0317] At step S1510B, the representation of the audio signal is input to the input stage of the reference-free DNN. A representation of the reference audio signal as in step S1510A is not required.
[0318] At step S1520, a representation of the indication of the listening score is determined based on an output of the output stage of the reference-free DNN. Conversion from pre-trained DNN to reference-free DNN
[0319] The following description is based on examples with Generative Machine Listeners.
[0320] Generative Machine Listener (GML) is an intrusive audio quality model learnt with reference and degraded signal pairs collected from Dolby internal listening tests and predicts a corresponding MUSHRA score and confidence interval (CI) for the given pair. The model is based on InSE-Net with adaptive input and output layers, which enables transfer learning to a reference-free model with minimal modification in the current architecture.Shape Horizontal Vertical Normal Average Layer Output Conv. Conv. Conv. Pooling (w x h / ch) (w x h / ch) (w x h / ch) (w x h) Input 4 x 32 x 2824 Inception A 208 x 16 x 180 3 x 7 / 64 7 x 3 / 64 l x l / 64 5 x 5 Inception A 224 x 16 x 90 3 x 7 / 64 7 x 3 / 64 1 x 1 / 64 5 x 5 SE 224 x 16 x 90 Inception B 256 x 16 x 45 3 x 5 / 64 5 x 3 / 64 1 x 1 / 64 5 x 5 SE 256 x 16 x 45 Inception C 256 x 14 x 22 3 x 3 / 64 5 x 5 / 64 l x l / 64 3 x 3 SE 256 x 14 x 22 Adaptive AvgPool 256 x 4 x 4 FCL 1 3200 x 1 FCL 2 512 x 1 FCL 3 2 x 1 Table 1: Architectures and parameters of an RF-GML
[0321] Table 1 shows example architectures and parameters of an RF-GML. Compared to its GML having at the input layer a shape of 8 x 32 x 2824, only the shape of the input layer has changed, in that the number of channels for inputting audio signals have been reduced to half of that of the GML since the reference audio signal is not required.
[0322] Fig.24 shows an example of conversion from intrusive GML to non-intrusive GML.
[0323] The conversion may comprise removing all the reference channels in the input layer of the intrusive GML 2401, and keeping only the degraded channels of the intrusive GML 2401; Accordingly, while in the intrusive GML 2401 a reference signal – degraded signal pair 2403 is required to be input, in the non-intrusive GML 2402 only degraded signals 2404 need to be provided.
[0324] The conversion may further comprise transferring the learnt weights 2405 from intrusive GML 2401 to non-intrusive GML 2402.
[0325] The conversion may further comprise re-training the new non-intrusive model with only a few iterations (epochs).
[0326] According to the present disclosure, the reference-free GML (rf-GML) has inherited the capabilities from its predecessor. i.e., the intrusive GML, and thus can evaluate: 48-kHz mono / stereo / binaural general audio signals.
[0327] Fig.25 shows examples for transferring or initializing the weights of the reference-GML based on weights of the intrusive GML. Therein, the blocks are merely given once their reference number, for simplicity purposes. - Transfer Inception1 degraded weights (TL1_deg):
[0328] As shown on the left-hand side of Fig.25, the weights of degraded input channels 2511 in rf-GML are initialized with the corresponding degraded input channels 2501 in intrusive GML. - Transfer Inception 1 average of reference and degraded weights (TL1_avg):
[0329] As shown in the middle of Fig.25, the weights of degraded input channels in rf-GML are initialized with the average of ref. and deg. channels in intrusive GML. - Transfer Inception 2 (TL2):
[0330] As shown on the right-hand side of Fig.25, the layer Inception 2 in rf-GML is kept unchanged and so the weights could be initialized in the same way. - Default Initialization:
[0331] Although not shown in Fig.25, the reference-free GML may be provided with default initialization of PyTorch without transfer learning from pre-trained intrusive model. Simulated Results on an Example Test Set
[0332] Two stereo listening tests were considered as a test set. One listening test tests low bitrate codecs, and the other tests high bitrate codecs. A strategy has been devised for selecting the best model out of several models from the trained epochs.
[0333] The following factors may be taken into consideration during model selection: • The stability of the training process. For example, in general, the model trained with logistic distribution shows a smoother training and validation loss decay than the Gaussian distribution under the same settings. But Gaussian still could be a promising option if the training process is fine-tuned with techniques like gradient clipping.• Pearson Correlation Coefficient (PCC) between predicted mean MUSHRA score and actual mean MUSHRA listening test score. Higher PCC (close to 1) is preferred. • The NLL loss of training and validation set. For example, a model trained with a Gaussian distribution shows much smaller NLL loss than a model trained with logistic distribution.
[0334] Taking the above-mentioned aspects into account, several models were selected (from different epochs, i.e., from different stages of training; for either Gaussian or logistic) with the highest PCC on the validation set, lower training NLL losses, and a moderately lower validation loss. For models with similar PCC scores, the models with lower training NLL loss were kept (which is typically the model produced at the later epochs), but a moderately lower validation NLL loss. The reason for the latter is because it has been found that the model with the least validation loss does not necessarily show the best performance in predicting the confidence interval on test sets. Figs.16 – 20 relate to intrusive schemes.
[0335] Fig.16 is a plot showing examples of the mean NLL loss on the two listening tests with Gaussian and logistic distributions. Overall, logistic shows higher NLL-loss than the Gaussian model on the test sets. However, the lower loss does not necessarily mean better models (e.g., visually the logistic model shows a closer fit to the ground truth).
[0336] For the plots shown in Fig.17 through Fig.20, each excerpt has a reference, 3.5 kHz anchor, and 7 kHz anchor, followed by different coded representations.
[0337] Fig.17 is a plot showing example results of the stereo low bitrate test. Specifically, the plot shows accuracy of predicting the mean MUSHRA score with the generative machine listener trained with logistic and Gaussian distribution for different classes of audio (3 speech, music, and mixed excerpts).
[0338] Fig.18 is another plot showing example results of the stereo low bitrate test Specifically, the plot shows accuracy of predicting the CI with the generative machine listener trained with logistic and Gaussian distribution for different classes of audio (3 speech, music, and mixed excerpts) with 44 listeners.
[0339] Fig.19 is a plot showing example results of the stereo high bitrate test. Specifically, the plot shows accuracy of predicting the mean MUSHRA score with thegenerative machine listener trained with logistic and Gaussian distribution for different classes of audio (3 speech, music, and mixed excerpts).
[0340] Finally, Fig.20 is a plot showing example results of the stereo high bitrate test. Specifically, the plot shows accuracy of predicting the CI with the generative machine listener trained with logistic and Gaussian distribution for different classes of audio (3 speech, music, and mixed excerpts) with 28 listeners.
[0341] Pearson correlation (PC) evaluates the linear relationship between two continuous variables. For the examples shown in the above plots of Fig.17 through Fig.20, the trend for predicting the CI is closer to ground truth (i.e., CI from listening test) with the model trained with logistic distribution (PC = 0.8190) than with Gaussian distribution (PC = 0.6733). For high bitrates, both logistic distribution (PC = 0.6461) and Gaussian distribution (PC = 0.6700) are on par.
[0342] For predicting the mean, both the models perform on par in the above examples. For the low bitrate test, the model trained with logistic distribution has PC = 0.9172, and Gaussian distribution has PC = 0.9183. For the high bitrate test, the model trained with logistic distribution has PC = 0.9316, and Gaussian distribution has PC = 0.9375.
[0343] Spearman correlation (SC) evaluates the monotonic relationship. The Spearman correlation coefficient is based on the ranked values for each variable rather than the raw data. It measures rank preservation. For the low bitrate test, the model trained with logistic distribution has SC = 0.8991, and Gaussian distribution has SC = 0.8716 in the above examples. For the high bitrate test, the model trained with logistic distribution has SC = 0.9297, and Gaussian distribution has SC = 0.9360.
[0344] It can be noted that in subjective scores, references, low-pass anchors, and in general bitrates associated with higher quality are rated with low CI, and bitrates associated with lower quality are rated with higher CI. Unlike the aforementioned bootstrapping approach, where higher CI is associated with more dense data points in the training set (and vice-versa), techniques according to the present disclosure model the diversity of the listener scores.
[0345] Fig.21 shows an example training loss plot comparing full-reference GML and three reference-free GMLs (RF-GMLs) on full-reference listening tests.
[0346] Therein, RF-GML No init refers to RF-GML with default random initialization; RF-GML TL1 deg init refers to RF-GML having weights of degraded input channels being initialized with the corresponding degraded input channels in full-reference GML; and RF-GML Transfer all init refers to an extreme case where everything from GML is transferred to RF-GML as initialization.
[0347] As can be seen from Fig.21, GML has the lowest NLL loss, which is naturally expected to be the best.
[0348] Although RF-GML TL1 deg init appears to have similar loss performance as RF-GML No init, it is significantly better in rating uncoded audio close to 100. This is because, when training the full-reference GML, it has seen many reference-reference pairs (as its reference-degraded input due to the nature of the MUSHRA listening test) in addition to reference-coded pairs. Thus, if the weights of the degraded input channels in RF-GML are initialized with the corresponding degraded input channels in the full-reference GML, it starts off with some inherent knowledge about rating the reference signal.
[0349] Transfer all init has the worst NLL loss and hence the worst performance.
[0350] Fig.22 shows an example scaling of RF-GML scores with bitrates for variants of RF-GML.
[0351] Two variants of RF-GML are compared: (1) RF-GML (def) was trained from scratch using the default PyTorch initializer [6], i.e., without transfer learning. (2) In RF- GML (deg), the weights of degraded input channels in the first inception (In-A) block were initialized with corresponding degraded input channels from a pre-trained GML.
[0352] As explained above, although RF-GML (deg) may have lower correlation scores at times, it still consistently ranks higher than RF-GML (def), which performs better in correlation metrics but overall is second best.
[0353] All the 24 stereo excerpts (from the test set) are coded for a wide range of bitrates using the (HE)-AAC family of codecs typically used in practice, and as demonstrated in Fig.22 the mean RF-GML quality score scales reasonably well with the bitrates, for both of the above two variants RF-GML (deg) and RF-GML (def).
[0354] However, the average quality predicted with RF-GML (deg) at 96 kb / s HE- AAC v1 is slightly higher than AAC at 192 kb / s. A possible reason could be that HE-AAC v1 at 96 kb / s has a higher overall bandwidth due to the usage of parametric spectral bandreplication [7]. Since it is expected the quality of both 192 kb / s AAC and 96 kb / s HE-AAC to be in the excellent quality range anyway, from a reference-free quality point of view, it would have been a tough task also for humans. On the other hand, the quality predicted with the RF- GML (def) saturates at around 80, and there is less separation in quality across bitrates. In addition, a slight increase in quality is observed when going down in bitrate from 32 to 20 kb / s. This observation, along with the strong performance in predicting the subjective quality (presented in Table I), makes RFGML (deg) also suitable for quality monitoring in clients in adaptive streaming scenarios.
[0355] Fig.23 shows an example of predicted quality scores versus audio bandwidth, for SESQA, RF-GML (deg), and RFGML(def).
[0356] It is further evaluated the model’s performance in rating uncoded content by running RF-GML and SESQA on 511 excerpts, including synthetic test tones and signal, and plotted scatter plots of predicted quality scores versus bandwidth. This content, curated internally over the years for testing the engineering implementation of codecs, was not fully known to before the experiments. Fig.23 shows almost no correlation (-0.08) between bandwidth and quality scores for SESQA, while RF-GML (deg) and RF-GML (def) show slight correlations of 0.44 and 0.51, respectively. This is expected, as RFGML is trained to evaluate coding artifacts, with audio bandwidth being a typical codec tuning parameter. Note that a better correlation with RF-GML (def) may not mean it is better because uncoded speech signals may have lower bandwidth. The scatter plot again shows that RF-GML (deg) rates uncoded audio closer to 100 than RF-GML (def). The five lowest-rated signals for RF-GML are the 3.5 kHz and 4.5 kHz low-pass filtered versions, while SESQA rated full bandwidth uncoded music signals the lowest. This is expected for RF-GML, trained on MUSHRA tests with low-pass filtered anchors, and for SESQA, which was not trained with music signals. Both SESQA and RF-GML (deg) rated the five lowest bandwidth signals (1 kHz sine tones and 500 Hz gong instrument) similarly, assigning high-quality scores. These results demonstrate that RF-GML (deg) consistently rates uncoded audio closer to 100 and could potentially be used to identify uncoded audio.
[0357] To summarize: RF-GML evaluates 48 kHz mono / stereo / binaural generic audio without aid of any form of reference signals;It achieves competitive results on mono / stereo signals at low bitrates as SESQA on speech signals, but much better accuracy on speech signals at high bitrates; On generic audio tests, the proposed model achieves on average 0.75-0.8 spearman ranking correlation and 0.8-0.9 pearson linear correlation; and Retraining and maintenance of rf-GML is efficient and as a by-product of GML, it could be easily updated updating / improving the GML. Binaural transformation
[0358] Fig.26 shows an example of binaural transformation of an audio signal which is a multi-channel signal or an object-based signal, to make it suitable for being provided to an RF-GML as described above.
[0359] As an example, a coded multichannel audio signal 2601 is transformed to a binaural signal 2602 for being further transformed into a binaural coded signal 2603. Further, the binaural coded signal 2603 is input into the RF-GML 2604 for providing an audio quality score (e.g., based on MUSHRA test or MUSHRA-like test) and the corresponding CI.
[0360] With binaural renderer as a pre-processor, rf-GML can be extended to also evaluate the quality of spatial audio. Apparatus for Implementing Methods According to the Disclosure
[0361] Finally, while reference above may be mainly made to methods according to the present disclosure, the present disclosure likewise relates to apparatus (e.g., computer- implemented apparatus) for performing methods and techniques described throughout the present disclosure. An example of such apparatus 2100 is schematically illustrated in Fig.27. Such an apparatus 2100 may implement, for example, the deep neural network (machine listener) described above. The apparatus 2100 comprises a processor 2110 and a memory 2120 coupled to the processor 2110. The memory 2120 may store instructions for the processor 2110. The processor 2110 may also receive, among others, suitable input data (e.g., suitable training data at the training stage, or suitable test and reference audio signals at inference), depending on use cases and / or implementations. The processor 2110 may be adapted to carry out the methods / techniques described throughout the present disclosure (e.g., method 1300 of Fig.13, method 1400A of Fig.14A, method 1400B of Fig.14B, method 1500A of Fig.15A, or method 1500B of Fig.15B) and to generate corresponding output data 1240 (e.g., an indication of a listening score), depending on use cases and / orimplementations. For example, the apparatus 2100 may implement a method of training the DNN described above, or it may implement the (trained) DNN described above, the DNN being non-intrusive or intrusive.
[0362] The present disclosure likewise relates to corresponding computer programs and computer-readable storage media. Interpretation
[0363] Aspects of the systems described herein may be implemented in an appropriate computer-based audio processing network environment (e.g., server or cloud environment) for processing digital or digitized audio files. Portions of the adaptive audio system may include one or more networks that comprise any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route the data transmitted among the computers. Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.
[0364] One or more of the components, blocks, processes or other functional components may be implemented through a computer program that controls execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware, firmware, and / or as data and / or instructions embodied in various machine-readable or computer-readable media, in terms of their behavioral, register transfer, logic component, and / or other characteristics. Computer-readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, physical (non-transitory), non- volatile storage media in various forms, such as optical, magnetic or semiconductor storage media.
[0365] Specifically, it should be understood that embodiments may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, one of ordinary skill in the art, and based on a reading of this detailed description, would recognize that, in at least one embodiment, the electronic-based aspects may be implemented in software (e.g., stored on non-transitory computer-readable medium) executable by one or more electronic processors, such as a microprocessor and / or application specific integrated circuits (“ASICs”). As such, it should be noted that a plurality of hardwareand software-based devices, as well as a plurality of different structural components, may be utilized to implement the embodiments. For example, the systems, services, clients, nodes, etc., described in the context of Fig.1A, Fig.1B, Fig.3A, Fig.3B, Fig.3C, Fig.5A, Fig.5B, Fig.8A, Fig.9A, Fig.8B, Fig.9B, Fig.10, Fig.12 and / or Fig.27 above can include or be implemented by one or more electronic processors, one or more computer-readable medium modules, one or more input / output interfaces, and various connections (e.g., a system bus) connecting the various components.
[0366] While one or more implementations have been described by way of example and in terms of the specific embodiments, it is to be understood that one or more implementations are not limited to the disclosed embodiments. To the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements.
[0367] Also, it is to be understood that the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. Enumerated Example Embodiments
[0368] Various aspects and implementations of the present disclosure may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims.
[0369] EEE1. A method for evaluating playout performance of adaptive streaming for a streaming client, wherein the method uses an objective quality assessment algorithm which emulates an intrusive quality test, wherein the method is implemented on a different network node than the client, wherein the method comprises: receiving playout related metadata from a client; using the metadata to reconstruct description of the test signal as required by the objective quality assessment algorithm; reconstructing description of the reference signal as required by the objective quality assessment algorithm; and computing the performanceestimates and distributing them to other network nodes than the node, where the method operates.
[0370] EEE2. The method of EEE1, where the objective quality assessment algorithm uses a set of pretrained models and performs selection of a model based on the metadata received from a client.
[0371] EEE3. The method of EEE1 or EEE2, where the metadata received from the client comprises information that facilitates reconstruction of the description of the test signal required by the objective quality assessment tool at the network node implementing the method, for example: segment name and bitrate; and / or representation of a set of spectrograms (required by the objective quality assessment algorithm) computed based on the segment content.
[0372] EEE4. The method of EEE3, where the metadata received from the client comprises information on the playout device and / or playout conditions, for example: indication of a playout device (headphone, soundbar or discrete speakers); characteristics of the playout conditions (e.g., Signal to Noise ratio); and / or other.
[0373] EEE5. The method of any of the preceding EEEs, where the estimates provided by the method are used to compare at least two different populations of clients.
[0374] EEE6. The method of any of the preceding EEEs, where the performance estimates computed by the method are provided to a service optimizing encoding / packaging of the content.
[0375] EEE7. The method of EEE6, where the performance estimates are used as an input to a utility function used for determining an optimal number of quality levels in the bitrate ladder to be distributed on Points of Presence (PoPs) of a CDN.
[0376] EEE8. The method of EEE6, where the performance estimates are used to determine a tuning (e.g. configuration, set of coding tools) of the content encoder.
[0377] EEE-A1. A method of configuring a deep neural network, DNN, for estimating an indication of a subjective listening score for an audio signal, wherein the listening score is a score according to a predefined listening test, the method comprising: providing an output stage of the DNN to generate the indication of the listening score; and training the DNN by, in a training epoch among a plurality of training epochs:inputting one or more training data items, each indicative of a respective value of the listening score; determining respective indications of the listening score based on the one or more training data items; determining respective loss values for the one or more training data items by evaluating a loss function, wherein the loss function depends on the indication of the listening score; and adjusting one or more internal parameters of the DNN based on the determined loss values.
[0378] EEE-A2. The method according to EEE-A1, wherein the indication of the listening score relates to a probability distribution of the listening score, with the output stage being adapted for generating the probability distribution of the listening score, wherein the probability distribution emulates listening scores obtained by a plurality of listening tests for the audio signal, and wherein the probability distribution is parameterized by two or more parameters of the probability distribution; wherein determining respective indications of the listening score based on the one or more training data items comprises determining respective parameters of the probability distribution based on the one or more training data items; and wherein the loss function depends on the parameters of the distribution.
[0379] EEE-A3. The method according to EEE-A1 or EEE-A2, wherein training the DNN is based on a maximum likelihood principle.
[0380] EEE-A4. The method according to any one of the preceding EEE-As, wherein the loss function relates to a negative log likelihood, NLL, loss.
[0381] EEE-A5. The method according to EEE-A4 when depending on EEE-A2,wherein the negative log likelihood loss is given by ^'(( = − log *+^^|-, / ^, where*+^^|-, / ^ is the probability distribution for test score ^ given a representation of the audiosignal / and a representation of a reference audio signal - for the audio signal / , and 0 indicates the internal parameters of the DNN.
[0382] EEE-A6. The method according to EEE-A2 or any one of EEE-A3 to EEE-A5 when depending on EEE-A2, wherein the probability distribution relates to a Gaussian distribution parameterized by a mean ^ and a variance ^^, and wherein the loss function^^^^^^ is given by ^^^^^^ = log ^ + + ^, where ^ is a constant and ^ is the subjectivelistening score.
[0383] EEE-A7. The method according to EEE-A2 or any one of EEE-A3 to EEE-A5 when depending on EEE-A2, wherein the probability distribution relates to a logistic distribution parameterized by a mean ^ and a scale ^, and wherein the loss function ^^^^^^^^^is given by ^^^^^^^^^ = log ^ + 2 log sech+ ^, where ^ is a constant and ^ is thesubjective listening score.
[0384] EEE-A8. The method according to any one of the preceding EEE-As, wherein the training data item is further indicative of a representation of the audio signal and a representation of a reference audio signal for the audio signal.
[0385] EEE-A9. The method according to EEE-A8, wherein the representation of the audio signal and the representation of the reference audio signal relate to Gammatone spectrograms.
[0386] EEE-A10. The method according to any one of the preceding EEE-As, wherein the predefined listening test is a Multi-Stimulus Test with Hidden Reference and Anchor, MUSHRA, listening test.
[0387] EEE-A11. The method according to any one of the preceding EEE-As, wherein the DNN implements a generative model.
[0388] EEE-A12. A method of estimating an indication of a subjective listening score for an audio signal using a deep neural network, DNN, wherein the listening score is a score according to a predefined listening test, wherein the DNN comprises: an input stage for receiving a representation of the audio signal and a representation of a reference audio signal for the audio signal; a plurality of layers for performing processing based on the representation of the audio signal and the representation of the reference audio signal; and an output stage for generating the indication of the listening score; and wherein the method comprises: inputting the representation of the audio signal and the representation of the reference audio signal to the input stage; and determining a representation of the indication of the listening score based on an output of the output stage.
[0389] EEE-A13. The method according to EEE-A12, wherein the indication of the listening score relates to a probability distribution of the listening score, with the output stage of the DNN being adapted for generating the probability distribution of the listening score, wherein the probability distribution emulates listening scores obtained by a plurality of listening tests for the audio signal, and wherein the probability distribution is parameterized by two or more parameters of the probability distribution.
[0390] EEE-A14. The method according to EEE-A13, wherein determining the representation of the probability distribution comprises determining at least one of a mean, a standard deviation, and a confidence interval from the output of the output stage.
[0391] EEE-A15. The method according to EEE-A14, wherein the confidence interval is determined based on the output of the output stage and a number of listeners to the listening test to be emulated.
[0392] EEE-A16. The method according to any one of EEE-A13 to EEE-A15, wherein the probability distribution relates to a Gaussian distribution parameterized by a mean ^ and a variance ^^or to a logistic distribution parameterized by a mean ^ and a scale ^.
[0393] EEE-A17. The method according to any one of EEE-A13 to EEE-A16, wherein the representation of the audio signal and the representation of the reference audio signal relate to Gammatone spectrograms.
[0394] EEE-A18. The method according to any one of EEE-A13 to EEE-A17, wherein the predefined listening test is a Multi-Stimulus Test with Hidden Reference and Anchor, MUSHRA, listening test.
[0395] EEE-A19. A deep neural network, DNN, for estimating an indication of a subjective listening score for an audio signal, wherein the listening score is a score according to a predefined listening test, the DNN comprising: an input stage for receiving a representation of the audio signal and a representation of a reference audio signal for the audio signal; a plurality of layers for performing processing based on the representation of the audio signal and the representation of the reference audio signal; and an output stage for generating the indication of the listening score.
[0396] EEE-A20. The DNN according to EEE-A19, wherein the DNN has been configured by training the DNN by, in a training epoch among a plurality of training epochs: inputting one or more training data items, each indicative of a respective value of the listening score; determining respective indications of the listening score based on the one or more training data items; determining respective loss values for the one or more training data items by evaluating a loss function, wherein the loss function depends on the indication of the listening score; and adjusting one or more internal parameters of the DNN based on the determined loss values.
[0397] EEE-A21. The DNN according to EEE-A19 or EEE-A20, wherein the indication of the listening score relates to a probability distribution of the listening score, with the output stage of the DNN being adapted for generating the probability distribution of the listening score, wherein the probability distribution emulates listening scores obtained by a plurality of listening tests for the audio signal, and wherein the probability distribution is parameterized by two or more parameters of the probability distribution.
[0398] EEE-A22. The DNN according to EEE-A21 when depending on EEE-A20, wherein determining respective indications of the listening score based on the one or more training data items comprises determining respective parameters of the probability distribution based on the one or more training data items; and wherein the loss function depends on the parameters of the distribution.
[0399] EEE-A23. The DNN according to EEE-A21 or EEE-A22, wherein the probability distribution relates to a Gaussian distribution parameterized by a mean ^ and a variance ^^or to a logistic distribution parameterized by a mean ^ and a scale ^.
[0400] EEE-A24. The DNN according to any one of EEE-A19 to EEE-A23, wherein the representation of the audio signal and the representation of the reference audio signal relate to Gammatone spectrograms.
[0401] EEE-A25. The DNN according to any one of EEE-A19 to EEE-A24, wherein the predefined listening test is a Multi-Stimulus Test with Hidden Reference and Anchor, MUSHRA, listening test.
[0402] EEE-A26. An apparatus comprising a processor and a memory coupled to the processor, and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of EEE-A1 to EEE-A18.
[0403] EEE-A27. An apparatus comprising a processor and a memory coupled to the processor, and storing instructions for the processor, wherein the processor is adapted to implement the DNN according to any one of EEE-A19 to EEE-A25.
[0404] EEE-A28. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of EEE-A1 to EEE-A18.
[0405] EEE-A29. A program comprising instructions that, when executed by a processor, cause the processor to implement the DNN according to any one of EEE-A19 to EEE-A25.
[0406] EEE-A30. A computer-readable storage medium storing the program of EEE- A28 or EEE-A29.
[0407] EEE-B1. A method of evaluating playout performance in an adaptive streaming environment, the method comprising: obtaining playout-related information from a streaming client; estimating a representation of a test audio signal based on the playout-related information, wherein the test audio signal is an audio signal played out by the streaming client; and determining, using an audio quality assessment algorithm, an estimate of an audio quality of the test audio signal based on the estimated representation of the test audio signal.
[0408] EEE-B2. The method according to EEE-B1, further comprising: generating a representation of a reference audio signal for the test audio signal.
[0409] EEE-B3. The method according to any one of the preceding EEE-Bs, further comprising: obtaining, from the streaming client, an indication of audio content processed by the streaming client.
[0410] EEE-B4. The method according to EEE-B3, wherein estimating the representation of the test audio signal is further based on the indication of the audio content.
[0411] EEE-B5. The method according to EEE-B3 or EEE-B4 when depending on EEE-B2, wherein generating the representation of the reference audio signal is based on the indication of the audio content.
[0412] EEE-B6. The method according to any one of the preceding EEE-Bs, wherein the playout-related information comprises bitrate information indicating a bitrate of the audio signal played out by the streaming client; and wherein estimating the representation of the test audio signal is based on the bitrate information.
[0413] EEE-B7. The method according to any one of the preceding EEE-Bs, wherein the audio quality assessment algorithm uses a set of pretrained models for audio quality assessment; and wherein generating the estimate of the audio quality comprises selecting a pretrained model among the set of pretrained models based on the playout-related information.
[0414] EEE-B8. The method according to EEE-B7, wherein the playout related information comprises information relating to a playout device associated with the streaming client; and wherein the pretrained model is selected based on the information relating to the playout device.
[0415] EEE-B9. The method according to any one of the preceding EEE-Bs, wherein the audio quality assessment algorithm is implemented by a deep neural network, DNN, for estimating an indication of a subjective listening score for the representation of a test audio signal as the estimate of the audio quality, wherein the listening score is a score according to a predefined listening test, the DNN comprising: an input stage for receiving the representation of the test audio signal and a representation of a reference audio signal for the test audio signal; a plurality of layers for performing processing based on the representation of the test audio signal and the representation of the reference audio signal; and an output stage for generating the indication of the listening score.
[0416] EEE-B10. The method according to EEE-B9, wherein the DNN has been configured by training the DNN by, in a training epoch among a plurality of training epochs:inputting one or more training data items, each indicative of a respective value of the listening score; determining respective indications of the listening score based on the one or more training data items; determining respective loss values for the one or more training data items by evaluating a loss function, wherein the loss function depends on the indication of the listening score; and adjusting one or more internal parameters of the DNN based on the determined loss values.
[0417] EEE-B11. The method according to any one of the preceding EEE-Bs, wherein the method is implemented at a different network node than the streaming client.
[0418] EEE-B12. The method according to any one of the preceding EEE-Bs, wherein the representation of the estimate of the test audio signal relates to one or more Gammatone spectrograms.
[0419] EEE-B13. The method according to EEE-B2 or any EEE-B depending on EEE-B2, wherein the representation of the reference audio signal relates to one or more Gammatone spectrograms.
[0420] EEE-B14. The method according to any one of the preceding EEE-Bs, further comprising: outputting the estimate of the audio quality of the test audio signal to a network node different from a network node associated with the streaming client.
[0421] EEE-B15. The method according to any one of the preceding EEE-Bs, wherein the estimate of the audio quality of the test audio signal is output to a network node for performing encoding and / or packaging of the audio content; and the method further comprises optimizing the encoding and / or packaging based on the estimate of the audio quality of the test audio signal.
[0422] EEE-B16. The method according to EEE-B15, further comprising determining an optimal number of quality levels in a bitrate ladder for distribution by a content delivery network, based on the estimate of the audio quality of the test audio signal.
[0423] EEE-B17. The method according to EEE-B15 or EEE-B16, further comprising determining a configuration and / or set of coding tools based on the estimate of the audio quality of the test audio signal.
[0424] EEE-B18. The method according to any one of the preceding EEE-Bs, further comprising: determining estimates of audio quality of test audio signal for streaming clients in each of a plurality of populations of streaming clients; and comparing the estimates of audio quality determined for the plurality of populations of streaming clients.
[0425] EEE-B19. A method of providing playout-related information at a streaming client that processes audio content in an adaptive streaming environment, the method comprising: generating the playout-related information by one or more of: analyzing a playout buffer associated with the streaming client for determining bitrate information indicating a bitrate of segments of an audio signal played out by the streaming client; analyzing manifest information associated with the audio content; and analyzing characteristics of a playout device associated with the streaming clients; and outputting the playout-related information to a network node different from a network node associated with the streaming client.
[0426] EEE-B20. An apparatus comprising a processor and a memory coupled to the processor, and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of EEE-B1 to EEE-B19.
[0427] EEE-B21. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of EEE-B1 to EEE-B19. EEE-B22. A computer-readable storage medium storing the program of EEE-B21.
[0428] EEE-C1. A method of an audio steaming device estimating an indication of a subjective listening score for an audio signal played out by the audio streaming device, wherein the listening score is a score according to a predefined listening test, the method comprising: obtaining metadata relating to the played out audio signal; providing, based on the obtained metadata, test data for being input into a trained reference- free deep neural network, DNN, for estimating the indication of the listening score;wherein the reference-free DNN comprises: an input stage for receiving the test data; a plurality of layers for performing processing based on the received test data; and an output stage for generating the indication of the listening score; and wherein the method comprises: inputting the test data into the input stage; and determining a representation of the indication of the listening score based on an output of the output stage.
[0429] EEE-C2. The method according to EEE-C1, wherein the method further comprises selecting, based on the obtained metadata and from a plurality of trained algorithms, an algorithm to be applied in the trained reference-free DNN for estimating the indication of the listening score.
[0430] EEE-C3. The method according to EEE-C2, wherein each of the plurality of trained algorithms corresponds to a respective playout mode for being used by the audio streaming device, wherein the metadata comprises information relating to the playout mode currently used on the audio streaming device.
[0431] EEE-C4. The method according to EEE-C3, wherein the playout modes comprise a headphone mode, a discrete speaker mode and a soundbar mode.
[0432] EEE-C5. The method according to any one of the preceding EEE-Cs, wherein the metadata comprises information relating to one or more of the following analyses: an analysis of the playout buffer provided in the audio streaming device, an analysis of a bitrate ladder used for providing the audio streaming device with streaming audio data, an analysis of the audio streaming device, and an analysis of playout conditions.
[0433] EEE-C6. The method according to any one of the preceding EEE-Cs, wherein the test data is indicative of a representation of a binaural transformation of the played out audio signal, wherein the played out audio signal relates to a multi-channel signal or an object-based signal.
[0434] EEE-C7. The method according to any one of the preceding EEE-Cs, wherein the test data is indicative of a representation of the audio signal and the representation of the audio signal relates to Gammatone spectrograms.
[0435] EEE-C8. The method according to EEE-C7, wherein the representation of the audio signal comprises segments of Gammatone spectrograms, each segment being of a respective pre-determined time length in the played out audio signal, wherein the test data comprises information relating to respective segment names and / or corresponding bitrates used for the respective segments.
[0436] EEE-C9. The method according to any one of the preceding EEE-Cs, wherein the method comprises estimating indications of respective listening scores for a plurality of audio streaming device clusters, each cluster corresponding to a respective content delivery infrastructure, and each cluster comprising one or more audio streaming devices, wherein the method further comprises: estimating for each cluster an indication of a listening score based on metadata collected from the one or more audio streaming devices included in the each cluster, and comparing the estimated listening scores of the clusters.
[0437] EEE-C10. The method according to any one of the preceding EEE-Cs, wherein the method further comprises providing the estimated indication of respective listening score to a content encoding and packaging infrastructure for adjusting a number of levels used in a bitrate ladder for providing the audio streaming device with streaming audio data and / or for adjusting bitrates applied in the content encoding and packaging infrastructure.
[0438] EEE-C11. The method according to any one of the preceding EEE-Cs, wherein the reference-free DNN has been configured by initializing training weights for at least one of the plurality of layers and training the reference-free DNN by, in a training epoch among a plurality of training epochs: inputting one or more training data items, each indicative of a representation of a training audio signal and further indicative of a respective value of a listening score for the training audio signal; determining respective indications of the listening score based on the one or more training data items; determining respective loss values for the one or more training data items by evaluating a loss function, wherein the loss function depends on the indication of the listening score; and adjusting one or more internal parameters of the reference-free DNN based on the determined loss values,wherein the training weights of the reference-free DNN are initialized based on a pre-trained DNN for estimating an indication of a listening score for the training audio signal based on the training audio signal and a reference audio signal for the training audio signal, wherein the pre-trained DNN comprises: an input layer for receiving a representation of the training audio signal and a representation of the reference audio signal for the training audio signal; a plurality of layers for performing processing based on the representation of the training audio signal and the representation of the reference audio signal; and an output stage for generating the indication of the listening score wherein the training weights of the reference-free DNN are initialized based on first weights of the input layer of the pre-trained DNN, which are weights associated with the representation of the training audio signal, and / or second weights of the input layer of the pre- trained DNN, which are weights associated with the representation of the reference audio signal.
[0439] EEE-C12. The method according to EEE-C11, wherein the initialized training weights of the reference-free DNN are weights of the input stage of the reference-free DNN.
[0440] EEE-C13. The method according to EEE-C11 or EEE-C12, wherein the initialized training weights of the reference-free DNN are initialized based on the first weights of the pre-trained DNN.
[0441] EEE-C14. The method according to any one of EEE-C11 to EEE-C13, wherein a value of each of the initialized training weights of the reference-free DNN is based on an average value of a respective first weight and a respective second weight.
[0442] EEE-C15. The method according to any one of the preceding EEE-Cs, wherein the indication of the listening score relates to a probability distribution of the listening score, with the output stage of the reference-free DNN being adapted for generating the probability distribution of the listening score, wherein the probability distribution emulates listening scores obtained by a plurality of listening tests for the audio signal, and wherein the probability distribution is parameterized by two or more parameters of the probability distribution.
[0443] EEE-C16. The method according to EEE-C15, wherein determining the representation of the probability distribution comprises determining at least one of a mean, a standard deviation, and a confidence interval from the output of the output stage.
[0444] EEE-C17. The method according to any one of the preceding EEE-Cs, wherein the predefined listening test is a Multi-Stimulus Test with Hidden Reference and Anchor, MUSHRA, listening test and / or MUSHRA-like listening test, but without presenting (or exposing the known) reference to the subjects.
[0445] EEE-C18. An apparatus comprising a processor and a memory coupled to the processor, and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of EEE-C1 to EEE-C17.
[0446] EEE-C19. A program comprising instructions that, when executed by a processor, cause the processor to implement the DNN according to any one of EEE-C1 to EEE-C17.
[0447] EEE-C20. A computer-readable storage medium storing the program of EEE- C19.
[0448] EEE-D1. A method of configuring a reference-free deep neural network, DNN, for estimating an indication of a subjective listening score for an audio signal, wherein the listening score is a score according to a predefined listening test, the method comprising: providing an output stage of the reference-free DNN to generate the indication of the listening score; initialize weights for at least one of a plurality of layers of the reference-free DNN; and after initializing said weights, training the reference-free DNN by, in a training epoch among a plurality of training epochs: inputting one or more training data items, each indicative of a respective value of the listening score and further indicative of a representation of the audio signal, into an input stage of the reference-free DNN; determining respective indications of the listening score based on the one or more training data items; determining respective loss values for the one or more training data items by evaluating a loss function, wherein the loss function depends on the indication of the listening score; and adjusting one or more internal parameters of the reference-free DNN based on the determined loss values, wherein the weights of the reference-free DNN are initialized based on a pre-trained DNN for estimating an indication of a subjective listening score for the audio signal based on the audiosignal and a reference audio signal for the audio signal, wherein the pre-trained DNN comprises: an input layer for receiving a representation of the audio signal and a representation of the reference audio signal for the audio signal; a plurality of layers for performing processing based on the representation of the audio signal and the representation of the reference audio signal; and an output stage for generating the indication of the listening score, wherein the weights of the reference-free DNN are initialized based on first weights of the input layer of the pre-trained DNN, which are weights associated with the representation of the audio signal, and / or second weights of the input layer of the pre-trained DNN, which are weights associated with the representation of the reference audio signal.
[0449] EEE-D2. The method according to EEE-D1, wherein the initialized weights of the reference-free DNN are weights of the input stage of the reference-free DNN.
[0450] EEE-D3. The method according to EEE-D1 or EEE-D2, wherein the initialized weights of the reference-free DNN are initialized based on the first weights of the pre-trained DNN.
[0451] EEE-D4. The method according to any one of EEE-D1 to EEE-D3, wherein a value of each of the initialized weights of the reference-free DNN is based on an average value of a respective first weight and a respective second weight.
[0452] EEE-D5. The method according to any one of the preceding EEE-Ds, wherein the one or more training data items are each indicative of a respective representation of a binaural transformation of the audio signal, wherein the audio signal is a multi-channel audio signal or an object-based audio signal.
[0453] EEE-D6. The method according to any one of the preceding EEE-Ds, wherein the indication of the listening score relates to a probability distribution of the listening score, with the output stage being adapted for generating the probability distribution of the listening score, wherein the probability distribution emulates listening scores obtained by a plurality of listening tests for the audio signal, and wherein the probability distribution is parameterized by two or more parameters of the probability distribution; wherein determining respective indications of the listening score based on the one or more training data items comprises determining respective parameters of the probability distribution based on the one or more training data items; andwherein the loss function depends on the parameters of the distribution.
[0454] EEE-D7. The method according to any one of the preceding EEE-Ds, wherein training the reference-free DNN is based on a maximum likelihood principle.
[0455] EEE-D8. The method according to any one of the preceding EEE-Ds, wherein the loss function relates to a negative log likelihood, NLL, loss.
[0456] EEE-D9. The method according to EEE-D6 or any one of EEE-D7 to EEE-D8 when depending on EEE-D6, wherein the probability distribution relates to a Gaussian distribution parameterized by a mean ^ and a variance ^^, and wherein the loss function^^^^^^ is given by ^^^^^^ = log ^ + ^^^ + ^, where ^ is a constant and ^ is the subjectivelistening score.
[0457] EEE-D10. The method according to EEE-D6 or any one of EEE-D7 to EEE- D9 when depending on EEE-D6, wherein the probability distribution relates to a logistic distribution parameterized by a mean ^ and a scale ^, and wherein the loss function ^^^^^^^^^constant and ^ is thesubjective listening score.
[0458] EEE-D11. The method according to any one of the preceding EEE-Ds, wherein the representation of the audio signal and the representation of the reference audio signal relate to Gammatone spectrograms.
[0459] EEE-D12. The method according to any one of the preceding EEE-Ds, wherein the predefined listening test is a Multi-Stimulus Test with Hidden Reference and Anchor, MUSHRA, listening test and / or MUSHRA-like listening test, but without presenting (or exposing the known) reference to the subjects.
[0460] EEE-D13. The method according to any one of the preceding EEE-Ds, wherein the reference-free DNN implements a generative model.
[0461] EEE-D14. The method according to any one of the preceding EEE-Ds, wherein the pre-trained DNN has been trained based on a negative log likelihood loss givenby ^'(( = − log *+^^|-, / ^, where *+^^|-, / ^ is the probability distribution for test score ^given a representation of the audio signal / and a representation of the reference audio signal - for the audio signal / , and 0 indicates internal parameters of the pre-trained DNN.
[0462] EEE-D15. A method of estimating an indication of a subjective listening score for an audio signal using a reference-free deep neural network, DNN, wherein the listening score is a score according to a predefined listening test, wherein the reference-free DNN comprises: an input stage for receiving test data items indicative of a representation of the audio signal; a plurality of layers for performing processing based on the received test data items; and an output stage for generating the indication of the listening score; and wherein the method comprises: inputting the test data items to the input stage; and determining a representation of the indication of the listening score based on an output of the output stage.
[0463] EEE-D16. The method according to EEE-D15, wherein the test data items are each indicative of a respective representation of a binaural transformation of the audio signal, wherein the audio signal is a multi-channel audio signal or an object-based audio signal.
[0464] EEE-D17. The method according to EEE-D15 or EEE-D16, wherein the method is performed by an audio streaming device, wherein the audio signal is played out by the audio streaming device.
[0465] EEE-D18. The method according to EEE-D15, wherein the indication of the listening score relates to a probability distribution of the listening score, with the output stage of the reference-free DNN being adapted for generating the probability distribution of the listening score, wherein the probability distribution emulates listening scores obtained by a plurality of listening tests for the audio signal, and wherein the probability distribution is parameterized by two or more parameters of the probability distribution.
[0466] EEE-D19. The method according to EEE-D18, wherein determining the representation of the probability distribution comprises determining at least one of a mean, a standard deviation, and a confidence interval from the output of the output stage.
[0467] EEE-D20. The method according to EEE-D19, wherein the confidence interval is determined based on the output of the output stage and a number of listeners to the listening test to be emulated.
[0468] EEE-D21. The method according to any one of EEE-D18 to EEE-D20, wherein the probability distribution relates to a Gaussian distribution parameterized by amean ^ and a variance ^^or to a logistic distribution parameterized by a mean ^ and a scale ^.
[0469] EEE-D22. The method according to any one of EEE-D15 to EEE-D21, wherein the representation of the audio signal relates to Gammatone spectrograms.
[0470] EEE-D23. The method according to any one of EEE-D15 to EEE-D22, wherein the predefined listening test is a Multi-Stimulus Test with Hidden Reference and Anchor, MUSHRA, listening test and / or MUSHRA-like listening test, but without presenting (or exposing the known) reference to the subjects.
[0471] EEE-D24. A reference-free deep neural network, DNN, for estimating an indication of a subjective listening score for an audio signal, wherein the listening score is a score according to a predefined listening test, the reference-free DNN comprising: an input stage for receiving test data items indicative of a representation of the audio signal; a plurality of layers for performing processing based on the received test data items; and an output stage for generating the indication of the listening score.
[0472] EEE-D25. The reference-free DNN according to EEE-D24, wherein the test data items are each indicative of a respective representation of a binaural transformation of the audio signal, wherein the audio signal is a multi-channel audio signal or an object-based audio signal.
[0473] EEE-D26. The reference-free DNN according to EEE-D24 or EEE-D25, wherein the reference-free DNN has been configured by initializing training weights for at least one of the plurality of layers and training the reference-free DNN by, in a training epoch among a plurality of training epochs: inputting one or more training data items, each indicative of a representation of a training audio signal and further indicative of a respective value of a listening score for the training audio signal, into the input stage of the reference-free DNN; determining respective indications of the listening score based on the one or more training data items; determining respective loss values for the one or more training data items by evaluating a loss function, wherein the loss function depends on the indication of the listening score; and adjusting one or more internal parameters of the reference-free DNN based on the determined loss values,wherein the training weights of the reference-free DNN are initialized based on a pre-trained DNN for estimating an indication of a subjective listening score for the training audio signal based on the training audio signal and a reference audio signal for the training audio signal, wherein the pre-trained DNN comprises: an input layer for receiving a representation of the training audio signal and a representation of the reference audio signal for the training audio signal; a plurality of layers for performing processing based on the representation of the training audio signal and the representation of the reference audio signal; and an output stage for generating the indication of the listening score, wherein the training weights of the reference-free DNN are initialized based on first weights of the input layer of the pre-trained DNN, which are weights associated with the representation of the training audio signal, and / or second weights of the input layer of the pre- trained DNN, which are weights associated with the representation of the reference audio signal.
[0474] EEE-D27. The reference-free DNN according to EEE-D26, wherein the initialized training weights of the reference-free DNN are weights of the input stage of the reference-free DNN.
[0475] EEE-D28. The reference-free DNN according to EEE-D26 or EEE-D27, wherein the initialized training weights of the reference-free DNN are initialized based on the first weights of the pre-trained DNN.
[0476] EEE-D29. The reference-free DNN according to any one of EEE-D26 to EEE- D28, wherein a value of each of the initialized training weights of the reference-free DNN is based on an average value of a respective first weight and a respective second weight.
[0477] EEE-D30. The reference-free DNN according to any one of EEE-D24 to EEE- D29, wherein the indication of the listening score relates to a probability distribution of the listening score, with the output stage of the reference-free DNN being adapted for generating the probability distribution of the listening score, wherein the probability distribution emulates listening scores obtained by a plurality of listening tests for the audio signal, and wherein the probability distribution is parameterized by two or more parameters of the probability distribution.
[0478] EEE-D31. The reference-free DNN according to EEE-D30 when depending on EEE-D29,wherein determining respective indications of the listening score based on the one or more training data items comprises determining respective parameters of the probability distribution based on the one or more training data items; and wherein the loss function depends on the parameters of the distribution.
[0479] EEE-D32. The reference-free DNN according to any one of EEE-D24 to EEE- D31, wherein the probability distribution relates to a Gaussian distribution parameterized by a mean ^ and a variance ^^or to a logistic distribution parameterized by a mean ^ and a scale ^.
[0480] EEE-D33. The reference-free DNN according to any one of EEE-D24 to EEE- D32, wherein the representation of the respective audio signal and the representation of the reference audio signal relate to Gammatone spectrograms.
[0481] EEE-D34. The reference-free DNN according to any one of EEE-D24 to EEE- D33, wherein the predefined listening test is a Multi-Stimulus Test with Hidden Reference and Anchor, MUSHRA, listening test and / or MUSHRA-like listening test, but without presenting (or exposing the known) reference to the subjects.
[0482] EEE-D35. An apparatus comprising a processor and a memory coupled to the processor, and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of EEE-D1 to EEE-D23.
[0483] EEE-D36. An apparatus comprising a processor and a memory coupled to the processor, and storing instructions for the processor, wherein the processor is adapted to implement the reference-free DNN according to any one of EEE-D24 to EEE-D35.
[0484] EEE-D37. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of EEE-D1 to EEE-D23.
[0485] EEE-D38. A program comprising instructions that, when executed by a processor, cause the processor to implement the reference-free DNN according to any one of EEE-D24 to EEE-D35.
[0486] EEE-D39. A computer-readable storage medium storing the program of EEE- D37 or EEE-D38.References [1] Y. Leng, X. Tan, S. Zhao, F. Soong, X. -Y. Li and T. Qin, "MBNET: MOS Prediction for Synthesized Speech with Mean-Bias Network," ICASSP 2021, pp.391-395 [2] W. -C. Huang, E. Cooper, J. Yamagishi, and T. Toda, "LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech," ICASSP 2022, pp.896-900 [3] G. Mittag, S. Zadtootaghaj, T. Michael, B. Naderi and S. Möller, "Bias-Aware Loss for Training Image and Speech Quality Prediction Models from Multiple Datasets," 202113th International Conference on Quality of Multimedia Experience (QoMEX), pp.97-102 [4] Zhi Li, Christos G. Bampis, Lucjan Janowski, Ioannis Katsavounidis, "A Simple Model for Subject Behavior in Subjective Experiments" in Proc. Int’l. Symp. on Electronic Imaging: Human Vision and Electronic Imaging, 2020, pp 131-1 - 131-14, https: / / doi.org / 10.2352 / ISSN.2470-1173.2020.11.HVEI-131 [5] Co-pending patent application “Robust Intrusive Perceptual Audio Quality Assessment based on Convolutional Neural Networks”, applicant’s docket No. D20118, filed as US provisional patent application 63 / 119,318 and international patent application PCT / EP2021 / 083531, published as WO / 2022 / 112594 [6] K. He et al., “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification, in 2015 IEEE International Conference on Computer Vision (ICCV), 2015, pp.1026–1034. [7] A. C. den Brinker et al., “An overview of the coding standard MPEG-4 audio amendments 1 and 2: HE-AAC, SSC, and HEAAC v2,” EURASIP Journal on Audio, Speech, and Music Processing, vol.2009, no.1, 2009.
Claims
CLAIMS 1. A method of an audio steaming device estimating an indication of a subjective listening score for an audio signal played out by the audio streaming device, wherein the listening score is a score according to a predefined listening test, the method comprising: obtaining metadata relating to the played out audio signal; providing, based on the obtained metadata, test data for being input into a trained reference-free deep neural network, DNN, for estimating the indication of the listening score; wherein the reference-free DNN comprises: an input stage for receiving the test data; a plurality of layers for performing processing based on the received test data; and an output stage for generating the indication of the listening score; and wherein the method comprises: inputting the test data into the input stage; and determining a representation of the indication of the listening score based on an output of the output stage.
2. The method according to claim 1, wherein the method further comprises selecting, based on the obtained metadata and from a plurality of trained algorithms, an algorithm to be applied in the trained reference-free DNN for estimating the indication of the listening score.
3. The method according to claim 2, wherein each of the plurality of trained algorithms corresponds to a respective playout mode for being used by the audio streaming device, wherein the metadata comprises information relating to the playout mode currently used on the audio streaming device.
4. The method according to claim 3, wherein the playout modes comprise a headphone mode, a discrete speaker mode and a soundbar mode.
5. The method according to any one of the preceding claims, wherein the metadata comprises information relating to one or more of the following analyses: an analysis of the playout buffer provided in the audio streaming device, an analysis of a bitrate ladder used forproviding the audio streaming device with streaming audio data, an analysis of the audio streaming device, and an analysis of playout conditions.
6. The method according to any one of the preceding claims, wherein the test data is indicative of a representation of a binaural transformation of the played out audio signal, wherein the played out audio signal relates to a multi-channel signal or an object-based signal.
7. The method according to any one of the preceding claims, wherein the test data is indicative of a representation of the audio signal and the representation of the audio signal relates to Gammatone spectrograms.
8. The method according to claim 7, wherein the representation of the audio signal comprises segments of Gammatone spectrograms, each segment being of a respective pre- determined time length in the played out audio signal, wherein the test data comprises information relating to respective segment names and / or corresponding bitrates used for the respective segments.
9. The method according to any one of the preceding claims, wherein the method comprises estimating indications of respective listening scores for a plurality of audio streaming device clusters, each cluster corresponding to a respective content delivery infrastructure, and each cluster comprising one or more audio streaming devices, wherein the method further comprises: estimating for each cluster an indication of a listening score based on metadata collected from the one or more audio streaming devices included in the each cluster, and comparing the estimated listening scores of the clusters.
10. The method according to any one of the preceding claims, wherein the method further comprises providing the estimated indication of respective listening score to a content encoding and packaging infrastructure for adjusting a number of levels used in a bitrate ladder for providing the audio streaming device with streaming audio data and / or for adjusting bitrates applied in the content encoding and packaging infrastructure.
11. The method according to any one of the preceding claims, wherein the reference- free DNN has been configured by initializing training weights for at least one of the pluralityof layers and training the reference-free DNN by, in a training epoch among a plurality of training epochs: inputting one or more training data items, each indicative of a representation of a training audio signal and further indicative of a respective value of a listening score for the training audio signal; determining respective indications of the listening score based on the one or more training data items; determining respective loss values for the one or more training data items by evaluating a loss function, wherein the loss function depends on the indication of the listening score; and adjusting one or more internal parameters of the reference-free DNN based on the determined loss values, wherein the training weights of the reference-free DNN are initialized based on a pre-trained DNN for estimating an indication of a listening score for the training audio signal based on the training audio signal and a reference audio signal for the training audio signal, wherein the pre-trained DNN comprises: an input layer for receiving a representation of the training audio signal and a representation of the reference audio signal for the training audio signal; a plurality of layers for performing processing based on the representation of the training audio signal and the representation of the reference audio signal; and an output stage for generating the indication of the listening score wherein the training weights of the reference-free DNN are initialized based on first weights of the input layer of the pre-trained DNN, which are weights associated with the representation of the training audio signal, and / or second weights of the input layer of the pre-trained DNN, which are weights associated with the representation of the reference audio signal.
12. The method according to claim 11, wherein the initialized training weights of the reference-free DNN are weights of the input stage of the reference-free DNN.
13. The method according to claim 11 or 12, wherein the initialized training weights of the reference-free DNN are initialized based on the first weights of the pre-trained DNN.
14. The method according to any one of claims 11 to 13, wherein a value of each of the initialized training weights of the reference-free DNN is based on an average value of a respective first weight and a respective second weight.
15. The method according to any one of the preceding claims, wherein the indication of the listening score relates to a probability distribution of the listening score, with the output stage of the reference-free DNN being adapted for generating the probability distribution of the listening score, wherein the probability distribution emulates listening scores obtained by a plurality of listening tests for the audio signal, and wherein the probability distribution is parameterized by two or more parameters of the probability distribution.
16. The method according to claim 15, wherein determining the representation of the probability distribution comprises determining at least one of a mean, a standard deviation, and a confidence interval from the output of the output stage.
17. The method according to any one of the preceding claims, wherein the predefined listening test is a Multi-Stimulus Test with Hidden Reference and Anchor, MUSHRA, listening test and / or MUSHRA-like listening test, but without presenting (or exposing the known) reference to the subjects.
18. An apparatus comprising a processor and a memory coupled to the processor, and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of claims 1 to 17.
19. A program comprising instructions that, when executed by a processor, cause the processor to implement the DNN according to any one of claims 1 to 17.
20. A computer-readable storage medium storing the program of claim 19.
Citation Information
Patent Citations
Random Access with Supplementary Uplink
US62631193P0
Dynamic speech enhancement component optimization
US20230419986A1
Robust intrusive perceptual audio quality assessment based on convolutional neural networks
WO2022112594A2
EP2021083531W