Audio settings of a set-top box according to the stream
The decoder box automatically optimizes audio playback by analyzing metadata, audio, and video signals to adapt settings based on stream genre, addressing the limitations of manual user intervention and ensuring reliable sound quality.
Patent Information
- Application Number
- FR2024006248
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-12
- Publication Date
- 2025-12-19
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Audio parameterization of a decoder box according to the stream
[0001] The invention relates to the field of decoder boxes.
[0002] BACKGROUND
[0003] A home multimedia system typically includes a set-top box (or STB), a television connected to the set-top box via an HDMI (High Definition Multimedia Interface) connection, and possibly additional audio playback equipment, such as satellite speakers, a soundbar, a subwoofer, headphones, etc. This additional audio playback equipment can be connected to the set-top box via wired or wireless communication methods (e.g., Bluetooth or Wi-Fi - registered trademarks).
[0004] Some recent set-top boxes are further enhanced with advanced audio functions, such as audio playback capabilities. These set-top boxes thus incorporate one or more speakers. For example, a set-top box is known to incorporate several midrange speakers (also called "medium" or "mid-range") and a bass speaker (also called a "woofer").
[0005] In an audio system, the use of several audio playback devices improves the quality of the sound rendering by enabling multichannel playback which uses the relative positions of the different devices and their particular audio characteristics.
[0006] The set-top box receives an input audio-video stream, which is, for example, an external stream from an external source: local network, satellite, cable, DVB-T (for Digital Video Broadcasting-Terrestrial), xDSL (which can be interpreted as "digital access line"), etc. The input audio-video stream is, for example, transmitted to the set-top box via a gateway. The input audio-video stream can also be an internal stream from a source internal to the set-top box, for example, from a hard disk drive (HDD).
[0007] The input audio-video stream comprises an input video signal and an input audio signal.
[0008] The decoder box distributes the input video signal by transmitting it (after decoding and appropriate processing) to the television. The decoder box distributes the input audio signal after decoding and processing by transmitting it to its own speakers if it is equipped with them, or to the television speakers, and possibly to other audio playback equipment of the audio system.
[0009] The aim is to optimize the sound output of the audio system integrating the set-top box and, in particular, to optimize the sound output according to the audio-video stream being broadcast. By optimizing the sound quality according to the content being broadcast, the user experience is significantly improved.
[0010] Audio playback equipment, particularly soundbars, is known to offer several "audio" modes. By selecting a specific audio mode, the user can adapt certain parameters of the audio playback chain to the content being played.
[0011] This system has two main disadvantages.
[0012] First of all, it requires manual intervention by the user, which, on the one hand, is relatively restrictive, and on the other hand, may deter some non-“experienced” users who may be reluctant to make their own adjustments.
[0013] Moreover, this system ultimately proves to be quite unreliable and not always suitable for the broadcast stream.
[0014] OBJECT
[0015] The invention aims to optimize the sound rendering of an audio playback device integrated into or connected to a decoder box, by adapting the sound rendering to the audio-video stream broadcast automatically, quickly and reliably.
[0016] SUMMARY
[0017] To achieve this goal, a decoder box is proposed, arranged to distribute an input stream including an input audio signal, the decoder box comprising a processing unit in which the following are implemented: - a configuration module designed for: • perform and / or control real-time analyses on at least two distinct data sources relating to the input stream, the data sources being chosen from metadata associated with the input stream, a current audio signal from the input audio signal, and, if the input stream also includes an input video signal, at least one target image from the input video signal; • define, based on the results of these analyses, a genre of the input stream, the genre being associated with audio parameters; - a configuration module, arranged to dynamically adapt, using audio parameters, a setting of at least one audio playback device integrated into or connected to the decoder box and including at least one speaker, so as to optimize a sound rendering of said audio playback device according to the type of input stream.
[0018] The parameterization module therefore performs and / or controls analyses on several distinct data sources to define the type of input audio-video stream, and configures the audio playback device (via the configuration module) to automatically adapt the sound output to the genre of the stream being played.
[0019] This multimodal analysis, which is made possible by the decoder box's access to multiple signals and data sources, allows the audio output to be adapted quickly and reliably to the broadcast stream.
[0020] A decoder box as previously described is further proposed, in which the analysis of each data source results in an estimation of the gender, and in which the parameterization module is arranged to implement a decision algorithm to define the gender of the input stream from the gender estimates.
[0021] A decoder box as previously described is also proposed, the configuration module being arranged to control an initial analysis of the metadata, which includes the following steps:
[0022] - grouping texts from metadata to produce an aggregated text;
[0023] - perform a single inference, for each program broadcast via the input stream, from a first classification model, by applying the aggregated text as input to said first classification model, to produce a first estimate of the genre.
[0024] A decoder box as previously described is also proposed, in which the first classification model uses a transformer.
[0025] A decoder box as previously described is also proposed, in which the execution of the single inference is carried out on a remote server.
[0026] A decoder box as previously described is also proposed, in which the parameterization module is arranged to perform a second analysis on the current audio signal, which includes the execution of at least one inference of a second classification model, by applying the current audio signal as input to said second classification model.
[0027] A decoder box as previously described is also proposed, in which the parameterization module is arranged to perform the second analysis on the current audio signal, to execute inferences of the second classification model repeated regularly.
[0028] A decoder box as previously described is also proposed, in which the second classification model is a convolutional neural network of the YAMNet or VGGish type.
[0029] A decoder box as previously described is also proposed, the configuration module being arranged to perform the second analysis and to perform and / or control at least one other analysis on at least one other data source, the configuration module being arranged so that, if the second analysis results in an estimate of the type which remains constant for a first predefined period, assign to the gender of the input stream, at the end of the first predefined duration, the value of said gender estimate regardless of the result of at least one other analysis.
[0030] A decoder box as previously described is also proposed, the parameterization module being arranged so that, if the second analysis results in a gender estimate which remains constant for a second predefined duration less than the first predefined duration, and if the gender estimate produced by at least one other analysis is identical to the gender estimate of the second analysis during the second predefined duration, it confers to the gender of the input stream, at the end of the second predefined duration, the value of said gender estimate.
[0031] A decoder box as previously described is also proposed, in which the input stream also includes an input video signal, the parameterization module is arranged to perform a third analysis, on at least one target image, which includes the execution of at least one inference of a third classification model, by applying at least one target image as input to said third classification model.
[0032] A decoder box as previously described is also proposed, in which the parameterization module is arranged to perform the third analysis on at least one target image, to execute inferences of the third classification model repeated regularly.
[0033] A decoder box as previously described is also proposed, in which the third classification model is a convolutional neural network of the MobileNet or CLIP type.
[0034] A decoder box as previously described is also proposed, the processing unit further implementing a control module arranged to define at least one control parameter intended to optimize the use of resources of the parameterization module and therefore of the decoder box, the parameterization module being arranged to acquire the control parameter and to adapt the execution of at least one analysis according to at least one control parameter.
[0035] A decoder box as previously described is further proposed, in which at least one analysis includes the execution of inferences of at least one previously trained classification model, and in which at least one control parameter includes a frequency of the execution of the inferences of said model.
[0036] A decoder box as previously described is also proposed, in which at least one control parameter includes a utilization rate of a processor of the processing unit.
[0037] A parameterization method is also proposed, implemented in the parameterization module of the decoder box's processing unit as previously described, and comprising the steps of: • perform and / or control real-time analyses on at least two distinct data sources relating to the input stream, the data sources being chosen from metadata associated with the input stream, a current audio signal from the input audio signal, and, if the input stream also includes an input video signal, at least one target image from the input video signal; • define, from the results of these analyses, a genre of the input stream, the genre being associated with audio parameters.
[0038] A computer program is also proposed comprising instructions which lead the parameterization module of the decoder box processing unit as previously described to execute the steps of the parameterization process as previously described.
[0039] A computer-readable recording medium is also proposed, on which the computer program as previously described is recorded.
[0040] The invention will be better understood in the light of the following description of a particular, non-limiting embodiment of the invention. Brief description of the drawings
[0041] Reference will be made to the attached drawings, among which:
[0042] [Fig. 1] [Fig. 1] represents a decoder box and a television;
[0043] [Fig.2] [Fig.2] represents the set of information sources, the module of control, the set of data sources, the parameterization module and the configuration module;
[0044] [Fig.3] [Fig.3] represents the parameterization module according to one embodiment;
[0045] [Fig.4] [Fig.4] represents sub-modules of the control module according to a mode of realization;
[0046] [Fig.5] [Fig.5] represents the interactions, according to one embodiment, between the set of information sources, the control module and the parameterization module;
[0047] [Fig.6] [Fig.6] represents steps in a process implemented by the control module to decide whether a flow has been started or stopped;
[0048] [Fig.7] [Fig.7] represents steps in a decision-making process implemented in the control module;
[0049] [Fig.8] [Fig.8] represents the analyses carried out and / or controlled by the parameterization module;
[0050] [Fig.9] [Fig.9] represents steps of the first analysis, and sub-modules carrying out said steps;
[0051] [Fig. 10] [Fig. 10] is a table which represents an example of the result of the first analysis;
[0052] [Fig. 11] [Fig. 11] represents steps of the second analysis, and sub-modules carrying out said steps;
[0053] [Fig. 12] the [Fig. 12] is a table which represents an example of the result of the second analysis;
[0054] [Fig. 13] [Fig. 13] represents steps of the third analysis, and sub-modules carrying out said steps;
[0055] [Fig. 14] the [Fig. 14] is a table which represents an example of the result of the third analysis;
[0056] [Fig. 15] [Fig. 15] represents the method for determining gender estimates from preliminary estimates. DETAILED DESCRIPTION
[0057] With reference to [Fig. 1], the decoder box 1 is here connected to a television 2 by an HDMI link 3.
[0058] The decoder box 1 incorporates an audio playback device which includes at least one, here two speakers 4. The decoder box 1 also includes audio components 5, which enable the shaping of digital audio signals, the transformation of them into analog audio signals, and the application of these analog audio signals to the input of the speakers 4.
[0059] The decoder box 1 includes communication means 6 which enable it to communicate with other equipment in the multimedia installation in which the decoder box 1 is integrated: television 2, gateway, satellite speakers, etc. The communication means 6 enable the decoder box 1 in particular to communicate with one or more remote servers 16 on a network such as a cloud.
[0060] The decoder box 1 broadcasts an input stream F.
[0061] The input stream F can be an external stream from a source external to the decoder box 1, which the decoder box 1 receives through the communication means 6. The input stream F can also be an internal incoming stream from a source internal to the decoder box 1. Known examples of external and internal sources have been cited earlier.
[0062] The input stream F is here an audio-video stream (but this is not mandatory: it could be an audio-only stream).
[0063] Here, "audio-video stream" means any signal comprising at least one video signal and at least one audio signal associated with the video signal, the signals being intended to be broadcast in a synchronized manner. The input audio-video stream therefore comprises an input video signal V and an input audio signal A. An "audio-video stream", As understood here, it can therefore correspond to objects that can be designated by a person skilled in the art using terms such as media, stream, multimedia stream, multimedia content, etc. The decoder box 1 also integrates a processing unit 7.
[0064] The processing unit 7 is an electronic and software unit. The processing unit 7 comprises at least one processing component 8, which is for example a "general-purpose" processor, a processor specializing in signal processing (or DSP, for Digital Signal Processor), a processor specializing in artificial intelligence algorithms (of the NPU type, for Neural Processing Unit), a microcontroller, or a programmable logic circuit such as an FPGA (for Field Programmable Gate Arrays) or an ASIC (for Application Specified Integrated Circuit).
[0065] The processing unit 7 also includes one or more memories 9, connected to or integrated into the processing component(s). At least one of these memories 9 forms a computer-readable storage medium on which is stored at least one computer program comprising instructions that lead the processing unit 7 to execute at least some of the steps of the parameterization and control processes that will be described.
[0066] The processing unit 7 performs all the functions of a conventional decoder box: acquisition of the input audio-video stream, decoding of the input audio signal and the input video signal, processing, encoding, transmission to the television and to the audio playback device(s), etc.
[0067] The processing unit 7 cooperates with the audio components 5 of the audio playback device to output the input audio signal. Thus, in this case, the speakers 4 of the decoder box 1 output the input audio signal A from the input audio / video stream F, the input video signal V of which is output by the television 2. The input audio signal A can be a multichannel audio signal. The processing unit 7 can manage multichannel playback and synchronization with the television 2. The multichannel audio signal can include at least one more audio channel than the audio system has speakers. Optionally, the additional channels can be dynamically generated from a reduced number of original channels by a virtualization system.
[0068] The processing unit 7 further implements a configuration module 10, a parameterization module 11, and a control module 12. As can be seen in [Fig.2], the parameterization module 11 cooperates with a set of data sources 14, and the control module 12 cooperates with a set of information sources 15.
[0069] The configuration module 10 is intended to configure the audio playback device of the decoder box 1. The configuration module 10 makes adjustments to the audio components 5, which notably allow the acoustic rendering to be adapted. Speakers 4. Audio parameters include mechanical protection treatments for speakers 4 (audio compressor), gain adjustments for bass and treble frequencies, creation of additional channels from other channels present in the source data (upmixing), etc. Configuration module 10 may include an equalizer configured to apply processing to the frequencies of audio signals.
[0070] The parameterization module 11 performs and / or controls real-time analyses of at least one data source and, advantageously, at least two distinct data sources relating to the input audio-video stream F, in order to define a genre for the input audio-video stream F. The genre belongs to a predefined list of genres. The predefined list includes, for example, the genres "Sport," "Music," and "Voice." Each genre is associated with audio parameters that form an audio profile.
[0071] Here, the data sources are chosen from the following sources: metadata 14a associated with the input audio-video stream F, a current audio signal 14b from the input audio signal A, and at least one target image 14c from the input video signal V. The at least one target image includes, for example, the current image (i.e., the image being broadcast at the present time), as well as possibly one or more past images.
[0072] The parameterization module 11 will therefore analyze several types of data from different data sources relating to the stream, in order to accurately recognize the genre of the input audio-video stream F. The audio parameters are defined by the parameterization module 11 according to the stream, and constitute an audio profile associated with the genre. The parameterization module 11 transmits the audio parameters to the configuration module 10. Alternatively, the parameterization module 11 transmits to the configuration module 10 an identifier of the audio profile to be taken into account.
[0073] The configuration module 10 then dynamically adapts, using the audio parameters defined by the parameterization module 11, the setting of the audio playback device integrated into the decoder box 1, so as to optimize the sound rendering of said audio playback device according to the type of the input audio-video stream F.
[0074] The parameterization module 11 therefore performs a "multimodal" analysis, drawing on as much information about the stream as possible and available in the decoder box 1. This allows the audio output configuration to be reliably adapted to the broadcast content in a minimum amount of time (fast convergence). This analysis is multimodal in that it uses several distinct data sources connected to the input audio-video stream F. Taking into account several data sources 14 accelerates convergence for determining the audio configuration parameters. The greater the number of sources used, the faster the convergence. The faster the audio output configuration, the more reliable and precise it can be. The data sources used therefore include at least two sources among the played audio, text (metadata of the stream, and for example the video title, artist, electronic program guide, etc.) and one or more images (for example a decoded image from the input video signal).
[0075] The parameterization module 11 can either carry out analyses itself, therefore using the resources of the decoder box 1, or control analyses which are carried out in an external equipment, for example in a server 16 of the cloud 17.
[0076] The control module 12, for its part, controls the parameterization module 11 according to events relating to the broadcasting of the input audio-video stream F, in order to ensure that the resources of the parameterization module 11 and therefore of the decoder box 1 are used more efficiently.
[0077] By "resources" we mean here computing resources, carried out for example by a processor of the processing unit 7 in which the parameterization module 11 is implemented, and / or memory resources.
[0078] Optimizing resource usage reduces the power consumption of the parameterization module 11 and therefore of the decoder box 1, and frees up resources for existing or new tasks.
[0079] These events are transitions between active / being activated and inactive / being deactivated states of stream broadcasting: reading, stopping reading, etc.
[0080] An example of the implementation of the parameterization module 11 is illustrated in [Fig. 3]. In this example, the parameterization module 11 is configured to obtain a decoded image 14c, a decoded audio channel 14b, and the description of the TV program 14a associated with the incoming audio-video stream being broadcast. The parameterization module 11 analyzes this data respectively so as to provide, for each source, a specific genre from among "Sport," "Music," and "Voice."
[0081] The parameterization module 11 therefore here controls a first analysis 18a of the metadata 14a, performs a second analysis 18b on a current audio signal 14b from the input audio signal, and performs a third analysis 18c on at least one target image 14c from the input video signal.
[0082] Each analysis 18 results in a gender estimate: the first analysis 18a results in a first gender estimate RI of the input audio-video stream, the second analysis 18b results in a second gender estimate R2(t), and the third analysis 18c results in a third gender estimate R3(t). The parameterization module 11 then implements a decision algorithm 20 to define the gender G of the input audio-video stream F based on the gender estimates.
[0083] As we have seen, three data sources are used here by the parameterization module. It would be possible to use only two data sources, or more than three data sources.
[0084] For the first analysis 18a, the information source considered includes metadata associated with the input audio-video stream. This metadata comes, for example, from the electronic program guide (EPG).
[0085] In the broadcast of the input audio-video stream (satellite, cable, IP, terrestrial), the EPG is standardized according to the DVB EN300468 standard and offers two descriptors contained in the EIT table: Short event descriptor and extended event descriptor. The set of these descriptors can contain: - the name of the program;
[0086] - the start and end time of the program;
[0087] - the type of program (for example, news, sports, film, etc.);
[0088] - a short description and a long description of a program;
[0089] - information about the producer, the names of the actors, the genre and other textual information.
[0090] In the case of applications such as YouTube and Spotify (registered trademarks), the media aggregator, which will be described below, can make available the title of the media stream, the name of the artist, the duration of the stream, a summary and other metadata related to the stream launched on the set-top box 1.
[0091] The first analysis 18a is performed only once per broadcast program. "Program" is understood to mean, for example, a film, an episode of a television series, or a particular sporting event (match, race, etc.). The first RI estimation of this type is therefore not time-dependent (even though it can be considered to be performed dynamically since it is repeated with each change of broadcast program).
[0092] For the second analysis 18b, the data source considered is a current audio signal derived from the input audio signal. By "current," we mean "currently being broadcast." The input audio-video stream F, originating from a source internal or external to the decoder box 1, is processed to extract the audio tracks of the current program. These audio tracks are generally encoded in a particular format (e.g., AC3, AAC, etc.). The audio tracks are decoded by the decoder box 1 to obtain audio tracks in PCM (Pulse-Code Modulation) format. These audio tracks form the current audio signal, derived from the input audio signal, on which the second analysis 18b is performed.
[0093] The second analysis 18b is performed at least once per broadcast program, and is here repeated regularly at a frequency which, as will be seen, can be adapted by control module 12. The second estimate R2(t) of the kind therefore depends on time.
[0094] For the third analysis 18c, the data source considered includes at least one target image of the input video signal.
[0095] The input audio-video stream F, possibly received via the communication means 6 of the decoder box 1, is processed to extract the video from the current program. This video is generally encoded in a particular format (e.g., H265, H264, VP9, MPEG, etc.). The decoder box 1 decodes this video to obtain at least one image, and for example, a sequence of images in raw ARGB or YUV format, on which the third analysis 18c is performed.
[0096] The third analysis 18c is carried out at least once per broadcast program, and is here repeated regularly at a frequency which, as we shall see, can be adapted by the control module 12. The third estimate of the kind R3(t) therefore depends on time.
[0097] According to a particular embodiment, the parameterization module 11 is not only configured to classify the input audio-video stream F according to several genres (for example, "Sport", "Music", "Voice"), but also to sub-categorize each genre into sub-genres. For example, for the "Music" genre, the parameterization module 11 is capable of estimating a music sub-genre selected from: Rock, Classical, Jazz, Blues, RnB / Pop.
[0098] The parameterization module 11 performs certain analyses entirely, and controls others (that is, it commands the external entity in charge of the analysis (for example, a server 16 in the cloud 17), transmits the signals to it, acquires the results, etc.). The parameterization module 11 can also perform only part of an analysis, with the remainder of the analysis being performed by the external entity.
[0099] The configuration module 11 uses potentially significant resources of the decoder box 1, which may result in significant power consumption.
[0100] The control module 12 will control the parameterization module 11 in order to optimize the power consumption of the parameterization module 11 and therefore of the decoder box 1. For this purpose, the control module 12 defines at least one control parameter Pc (visible on the [Fig.2]) intended to control the power consumption of the parameterization module 11, and the parameterization module 11 acquires each control parameter and adapts the execution of at least one analysis according to said control parameter Pc.
[0101] To control the power consumption of the parameterization module 11, the control module 12 can, for example, control the frequency of the analyses performed by the parameterization module 11. The control parameter Pc is then the value of this frequency. As will be seen, the parameterization module 11 is arranged to to perform classification model inferences. In this case, the frequency of analyses is the frequency of execution of said inferences.
[0102] To control the power consumption of the parameterization module 11, the control module 12 can also control a usage rate of a processor 8 of the processing unit 7, in which the parameterization module 11 is implemented. The control parameter is therefore the usage rate of the processor 8. The parameterization module acquires this rate, which is a maximum processor usage setpoint, and adapts its analyses according to said setpoint.
[0103] The control module 12 therefore makes it possible to optimize the use of the hardware resources (e.g. processor(s), memory(s)) of the processing unit 7 which implements the parameterization module 11, which makes it possible to reduce the power consumption of the parameterization module 11 and the decoder box 1, and to avoid the undesirable slowing down of the other software layers of the decoder box 1.
[0104] The control module 12 detects for this purpose the occurrence of at least one current event among a set of predefined events, relating to the broadcasting of the input audio-video stream F, and controls the parameterization module 11 according to said current event in order to optimize the power consumption of the parameterization module 11 and therefore of the decoder box 1. The control module 12 therefore adapts the control parameter Pc according to the current event.
[0105] The events are therefore detectable on the decoder box 1 and are of internal or external origin, that is to say that they can either be generated by the decoder box 1, or received by the decoder box 1 but from an entity external to the decoder box 1, such as for example data from the electronic program guide sent by operators on radio transmissions.
[0106] With reference to [Fig.4], the control module 12 comprises three sub-modules: a listening sub-module 12a, an analysis sub-module 12b and a configuration sub-module 12c.
[0107] During an initialization step E0, the control module 12 subscribes to the information sources 15. Preferably, these sources 15 are written in a predefined configuration file in the source code of the control module 12.
[0108] The listening sub-module 12a is configured to continuously listen and detect events from the set of information sources 15. As soon as an event is detected, the listening sub-module 12a transmits it to the analysis sub-module 12b and resumes waiting for new events.
[0109] The analysis sub-module 12b is configured to analyze one or more events previously detected and provided by the listening sub-module 12a.
[0110] The configuration sub-module 12c is configured to determine, based on the result of the analysis of the analysis sub-module 12b, the configuration instructions to be applied to the input of the parameterization module 11 in order to configure it.
[0111] As already mentioned, the control module 12 continuously listens for events Ev detected by the set of information sources 15, by means of the listening sub-module 12a. These events can come from several distinct information sources.
[0112] To obtain a robust decision, these sources must be as varied as possible, in terms of origin (i.e., internal or external to the decoder box 1) and in terms of "software level" (e.g., system level, driver, etc.). Thus, the listening sub-module 12a is configured to listen to a diverse set of events from different information sources.
[0113] In the present embodiment, three information sources 15 are considered:
[0114] - media session aggregator 15a of the decoder box operating system;
[0115] - electronic guide to programs 15b;
[0116] - audio driver and / or video driver 15c of the decoder box 1.
[0117] The media session aggregator 15a is a particular software component, which is present in the decoder box operating system 1.
[0118] For example, this software component is the MediaSessionService available in the Android TV operating system (registered trademark).
[0119] This software component is particularly advantageous for playing the role of information aggregator for obtaining information related to the input audio-video stream F, insofar as it is able to provide information on the playback state of the streams (for example, states "Pause", "Play") as well as metadata relating to the content of the input audio-video stream (for example, the title of the content, the name of the artist, etc.).
[0120] The EPG 15b is an external source of information to the set-top box 1. It is provided by an operator and contains textual information relating to the TV program being played from the broadcast source (e.g. satellite, cable or DTT). For example, this information includes the name of the program being watched.
[0121] In a known manner, the EPG processes information from: - DVB EIT (Digital Video Broadcasting - Event Information Table) tables, if the EPG is broadcast using a broadcast signal on media such as satellite, cable or terrestrial network, and / or - from a server on an IP (Internet Protocol) network
[0122] This information relates to a digital television program. It indicates, for example, the start and end of the program.
[0123] In this embodiment, a local TV program database is implemented in the set-top box 1 and is continuously updated by the EPG. Advantageously, this database includes information on the current event (i.e., E1T Present) and the next event (i.e., E1TEollowing) of a television channel. Preferably, this database is queryable to retrieve program-related information. Software entities external to this database (e.g., processes or lightweight processes (threads)) can subscribe to events, such as the transition from a current event to the next event for a given channel.
[0124] The program start and end information can be used to instantly apply a control configuration of parameter module 11, corresponding to a start and end of an audio stream.
[0125] With regard to audio / video driver information 15c (d river), it is known that, in order to start audio or video content on the decoder box 1, the application in charge of this start communicates directly or indirectly (via the operating system of the decoder box 1) with the "driver layer" (driver) to allocate resources and start decoding and display.
[0126] By querying or monitoring the audio driver and / or video driver of the decoder box 1, it is possible to detect the launch of content on the decoder box 1.
[0127] For example, an audio driver notification provides information indicating the launch of an audio-only program. This notification can be associated with a driver notification related to the set-top box 1.
[0128] In all cases, it is possible to detect the start of an audio or video program (including a sound component) on the basis of driver notifications.
[0129] In order to perform the reading of the input audio-video stream F, a master application is required to control all the software actors (for example the graphics part, the audio decoding and the video decoding).
[0130] It is possible to differentiate a so-called Broadcast application, that is to say powered by a broadcast source carrying media based on standards such as DVB EN 300 468 and ISO / IEC 13818, from a so-called OTT (Over The Top) application based on streaming technologies - which can be translated as "streaming playback" (for example HTTP, MPEG DASH, Microsoft Smooth Streaming).
[0131] The distinction between a Broadcast application and an OTT application can be used to define the information sources taken into account.
[0132] Thus, the control module 12 selects at least one information source 15, to detect the occurrence of the current event, according to a source of the input audio-video stream F.
[0133] For example, for a Broadcast application, EPG 15b is more likely to be available and used, whereas for an OTT application, such as Netflix, information from the system's media aggregator 15a will be preferred.
[0134] The control module 12 therefore detects the occurrence of a current event Ev among a set of predefined events, relating to the broadcast of the input audio-video stream F.
[0135] The predefined set of events includes transitions between states relating to the reading of the input audio-video stream F.
[0136] The predefined set of events includes at least:
[0137] - a first transition, from an active state or activation of the input audio-video stream, to an inactive or deactivated state, and / or
[0138] - a second transition, from an inactive state or deactivation of the audio-video stream input, to an active or activation state, and / or
[0139] - a third transition between a first active state, in which the input flow contains a first broadcast program, to a second active state, in which the input stream contains a second broadcast program.
[0140] Here, the states relating to the reading of the input audio-video stream are as follows: - "Reading stopped" state (inactive state): stream reading is stopped, no hardware resources are used; - "Stopping" state (deactivation state): the flow is being stopped; - "Startup" state (activation state): the stream reading is starting and the allocations to hardware resources are being committed; - "reading" state (active state): the stream is being read, all hardware resources are correctly allocated and used; - "Pause" state (active state without inference): reading the stream is momentarily stopped, all hardware resources remain active and allocated.
[0141] Each transition between these states can be an event which leads the control module 12 to issue a command to the parameterization module 11, and thus to modify the control parameter Pc.
[0142] If the control parameter Pc used is the frequency of analyses, the control module 12 reduces the frequency of analyses performed by the parameterization module 11 when the first transition occurs, and increases said frequency when the second or third transition occurs.
[0143] The control module 12 stops the analyses 18 when the input flow goes into the inactive state.
[0144] If the control parameter Pc used is the CPU utilization rate, the control module 12 reduces the CPU utilization rate setpoint when the first transition occurs, and increases said setpoint when the second or third transition occurs.
[0145] The control module 12 assigns a zero value to said setpoint when the input flow passes into the inactive state.
[0146] It is noted that here, the control module 12 is also arranged to control the parameterization module 11, so as to optimize the use of the resources of the parameterization module 11 and therefore of the decoder box 1, depending on a convergence or divergence of the analyses carried out by the parameterization module 11.
[0147] If the analysis of the different data sources 14 converges towards the same gender estimate, the control module 12 decreases the control parameter. Conversely, if the analysis diverges, because the data sources 14 provide gender estimates that vary over time, the control module 12 increases the control parameter.
[0148] With reference to [Fig.5], we now consider a particular embodiment of the decoder box 1. In this embodiment, the operating system of the decoder box 1 is Android TV.
[0149] When the set-top box 1 is started, the control module 12 is started and begins subscribing to the available services of the information sources 15 that provide the information necessary for detecting the start or stoppage of the input audio-video stream. Preferably, the listening sub-module 12a of the control module 12 listens for events from the three sources described above: media session aggregator, electronic program guide, and audio and / or video drivers of the set-top box.
[0150] Regarding the media session aggregator, MediaSessionService provides an "asynchronous return function" which is triggered when an application starts an input audio-video stream.
[0151] This asynchronous return function returns a list of MediaController objects. Each MediaController represents one of the currently active audio-video streams. Each audio-video stream started on Android ZV therefore has its own MediaController.
[0152] This MediaController also makes available events on the current stream, such as the change of the playback state (for example, from Play state to Pause state).
[0153] Regarding driver information, when content starts, the driver layer of the decoder box 1 reserves access to the hardware to perform audio-video decoding. This access is stored through a memory reference for each resource used. It is possible to query or subscribe to this reference to obtain information about the stream being decoded. For example, it is possible to query the video driver of the decoder box 1 via its reference to obtain information about the encoded video being decoded.
[0154] As described above, the TV program database can notify a transition between the "current" and "next" events of a TV program in progress. It can also, if requested, notify such transitions on a predefined subset of channels belonging to the TV service plan of the set-top box 1 to which the user is subscribed.
[0155] The Ev events detected by the listening sub-module 12a are sent to the analysis sub-module 12b to analyze them and decide whether it is a start or a stop of an audio-video stream on the decoder box 1.
[0156] With reference to [Fig. 6], in the case of a system event, the analysis submodule 12b determines the type of event (step E1). If the event originates from the MediaSession service of the Android TV operating system, submodule 12b first checks the size of the resulting MediaController list. If it is empty, submodule 12b assumes that there is no longer an active stream on set-top box 1 (step E3). Otherwise, submodule 12b scans the list and counts the number of active objects in the list, i.e., the number of MediaControllers whose state is "playing" (step E4). Analysis submodule 12b compares this number with 0 (step E5). If this number is equal to 0, it assumes that there is no audio / video stream in the playback state on set-top box 1 (step E6). Otherwise, it considers that a flow has started (step E7).
[0157] At step El, if the received event is from a MediaController, the analysis submodule 12b checks the nature of the asynchronous call.
[0158] Analysis submodule 12b checks if the active MediaController is being destroyed (step E8). If so, this means that the stream attached to that controller has stopped (step E9). Otherwise, analysis submodule 12b checks if a state change has occurred (step E10). A state change to the "reading" state means that the stream is being read (step E10). A different state change means that the stream has stopped (step E10).
[0159] In both cases, the analysis submodule 12b checks whether a "minimum" of metadata is available; otherwise, the event will be ignored. By "minimum" is to understand at least the title and duration of the content being streamed from the input audio-video feed.
[0160] For events from the EPG 15b database, the EPG makes available transition events from one current program to the next for all EPG channels.
[0161] Thus, after receiving an EPG event, the analysis sub-module 12b checks if the user is currently watching the channel concerned by this event by querying the operating system of the set-top box 1 (here Android TV): step E13. If not, the event is ignored. Otherwise, sub-module 12b considers that a program, therefore an audio-video stream, has ended and that a new one has started (step E14).
[0162] Regarding the "pilot" events 15c, the memory reference of the decoder box 1 contains information on the use of the decoder box 1. The sub-module 12b checks the type of event (step E15). If the event received from this reference is a start-up of the audio-video decoder hardware block of the decoder box 1, this means that a stream is being played (step E16); a release of the audio-video decoder, indicated as "Stop", means that the stream has stopped (step E17).
[0163] We are now interested, with reference to [Fig.7], in the decision-making implemented in control module 12.
[0164] The listening sub-module 12a detects Ev events.
[0165] The analysis submodule 12b checks for each event whether it should be ignored or not (step E20).
[0166] The analysis submodule 12b of the control module 12 begins by accumulating a set of non-ignored events (e.g., Ev1, Ev2, and Ev3). If, after a predefined time Tl (e.g., Tl = 500 ms), no further events are received, the analysis submodule 12b checks the accumulated events to make a decision.
[0167] Several decision-making methods are conceivable, which are for example based on information from the last event received.
[0168] As we have seen, depending on the system context (for example, in "Application" mode with a dedicated application for playing audio-video content, or in "Direct" mode in the context of a DVB-type audio-video source), one information source may be preferred over another. This choice is justified by the fact that certain information is more frequently available in one context than in another. For example, when using a streaming application, events from MediaSession may be preferred because they are more frequently available than other information such as the EPG. In "Live Program" mode, the system may use information from the EPG. This configuration, depending on the system context, may be predefined or specified by the user via a menu in a graphical interface.
[0169] Optionally, the predefined time Tl is such that Tl = 0. This means that the analysis submodule 12b makes a decision for each event received (i.e., without waiting for several events to accumulate).
[0170] Optionally, and in the case of a MediaSession / MediaController event, the analysis submodule 12b can perform additional checks on the metadata present in these events to detect whether or not it is an advertisement. The distinction between a program deemed "interesting" for the viewer and a less important "advertising" program can be made to apply a different configuration depending on the type of program. In the case of an advertisement, the analysis submodule 12b can, for example, consider that there is no stream currently playing. These checks also depend on the context of the system being used. For example, in an "application" mode with the Spotify application (registered trademark), a "flag" (or tag) named "ADVERTISEMENT" is present in the metadata to indicate that it is an advertisement.In this case, if it is the last Media Session event in the history, the system ignores the other events.
[0171] The control module 12 then configures the parameterization module 11, using the consumption parameters mentioned above, which are deduced from the results of analyses carried out by the analysis sub-module 12b described above.
[0172] If the analysis submodule 12b of the control module 12 determines that a flow is active, the configuration submodule 12c increases the number of analyses per second performed by the parameterization module 11, for example, by setting it to 2. Otherwise, the configuration submodule 12c configures the parameterization module 11 to 0.1 analyses per second (i.e., one analysis every ten seconds).
[0173] The control module 12 can also define a maximum CPU (Central Processing Unit) usage target. For example, this target is predetermined in submodule 12c as being equal to 5% in the case of an ongoing flow, and 1% otherwise.
[0174] We are now more particularly interested in the analyses carried out and controlled by the parameterization module 11.
[0175] Here, parameterization module 11 classifies each input audio-video stream F according to a genre from among several predefined genres, namely "Sport", "Voice", and "Music". Parameterization module 11 can also sub-categorize each genre into subgenres. For example, for the "Music" genre, the parameterization module determines a music subgenre from among: Rock, Classical, Jazz, Blues, RnB / Pop.
[0176] With reference to [Fig.8], the analysis begins with an initialization step E30 which includes: - retrieval of the PC control parameter(s) from control module 12 (here the number of inferences per second and / or the CPU usage rate); - loading the reference values into volatile memory for the steps to follow (see sub-modules lia, 11b, 11c described below), from non-volatile memory; - the initialization of the algorithms of the different sub-modules of module 11; - optionally, an initialization of the process to limit CPU consumption by available means, for example by the operating system (e.g. the cpulimit application).
[0177] In order to obtain a preliminary estimate of the genre of the input audio-video stream being played, the lia submodule begins by analyzing the text of the metadata 14a of this stream. The lia submodule produces a first preliminary estimate Rbl of the genre of the input audio-video stream F. The first preliminary estimate Rbl includes probabilities of belonging to the different classes (i.e., the different genres).
[0178] Then, the analyses of the "audio" source 14b and the "image" source 14c are performed in parallel by sub-modules 11b and 11c, respectively. As output from these analyses, two new preliminary estimates, Rb2(t) and Rb3(t), of the stream genus are obtained and are detailed below. These preliminary estimates are again probabilities of belonging to the different classes.
[0179] For obtaining the preliminary estimates Rb2(t) and Rb3(t), the analysis is done continuously, as long as the stream is being read, unlike the Rbl estimate, which is obtained by an analysis performed only once per program being read.
[0180] As we shall see, the first analysis 18a, the second analysis 18b and the third analysis 18c use machine learning models, which here are classification models.
[0181] The classifications of the first analysis 18a and the third analysis 18c are possibly so-called Zero-Shot classifications. The classification problem is one of the classics of machine learning. Classification consists of training a neural network to predict the type of a new instance. The type predicted by the network is a class from a fixed set of classes specific to the network. Thus, a network trained to recognize cats and dogs from an input image is not able to recognize a turtle (unpredictable behavior).
[0182] In the case of Zero-Shot classifications, the classes and the instance to be classified are applied as inputs to the network. The network is therefore able to predict the probability distribution of this instance's membership in these classes. The network here is not "theoretically" limited to a fixed set of classes. The model can perform the classification with instances or classes not encountered during training.
[0183] In this implementation, an instance represents a text in the case of metadata analysis, and an image for the image analysis part.
[0184] We are now interested in the first analysis 18a.
[0185] After an input audio-video stream F is launched on the decoder box 1, the lia sub-module checks for the presence of metadata (source 14a) associated with this stream. If this metadata data is present, the parameterization module 11 performs an initial analysis 18a on this metadata to obtain a preliminary assessment of the type of input audio-video stream F.
[0186] The first analysis 18a is a text analysis that is applied to these data to deduce the first preliminary estimate Rbl. The text analysis can be performed in a cloud instance or in the decoder box 1. According to a cloud instance-based embodiment, the lia submodule uses transformer-based neural networks such as BART or Gemma.
[0187] Thus, in one embodiment, a large part of this first analysis 18a, and in particular the neural network inference, is carried out not in the processing unit 7 of the decoder box 1 but in a server 16 of the cloud 17.
[0188] Figure 9 illustrates an example of an implementation to analyze the texts of metadata 14a and make a decision on the genre of the stream within the framework of this embodiment.
[0189] This implementation is based on BART-type models, which is a transformer launched by Meta in 2019. This transformer can be trained on several Sequence to Sequence tasks (e.g. translation, text summarization, etc.).
[0190] The optional submodule 1 lal detects the language of the text. The languages used in the metadata fields are detected, for example, using a Mediapipe model (Google). The output of block 1 lal contains ak texts detected according to k data fields carried in the source 14a. The submodule liai detects the language of each text included in the metadata and checks if that language is English (step E30).
[0191] Among these k texts there are kl texts in English and k2 non-English texts, so that k = kl + k2.
[0192] If the language detected in the first step is not English, the corresponding text is translated into English in submodule 1 la2, for example by means of a BART network capable of translating between 50 different languages.
[0193] Optionally, submodule 1 la3 is configured to reduce the size of the text, for example by means of another BART network suitable for summarizing a text.
[0194] Submodule 1 la4 is configured to group the texts of the different metadata fields into a single large aggregated text, in English, for example using a predefined pattern (t employs).
[0195] Submodule 1 la5 is configured to classify the text constructed by submodule 1 la4 and recognize if the metadata corresponds to content of genre "Sport", "Music", etc.
[0196] The first analysis 18a on the metadata 14a therefore includes the step of executing a single inference, for each broadcast program, of a first classification model 30 previously trained, by applying the aggregated text as input to said first classification model, to produce a first preliminary estimate of the genre Rbl.
[0197] Text contraction by summary, carried out by submodule 1 la3, improves the results of the classification step by module 1 la5.
[0198] The first classification model 30 uses a transformer. This classification can be performed by a B ART network of Zero-Shot classification. According to other embodiments, it is also possible to use an Open Source LLM (Large Language Model) such as Gemma, or a paid service such as Gemini Pro (and the added benefit of sub-module 1 la3 "text summarization", since billing is based on the size of the text processed) to predict the genre of the stream corresponding to the analyzed metadata.
[0199] According to the present embodiment under Android TV, in the case of a Broadcast stream, the metadata 14a is obtained by concatenating the title of the TV program and Y extended descriptor present in the EIT table.
[0200] Optionally, another text analysis can be carried out using the Short descriptor instead of the extended descriptor.
[0201] Optionally, the two analyses can be performed to obtain two separate preliminary estimates Rbl.
[0202] In the case of an application-origin OTT stream (e.g., YouTube, Spotify), the metadata is obtained by concatenating all the information available in Android MediaController Metadata, preferably in the form of a predefined template. For example, in the case of a YouTube stream where the information is the title and the channel name, the text to be parsed is created using the following template:
[0203] « title : extracted title>, channel : extracted channel>”.
[0204] According to this embodiment, language detection, translation and summarization are applied to each metadata field independently of the others and before concatenation in the pattern.
[0205] According to the present embodiment, using a BART network, a set of predefined texts (classes) is used to measure their similarity to the metadata. For example, after the construction of the text to be analyzed by module 1 la4, the latter is passed to the Zero-Shot classification network of module 1 la5 along with the following expressions: "A Sports Event", "A sports match", "A News Show", "A Talkshow", "A music event", "A music video".
[0206] According to this example, this results in two expressions per audio class.
[0207] Optionally, and according to this embodiment, expressions relating to the subgenre can also be transmitted in a second step to module 1 la5 to measure similarity, such as "Rock Music", "Blues Music", etc.
[0208] According to this example, this results in an expression by subcategory.
[0209] It should be noted that the number of expressions per class is not limited and that it is possible to use a different number for each category / subcategory.
[0210] The Rbl output contains the similarity values between the metadata text and these expressions.
[0211] Figure 10 shows a table representing the Rbl output of the first analysis. The subcategories (subgenera) are classified independently of the classification of the genera.
[0212] The second analysis 18b includes an execution of inferences of a second previously trained classification model, by applying the current audio signal as input to said second classification model.
[0213] The parameterization module 11 analyzes the current audio signal "n" times per second, "n" being defined by the control module 12.
[0214] The second classification model is configured and trained to estimate the genre and / or subgenre of the current audio signal stream A and therefore of the input audio-video stream F.
[0215] The second classification model is here for example a convolutional neural network of the YAMNet type, or of the VGGish type.
[0216] The processed audio signal, decoded by the decoder box 1, is applied as input to the second classification model.
[0217] The YAMNet model is a neural network introduced and trained by researchers at Google. For example, this network is configured to take as input a single-channel PCM audio signal in 32-bit floating point, sampled at 16 kHz and with a size equal to 15600 samples (which is equivalent to a duration of 0.975 seconds).
[0218] The neural network is configured here to classify the genre of audio content among a set of 521 classes.
[0219] With reference to [Fig. 11], submodule 1 Ibl acquires the processed audio signal 14b which is applied as input to the second classification model 31, here for example the
[0220]
[0221]
[0222]
[0223]
[0224]
[0225]
[0226]
[0227]
[0228]
[0229]
[0230]
[0231]
[0232]
[0233]
[0234]
[0235]
[0236]
[0237]
[0238] YAMNet network. This analysis provides a probability distribution P over the 521 classes. In submodule 1 lb2, another list of final probability values P' is calculated as a function of P. For example, the final probabilities P' of the genres Sport, Music and Voice are calculated as follows: P Sport (Pcheering 4“ PbuII Sound 4“ Pscream-^- • • • )xKl (eg Kl = 100) P Voice P Voice / P Total P Silence Psilence / Protal P Music — PmusIc / Plotal Plotal Pvoice+Psilence-^- ^Music Optionally, smoothing of the P' values can be performed to avoid drifts. Optionally, submodule 1 lb2 also estimates the probabilities of subgenres (e.g., Rock, Blues, Jazz, etc.). To find these probabilities, a specific mapping for each genre is applied to the output of neural network 31. Submodule 1 lb2 first calculates a value P' l_ <genre>for each genre. For the Rock subgenre for example, P' l_rock(t) is the sum of all network outputs whose subgenre is Rock such as Metal, RockNRoll, etc. A similar value is then calculated for the Classic, Blues, RnB / Pop, Disco and Vocal genres. These values are then normalized over the sum of the P' l_ <genre>to obtain a probability distribution. For example, for a given genus "i": P' l_rock_norme(t) = P' l_rock(t) / Sum(P' l_i(t)) After normalization, a probability P'2_ <genre>is calculated using the following formula: P'2_ <genre>(t) = (P’2_ <genre>(t-1) + PMusic (t)* P’ l_ <genre>_norm(t)) / (1 + PMusic(t)) This formula means that the value of P' l_ <genre>_norm is reliable only when PMusic is high, therefore when it is very likely that it is really music. After finding the values P' and P'2_ <genre>, submodule 1 lb3 starts to build the response Rb2(t) as illustrated in the table in [Fig. 12]. Again, the subcategories (subgenres) are classified independently of the classification of genres. We now turn our attention to the third analysis, 18c.
[0239] The parameterization module 11 performs a third analysis 18c on at least one target image 14c from the input video signal V, said third analysis comprising an execution of inferences of a third previously trained classification model, by applying the input images of said third classification model.
[0240] The parameterization module 11 performs "n" analyses per second on at least one target image, "n" being defined by the control module 12. For each analysis, the parameterization module 11 analyzes the current image corresponding to the present time and possibly one or more past images.
[0241] The third classification model is, for example, a convolutional neural network of the MobileNet or CLIP (for Contrastive Language-Image Pretrained) type.
[0242] The MobileNet network performs image-to-image comparisons. MobileNet is a convolutional neural network architecture optimized for running on edge devices. This architecture can be trained on several tasks, including image vectorization. This task consists of transforming two similar images into two vectors that are close (for example, along the cosine distance).
[0243] The CLIP network is a neural network trained on Image / Text pairs. This network is capable of measuring the similarity between a text and an image. This network can be used for Zero-Shot classification.
[0244] According to one embodiment, a database comprising several image vectors by genre (or class) is embedded in one of the memories 9 of the processing unit 7 of the decoder box 1. These are, for example, vectors linked to images of stadiums, swimming pool, Formula 1, etc. for the genre "Sport", as well as images of concerts for the genre "Music" and images of television programs such as talk shows, news broadcasts, for the genre "Voice".
[0245] The images used for this third analysis are screenshots of the content that the user is currently viewing.
[0246] With reference to [Fig. 13], each target image 14c (screenshot for example) is first transformed into a vector by the third classification model 32 (submodule llcl). This vector is then compared to the vectors stored in one of the memories 9 of the processing unit 7 of the decoder box 1 (submodule 1 lc2) to construct the output R3b(t) (submodule 1 lc3).
[0247] According to this embodiment, if the user watches a football match, a decoded image capture, along with three texts "Sports Event", "Music Video", "News Studio", are transmitted to the network to calculate the similarities between the decoded image capture and the three texts. It is possible to use more than one text per category, and therefore instead of using "Sports Event", it is possible to use Football Match, Basketball Match, Fl Race, etc. These similarity values will constitute in the Rb3(t) sequence, as illustrated in the table in [Fig. 14]. Again, the subcategories (subgenera) are classified independently of the classification of the genera.
[0248] We have therefore explained how the parameterization module 11 obtains the preliminary estimates of the genre of the audio-video stream: Rbl, Rb2(t), Rb3(t). We now turn our attention, with reference to [Fig. 15], to how the parameterization module 11 determines the genre estimates from the preliminary genre estimates.
[0249] As already mentioned, the Rbl output contains the similarity values between expressions representing genres and the text constructed from the metadata. Submodule 1 la4 is configured to equate the first estimate of the RI genre with the genre of the expression having the highest similarity value to the metadata.
[0250] Optionally, if this genre is "Music", submodule 1 la4 then checks for similarities with musical subgenres. It applies the same logic to find RI.
[0251] The output Rb2(t) of submodule 11b includes a probability list P' ("Music", "Voice", "Silence") as well as a value P'Sport.
[0252] In order to find the value of the second genus estimate R2(t), the submodule 1 lb4 applies the following steps: 1. Identify the highest probability among "Music", "Voice" and "Silence". " 2. If it is "Music", with a probability greater than a predetermined threshold Cl (for example, Cl = 0.3), the value R2(t) will be "Music". 3. If it is "Voice" with P'Voice > C2 (e.g., C2 = Cl), submodule 1 lb4 checks the value P'sport- a. If P'sSport > C3 (e.g., C3 = 1), the value R2(t) will be "Sport", b. Otherwise R2(t) will be "Voice". 4. In the case where the highest probability value is "Silence", the submodule llb4 checks P'sport- a. If P'spOrt > C3, the value R2(t) will be "Sport", b. Otherwise, the value R2(t) = R2(t-1). 5. Otherwise, R2(t) = Unknown.
[0253] Submodule 1 lc4, which determines the third estimate of the genus R3(t), corresponds mutatis mutandis to that implemented for text classification. The third estimate of the genus R3(t) corresponds to the class with the highest similarity value.
[0254] As we have just seen, the parameterization module therefore controlled three analyses 18 (and fully carried out the second analysis 18b and the third analysis 18c), and a thus obtained three estimates of the type of audio-video stream F (the values RI, R2(t) and R3(t)).
[0255] The parameterization module 11 then implements the decision algorithm 20 to define the genre G of the input audio-video stream from these estimates.
[0256] It is the submodule 1 Id which implements this algorithm and which makes the decision on the genre of the stream decoded on the decoder 11.
[0257] This decision-making process is carried out using, for example, the following algorithm, the purpose of which is to calculate a confidence index in order to deduce the final genre G of the stream and therefore the audio profile to be sent to the configuration module 10:
[0258] Initialization: confidence = 0, R2Last=null
[0259] Algorithm:
[0260] If R2(t) == R2(t-1) and R2(t) != Unknown then confidence += a N
[0261] If R2(t) == R3(t) then confidence += a M (M < N)
[0262] If R2(t) == RI then confidence += a E (E < M)
[0263] R2Last = R2(t)
[0264] If R2(t) == Unknown and R3(t) == R3(t-1) and R3(t) == R2Last
[0265] If R3(t) == RI then confidence += a E (E < M)
[0266] Otherwise, confidence = 0
[0267] If confidence > 0.95 then G = R2(t) or R2Last.
[0268] Thus, the parameterization module 11 performs the second analysis 18b (on the audio signal 14b) and performs and / or controls at least one other analysis on another data source (here, two other analyses: on the metadata 14a and the images 14c). It can be seen that if the second analysis 18b results in a genre estimate that remains constant for a first predefined duration, the parameterization module 11 assigns the genre of the input audio-video stream, at the end of the first predefined duration, the value of said genre estimate regardless of the result of the at least one other analysis.
[0269] Furthermore, if the second analysis 18b results in a genre estimate that remains constant for a second predefined duration shorter than the first predefined duration, and if the genre estimate produced by at least one other analysis is identical to the genre estimate of the second analysis during the second predefined duration, the parameterization module 11 assigns to the genre of the input audio-video stream, at the end of the second predefined duration, the value of said genre estimate.
[0270] We can therefore see that the audio signal is the main data source for determining the type of stream F, and that the other sources help the decision-making process, and speed it up in the case where the image and text correspond to the audio.
[0271] Taking into account several sources of information relating to the broadcast audio content therefore allows for a faster convergence in determining the audio parameters to be applied.
[0272] Note that the value a is inversely proportional to the time during which the value "G" has not changed. This increases the stability of the process.
[0273] Optionally, a security measure is put in place to avoid frequent changes in the G value of the type.
[0274] According to one embodiment, this measure corresponds to the following algorithm:
[0275] Init: stability = 10, stable = true.
[0276] If G(t) == G(t-1):
[0277] stability += 1
[0278] if stability >=10 then stability =10
[0279] Otherwise:
[0280] stability -= 3
[0281] if stability < 0 then stability = 0
[0282] If stable:
[0283] If stability < 4 then stable = false
[0284] If ! stable:
[0285] If stability > 7 then stable = true
[0286] If stable then G is the value to send to module A3
[0287] Otherwise, a default value is sent to module A3, for example G=Music
[0288] It can therefore be seen that a gender value for the stream is taken into account only if this value remains constant for a certain time (or more precisely, if the result of a certain number of consecutive analyses is constant). Otherwise, a default value for the gender is used.
[0289] Of course, the invention is not limited to the embodiment described but encompasses any variant falling within the scope of the invention as defined by the claims.
[0290] The input stream is not necessarily, as we have seen, an audio-video stream. It may be a stream that includes only an input audio signal. In this case, the analysis does not focus on the images but only on the audio signal and possibly on the metadata to determine the type of input stream.
[0291] The invention can be implemented in a set-top box that does not include an audio playback device (and therefore no speaker), but which is connected to one or more external audio playback devices (satellite speakers, TV speakers, etc.). In this case, the configuration module sets up said device(s) by transmitting appropriate parameters to them via the set-top box's communication means.
[0292] The set-top box operating system is not necessarily Android TV.
[0293] The genres of the input audio-video stream could be different from those described here.
[0294] The third analysis could be carried out on videos (i.e. on successive image sequences), using a suitable model.
[0295] Classification models are not necessarily pre-trained. It would be possible to use, for at least one of the models, a classification model that does not require training (algorithmic classifier).< / genre> < / genre> < / genre> < / genre> < / genre> < / genre> < / genre> < / genre>
Claims
Demands
1. Decoder box (1), arranged to distribute an input stream (F) comprising an input audio signal (A), the decoder box comprising a processing unit (7) in which are implemented: - a parameterization module (11) arranged to: • perform and / or control in real time analyses (18; 18a, 18b, 18c) on at least two distinct data sources (14) relating to the input stream (F), the data sources being chosen from metadata (14a) associated with the input stream, a current audio signal (14b) from the input audio signal, and, if the input stream also includes an input video signal, at least one target image (14c) from the input video signal; • define, from the results of these analyses, a genre (G) of the input stream (F), the genre being associated with audio parameters;- a configuration module (10), arranged to dynamically adapt, using audio parameters, a setting of at least one audio playback device integrated into or connected to the decoder box (1) and comprising at least one loudspeaker (4), so as to optimize a sound rendering of said audio playback device according to the type (G) of the input stream (F).;
2. Decoder box according to claim 1, wherein the analysis of each data source (14) results in an estimation of the gender (G), and wherein the parameterization module (11) is arranged to implement a decision algorithm to define the gender of the input stream from the gender estimates.
3. Decoder box according to claim 2, the parameterization module (11) being arranged to drive a first analysis (18a), on the metadata (14a), which includes the steps of: - grouping texts from the metadata to produce an aggregated text; - performing a single inference, for each program broadcast via the input stream, from a first classification model (30), by applying the aggregated input text of said first classification model, to produce a first estimate of the genre (RI).
4. Decoder box according to claim 3, wherein the first classification model (30) uses a transformer.
5. Decoder box according to claim 3 or 4, wherein the execution of the single inference is carried out on a remote server (16).
6. Decoder box according to any one of claims 2 to 5, wherein the parameterization module (11) is arranged to perform a second analysis (18b) on the current audio signal (14b), which includes performing at least one inference of a second classification model (31), by applying the current audio signal as input to said second classification model.
7. Decoder box according to claim 6, in which the parameterization module (11) is arranged, to perform the second analysis (18b) on the current audio signal (14b), to execute inferences of the second classification model (31) repeated regularly.
8. Decoder box according to claim 6 or 7, wherein the second classification model (31) is a convolutional neural network of the YAMNet or VGGish type.
9. Decoder box according to any one of claims 6 to 8, the parameterization module (11) being arranged to perform the second analysis (18b) and to perform and / or control at least one other analysis on at least one other data source, the parameterization module being arranged so that, if the second analysis (18b) results in a gender estimate which remains constant for a first predefined duration, it confers to the gender of the input stream, at the end of the first predefined duration, the value of said gender estimate regardless of the result of the at least one other analysis.
10. Decoder box according to claim 9, the parameterization module (11) being arranged so that, if the second analysis (18b) results in a gender estimate that remains constant for a second predefined duration less than the first predefined duration, and if the gender estimate produced by at least one other analysis is identical to the gender estimate of the second analysis during the second predefined duration, it confers to the gender of the input stream, at the end of the second predefined duration, the value of said gender estimate.
11. Decoder box according to any one of claims 2 to 10, wherein the input stream also includes an input video signal, the parameterization module (11) is arranged to perform a third analysis (18c), on at least one target image (14c), which includes performing at least one inference of a third classification model (32), by applying the at least one target image as input to said third classification model (32).
12. Decoder box according to claim 11, in which the parameterization module (11) is arranged, to perform the third analysis (18c) on at least one target image (14c), to execute regularly repeated inferences of the third classification model (32).
13. Decoder box according to claim 11 or 12, wherein the third classification model (32) is a convolutional neural network of the MobileNet or CLIP type.
14. Decoder box according to any one of the preceding claims, the processing unit (7) further implementing a control module (12) arranged to define at least one control parameter (Pc) intended to optimize resource use of the parameterization module (11) and thus of the decoder box, the parameterization module being arranged to acquire the control parameter and to adapt the performance of at least one analysis as a function of at least one control parameter.
15. Decoder box according to claim 14, wherein at least one analysis comprises the execution of inferences of at least one previously trained classification model (30, 31, 32), and wherein at least one control parameter (Pc) comprises a frequency of the execution of the inferences of said model.
16. Decoder box according to claim 14 or 15, wherein at least one control parameter (Pc) includes a utilization rate of a processor of the processing unit (7).
17. A parameterization method, implemented in the parameterization module (11) of the processing unit (7) of the decoder box (1) according to any one of the preceding claims, and comprising the steps of: • performing and / or controlling real-time analyses (18; 18a, 18b, 18c) on at least two distinct data sources (14) relating to the input stream (F), the data sources being t chosen from metadata (14a) associated with the input stream, a current audio signal (14b) from the input audio signal, and, if the input stream also includes an input video signal, at least one target image (14c) from the input video signal; • define, from the results of these analyses, a genre (G) of the input stream (F), the genre being associated with audio parameters.
18. Computer program comprising instructions that cause the parameterization module (11) of the processing unit (7) of the decoder box (1) according to any one of claims 1 to 16 to perform the steps of the parameterization process according to claim 16.
19. Computer-readable recording medium on which the computer program according to claim 18 is recorded.
Citation Information
Patent Citations
Display apparatus and control method thereof
EP2916557A1
System and method for adaptive automated preset audio equalizer settings
US20220197588A1
Systems, devices and methods for distributed hierarchical video analysis
US20220222469A1
Machine-control of a device based on machine-detected transitions
US20240169960A1