Radio reception method and device comprising reinforcement learning based on user feedback

The radio reception method employs reinforcement learning to adapt processing operations based on user feedback, addressing suboptimal user experiences in recurring scenarios and improving audio quality over time.

WO2025131434A1PCT designated stage expired Publication Date: 2025-06-26CONTINENTAL AUTOMOTIVE TECHNOLOGIES GMBH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/082057
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-11-12
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing radio reception devices struggle to adapt optimally to recurring scenarios due to limitations in processing parameter settings, which can lead to suboptimal user experiences.

Method used

A radio reception method utilizing reinforcement learning based on user feedback, where a selection model and a reward allocation model are evolved over usage phases to improve audio content extraction and broadcasting.

Benefits of technology

This approach enhances user experience by adapting processing operations to recurring scenarios without increasing memory capacity, and allows for progressive improvement based on user validation of audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024082057_26062025_PF_FP_ABST
    Figure EP2024082057_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a radio reception method (40), implemented by a radio reception device (20) comprising a selection model (223) configured to select an action for controlling the processing operations performed in order to extract audio content from radio signals received by the radio reception device, and a reward allocation model (224), wherein the method comprises, during at least one phase of use of the radio reception device: - at the beginning of the at least one phase of use: modifying (S40) the reward allocation model; - during the duration of the at least one phase of use: performing reinforcement learning (S41) on the selection model by using the modified reward allocation model, and broadcasting (S42) extracted audio content; - after the end of the at least one phase of use: transmitting (S43), to a user of the radio reception device, a request to evaluate the quality of the broadcast audio content; - when the quality of the broadcast audio content is validated by the user: storing (S44) the modified selection model and the reward allocation model for a subsequent phase of use.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for radio reception comprising reinforcement learning based on user feedback Technical field

[0001] The present invention belongs to the field of radiophony, and more particularly relates to a method and a device for radio reception comprising reinforcement learning based on user feedback. State of the art

[0002] A radio reception device is a piece of equipment, for example one installed in a motor vehicle, which allows a user to access audio content from different radio stations.

[0003] A "radio station" here means a source generating audio content. Audio content generated by radio stations is broadcast to users via a network of geographically distributed transmitting antennas that emit this audio content in the form of radio signals.

[0004] Nowadays, many radio signal processing algorithms are implemented by a radio receiving device, in order to improve the listening experience of users of said radio receiving device. Generally, different processing operations can be applied to the received radio signals, in order to extract audio content to be broadcast to the user, and the control of the applied processing operations is done according to a measurement of a radio state of the radio receiving device and predetermined processing parameters of said radio receiving device.

[0005] The values ​​of these processing parameters are generally set during the development of the radio reception device, and are then not changed once the production of the radio reception device is launched. During development, the setting of the values ​​of the processing parameters is generally carried out on the basis of reference scenarios (a reference scenario corresponding to a predefined radio state or a predefined sequence of radio states), and said values ​​of the processing parameters are adjusted to ensure that the radio reception device applies predefined reference processes associated with these reference scenarios.

[0006] Generally speaking, it is not possible to foresee all possible scenarios to which the radio reception device may be exposed during its use, so that such an approach does not guarantee that the values ​​of the processing parameters thus set will be optimal, from a user experience point of view, when using said radio reception device, after it has been produced.

[0007] Furthermore, a radio reception device, for example installed in a motor vehicle, will most often be exposed to the same scenarios repeating themselves recurrently, given that it is not uncommon for a motor vehicle to regularly perform the same journeys, and that the user of this motor vehicle also makes the same choices in terms of the radio station(s) listened to. If these recurring scenarios are poorly covered by the reference scenarios, then the user experience will most often be suboptimal.

[0008] International application PCT WO 2012 / 059782 A1 proposes to provide several sets of processing parameter values, and to select a set of values ​​to be used based on a measurement of a radiophonic state of the radiophonic reception device. International application PCT WO 2012 / 059782 A1 also proposes to request real-time feedback from the user on the listening quality obtained by using the selected set of values.

[0009] PCT international application WO 2012 / 059782 A1 therefore allows for improved coverage of a larger number of different scenarios. However, such an approach also does not allow for correct adaptation to all possible scenarios. Increasing the number of different scenarios that can be supported by the radio reception device requires increasing the number of different sets of processing parameter values, which prohibitively increases the memory capacity requirements. Furthermore, when the user is the driver of the motor vehicle, it is not possible for the user, from a road safety point of view, to provide real-time feedback on each set of processing parameter values ​​applied. Statement of the invention

[0010] The present invention aims to remedy all or part of the limitations of the solutions of the prior art, in particular those set out above, by proposing a solution which makes it possible to develop the processing applied by a radio reception device in a manner which improves the user experience, in particular for scenarios of use of the radio reception device which occur repeatedly.

[0011] To this end, it proposes according to a first aspect a radio reception method, implemented by a radio reception device comprising a selection model configured to select, as a function of a radio state of the radio reception device, a control action of processing carried out to extract audio content from radio signals received by the radio reception device, the selection model being a machine learning model, the radio reception device further comprising a reward allocation model as a function of a radio state and a selected control action, said method comprising, during at least one phase of use of the radio reception device: - at the start of the at least one phase of use: a modification of the reward allocation model, so as to obtain a modified reward allocation model, the modified reward allocation model being obtained by a random modification of parameter values ​​of the reward allocation model, the random modification of a value of a parameter of the reward allocation model following a law of zero-mean probabilities, - during the duration of the at least one phase of use: reinforcement learning of the selection model using the modified reward allocation model, so as to modify the selection model, and sound broadcasting of audio content extracted from radio signals processed according to control actions selected by the modified selection model, - after the end of at least one phase of use: a transmission, to a user of the radio reception device, of a request for evaluation of the quality of the audio content broadcast, - when the quality of the broadcast audio content is validated by the user: a storage of the modified selection model as a selection model and of the modified reward allocation model as a reward allocation model, for a subsequent phase of use.

[0012] Thus, it is proposed to develop the processing applied by the radio reception device by automatic learning ("machine learning" in Anglo-Saxon literature), and more particularly by reinforcement learning based on user feedback ("reinforcement learning via human feedback").

[0013] Conventionally, reinforcement learning consists of an autonomous agent (here the radio reception device) learning the best action or sequence of actions to perform through experimentation. More precisely, the autonomous agent observes an environment (here the radio state in which the radio reception device operates), and decides on a control action or sequence of control actions to perform based on its current radio state. A control action here corresponds to an action controlling the processing performed on the radio signals, for example an action modifying the values ​​of processing parameters.In response to a control action, the autonomous agent receives a reward that can be positive if the control action had a positive impact on the extraction of the audio content, or negative if the control action had a negative impact on the extraction of the audio content, or zero. These steps of selecting a control action, transitioning to a new radio state and receiving a reward are iterated as long as a stopping criterion is not verified, and allow the autonomous agent to learn or determine a behavior, also referred to as a "strategy" or "policy", which is optimal in that it tends to maximize not the reward perceived in the short term but the sum (possibly weighted) of the cumulative rewards in the medium or long term.

[0014] Such reinforcement learning is therefore based on a selection model, which implements the selection policy to select a control action based in particular on the radiophonic state of the radiophonic reception device, and a model of attribution of reward, which assigns a reward to a selected control action. The selection model is a machine learning model such as an (artificial) neural network.

[0015] It is proposed to evolve the selection model (and therefore to evolve the processing applied by the radio reception device) during at least one phase of use of the radio reception device. A phase of use corresponds to a phase beginning with the activation of said radio reception device (for example when the engine of the motor vehicle is started) and ending with the deactivation of said radio reception device (for example when the engine of the motor vehicle is stopped). At the start of the at least one phase of use, the reward allocation model is modified, in order to modify the reward function implemented, then a reinforcement learning algorithm is implemented to evolve the selection model over the duration of the at least one phase of use according to the modified reward allocation model.After the end of the at least one usage phase (i.e., in delayed time), the user is asked to evaluate the quality of the streamed audio content. If the user indicates that he is satisfied with the quality of the streamed audio content, then the changes made to the reward allocation model (at the beginning of the at least one usage phase) and to the selection model (by reinforcement learning) are considered to improve the user experience, and are retained for the next usage phase.

[0016] In this way, the user experience can be improved over successive phases of use, in particular for scenarios of use of the radio reception device that occur repeatedly (such as those encountered during daily journeys in a motor vehicle), on the basis of user feedback that is not in real time (only after the end of a phase of use) and is therefore not a source of inattention for the user / driver during the journey of the motor vehicle. Such an evolution of the applied processing operations is also done without having to increase the required memory capacity, since it is not necessary to increase the number of different sets of processing parameter values ​​that must be stored.

[0017] In particular embodiments, the radio reception method may further comprise one or more of the following optional features, taken individually or in any technically possible combination.

[0018] In particular implementations, when the quality of the streamed audio content is not validated by the user, the modified reward allocation model is not stored as the reward allocation model for the next usage phase. In other words, the modifications made to the reward allocation model are discarded, and the initial reward allocation model, before its modification at the start of the at least one usage phase, will be reused (and possibly re-modified) during the next usage phase.

[0019] In particular embodiments, when the quality of the broadcast audio content is not validated by the user, the modified selection model is not stored as a selection model for the next usage phase. In other words, the modifications made to the selection model are discarded, and the initial selection model, before its modification by reinforcement learning during the at least one usage phase, will be reused (and possibly re-modified) during the next usage phase.

[0020] In particular implementations, the variance of the probability distribution used to modify the value of a parameter of the reward allocation model is determined from the starting value of the parameter before modification.

[0021] In particular embodiments, the variance of the probability distribution used to modify the value of a parameter of the reward allocation model is equal to s ■ x0, expression in which x0 is the starting value of said parameter of the reward allocation model and O < e < l or O < £ < 0.1.

[0022] In particular implementations, the selection model is an (artificial) neural network, for example a deep neural network. For example, the selection model is a feedforward neural network (FNN).

[0023] In particular implementation modes, the selection model is previously trained, for example by supervised learning.

[0024] In particular embodiments, the reward assignment model is a machine learning model, for example, a neural network such as a deep neural network. For example, the reward assignment model is a feedforward neural network (FNN).

[0025] In particular implementation modes, the reward allocation model is previously trained by supervised learning.

[0026] In particular embodiments, the control action selected by the selection module controls at least one of: a filter to be applied to a radio signal, a filter to be applied to audio content extracted from a radio signal, a decision to switch between a first reception frequency of a radio station and a second reception frequency of the same radio station, - a decision to switch between an analog radio broadcasting system and a digital radio broadcasting system.

[0027] In particular modes of implementation, the radio reception device being on board a motor vehicle, a phase of use corresponds to an activation of the radio reception device on a journey of the motor vehicle.

[0028] In particular modes of implementation, the probability law used to modify the value of a parameter of the reward allocation model is a Normal distribution.

[0029] According to a second aspect, there is provided a computer program product comprising instructions which, when executed by a radio reception device comprising at least one tuner, a processing circuit, a sound broadcasting module and a human-machine interface, configure said radio reception device to implement a radio reception method according to any one of the embodiments of the present disclosure.

[0030] According to a third aspect, there is provided a computer-readable recording medium on which is recorded a set of instructions which, when executed by a radio reception device comprising at least one tuner, a processing circuit, a sound broadcasting module and a human-machine interface, configure said radio reception device to implement a radio reception method according to any of the embodiments of the present disclosure.

[0031] According to a fourth aspect, a radio reception device is proposed comprising at least one tuner, a processing circuit, a sound broadcasting module and a human-machine interface, configured to implement a radio reception method according to any one of the embodiments of the present disclosure.

[0032] According to a fifth aspect, there is provided a motor vehicle comprising a radio reception device according to any of the embodiments of the present disclosure. Presentation of figures

[0033] The invention will be better understood by reading the following description, given as a non-limiting example, and made with reference to the figures which represent: Figure 1: A schematic representation of an example of the implementation of a radio broadcasting system, Figure 2: a schematic representation of an exemplary embodiment of a radio reception device, Figure 3: A schematic representation of a reinforcement learning system based on user feedback from the radio reception device, Figure 4: A diagram illustrating the main steps of an example implementation of a radio reception method.

[0034] In these figures, like references from one figure to another designate identical or similar elements. For reasons of clarity, the elements shown are not to scale, unless otherwise stated.

[0035] Furthermore, the order of steps shown in these figures is given solely as a non-limiting example of the present disclosure which may be applied with the same steps performed in a different order and / or with steps performed in parallel and / or jointly. Description of the embodiments

[0036] Figure 1 schematically represents an example of a radio broadcasting system comprising a transmitting antenna 11, which transmits radio signals in a coverage area 12.

[0037] It should be noted that the present disclosure relates to radio reception in general, and is applicable to all types of radio broadcasting systems, in particular to all types of analog radio broadcasting systems ("Amplitude Modulation", AM, "Frequency Modulation, FM, etc.) and / or digital radio broadcasting systems ("Digital Audio Broadcasting", DAB, etc.), including in the case of a switch from an analog radio broadcasting system to a digital radio broadcasting system, and vice versa, for example to perform radio station tracking.

[0038] Figure 1 also shows a motor vehicle 13 in which a radio reception device 20 is mounted. The radio reception device 20 broadcasts audio content from a radio station selected by a user of the motor vehicle 13. This audio content is extracted from radio signals received from the transmission antenna 11.

[0039] Figure 2 schematically represents an exemplary embodiment of a radio reception device 20.

[0040] As illustrated in Figure 2, the radio reception device 20 comprises at least one tuner 21. In the example illustrated in Figure 2, the radio reception device 20 comprises a single tuner. In other examples not illustrated by figures, the radio reception device 20 may comprise two tuners, or even more.

[0041] In a known manner, a tuner 21 makes it possible to bring a radio signal, received in a channel associated with a frequency, back into baseband, the frequency being for example selected by a user. For example, a tuner 21 provides as output, in analog or digital form, a baseband signal which corresponds to the selected radio signal brought back to a zero or intermediate frequency. Such a tuner 21 comprises equipment (amplifier(s), local oscillator(s), mixer(s), analog and / or digital filter(s), analog / digital converter(s), etc.) considered to be known to the person skilled in the art. Such a tuner 21 may also comprise and / or be connected to other equipment (antenna(s), etc.) considered to be known to the person skilled in the art.

[0042] As illustrated in Figure 2, the radio reception device 20 also comprises a processing circuit 22. The processing circuit 22 comprises, for example, one or more processors 220 (CPU, DSP, GPU, FPGA, ASIC, etc.). In the case of several processors 220, these may be integrated into the same equipment and / or integrated into materially distinct equipment in communication with each other. The processing circuit 22 also comprises one or more memories 221 (magnetic hard disk, electronic memory, optical disk, etc.) in which is for example stored a computer program product 222, in the form of a set of program code instructions to be executed by the processor(s) to implement all or part of the steps of a radio reception method 40 which will be described below.

[0043] As illustrated in FIG. 2, the radio reception device 20 also comprises a sound broadcasting module 23, making it possible to deliver audio content to users. The sound broadcasting module 23 may, for example, comprise one or more loudspeakers and / or one or more connection modules for headphones. A connection module may, for example, allow a wired connection (jack plug, USB port, etc.) or a wireless connection (Bluetooth communication module, WiFi, etc.) with headphones.

[0044] As illustrated in FIG. 2, the radio reception device 20 also includes a human-machine interface 24 (HMI), allowing a user of the radio reception device 20 to interact therewith. The HMI 24 may take any suitable form allowing the radio reception device 20 to transmit information to a user and allowing a user to transmit information to the radio reception device 20. For example, the HMI 24 includes a touchscreen and, optionally, a microphone. It should also be noted that, in certain exemplary embodiments, the HMI 24 may use the sound broadcasting module 23 to transmit information to the user.

[0045] In certain exemplary embodiments, and as illustrated in the non-limiting example of FIG. 2, the radio reception device 20 may optionally comprise a location module 25, making it possible to determine a geographical position of the radio reception device 20. Any type of location module 25 may be considered within the scope of the present disclosure and the choice of a particular type of location module 25 constitutes only a variant implementation of the present disclosure. In preferred embodiments, the location module 25 corresponds to a reception module of a global navigation satellite system (GNSS), for example a GPS reception module (global positioning system).

[0046] As indicated previously, the present disclosure is based on a reinforcement learning system. Figure 3 schematically represents the main components implemented in such a reinforcement learning system. As illustrated by Figure 3, the radio reception device 20 comprises a selection model 223, which implements a selection policy for selecting a control action based in particular on the radio state of the radio reception device 20, and a reward allocation model 224, which allocates a reward to a selected control action. As illustrated by Figure 3, the selection model 223 and the reward allocation model 224 reward are for example implemented by the processing circuit 22. The radiophonic state of the radiophonic reception device 20 is for example measured and / or determined from measurements carried out by means of the tuner 21 (and possibly by means of the location module 25), and the control action selected by the selection model 223 is for example executed by said tuner 21 and / or by the processing circuit 22.

[0047] In a non-limiting manner, O denotes a set of radio states of the radio reception device 20, A a set of control actions that can be implemented to control the processing of radio signals by the radio reception device 20, and Y a set of rewards that can be attributed by the reward allocation module 224. The selection model 223 receives at each instant t an observation o te O of the radiophonic state in which the radiophonic reception device 20 is located (for example provided by the tuner 21), and determines a control action a t e A to be carried out on the basis of a selection policy n: OA, that is to say a t = n( t). Following the performance of this control action a t , the radio reception device 20 enters a radio state o t+1 and the reward allocation model 224 generates a reward r t e Y based on a reward function f : O x 4 -> Y, i.e. r t = r(o t , has t). The selection model 223 is a machine learning model, which can therefore be adjusted by reinforcement learning to select a control action that tends to optimize a predicted (possibly weighted) sum of rewards that can be accumulated over time following the implementation of this control action.

[0048] Generally, any reinforcement learning algorithm can be implemented to optimize the selection policy implemented by the selection module 223. According to a non-limiting example, the reinforcement learning algorithm is a “Q-learning” type algorithm (see for example the document “Reinforcement Learning, An Introduction, Second Edition”, Richard S. Sutton and Andrew G. Barto, The MIT Press, Cambridge, Massachusetts, 2018), preferably a deep Q-learning algorithm. In such a case, the selection model 223 includes a deep neural network (usually referred to as a deep Q-network in the English literature) to model a Q function that predicts a future gain (or Q-value) resulting from a control action a t selected at a time t. For example, for a selection policy that must select a sequence of control actions a , a2, - -- , a T ~), the function Q is for example defined according to the following expression: [Math. 1] expression in which corresponds to the expectation with respect to the decision policy n.

[0049] Following another example, the function Q is defined according to the following expression: [Math. 2] expression in which 0 < y < 1 is a weighting factor used to give more weight to rewards obtained in the long term than to rewards obtained in the short term.

[0050] The reinforcement learning algorithm aims to automatically learn the Q function, so that the optimal n selection policy can be constructed by selecting the control action which, from a given radio state, maximizes the Q function. Any reinforcement learning method can be implemented to perform the automatic learning of the Q function, and the choice of a particular reinforcement learning method is only one implementation variant of the present disclosure.

[0051] As indicated above, the selection model 223 determines a control action as a function of a radio state of the radio reception device 20. The radio state is for example defined by one or more state parameters such as: - a temporal or frequency representation of a radio signal received by the radio reception device 20, or parameters determined from an analysis of such a temporal or frequency representation (for example an evaluation of variability of the amplitude of the radio signal, in particular in the case of a modulation assumed to be at constant amplitude, or even an evaluation of a width of a frequency band occupied by the radio signal, etc.), - an estimate of a quantity representative of a reception quality of a radio signal (for example an estimate of a reception power of the radio signal, of a frequency drift affecting the radio signal, of a signal-to-noise ratio of the radio signal, of a propagation channel between the transmission antenna 11 and the radio reception device 20, etc.), - one or more reception frequencies on which audio content from the same given radio station is received in the geographical area in which the radio reception device 20 is located, - one or more identifiers of radio stations whose audio content is received in the geographical area in which the radio reception device 20 is located, and optionally the associated reception frequencies, - a parameter determined by analysis (for example temporal or frequency) of audio content extracted from a radio signal (for example an estimation of the type of audio content - for example among speech, music, etc. -, or an estimation of the type of music corresponding to the audio content - for example among classical music, rock, pop, etc. -, etc.), - a geographical position of the radio reception device 20, for example provided by the location module 25, - etc.

[0052] The control action makes it possible to control all or part of the processing applied to the radio signals by the radio reception device 20. For example, the control action can modify the values ​​of processing parameters applied by the radio reception device 20. For example, a control action selected by the selection module 223 can for example provide one or more of the following processing parameters: - the coefficients of a filter to be applied to a radio signal (for example by calculating the filter coefficients, or by selecting a filter from among several predefined filters, etc.), - the coefficients of a filter to be applied to an audio content extracted from a radio signal (for example by calculating the coefficients of the filter, or by selecting a filter from several predefined filters, etc.), a decision to switch from a first reception frequency to a second reception frequency in the context of monitoring a radio station whose audio content is received, in the geographical area in which the radio reception device 20 is located, both on the first reception frequency and the second reception frequency (such a decision corresponds for example to an indication of the reception frequency to be considered, or to a determination of threshold value(s) to be used to make the decision whether or not to switch from one reception frequency to another, etc.), a decision to switch between an analogue radio broadcasting system and a digital radio broadcasting system (such a decision corresponds, for example, to an indication of the radio broadcasting system to be considered, or to a determination of threshold value(s) to be used to make the decision whether or not to switch between an analogue radio broadcasting system and a digital radio broadcasting system, etc.),. - etc.

[0053] For example, a filter to be applied to a radio signal (e.g., brought back to an intermediate or zero frequency), which can be determined by the control action, may correspond to a dynamic propagation channel equalization filter (e.g., to obtain a substantially constant amplitude, in the case of a modulation assumed to be of constant amplitude, or to obtain a substantially constant phase, in the case of a modulation assumed to be of constant amplitude, or more generally to reduce the negative effects of multiple paths of the propagation channel, etc.). In such a case, the control action may, for example, adapt the coefficients of the dynamic propagation channel equalization filter propagation as a function of the observed radiophonic state, for example as a function of an evaluation of amplitude or phase variability of the radiophonic signal, and / or as a function of an estimation of the propagation channel, etc.

[0054] Another non-limiting example of a filter to be applied to the radio signal, which can be determined by the control action, corresponds for example to a band-pass filter, the control action being able for example to adapt the bandwidth and / or cut-off frequencies of said band-pass filter as a function of the observed radio state, for example as a function of a frequency representation of the radio signal or of parameters determined from such a frequency representation (such as for example an evaluation of the width of the frequency band occupied by the radio signal), etc.

[0055] For example, a filter to be applied to audio content extracted from a radio signal, which can be determined by the control action, may correspond to an audio equalization filter (e.g. to reinforce or attenuate the low frequencies or the high frequencies of the audio spectrum, etc.). In such a case, the control action may for example adapt the coefficients of the dynamic audio equalization filter according to the observed radio state, for example according to an identifier of the radio station whose audio content is broadcast to the user, and / or according to a parameter determined by time and / or frequency analysis of the extracted audio content (e.g. the type of music), etc.

[0056] For example, the decision to switch between a first reception frequency and a second reception frequency of the same radio station may be taken by the selection module 223 as a function of state parameters such as the estimated reception qualities (for example the estimated reception power, the estimated signal-to-noise ratio, etc.) on the first reception frequency and the second reception frequency respectively, and / or the geographical position of the radio reception device 20, etc.

[0057] For example, the decision to switch between an analog radio broadcasting system and a digital radio broadcasting system may be made by the selection module 223 based on state parameters such as the estimated reception qualities (e.g., the estimated reception power, the estimated signal-to-noise ratio, etc.) for the analog radio broadcasting system and the digital radio broadcasting system, respectively, and / or the geographical position of the radio reception device 20, etc.

[0058] It should be noted that the reinforcement learning algorithm is implemented, to evolve the selection model 223, during the use of the radio reception device 20, that is to say after the latter has been produced and installed in a motor vehicle purchased by a user.

[0059] Several strategies are possible to initialize the selection model 223, i.e. to define the selection model 223 (for example a deep neural network such as a FNN) which will be implemented during the very first use of the radio reception device 20 by a user. For example, the selection model 223 may be initialized randomly. However, such random initialization could result in a degraded user experience during the first uses of the radio reception device 20, even if the user experience will be gradually improved thanks to the reinforcement learning algorithm.

[0060] In preferred embodiments, the selection model 223 is previously trained, during the development of the radio reception device 20, by supervised learning. Conventionally, such supervised learning is based on a training data set comprising reference data (reference scenarios) to be provided as input to the selection model 223, as well as the outputs (control actions) expected in response to these reference data, and the supervised learning aims to obtain a selection model 223 which tends to select the expected control actions for the reference scenarios considered. Such a training data set is for example determined by experts and / or by calibration or simulation.The training data set may also be supplemented over time with information received from radio reception devices 20 in use, after their selection models 223 have been optimized by reinforcement learning according to the present disclosure. Such supervised learning therefore makes it possible to initialize the selection model 223 in such a way that the initial user experience provided by the radio reception device 20 is comparable to that provided by the prior art solutions. This user experience will, however, in the present case, be progressively improved thanks to the reinforcement learning algorithm.

[0061] As indicated above, the reinforcement learning algorithm further relies on user feedback. For this purpose, the reward allocation model 224 may also be required to evolve during use of the radio reception device 20, in order to attempt to progressively obtain an allocated reward (or gain / Q-value) that is more representative of the listening quality as perceived by the user of the radio reception device 20. The reward allocation model 224 must therefore be able to be modified by the radio reception device 20. The reward function implemented by the reward allocation model 224 is for example defined by different parameters, at least some of which have values ​​that can be modified.In preferred embodiments, the reward allocation model 224 is a machine learning model, such as a neural network, preferably a deep neural network (e.g., a FNN). Where applicable, everything said above for the initialization of the selection model 223 is also applicable for the initialization of the reward allocation model 224. In particular, in preferred embodiments, the. reward allocation model 224 is previously trained, during the development of the radio reception device 20, by supervised learning. Conventionally, such supervised learning is based on a training data set comprising reference data (reference scenarios) to be provided as input to the reward allocation model 224, as well as the outputs (rewards) expected in response to these reference data, and the supervised learning aims to obtain a reward allocation model 224 which tends to allocate the expected rewards for the reference scenarios considered. Such a training data set is for example determined by experts and / or by calibration or simulation.The training data set may also be supplemented over time with information received from radio reception devices 20 in use, after their reward allocation models 224 have been optimized by reinforcement learning based on user feedback according to the present disclosure.

[0062] Reinforcement learning is carried out after the production of the radio reception device 20, the latter being embedded in a motor vehicle, during phases of use of the radio reception device 20.

[0063] A phase of use of the radio reception device 20 corresponds to a phase beginning with the activation of said radio reception device 20 (such activation being able to be simultaneous with the starting of the engine of the motor vehicle 13) and ending with the deactivation of said radio reception device 20 (such deactivation being able to be simultaneous with the stopping of the engine of the motor vehicle 13).

[0064] To be able to perform user feedback reinforcement learning, the reward allocation model 224 is modified at the beginning of a usage phase, so as to obtain a modified reward allocation model 224, and this modified reward allocation model 224 is used for reinforcement learning of the selection model 223 over the duration of this usage phase (during which the reward allocation model 224 was modified), and possibly also over the duration of one or more subsequent usage phases. In general, reinforcement learning can be performed (before requiring user feedback as described below) over a single usage phase.However, if this is considered too short (for example, of a duration less than a first predetermined threshold value), then it is possible to continue the reinforcement learning over one or more subsequent phases of use, for example until a cumulative duration considered sufficient is obtained (for example, greater than a second predetermined threshold value). In certain examples, it is also possible to carry out the reinforcement learning over a predetermined number of phases of use. In summary, depending on the case, the reinforcement learning, using the modified reward allocation model 224, can be carried out (before requiring user feedback as described below) on. the duration of one or more phases of use.

[0065] Figure 4 schematically represents the main steps of an example of implementation of a radio reception method 40.

[0066] As illustrated by FIG. 4, and as indicated previously, the radio reception method 40 comprises, at the start of a use phase, a step S40 of modifying the reward allocation model 224, so as to obtain a modified reward allocation model.

[0067] Such a modification of the reward allocation model 224 is for example carried out by modifying the parameter values ​​of the reward allocation model 224, for example randomly according to a zero-mean probability distribution. For example, the variance of the probability distribution used is determined from the starting value (at the start of step S40) of the modified parameter of the reward allocation model 224. For example, for a parameter X whose starting value is x0, the random modification introduced follows for example a zero-mean probability distribution and variance s ■ x0, expression in which 0 < s < 1 (and preferably 0 < £ < 0.1, for example £ = 0.05).

[0068] Generally, different probability laws may be considered for modifying the parameter values ​​of the reward allocation model 224. For example, the random modification of a value of a parameter of the reward allocation model 224 follows a uniform distribution with zero mean. In preferred embodiments, the random modification of a value of a parameter of the reward allocation model 224 follows a Normal distribution. For example, the Normal distribution used has a zero mean, and its variance is for example determined from the starting value (at the beginning of step S40) of the modified parameter of the reward allocation model 224. For example, for a parameter X whose starting value is x0, the random modification introduced follows for example a Normal law J\r(O, £ ■ x0), expression in which 0 < £ < 1 (and preferably 0 < £ < 0.1, for example £ = 0.05), so that X~N(x0,£ ■ x0).

[0069] As illustrated in FIG. 4, the radio reception method 40 comprises, during the duration of one or more phases of use of the radio reception device 20, a step S41 of reinforcement learning of the selection model 223, which uses the modified reward allocation model 224 obtained during the step S40. As indicated above, the reinforcement learning step S41 can implement any type of suitable reinforcement learning algorithm, and the choice of a particular algorithm constitutes only a variant implementation of the present disclosure.By this reinforcement learning, the selection model 223 is iteratively modified, and the selection model 223, thus iteratively modified, is used to select the control actions of the processing applied to continuously extract audio content from the radio signals received by the radio reception device 20. The audio content thus. extract is broadcast to the user of the motor vehicle 13, by means of the sound broadcasting module 23, during a step S42.

[0070] As illustrated in FIG. 4, after the end of one or more phases of use implementing reinforcement learning using the modified reward allocation model 224, the radio reception method 40 comprises a step S43 of transmitting, by means of the HMI 24 and to the user of the radio reception device 20, a request for evaluating the quality of the audio content broadcast during the duration of the phase of use which has just ended. It should be noted that the execution of the transmission step S43 may also, optionally, be subject to other conditions in addition to that relating to the completion of a phase of use (i.e. the fact that the radio reception device 20 is deactivated, i.e. no longer broadcasts audio content).In particular, in certain examples, the transmission step S43 can be executed only if the motor vehicle 13 is stopped (stationary), or even only if the engine of said motor vehicle 13 is stopped (if the radio reception device 20 is deactivated before the engine of said motor vehicle 13 is stopped). Such provisions make it possible to ensure that the user's return is requested when the motor vehicle 13 is stopped, that is to say in good conditions from a road safety point of view.

[0071] As illustrated in FIG. 4, when the quality of the broadcast audio content is validated by the user (reference S43a in FIG. 4), the radio reception method 40 comprises a step S44 of storing the modified selection model 223 as selection model 223 and the modified reward allocation model 224 as reward allocation model 224, for a subsequent use phase. In other words, a positive feedback from the user amounts to validating the modifications made to the selection model 223 and the reward allocation model 224, and these modifications are consequently retained for subsequent use phases.

[0072] When the quality of the broadcast audio content is not validated by the user (reference S43b in FIG. 4), the modified reward allocation model 224 is not stored as the reward allocation model 224 for the next use phase, and the radio reception method 40 comprises a step S45 during which the modifications made to the reward allocation model 224 are discarded. In this way, it is the reward allocation model 224 in force before the modification step S40 which is used for the next use phase.

[0073] In some cases, it is also possible, when the quality of the broadcast audio content is not validated by the user, not to memorize the modified selection model 223 as the selection model 223 for the next usage phase. In this case, the modifications made to the selection model 223 (by reinforcement learning using the modified reward allocation model 224) are also discarded during of step S45. Where appropriate, it is the selection model 223 in force before the modification step S40 which is used for the following use phase. However, it is also possible, in certain examples not illustrated by figures, to retain the modifications made to the selection model 223 even when the quality of the broadcast audio content is not validated by the user, because the reinforcement learning of said selection model 223 will continue during the following use phases.

[0074] It should be noted that the quality of the audio content is considered not to be validated if the radio reception device 20 receives, via the HMI 24, negative feedback from the user (i.e., if the user indicates that he is not satisfied with the quality of the broadcast audio content). However, the quality of the audio content may also be considered not to be validated if the radio reception device 20 does not receive feedback from the user. For example, the quality of the audio content may be considered not to be validated if the user does not respond within a predetermined response time.

[0075] More generally, it should be noted that the methods of implementation and embodiment considered above have been described as non-limiting examples, and that other variants are consequently conceivable.

[0076] In particular, the invention has been described by considering mainly the case of a radio reception device 20 on board a motor vehicle 13. It should however be noted that it is also possible, according to other examples, to have a radio reception device 20 which is not on board a motor vehicle 13, but which is for example carried by a user.

[0077] Furthermore, the invention has been described by considering mainly a selection module 223 and a reward allocation module 224. It should however be noted that it is also possible, according to other examples, to consider several selection module 223 / reward allocation module 224 pairs to control different treatments. For example, it is possible to provide: - a first selection module 223, associated with a first reward allocation module 224, for selecting a filter to be applied to a radio signal or to audio content extracted from this radio signal, - a second selection module 223, associated with a second reward allocation module 224, to decide whether to switch between different reception frequencies or between different radio broadcasting systems, etc.

Claims

Claims 1. Method (40) for radio reception, implemented by a radio reception device (20) comprising a selection model (223) configured to select, as a function of a radio state of the radio reception device, a control action of processing carried out to extract audio content from radio signals received by the radio reception device, the selection model being a machine learning model, the radio reception device (20) further comprising a reward allocation model as a function of a radio state and a selected control action, said method comprising, during at least one phase of use of the radio reception device: - at the start of the at least one phase of use: a modification (S40) of the reward allocation model, so as to obtain a modified reward allocation model, the modified reward allocation model being obtained by a random modification of parameter values ​​of the reward allocation model, the random modification of a value of a parameter of the reward allocation model following a zero-mean probability distribution, - during the duration of the at least one phase of use: reinforcement learning (S41) of the selection model using the modified reward allocation model, so as to modify the selection model, and sound broadcasting (S42) of audio content extracted from radio signals processed according to control actions selected by the modified selection model, - after the end of the at least one phase of use: a transmission (S43), to a user of the radio reception device, of a request for evaluation of the quality of the audio content broadcast during the duration of the at least one phase of use, - when the quality of the broadcast audio content is validated by the user: a storage (S44) of the modified selection model as a selection model and of the modified reward allocation model as a reward allocation model, for a following use phase.

2. Method (40) according to claim 1, wherein, when the quality of the broadcast audio content is not validated by the user, the modified reward allocation model is not stored as a reward allocation model for the next phase of use.

3. Method (40) according to any one of the preceding claims, wherein, when the quality of the broadcast audio content is not validated by the user, the modified selection model is not stored as a selection model for the next phase of use.

4. Method (40) according to any one of the preceding claims, in which the variance of the probability law used to modify the value of a parameter of the model reward allocation is determined from the starting value of the parameter before modification.

5. Method (40) according to claim 4, in which the variance of the probability law used to modify the value of a parameter of the reward allocation model is equal to s ■ x0, expression in which x0 is the starting value of said parameter of the reward allocation model and 0 < s < 1 or 0 < s < 0.

1.

6. Method (40) according to any one of the preceding claims, wherein the selection model is a neural network.

7. Method (40) according to any one of the preceding claims, in which the selection model is previously trained by supervised learning.

8. A method (40) according to any preceding claim, wherein the reward allocation model is a machine learning model, for example a neural network.

9. The method (40) of claim 8, wherein the reward allocation model is pre-trained by supervised learning.

10. Method (40) according to any one of the preceding claims, in which the control action selected by the selection module controls at least one of: - a filter to be applied to a radio signal, - a filter to be applied to audio content extracted from a radio signal, - a decision to switch between a first reception frequency of a radio station and a second reception frequency of the same radio station, - a decision to switch between an analog radio broadcasting system and a digital radio broadcasting system.

11. Method (40) according to any one of the preceding claims, in which, the radio reception device being on board a motor vehicle, a phase of use corresponds to an activation of the radio reception device on a journey of the motor vehicle.

12. Method (40) according to any one of the preceding claims, in which the probability law used to modify the value of a parameter of the reward allocation model is a Normal law.

13. Computer program product (222) comprising instructions which, when executed by a radio reception device (20) comprising at least one tuner (21), a processing circuit (22), a sound broadcasting module (23) and a human-machine interface (24), configure said radio reception device to implement a radio reception method (40) according to any one of the preceding claims.

14. Computer-readable recording medium on which is recorded a set of instructions which, when executed by a radio reception device (20) comprising at least one tuner (21), a processing circuit (22), a sound broadcasting module (23) and a human-machine interface (24), configure said radio reception device to implement a radio reception method (40) according to any one of claims 1 to 12.

15. Radio reception device (20) comprising at least one tuner (21), a processing circuit (22), a sound broadcasting module (23) and a human-machine interface (24), configured to implement a radio reception method (40) according to any one of claims 1 to 12.

16. Motor vehicle (13) comprising a radio reception device (20) according to claim 15.

Citation Information

Patent Citations

  • Radio receiver with adaptive tuner

    WO2012059782A1

  • Robust learning-based radio stabilization in a vehicle

    DE102021127525A1

  • Processing of communications signals using machine learning

    US10396919B1