METHOD FOR RECEIVING AND PROCESSING AUDIO DATA AND TRIGGERING ASSOCIATED SOUND EFFECTS BASED ON PROSODY AND / OR MOVEMENT

The method synchronizes sound effects with real-time audio data processing, addressing errors and limitations in existing technologies by using phoneme detection and movement recognition for adaptive and immersive reading experiences.

FR3152077B1Active Publication Date: 2025-10-24POÉTIE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2023008565
Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-08-08
Publication Date
2025-10-24
Estimated Expiration
2043-08-08

AI Technical Summary

Technical Problem

Existing audio data processing methods for triggering sound effects during reading are prone to errors, lack adaptability to speakers' emotions and reading styles, and do not allow for interactive movements, resulting in a limited and non-immersive reading experience.

Method used

A method that utilizes real-time audio data processing to synchronize sound effects with prosody and movement, employing preconfigured algorithms to detect phonemes and movement, ensuring accurate and adaptive sound effect triggering without requiring internet connectivity.

Benefits of technology

Enhances the reading experience by providing a personalized and immersive interaction through adaptive sound effects that respond to the speaker's emotions and movements, improving accuracy and reducing errors in sound effect timing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000028_0000
    Figure 00000028_0000
  • Figure 00000028_0001
    Figure 00000028_0001
  • Figure 00000028_0002
    Figure 00000028_0002
Patent Text Reader

Abstract

METHOD FOR RECEIVING AND PROCESSING AUDIO DATA AND TRIGGERING ASSOCIATED SOUND EFFECTS BASED ON PROSODY AND / OR MOVEMENT The invention relates to a method for receiving and processing audio data comprising words corresponding to the real-time reading of a source text for triggering sound effects synchronized with said reading of the text, characterized in that it comprises a step (110) of determining a position index (12) of the speaker in the source text, by detecting a correspondence between the received audio data and the source text, a step (112) of receiving at least one piece of data representative of a prosody piece of data (10), and / or of receiving at least one piece of movement data, a step (140) of determining the sound effect to be triggered as a function of the position index (12), and as a function of the prosody piece of data and / or the movement piece of data, movement,a step (150) of triggering the determined sound effect. Figure for the abstract: figure 1,
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: METHOD FOR RECEIVING AND PROCESSING AUDIO DATA AND TRIGGERING ASSOCIATED SOUND EFFECTS FUNCTION OF PROSODY AND / OR MOVEMENT Technical field of the invention

[0001] The invention relates to a method for receiving and processing audio data, and to a system for receiving and processing audio data. In particular, the invention relates to a method for triggering sound effects associated with a source text during reading of the source text by a human speaker, taking into account the prosody of the reading and / or movement during reading. Technological background

[0002] The invention is placed in the field of reading and proposes to accompany the reading of a book by triggering sound effects linked to the source text which is read.

[0003] Various interactive reading techniques have already been proposed to enable the triggering of sound effects during reading.

[0004] Some techniques, for example, offer an estimation of the playback speed to allow the triggering and pre-loading of sound effects at the time that seems appropriate. These techniques are simple but the risk of triggering a sound effect at the wrong time is high.

[0005] Other techniques rely on recognizing the reading of predetermined keywords in the text, to trigger the associated sound effect. These word detection techniques are, however, more subject to recognition and reading tracking errors.

[0006] Furthermore, current techniques generally only offer one sound effect intended for each passage of the text and do not allow for varying the reading experience with each new reading aloud of the source text. They also do not allow for adapting the sound effect to the speaker's emotions and reading style. They also do not allow for triggering sound effects in the event of movements of the book made by the speaker for interaction with the book.

[0007] In particular, the sound effects used are based on a principle of “linear music” which uses sounds played in a loop and a transition from one loop to another by cross-fading at the end of predetermined sequences.

[0008] The inventors have thus sought to provide a method overcoming these drawbacks, by allowing the integration of a method similar to the principle of adaptive music when reading a text, in particular a book. Objectives of the invention

[0009] The invention aims to provide a method, a computer program product and a system for receiving and processing audio data comprising words corresponding to the real-time reading of a source text by a speaker, for triggering sound effects synchronized with said reading of the text.

[0010] The invention aims in particular to provide, in at least one embodiment, a method, a computer program product and a system for receiving and processing audio data allowing the taking into account of interactions and variations in emotions when reading the text, for a better renewal of the reading experience, in particular to adapt the sound effect to the emotion that the speaker wants to convey.

[0011] The invention also aims to provide, in at least one embodiment, a method, a computer program product and a system for receiving and processing audio data allowing better monitoring of the reading of the source text to guarantee appropriate triggering of the sound effects at the right time.

[0012] The invention also aims to provide, in at least one embodiment, a method, a computer program product and a system for receiving and processing audio data which can be embedded in a small portable device and which can operate without an Internet connection. Statement of the invention

[0013] To this end, the invention relates to a method for receiving and processing, in a receiving and processing system, audio data comprising words corresponding to the real-time reading of a source text by at least one speaker, for triggering sound effects synchronized with said reading of the text,

[0014] characterized in that it comprises: - a step of determining a position index of the speaker in the source text, by detecting a correspondence between the received audio data and the source text recorded in a storage module of the reception and processing system, - a step of receiving at least one piece of data representative of a prosody value of the audio data, called prosody data, and / or of receiving at least one piece of data representative of the presence or absence of a movement from a movement recognition device during playback of the text and representative of characteristics of a present movement, called movement data, - a step of determining the sound effect to be triggered based on the position index, and based on the prosody data and / or the movement data, - a step of triggering the determined sound effect.

[0015] A reception and processing method according to the invention therefore allows prosody and / or movement data to be taken into account to determine whether a sound effect should be triggered when the position index corresponds to a sound effect triggering position. Taking this additional data into account allows for a renewal of the reading experience and a more immersive and personalized reading experience, a sound effect being able to be triggered or selected, or modified or not, taking into account reading variations. The objective is to adapt the sound effect to the emotion that the speaker of the source text wants to convey.

[0016] The sound effect triggering is thus adaptive, and can evolve in real time at each position index according to the movements and / or prosody, which are respectively representative of actions and emotions. The operation is thus similar to the principle of adaptive music, used in particular in video games where the music changes dynamically according to the player's activity. The sound effects are divided into fragments and several layers which can be selected and possibly combined according to the movements and prosody.

[0017] In particular, taking into account the prosody of reading makes it possible to recognize the emotions of the speaker, for example adapting the sound effects triggered to the sound intensity of reading, to the speed of reading, the tone of reading, etc. A combination of these characteristics of the words constituting the audio data constitutes prosody. More generally, prosody is defined as the set of non-verbal characteristics of a spoken word, in particular variations in rate, pitch (tone and intonation) and variation in duration (accentuation and rhythm), which can be representative of an emotion transmitted in the words when reading the text. The same text can transmit different emotions depending on the prosody of the speech during its reading by a speaker without considering the meaning of the words and sentences spoken.Prosody can also characterize a sociolinguistic accent such as a regional accent and can also add context and meaning to a reading.

[0018] Prosody is detected and recognized through the use of at least one preconfigured dedicated algorithm, for example implementing at least one pre-trained learning model with annotated prosody data, or for example using predetermined prosody thresholds, adjusted using the audio data. speaker-specific. The algorithm(s) take into account the combination of prosody characteristics of the audio data to determine a prosody of the audio data, which can be associated with an emotion conveyed when reading the text.

[0019] The algorithm is preferably fast to execute and self-contained, in particular requiring few computing resources and not requiring a connection to the Internet or any other external database.

[0020] Taking movement into account makes it possible to associate physical interactions with the reading on the movement recognition device during the reading of the text, and to be able to link the reading and the movement to the triggering of a sound effect. The movement data may be data indicating the absence of movement. The movement data may also include information representative of the speed and / or acceleration of the movement.

[0021] The movement is detected and recognized through the use of at least one preconfigured dedicated algorithm, for example implementing at least one pre-trained learning model with annotated movement data, or for example using predetermined movement models, adjusted using the speaker-specific movement data and / or relative to the initial position of the movement recognition device. The algorithm(s) take into account the combination of the movement characteristics detected by the movement recognition device to determine the recognized movement.

[0022] Advantageously and according to the invention, the source text is recorded in the processing device in the form of a list of phonemes, and the processing device is configured to detect in the audio data the presence of a phoneme corresponding to the source text.

[0023] According to this aspect of the invention, the decomposition of the text into phonemes makes it possible to improve the accuracy of the position index as well as the calculation speed and thus to improve the triggering conditions. Phoneme detection allows a better result than word detection and is more compatible with chopped reading, backtracking in reading, pauses in reading, etc.

[0024] Furthermore, phoneme detection requires fewer resources than word detection and can thus be implemented by a portable and / or embedded device, even in the absence of an Internet connection. In particular, the number of different phonemes to be detected is much lower than the number of words that can be detected in methods using word detection and phoneme detection does not require deduction of the next word. The phoneme detection speed is thus of the order of 200ms.

[0025] Advantageously and according to the invention, the position index of the speaker corresponds to a phoneme from the phoneme list of the source text, and the step of determining the position index of the speaker in the source text comprises: - a sub-step of receiving the current position of the index in the list of phonemes of the source text, - a sub-step of comparing at least one phoneme detected in the audio data with at least one phoneme expected among the following phonemes in the list of phonemes of the source text, - if no phoneme detected in the audio data corresponds with the following phonemes, a sub-step of receiving a plurality of phonemes detected in the audio data and a sub-step of searching in at least part of the source text for a sequence of phonemes in the list of phonemes of the source text corresponding to said plurality of detected phonemes, - if the plurality of phonemes detected in the audio data do not correspond with any phoneme sequence in the phoneme list of the source text, a sub-step of searching in the entire source text for a phoneme sequence in the phoneme list of the source text corresponding to said plurality of detected phonemes, - in case of detection of a correspondence of a phoneme or of a plurality of detected phonemes with the list of phonemes of the source text, a sub-step of updating the index with said corresponding phoneme or the last phoneme of the corresponding sequence of phonemes.

[0026] According to this aspect of the invention, tracking the correspondence of the phonemes as described allows efficient determination of the position index for triggering the sound effects.

[0027] Advantageously, if no phoneme detected in the audio data corresponds with the following phonemes, the step of determining the speaker's position index in the source text also comprises a step of advancing the index to the next phoneme. Thanks to this step, the method can continue to advance in the list of phonemes to avoid losing the correspondence between the audio data and the list of phonemes.

[0028] Advantageously and according to the invention, the comparison between the audio data and the phoneme list comprises a determination of an index representative of the resemblance between the phoneme detected in the audio data and the phoneme in the phoneme list. This index representative of the resemblance is also representative of the distance between these phonemes according to its value, for example a value equal to zero indicates a perfect match, a value close to zero indicates a resemblance and a higher positive value indicates a distance between the phonemes. For example, in the French language, phonemes beginning with the letters "p" and "b" are considered to be similar and therefore with a low distance, in the same way as the letters "t" and "d". If the value of the index is lower than a predetermined threshold, it can be considered that there is a correspondence of the phonemes and the index can be advanced. In the variants of the invention where a confidence index as described below is implemented, this confidence index can be lowered if the correspondence of the phonemes is not perfect (resemblance index not zero but lower than the predetermined threshold).Thus, advantageously and according to the invention, the sub-step of comparing at least one phoneme detected in the audio data with at least one phoneme expected among the following phonemes in the list of phonemes of the source text comprises a determination of an index representative of the resemblance between the phoneme detected in the audio data and the phoneme in the list of phonemes, and if the index representative of the resemblance is lower than a predetermined threshold, a sub-step of updating the index with said detected phoneme and a sub-step of reducing the confidence index.

[0029] Advantageously and according to the invention, the step of triggering the sound effect comprises a sub-step of verifying the value of a confidence index and in that the sound effect is triggered only if the confidence index is greater than a predetermined threshold.

[0030] According to this aspect of the invention, the confidence index thus makes it possible to avoid triggering a sound effect in the event of a lack of confidence in the position tracking. In particular, this makes it possible to avoid triggering a sound effect at an inopportune moment due to an incorrect determination of the position index.

[0031] Each sound effect can be linked to its own predetermined threshold.

[0032] Advantageously and according to the invention, the method comprises a step of increasing the value of the confidence index in the event of detection of a correspondence of a phoneme or of a plurality of phonemes detected with the list of phonemes of the source text and comprises a step of decreasing the value of the confidence index if no phoneme detected in the audio data corresponds with the following phonemes.

[0033] According to this aspect of the invention, the increase and decrease of the confidence index is linked to the detection of phoneme correspondence, in particular the confidence index thus makes it possible to avoid the triggering of a sound effect if the detection of phoneme correspondence is not carried out with sufficient confidence.

[0034] Advantageously and according to the invention, all of the sound effects associated with the source text are distributed in a plurality of time sequences associated with portions of the source texts, each time sequence being associated with a sound effect or with a sound effect group comprising multiple sound effects, and when the position index corresponds to a time sequence associated with a sound effect group, the step of determining the sound effect to be triggered comprises: - a sub-step of retrieving, as a function of the position index, the group of sound effects associated with the time sequence in which the position index is located, - a sub-step of determining, based on the prosody data and / or the movement data, the sound effect to be triggered from among the sound effects of the sound effects group.

[0035] According to this aspect of the invention, the time sequences are each associated with one or more sound effects and the triggered sound effects are adapted to the determined prosody and / or movement data. The objective is to propose, for at least certain time sequences, sound effect variants to be triggered for the same position index.

[0036] Advantageously and according to the invention, the sound effects that can be triggered in each group of sound effects can also depend on parameters predefined in advance by the speaker, or can vary according to the prosody data and the confidence index in the associated variants.

[0037] For example, the speaker can select a “day” mode or a “night” mode which makes it possible to prevent the triggering of certain sound effects in the sound effect group. Taking into account the confidence index also makes it possible to limit the risks of untimely triggering of a sound effect at the wrong time, for example by delaying the triggering and / or by not triggering a specific sound effect.

[0038] The speaker can also enable or disable motion or prosody detection for reading.

[0039] Advantageously and according to the invention, the prosody data comprises data or a combination of data from among different types of data in the following list: - data on the intensity of speech in audio data, - data on the frequency and / or fundamental frequency of the speech of the audio data, - data on the intensity and / or frequency and / or fundamental frequency of the consonants pronounced in the audio data, - data on the intensity and / or frequency and / or fundamental frequency of the vowels pronounced in the audio data, - data on the length of vowels and / or spoken words in the audio data,

[0040]

[0041]

[0042]

[0043]

[0044]

[0045] - data on the length of consonants pronounced in the audio data, - data on the speech rate of audio data. According to this aspect of the invention, the prosody depends on several factors that make it possible to symbolize the emotion of the speaker of the source text during the reading aloud of the text. The prosody data can be based on an absolute value of each data in the list or on a temporal variation of the values ​​thereof, for example a gradual increase or decrease in sound intensity. Advantageously and according to the invention, the movement data comprises data relating to a movement and / or a rotation in space and / or to the speed of movement of the movement recognition device, and / or data relating to the correspondence of a movement or a combination of movements detected by the movement recognition device with a predetermined movement from a list of predetermined movements recorded in a storage module of the processing system. According to this aspect of the invention, different types of movement may be taken into account for the selection of the sound effect to be triggered. In particular, the method may comprise detecting a predetermined movement, for example if the user makes a circle with the movement recognition device. Movement in space is known to consist of one or a combination of the following movements: - vertical translation, - lateral translation, - longitudinal translation, - pitching, - lace, - roll - shaking movement, - brief shocks. The speed and / or acceleration of the movement can also be integrated into the movement data. Advantageously and according to the invention, the sound effect triggering step is executed when the position index is at a position equal to or after a predetermined trigger index, or in a predetermined trigger window, and in that the sound effect can be immediately emitted as soon as the position index reaches or exceeds the trigger index or reaches the trigger window, or emitted in a delayed manner after the end of a sound effect currently being emitted.

[0046] According to this aspect of the invention, the triggering of a sound effect may be immediate or delayed, for example the determined sound effect may be triggered at the end of a time sequence. The temporality of the triggering depends in particular on the position or trigger window in which the position index is located.

[0047] The invention also relates to a computer program product for receiving and processing audio data corresponding to the real-time reading of a source text by at least one speaker, for triggering sound effects synchronized with said reading of the text, said computer program product comprising program code instructions for executing, when said computer program product is executed on a computer device, the steps of the method according to the invention.

[0048] The computer program product is advantageously stored in the processing device, in particular in a portable computing device, preferably a smartphone, a digital tablet or a smart watch, for example in the form of an application.

[0049] The invention also relates to a system for receiving and processing audio data corresponding to the real-time reading of a source text by at least one speaker, for triggering sound effects synchronized with said reading of the text, characterized in that it comprises: - a source text storage module, - an audio data reception module, - an audio data processing module, configured to determine a speaker position index in the source text, by detecting a correspondence between the received audio data and the source text recorded in the storage module, - a motion recognition device configured to provide at least one piece of data representative of the presence or absence of a movement of a motion recognition device during the reading of the text and representative of characteristics of a movement present, called motion data, and / or a module for analyzing the variation in the prosody of the audio data configured to provide at least one piece of data representative of a variation in the prosody of the audio data - a module for determining the sound effect to be triggered based on the position index, and based on the prosody data and / or the movement data, - a sound emitting device configured to emit the determined sound effect.

[0050] Advantageously, the reception and processing system according to the invention is configured to implement the reception and processing method according to the invention.

[0051] Advantageously and according to the invention, the reception and processing method according to the invention is configured to be implemented by a reception and processing system according to the invention.

[0052] Advantageously and according to the invention, the system comprises a portable computer device, preferably a smartphone, a digital tablet or a smartwatch, comprising the source text storage module, the audio data reception module, the audio data processing module, the motion recognition device and / or the prosody variation analysis module, the sound effect determination module and the sound emission device.

[0053] According to this aspect of the invention, a major part or all of the components of the system can be integrated into a portable computing device such as a smart phone (commonly called a smartphone), so as to bring together the functionalities in a single system. The portable computing device can also display the source text or preferably be attached to a physical support of the source text, in particular where the source text is printed, such as a book.

[0054] The computing device is configured to implement the reception and processing method according to the invention in particular thanks to a combination of one or more computing components such as a processor (CPU for Central Processing Unit in English), and / or a graphics processor (GPU for Graphics Processing Unit in English), and / or a digital signal processor (DSP for Digital Signal Processor in English, and / or one or more memories, and / or an analog / digital converter, a microphone, an accelerometer, a gyroscope, etc.

[0055] Advantageously and according to the invention, the system comprises a motion recognition device and a device for attaching the motion recognition device, configured to allow the motion recognition device to be attached to a medium on which the source text is printed, so that the motion recognition device is mechanically secured to said medium when reading the medium.

[0056] According to this aspect of the invention, the motion recognition device can be directly associated with the physical medium on which the source text is printed, so that a movement of the physical medium causes a movement of the motion recognition device. Thus, a movement of the physical medium is detected and is taken into account in the motion data.

[0057] When the motion recognition device is integrated into a portable computing device, the tethering device is configured to tether the portable computing device to the physical medium.

[0058] Advantageously and according to the invention, the attachment device is formed from an element or a combination of elements from the elements in the following list: - elastic elements connected to the support on which the source text is printed, configured to hold the motion recognition device in position, - a pocket arranged in the cover of said support, configured to accommodate the motion recognition device, - a compartment arranged in the cover of said support, configured to accommodate the motion recognition device, - a permanent magnet arranged on said support and configured for magnetization with a magnetic element arranged on the motion recognition device, or a permanent magnet arranged on the motion recognition device and configured for magnetization with a magnetic element arranged on said support, - a pocket clipped onto said support, configured to accommodate the motion recognition device.

[0059] According to this aspect of the invention, these different attachment variants allow compatibility with different types of motion recognition device, in particular when the latter is integrated into a portable computing device.

[0060] The invention also relates to a reception and processing method, a computer program product and a reception and processing system characterized in combination by all or part of the characteristics mentioned above or below. List of figures

[0061] Other aims, characteristics and advantages of the invention will appear on reading the following description given solely for non-limiting purposes and which refers to the appended figures in which:

[0062] [Fig. 1] is a schematic view of a method for receiving and processing audio data according to one embodiment of the invention,

[0063] [Fig.2] is a schematic view of a source text and a list of phonemes as recorded in a storage module of a reception and processing system according to an embodiment of the invention,

[0064] [Fig.3] is a graph schematically illustrating a variation in prosody of audio data processed by a reception and processing method according to an embodiment of the invention,

[0065] [Fig.4] is a graph schematically illustrating three audio tracks comprising sound effects that can be triggered when executing a receiving and processing method according to an embodiment of the invention,

[0066] [Fig.5] is a schematic view of a system for receiving and processing audio data according to one embodiment of the invention and according to a front and back view of a book forming a support for the source text.

[0067] Detailed description of an embodiment of the invention

[0068] In the figures, the scales and proportions are not strictly respected, for the purposes of illustration and clarity.

[0069] Furthermore, identical, similar or analogous elements are designated by the same references in all the figures.

[0070] [Fig.l] schematically represents a method 100 for receiving and processing audio data according to an embodiment of the invention. The method allows the reception and processing of audio data comprising words corresponding to the real-time reading of a source text by a speaker, for the triggering of sound effects synchronized with said reading of the text. The method is implemented here in a system for receiving and processing audio data, one embodiment of which is described below with reference to [Fig.5].

[0071] The method comprises a 110 step of determining a position index of the speaker in the source text, by detecting a correspondence between the received audio data and the source text recorded in a storage module of the reception and processing system.

[0072] The source text is recorded in the processing device in the form of a phoneme list, and the processing device is configured to detect in the audio data the presence of a phoneme corresponding to the source text. The phoneme list is created beforehand, automatically, manually or semi-automatically (for example with manual correction of an automatic pre-processing). The creation of the phoneme list is generally carried out outside the system implementing the method according to the invention.

[0073] Figure 2 schematically represents an example 200 of a sentence in French from the source text as recorded in the storage module “In a flash, the cat climbed the tree.”, its transposition into phonemes “ eu ab é k le a le f æt gr pa d æ nz la rb re » and the index assigned to each phoneme associated with this example sentence, the index being in this example between 1 and 19. Other index numbering can be used, for example the first index can correspond to a zero index, etc.

[0074] The speaker position index corresponds to the index of the corresponding phoneme in the phoneme list of the source text, and step 110 of determining the speaker position index in the source text comprises: - a sub-step 112 of receiving the current position of the index in the list of phonemes of the source text. Initially, the position index corresponds to the index of the first phoneme, for example index 1 in the example sentence of [Fig.2]. - a sub-step 114 of comparing at least one phoneme detected in the audio data with at least one phoneme expected from among the following phonemes in the list of phonemes of the source text. The following phonemes in the list of phonemes of the example sentence are the phonemes “”, “ é”, “k” and “le”. - if no phoneme detected in the audio data corresponds with the following phonemes, a sub-step 116 of receiving a plurality of phonemes detected in the audio data and a sub-step 118 of searching in at least part of the source text for a sequence of phonemes in the list of phonemes of the source text corresponding to said plurality of detected phonemes. For example, in the example sentence of Figure 2, the phonemes “le tf æt gr im pa” are detected in the audio data and the position index is then placed at the end of this sequence of phonemes in the source text, i.e. at index 12 which corresponds to the last phoneme of the sequence. - if the plurality of phonemes detected in the audio data do not correspond to any phoneme sequence in the list of phonemes of the source text, a sub-step 120 of searching in the entire source text for a sequence of phonemes in the list of phonemes of the source text corresponding to said plurality of detected phonemes. This sub-step allows the search for a phoneme sequence in the entire source text, for example the sentences following the example sentence or the sentences preceding the example sentence, not visible in the figures. - in case of detection of a correspondence of a phoneme or of a plurality of detected phonemes with the list of phonemes of the source text, a sub-step 122 of updating the index with said corresponding phoneme or the last phoneme of the corresponding sequence of phonemes. This final step of the cycle allows the updating of the index according to the correspondence which was established in the preceding sub-steps. The updated index is retrieved during a new execution of the sub-step 112 of receiving the current position of the index.

[0075] The method may also comprise other sub-steps not described for ensuring the matching of the phonemes detected in the audio text and the phonemes as recorded by the source text.

[0076] The method then comprises a step 130 of receiving at least one piece of data representative of a prosody value of the audio data, called prosody data 10, and / or of receiving at least one piece of data representative of the presence or absence of a movement of a movement recognition device during the reading of the text and representative of characteristics of a present movement, called movement data 12.

[0077] Prosody is representative of the emotion given to the speaker of the source text when reading it. Taking movement into account allows for interactivity when reading the source text.

[0078] Prosody and movement are taken into account independently or in combination in the remainder of the method. This choice can be made by a user of the method, for example using an interface allowing them to select whether prosody and / or movement are taken into account in future readings.

[0079] The prosody data includes data or a combination of data from among different data types in the following list: - data on the intensity of speech in audio data, - data on the frequency and / or fundamental frequency of the speech of the audio data, - data on the intensity and / or frequency and / or fundamental frequency of the consonants pronounced in the audio data, - data on the intensity and / or frequency and / or fundamental frequency of the vowels pronounced in the audio data, - data on the length of vowels and / or spoken words in the audio data, - data on the length of consonants pronounced in the audio data, - data on the speech rate of audio data.

[0080] A combination of these data makes it possible to better characterize the emotion and adapt the triggering of sound effects accordingly.

[0081] [Fig. 3] is an illustrative graph 300 showing an example of prosody versus time, as detected from the audio data. For reasons of simplification and for illustrative purposes only, the prosody in graph 300 takes into account only the intensity and fundamental frequency of the audio data. The value assigned to the prosody data here evolves in two dimensions but in In practice, the value of the prosody data is characterized by the variation of several data, for example, gathered in a vector or a matrix.

[0082] During a first time interval, called calibration interval 310, the intensity and the fundamental frequency are measured to obtain a value which will be considered as average for the speaker of the audio data. The calibration interval 310 can be defined by a trigger, for example from a first predefined position index, and end from a second predefined position index.

[0083] The average prosody value determined during the calibration interval 310 is assigned to a zone called the average zone 312. In the event of variation in the prosody value over time, this value is likely to be found in other zones, a high zone 314 defined by the values ​​above a high threshold 315 and a low zone 316 defined by the values ​​below a low threshold 317. The variation in the value is for example representative of a percentage variation in the fundamental frequency and a percentage variation in the intensity, these two values ​​being able to be weighted differently to give the total value. More zones can be defined depending on the data taken into account in the prosody, for example four zones, five zones, or more zones.The areas thus described are mainly for illustrative purposes and interpretations of prosody values ​​may be based on other classification methods, for example decision trees, etc.

[0084] Each zone then corresponds to an audio track comprising the sound effects to be triggered. When a sound effect must be triggered, for example at times 318a, 318b, 318c, 318d represented by circles, the sound effect of the audio track associated with the zone in which the prosody value is located can be triggered. Here, times 318a, 318c and 318d are linked to the high zone and time 318b is linked to the low zone.

[0085] The triggering of the sound effect can of course be subject to other conditions, in particular to the movement data.

[0086] The movement data comprises data relating to a movement and / or a rotation in space and / or to the speed of movement of the movement recognition device, and / or data relating to the correspondence of a movement or a combination of movements detected by the movement recognition device with a predetermined movement from a list of predetermined movements recorded in a storage module of the processing system. The absence of movement is also movement data which can be interpreted by the method.

[0087] Further details of the motion data are described below with reference to [Fig.5] illustrating a system for receiving and processing audio data.

[0088] Returning to [Fig.l], the method then comprises a step 140 of determining the sound effect to be triggered as a function of the position index, and as a function of the prosody data and / or the movement data. The sound effect triggered is thus adapted as a function of the reading context provided by the prosody and / or the movements.

[0089] In particular, in one embodiment of the invention, all of the sound effects associated with the source text are distributed in a plurality of time sequences associated with portions of the source texts, each time sequence being associated with a sound effect or with a group of sound effects comprising several sound effects, and in that when the position index corresponds to a time sequence associated with a group of sound effects, the step 140 of determining the sound effect to be triggered comprises: - a sub-step 142 of recovering, as a function of the position index, the group of sound effects associated with the time sequence in which the position index is located, - a sub-step 144 of determining, as a function of the prosody data and / or the movement data, the sound effect to be triggered from among the sound effects of the group of sound effects.

[0090] [Fig.4] schematically represents a time graph 400 comprising three audio tracks, a medium track 412, a high track 414 and a low track 416, corresponding respectively to the sound effects to be triggered according to the medium, high and low zones described previously. As described previously, these tracks are mainly for illustrative purposes and in practice the triggering of sound effects may not be associated with particular tracks.

[0091] The audio tracks are played from left to right as the position index and playback progress. The sound effects of the audio tracks are grouped into groups of one to three sound effects associated with time sequences 421, 422, 423, 424, 425, 426, 427, 428, 429.

[0092] The first time sequence 421 is associated with a group which comprises for example a sound effect 421a which is applied regardless of the area reached by the prosody value. The second time sequence 422 is associated with a group which also comprises a sound effect 422a, in particular because these first two time sequences correspond to the calibration interval as described previously.

[0093] The following two sequences 423 and 424 are associated with groups each comprising two sound effects, 423a and 423b on the one hand and 424a and 424b on the other hand. Thus, if the prosody value is in the high zone, the sound effects 423b and 424b will be played while if the prosody value is in the middle or low zone, sound effects 423a and 424a will be played.

[0094] Finally, the following sequences 425 to 429 are associated with groups comprising a sound effect for each zone, referenced respectively 425a, 425b and 425c, 426a, 426b and 426c, 427a, 427b and 427c, 428a, 428b and 428c, 429a, 429b and 429c. Thus, a sound effect is available for the high zone, the middle zone and the low zone reached by the prosody value.

[0095] The sound effect triggering step 150 is executed when the position index is at a position equal to or after a predetermined trigger index, or within a predetermined trigger window, and the sound effect may be immediately emitted as soon as the position index reaches or exceeds the trigger index or reaches the trigger window, or emitted in a delayed manner after the end of a sound effect currently being emitted. If the position index reaches, for example, the marker 430, and the prosody value has changed since the start of the time sequence 424, the sound effect played may be modified immediately. Alternatively, if the position index reaches, for example, the marker 431, and the prosody value has changed since the start of the time sequence 426, the sound effect played may be modified only at the start of the time sequence 427.

[0096] According to a non-represented embodiment, these audio tracks correspond to a sound background and additional sound effects can be added during playback depending on the prosody data and / or the movement data and / or the position index. Furthermore, several sound effects can be combined, depending on the prosody and / or the movement. For example, for a reading of a text, a first sound effect comprising the sound of a single musical instrument can be played and if the sound intensity detected by determining the prosody increases, one or more other sound effects adding musical instruments can be combined with the first sound effect.

[0097] In addition, the number of soundtracks and sound zones may be greater than shown.

[0098] The method finally comprises a step 150 of triggering the determined sound effect. The sound effect is in particular emitted by a sound emission device of the reception and processing system. This sound emission device is for example a loudspeaker or an audio headset.

[0099] The emission of the sound effect thus makes it possible to increase immersion when reading the source text.

[0100] Step 150 of triggering the sound effect comprises a sub-step 152 of verifying the value of a confidence index and the sound effect is triggered only if the confidence index is greater than a predetermined threshold.

[0101] The confidence index is managed by a sub-process 160 for managing the confidence index, which comprises a step 162 for increasing the value of the confidence index in the event of detection of a correspondence of a phoneme or of a plurality of phonemes detected with the list of phonemes of the source text and it comprises a step 164 for decreasing the value of the confidence index if no phoneme detected in the audio data corresponds with the following phonemes, or if the detected phoneme is sufficiently close to the expected phoneme to allow the advancement of the position index but does not have an exact correspondence. The confidence index thus makes it possible not to trigger the sound effect in the event of excessive uncertainty about the accuracy of tracking the phonemes. The confidence index can also be impacted by other parameters.

[0102] Furthermore, a confidence index can also be applied independently to the prosody data.

[0103] [Fig.5] schematically represents a system 500 for receiving and processing audio data according to an embodiment of the invention. The system allows the reception and processing of audio data corresponding to the real-time reading of a source text by at least one speaker, for the triggering of sound effects synchronized with said reading of the text, by implementing the steps of the reception and processing method described previously.

[0104] The reception system 500 here comprises a portable computer device formed by a smart phone 510 (better known as a smartphone in English) comprising a source text storage module, the audio data reception module, an audio data processing module, a motion recognition device and a prosody variation analysis module, a sound effect determination module and a sound emission device, adapted to implement the associated steps of the audio data reception and processing method. The motion recognition device is for example an accelerometer, a gyrometer, a magnetometer or a combination. The different modules are managed for example by the processor and the storage memory(s) of the smart phone 510.

[0105] The receiving system 500 also comprises a device for attaching the motion recognition device, more particularly in this embodiment a device for attaching the smart phone 510 comprising this motion recognition device, configured to allow the motion recognition device to be attached to a support on which the source text is printed, so that the motion recognition device is mechanically secured to said support when reading the support. The support on which the source text is printed the source text here is a book of which [Fig.5] represents a view 590a of the inside of the book and a view 590b of the outside of the book.

[0106] In particular, the attachment device here comprises an elastic band 512 passing through the back cover of the book through notches 514 of the attachment device. The smartphone 510 may for example be arranged at the last endpaper 592a of the book (also called the third cover, third cover or back cover). Alternatively, not shown, the smartphone 510 may be arranged with the same attachment device on the back cover 592b (also called the fourth cover).

[0107] Other types of attachment devices can be implemented: - elastic elements connected to the support on which the source text is printed, configured to hold the motion recognition device in position, - a pocket arranged in the cover of said support, configured to accommodate the motion recognition device, - a compartment arranged in the cover of said support, configured to accommodate the motion recognition device, - a permanent magnet arranged on said support and configured for magnetization with a magnetic element arranged on the motion recognition device, or a permanent magnet arranged on the motion recognition device and configured for magnetization with a magnetic element arranged on said support, - a pocket clipped onto said support, configured to accommodate the motion recognition device - etc.

[0108] The link between the motion recognition device and the support makes it possible to detect a movement of the support, here the book, during reading and to propose a triggering of an associated sound effect. A movement can for example consist of: - a rotation of the book in the manner of a vehicle steering wheel, and trigger associated sound effects if the reading position index is associated with such a triggering, - tapping the book to mime knocking on a door, or playing a percussion instrument, - shake the book to mime the shaking of maracas, a movement of a fan, - lift the book, - left / right pitching to mime the rocking of an animal, a baby, for the movement of a rain stick, - roll forward / backward to speed up / slow down the music, - etc.

[0109] The movement may also have to be executed at a minimum number of occurrences or with a minimum rotation angle to validate the triggering of the sound effect.

[0110] The invention is not limited to the embodiments described. In particular: - sound effects can be predetermined or generated in real time, for example via generative artificial intelligence, - the motion recognition device may not be attached to the source text carrier if it is moving at the same time as the source text (if it is integrated into a smart watch worn by the speaker for example), - detection of certain words in addition to phonemes can be implemented for the detection of special cases, in particular for the detection of personalized words replacing or completing a portion of the source text, - multiple speakers can be detected while reading the source text and their respective prosody taken into account.

Claims

1. Claims Method for receiving and processing, in a reception and processing system, audio data comprising speech corresponding to the real-time reading of a source text by at least one speaker, for triggering sound effects synchronized with said reading of the text, characterized in that the source text is recorded in the processing device in the form of a list of phonemes, and the processing device is configured to detect in the audio data the presence of phonemes corresponding to the source text, and in that the reception and processing method comprises: - a step (110) of determining an index (12) of the speaker's position in the source text corresponding to a phoneme in the phoneme list of the source text, by detecting a correspondence between the received audio data and the source text recorded in a storage module of the reception and processing system, said step (110) of determining the index (12) comprising: • a sub-step (112) of receiving the current position of the index (12) in the list of phonemes of the source text, • a sub-step (114) of comparing at least one phoneme detected in the audio data with at least one phoneme expected among the following phonemes in the list of phonemes of the source text, • if no phoneme detected in the audio data corresponds with the following phonemes, a sub-step (116) of receiving a plurality of phonemes detected in the audio data and a sub-step (118) of searching in at least part of the source text for a sequence of phonemes in the list of phonemes of the source text corresponding to said plurality of detected phonemes, • in the event of detection of a correspondence of a phoneme or a plurality of phonemes detected with the list of phonemes of the source text, a sub-step (122) of updating the index with said corresponding phoneme or the last phoneme of the corresponding sequence of phonemes, - a step (112) of receiving at least one data item representative of a prosody value of the audio data, called prosody data (10), and / or of receiving at least one data item (12) representative of the presence or absence of a movement of a movement recognition device during the reading of the text and representative of characteristics of a present movement, called movement data, - a step (140) of determining the sound effect to be triggered as a function of the position index (12), and as a function of the prosody data and / or the movement data, - a step (150) of triggering the determined sound effect.

2. A receiving and processing method according to claim 1, characterized in that the step (110) of determining the speaker's position index in the source text further comprises, if the plurality of phonemes detected in the audio data do not correspond to any phoneme sequence in the phoneme list of the source text, a sub-step (120) of searching in the entire source text for a phoneme sequence in the phoneme list of the source text corresponding to said plurality of detected phonemes.

3. Reception and processing method according to one of claims 1 to 2, characterized in that the step of triggering the sound effect comprises a sub-step (152) of verifying the value of a confidence index and in that the sound effect is triggered only if the confidence index is greater than a predetermined threshold.

4. Reception and processing method according to claim 3, characterized in that the sub-step (114) of comparing at least one phoneme detected in the audio data with at least an expected phoneme among the following phonemes in the list of phonemes of the source text comprises a determination of an index representative of the resemblance between the phoneme detected in the audio data and the phoneme of the list of phonemes, and in that if the index representative of the resemblance is lower than a predetermined threshold, a sub-step (122) of updating the index with said detected phoneme and a sub-step of decreasing the confidence index.

5. Reception and processing method according to one of claims 1 to 4, characterized in that it comprises a step (162) of increasing the value of the confidence index in the event of detection of a correspondence of a phoneme or of a plurality of phonemes detected with the list of phonemes of the source text and in that it comprises a step (164) of decreasing the value of the confidence index if no phoneme detected in the audio data corresponds with the following phonemes.

6. Reception and processing method according to one of claims 1 to 5, characterized in that all of the sound effects associated with the source text are distributed in a plurality of time sequences associated with portions of the source texts, each time sequence being associated with a sound effect or a group of sound effects comprising several sound effects, and in that when the position index corresponds to a time sequence associated with a group of sound effects, the step of determining the sound effect to be triggered comprises: - a sub-step (142) of retrieving, as a function of the position index, the group of sound effects associated with the time sequence in which the position index is located, - a sub-step (144) of determining, as a function of the prosody data and / or the movement data, the sound effect to be triggered from among the sound effects of the group of sound effects.

7. Receiving and processing method according to one of claims 1 to 6, characterized in that the prosody data comprises data or a combination of data from different types of data in the following list: - data on the intensity of the speech in the audio data, - data on the frequency and / or fundamental frequency of the speech in the audio data, - data on the intensity and / or frequency and / or fundamental frequency of the consonants spoken in the audio data, - data on the intensity and / or frequency and / or fundamental frequency of the vowels spoken in the audio data, - data on the length of the vowels and / or speech spoken in the audio data, - data on the length of the consonants spoken in the audio data, - data on the speech rate of the audio data.

8. Reception and processing method according to one of claims 1 to 7, characterized in that the movement data comprises data relating to a movement and / or a rotation in space and / or to the speed of movement of the movement recognition device, and / or data relating to the correspondence of a movement or a combination of movements detected by the movement recognition device with a predetermined movement from a list of predetermined movements recorded in a storage module of the processing system.

9. A receiving and processing method according to one of claims 1 to 8, characterized in that the step (150) of triggering a sound effect is executed when the position index (12) is at a position equal to or after a predetermined trigger index, or in a predetermined trigger window, and in that the sound effect can be immediately emitted as soon as the position index reaches or exceeds the trigger index or reaches the trigger window, or emitted in a delayed manner after the end of a sound effect currently being emitted.

10. Computer program product for receiving and processing audio data corresponding to the real-time reading of a source text by at least one speaker, for triggering sound effects synchronized with said reading of the text, said computer program product comprising code instructions

11. program for executing, when said computer program product is executed on a computer device, the steps of the method according to one of claims 1 to 9. System for receiving and processing audio data corresponding to the real-time reading of a source text by at least one speaker, for triggering sound effects synchronized with said reading of the text, characterized in that the source text is recorded in the processing device in the form of a list of phonemes, and the processing device is configured to detect in the audio data the presence of phonemes corresponding to the source text and in that the receiving and processing system comprises: a source text storage module, an audio data reception module, an audio data processing module, configured to determine a speaker position index in the source text corresponding to a phoneme in the phoneme list of the source text, by detecting a correspondence between the received audio data and the source text recorded in the storage module, and configured to: • a reception of the current position of the index (12) in the list of phonemes of the source text, • a comparison of at least one phoneme detected in the audio data with at least one phoneme expected among the following phonemes in the phoneme list of the source text, • if no phoneme detected in the audio data corresponds with the following phonemes, receiving a plurality of phonemes detected in the audio data and searching in at least part of the source text for a sequence of phonemes in the list of phonemes of the source text corresponding to said plurality of detected phonemes, • in the event of detection of a correspondence of a phoneme or a plurality of detected phonemes with the list of phonemes of the source text, an update of the index with said corresponding phoneme or the last phoneme of the corresponding sequence of phonemes, - a motion recognition device configured to provide at least one data item representative of the presence or absence of a movement of a motion recognition device during the reading of the text and representative of characteristics of a present movement, called motion data, and / or a module for analyzing the variation of the prosody of the audio data configured to provide at least one data item representative of a variation of prosody of the audio data - a module for determining the sound effect to be triggered as a function of the position index, and as a function of the prosody data and / or the motion data, - a sound emission device configured for the emission of the determined sound effect.

12. System according to claim 11, characterized in that it comprises a portable computing device (510), preferably a smartphone, a digital tablet or a smart watch, comprising the source text storage module, the audio data reception module, the audio data processing module, the motion recognition device and / or the prosody variation analysis module, the sound effect determination module and the sound emission device.

13. System according to one of claims 11 or 12, characterized in that it comprises a motion recognition device and a device (512) for attaching the motion recognition device, configured to allow the motion recognition device to be attached to a support on which the source text is printed, so that the motion recognition device is mechanically secured to said support when reading the support.

14. System according to claim 13, characterized in that the attachment device is formed from an element or a combination of elements from the elements of the following list: elastic elements (512) connected to the support on which the source text is printed, configured to hold the motion recognition device in position, a pocket arranged in the cover of said support, configured to accommodate the motion recognition device, a compartment arranged in the cover of said support, configured to accommodate the motion recognition device, a permanent magnet arranged on said support and configured for magnetization with a magnetic element arranged on the motion recognition device, or a permanent magnet arranged on the motion recognition device and configured for magnetization with a magnetic element arranged on said support, a pocket clipped onto said support, configured to accommodate the motion recognition device.