Intelligent voice synthesis

EP4584779A1Pending Publication Date: 2025-07-16ORANGE SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2023762542
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-08
Filing Date
2023-09-06
Publication Date
2025-07-16

AI Technical Summary

Technical Problem

Existing speech synthesis systems lack true interactivity and adaptability during presentations, relying on pre-established scenarios or human intervention for alternating between human speech and voice synthesis, limiting audience engagement and flexibility.

Method used

A method for real-time analysis of a speaker's speech to automatically select and transition between groups of words in a text, enabling intelligent speech synthesis that adapts to the progress of a presentation without pre-defined scenarios, using voice recognition and analysis to detect interruptions and resume speech, ensuring seamless integration with the speaker's narrative.

Benefits of technology

This approach allows for dynamic and context-aware audio alternation between human speech and voice synthesis, enhancing presentation experience by providing continuous and adaptive speech synthesis that mirrors the speaker's progress, maintaining audience engagement and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 1.1
    Figure 1.1
Patent Text Reader

Abstract

The invention relates to a method for automatically reading a continuous text made up of several groups of words, and to a corresponding computer program, a recording medium, an automatic reader and a user terminal. The method comprises providing (7) in real time a sound stream corresponding to the text. The sound stream starts from a second selected group of words, also referred to as a second group of words (6) in the text as a function of at least one result of a real-time analysis (2) of captured speech (1). The result of the analysis is indicative of a first group of words being verbalised by a speaker, the first group of words and the second group of words being distinct groups of words.
Need to check novelty before this filing date? Find Prior Art

Description

Intelligent speech synthesis

[0001] This disclosure relates to the field of speech synthesis.

[0002] More particularly, the present disclosure relates to a method for automatically reading a text and to a corresponding computer program, recording medium, automatic reader and user terminal.

[0003] Text-to-speech (TTS) is the transformation or transcription of written text into an audio rendition of the same content. The voice type and speaking speed can be adjusted.

[0004] If we want to create a synchronized audio mix between oral interventions of a user who is reading or presenting a text and speech synthesis interventions relating to this same text, a known possibility is to allow the user to trigger interruptions and restarts of the speech synthesis at desired points. The management of the audio alternation between human speech and speech synthesis linked to the same content can be carried out by human intervention. These interventions using manual or vocal interactions for example can trigger various functions of playing, pausing, stopping, or even moving to the next or previous chapter.

[0005] Another known possibility is to implement a pre-established setup relating to a pre-prepared scenario. Such setup can be described as semi-automated in that the setup is carried out by a human before the presentation, but no human intervention is then required during the presentation to activate the play, pause, stop, or other functions. A disadvantage of the pre-established setup is the limited interactivity offered with the audience, as the speaker is forced to adhere to the pre-prepared scenario.

[0006] There is therefore a need for a truly automatic, even contextual, implementation of audio alternation between human speech and voice synthesis relating to the same text, i.e. without human intervention and without relying on any pre-prepared scenario. Summary

[0007] This disclosure improves the situation.

[0008] A method for automatically reading a continuous text composed of several groups of words is proposed, the method comprising a real-time supply of a sound stream corresponding to the text, the sound stream starting from a second group of words chosen, in the text, according to at least one result of a real-time analysis of captured speech, the result of the analysis being indicative of a first group of words being verbalized by a speaker, the first group of words and the second group of words being distinct groups of words.

[0009] Continuous text can be a presentation, speech, narrative, or other support. It can be a text prepared in advance and written, for example, using a word processor. Continuous text can also result from automatic processing of a screenshot or a photographic capture of a slide presented by a speaker, such automatic processing involving, for example, character recognition. A group of words can designate, for example, one or more sentences or one or more constituents of a sentence, for example, one or more clauses.

[0010] It is understood that, according to the proposed method, the chosen group of words, also called second group of words, is the result of an automatic choice in the continuous text.

[0011] The audio stream can be a simple or enriched transcription of a portion of the continuous text beginning with the chosen group of words, also called the second group of words. According to an example of enriched transcription, the audio stream can include introductory words such as "let's resume", "a little flashback" or "I introduce myself, I am the Text-To-Speech assistant...".

[0012] The proposed method provides intelligent speech synthesis rendering that automatically adapts to the flow of a speech or presentation. This intelligent rendering results from the selection of a second relevant group of words as the starting point of the audio stream, this choice resulting from the real-time analysis of a user's current speech.

[0013] The features set out in the following paragraphs may optionally be implemented. They may be implemented independently of each other or in combination with each other.

[0014] In one example, the provision of the audio stream is triggered if a speech interruption by the speaker is detected. Speech interruption detection refers to the detection of any explicit or implicit interaction by the speaker, or any combination of such interactions, that reflects a temporary cessation of speech. Silence, hesitation, or a particular posture are all examples of implicit interactions that may be captured and interpreted for the purpose of such detection.

[0015] In one example, the provision of the audio stream is interrupted if a resumption of speech by the speaker is detected. Resumption of speech detection refers to the detection of any explicit or implicit interaction by the speaker, or any combination of such interactions, that reflects a resumption of speech or a cessation of a speech interruption. Real-time analysis of captured speech, alone or in combination with other real-time analyses, may, for example, be used to detect speech interruptions and resumptions.

[0016] When the two examples above are combined, speech synthesis is likely to automatically take over in the event of an impromptu and temporary interruption in speech until the speaker resumes speaking later.

[0017] In one example, the chosen group of words, also called the second group of words, is identical or consecutive, in the text, to the group of words currently being verbalized by the speaker, also called the first group of words.

[0018] Real-time analysis of captured speech can, for example, make it possible to determine not only a group of words currently being verbalized. When the group of words currently being verbalized includes several words, the analysis also makes it possible to indicate whether this group of words becomes fully verbalized or whether, on the contrary, it remains only partially verbalized. By fully verbalized is meant that the user has verbalized all the words of this first group of words, and by partially verbalized is meant that the user has verbalized at least one word of this second group of words but not all the words of this second group of words.Such an indication can have an impact both on the result of the analysis, the first group of words of which will be respectively the group of words currently being verbalized, fully verbalized, or the group of fully verbalized words preceding the group of partially verbalized words, and on the choice of the second group of words with which to begin the voice synthesis.

[0019] To illustrate this point, the example of triggering speech synthesis following the detection of a speech interruption is now repeated. If the speech interruption occurs during the verbalization, which remains partial, of a group of words comprising several words, it may be desirable for the analysis to indicate that the first group of words is the group of words preceding the partially verbalized group of words and to begin the speech synthesis with a complete repetition of this same partially verbalized group of words then constituting the second group of words. If, conversely, the speech interruption occurs just after the complete verbalization of a first group of words and just before the start of the verbalization of a second immediately consecutive group of words, then it may be desirable to begin the speech synthesis directly with the utterance of this second group of words.

[0020] In one example, the result of the real-time analysis is indicative of several first groups of words successively verbalized by the speaker, and the chosen group of words, also called second group of words, is identical or consecutive to the group of words closest to the end of the text among the first groups of words having been verbalized or being verbalized by the speaker.

[0021] For example, it is common for identical or similar clauses to be repeated in different sentences, or for identical or similar sentences to be repeated in different passages of the same text. Choosing to start the speech synthesis with the second group of words following the last group of words similar to the first group of words being verbalized, among those already verbalized by the speaker, makes it possible to avoid repetitions that could annoy the audience.

[0022] In one example, the method is implemented during a session and the chosen group of words, also called the second group of words, is a group of words not appearing in the speech captured during the session and / or not appearing in an audio stream provided during the session prior to the implementation of the method.

[0023] Thus, it is possible, for example, to start the voice synthesis with the group of words positioned first in the text that has neither been verbalized by the speaker nor been the subject of a previous voice synthesis during the session. This makes it possible to restore the entire content of the text while avoiding any repetition.

[0024] Also provided is a computer program comprising instructions for implementing the above method when this program is executed by a processor.

[0025] Also provided is a non-transitory computer-readable recording medium on which is recorded a program for implementing the above method when this program is executed by a processor.

[0026] An automatic reader is also proposed comprising a real-time provider of sound streams, the sound stream corresponding to a continuous text composed of several groups of words, the sound stream starting from a chosen group of words, also called a second group of words, in the text, based on at least one indication of a first group of words being verbalized by a speaker, the indication coming from a real-time analyzer of captured speech.

[0027] A user terminal is also provided comprising a real-time sound stream provider and a sound card, the provider being connected to the sound card and capable of providing a sound stream to the sound card, the sound stream corresponding to a continuous text composed of several groups of words, the sound stream starting from a chosen group of words, also called a second group of words, in the text, based at least on a result indicative of a first group of words being verbalized by a speaker, the result coming from a real-time analyzer of captured speech.

[0028] In one example, the sound card is connected to one or more of the following speakers: a speaker of the user terminal, a speaker of a device connected to the user terminal via a local network.

[0029] The connections between the sound card and the speaker(s) can be wired or via radio communication.

[0030] In one example, the user terminal further comprises a display of the text.

[0031] In one example, the user terminal further comprises a real-time word processing device adapted to highlight a group of words of the text based on the result and provide the text with the highlighted group of words to the display.

[0032] Providing both the audio stream and the text in real time with the highlighted group of words enhances the accessibility of the presentation.

[0033] Other features, details and advantages will become apparent upon reading the detailed description below, and upon analyzing the attached drawings, in which: Fig. 1

[0034] represents a sequence of manually triggered audio alternation between human speech and voice synthesis related to the same content. Fig. 2

[0035] illustrates by a flowchart a method of automatic reading of a text, according to an example of implementation. Fig. 3

[0036] represents a set of data considered successively to operate an automatic audio transition from human words to a voice synthesis linked to the same content, according to a particular exemplary embodiment. Fig. 4

[0037] represents a sequence of automatic audio alternation between human speech and voice synthesis linked to the same content, according to the particular embodiment of. Fig. 5

[0038] represents a set of data considered successively to operate an automatic audio transition from human speech to a voice synthesis linked to the same content, according to a set of particular exemplary embodiments. Fig. 6 Fig. 7

[0039] and each represent a sequence of an automatic audio alternation between human speech and voice synthesis linked to the same content, according to two examples from the set of particular embodiment examples of.

[0040] It is known to control a speech synthesis method by means of manual actions. This is an illustrative example of the prior art where a positioning action (102) in the text can be combined with a launching action (104) of the speech synthesis in order to start a broadcast of an audio signal from a desired location in the text. A pausing or stopping action (106) of the speech synthesis can subsequently make it possible to stop the broadcast of the audio signal at another desired location.

[0041] The invention differs from the prior art and aims to intelligently mix the speech of the speaker who reads or presents from a text medium with appropriate parts of the same text reproduced in voice synthesis.

[0042] Automatic and live accompaniment during audio presentations allows for voice synthesis relays based on the instantaneous flow of the presentation.

[0043] These relays offer various benefits to the experience shared by the speaker and his audience.

[0044] For example, choosing a synthesized voice distinct from that of the speaker makes it possible to simulate interventions by a second speaker and thus obtain a two-voice effect.

[0045] The speaker can also be replaced in case of difficulty speaking for a long time, in case of forgetting the text, stress, shortness of breath, external disturbance such as a telephone call, etc. The choice of a synthetic voice identical to that of the speaker can allow the audience not to perceive the substitution.

[0046] A particular example of embodiment is now described with reference to which visually represents an algorithm corresponding to a method of automatic reading of a text.

[0047] During a session corresponding to a presentation, a speech or any other event involving an audio reproduction of a text medium, the words of one or more human speakers are captured (1) by means of one or more microphones.

[0048] These words are analyzed (2) in real time by an analyzer implementing a voice recognition algorithm. Such algorithms are well known to those skilled in the art and are not detailed here.

[0049] Real-time analysis of captured speech makes it possible to determine (3), at any time, a first group of words being verbalized by a speaker. The first group of words being verbalized can be found literally in the text support. It can also be a variation that can be assimilated to a first group of words present in the text support. Finally, it can be a digression initiated by the speaker, that is to say at least one group of words accompanying the audio reproduction of the text but which cannot be linked to any particular group of words in the text support.

[0050] The first group of words being verbalized can be stored in memory. Storing in memory the successive groups of words being verbalized throughout a speaker's intervention corresponds to forming a history of the verbalized groups of words. When the speaker's intervention deviates from the text support, it can be useful to automatically process the history by comparing it with the text support in such a way as to consider, among the verbalized groups of words, only groups of words which either actually appear in the text or are equivalent to groups of words which actually appear in the text. Obtaining (8) such a history therefore makes it possible to identify, at any time during a speaker's intervention, the groups of words in the text which have already been verbalized, literally or not, by the speaker, the one currently being verbalized by the speaker and finally those in the text which remain to be verbalized.

[0051] The result of the real-time analysis of the captured speech is used to choose (6) a position in the text, i.e. a second group of words in the text from which to start a speech synthesis of the rest of the text. The logical link between the result of the analysis of the captured speech and the chosen group of words, also called second group of words, is explained through several examples in the rest of this document.

[0052] Voice synthesis can then be implemented, and a sound stream corresponding to the result of the voice synthesis can be provided (7) for example in the form of a digital signal intended to be reproduced by one or more loudspeakers.

[0053] In addition, the groups of words in the text that have been the subject of voice synthesis can be identified as such and can be stored in the history of verbalized groups of words. Obtaining (8) such a history thus makes it possible to identify, at any time during the session, the groups of words in the text that have already been verbalized or are in the process of being verbalized either by the speaker or by voice synthesis and those that remain to be verbalized.

[0054] In the example of the, it is provided, optionally, not to implement automatic reading while the speaker is speaking and to trigger (5) automatic reading when an interruption in the speaker's speech is detected (4).

[0055] In general, it is possible to define pre-established situations and to provide for triggering, or interrupting, automatic reading upon detection of such a pre-established situation. Speech interruption represents here a particular example of a pre-established situation usable as a trigger for automatic reading. Correspondingly, a resumption of speech can represent an example of a pre-established situation which, when detected, causes an interruption of automatic reading.

[0056] A pre-established situation can be detected (4) by interpreting data from one or more sensors. These data may be indicative of an interaction or a set of interactions of the speaker. These interactions can be explicit or implicit.

[0057] Various examples of data that can be captured and interpreted in such a way as to lead to the detection of a pre-established situation are now provided.

[0058] Background noise, a technical failure of the speaker's microphone, or a loss of connection are examples of speech capture incidents. Such incidents are detectable by various known technical means and correspond to an inability to reproduce the speaker's words, which may constitute an example of a pre-established situation.

[0059] Silence or a significant slowdown in speech rate are examples of implicit speaker interactions that can be detected by low-level analysis of captured speech. These examples of implicit interactions are indicative of a time period during which no group of words is being verbalized by the speaker, which corresponds to a literal interruption of speech by the speaker. Speech synthesis can be triggered, for example, by comparing the duration of this time period with a configurable threshold, for example, of the order of a few seconds. Below this threshold, the speech interruption is considered a normal pause in the speech that does not justify a relay in speech synthesis, and conversely, above this threshold, the speech interruption is considered too long and a relay in speech synthesis is automatically provided.

[0060] Other thresholds for triggering or interrupting speech synthesis may be defined, on a case-by-case basis, depending on the nature of the data captured and / or the results of analysis of the data captured. These thresholds may be set manually or automatically.

[0061] For example, the setting of a threshold relating to the duration of a pause in speech, determined by analysis of the captured speech, may be a function of past analysis results of the speech of the speaker in question and / or based on criteria relating to a desired audio reproduction quality.

[0062] A stutter, a hesitation or more generally an indication of fatigue or lack of intelligibility, as well as a digression are other examples of implicit interactions of the speaker. These examples of implicit interactions can be detected by speech recognition and can be interpreted as proven or desired interruptions of the oral restitution of the text support by the speaker. When, for example, detected hesitations exceed a certain frequency threshold during a given time period, then it can be automatically planned to ensure a relay in speech synthesis to spare the speaker.

[0063] In addition to the speaker's words, other types of data can be captured in real time. Images from a video capture of the speaker by a camera during the session are an example of data that can be analyzed in real time, and the result of such analysis can be used to detect events corresponding to predetermined situations. Event detection can be based, for example, on indications relating to a movement of the speaker, such as a lip movement, a change in gaze direction, a rotation of the head, a gesture, a change in posture, a movement, etc.

[0064] Certain predetermined situations may correspond simply to a reception of one or more explicit instructions from the speaker, for example by interaction of the speaker with a display element or a button provided for this purpose, or by a gesture of the speaker detectable for example by a motion sensor, or even by a vocal instruction from the speaker detectable by voice recognition.

[0065] It is understood that the proposed technique is not limited to embodiments where automatic playback is triggered from an event occurring during the session.

[0066] To illustrate this point, in one example, the audio stream corresponding to the captured speech and that corresponding to the speech synthesis can be automatically provided continuously throughout the duration of the session, for example in the form of two separate tracks each intended to be played exclusively. No triggering of automatic playback is therefore imposed in this example. It should be noted, however, that the provision of the speech synthesis track requires an underlying mechanism for automatically synchronizing the speech read by speech synthesis with those read by the speaker to preserve harmony and fidelity to the speech in real time. The details of such a mechanism are not discussed in this document.

[0067] The possibility of switching from one track to another can be provided for example by means of manual interactions and / or automatically depending on the progress of the session.

[0068] The sound stream corresponding to the voice synthesis can also be modified in real time based on the result of the analysis of the captured words. The modification may notably include a choice, in the text, of a second group of words to be reproduced by voice synthesis corresponding to that currently being verbalized by the speaker. This is therefore an adaptation of the voice synthesis track by groups of words consistent with the groups of words successively being read by the speaker.

[0069] The aim in such an example is to provide automatic, real-time voice synthesis of the speaker's intervention while ensuring that the groups of words thus synthesized are consistent with those of the text support.

[0070] Reference is now made to Figures 3 and 4, which refer to the same particular example. They illustrate a logical path for choosing a second group of words with which to begin a speech synthesis. They illustrate a sequence of automatic audio alternation between a speaker's words and a speech synthesis starting with the second group of words thus chosen.

[0071] In this example, we consider that a speaker took the floor during a session to vocally reproduce, at least, the content of a text medium "c". The text medium is conceptually divided into consecutive parts noted "Txt A", "Txt B"... each formed of one or more groups of words, the parts "Txt A", "Txt B"... of the text medium thus corresponding to propositions, sentences, or passages composed of several sentences.

[0072] The speaker's words (100), denoted "Audio A'", are captured (1) and analyzed (2) in real time. At a given moment, the analysis of the captured words includes a real-time transcription of a first group of words being verbalized, the result of which is a piece of text denoted "Txt A'" (200) and an interpretation of the transcription thus obtained.

[0073] The analysis makes it possible to establish (3) a correspondence between the captured words “Audio A'” and at least one part “Txt A” of the text support “c”.

[0074] Ideally, when the speaker reads his text strictly, the correspondence is easy and quick. In other cases, such as during presentations on a given topic, the speaker may use synonyms, add or remove words, add or remove details or clarifications.

[0075] The correspondence can be obtained by comparing the transcription result with the text support. A given piece of text "Txt A'" can for example be associated with a given part "Txt A" of the text support by detecting similarity or by detecting the inclusion of one in the other (either the inclusion of "Txt A'" in "Txt A" or conversely the inclusion of "Txt A" in "Txt A'").

[0076] When a speech interruption, i.e. a pause by the speaker, is detected (4) at a given moment, the established correspondence makes it possible to determine (6) a place (600) in the text at which the speaker has arrived. In other words, the established correspondence makes it possible to identify the next group of words in the text to be uttered in order to continue the speech in a coherent manner.

[0077] If the pause occurred abruptly in speech, for example, in the middle of a sentence, the next group of words to be uttered, also called the second group of words, may be the group of words partially verbalized by the speaker at the time of the pause. If the pause occurred more smoothly in speech, for example, after the end of a sentence, the next group of words to be uttered, also called the second group of words, may be the group of words following the first group of words last verbalized by the speaker.

[0078] To ensure a relay following the speaker's pause, an audio stream (700) is provided (7), this audio stream beginning with the "Txt B" part of the text medium comprising the next group of words to be spoken, also called the second group of words. It may be provided that, by default, this audio stream continues automatically until the end of the text medium. It may also be provided that the audio stream is automatically interrupted if a resumption of speech by the speaker is detected.

[0079] Reference is now made to Figures 5, 6 and 7 which illustrate a more complex set of specific examples where a text medium contains repetitions of the same group of words being verbalized.

[0080] This illustrates a logical path for choosing a second group of words with which to begin the speech synthesis in these more complex cases. Figures 6 and 7 each illustrate a sequence of an automatic audio alternation between a speaker's words and a speech synthesis starting with a second group of words thus chosen.

[0081] As in the example of figures 3 and 4, the words (100) of the speaker, noted "Audio A", are captured (1) and analyzed (2) in real time.

[0082] At a given, current moment, the analysis of the captured words includes a real-time transcription of a first group of words being verbalized, the result of which is a piece of text noted "Txt A'" (200) and an interpretation of the transcription thus obtained.

[0083] To implement an automatic relay by voice synthesis starting for example from the current moment, it is necessary to automatically choose the next group of words to be spoken, also called the second group of words, and different settings can be chosen for this purpose.

[0084] In the set of examples in Figures 5, 6, and 7, the piece of text "Txt A'" (200) is first associated (3), by similarity or inclusion, with several parts of the text medium, for example, three parts denoted "Txt A1" (302), "Txt A2" (304), and "Txt A3" (306). It is also assumed, in each of these examples, that the speaker does not read the content, also called text medium, "c" in a linear manner. Thus, the parts "Txt A1", "Txt A2," and "Txt A3" are included in this order in the person's oratory, that is, the speaker first reads the part "Txt A1," then "Txt A2," and finally "Txt A3." On the other hand, the order of appearance of the parts in the content "c" is different.Thus, the parts "Txt A1", "Txt A3" and "Txt A2" appear in this order in the content c, i.e. a reader such as the speaker or the automatic reader reading linearly the content "c" would first read the part "Txt A1" then "Txt A3" and finally "Txt A2". .,.

[0085] The parts "Txt A1" (302), "Txt A2" (304), and "Txt A3" (306) are distinct and distributed discontinuously in the text medium, that is to say they cannot be merged into a single continuous part of the text medium. In this case, to ensure a relay in particular following a detected pause (4) of the speaker, a sound stream (700) is provided, this sound stream starting with the part "Txt B3" of the text medium comprising the next group of words to be spoken, also called the second group of words, following the part "Txt A3", also called the first group of words, associated with the text "Txt A" verbalized by the speaker. According to this definition, the parts "Txt A3" (first group of words) and "Txt B3" (second group of words) can be contiguous.Alternatively, the "Txt A3" and "Txt B3" sections may overlap very slightly, i.e. include a common group of words corresponding to a group of words whose verbalization was interrupted by the speaker's pause. It may be provided that, by default, this sound stream continues automatically until the end of the text support. It may also be provided that the sound stream is automatically interrupted if a resumption of speech by the speaker is detected.

[0086] This association can be related to two other different scenarios. In these two other cases, the result of the association does not allow us to identify with certainty the part of the text support being orally restored by the speaker but only allows us to identify several candidates which are, in this example, the three distinct parts "Txt A1" (302), "Txt A2" (304), and "Txt A3" (306) of the text support "c". In these two cases, the words "Txt A'" of the speaker were uttered in the following temporal order: "Txt A1" followed by "Txt A2" and finally "Txt A3". Analysis (2) therefore finds from "Txt A'" the 3 groups of words "Txt A1", "Txt A2", and "Txt A3" forming part of the reference discourse (of the text support "c").

[0087] Note, as already indicated above: - "Txt A2" corresponds to the group of words furthest in position in the reference text or support text "c" but does not correspond to the first group of words last spoken by the speaker; - "Txt A3" corresponds to the first group of words last spoken by the speaker but is positioned upstream in the reference text or support text "c". This may correspond to the fact that the speaker forgot (skipped) the group of words "Txt A3" and went from "Txt A1" to "Txt A2" then realized his forgetfulness and continued orally with "Txt A3" which does not correspond to the order of the reference text "c".

[0088] In a first case illustrated on the, the choice of the next group of words to be vocally synthesized, also called second group of words, can be the group of words following first the part closest to the end of the text medium, here "Txt A2". This choice makes it possible to avoid repetitions even if it means not reproducing the entire text medium. For example, the speaker reads the content "c", sensors such as microphones provide a captured audio signal 100, a real-time transformation of speech into text, in particular voice recognition, generates the text 200 corresponding to the captured audio 100.An analysis of the content "c" makes it possible to determine that the text "Txt A" spoken by the speaker potentially corresponds to one or more parts of the content "c", in this case in the spoken order to parts 302, 304 and 306, since the speaker does not read the content c in the writing order but first parts 302 followed by 304 and returns to part 306 (placed before 304 in the text support c). In the example of the, the interruption of reading by the speaker is estimated to correspond to the end of the most distant part in the text support c, in this case part 304 triggering the start of the voice synthesis with the beginning of part B2. Eventually, at a given moment during the voice synthesis of the content "c", the speaker can resume reading, thus interrupting the voice synthesis. This marks the end of part B2.

[0089] In a second case illustrated on the, the choice of the next group of words to be stated, also called second group of words, can be the group of words appearing first after the last part 306 associated with the text support being orally rendered by the speaker, here "Txt A3". This choice ensures continuity of the speech at the risk of causing repetitions. For example, the speaker reads the content “c”, sensors such as microphones provide a captured audio signal 100, a real-time speech-to-text transformation, in particular voice recognition, generates the text 200 corresponding to the captured audio 100. An analysis of the content “c” makes it possible to determine that the text “Txt A'” spoken by the speaker potentially corresponds to one or more parts of the content “c”, in this case in the spoken order to parts 302, 304 and 306 because the speaker having skipped passage 306 before reading passage 304, will read it after.In the example of the, the interruption of reading by the speaker is estimated to correspond to the end of part 306 triggering the start of speech synthesis with the beginning of part B3. Optionally, at a given time during the speech synthesis of the content "c", the speaker can resume reading thus interrupting the speech synthesis. This marks the end of part B3, which can then possibly overlap or include part 304. .

[0090] It is also possible to take into account all the parts of text already presented, by means of a history of captured speech and / or content previously provided by voice synthesis, in order to choose the next group of words to be spoken, also called the second group of words.

[0091] Three specific examples of applications of the proposed technique are now described for illustrative purposes.

[0092] In a first example, Pierre planned to deliver a presentation with his colleague Paul that they had prepared together, alternating their speaking engagements for better dynamics but also because each is a little more specialized in certain aspects than the other. Unfortunately, at the last moment, Paul was unable to be present and accompany him. Pierre provided the presentation support in the form of a text file to an automatic reading service implementing a realization of the proposed automatic reading technique. Pierre thus felt reassured and would not hesitate to take breaks at any time, knowing that the service would take over.

[0093] In a second example, Jeanne is orally accompanying, using a microphone, a presentation of her latest tutorial video in a meeting room with her colleagues. During the presentation, she receives a message or a call via her phone requiring an urgent response. She cannot interrupt the current video, and it is obviously preferable that the speech not be interrupted. She steps away for a moment into the next room to make a brief phone call. During this time, according to one implementation of the proposed technique, a service automatically detected that Jeanne was no longer speaking into the microphone and activated a text-to-speech module to take over by broadcasting the rest of the planned speech. Thus, the listeners captivated by the video were practically unaware of the replacement, especially since Jeanne had configured the synthesized voice to clone her own.As soon as she returns and picks up the microphone, the voice synthesis automatically stops, and Jeanne continues her explanations.

[0094] In a third example, Rose gives a presentation despite a sore throat, having previously activated a service in the background that implements a realization of the proposed technique. For the first 15 minutes, everything goes well, then her throat starts to irritate her, and she can no longer express herself as easily as she would like. With a click, she activates the voice synthesis while she recovers. She feels less embarrassed and can resume whenever she wants.

Claims

Method for automatically reading a continuous text composed of several groups of words, the method comprising a real-time supply (7) of a sound stream corresponding to the text, the sound stream starting from a second chosen group of words, also called a second group of words (6), in the text, based at least on a result of a real-time analysis (2) of captured words (1), the result of the analysis being indicative of a first group of words being verbalized by a speaker, the first group of words and the second group of words being distinct groups of words. Method according to claim 1, the provision (7) of the sound stream being triggered (5) if an interruption in the speaker's speech is detected (4). Method according to claim 2, the supply (7) of the sound stream being interrupted if a resumption of speech by the speaker is detected. Method according to one of claims 1 to 3, the chosen group of words, also called the second group of words (6) being identical or consecutive, in the text, to the group of words currently being verbalized by the speaker, also called the first group of words. Method according to one of claims 1 to 3, in which the result of the analysis (2) in real time is indicative of several groups of words successively verbalized by the speaker, and the chosen group of words, also called second group of words (6) is identical or consecutive to the group of words closest to the end of the text among the groups of words having been verbalized or being verbalized by the speaker. Method according to one of claims 1 to 5, in which the method is implemented during a session and the chosen group of words, also called second group of words (6) is a group of words not appearing in the words captured during the session and / or not appearing in a sound stream provided during the session prior to the implementation of the method. Computer program comprising instructions for implementing the method according to one of claims 1 to 6 when this program is executed by a processor. Non-transitory recording medium readable by a computer on which is recorded a program for implementing the method according to one of claims 1 to 6 when this program is executed by a processor. Automatic reader comprising a real-time provider of sound streams, the sound stream corresponding to a continuous text composed of several groups of words, the sound stream starting from a chosen group of words, also called a second group of words, in the text, based on at least one indication of a first group of words being verbalized by a speaker, the indication coming from a real-time analyzer of captured speech. User terminal comprising a real-time sound stream provider and a sound card, the provider being connected to the sound card and capable of providing a sound stream to the sound card, the sound stream corresponding to a continuous text composed of several groups of words, the sound stream starting from a chosen group of words, also called a second group of words, in the text, based at least on a result indicative of a first group of words being verbalized by a speaker, the result coming from a real-time analyzer of captured speech. User terminal according to claim 10, wherein the sound card is connected to one or more speakers among the following: a speaker of the user terminal, a speaker of a device connected in a local network to the user terminal. A user terminal according to claim 10 or 11, further comprising a text display. The user terminal of claim 12, further comprising a real-time word processing device adapted to highlight a group of words of the text based on the result and to provide the text with the highlighted group of words to the display.