METHOD AND SYSTEM FOR CONVERSING TEXT TO SPEECH
The system addresses voice switching and semantic adaptation in text-to-speech by using semantic analysis and deep learning to produce a natural and expressive audio output with varied voice and prosodic adjustments.
Patent Information
- Application Number
- FR2024008069
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-22
- Publication Date
- 2026-01-23
AI Technical Summary
Existing text-to-speech systems fail to automatically switch between different voices embedded in press articles, adapt voices according to semantics, and handle variations in prosody, style, and emotion, and poorly manage loanwords and foreign language passages.
A system and method utilizing semantic analysis, broad language models, and deep machine learning to detect speakers, phonostyles, and prosodic parameters, enabling automatic voice selection and synthesis that mimics human speech by incorporating intonation, rhythm, and emotional variations.
The system produces a natural and expressive audio output that faithfully reproduces the intended phonostyle and emotion of the text, maintaining listener engagement through varied voice alternations and prosodic adjustments.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: METHOD AND SYSTEM FOR CONVERSING TEXT TO SPEECH STATE OF THE ART
[0001] Systems or software are known for converting text into audio, particularly for reading newspaper articles distributed in paper or digital format.
[0002] However, known systems do not allow for automatic switching between the different voices embedded in press articles, nor for configuring and adapting the voices according to the semantics of the texts. To do this, it is necessary to distinguish the journalist's text (the primary speaker) from one or more quotations referring to one or more other people. In interviews, it is necessary to distinguish the turns of speech between the journalists and the interviewees. Known systems also do not allow for significant variations in prosody, style, and emotion related to the content and structure of the articles, in order to automatically produce a speech synthesis blending several voices and thus avoid a monotonous reading with a single voice. Loanwords and passages in a foreign language within a sentence are also poorly handled by speech synthesis models.
[0003] While much research focuses on extracting quotations, particularly from the print media, there is little data on the various subgenres of press articles (interviews, columns, editorials, etc.), where markers of speaker alternation vary throughout the text and are sometimes subtle. No research exists on characterizing speakers to determine the selection of the best speech synthesis parameters.
[0004] The invention aims to remedy these problems.
[0005] For the remainder of the description, certain terms will be used, each term being associated with a definition. A list is added below.
[0006] DEFINITIONS:
[0007] Semantic Analysis: The study of a language or languages considered from the point of view of meaning; by the analysis of the structures and phenomena of meaning in a language or in language, in particular the analysis of symbolic rules;
[0008] Symbolic rules: detection of quotations by means of typographical clues;
[0009] Broad language model: deep neural network trained on large quantities of unlabeled texts aimed at predicting, from a sequence of words initial the following words. These models are used for a wide variety of tasks (text generation, machine translation, text classification, etc.);
[0010] Machine Learning or deep machine learning is a field of study of artificial intelligence which aims to give machines the ability to "learn" from data, via mathematical models;
[0011] Phonostyle: Style associated with a way of speaking. It is the repertoire of sound styles, characteristic of an individual (young, old, man, woman), of a social group or of a particular circumstance of enunciation (political speech, sermon, radio genre, etc.). SUBJECT OF THE INVENTION
[0012] To this end, and according to a first aspect, the invention refers to a system for converting written data into audio data, in particular speech, comprising: - a semantic analysis module for written data, to process written data in order to extract data that can be transposed into audio, - a module for performing an audio transposition of the extracted written data, so as to obtain so-called transposed audio data, characterized in that it further comprises a module for exchanging and analyzing the extracted written data and / or the transposed audio data, arranged to cooperate via instructions with several Large Language Models trained for solving semantic analysis tasks.
[0013] Thus, the system makes it possible to offer a reading that takes into account several speakers, different languages and prosodic parameters automatically according to the semantics of the texts and to vary the style and emotion through modulations of intonation, accentuation, rhythm of speech, or even by means of sound and musical embellishment.
[0014] According to different embodiments of the conversion system, which may or may not be combinable with each other:
[0015] - the semantic analysis module includes the analysis and characterization of written data;
[0016] - broad language models can be specialized for semantic analysis;
[0017] - broad language models can be different;
[0018] - broad language models can be trained from datasets;
[0019] - large language models can be supervised by a neural network;
[0020] - Broad language models make it possible to detect and characterize speakers and the phonostyle associated with their statements;
[0021] - broad language models can implement artificial intelligence at deep machine learning, preferably generative or non-generative;
[0022] - the system may further include a query or search module terms arranged to query one or more digital libraries of languages and / or general knowledge.
[0023] It is possible to process different article or document formats, in different languages, in order to identify speaker changes, characterize them, and resolve coreference chains in cases where the same person's statements are quoted repeatedly. The integration of this analysis component aims to determine the best voices to apply, as well as their parameters, prior to synthesis.
[0024] - the audio transposition parameterization depends on the semantic analysis (choice of the voice(s) to take into account turn-taking and speaker alternations, setting the enunciation style according to the detected phonostyle).
[0025] - the audio transposition can be achieved at least in part thanks to the marking of the written data extracted according to a speech synthesis language called SSML corresponding to “Speech Synthesis Markup Language”;
[0026] - the audio transposition is achieved using speech synthesis or TTS models corresponding to “Text-to-speech”.
[0027] According to a second aspect, the invention relates to a method for converting written data into audio data by a conversion system according to one or more of the characteristics of the first aspect, the method being characterized by the following steps: - a semantic analysis of the written data, to extract written data to be transposed into audio data extracted for transposition into audio, - a transformation of the extracted written data into audio data, - a sending of the written data and / or audio data and at least one set of instructions to specialized wide language models.
[0028] According to an alternative embodiment, the method of converting written data into audio data by a conversion system according to one or more of the characteristics of the first aspect, the method can be characterized by the following steps: - a semantic analysis of the written data, to identify the different speakers and the phonostyle to be associated with each of their statements (formal and controlled speech, spontaneous speech, etc.).
[0029] For example, the analysis is carried out by sending a series of instructions to different specialized broad language models to solve the different semantic analysis tasks: detection of the style and theme of written data, detection of article genre (column, editorial, analytical article, interview, etc.), detection of quotations, detection of different speakers, characterization of speakers (gender, age, name, function), detection of turns of speech; the objective of semantic analysis being to detect all the parameters likely to influence the parameterization of the audio transposition.
[0030] - a selection and parameterization of the synthesized voices resulting from the analysis semantics, for example male or female voice, voice reflecting the speaker's age, selection of a phonostyle adapted to each utterance (according to a gradient ranging from spontaneous speech, including disfluencies and hesitations, to controlled speech of the type of reading aloud as in the context of a speech), - a transformation of the extracted written data into audiophonic data, in particular the utterances of the different speakers are then transformed into audios,
[0031] - an aggregation of audio segments for the production of the final audio, in in particular the removal of any unwanted noise, removal of any final breaths, harmonization of volumes, musical arrangement.
[0032] The conversion process offers the same advantages as the conversion system. Furthermore, both the system and the conversion process allow prosodic parameters to be varied automatically according to the semantics of the texts, in order to reproduce as faithfully as possible the phonostyle that can be associated with each utterance.
[0033] According to different embodiments of the conversion process, which may or may not be combinable with each other:
[0034] - the method may include a step for detecting speaker alternations (turn-taking detection, extraction of quotations via typographical cues and via instructions sent to specialized broad language models for solving different semantic analysis tasks);
[0035] - the sending step can be repeated in at least one series of several sendings of written data and / or audio data and at least one series of several instructions to broad language models;
[0036] - preferably a series of several written data transmissions corresponding to several parameters or several tasks or written data sent several times depending on or with regard to several parameters and / or tasks;
[0037] - several instructions can be sent to the broad language models;
[0038] - the method may include a step of selecting one or more voices, of preference with phonostyle settings;
[0039] - the method may include at least one query sent to at least one digital library of languages and / or general knowledge;
[0040] - the method may further comprise a generation of at least one illustration or at least one image or at least one video piece of content;
[0041] - generation can be performed based on written data and / or data audio;
[0042] - the method may further include at least one supervisory step, in particular audio quality, of at least one step of the process.
[0043] DESCRIPTION OF THE INVENTION
[0044] Other features and advantages of the invention will become apparent from the detailed description of the invention which will follow with reference to the attached figure:
[0045] [Fig-1] Fig. 1 is a diagram of the process for converting written data into audio data.
[0046] According to a particular embodiment, the following steps are planned:
[0047] Natural language processing components
[0048] Via connectors and a custom content collection module (REST API, FTP, Scraping, ...), the article or press articles, referred to as input articles, are converted into raw written data.
[0049] Next, the raw written data is analyzed according to one or more of the following steps, in this order or in another order:
[0050] - Analysis of subject, tone and style using specialized language models (theme and style detection or extraction);
[0051] - Detection of speakers, quotations and turns of speech; identification of speakers and identifying the gender of speakers using specialized language models,
[0052] - Detection or analysis of the phonostyles specific to each utterance of the speaker(s)
[0053] - Detection of interlingual passages such as Anglicisms, Italianisms, etc.;
[0054] - Identification of the problematic lexicon with a specific pronunciation (detection or lexicon extraction); an enrichment step can be carried out using a monitoring tool;
[0055] - Selection and configuration of multilingual synthesized voices adapted to different speakers and the phonostyle specific to each utterance of the different speakers (translation);
[0056] Then, the extracted written data can be transcribed into audio.
[0057] Multimedia file production components
[0058] The extracted written data can be transposed into audio data via systems, software, and protocols known to a person skilled in the art, for example, the SSML protocol and / or language, and TTS input files, in order to select and configure voices and generate audio segments. One or more voices can be selected.
[0059] One or more audio post-processing steps can be carried out, for example one or more of the following steps: normalization, merging of audio fragments, optimization of pauses, standardization of volume, removal of noise and final breaths, sound and musical dressing, integration of musical effects: background music, jingles, etc.
[0060] Optionally, one or more images can be generated, for example the generation of one or more illustrations related to the article if there are no existing images for it.
[0061] Optionally, subtitles can be generated. According to one variant, a step for calculating the synchronization of text, audio, and video to produce subtitles can be provided.
[0062] Optionally, one or more videos can be generated.
[0063] One or more post-processing steps may be carried out, for example one or more of the following steps: transitions, graphic design, etc.
[0064] At the end of the process, the generated audio, images and subtitles are combined to generate a video file, for example of the MP4 type.
[0065] For the remainder of the description, the step(s) of analysis of the written data are developed, at least one of the steps of the language processing of [Fig.1].
[0066] The proposed analysis combines linguistic rule analysis with deep machine learning, comprising one or more of the following steps:
[0067] - explication and / or detection of the discursive segmentation of texts, for example pacing, transitions and rest periods,
[0068] - speaker detection,
[0069] - detection of the article's theme,
[0070] - detection of turns of speech and quotations,
[0071] - detection of terms likely to be mispronounced (iterative feeding of the thesaurus),
[0072] - detection of interlingual passages, for example anglicisms,
[0073] - explication and / or detection of coreferences, for example of speakers cited in several times.
[0074] Furthermore, the conversion process allows for the management of prosody, using for example the personalization of neural voices, comprising one or both of the following steps:
[0075] - management of journalistic prosody, in particular over-articulation, over-articulation segmentation, melodic hyperactivity, initial accents,
[0076] - management of the prosodic curve, in particular questions, exclamations, emotions, the spontaneity of the speech or, on the contrary, its formal and controlled aspect.
[0077] The various parameters mentioned above are the subject of instructions transmitted to broad language models using one or more techniques that fall under Natural Language Processing and deep learning: specialization of deep learning models, trained from examples to reproduce semantic analysis tasks (classification, detection and characterization of speakers, resolution of coreference chains, etc.).
[0078] In order to best guide speech synthesis and produce an audio transcript capable of capturing and maintaining listeners' attention, significant preliminary work on the texts themselves must be carried out upstream of the processing chain. This preprocessing aims to identify all elements in the texts that are likely to impact the synthesis downstream, either to associate a voice with the text or a portion of text, to parameterize the voice style (emotion, pitch, etc.), or to force a particular pronunciation.For example, detecting the theme of articles can lead to selecting one voice over another (the voice associated with reading a sports article may differ from that associated with reading a political news article); identifying quotations and speaker alternations in texts can lead to interweaving several voices; determining the characteristics of speakers can again lead to choosing one voice or another (feminine or masculine), and so on. It may also be necessary to associate a particular pronunciation with terms or names of foreign origin, as some phonemes involved in their pronunciation do not belong to the general phonological system of the source language. This set of preprocessing steps falls within the field of text analysis, with analysis conducted from different perspectives: syntactic, semantic, discursive, or stylistic.
[0079] Speech synthesis, on the one hand, and text analysis on the other, form the two sides of the broader field of Natural Language Processing (NLP), which is the subject of the conversion process and system.
[0080] Semantic analysis unfolds in three main operations:
[0081] - Operation 1: Polyphonic speech synthesis
[0082] - Operation 2: Audio rendering improvement
[0083] - Operation 3: Improvement of the pronunciation of specific lexicons, words borrowing and interlingual passages.
[0084] These various operations allow for optimal guidance of speech synthesis. They all have one thing in common: they aim to determine the characteristics of the texts that are likely to influence the reading process—either the choice of the selected voice (male or female voice, young or old, voice capable of expressing a particular emotion), or their alternation (changing speakers during long quotations or to reflect turn-taking in a interview), either on the settings of the voices for the reading of a given passage (choice of style, emotion, pace, language, etc.), or on the pronunciation of certain words, or even for determining the pause times to mark in the flow of speech, in order to give it a coherent and non-monotonous rhythm.
[0085] Regarding operation 1: polyphonic speech synthesis via instructions with a broad language model supervised by a neural network makes it possible to automatically determine in a text the changes of speakers leading to alternating the synthetic voices. This operation brings together a set of diverse tasks known as such: extraction of quotations (Pouliquen et al., 2007; Scheible et al., 2016; Papay & Padô, 2019; Papay et al. 2019; Vaucher et al., 2021), extraction of quotation brackets (Bonami & Godard, 2008; Danlos et al., 2010), determination and characterization of speakers (Almeida et al., 2014; Zhang & Liu, 2021), resolution of coreferences to determine if the speaker of a quotation has already been quoted previously in the text, which should lead to reusing the same voice (Almeida et al., 2014; Sukhanter et al., 2020; Ferreira et al., 2020).
[0086] Regarding operation 2: the audio rendering generated via instructions with a broad language model supervised by a neural network makes it possible to imitate human speech in long texts (articles, books), and to introduce richer variations in intonation and rhythm, according to a model specific to journalistic style, such as radio or television news, where prosody varies depending on whether the journalist is announcing the main news headlines or presenting a specific story, regularly playing with rhythmic variations to capture and maintain the listener's attention. The theme or section of an article (sports, politics, science, etc.), the nature of the news (casualty count, surprising discovery), or even the genre of the article (opinion piece, interview, etc.) can also influence the speech rendering.
[0087] Regarding operation 3: the pronunciation of specific lexicons of loanwords and interlingual passages generated via instructions with a broad language model supervised by a neural network can be improved.
[0088] The process and conversion system makes it possible to produce a speech synthesis capable of exploiting the richest and most varied expressive acoustic palette possible (variations in intonation, emotions, rhythms, alternations and changes of voice, introduction of breaths, soundscapes, musical arrangements), in order to produce audio content capable of capturing and maintaining the attention of listeners over time.
[0089] They allow prosodic parameters to be varied automatically according to the semantics of the texts, and thus produce in the most automatic way possible an audio format that is pleasant, natural and expressive.
Claims
Demands
1. A system for converting written data into audio data comprising: - a semantic analysis module for written data, to process written data in order to extract data to be transposed into audio, - a module for performing an audio transposition of the extracted written data, characterized in that it further comprises a module for exchanging and analyzing the extracted written data and / or the transposed audio data, arranged to cooperate via instructions with several large language models trained for solving semantic analysis tasks.
2. A system according to the preceding claim in which the broad language models implement deep machine learning artificial intelligence, preferably generative or non-generative.
3. System according to claim 1 or 2 further comprising a query or term search module arranged to query one or more digital libraries of languages and / or general knowledge.
4. A method for converting written data into audio data by a conversion system according to one of the preceding claims, the method being characterized by the following steps: - a semantic analysis of the written data, to extract written data to be transposed into audio extracted for transposition into audio, - a transformation of the extracted written data into audio data, - a sending of the written data and / or audio data and at least one series of instructions to specialized wide language models.
5. A method according to the preceding claim, wherein the sending step is repeated in at least one series of several sendings of written data and / or audio data and at least one series of several instructions to the broad language models.
6. A method according to claim 4 or 5, wherein several instructions are sent to the broad language model.
7. A method according to any one of claims 4 to 6, wherein at least one query is sent to at least one digital library of languages and / or general knowledge.
8. A method according to any one of claims 4 to 7, further comprising a generation of at least one illustration or at least one image or at least one video content.
9. Method according to the preceding claim, wherein the generation is carried out based on written data and / or audio data.
10. A method according to any one of claims 4 to 9, further comprising at least one supervisory step of at least one step of the method.
Citation Information
Patent Citations
Intelligent voice service management system and method driven by natural language understanding
CN118194875A
Systems for controllable summarization of content
US12008332B1
Using large language model(s) in generating automated assistant response(s
US20230074406A1
Real-time system for spoken natural stylistic conversations with large language models
WO2024112393A1