Speech synthesis process, associated device and vehicle

The method converts synthetic speech to simulated natural speech using AI, addressing user discomfort and cost issues in SIVEs by enhancing user experience without system replacement.

FR3162548A1Pending Publication Date: 2025-11-28SNCF VOYAGEURS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2024005404
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-27
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing on-board passenger information systems (SIVEs) in rail transport use artificial voices, which induce feelings of loneliness and insecurity among users, and replacing them with natural voices is costly and complex.

Method used

A method involving conversion of synthetic speech sequences into simulated natural speech sequences using an AI model trained on voice pairs, extracting vocal features, and parameterizing a speech synthesis engine to simulate natural voices without replacing the existing system.

Benefits of technology

Simulates natural voice characteristics to enhance user comfort while maintaining system simplicity and cost-effectiveness, allowing natural voice announcements without replacing existing SIVEs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a speech synthesis method comprising the steps: converting at least one synthetic speech sequence into a corresponding simulated natural speech sequence; extracting a set of speech features from each generated simulated natural speech sequence; and configuring a speech synthesis engine based on the extracted set of speech features. Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Speech synthesis method, associated device and vehicle. Technical field

[0001] The present invention relates to a computer-implemented speech synthesis method.

[0002] The invention also relates to a computer program, a device implementing such a method and a vehicle incorporating such a device.

[0003] The invention applies to the field of information dissemination systems for users, in particular for users of a transport network. Prior art

[0004] It is known to use information dissemination systems for users of a transport network. In particular, in the field of rail transport, it is known to use systems called "on-board passenger information systems" (or IPS).

[0005] Such SIVEs are commonly carried on board the railway vehicle, and include a computer and display screens and / or speakers controlled by the computer.

[0006] Such SIVEs are, in particular, configured to produce, for users, contextual information of a commercial mission (defined as the journey of a railway vehicle from a departure station A to a terminus station B), in particular a commercial mission in progress.

[0007] In particular, the computer is configured to use an embedded database, as well as GPS (for "Global Positioning System") and odometric data, in order to generate the information to be disseminated and determine its time of dissemination.

[0008] Such a database is, in particular, configured to store, for each commercial mission, information such as a mission code, a list of stations served and their respective GPS coordinates, as well as a list of distances between successive stations.

[0009] Typically, during a commercial mission, a SIVE first loads the data for the upcoming commercial mission from the database (for example, the list of stations to be served and their characteristics). Such a commercial mission is, in particular, determined based on a mission code entered by the train crew.

[0010] The SIVE also stores text sequences containing, in particular, announcements corresponding to different operating contexts that the rail vehicle may encounter. In this case, the text sequences (including station names) are transformed into synthetic speech sequences by a speech synthesis engine configured with an artificial voice.

[0011] The SIVE also stores sound announcements generated from text sequences by a speech synthesis engine, also configured with an artificial voice.

[0012] Then, iteratively during the commercial mission, the SIVE determines the name of the current station and the next station in order to make a visual and audio announcement to the passengers.

[0013] In addition, on occasion, in specific operating contexts, the driver may trigger visual and / or audible announcements known as "contingency announcements", by means of dedicated controls located in the driver's cab.

[0014] For example, the SIVE is capable of broadcasting a visual and audible announcement when the train driver presses a specific button indicating that the train is forced to stop on the track. Such an announcement is, for example: "Our train is stopped on the track. For your safety, do not attempt to open the doors."

[0015] However, such a system does not give complete satisfaction.

[0016] Indeed, it has been observed that the use of an artificial voice reinforces feelings of loneliness and insecurity among users. In fact, the physical absence of cabin crew combined with a voice unrelated to a human induces a feeling of abandonment in the traveler.

[0017] However, the delivery of voice announcements cannot be entrusted to the cabin crew, as this would result in an undesirable additional workload. This constraint is even greater in a so-called "single-agent operation" (SA) mode, in which no sales representative accompanies the flight attendant.

[0018] Furthermore, for reasons of cost and complexity, it cannot be envisaged to replace the existing SIVEs of a fleet of railway vehicles in favor of SIVEs capable of performing speech synthesis with a natural voice.

[0019] One object of the present invention is to remedy at least one of the drawbacks of the prior art.

[0020] Another objective of the invention is to propose a speech synthesis method whose deployment is not very complex and not very expensive. Description of the invention

[0021] To this end, the invention relates to a method of the aforementioned type, comprising the following steps: • conversion of at least one synthetic speech sequence into a corresponding simulated natural speech sequence, the conversion including the implementation of an artificial intelligence model previously trained on the basis of a training dataset comprising at least one pair of voice sequences, each pair of voice sequences comprising a synthetic voice sequence and a natural voice sequence corresponding to the same text sequence, each synthetic voice sequence forming an input to the artificial intelligence model, the respective natural voice sequence forming an expected output of the artificial intelligence model, the conversion including the provision, as input to the trained artificial intelligence model, of each synthetic speech sequence, and, for each synthetic speech sequence, the corresponding output of the trained artificial intelligence model forming the simulated natural speech sequence; • extraction of a set of vocal features from each generated simulated natural vocal sequence; and • Parameterization of a speech synthesis engine based on the extracted set of vocal characteristics.

[0022] Indeed, thanks to the conversion step, a set of vocal sequences (called "simulated natural vocal sequences") with a voice comparable to a natural voice is simulated.

[0023] In this way, vocal characteristics representative of natural voices can be extracted from the simulated natural vocal sequences, and used to parameterize the speech synthesis engine.

[0024] This is all the more advantageous when the speech synthesis engine to be configured is the native speech synthesis engine of a given vehicle's information system: with such configuration (which is easy to implement), the replacement of the information system is no longer required, and voice announcements with a natural voice can be broadcast to vehicle users.

[0025] Advantageously, the process according to the invention has one or more of the following characteristics, taken individually or in any technically feasible combination:

[0026] the method further comprises a synthesis of a voice announcement from a textual sequence, the synthesis comprising the provision of the textual sequence as input to the parameterized speech synthesis engine, a corresponding output of the parameterized speech synthesis engine forming the synthesized voice announcement;

[0027] the synthesis further includes the detection of a current operational situation, the text sequence being selected according to the detected current operational situation;

[0028] the detection of the current operational situation includes geolocation of a target vehicle;

[0029] The method further comprises broadcasting the synthesized voice announcement in and / or near a target vehicle.

[0030] According to another aspect of the invention, a computer program is proposed comprising executable instructions which, when executed by computer, implement the steps of the process as defined above.

[0031] The computer program can be in any computer language, such as for example in machine language, in C, C++, JAVA, Python, etc.

[0032] According to another aspect of the invention, an information device is proposed comprising: • a conversion module configured to perform a conversion of at least one synthetic speech sequence into a corresponding simulated natural speech sequence, the conversion including the implementation of an artificial intelligence model previously trained on the basis of a training dataset comprising at least one pair of voice sequences, each pair of voice sequences comprising a synthetic voice sequence and a natural voice sequence corresponding to the same text sequence, each synthetic voice sequence forming an input to the artificial intelligence model, the respective natural voice sequence forming an expected output of the artificial intelligence model, the conversion including the provision, as input to the trained artificial intelligence model, of each synthetic speech sequence, and, for each synthetic speech sequence, the corresponding output of the trained artificial intelligence model forming the simulated natural speech sequence; • an extraction module configured to extract a set of vocal features from each generated simulated natural vocal sequence; and • a speech synthesis module configured to set up a speech synthesis engine based on the extracted set of voice characteristics, to obtain a parameterized speech synthesis engine.

[0033] Advantageously, the information device according to the invention has the following characteristic:

[0034] The speech synthesis module is further configured to perform the synthesis of a speech announcement from a text sequence, the synthesis comprising the provision of the text sequence as input to the parameterized speech synthesis engine, a corresponding output of the parameterized speech synthesis engine forming the synthesized speech announcement.

[0035] The device according to the invention can be any type of device such as a server, a computer, a tablet, a calculator, a processor, a computer chip, programmed to implement the method according to the invention, for example by executing the computer program according to the invention.

[0036] According to another aspect of the invention, a vehicle is proposed, in particular a railway transport vehicle, comprising an information device as defined above and at least one sound diffusion device connected to the information device to broadcast each synthesized voice sequence in the vehicle and / or in the vicinity of the vehicle. Brief description of the figures

[0037] The invention will be better understood upon reading the following description, given solely by way of non-limiting example and made with reference to the accompanying drawings in which:

[0038] [Fig. 1] is a schematic representation of a vehicle according to the invention; and

[0039] [Fig. 2] is a flowchart of a speech synthesis process implemented by a vehicle information device of the [Fig.l].

[0040] It is understood that the embodiments described below are in no way limiting. In particular, variants of the invention may be conceived comprising only a selection of the features described below, isolated from the other features described, if this selection of features is sufficient to confer a technical advantage or to differentiate the invention from the prior art. This selection includes at least one preferably functional feature without structural details, or with only a portion of the structural details if this portion alone is sufficient to confer a technical advantage or to differentiate the invention from the prior art.

[0041] In particular, all the variants and embodiments described are combinable with each other if there is no technical obstacle to this combination.

[0042] In the figures and in the rest of the description, elements common to several figures retain the same reference. Detailed description

[0043] A vehicle 2 according to the invention is illustrated by [Fig. 1].

[0044] Vehicle 2 is, in particular, a passenger transport vehicle. Preferably, vehicle 2 is a rail transport vehicle such as a train, in particular a rail transport vehicle intended for the transport of passengers.

[0045] The vehicle 2 includes at least one sound diffusion device 4 and one information device 6.

[0046] Preferably, for their communication, the information device 6 and each sound diffusion device 4 are connected to a corresponding port of a network switch (also called "switch") of a vehicle communication network 2, such as a wired Ethernet IP communication network with a ring topology.

[0047] Preferably, the vehicle 2 also includes a geolocation device 8 and / or an input interface 10. In this case, the information device 6 is also capable of communicating with the geolocation device 8 and the input interface 10, in particular via the communication network of the vehicle 2.

[0048] Sound diffusion device 4

[0049] Each sound diffusion device 4 includes at least one loudspeaker configured to diffuse an acoustic wave representative of a signal received at the input of said sound diffusion device 4.

[0050] Preferably, the sound diffusion device 4 also includes a processing unit configured, for example, to shape the received signal, in particular by filtering and / or amplification, before its diffusion by the corresponding loudspeaker or loudspeakers.

[0051] Geolocation device 8

[0052] The geolocation device 8 is configured to store information relating to each journey that the vehicle 2 is likely to make.

[0053] In particular, for each journey (also called a "commercial mission"), the geolocation device 8 is configured to store a list of corresponding stops (i.e., stations, in the case of a rail vehicle) served. Preferably, the geolocation device 8 is also configured to store the distances between successive stops and / or the coordinates of each stop, including the GPS coordinates of each stop.

[0054] In addition, the geolocation device 8 is configured to determine, over time, a current position of the vehicle 2. Such a position includes, for example, coordinates of the vehicle 2 in a reference frame and / or a distance traveled by the vehicle 2 from a predetermined position, for example a previous stop.

[0055] For example, the geolocation device 8 is configured to receive signals from one or more beacons (for example, GPS signals), and to determine the current coordinates of the vehicle 2 from the received signals.

[0056] In this case, the geolocation device 8 is configured to compare the coordinates of vehicle 2 with the stored coordinates of each stop. In addition, the geolocation device 8 is configured to send an alert to the information device 6 when the current position of vehicle 2 meets a predetermined condition, for example, when the distance between the current position of vehicle 2 and a stop on the current commercial route becomes less than or equal to a predetermined threshold (generally a few hundred meters, in order to anticipate the announcement made to passengers).

[0057] Alternatively, or in addition, the geolocation device 8 includes at least one odometer configured to calculate a distance travelled by the vehicle 2 from a predetermined position.

[0058] In this case, the geolocation device 8 is configured to send an alert to the information device 6 when the distance traveled since the last stop meets a predetermined condition, for example when the difference between the distance traveled since the last stop and the distance memorized between said stop and the next arrival for the current commercial mission becomes less than or equal to the predetermined threshold.

[0059] Input interface 10

[0060] The input interface 10 is a human / machine interface adapted to allow an operator to input instructions and / or indications relating to an operational situation of the vehicle 2.

[0061] Preferably, the input interface 10 includes at least one button and / or at least one touch surface which has been previously associated with a predetermined operational situation.

[0062] In this case, the input interface 10 is configured to send, to the information device 6, an alert representative of the operational situation associated with the button (respectively, with the touch surface) when it is used by the operator.

[0063] Preferably, the input interface 10 includes a keyboard allowing the operator to enter a text sequence.

[0064] In this case, the input interface 10 is configured to send, to the information device 6, the text sequence entered by the operator.

[0065] Information device 6

[0066] The information device 6 is intended for the dissemination of information to passengers of vehicle 2, or to persons in the vicinity of vehicle 2. In particular, the information device 6 is configured to control the dissemination of sound information by the sound dissemination device 4.

[0067] Preferably, the information device 6 has a compact form ensuring a small footprint in the vehicle 2, and is, in particular, compliant with railway standards for electromagnetic compatibility, fire and smoke resistance, and shock and vibration resistance.

[0068] The information device 6 includes a memory 12, a voice conversion module 14, a voice feature extraction module 16 (called "extraction module") and a speech synthesis module 18.

[0069] Memory 12

[0070] Memory 12 is configured to store information relating to each journey that vehicle 2 is capable of making. In particular, for each journey, memory 12 is configured to store the list of corresponding stops served.

[0071] In addition, memory 12 is configured to store at least one text sequence. In this case, each text sequence is associated with a corresponding operational situation.

[0072] For example, at least one stored text sequence is a notification of arrival of vehicle 2 at a stop, or a notification of departure of vehicle 2 from a stop.

[0073] According to another example, at least one memorized text sequence corresponds to a statement of the stops served by vehicle 2 during a given journey.

[0074] According to another example, at least one stored text sequence is a warning message, such as: • a warning message to inform passengers of an ongoing event; • a warning message inviting users near or on board vehicle 2 to move away from the doors of said vehicle 2; • a warning message to alert passengers of an exceptional stop of vehicle 2 between two stops, etc.

[0075] In addition, memory 12 is configured to store at least one synthetic speech sequence.

[0076] Each synthetic speech sequence has been previously generated using a speech synthesis engine, and is associated with a corresponding operating context (i.e., an operational situation).

[0077] Voice conversion module 14

[0078] The speech conversion module 14 is configured to convert each synthetic speech sequence stored in memory 12 into a corresponding simulated natural speech sequence.

[0079] In order to carry out such a conversion, the voice conversion module 14 is configured to implement a pre-trained artificial intelligence model.

[0080] In particular, to carry out such a conversion, the speech conversion module 14 is configured to provide, as input to the trained artificial intelligence model, each synthetic speech sequence stored in memory 12. In this case, for each synthetic speech sequence received at its input, the corresponding output of the trained artificial intelligence model forms the corresponding simulated natural speech sequence.

[0081] The artificial intelligence model was previously trained on a training dataset comprising at least one pair of speech sequences. In particular, each pair of speech sequences includes a synthetic speech sequence and a natural speech sequence, both corresponding to the same text sequence. In this case, during the training of the artificial intelligence model, each synthetic speech sequence forms an input to the artificial intelligence model, while the respective natural speech sequence forms an expected output of the artificial intelligence model.

[0082] Preferably, the artificial intelligence model is a sequence-to-sequence model featuring, in particular, an encoder-decoder architecture.

[0083] Extraction Module 16

[0084] The extraction module 16 is configured to extract a set of voice features from each simulated natural voice sequence generated by the voice conversion module 14.

[0085] Preferably, to extract the set of voice features, the extraction module 16 is configured to implement an artificial intelligence model with a so-called GE2E architecture (from the English "Generalized End-to-End loss", or generalized end-to-end cost).

[0086] Such an architecture is described by Li Wan et al. in the digital preprint "Generalized End-to-End Loss for Speaker Verification", referenced arXiv:1710.10467.

[0087] Such a set of vocal characteristics includes values ​​that can be likened to a vocal signature of a corresponding speaker.

[0088] For example, the set of vocal features provided as output by the extraction module 16 has a value matrix structure. In this case, the values ​​taken by the matrix coefficients represent a decomposition of the speaker's vocal signature in a given tensor space. Such a space is such that, for two speakers with similar voices in terms of sound, the distance between the vectors representing their respective vocal signatures is small.

[0089] Preferably, the extraction module 16 is configured to store, in a predetermined location in memory 12, the extracted set of voice features, in the form of a binary file with a predefined name.

[0090] Speech synthesis module 18

[0091] The speech synthesis module 18 is configured to implement a speech synthesis engine in order to synthesize, from a given text sequence, a corresponding speech announcement (called a "synthesized speech announcement").

[0092] Preferably, the speech synthesis engine is a sequence-to-sequence artificial intelligence model featuring, in particular, an encoder-decoder architecture.

[0093] More specifically, the speech synthesis module 18 is configured to, initially, perform a parameterization of the speech synthesis engine according to the set of voice characteristics extracted by the extraction module 16, in order to obtain a parameterized speech synthesis engine.

[0094] For example, to achieve such a setting, the speech synthesis module 18 is configured to replace all or part of the speech synthesis engine parameters with the set of speech features previously extracted by the extraction module 16 (or with the result of a predetermined function applied to said set of speech features).

[0095] In particular, the speech synthesis module 18 is configured to search, at the predetermined location of memory 12 mentioned above, for a binary file having the predetermined name.

[0096] In addition, the speech synthesis module 18 is preferably configured to detect the current operational situation.

[0097] Preferably, the speech synthesis module 18 is configured to detect the current operational situation from the current position of the vehicle 2 determined by the geolocation device 8, and / or alerts issued by the geolocation device 8.

[0098] Alternatively, or in addition, the speech synthesis module 18 is configured to detect the current operational situation from an input performed by the operator on the input interface 10, such as the activation of a button specific to a given operational situation.

[0099] In addition, the speech synthesis module 18 is configured to load a text sequence for which a corresponding speech announcement is to be synthesized.

[0100] Where the speech synthesis module 18 is configured to detect the current operational situation, the speech synthesis module 18 is advantageously configured to load, from memory 12, the text sequence corresponding to said detected current operational situation, in order to synthesize the voice announcement. Such a feature is advantageous insofar as it ensures consistency between the current context and the voice announcement issued to users.

[0101] Alternatively, or in addition, where the input interface 10 is adapted to allow the operator to enter a text sequence, the speech synthesis module 18 is configured to load the text sequence entered by the operator as a text sequence to be used for the synthesis of a voice announcement.

[0102] In addition, the speech synthesis module 18 is configured to provide the loaded text sequence as input to the parameterized speech synthesis engine. In this case, the output of the parameterized speech synthesis engine forms the synthesized speech announcement to be broadcast.

[0103] In addition, the speech synthesis module 18 is configured to transmit the synthesized speech announcement to all or part of the sound broadcasting devices 4 for broadcasting said synthesized speech announcement in the vehicle 2 and / or in the vicinity of the vehicle 2.

[0104] Operation

[0105] The operation of the information device 6 will now be described with reference to the figures.

[0106] During an initial memorization step, information relating to each journey that the vehicle 2 is likely to make is stored in memory 12.

[0107] In addition, at least one text sequence is stored in memory 12. At least one synthetic speech sequence is also stored in memory 12.

[0108] In addition, information relating to each journey that vehicle 2 is likely to make is also stored in the geolocation device 8.

[0109] Then, during an initialization step, an operator (for example, a driver) initializes the information device 6, in particular by entering a code for the upcoming journey.

[0110] Then, the speech synthesis process 20 according to the invention is implemented.

[0111] More specifically, during a conversion step 22, the speech conversion module 14 converts each synthetic speech sequence stored in memory 12 into a corresponding simulated natural speech sequence.

[0112] Then, during an extraction step 24, the extraction module 16 extracts a set of vocal features from each simulated natural vocal sequence previously generated by the vocal conversion module 14 during the conversion step 22.

[0113] Then, during a parameterization step 26, the speech synthesis module 18 performs a parameterization of the speech synthesis engine according to the set of voice characteristics extracted by the extraction module 16.

[0114] This results in a parameterized speech synthesis engine.

[0115] Preferably, during a synthesis step 28, when the speech synthesis module 18 detects a given current operational situation requiring the broadcast of a voice announcement, the speech synthesis module 18 loads a text sequence corresponding to said detected current operational situation.

[0116] For example, the speech synthesis module 18 detects the current operational situation based on alerts issued by the geolocation device 8 or alerts issued by the input interface 10.

[0117] Alternatively, or in addition, the speech synthesis module 18 detects the current operational situation as the entry of a text sequence by an operator, via the input interface 10. In this case, the text sequence loaded is the text sequence entered by the operator.

[0118] Then, the speech synthesis module 18 provides the text sequence loaded as input to the parameterized speech synthesis engine. In this case, the output of the parameterized speech synthesis engine forms the synthesized speech announcement to be broadcast.

[0119] Then, during a broadcasting step 30, the speech synthesis module 18 transmits the synthesized speech announcement to all or part of the sound broadcasting devices 4 for the broadcasting of said synthesized speech announcement in the vehicle 2 and / or in the vicinity of the vehicle 2.

[0120] Of course, the invention is not limited to the examples just described.

Claims

Demands

1. A computer-implemented speech synthesis method (20) comprising the following steps: • conversion (22) of at least one synthetic speech sequence into a corresponding simulated natural speech sequence, the conversion comprising the implementation of an artificial intelligence model previously trained on the basis of a training dataset comprising at least one pair of speech sequences, each pair of speech sequences comprising a synthetic speech sequence and a natural speech sequence corresponding to the same text sequence, each synthetic speech sequence forming an input to the artificial intelligence model, the respective natural speech sequence forming an expected output of the artificial intelligence model, the conversion comprising providing, as input to the trained artificial intelligence model, each synthetic speech sequence, and, for each synthetic speech sequence,the corresponding output of the trained artificial intelligence model forming the simulated natural speech sequence; • extraction (24) of a set of speech features from each generated simulated natural speech sequence; and • parameterization (26) of a speech synthesis engine based on the extracted set of speech features.

2. A method according to claim 1, further comprising a synthesis (28) of a voice announcement from a text sequence, the synthesis comprising supplying the text sequence as input to the parameterized speech synthesis engine, a corresponding output of the parameterized speech synthesis engine forming the synthesized voice announcement.

3. A method according to claim 2, wherein the synthesis (28) further comprises a detection of a current operational situation, the text sequence being selected according to the detected current operational situation.

4. Method according to claim 3, wherein the detection of the current operational situation includes a geolocation of a target vehicle (2).

5. A method according to any one of claims 2 to 4, further comprising broadcasting (30) the synthesized voice announcement in and / or near a target vehicle (2).

6. A computer program comprising executable instructions which, when executed by a computer, implement the steps of the process according to any one of claims 1 to

7. J. Information device (6) comprising: • a conversion module (14) configured to perform a conversion of at least one synthetic speech sequence into a corresponding simulated natural speech sequence, the conversion comprising the implementation of an artificial intelligence model previously trained on the basis of a training dataset comprising at least one pair of speech sequences, each pair of speech sequences comprising a synthetic speech sequence and a natural speech sequence corresponding to the same text sequence, each synthetic speech sequence forming an input to the artificial intelligence model, the respective natural speech sequence forming an expected output of the artificial intelligence model, the conversion comprising providing, as input to the trained artificial intelligence model, each synthetic speech sequence, and, for each synthetic speech sequence,the corresponding output of the trained artificial intelligence model forming the simulated natural speech sequence; • an extraction module (16) configured to extract a set of speech features from each generated simulated natural speech sequence; and • a speech synthesis module (18) configured to parameterize a speech synthesis engine based on the extracted set of speech features, to obtain a parameterized speech synthesis engine.

8. Information device (6) according to claim 7, wherein the speech synthesis module (16) is further configured to perform the synthesis of a speech announcement from a text sequence, the synthesis comprising the provision of the text sequence as input to the parameterized speech synthesis engine, a corresponding output of the parameterized speech synthesis engine forming the synthesized speech announcement.

9. Vehicle (2), in particular a railway transport vehicle, comprising an information device (6) according to claim 8 and at least one sound diffusion device (4) connected to the information device (6) for broadcasting each synthesized voice sequence in the vehicle (2) and / or in the vicinity of the vehicle (2).