Announcement synthesis process, associated device and vehicle

By parameterizing speech synthesis engines with individual vocal characteristics and generating synchronized facial animations, the method addresses user discomfort in SIVEs, providing a human-like experience without system replacement or additional workload.

FR3162547A1Pending Publication Date: 2025-11-28SNCF VOYAGEURS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
FR2024005410
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-27
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing on-board passenger information systems (SIVEs) in rail transport induce feelings of loneliness and insecurity among users due to the use of artificial voices, which cannot be replaced easily due to cost and complexity, and entrusting announcements to cabin crew adds undesirable workload, especially in single-agent operations.

Method used

A method that parameterizes a speech synthesis engine with vocal characteristics of a target individual, synthesizes voice announcements, and generates an animated facial visual to create a human-like experience, reducing discomfort by correlating voice and facial attributes.

Benefits of technology

The method provides voice announcements that mimic human speakers, enhancing user comfort by creating a consistent and familiar experience through synchronized voice and facial animations, without requiring system replacement or additional workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method for synthesizing announcements, comprising the following steps: configuring a speech synthesis engine based on a set of vocal characteristics associated with a target individual; synthesizing a synthesized voice announcement from a text sequence, the synthesis comprising providing the text sequence as input to the configured speech synthesis engine, and a corresponding output from the configured speech synthesis engine forming the synthesized voice announcement; and generating an animated facial visual, comprising providing, as input to a facial animation model, the synthesized voice announcement and a portrait of the target individual. Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Announcement synthesis method, associated device and vehicle. Technical field

[0001] The present invention relates to a computer-implemented method for synthesizing advertisements.

[0002] The invention also relates to a computer program, a device implementing such a method and a vehicle incorporating such a device.

[0003] The invention applies to the field of information dissemination systems for users, in particular for users of a transport network. Prior art

[0004] It is known to use information dissemination systems for users of a transport network. In particular, in the field of rail transport, it is known to use systems called "on-board passenger information systems" (or IPS).

[0005] Such SIVEs are commonly carried on board the railway vehicle, and include a computer and display screens and / or speakers controlled by the computer.

[0006] Such SIVEs are, in particular, configured to produce, for users, contextual information of a commercial mission (defined as the journey of a railway vehicle from a departure station A to a terminus station B), in particular a commercial mission in progress.

[0007] In particular, the computer is configured to use an embedded database, as well as GPS (for "Global Positioning System") and / or odometric data, in order to generate the information to be disseminated and determine its time of dissemination.

[0008] Such a database is, in particular, configured to store, for each commercial mission, information such as a mission code, a list of stations served and their respective GPS coordinates, as well as a list of distances between successive stations.

[0009] Typically, during a commercial mission, a SIVE first loads the data for the upcoming commercial mission from the database (for example, the list of stations to be served and their characteristics). Such a commercial mission is, in particular, determined based on a mission code entered by the train crew.

[0010] The SIVE also stores text sequences containing, in particular, announcements corresponding to different operating contexts that the rail vehicle may encounter. In this case, the text sequences (including station names) are transformed into synthetic speech sequences by a speech synthesis engine configured with an artificial voice.

[0011] The SIVE also stores sound announcements generated from text sequences by a speech synthesis engine, also configured with an artificial voice.

[0012] Then, iteratively during the commercial mission, the SIVE determines the name of the current station and the next station in order to make a visual and audio announcement to the passengers.

[0013] In addition, on occasion, in specific operating contexts, the driver may trigger visual and / or audible announcements known as "contingency announcements", by means of dedicated controls located in the driver's cab.

[0014] For example, the SIVE is capable of broadcasting a visual and audible announcement when the train driver presses a specific button indicating that the train is forced to stop on the track. Such an announcement is, for example: "Our train is stopped on the track. For your safety, do not attempt to open the doors."

[0015] However, such a system does not give complete satisfaction.

[0016] Indeed, it has been observed that the use of an artificial voice reinforces feelings of loneliness and insecurity among users. In fact, the physical absence of cabin crew, combined with a voice unrelated to a human, induces a feeling of abandonment in the traveler.

[0017] However, the delivery of voice announcements cannot be entrusted to the cabin crew, as this would result in an undesirable additional workload. This constraint is even greater in a so-called "single-agent operation" (SA) mode, in which no sales representative accompanies the flight attendant.

[0018] Furthermore, for reasons of cost and complexity, it cannot be envisaged to replace the existing SIVEs of a fleet of railway vehicles in favor of SIVEs capable of performing speech synthesis with a natural voice.

[0019] One object of the present invention is to remedy at least one of the drawbacks of the prior art.

[0020] Another object of the invention is to propose a method of announcement synthesis whose deployment is not very complex and not very expensive.

[0021] Another object of the invention is to propose a method of disseminating information that does not diminish the feeling of unease among users. Description of the invention

[0022] To this end, the invention relates to a method of the aforementioned type, comprising the steps: • parameterization of a speech synthesis engine according to a set of vocal characteristics associated with a target individual; • synthesis of a synthesized voice announcement from a text sequence, the synthesis comprising the provision of the text sequence as input to the parameterized speech synthesis engine, a corresponding output from the parameterized speech synthesis engine forming the synthesized voice announcement; • generation of an animated facial visual, including the provision, as input, of a facial animation model, the synthesized voice announcement and a portrait of the target individual.

[0023] Indeed, thanks to the parameterization step, synthesized vocal sequences similar, in terms of sound, to those that would have been spoken by a human speaker (here, the target individual) are produced.

[0024] In particular, by constraining the speech synthesis model with a set of vocal characteristics of a target individual who is a crew member of a vehicle, the synthesized voice announcements are close, in terms of sound, to those that would have been made by said crew member, and that the vehicle's passengers would expect to hear, especially after having seen said crew member.

[0025] In addition, such a method is advantageous, in particular when the speech synthesis engine is the native speech synthesis engine of a given vehicle's information system: by such a parameterization (which is easy to implement), the replacement of the information system is no longer required, and voice announcements with a natural voice can be broadcast to vehicle users.

[0026] Moreover, by using a facial animation model based on a portrait of the target individual, an animated facial image is generated whose features evolve over time in accordance with the given synthesized voice sequence. Such a human-like facial image, combined with the synthesized voice announcement, reduces the user's discomfort. In particular, the user perceives a consistency between the sound characteristics of the synthesized voice announcements and the physical attributes of the animated facial image. Indeed, there is a correlation between a speaker's voice and their physical attributes (such as their sex, age, and ethnicity).

[0027] Advantageously, the process according to the invention has one or more of the following characteristics, taken individually or in any technically feasible combination:

[0028] the generation of the animated facial visual further includes the provision, as input to the facial animation model, of an emotional marker of the textual sequence associated with the synthesized voice announcement;

[0029] The method includes, prior to the generation of the animated facial visual, a provision of the textual sequence as input to a previously trained language model, an output of the language model including the emotional marker corresponding to the textual sequence;

[0030] The process includes an extraction of vocal characteristics from at least one natural vocal sequence previously uttered by the target individual;

[0031] the synthesis further includes the detection of a current operational situation, the text sequence being selected according to the detected current operational situation;

[0032] the detection of the current operational situation includes geolocation of a target vehicle;

[0033] The method further comprises a transmission of the synthesized voice announcement and the corresponding generated animated facial visual for their synchronous broadcast in and / or near a target vehicle.

[0034] According to another aspect of the invention, a computer program is proposed comprising executable instructions which, when executed by computer, implement the steps of the process as defined above.

[0035] The computer program can be in any computer language, such as for example in machine language, in C, C++, JAVA, Python, etc.

[0036] According to another aspect of the invention, an information device is proposed configured to: • configure a speech synthesis engine based on a set of vocal characteristics associated with a target individual; • synthesize a synthesized voice announcement from a text sequence, the synthesis comprising the provision of the text sequence as input to the parameterized speech synthesis engine, a corresponding output from the parameterized speech synthesis engine forming the synthesized voice announcement; • generate an animated facial visual, including the provision, as input, of a facial animation model, the synthesized voice announcement and a portrait of the target individual.

[0037] The device according to the invention can be any type of device such as a server, a computer, a tablet, a calculator, a processor, a computer chip, programmed to implement the method according to the invention, for example by executing the computer program according to the invention.

[0038] According to another aspect of the invention, a vehicle, in particular a railway transport vehicle, is proposed, comprising an information device as defined above, at least one sound diffusion device connected to the information device to broadcast each synthesized voice announcement in the vehicle and / or in the vicinity of the vehicle, and at least one display support to display each animated facial visual, the information device being configured to control each sound diffusion device and each display support so as to broadcast each synthesized voice announcement and the corresponding animated facial visual synchronously.

[0039] Advantageously, the vehicle according to the invention has the following characteristic:

[0040] The vehicle further comprises: • at least one camera connected to the information device, each camera being configured to acquire an image of a predetermined volume of the vehicle, and to transmit the acquired image to the information device; and / or • at least one microphone connected to the information device, each microphone being adapted to acquire a natural speech sequence, and to transmit the acquired natural speech sequence to the information device, the information device further comprising an extraction module configured to extract the set of speech features from the acquired natural speech sequence. Brief description of the figures

[0041] The invention will be better understood upon reading the following description, given solely by way of non-limiting example and made with reference to the accompanying drawings in which:

[0042] [Fig. 1] is a schematic representation of a vehicle according to the invention; and

[0043] [Fig.2] is a flowchart of an announcement synthesis process implemented by a vehicle information device of the [Fig.1].

[0044] It is understood that the embodiments described below are in no way limiting. In particular, variants of the invention may be conceived comprising only a selection of the features described below, isolated from the other features described, if this selection of features is sufficient to confer a technical advantage or to differentiate the invention from the prior art. This selection includes at least one preferably functional feature without structural details, or with only a portion of the structural details if this part alone is sufficient to confer a technical advantage or to differentiate the invention from the prior art.

[0045] In particular, all the variants and embodiments described are combinable with each other if there is no technical obstacle to this combination.

[0046] In the figures and in the rest of the description, elements common to several figures retain the same reference. Detailed description

[0047] A vehicle 2 according to the invention is illustrated by [Fig. 1].

[0048] Vehicle 2 is, in particular, a passenger transport vehicle. Preferably, vehicle 2 is a rail transport vehicle such as a train, in particular a rail transport vehicle intended for the transport of passengers.

[0049] The vehicle 2 includes at least one sound diffusion device 4, at least one display support 5 and one information device 6.

[0050] Preferably, for their communication, the information device 6 and each sound diffusion device 4 are connected to a corresponding port of a network switch (also called a "switch") of a vehicle communication network 2, such as a wired Ethernet IP communication network with a ring topology.

[0051] Similarly, each display support 5 is preferably connected to the information device 6 via the vehicle's communication network 2.

[0052] Advantageously, the vehicle 2 also includes at least one camera 8. In this case, each camera 8 is connected to the information device 6, in particular via the communication network of the vehicle 2, to transmit at least one image acquired by said camera 8 to the information device 6.

[0053] Such a feature is advantageous, insofar as, as will be described later, the presence of camera(s) allows the acquisition of images of target individuals, with a view to generating an animated facial visual.

[0054] Advantageously, the vehicle 2 also includes at least one microphone 9. In this case, each microphone 9 is connected to the information device 6, in particular via the communication network of the vehicle 2, to transmit, to the information device 6, at least one sound sequence acquired by said microphone 9.

[0055] Such a feature is advantageous, insofar as, as will be described later, the presence of microphone(s) allows the acquisition of natural vocal sequences uttered by said target individuals, for the purpose of synthesizing a synthesized voice announcement.

[0056] Preferably, the vehicle 2 also includes a geolocation device 10 and / or an input interface 12. In this case, the information device 6 is also capable of communicating with the geolocation device 10 and the input interface 12, in particular via the communication network of the vehicle 2.

[0057] Sound diffusion device 4 and display support 5

[0058] Each sound diffusion device 4 includes at least one loudspeaker configured to diffuse an acoustic wave representative of a signal received at the input of said sound diffusion device 4.

[0059] Preferably, the sound diffusion device 4 also includes a processing unit configured, for example, to shape the received signal, in particular by filtering and / or amplification, before its diffusion by the corresponding loudspeaker or loudspeakers.

[0060] In addition, each display support 5 is configured to broadcast a video sequence received at its input (in particular the animated facial visual mentioned above). For example, each display support 5 includes a screen.

[0061] Camera 8

[0062] In a conventional manner, each camera 8 is configured to acquire at least one image of a scene included in a field of view 13 of said camera 8.

[0063] Preferably, each camera 8 is mounted on board the vehicle 2. In this case, each camera 8 is positioned so that a predetermined volume of the vehicle 2 is included in a field of vision of said camera.

[0064] In particular, each camera 8 is positioned so that, when a target individual 14 is at a predetermined location, at least the face of said target individual 14 is within the predetermined volume included in the field of vision of said camera 8. In other words, when the individual is at the predetermined location, at least his face is likely to be imaged by the camera 8.

[0065] In particular, at least one camera 8 is positioned so as to image a driver of the vehicle 2 when said driver is at his driving position. In this case, said camera 8 is preferably mounted on a console in a driver's cab of the vehicle 2, or even on a front windshield of the driver's cab.

[0066] Alternatively, or in addition, at least one camera 8 is positioned to image a target individual 14 distinct from the driver of vehicle 2, such as a member of the vehicle 2 crew. Such an alternative is easy to implement, since it is common practice for railway vehicle controllers, for example, to enter a driver's cab to initiate a commercial service. The target individual could also be a staff member from a station where vehicle 2 departs, who temporarily enters the driver's cab to initiate the commercial mission, before letting vehicle 2 carry out its commercial mission in "single agent operation" (SAO) mode.

[0067] Alternatively, or in addition, at least one camera 8 is not mounted on board the vehicle 2. In this case, the camera 8 is preferably positioned so as to image at least one target individual 14, in particular the driver of the vehicle 2, during his journey to board the vehicle 2. For example, the camera 8 is positioned so as to image a vicinity of a door intended to be used by the target individual to board the vehicle 2.

[0068] Microphone 9

[0069] Each microphone 9 is configured to acquire at least one sound sequence representative of an acoustic wave detected by said microphone 9.

[0070] Each microphone 9 is positioned within a predetermined perimeter around a corresponding predetermined location.

[0071] Preferably, each microphone 9 is mounted on board the vehicle 2.

[0072] In particular, at least one microphone 9 is positioned so as to acquire a sound sequence (forming a natural vocal sequence) uttered by the target individual 14 located at a corresponding predetermined location on board the vehicle 2.

[0073] For example, such a voice sequence is a natural voice sequence uttered by the driver of vehicle 2 when said driver is at the driver's seat. In this case, said microphone 9 is preferably mounted on a console in the driver's cab of vehicle 2, or even on a front windshield of the driver's cab. Alternatively, at least one microphone 9 is a microphone shared by a plurality of systems in vehicle 2. In the field of rail transport, this may, for example, be the internal communication handset (also called a "gooseneck" microphone) used by the driver to communicate, inside the rail vehicle, with other staff members or with passengers via the public address and intercom systems.It may also be the microphone of the ground-to-train radio intended for communications between the driver and the traffic management staff of the infrastructure manager.

[0074] In such a configuration, the natural voice sequence is also likely to be a voice sequence uttered by a target individual 14 distinct from the driver, such as a member of the vehicle 2 crew, or a staff member from a vehicle 2 departure station who temporarily gets into the driver's cab.

[0075] Alternatively, or as a complement, at least one microphone 9 is not carried on board the vehicle 2. In this case, the microphone 9 is preferably arranged in a centralized control station of a transport network of the vehicle 2, or in a control station of a stop of the vehicle 2, so as to acquire a voice sequence spoken by a predetermined operator.

[0076] Geolocation device 10

[0077] The geolocation device 10 is configured to store information relating to each journey that the vehicle 2 is likely to make.

[0078] In particular, for each journey (also called a "commercial mission"), the geolocation device 10 is configured to store a list of corresponding stops (i.e., stations, in the case of a rail vehicle) served. Preferably, the geolocation device 10 is also configured to store the distances between successive stops and / or the coordinates of each stop, including the GPS coordinates of each stop.

[0079] In addition, the geolocation device 10 is configured to determine, over time, a current position of the vehicle 2. Such a position includes, for example, coordinates of the vehicle 2 in a reference frame and / or a distance traveled by the vehicle 2 from a predetermined position, for example a previous stop.

[0080] For example, the geolocation device 10 is configured to receive signals from one or more beacons (for example, GPS signals), and to determine the current coordinates of the vehicle 2 from the received signals.

[0081] In this case, the geolocation device 10 is configured to compare the coordinates of vehicle 2 with the stored coordinates of each stop. In addition, the geolocation device 10 is configured to send an alert to the information device 6 when the current position of vehicle 2 meets a predetermined condition, for example, when the distance between the current position of vehicle 2 and a stop on the current commercial route becomes less than or equal to a predetermined threshold (generally a few hundred meters, in order to anticipate the announcement made to passengers).

[0082] Alternatively, or in addition, the geolocation device 10 includes at least one odometer configured to calculate a distance travelled by the vehicle 2 from a predetermined position.

[0083] In this case, the geolocation device 10 is configured to send an alert to the information device 6 when the distance traveled since the last stop meets a predetermined condition, for example when the difference between the distance traveled since the last stop and the distance memorized between said stop and the next arrival for the current commercial mission becomes less than or equal to the predetermined threshold.

[0084] Input interface 12

[0085] The input interface 12 is a human / machine interface adapted to allow an operator to input instructions and / or indications relating to an operational situation of the vehicle 2.

[0086] Preferably, the input interface 12 includes at least one button and / or at least one touch surface which has been previously associated with a predetermined operational situation.

[0087] In this case, the input interface 12 is configured to send, to the information device 6, an alert representative of the operational situation associated with the button (respectively, with the touch surface) when it is used by the operator.

[0088] Preferably, the input interface 12 includes a keyboard allowing the operator to enter a text sequence.

[0089] In this case, the input interface 12 is configured to send, to the information device 6, the text sequence entered by the operator.

[0090] Information device 6

[0091] The information device 6 is intended for the dissemination of information to passengers of the vehicle 2, or to persons located near the vehicle 2. In particular, the information device 6 is configured to control the dissemination of sound information by the sound dissemination device 4 and of corresponding animations by the display support 5.

[0092] Hereafter, the pair comprising a sound announcement and the corresponding animation is designated by the expression "audiovisual announcement".

[0093] Preferably, the information device 6 has a compact shape ensuring a small footprint in the vehicle 2, and is, in particular, compliant with railway standards for electromagnetic compatibility, fire and smoke resistance, and shock and vibration resistance.

[0094] The information device 6 includes a memory 15, an optional voice feature extraction module 16 (called "extraction module"), an optional text analysis module 18 and an audiovisual synthesis module 20.

[0095] Memory 15

[0096] Memory 15 is configured to store information relating to each journey that vehicle 2 is capable of making. In particular, for each journey, memory 15 is configured to store the list of corresponding stops served.

[0097] In addition, memory 15 is configured to store at least one text sequence. In this case, each text sequence is associated with a corresponding operational situation.

[0098] For example, at least one stored text sequence is a notification of arrival of vehicle 2 at a stop, or a notification of departure of vehicle 2 from a stop.

[0099] According to another example, at least one memorized text sequence corresponds to a statement of the stops served by vehicle 2 during a given journey.

[0100] According to another example, at least one stored text sequence is a warning message, such as: • a warning message to inform passengers of an ongoing event; • a warning message inviting users near or on board vehicle 2 to move away from the doors of said vehicle 2; • a warning message to alert passengers of an exceptional stop of vehicle 2 between two stops, etc.

[0101] Optionally, the memory 15 is configured to store at least one representative image of the face of a target individual 14 (i.e. a portrait of the target individual), in particular a predetermined member of the vehicle 2 crew, for example a predetermined driver of vehicle 2.

[0102] Such a portrait has preferably been previously acquired using a camera 8.

[0103] Optionally, memory 15 is configured to store at least one natural vocal sequence uttered by a target individual, in particular a predetermined member of the vehicle 2 crew, for example a predetermined driver of vehicle 2. In this case, each vocal sequence is preferably uniquely associated with the portrait of the respective target individual.

[0104] Such a voice sequence has preferably been previously acquired using a microphone 9.

[0105] Extraction Module 16

[0106] The extraction module 16 is configured to extract a set of vocal features from each natural vocal sequence associated with a given target individual 14.

[0107] Preferably, to extract the set of voice features, the extraction module 16 is configured to implement an artificial intelligence model with a so-called GE2E architecture (from the English "Generalized End-to-End loss", or generalized end-to-end cost).

[0108] Such an architecture is described by Li Wan et al. in the digital preprint "Generalized End-to-End Loss for Speaker Verification", referenced arXiv:1710.10467.

[0109] Such a set of vocal characteristics includes values ​​that can be likened to a vocal signature of the target individual.

[0110] For example, the set of voice features produced as output by the extraction module 16 has a matrix structure. In this case, the values ​​taken by the matrix coefficients represent a decomposition of a speaker's (e.g., the target individual's) voice signature in a given tensor space. Such a space is such that, for two speakers with similar voices on a sonic plane, the distance between the vectors representing the respective vocal signatures is small.

[0111] Preferably, the extraction module 16 is configured to store the extracted feature set in a predetermined location in memory 15, in the form of a binary file with a predefined name. Even more preferably, the extracted voice feature set is stored in association with an identifier of the respective target individual.

[0112] Text analysis module 18

[0113] Preferably, the text analysis module 18 is configured to provide, as input to a previously trained language model, a text sequence from among the text sequences stored in memory 15. In this case, an output of the language model includes the emotional marker corresponding to said text sequence.

[0114] Preferably, the language model implemented by the text analysis module 18 is the RoBERTa language model, as described by Yinhan Liu et al. in the digital preprint "RoBERTa: A Robustly Optimized BERT Pretraining Approach", referenced arXiv:1907.11692.

[0115] More specifically, the RoBERTa model is subjected to non-specific pre-training for emotion marker recognition, based on several datasets (such as BookCorpus, English Wikipedia, CC-News, OpenWebText, and Stories, which are known to those skilled in the art). Then, the trained model is subjected to fine-tuning training based on a dataset specific to emotion marker recognition, such as Sentiment 140.

[0116] Within the framework of the present invention, the improvement training is advantageously carried out on the basis of a dataset comprising textual sequences of traveler information, each textual sequence being associated with a label representing the corresponding emotional marker.

[0117] In this case, each emotional marker produced as output by the language model corresponds to a class of said language model.

[0118] Such an emotional marker is representative of an emotion associated with the textual sequence, for example joy (for example, in the case of an announcement that is pleasing to users, such as the announcement of an arrival at a terminus station of the current commercial mission), sadness (for example, in the case of an unexpected parking of vehicle 2), or even a neutral emotion (for example, in the case of an announcement of the name of a current stop).

[0119] The text analysis module 18 is further configured to transmit the determined emotional marker to the audiovisual synthesis module 20.

[0120] Alternatively, for each text sequence stored in memory 15, memory 15 is also configured to store a corresponding emotional marker. In this case, the implementation of a text analysis module as described above is not required.

[0121] Audiovisual synthesis module 20

[0122] The audiovisual synthesis module 20 is configured to synthesize each audiovisual announcement to be broadcast.

[0123] More specifically, the audiovisual synthesis module 20 is configured to implement a speech synthesis engine 22 in order to synthesize, from a given textual sequence, a corresponding synthesized speech announcement.

[0124] In addition, the audiovisual synthesis module 20 is configured to implement a speech-based facial animation model 24 (referred to as the "animation model") in order to generate, from said synthesized voice announcement and a portrait of a target individual, a corresponding animated facial visual.

[0125] More specifically, the audiovisual synthesis module 20 is configured to, initially, perform a parameterization of the speech synthesis engine 22 according to the set of voice characteristics extracted by the extraction module 16, in order to obtain a parameterized speech synthesis engine.

[0126] For example, to achieve such a setting, the audiovisual synthesis module 20 is configured to replace all or part of the parameters of the speech synthesis engine 22 with the set of speech features previously extracted by the extraction module 16 (or with the result of a predetermined function applied to said set of speech features).

[0127] In particular, the audiovisual synthesis module 20 is configured to search, at the previously mentioned predetermined location in memory 15, for a binary file having the predetermined name.

[0128] Preferably, the speech synthesis engine 22 is a sequence-to-sequence artificial intelligence model featuring, in particular, an encoder-decoder architecture.

[0129] Preferably, the audiovisual synthesis module 20 is also configured to detect a current operational situation.

[0130] Preferably, the audiovisual synthesis module 20 is configured to detect the current operational situation from the current position of the vehicle 2 determined by the geolocation device 8, and / or alerts issued by the geolocation device 8.

[0131] Alternatively, or in addition, the audiovisual synthesis module 20 is configured to detect the current operational situation from an input performed by an operator on the input interface 10, such as the activation of a button specific to a given operational situation.

[0132] In addition, the audiovisual synthesis module 20 is configured to load the text sequence for which a corresponding audiovisual advertisement is to be synthesized.

[0133] Where the audiovisual synthesis module 20 is configured to detect the current operational situation, the audiovisual synthesis module 20 is advantageously configured to load, from memory 15, the text sequence corresponding to said detected current operational situation, in order to synthesize the audiovisual announcement. Such a feature is advantageous insofar as it ensures consistency between the current context and the audiovisual announcement broadcast to users.

[0134] Alternatively, or in addition, in the case where the input interface 10 is adapted to allow the operator to enter a text sequence, the audiovisual synthesis module 20 is configured to load the text sequence entered by the operator as a text sequence to be used for the synthesis of an audiovisual announcement.

[0135] In addition, the audiovisual synthesis module 20 is configured to provide the text sequence loaded as input to the parameterized speech synthesis engine. In this case, the output of the parameterized speech synthesis engine forms the synthesized speech announcement to be broadcast.

[0136] In addition, as previously stated, the audiovisual synthesis module 20 is configured to implement the animation model 24 in order to generate, from the synthesized voice announcement and a portrait of the target individual, the corresponding animated facial visual.

[0137] Such an animated facial visual is a simulated representation of the face of the target individual uttering the synthesized voice announcement.

[0138] Advantageously, in addition to the synthesized voice announcement, the audiovisual synthesis module 20 is also configured to provide, as input to the animation model 24, the emotional marker of the textual sequence corresponding to the synthesized voice announcement.

[0139] Such a feature is advantageous, insofar as it allows the generation of an animated portrait (i.e. an animated facial visual) whose facial expression is in line with the synthesized voice announcement to be broadcast.

[0140] Preferably, the animation model 24 implemented by the audiovisual synthesis module 20 is the Emotalkingface model, as described by Sefik Emre Eskimez et al. in the digital preprint "Speech Driven Talking Face Generation from a Single Image and an Emotion Condition", referenced arXiv:2008.03592.

[0141] In addition, the audiovisual synthesis module 20 is configured to transmit, on the one hand, the synthesized voice announcement to all or part of the sound broadcasting devices 4 and, on the other hand, the animated facial visual to all or part of the display media 5, for their synchronous broadcast in the vehicle 2 and / or in the vicinity of the vehicle 2.

[0142] Operation

[0143] The operation of the information device 6 will now be described with reference to the figures.

[0144] During an initial memorization step, information relating to each journey that the vehicle 2 is likely to make is stored in memory 15.

[0145] In addition, at least one text sequence is stored in memory 15.

[0146] In addition, information relating to each journey that vehicle 2 is likely to make is also stored in the geolocation device 10.

[0147] Then, during an initialization step, an operator (for example, a driver) initializes the information device 6, in particular by entering a code for the upcoming journey.

[0148] Then, the audiovisual announcement synthesis process 30 according to the invention is implemented.

[0149] Preferably, during an optional extraction step 32, the extraction module 16 extracts a set of vocal features from at least one natural vocal sequence associated with a given target individual 14. Such a vocal sequence has preferably been previously acquired using a microphone 9.

[0150] Then, during a parameterization step 34, the audiovisual synthesis module 20 performs a parameterization of the speech synthesis engine 22 according to the set of voice characteristics extracted by the extraction module 16. This results in a parameterized speech synthesis engine.

[0151] Alternatively, the set of vocal characteristics associated with the target individual 14 has been previously determined and stored in memory 15. In this case, during the parameterization step 34, the audiovisual synthesis module 20 performs the parameterization of the speech synthesis engine 22 from said set of vocal characteristics stored in memory 15.

[0152] Preferably, when the audiovisual synthesis module 20 detects a given current operational situation requiring the broadcast of an audiovisual announcement, the audiovisual synthesis module 20 loads a text sequence corresponding to said detected current operational situation.

[0153] For example, the audiovisual synthesis module 20 detects the current operational situation based on alerts issued by the geolocation device 10 and / or by the input interface 12.

[0154] Alternatively, or in addition, the audiovisual synthesis module 20 detects the current operational situation as the entry of a text sequence by an operator, via the input interface 12. In this case, the text sequence loaded is the text sequence entered by the operator.

[0155] Then, during a speech synthesis step 36, the audiovisual synthesis module 20 provides the text sequence loaded as input to the parameterized speech synthesis engine 22. In this case, the output of the parameterized speech synthesis engine forms a synthesized voice announcement to be broadcast.

[0156] Then, preferably, during an optional text analysis step 38, the text analysis module 18 provides, as input to the trained language model, the text sequence loaded by the audiovisual synthesis module. The resulting output of the language model is an emotional marker corresponding to said text sequence.

[0157] In this case, the text analysis module 18 transmits the determined emotional marker to the audiovisual synthesis module 20.

[0158] Then, during an animated facial image generation step 40 (referred to as the "generation step"), the audiovisual synthesis module 20 implements the animation model 24 in order to generate, from the synthesized voice announcement and a portrait of the target individual, the corresponding animated facial image. Such a portrait has preferably been previously acquired using a camera 8.

[0159] Advantageously, in the case where the text analysis step 38 has been implemented, the audiovisual synthesis module 20 also provides, as input to the animation model 24, in addition to the synthesized voice announcement, the emotional marker of the textual sequence corresponding to the synthesized voice announcement.

[0160] Alternatively, the emotional marker of the text sequence corresponding to the synthesized voice announcement has been previously stored in memory 15, in association with said text sequence. In this case, the audiovisual synthesis module 20 implements the animation model based on the synthesized voice announcement, the portrait of the target individual, and the emotional marker of the text sequence associated with said synthesized voice announcement, which has been loaded from memory 15.

[0161] Then, during a broadcasting step 42, the audiovisual synthesis module 20 transmits, preferably, on the one hand, the synthesized voice announcement to all or part of the sound broadcasting devices 4 and, on the other hand, the animated facial visual to all or part of the display media 5, for their synchronous broadcast in the vehicle 2 and / or in the vicinity of the vehicle 2.

[0162] Of course, the invention is not limited to the examples just described.

Claims

Demands

1. A computer-implemented method (30) for synthesizing announcements comprising the steps: • setting up (34) a speech synthesis engine according to a set of voice characteristics associated with a target individual; • synthesizing (36) a synthesized voice announcement from a text sequence, the synthesis comprising providing the text sequence as input to the parameterized speech synthesis engine, a corresponding output from the parameterized speech synthesis engine forming the synthesized voice announcement; • generating (40) an animated facial visual, comprising providing, as input to a facial animation model, the synthesized voice announcement and a portrait of the target individual.

2. A method according to claim 1, wherein the generation (40) of the animated facial visual further comprises providing, as input to the facial animation model, an emotional marker of the textual sequence associated with the synthesized voice announcement.

3. Method according to claim 2, comprising, prior to the generation (40) of the animated facial visual, a supply (38) of the textual sequence as input to a previously trained language model, an output of the language model comprising the emotional marker corresponding to the textual sequence.

4. A method according to any one of claims 1 to 3, comprising an extraction (32) of vocal features from at least one natural vocal sequence previously uttered by the target individual.

5. A method according to any one of claims 1 to 4, wherein the synthesis (36) further comprises a detection of a current operational situation, the text sequence being selected according to the detected current operational situation.

6. Method according to claim 5, wherein the detection of the current operational situation includes a geolocation of a target vehicle (2).

7. A method according to any one of claims 1 to 6, further comprising a transmission (42) of the synthesized voice announcement and the corresponding generated animated facial visual for their synchronous broadcast in and / or near a target vehicle (2).

8. A computer program comprising executable instructions which, when executed by a computer, implement the steps of the process according to any one of claims 1 to 7

9. / . Information device (6) configured to: • configure a speech synthesis engine based on a set of voice characteristics associated with a target individual; • synthesize a synthesized voice announcement from a text sequence, the synthesis comprising the provision of the text sequence as input to the parameterized speech synthesis engine, a corresponding output from the parameterized speech synthesis engine forming the synthesized voice announcement; • generate an animated facial visual, comprising the provision, as input to a facial animation model, of the synthesized voice announcement and a portrait of the target individual.

10. Vehicle (2), in particular a railway transport vehicle, comprising an information device (6) according to claim 9, at least one sound broadcasting device (4) connected to the information device (6) for broadcasting each synthesized voice announcement in the vehicle and / or in the vicinity of the vehicle, and at least one display support (5) for displaying each animated facial visual, the information device (6) being configured to control each sound broadcasting device (4) and each display support (5) so as to broadcast each synthesized voice announcement and the corresponding animated facial visual synchronously.

11. Vehicle (2) according to claim 10, further comprising: • at least one camera connected to the information device, each camera being configured to acquire an image of a predetermined volume of the vehicle, and to transmit the acquired image to the information device; and / or • at least one microphone connected to the information device, each microphone being adapted to acquire a natural speech sequence, and to transmit the acquired natural speech sequence to the information device, the information device includes, in addition, an extraction module configured to extract the set of vocal features from the acquired natural vocal sequence.

Citation Information

Patent Citations

  • Processing speech to drive animations on avatars

    US10521946B1