Speech synthesis process, associated device and vehicle
By training a generative AI model on natural speech with physical attributes to simulate crew voices, the method enhances user comfort in rail transport systems by producing natural-sounding announcements without costly system replacements.
Patent Information
- Application Number
- FR2024005409
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-27
- Publication Date
- 2025-11-28
AI Technical Summary
Existing on-board passenger information systems (SIVEs) in rail transport vehicles use artificial voices, which induce feelings of loneliness and insecurity among users, and replacing them with natural voices is costly and complex.
A method involving a generative artificial intelligence model trained on natural speech sequences with physical attributes to simulate natural vocal sequences, parameterizing a speech synthesis engine to produce voice announcements that mimic crew members, eliminating the need for system replacement.
Provides voice announcements that are sonically similar to those of human crew members, enhancing user comfort without significant complexity or cost, addressing the emotional disconnect caused by artificial voices.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Speech synthesis method, associated device and vehicle. Technical field
[0001] The present invention relates to a computer-implemented speech synthesis method.
[0002] The invention also relates to a computer program, a device implementing such a method and a vehicle incorporating such a device.
[0003] The invention applies to the field of information dissemination systems for users, in particular for users of a transport network. Prior art
[0004] It is known to use information dissemination systems for users of a transport network. In particular, in the field of rail transport, it is known to use systems called "on-board passenger information systems" (or IPS).
[0005] Such SIVEs are commonly carried on board the railway vehicle, and include a computer and display screens and / or speakers controlled by the computer.
[0006] Such SIVEs are, in particular, configured to produce, for users, contextual information of a commercial mission (defined as the journey of a railway vehicle from a departure station A to a terminus station B), in particular a commercial mission in progress.
[0007] In particular, the computer is configured to use an embedded database, as well as GPS (for "Global Positioning System") and / or odometric data, in order to generate the information to be disseminated and determine its time of dissemination.
[0008] Such a database is, in particular, configured to store, for each commercial mission, information such as a mission code, a list of stations served and their respective GPS coordinates, as well as a list of distances between successive stations.
[0009] Typically, during a commercial mission, a SIVE first loads the data for the upcoming commercial mission from the database (for example, the list of stations to be served and their characteristics). Such a commercial mission is, in particular, determined based on a mission code entered by the train crew.
[0010] The SIVE also stores text sequences containing, in particular, announcements corresponding to different operating contexts that the rail vehicle may encounter. In this case, the text sequences (including station names) are transformed into synthetic speech sequences by a speech synthesis engine configured with an artificial voice.
[0011] The SIVE also stores sound announcements generated from text sequences by a speech synthesis engine, also configured with an artificial voice.
[0012] Then, iteratively during the commercial mission, the SIVE determines the name of the current station and the next station in order to make a visual and audio announcement to the passengers.
[0013] In addition, on occasion, in specific operating contexts, the driver may trigger visual and / or audible announcements known as "contingency announcements", by means of dedicated controls located in the driver's cab.
[0014] For example, the SIVE is capable of broadcasting a visual and audible announcement when the train driver presses a specific button indicating that the train is forced to stop on the track. Such an announcement is, for example: "Our train is stopped on the track. For your safety, do not attempt to open the doors."
[0015] However, such a system does not give complete satisfaction.
[0016] Indeed, it has been observed that the use of an artificial voice reinforces feelings of loneliness and insecurity among users. In fact, the physical absence of cabin crew, combined with a voice unrelated to a human, induces a feeling of abandonment in the traveler.
[0017] However, the delivery of voice announcements cannot be entrusted to the cabin crew, as this would result in an undesirable additional workload. This constraint is even greater in a so-called "single-agent operation" (SA) mode, in which no sales representative accompanies the flight attendant.
[0018] Furthermore, for reasons of cost and complexity, it cannot be envisaged to replace the existing SIVEs of a fleet of railway vehicles in favor of SIVEs capable of performing speech synthesis with a natural voice.
[0019] One object of the present invention is to remedy at least one of the drawbacks of the prior art.
[0020] Another objective of the invention is to propose a speech synthesis method whose deployment is not very complex and not very expensive. Description of the invention
[0021] To this end, the invention relates to a method of the aforementioned type, comprising the steps: • generation of at least one simulated natural speech sequence, the generation comprising the provision of at least one physical speaker attribute as input to a generative artificial intelligence model, an output of the generative artificial intelligence model forming said at least one simulated natural speech sequence, the generative artificial intelligence model having been previously trained on the basis of a training dataset comprising at least one natural speech sequence, each natural speech sequence being associated with at least one physical attribute of the corresponding locator; • extraction of a set of vocal features from each generated simulated natural vocal sequence; and • Parameterization of a speech synthesis engine based on the extracted set of vocal characteristics.
[0022] Indeed, thanks to the generation step, at least one vocal sequence (called a "simulated natural vocal sequence") with a voice comparable to a natural voice is simulated. In particular, due to the way in which the generative artificial intelligence model was trained, providing a constraint to said model, in the form of at least one physical attribute, results in the generation of simulated natural vocal sequences that are sonically similar to those that would have been uttered by a (human) speaker exhibiting said physical attribute(s).
[0023] In this way, vocal characteristics representative of natural voices associated with specific physical attributes are likely to be extracted from the generated simulated natural vocal sequences, and used to parameterize the speech synthesis engine.
[0024] This is all the more advantageous when the speech synthesis engine is the native speech synthesis engine of a given vehicle's information system: with such a setting (which is easy to implement), the replacement of the information system is no longer required, and voice announcements with a natural voice can be broadcast to vehicle users.
[0025] Furthermore, by appropriately selecting the physical attributes used to generate the simulated natural speech sequences, the speech synthesis engine, once configured, is capable of synthesizing voice announcements that are sonically similar to announcements expected by listeners of said announcements. Indeed, there is a correlation between a speaker's voice and their physical attributes (such as their sex, age, and ethnicity).
[0026] In particular, by constraining the generative artificial intelligence model with attributes of a vehicle crew member, the synthesized voice announcements are close, in terms of sound, to those that would have been made by said crew member, and that the vehicle's passengers would expect to hear after seeing said crew member.
[0027] Advantageously, the process according to the invention has one or more of the following characteristics, taken individually or in any technically feasible combination:
[0028] The process includes, prior to the generation step, a determination of each speaker's physical attribute, said determination including the provision of an image of a target individual as input to a classification model, each class provided as output of the model forming a corresponding speaker's physical attribute;
[0029] the method further comprises a synthesis of a synthesized voice announcement from a textual sequence, the synthesis comprising the provision of the textual sequence as input to the parameterized speech synthesis engine, a corresponding output of the parameterized speech synthesis engine forming the synthesized voice announcement;
[0030] the synthesis further includes the detection of a current operational situation, the text sequence being selected according to the detected current operational situation;
[0031] the detection of the current operational situation includes geolocation of a target vehicle;
[0032] The method further comprises a transmission of the synthesized voice announcement for its dissemination in and / or in the vicinity of a target vehicle.
[0033] According to another aspect of the invention, a computer program is proposed comprising executable instructions which, when executed by computer, implement the steps of the process as defined above.
[0034] The computer program can be in any computer language, such as for example in machine language, in C, C++, JAVA, Python, etc.
[0035] According to another aspect of the invention, an information device is proposed comprising: • a first speech synthesis module configured to generate at least one simulated natural speech sequence, the generation comprising the provision of at least one physical speaker attribute as input to a generative artificial intelligence model, an output of the generative artificial intelligence model forming said at least one simulated natural speech sequence, the generative artificial intelligence model having been previously trained on the basis of a training dataset comprising at less one natural voice sequence, each natural voice sequence being associated with at least one physical attribute of the corresponding locator; • a voice feature extraction module configured to extract a set of voice features from each generated simulated natural voice sequence; and • a second speech synthesis module configured to set up a speech synthesis engine based on the extracted set of voice characteristics.
[0036] Advantageously, the information device according to the invention has the following characteristic:
[0037] the second speech synthesis module is, furthermore, configured to perform the synthesis of a speech announcement from a text sequence, the synthesis comprising the provision of the text sequence as input to the parameterized speech synthesis engine, a corresponding output of the parameterized speech synthesis engine forming the synthesized speech announcement.
[0038] The device according to the invention can be any type of device such as a server, a computer, a tablet, a calculator, a processor, a computer chip, programmed to implement the method according to the invention, for example by executing the computer program according to the invention.
[0039] According to another aspect of the invention, a vehicle is proposed, in particular a railway transport vehicle, comprising an information device as defined above and at least one sound diffusion device connected to the information device to broadcast each synthesized voice announcement in the vehicle and / or in the vicinity of the vehicle.
[0040] Advantageously, the vehicle according to the invention has the following characteristic:
[0041] the vehicle further comprises at least one camera connected to the information device, each camera being configured to acquire an image of a predetermined volume of the vehicle, and to transmit the acquired image to the information device, the information device further comprising an image processing module configured to determine, for each individual represented on the acquired image, at least one physical attribute of said individual, and to provide each determined physical attribute as input to the first speech synthesis module. Brief description of the figures
[0042] The invention will be better understood upon reading the following description, given solely by way of non-limiting example and made with reference to the accompanying drawings in which:
[0043] [Fig. 1] is a schematic representation of a vehicle according to the invention; and
[0044] [Fig. 2] is a flowchart of a speech synthesis process implemented by a vehicle information device of the [Fig.l].
[0045] It is understood that the embodiments described below are by no means limiting. In particular, variants of the invention may be conceived comprising only a selection of the features described below, isolated from the other features described, if this selection of features is sufficient to confer a technical advantage or to differentiate the invention from the prior art. This selection includes at least one preferably functional feature without structural details, or with only a portion of the structural details if this portion alone is sufficient to confer a technical advantage or to differentiate the invention from the prior art.
[0046] In particular, all the variants and embodiments described are combinable with each other if there is no technical obstacle to this combination.
[0047] In the figures and in the rest of the description, elements common to several figures retain the same reference. Detailed description
[0048] A vehicle 2 according to the invention is illustrated by [Fig.1].
[0049] Vehicle 2 is, in particular, a passenger transport vehicle. Preferably, vehicle 2 is a rail transport vehicle such as a train, in particular a rail transport vehicle intended for the transport of passengers.
[0050] The vehicle 2 includes at least one sound diffusion device 4 and one information device 6.
[0051] Preferably, for their communication, the information device 6 and each sound diffusion device 4 are connected to a corresponding port of a network switch (also called a "switch") of a vehicle communication network 2, such as a wired Ethernet IP communication network with a ring topology.
[0052] Advantageously, the vehicle 2 also includes at least one camera 8. In this case, each camera 8 is connected to the information device 6, in particular via the communication network of the vehicle 2, to transmit at least one image acquired by said camera 8 to the information device 6.
[0053] Such a feature is advantageous, insofar as, as will be described later, the presence of camera(s) allows the acquisition of images of target individuals, with a view to determining corresponding physical attributes, which may be implemented during the process according to the invention.
[0054] Preferably, the vehicle 2 also includes a geolocation device 10 and / or an input interface 12. In this case, the information device 6 is also capable of communicating with the geolocation device 10 and the input interface 12, in particular via the communication network of the vehicle 2.
[0055] Sound diffusion device 4
[0056] Each sound diffusion device 4 includes at least one loudspeaker configured to diffuse an acoustic wave representative of a signal received at the input of said sound diffusion device 4.
[0057] Preferably, the sound diffusion device 4 also includes a processing unit configured, for example, to shape the received signal, in particular by filtering and / or amplification, before its diffusion by the corresponding loudspeaker or loudspeakers.
[0058] Camera 8
[0059] In a conventional manner, each camera 8 is configured to acquire at least one image of a scene included in the field of view 13 of said camera 8.
[0060] Preferably, each camera 8 is mounted on board the vehicle 2. In this case, each camera 8 is positioned so that a predetermined volume of the vehicle 2 is included in a field of vision of said camera.
[0061] In particular, each camera 8 is positioned so that, when an individual is at a predetermined location, at least the upper body (including, preferably, the face) of said individual is within the predetermined volume within the field of view of said camera 8. In other words, when the individual is at the predetermined location, at least the upper part of his body is likely to be imaged by the camera 8.
[0062] In particular, at least one camera 8 is positioned so as to image a driver 14 of the vehicle 2 when said driver is at his driving position. In this case, said camera 8 is preferably mounted on a console in a driver's cab of the vehicle 2, or even on a front windshield of the driver's cab.
[0063] Alternatively, or in addition, at least one camera 8 is positioned to image a target individual distinct from the driver 14, such as a member of the vehicle 2 crew. Such an alternative is easy to implement, since it is common practice for train controllers to enter the driver's cab to initiate a commercial operation. The target individual could also be a staff member from a station where vehicle 2 departs, who temporarily enters the driver's cab to initiate the commercial operation before allowing vehicle 2 to carry out its commercial operation in single-agent operation (SA) mode.
[0064] Alternatively, or additionally, at least one camera 8 is not mounted on board the vehicle 2. In this case, the camera 8 is preferably, positioned so as to image at least one target individual, in particular a driver of vehicle 2, during their journey to board vehicle 2. For example, camera 8 is positioned so as to image a vicinity of a door intended to be used by the target individual to board vehicle 2.
[0065] Geolocation device 10
[0066] The geolocation device 10 is configured to store information relating to each journey that the vehicle 2 is likely to make.
[0067] In particular, for each journey (also called a "commercial mission"), the geolocation device 10 is configured to store a list of corresponding stops (i.e., stations, in the case of a rail vehicle) served. Preferably, the geolocation device 10 is also configured to store the distances between successive stops and / or the coordinates of each stop, including the GPS coordinates of each stop.
[0068] In addition, the geolocation device 10 is configured to determine, over time, a current position of the vehicle 2. Such a position includes, for example, coordinates of the vehicle 2 in a reference frame and / or a distance traveled by the vehicle 2 from a predetermined position, for example a previous stop.
[0069] For example, the geolocation device 10 is configured to receive signals from one or more beacons (for example, GPS signals), and to determine the current coordinates of the vehicle 2 from the received signals.
[0070] In this case, the geolocation device 10 is configured to compare the coordinates of vehicle 2 with the stored coordinates of each stop. In addition, the geolocation device 10 is configured to send an alert to the information device 6 when the current position of vehicle 2 meets a predetermined condition, for example, when the distance between the current position of vehicle 2 and a stop on the current commercial route becomes less than or equal to a predetermined threshold (generally a few hundred meters, in order to anticipate the announcement made to passengers).
[0071] Alternatively, or in addition, the geolocation device 10 includes at least one odometer configured to calculate a distance travelled by the vehicle 2 from a predetermined position.
[0072] In this case, the geolocation device 10 is configured to send an alert to the information device 6 when the distance traveled since the last stop meets a predetermined condition, for example when the difference between the distance traveled since the last stop and the distance memorized between said stop and the next arrival for the current commercial mission becomes less than or equal to the predetermined threshold.
[0073] Input interface 12
[0074] The input interface 12 is a human / machine interface adapted to allow an operator to input instructions and / or indications relating to an operational situation of the vehicle 2.
[0075] Preferably, the input interface 12 includes at least one button and / or at least one touch surface which has been previously associated with a predetermined operational situation.
[0076] In this case, the input interface 12 is configured to send, to the information device 6, an alert representative of the operational situation associated with the button (respectively, with the touch surface) when it is used by the operator.
[0077] Preferably, the input interface 12 includes a keyboard allowing the operator to enter a text sequence.
[0078] In this case, the input interface 12 is configured to send, to the information device 6, the text sequence entered by the operator.
[0079] Information device 6
[0080] The information device 6 is intended for the dissemination of information to passengers of vehicle 2, or to persons located near vehicle 2. In particular, the information device 6 is configured to control the dissemination of sound information by the sound dissemination device 4.
[0081] Preferably, the information device 6 has a compact form ensuring a small footprint in the vehicle 2, and is, in particular, compliant with railway standards for electromagnetic compatibility, fire and smoke resistance, and shock and vibration resistance.
[0082] The information device 6 includes a memory 15, an optional module 16 image processing, a first speech synthesis module 18, a module 20 for extracting voice features (called "extraction module") and a second speech synthesis module 22.
[0083] Memory 15
[0084] Memory 15 is configured to store information relating to each journey that vehicle 2 is capable of making. In particular, for each journey, memory 15 is configured to store the list of corresponding stops served.
[0085] In addition, memory 15 is configured to store at least one text sequence. In this case, each text sequence is associated with a corresponding operational situation.
[0086] For example, at least one stored text sequence is a notification of the arrival of vehicle 2 at a stop, or a notification of the departure of vehicle 2 from a stop.
[0087] According to another example, at least one memorized text sequence corresponds to a statement of the stops served by vehicle 2 during a given journey.
[0088] According to another example, at least one stored text sequence is a warning message, such as: • a warning message to inform passengers of an ongoing event; • a warning message inviting users near or on board vehicle 2 to move away from the doors of said vehicle 2; • a warning message to alert passengers of an exceptional stop of vehicle 2 between two stops, etc.
[0089] Optionally, memory 15 is configured to store at least one physical speaker attribute. In particular, memory 15 is configured to store at least one physical attribute of at least one predetermined individual, including a predetermined crew member of vehicle 2, for example, a predetermined driver of vehicle 2. In this case, the at least one physical attribute is stored in memory 15 in association with a unique identifier of the respective predetermined individual.
[0090] Image processing module 16
[0091] The image processing module 16 is configured to determine at least one physical attribute of a target individual represented in an image, in particular an image from a camera 8.
[0092] In order to carry out such processing, the image processing module 16 is configured to implement a previously trained classification model.
[0093] In particular, to perform such processing, the image processing module 16 is configured to provide, as input to the classification model, the representative image of the target individual. In this case, the corresponding output of the classification model comprises a set of physical attributes associated with the target individual. As a result, each physical attribute provided as output by the classification model corresponds to a class of said classification model.
[0094] It should be noted that such a set of physical attributes is likely to include only one physical attribute.
[0095] The classification model was previously trained on the basis of an auxiliary training dataset comprising at least one image of an individual, associated with a set of corresponding physical attributes.
[0096] In this case, during the training of the classification model, each image of the auxiliary training dataset forms an input of said classification model, while the respective set of physical attributes forms an expected output of the classification model.
[0097] Such physical attributes of the auxiliary training dataset include at least one of a gender class (e.g., male, female), an age group class (e.g., 15-24 years, 25-39 years, 40-49 years, 50-70 years, over 70 years) and an ethnicity class (e.g., Asian, African American, Caucasian, Indian, Middle Eastern, Latino-Hispanic).
[0098] Preferably, the classification model is a convolutional neural network, for example featuring the EfficientNet architecture.
[0099] Such an architecture is described by Mingxing Tan et al. in the digital preprint "EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks", referenced arXiv: 1905.11946.
[0100] First speech synthesis module 18
[0101] The first speech synthesis module 18 is configured to generate at least one simulated natural speech sequence from a set of physical attributes (said set being likely to include a single physical attribute).
[0102] More specifically, to generate each simulated natural speech sequence, the first speech synthesis module 18 is configured to implement a generative artificial intelligence model, and more specifically a conditional generative artificial intelligence model.
[0103] In this case, the first speech synthesis module 18 is configured to provide, as input to the generative artificial intelligence model, the set of physical attributes determined by the image processing module 16. Alternatively, when the memory 15 is configured to store each set of physical attributes, the first speech synthesis module 18 is configured to provide, as input to the generative artificial intelligence model, a set of physical attributes memorized and associated with a target individual, in particular the set of physical attributes corresponding to one of the current crew members of the vehicle 2.
[0104] In both cases, an output of the generative artificial intelligence model forms at least one simulated natural speech sequence.
[0105] Preferably, but not exclusively, the generative artificial intelligence model is a Conditional Generative Adversarial Network (CGAN). Such a Conditional Generative Adversarial Network comprises, on the one hand, a generator network, configured to generate data from rules developed during a training phase and from labels, and, on the other hand, a discriminator network, configured to evaluate the conformity of the generated data with said rules and said labels.
[0106] The generative artificial intelligence model implemented by the first speech synthesis module 18 was previously trained on the basis of a training dataset.
[0107] In particular, the training dataset comprises a set of natural speech sequences (i.e., each uttered by a corresponding human), each associated with physical attributes of the corresponding speaker.
[0108] Preferably, the physical attributes used during the training of the generative artificial intelligence model are the classes likely to be produced as output from the image processing module 16, or the physical attributes stored in memory 15.
[0109] Extraction Module 20
[0110] The extraction module 20 is configured to extract a set of vocal features from each simulated natural speech sequence generated by the first speech synthesis module 18.
[0111] Preferably, to extract the set of voice features, the extraction module 20 is configured to implement an artificial intelligence model with a so-called GE2E architecture (from the English "Generalized End-to-End loss", or generalized end-to-end cost).
[0112] Such an architecture is described by Li Wan et al. in the digital preprint "Generalized End-to-End Loss for Speaker Verification", referenced arXiv:1710.10467.
[0113] Such a set of vocal characteristics includes values that can be likened to a vocal signature of a corresponding speaker.
[0114] In particular, such a set of vocal characteristics includes values that can be likened to a vocal signature of a speaker whose physical attributes would be similar to the physical attributes provided as input to the first speech synthesis module 18.
[0115] For example, the set of vocal features produced as output by the extraction module 20 has a matrix structure. In this case, the values taken by the matrix coefficients represent a decomposition of a speaker's vocal signature in a given tensor space. Such a space is such that, for two speakers with similar voices on a sonic plane, the distance between the vectors representing their respective vocal signatures is small.
[0116] Preferably, the extraction module 20 is configured to store, in a predetermined location in memory 15, the extracted set of voice features, in the form of a binary file with a predefined name.
[0117] Second speech synthesis module 22
[0118] The second speech synthesis module 22 is configured to implement a speech synthesis engine in order to synthesize, from a given text sequence, a corresponding speech announcement (called a "synthesized speech announcement").
[0119] Preferably, the speech synthesis engine implemented by the second speech synthesis module 22 is a sequence-to-sequence artificial intelligence model featuring, in particular, an encoder-decoder architecture.
[0120] More specifically, the second speech synthesis module 22 is configured to, initially, perform a parameterization of the speech synthesis engine according to the set of voice characteristics extracted by the extraction module 20, in order to obtain a parameterized speech synthesis engine.
[0121] For example, to achieve such a setting, the second speech synthesis module 22 is configured to replace all or part of the speech synthesis engine parameters with the set of speech features previously generated by the extraction module 20 (or with the result of a predetermined function applied to said set of speech features).
[0122] In particular, the second speech synthesis module 22 is configured to search, at the predetermined location of memory 15 mentioned above, for a binary file having the predetermined name.
[0123] In addition, the second speech synthesis module 22 is preferably configured to detect the current operational situation.
[0124] Preferably, the second speech synthesis module 22 is configured to detect the current operational situation from the current position of the vehicle 2 determined by the geolocation device 10, and / or alerts issued by the geolocation device 10.
[0125] Alternatively, or in addition, the second speech synthesis module 22 is configured to detect the current operational situation from an input made by the operator on the input interface 12, such as the actuation of a button specific to a given operational situation.
[0126] In addition, the second speech synthesis module 22 is configured to load a text sequence for which a corresponding speech announcement is to be synthesized.
[0127] Where the second speech synthesis module 22 is configured to detect the current operational situation, the second speech synthesis module 22 is advantageously configured to load, from memory 15, the text sequence corresponding to said detected current operational situation, in order to synthesize the voice announcement. Such a feature is advantageous insofar as it ensures consistency between the current context and the voice announcement issued to users.
[0128] Alternatively, or additionally, where the input interface 12 is adapted to allow the operator to enter a text sequence, the second speech synthesis module 22 is configured to load the entered text sequence. by the operator as a text sequence to be used for the synthesis of a voice announcement.
[0129] In addition, the second speech synthesis module 22 is configured to provide the loaded text sequence as input to the parameterized speech synthesis engine. In this case, the output of the parameterized speech synthesis engine forms the synthesized speech announcement to be broadcast.
[0130] In addition, the second speech synthesis module 22 is configured to transmit the synthesized speech announcement to all or part of the sound broadcasting devices 4 for broadcasting said synthesized speech announcement in the vehicle 2 and / or in the vicinity of the vehicle 2.
[0131] Operation
[0132] The operation of the information device 6 will now be described with reference to the figures.
[0133] During an initial memorization step, information relating to each journey that the vehicle 2 is likely to make is stored in memory 15.
[0134] In addition, at least one text sequence is stored in memory 15.
[0135] In addition, information relating to each journey that vehicle 2 is likely to make is also stored in the geolocation device 10.
[0136] Then, during an initialization step, an operator (for example, a driver) initializes the information device 6, in particular by entering a code for the upcoming journey.
[0137] Then, the speech synthesis process 30 according to the invention is implemented.
[0138] Preferably, during an optional image processing step 32, the image processing module 16 determines a set of physical attributes of a target individual represented on an image, in particular an image previously acquired using a camera 8.
[0139] Then, during a first speech synthesis step 34, the first speech synthesis module 18 generates at least one simulated natural speech sequence from the set of physical attributes determined by the image processing module 16.
[0140] Alternatively, the set of physical attributes of the target individual has been previously stored in memory 15. In this case, during the first speech synthesis step 34, the first speech synthesis module 18 loads the set of physical attributes corresponding to the target individual from memory 15, and generates at least one simulated natural speech sequence from the loaded set of physical attributes.
[0141] Then, during an extraction step 36, the extraction module 20 extracts a set of vocal features from each simulated natural vocal sequence generated by the first speech synthesis module 18.
[0142] Then, during a parameterization step 38, the second speech synthesis module 22 performs a parameterization of the speech synthesis engine according to the set of voice characteristics extracted by the extraction module 20.
[0143] This results in a parameterized speech synthesis engine.
[0144] Preferably, during a second speech synthesis step 40, when the second speech synthesis module 22 detects a given current operational situation requiring the broadcast of a voice announcement, the second speech synthesis module 22 loads a text sequence corresponding to said detected current operational situation.
[0145] For example, the second speech synthesis module 22 detects the current operational situation based on alerts issued by the geolocation device 10 and / or alerts issued by the input interface 12.
[0146] Alternatively, or in addition, the second speech synthesis module 22 detects the current operational situation as the entry of a text sequence by an operator, via the input interface 12. In this case, the text sequence loaded is the text sequence entered by the operator.
[0147] Then, the second speech synthesis module 22 provides the text sequence loaded as input to the parameterized speech synthesis engine. In this case, the output of the parameterized speech synthesis engine forms the synthesized speech announcement to be broadcast.
[0148] Then, during a broadcasting step 42, the speech synthesis module 18 transmits, preferably, the synthesized speech announcement to all or part of the sound broadcasting devices 4 for the broadcasting of said synthesized speech announcement in the vehicle 2 and / or in the vicinity of the vehicle 2.
[0149] Of course, the invention is not limited to the examples just described.
Claims
Demands
1. A computer-implemented speech synthesis method comprising the steps: • generating (34) at least one simulated natural speech sequence, the generation comprising providing at least one speaker physical attribute as input to a generative artificial intelligence model, an output of the generative artificial intelligence model forming said at least one simulated natural speech sequence, the generative artificial intelligence model having been previously trained on the basis of a training dataset comprising at least one natural speech sequence, each natural speech sequence being associated with at least one corresponding speaker physical attribute; • extracting (36) a set of speech features from each generated simulated natural speech sequence; and • parameterizing (38) a speech synthesis engine based on the extracted set of speech features.
2. A method according to claim 1, comprising, prior to the generation step (34), a determination (32) of each speaker physical attribute, said determination comprising providing an image of a target individual as input to a classification model, each class provided as output of the model forming a corresponding speaker physical attribute.
3. A method according to claim 1 or 2, further comprising a synthesis (40) of a synthesized speech announcement from a text sequence, the synthesis comprising supplying the text sequence as input to the parameterized speech synthesis engine, a corresponding output of the parameterized speech synthesis engine forming the synthesized speech announcement.
4. A method according to claim 3, wherein the synthesis further comprises a detection of a current operational situation, the text sequence being selected according to the detected current operational situation.
5. A method according to claim 4, wherein the detection of the current operational situation includes a geolocation of a target vehicle.
6. A method according to any one of claims 3 to 5, further comprising a transmission of the synthesized voice announcement for broadcast in and / or near a target vehicle.
7. A computer program comprising executable instructions which, when executed by a computer, implement the steps of the process according to any one of claims 1 to A
8. U. Information device (6) comprising: • a first speech synthesis module (18) configured to generate at least one simulated natural speech sequence, the generation comprising the provision of at least one speaker physical attribute as input to a generative artificial intelligence model, an output of the generative artificial intelligence model forming said at least one simulated natural speech sequence, the generative artificial intelligence model having been previously trained on the basis of a training dataset comprising at least one natural speech sequence, each natural speech sequence being associated with at least one corresponding speaker physical attribute; • a speech feature extraction module (20) configured to extract a set of speech features from each generated simulated natural speech sequence;and • a second speech synthesis module (22) configured to set up a speech synthesis engine according to the extracted set of speech features.;
9. Information device (6) according to claim 8, wherein the second speech synthesis module (22) is further configured to perform the synthesis of a speech announcement from a text sequence, the synthesis comprising providing the text sequence as input to the parameterized speech synthesis engine, a corresponding output of the parameterized speech synthesis engine forming the synthesized speech announcement.
10. Vehicle (2), in particular a rail transport vehicle, comprising an information device (6) according to claim 9 and at least one sound broadcasting device (4) connected to the information device to broadcast each synthesized voice announcement in the vehicle and / or in the vicinity of the vehicle.
11. Vehicle according to claim 10, further comprising at least one camera (8) connected to the information device (6), each camera (8) being configured to acquire an image of a predetermined volume of the vehicle, and to transmit the acquired image to the information device, the information device (6) further comprising an image processing module configured to determine, for each individual represented on the acquired image, at least one physical attribute of said individual, and to provide each determined physical attribute as input to the first speech synthesis module.