METHOD FOR DELIVERING DIGITAL MULTIMEDIA CONTENT IN REAL TIME FROM AN ADDRESSING FUNCTION AND TRANSLATION EQUIPMENT
The method and system address the challenge of simultaneous multimedia translation for multiple users by using software translation channels with language detection and routing, ensuring efficient and low-latency translation distribution.
Patent Information
- Application Number
- FR2022014274
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-12-22
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2042-12-22
AI Technical Summary
Existing technologies face challenges in simultaneously translating multimedia content for multiple users speaking different languages, due to complex routing configurations and translation latency, especially in setups with multiple speakers and interpreters.
A method and system for real-time multimedia content delivery using software translation processing channels that detect input language through learning functions, automatic transcription, and speech synthesis, combined with an addressing function to route content based on user language preferences, enabling seamless and instantaneous translation distribution.
The solution allows for almost instantaneous and accurate translation of multimedia content to multiple users by detecting input languages and routing streams efficiently, reducing latency and complexity in multi-speaker environments.
Smart Images

Figure 00000028_0000 
Figure 00000028_0001 
Figure 00000029_0000
Abstract
Description
Title of the invention: METHOD FOR DELIVERING DIGITAL MULTIMEDIA CONTENT IN REAL TIME FROM AN ADDRESSING FUNCTION AND TRANSLATION EQUIPMENT Field of the invention
[0001] The field of the invention is that of methods for routing digital content within a data network to a community of users. More particularly, the field of the invention relates to methods and systems for routing a simultaneous interpretation of digital content to a plurality of users in a plurality of languages. The field of the invention relates to methods comprising software means for processing dynamic routing of multimedia data according to a predefined user configuration. Prior art
[0002] There are various techniques for delivering multimedia content corresponding to an oral translation or interpretation of original content according to the user's language.
[0003] Among existing solutions, there is a first family of techniques aimed at producing automatic translation of digital content using a translation engine. According to these techniques, algorithms using dictionaries, ontologies, and thesauri can be used to perform automatic translations. Artificial intelligence can be configured so that learning improves the accuracy of automatic linguistic translations. However, these techniques are not satisfactory both in terms of the quality of the automatic translation produced and the limitations encountered in generating multilingual translations for multiple users simultaneously.
[0004] Among the existing solutions, there is a second family of techniques aimed at the oral translation or interpretation by one or more human translators, such as simultaneous interpreters, of audio content, and for which the data stream corresponding to the translated data is routed to the listener(s). These techniques generally involve an interpretation booth in which an interpreter receives an incoming audio stream and simultaneously produces an outgoing audio stream corresponding to the linguistic translation of the incoming content in the specified target language. When different audio sources of different languages are used When content is produced, for example, during a debate with different participants speaking different languages, there is a real difficulty in simultaneously producing translated content for different users who have chosen different listening languages. Current technologies do not allow for configuring a large number of listeners, each choosing their listening language, in a setup where multiple speakers can produce content in different languages.
[0005] One difficulty arises from the fact that the configuration is challenging to implement from the perspective of routing digital audio content, depending on the speaker, the interpreters involved, and the routing of translated content. Current solutions have the following main drawbacks: the complexity of the routing configuration to be implemented depending on the operational configuration of the speakers' languages, the listening languages, and the translators, and the translation latency introduced by the complexity of the systems routing the audio content.
[0006] There is a need to overcome the aforementioned disadvantages. Summary of the invention
[0007] According to a first aspect, the invention relates to a method for transmitting real-time multimedia digital content comprising: • Access by a first user to a digital service hosted on at least one data server from a first user terminal, said access including the selection of a first user language; • A first transmission of at least one first multimedia content via an addressing function to a first set of software translation processing channels to generate at least one second multimedia content corresponding to a first translation of said first multimedia content into the first user language, each software translation processing channel being associated with an input language and an output language; • at least one software translation processing channel comprising at least one computer and memory for implementing the following steps: • Input language detection; • Automatic transcription of the first multimedia stream to generate a first set of real-time data corresponding to the transcription according to a predefined alphabet and a given language; • Automatic translation of the transcribed and encoded text in the first machine-interpretable dataset, the automatic translation being carried out from the input language to the output language through the implementation of a learning function, said translation of the text being encoded in a second machine-interpretable dataset; • Speech synthesis of the text encoded in the second set of data interpretable by a machine to generate an audio stream corresponding to the second multimedia content; • The production of an activation signal by at least one software translation processing channel when a second multimedia content is generated by said at least one software translation processing channel; • Detection by the addressing function of at least one activation signal produced and determination of the language of the first multimedia content based on the activation signals detected; • A second transmission of the first multimedia content via the addressing function to the first user terminal if the language of the first multimedia content determined is the first user language; • A third transmission of the second multimedia content via the addressing function to the first user terminal if the language of the first determined multimedia content is the first user language.
[0008] One advantage is that it allows for almost instantaneous addressing of data streams due to the detection of a signal emitted from translation equipment. A key benefit of the solution is the ability to analyze translation booths or software translation processing channels to deduce the speakers' languages.
[0009] According to one embodiment, the detection of the input language is carried out from a probability calculation performed by means of a learning software function; said detection of a given input language activating an automatic transcription step.
[0010] According to one embodiment, the detection of the input language is carried out from a decoding of a metadata in the first multimedia stream received by the first software block, said metadata comprising a language selected by the sender of the first multimedia stream.
[0011] According to one embodiment, the detection of the input language is carried out from a second activation signal emitted by a rephrasing booth delivering a first multimedia stream repeated by an individual from the first multimedia stream, the generation of said first multimedia stream repeated by said individual allowing the generation of an input language identification data defining said second activation signal.
[0012] According to one embodiment, the detection of the presence or absence of each activation signal by the addressing function makes it possible to determine a first value of the first multimedia content, said first value corresponding to the identification of a language among a plurality of languages.
[0013] One advantage is that knowledge of the configuration of the booths or the software translation processing channels is sufficient to know the languages of the streams to be routed.
[0014] According to one embodiment, the deterministic position of a microphone in a reframing booth determines the generation of an activation signal, said first value being calculated from the deterministic position of a microphone in each set of software translation processing channels.
[0015] According to one embodiment, the method includes an automatic deactivation of software translation processing channels whose input language is not the language identified by the first value.
[0016] According to one embodiment, the addressing function comprises reading from a memory location each first user language associated with a terminal identifier, said identifiers being associated with a first user language when a user accesses the service, the routing of the first multimedia content or the routing of the second multimedia content being determined according to: • of the language of the first determined multimedia content and / or; • the output language of a software translation processing channel whose An activation signal has been generated.
[0017] One advantage is knowing the language configuration of each user connecting to the service in order to automatically send them content in the language that matches their configuration.
[0018] According to one embodiment, the addressing function is implemented by a virtual switch implementing the WebRTC, SIP or any other audio-visual-text communication protocol.
[0019] According to one embodiment, the method comprises recording the configurations of each software translation processing channel connected to the data network, each configuration of each software translation processing channel comprising an input language, an output language, a channel identifier Software-based translation processing, with the recording being performed by a software function known as a digital interpreting system. One advantage is that the language settings of each translation device or between each software-based translation processing channel also allow for the configuration of routing between the booths themselves or between the channels themselves to route all the data streams to all the booths, such as inputs and outputs that can be routed between the booths themselves.
[0020] In the case where the invention includes the implementation of a rephrasing step preceding the processing of the audio stream by the first software block, the input language configuration can be associated with the rephrasing booth. That is to say, a translator has selected an input language which has been stored in memory. This data can then be retrieved by the software processing chain.
[0021] According to one embodiment, the language detection block is located upstream of the software processing blocks so as to route the first multimedia stream to a given software processing chain comprising all the processing blocks. One advantage is to detect the input language in real time in order to route the first multimedia stream to the correct software processing chain.
[0022] According to one embodiment, the method comprises association data for two translation software processing channels stored in memory by means of the digital interpretation control system, said association data defining a translation booth, each translation software processing channel of the same booth having a communication interface for receiving a graphical interface on a display, said graphical interface representing a value of a digital signal indicator transmitted by one of the translation software processing channels to the other translation software processing channel of the same booth or by one of the translation software processing channels to another translation software processing channel. One advantage is to streamline the transitions between interpreters using the translation booths without latency.
[0023] According to one embodiment, a second user transmits the first multimedia content by means of a microphone and communication equipment.
[0024] According to one embodiment, the first multimedia content is transferred by means of a physical interpretation control room receiving a plurality of multimedia content from a plurality of secondary users and generating, at the output of said physical interpretation control room, the first multimedia content for a digital interpretation control room. An advantage is the use of the invention in a context of multi-audience conferences involving the implementation of a physical control room.
[0025] According to one embodiment, when an activation signal is not generated by a given software translation processing channel, the method comprises a selection of a second multimedia content corresponding to an output language of another software translation processing channel that generated the activation signal, said selection being carried out by means of a multimedia flow controller of said given software translation processing channel.
[0026] According to one embodiment, the calculation of said first value is, moreover, a function of: • the selection of at least one software processing channel for translating a second multimedia content via the controller or; • the presence of an activation signal of at least one software processing channel for translating a second multimedia content.
[0027] According to one embodiment, when a new user terminal requests access from the remote server, said method includes a memory in which is stored an identifier of said new terminal and a language selected via the new user terminal.
[0028] According to one embodiment, the microphone of the first user's terminal is controlled according to a broadcast configuration of the first transmission, the first transmission being of the unidirectional or bidirectional conference type.
[0029] According to one embodiment, the process includes a step of defining a new conference, a step of assigning a type of conference to said conference and a step of associating a set of second users with said conference, said second users having rights to broadcast a first multimedia content.
[0030] According to one embodiment, the method includes an actuator for disabling the software translation processing blocks in order to create group communication between users who have chosen as listening language a language selected by a given speaker of the user group.
[0031] According to another aspect, the invention relates to a system for transmitting real-time multimedia digital content comprising: • A plurality of multimedia acquisition terminals for a plurality of users, each multimedia acquisition terminal being associated with a predefined language; • A plurality of user terminals, each with an identifier and the choice of a predefined language; • Network equipment and a data server configured to implement the process of the invention. Brief description of the figures
[0032] Other features and advantages of the invention will become apparent from the following detailed description, with reference to the accompanying figures, which illustrate:
[0033] [Fig-1]: an embodiment of a system of the invention implementing a addressing function and translation equipment;
[0034] [Fig.2]: an embodiment of a system of the invention implementing an addressing function and a plurality of simultaneous interpretation equipment;
[0035] [Fig.3A]: a first embodiment of the process of the invention according to a first value of the activation signal;
[0036] [Fig.3B]: the first embodiment of the process of the invention according to a second value of the activation signal;
[0037] [Fig.3C]: the first embodiment of the process of the invention according to a third value of the activation signal;
[0038] [Fig.4]: a second embodiment of the method of the invention according to a configuration of simultaneous interpretation equipment,
[0039] [Fig.5]: an example of a graphical interface of a simultaneous interpretation equipment according to an embodiment of the invention.
[0040] In the remainder of this description, an "interpreter" or an "oral translator" is referred to interchangeably as a translator, an oral translator, or an interpreter. Interpretation equipment or translation equipment is designated by the notation "CAB". The terms "translation equipment" and "interpretation equipment" are used interchangeably. Speakers
[0041] Figure 1 represents an embodiment of the system of the invention in which a set of microphones MIC1 and MIC2 correspond to speaker microphones UA and UB speaking respectively in a first language FR (French) and a second language EN (English). The speakers are users such as on-site presenters or people participating in remote meetings.
[0042] According to another embodiment, UA, UB users are users such as persons communicating within a bidirectional network by means of electronic terminals such as a computer, a tablet, a smartphone or any other equipment comprising a user interface and a communication interface enabling it to be connected to a Net data network.
[0043] According to another embodiment, the FR and EN input streams are multimedia content of audio, text, or video type. The multimedia content can be encapsulated within data frames of a transfer protocol of real-time data suitable for data exchange over a data network, for example the Internet.
[0044] The multimedia content delivered by a user device, such as the MICi or MIC2 microphone of a speaker UA or UB, is denoted FL, denoting the term "FLOOR" commonly used to define the default selected channel of an audio device transmitting an audio stream to a network or a physical control room. In the context of the invention, the multimedia stream FL refers to a multimedia stream transmitted by a device such as a MICi microphone, a stream from an electronic terminal such as a smartphone or computer, or a stream from a device centralizing different channels such as a physical control room collecting a plurality of audio channels and delivering an output stream.
[0045] In the example of [Fig. 1], the FL multimedia stream is transmitted across a NET data network, which can be a local area network (LAN) or an open network such as the internet. The FL stream can be transmitted using any communication protocol that allows the exchange of data frames over a data network. Preferably, a real-time protocol is implemented so that the data frames are transmitted with minimal latency. Access to the service by an auditor
[0046] In the context of the invention, a user Ui wishing to access the service implemented by means of the invention registers with an identification service, for example, supported by a network infrastructure comprising equipment such as a SERVi data server. The SERVi server includes means such as a computer and memory to allow a user to define, at a minimum, an identifier and a listening language Lb. According to one implementation, the SERVi server is the same server as the server hosting the addressing function and / or the digital interpretation control function.
[0047] A user Ui will, a priori, retain the listening language Li throughout the service regardless of the speaker UA, UB, and their respective speaking language FR or EN on [Fig. 1]. The invention notably allows for a smooth and instantaneous switching of the routing of translated content for each speaker. Digital content
[0048] The listening language corresponds to the language in which the user Ui wishes to receive the multimedia content corresponding to a given digital content.
[0049] The determined digital content may be, for example, in the case of audio content: • a conference or debate with one or more speakers located in the same place and broadcast in real time; • a conference or debate with one or more speakers located in different geographical locations and broadcast in real time; • a conference or debate by one or more speakers, recorded and broadcast on demand; • a telephone conversation with a pre-defined person, • a telephone conversation with a plurality of pre-defined people, • etc.
[0050] The invention also applies to a specific digital content, whether textual or video.
[0051] Figure 1 represents a user's device Ti in the form of an electronic terminal, more broadly referring to a smartphone, personal computer, tablet, or any other device that enables digital communication over a data network. The device is configured to receive a digital stream, denoted TR or FL. However, the invention is not limited to configuring a unidirectional stream but can be applied so that the user Ui is also a provider of multimedia digital content. For example, the stream transmitted by the user Ui could be an audio stream, text, or video, or a combination thereof.
[0052] Oral translation or interpretation equipment
[0053] The invention includes CAB means for implementing the generation of a second translated multimedia content TR from a first initial multimedia stream FL. In the following description, such a means is referred to as CAB translation equipment. Generally, in the description, the stream produced by a speaker is denoted as a first multimedia stream FL or first multimedia content, and the stream produced by CAB translation equipment is denoted as a second multimedia stream TR or TRL [or second multimedia content TR or TRLi]. The notation TR is used to designate the second multimedia stream in the general case, and the notation TRM is used to designate the second multimedia stream in the case where the Li language is also designated.
[0054] Each CAB translation device is associated with an input language LE and an output language Ls. Each CAB translation device produces an activation signal SCab or not indicating the generation of multimedia content TR corresponding to the translation of input multimedia content FL.
[0055] In its simplest configuration, the SCab activation signal is the multimedia content TR generated at the output of the equipment, which is routed through the data network to an addressing function SW. In this latter case, a translation equipment CAB producing no output signal, and therefore no translation, produces no SCab- activation signal.
[0056] In other configurations, a SCab activation signal different from the TR multimedia content is produced, for example, from a separate signal corresponding, for example, to the switching on of electronic equipment such as a microphone. According to another example, the SCab activation signal is generated from a metadata frame embedded in a message header encapsulating the digital encoding of the translated TR digital content.
[0057] The SCab activation signal is therefore generated each time a CAB translation device produces an output signal corresponding to a translation of an input signal. Thus, the activation signal can be generated each time an electronic device, such as a microphone in the translation device, is activated, or each time a stream is produced.
[0058] These means may be, for example: - a physical booth in which a human translator receives, via a communication interface, an incoming audio stream and provides an outgoing audio stream rerouted within the NET network from a communication interface; the physical booth can also be understood as a simple audio system for delivering sound to a loudspeaker and acquiring sound via a microphone; - an automatic translation machine comprising a software component configured to perform automatic audio translation from speech recognition software and instant translation software;
[0059] This solution can be combined with a speech synthesizer to reconstruct a translated audio stream. It is called a "software translation processing channel", - an automatic translation machine comprising a software component configured to perform automatic textual translation from instant translation software.
[0060] Each CAB translation equipment can therefore be designed to operate automatically using software and hardware means or to operate with a translator or interpreter and a human-machine interface.
[0061] When a plurality of CAB translation equipment is configured within the framework of the invention, different embodiments can be implemented.
[0062] According to a first embodiment, the CAB translation equipment can be co-located in the same location. This equipment can then include an interface with a physical RPI interpretation control room allowing the different streams of each piece of equipment to be referenced according to the input and output languages of each piece of equipment.
[0063] According to a second embodiment, each CAB translation unit is connected directly to the NET data network by means of an interface communication. In this last embodiment, a digital RDI interpretation control system allows the different streams of each translation equipment to be referenced according to the input languages LE and output languages Ls of each of these equipment.
[0064] According to a third embodiment, a first set of CAB translation equipment is directly connected to the NET data network, and a second set of CAB translation equipment is co-located together in the same geographical location. Each CAB translation device in this second set has a physical interface with an RPI physical interpretation control room in order to reference the various data streams from each CAB translation device co-located in the same location. In this third embodiment, the RPI physical interpretation control room exchanges data with the RDI digital interpretation control room so that the RDI digital interpretation control room references all the CAB translation equipment.
[0065] According to one embodiment, at least one CAB translation equipment comprises three software blocks and computing means such as a computer, memory and communication interface enabling the processing of an incoming audio stream in a first language LE and the generation of an outgoing audio stream in an output language Ls.
[0066] A first software block includes a transcription function that processes the incoming audio stream LE to generate text. The text can be generated as a real-time data stream as the voice transcription is processed. The text is preferably formatted in a machine-interpretable language, such as a computer-implemented software function capable of storing said text in memory, for example, in a temporary file, in order to perform operations on said data.
[0067] According to one embodiment, the first block comprises an acoustic processing step to extract acoustic vectors from the speech signal, corresponding to 20- to 30-ms signal segments. The signal is digitized and parameterized using a frequency analysis technique based on the Fourier transform. A machine learning step establishes an association between the elementary speech segments and lexical elements. Various techniques can be implemented. For example, this association can be achieved through statistical modeling using artificial neural networks. The first software block further comprises a decoding step that concatenates the previously learned elementary segments to reconstruct the most probable text. A dynamic temporal warping algorithm can be used for this purpose.
[0068] According to one embodiment, the first software processing block is preceded by a rephrasing booth. The rephrasing booth is a booth in which a An individual reads back the incoming audio stream using a microphone in the input language (IL). The audio stream processed by the individual is then spoken back into the same input language (IL) and fed back into the first processing block. One advantage of this rephrasing process is that it reduces the number of errors related to differences in reading or speaking rate, accents, timbre, articulation, etc.
[0069] According to one embodiment, the rephrasing booth is an audio natural language processing booth that regenerates an audio stream with a normalized voice for processing by the first software block. The automated processing can be performed using dictation software that reads back the content of an input decoded audio text in the same language.
[0070] A second software block includes the implementation of a function aimed at taking as input the text output generated by the first block and applying an operation implemented by an artificial intelligence algorithm to generate a translation of said text into an output language.
[0071] The artificial intelligence algorithm is preferably a learning function that has been learned by means of a training dataset. The algorithm can be a network model of the LSTM, RNN type or a Transformer such as a Generative Pre-trained Transformer 3, whose acronym is GPT-3, implementing an autoregressive language model that uses deep learning to produce a text.
[0072] Within the framework of the invention, a learning function is configured with a memory and a memory depth enabling it to learn on ordered data sequences such as language.
[0073] The second software block therefore makes it possible to generate in real time an automatic translation of the input text, resulting in the production of a second translated text as output from the second software block.
[0074] The third software block processes as input text translated in real time by the second software block. The input text is in the form of a data stream produced by the second block and addressed to resources capable of executing the third software block. These resources are preferably computer-implemented software, computers, and memory.
[0075] The third software block also relies on a speech synthesizer-type device capable of generating an audio stream from text. In the implementation of this third block with the second software block, the speech synthesizer generates an audio stream from the translated text produced by the second software block. The speech synthesizer implements a component enabling linguistic processing, in particular to transform the orthographic text into a phonetic version that can be pronounced without ambiguity, and components allowing this phonetic version to be transformed into a digitized sound that can be listened to on a loudspeaker.
[0076] According to one example, the first software block includes a function for automatically detecting the input language (IL) of the translation equipment. This function can be implemented using software means such as a computer and memory. The function can be a machine learning algorithm, such as a neural network or any other architecture of a learning function learned with a training dataset. The function generates a probability that the input audio stream is a given language. A classifier or a statistical distribution of the detection indexes a probability for each language. This learning function is preferably learned with a set of languages and with samples of different subjects / individuals.The phoneme or phrase samples used during training allow for a heterogeneous training dataset. The samples may consist of phrases or phonemes from individuals of different ages, genders, timbres, or speech rates, etc.
[0077] When the input language LE is automatically detected by calculating the probabilities of each language and selecting the language with the highest probability, the first software block automatically activates the transcription of the incoming audio stream. Thus, each software block is automatically activated to generate an output audio stream corresponding to a translation of the incoming audio stream. Since the output language Ls is already configured in the translation equipment, detecting the activation of the software blocks allows the input language LE to be deduced.
[0078] The output language Ls acts as an activation signal that can be transmitted to other entities such as servers or computers. The activation signal is associated with an input language LE and an output language {LE, Ls} pair. This pair allows another component to know the input language LE of the first software block.
[0079] This configuration not only allows the input language LE information to be transmitted to other equipment but also allows other automatic translations to be synchronized with the information which is the input language LE.
[0080] According to one embodiment, this input language detection information from a translation equipment can be communicated to other translation equipment automatically in order to corroborate the statistical analysis of language detection.
[0081] The input language LE detection information also allows the Floor, i.e. the input language, to be automatically routed to users who have configured their audio input to listen to the stream in the same language as the input language LE.
[0082] The output stream from the third software block can also be routed to the input of another translation equipment which has a configuration or input language of that equipment which corresponds to the output language of the translation equipment which produced the translated audio stream as output.
[0083] According to another embodiment, the input language LE can be decoded in the header of an incoming data stream including the audio stream. In this case, the method of the invention makes it possible to directly activate a translation into a predefined output language if the equipment and / or the software configured on the equipment is configured for translation of said input language and the output language.
[0084] According to another embodiment, the input language (LE) can be selected from an indicator. The indicator can be, for example, metadata from the incoming stream in the first processing block or in the rephrasing booth. The metadata can, for example, be inserted from a language selection by a user generating the first FL multimedia content. One advantage is that the input language can be immediately decoded within the CAB software translation processing channel.
[0085] According to another embodiment, the input language is determined by the activation of the rephrasing booth. Indeed, if an individual regenerates an incoming audio stream in the input language LE from the first FL multimedia content entering their rephrasing booth, the input language LE can be automatically detected. Addressing function
[0086] The invention comprises the implementation of an addressing function SW. This latter function SW can be implemented by any type of machine such as a data server, or as a distributed function implemented by a plurality of machines. In the example of [Fig. 1], a server denoted SWi is configured to implement the addressing function SW and also the RDI digital interpretation control function.
[0087] The SW addressing function is configured to transmit at least one first FL multimedia content to a first set of CAB translation equipment to generate at least one second TRL multimedia content corresponding to a first translation of said first FL multimedia content into the first user language Lp. The SW addressing function, in combination with the RDI digital interpretation control system, automatically routes the translated TR content. to the equipment of the first user Ui whose language Li is stored in memory and known to the SW addressing function during its implementation.
[0088] The SW addressing function is configured to detect the presence or absence of each SCab activation signal from each of the CAB translation equipment in order to initiate TR translated content routing actions to users who have chosen an output Ls language corresponding to the language of the translated content.
[0089] To this end, the SW addressing function is configured to detect the presence of an SCab activation signal emitted by each CAB translation device. When the SW activation function detects an outgoing stream from a first CABi translation device by the presence of an SCab activation signal, a calculator is configured to determine the input language LE of the CABi translation device from the RDI digital interpreting system. Indeed, the RDI digital interpreting system is aware of each CAB translation device connected to the NET network and its translation configuration, which includes, in particular, the definition of an input language LE and an output language Ls. The detection of an SCab activation signal makes it possible to determine the input language LE of the first active CABi translation device and therefore the language of the multimedia stream FL emitted by a user UA or UB.
[0090] The method of the invention therefore makes it possible to convey: - the outgoing TR stream to users who have chosen a listening language Li corresponding to the output language Ls of the active CABi translation equipment; - the FL stream sent to users who have chosen the LE input language of the active CABi translation equipment; - the outgoing TR stream from the first active CABi equipment to at least one other CAB2 translation equipment whose input language LE corresponds to the output language Ls of the first active CABi translation equipment.
[0091] Figure 2 illustrates an embodiment in which different CABI devices, CAB2, CAB3, CAB4, and CAB5, are represented. Functionally, the set of translation devices is denoted {CABi}, with each CAB device connected to the NET data network from any network access point.
[0092] Thus, the UER user who has chosen a French language receives: • either directly the FL stream when a speaker UA activates their microphone MICie and produces a stream in the French language identical to that of the user UER. In this latter case, the SW addressing function, by detecting the active CAB3 translation equipment because TR content is being emitted by it, deduces that the input language LE of this CAB3 equipment is French thanks to the RDI digital interpretation agency which has knowledge of the configuration of the translation equipment {CAB}; • either the TRfr stream generated by the CABi translation equipment, which corresponds to a translation of the first FL multimedia content originating from the MIC2 microphone of a speaker producing initial multimedia content in the GB language. In this latter case, the SW addressing function, by detecting the activity of the CAB3 translation equipment, which is active because TR content is being emitted by it, deduces that the input language LE of this CAB3 equipment is FR. This deduction is also made thanks to the RDI digital interpretation control room, which has knowledge of the configuration of the {CAB} translation equipment.
[0093] Thus, for each UFR and UFN user, the two streams FL and TRfr for the UFR user and respectively FL and TRFN for the UFN user are represented on [Fig.2], although these streams are not received simultaneously, but according to a sequence depending on the language of the speaker UA or UB.
[0094] Figure 2 illustrates the case where a user UCn with a T3 terminal, having configured a CN listening language when accessing the service from the SERVi server, receives a TRCN stream from a CAB5 translation device. In this scenario, there is no translation device configured with a FR input language and a CN output language. The invention allows for the automatic routing, using the SW addressing function, of either the content produced by the CAB5 translation device when the speaker transmits a FL multimedia stream in GB, or the stream produced by the CAB5 translation device receiving the translated stream from the CAB3 device itself receiving the stream from the first speaker UA when the speaker's language, corresponding to the FL content, is FR.
[0095] The digital platform and equipment enabling the implementation of the addressing function and the functions allowing the use of the invention's process by multiple users are referred to as the "iBP digital system". The addressing module is a software component enabling the implementation of the addressing function.
[0096] The addressing function inherent in the system also allows the redistribution of multiple streams in a multitude of cases:
[0097] 1 / In the first case, it allows interconnection with physical platforms. In this first case, the SW addressing function interconnects the digital platform with a physical interpretation platform via a single PC connected to the internet. Within this PC, a flow addressing module enables the interconnection of the iBP digital system with an external physical system, via a card. An external sound card and a virtual sound card installed in this PC. This virtual sound card addresses the various digital channels carried by the stream addressing module to the different ports of the external sound card. Each stream is addressed in both directions: input and output.
[0098] 2 / In a second case, the IBP platform's flow addressing module allows the interconnection of streams from other digital video conferencing platforms.
[0099] 3 / In a third case, the incoming flow addressing module, also called "Switcher" allows, from a single sound source, the distribution of the stream to all booths or to all CAB translation software processing channels and all users, via a single PC connected to the room's sound system and the Internet.
[0100] 4 / In a fourth case, the central system sends information from all the Streams corresponding to multiple conferences of the same event, on a single terminal. This addressing allows the selection of the stream by room, by day and by language, as well as its delayed listening.
[0101] Thus, the invention allows a large capacity for configuration according to the language of the speakers and the languages chosen by users wishing to access the service.
[0102] Fig. 3A represents an embodiment in which the addressing function SW directly routes the first FL multimedia stream produced by a speaker to the terminal Ti of the user Up. This routing of the first FL multimedia content is carried out following the non-detection of output streams from the CABi to CABn translation equipment.
[0103] In the case of [Fig. 3B], the routing of the second multimedia stream TR to a terminal Ti of a user who has selected the French language (FR) is determined by the addressing function through the detection of an outgoing stream from the CABi translation equipment representing the activation signal SCab i of the CABI equipment. The addressing function SW, through the RDI digital interpretation control system, knows that the CABI translation equipment is activated and that, consequently, the language of the first multimedia content FL is English (GB). This detection ensures: • routing of the second multimedia stream TRfr from the CABI translation equipment to users who have chosen the FR language as well as to CABi translation equipment configured with an input language Le corresponding to the GB language and;
[0104] routing of the first FL multimedia content directly to users who have chosen the GB language.
[0105] Figures 3A and 3B illustrate that the outgoing streams from the other CAB translation equipment are cut off and that, consequently, the SW addressing function does not detect any SCab activation signal at the output of said translation equipment. CAB. This detection also allows us to deduce that the input language LE of each disabled CAB translation device is not the language of the first FL multimedia content.
[0106] Figure 3C illustrates a mode in which a new translation device CAB2 produces an output stream in French (FR). The addressing function SW then determines the input language LE of the translation device CAB2, which corresponds to German (DE). The addressing function SW then initiates the routing of the first multimedia content FL directly to all users who have selected German (DE) and initiates the routing of the output stream from the translation device CAB2 to devices that have selected French (FR) and, if necessary, to CAB translation devices whose input language is French (FR).
[0107] Figure 4 illustrates a case in which a CABi translation device configured with an input language LE = FIN and an output language Ls = DE, whose output content TRDE is routed to the input of a CAB2 translation device configured with an input language LE = DE and an output language Ls = FR. The SW addressing function thus allows the multimedia content to be automatically routed from the output of the translation devices based on the analysis of the active languages of the first FL multimedia content, deduced from the deterministic states of each CAB translation device, and according to the RDI digital translation control system, which is aware of the configurations of the CAB translation devices.
[0108] Translation booth, translation software processing channel
[0109] The method of the invention also includes software means for configuring translation booths or software translation processing channels. When translators occupy booths, these booths can be formed from a plurality of CABi translation devices, each configured with the same input language (LE) and the same output language (Ls). When these booths are software translation processing channels, software means comprising at least one computer and memory can be configured to acquire the first multimedia stream (FL) and process it using different software blocks.
[0110] A "translation booth" or translation equipment is more commonly referred to in this description as: - a physical booth in which a translator can generate a new audio stream from the first FL or multimedia stream; - a software translation processing channel that automatically generates an audio stream in an output language corresponding to a translation of the incoming stream received in an input language.
[0111] Consequently, in this description the term "translation booth" or the term "translation equipment" may be replaced interchangeably by "software translation processing channel".
[0112] According to a preferred mode, a translation booth comprises two translation units [CABi, CABk] configured with the same input language Le and output language Ls settings. One advantage of such a configuration is that it allows the translator to be handed over after a certain time when the CAB translation units are intended to generate an audio multimedia stream from an incoming audio multimedia stream in a first language. The CAB translation unit advantageously includes a sound card, a speaker, and a microphone to provide the necessary interfaces for a translator to listen to and record audio multimedia streams. Numerical indicator of transfer
[0113] The invention makes it possible to implement a digital INDP handover indicator between the two translation units [CABi, CABk] of the same booth, which is automatically generated following an action by one of the translators. One advantage is the ability to generate a second multimedia stream [TR] corresponding to a translation of a first multimedia content [FL] during periods requiring a seamless and continuous handover of translators without any deterioration in the quality of service, particularly with regard to the transition time due to a change of translators.
[0114] The INDP digital handover indicator allows a given translator in a booth to anticipate the continuation of a translation already started by a first translator. The second outgoing multimedia streams TR from the CAB translation equipment are naturally routed to users via the SW addressing function. Indeed, when the new CAB translation equipment becomes active, it generates an SCab activation signal that allows the second output multimedia content TR to be routed to users speaking the same language.
[0115] The INDP digital handover indicator allows a signal to be routed between translators in the same booth so that they can, for example, start producing a translation before turning on their microphone so that the transition has no consequences for the listeners.
[0116] The transmission of the INDP digital transmission indicator can be ensured, for example, by the SW addressing function and the RDI digital interpretation control system, which is configured to know the configurations of translation equipment and therefore their association within booths. According to another example, a digital indicator INDP transmission is issued via a dedicated messaging system between translators in the same booth.
[0117] The INDP digital handover indicator can take different values depending on which cabin translator generates it and depending on the activation signal status of the CAB translation equipment considered.
[0118] The CAB translation equipment may include a communication interface and a display for showing the different values of the INDP digital transmission indicator. Figure 5 shows an example of a generated graphical interface for displaying the different values of the INDP digital transmission indicator for a translator in a booth. The graphical interface can be generated from a smartphone, tablet, or computer, or even from dedicated electronic equipment. In this example, it is assumed that the first translator is translating and generating a second outgoing multimedia stream, TR, and that the second translator takes over the ongoing translation to replace the first translator.
[0119] A first value of 12 informs the first translator in the booth that the second translator is ready to take over. A value of 13, on the other hand, informs the second translator that they are not yet ready. This message is preferentially sent by the second translator from their interface. When the first translator receives the active value of 12 from the first indicator sent by the second translator, they can initiate the translation switch by changing the value of 10 in the INDP digital handover indicator. The second translator is then notified and begins translating the audio content of the incoming FL stream to generate an outgoing TR stream.
[0120] According to one embodiment, the INDP numerical passing indicator is a composite indicator insofar as it allows different values to be exchanged between translators in a bidirectional manner.
[0121] According to one embodiment, the graphical interface provided within the CAB translation equipment of a second translator includes a dual listening function that allows said second translator to listen to the first multimedia stream FL simultaneously with the stream of the first translator sharing the same booth. This function can be activated by means of a selector button 20, which can be activated from the graphical interface. This function allows for a smoother transition during the handover from the first to the second translator and therefore better routing of the second multimedia stream TR transmitted to the user without degradation of the transition delay.
[0122] According to one embodiment, a user of the method and system of the invention may have an interface for communicating with other users. This interface This allows the user to choose the translation language. This choice of translation language allows the user to select a CAB translation software processing channel whose output language is the desired language.
[0123] Thus, in this embodiment, when the input language is detected, the multimedia content processed by a CAB software translation processing channel can be routed to user terminals that have selected the translation language. According to one embodiment, the detection of an activation signal allows the content produced by the software translation processing channel to be routed directly to terminals that have selected the corresponding output language. In this case, the activation signal corresponds, for example, to the simple fact that a processing channel is active.
[0124] According to one embodiment, the user interface includes an actuator to allow the definition of a subgroup of users. The actuator can be a graphic symbol displayed on a computer, tablet, or smartphone screen, activated by a pointer or touch control.
[0125] In this embodiment, a user can define a subgroup of users who have all chosen the same listening language, for example French, corresponding to the language of the multimedia stream's sender. One advantage is to move away from the existing global user group exchanging information in multiple languages with automatic translations of the exchanged streams, and instead enable communication within a subgroup of users sharing the same language.
[0126] This feature allows direct communication within a subgroup of users without implementing software translation processing channels. This mode automatically disables translation of the stream emitted by the speaker.
[0127] When this mode is activated, the user interface includes, according to an example embodiment, a button allowing the user to return to a communication mode with all users listening to the stream transmitted by one of the speakers. This mode automatically activates the translation of the stream transmitted by the speaker.
[0128] More generally, the user interface may include an actuator such as a numeric button to activate or deactivate automatic translations of the stream generated by a speaker according to each user's language configuration.
Claims
1. Demands A method for delivering real-time digital multimedia content (TRLi FL) comprising: Access (ACCESS) by a first user (Ui) to a digital service hosted on at least one data server (SERVi, RDI) from a first user terminal (Ti), said access including the selection of a first user language (LJ; • A first transmission (DIFFi) of at least one first multimedia content (FL) by an addressing function (SW) to a first set of software translation processing channels (CAB) to generate at least one second multimedia content (TR, TRL[) corresponding to a first translation of said at least one first multimedia content (FL) into the first user language (LJ, each software translation processing channel (CAB) being associated with an input language (LE) and an output language (Ls); • the following steps implemented by at least one software translation processing channel (STC) comprising at least one computer and one memory: • Input language (IL) detection; • Automatic transcription of the first multimedia content (FL) to generate a first set of real-time data corresponding to the transcription according to a predefined alphabet and a given language; • Automatic translation of the transcribed and encoded text in the first machine-interpretable dataset, the automatic translation being performed from the input language (Le) to the output language (Ls) through the implementation of a learning function, said translation of the text being encoded in a second dataset
2. machine-interpretable data; • Speech synthesis of the text encoded in the second set of data interpretable by a machine to generate an audio stream corresponding to the second multimedia content (TR, TRLi); • Production of an activation signal (SCab) by at least one software translation processing channel (CAB) when a second multimedia content (TR, TRLI) is generated by said at least one software translation processing channel (CAB); • Detection by the addressing function (SW) of at least one activation signal (SCab) produced and determination of the language (L1D) of the first multimedia content (FL) as a function of at least one activation signal detected; • A second transmission (DIFF2) of the first multimedia content (FL) by the addressing function (SW) to the first user terminal (TJ if the language (L1D) of the first multimedia content (FL) determined is the first user language (Li); • A third transmission (DIFF3) of the second multimedia content (TR, TRL[) produced by a software translation processing channel (CAB) whose activation signal (SCab) has been generated and detected by the addressing function (SW) to the first user terminal (TJ if the output language (Ls) of the second multimedia content (TR, TRLi) is the first user language (Li). A method according to claim 1, characterized in that the detection of the input language (IL) is carried out from a probability calculation performed by means of a learning software function; said detection of a given input language (IL) activating an automatic transcription step.
3. Method according to claim 1, characterized in that the detection of the input language (IL) is carried out from a decoding of a metadata in the at least one first multimedia content (ML) received by a first software block, said metadata comprising a language selected by a transmitter of the at least one first multimedia content (ML).
4. A method according to claim 1, characterized in that the detection of the input language (IL) is carried out from a second activation signal emitted by a rephrasing booth delivering a first multimedia stream repeated by an individual from the first multimedia content, the generation of said first multimedia content repeated by said individual allowing the generation of an input language (IL) identification data defining said second activation signal.
5. Method according to claim 1, characterized in that the detection of each activation signal (SCab) by the addressing function (SW) makes it possible to determine a first value of at least one first multimedia content (FL), said first value corresponding to the identification of the language (Lu,) of at least one first multimedia content (FL) among a plurality of languages.
6. Method according to claim 5, characterized in that the deterministic position of a microphone in a reframing booth determines the generation of an activation signal, said first value being calculated from the deterministic position of a microphone in each set of software translation processing channels (CAB).
7. A method according to any one of claims 5 or 6, characterized in that it comprises an automatic deactivation (DESACT) of the software translation processing channels (CAB) of the first set of software translation processing channels (CAB) whose input language (LE) is not the language identified (L1D) by the first determined value.
8. The method of claim 1, characterized in that the addressing function (SW) comprises reading from a memory (RDC, SERVi) each first user language (Li, L2) associated with a terminal identifier (Tb T2), said identifiers being associated with a first user language (Lb L2) upon access to the service (ACCESS) by a user (Ui), a routing due to the less a first multimedia content (FL) or a routing of the second multimedia content (TR, TRLi) being determined according to: • the language (L1D) of the at least one first multimedia content (FL) determined and / or; • the output language (Ls) of a software translation processing channel (CAB) whose activation signal (SCab) has been generated.
9. Method according to claim 1, characterized in that the addressing function (SW) is implemented by a virtual switch implementing the WebRTC, SIP or any other audio-visual-text communication protocol.
10. A method according to claim 1, characterized in that the method comprises recording configurations of each software translation processing channel (STC) connected to a data network (NET), each configuration of each software translation processing channel comprising an input language (IL), an output language (IL), an identifier of the software translation processing channel (STC), said recording being carried out by a software function called a digital interpretation control (DTC).
11. Method according to claim 1, characterized in that a second user (UA UB) transmits at least one first multimedia content (FL) by means of a microphone and communication equipment.
12. A method according to claim 1, characterized in that when an activation signal (SCab) is not generated by a given software translation processing channel (CABi), the method comprises a selection of a second multimedia content (TRL[, TRL2) corresponding to an output language (Ls) from another software translation processing channel (CAB2) which generated the activation signal (Scab), said selection being carried out by means of a multimedia flow controller (TRL[, TRL2) of said given software translation processing channel (CABi).
13. A method according to claim 5, characterized in that the calculation of said first value is, furthermore, a function of: • the selection of at least one software translation processing channel (CAB2, CAB2) of a second multimedia content (TRLi, TRl2) via a controller or; • the presence of an activation signal (SCab) of at least one software translation processing channel (CAB2, CAB2) of a second multimedia content (TRM, TRL2).
14. A method according to claim 1, characterized in that when a new user terminal (T2) requests access from a remote server (SERVi), said method comprises a memory in which is stored an identifier of said new terminal (T2) and a language selected via the new user terminal (T2).
15. Method according to claim 1, characterized in that a microphone of the first user's terminal (Ui) is driven according to a broadcast configuration of the first transmission (DIFF1), the first transmission (DIFF1) being able to be of the unidirectional or bidirectional conference type.
16. A method according to claim 1, characterized in that it comprises an actuator for disabling software translation processing blocks in order to create group communication between users who have chosen as their listening language a language selected by a given speaker in the user group.
17. A system for delivering real-time digital multimedia content comprising: • A plurality of multimedia acquisition terminals (MICi, MIC2) of a plurality of users, each multimedia acquisition terminal (MICi, MIC2) being associated with a predefined language; • A plurality of user terminals (Tb T2) having an identifier and the choice of a predefined language; • Network equipment (SW) and a data server (SERVI) configured to implement the method of any one of claims 1 to 15.