A computer-implemented method for communicating inherent meaning between users
The method optimizes language model selection based on user profiles to enhance real-time multilingual communication by adapting translation to user characteristics, addressing latency and resource inefficiency in current machine translation systems.
Patent Information
- Application Number
- PCT/GB2025/050371
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-12
- Filing Date
- 2025-02-26
- Publication Date
- 2025-10-16
AI Technical Summary
Current machine translation systems suffer from high latency, resource inefficiency, cultural insensitivity, and inability to adapt to user-specific comprehension characteristics, leading to miscommunication and misunderstandings in real-time multilingual interactions.
A computer-implemented method that selects an optimal language model based on user profiles and characteristics to translate input phrases in real-time, using bespoke or off-the-shelf transformer-based language models, optimizing computational resources and reducing latency.
Enables efficient, real-time multilingual communication that adapts translation to user cognitive and linguistic profiles, improving accuracy and reducing latency by selecting the most suitable language model for each interaction.
Smart Images

Figure GB2025050371_16102025_PF_FP_ABST
Abstract
Description
[0001] A COMPUTER-IMPLEMENTED METHOD FOR COMMUNICATING INHERENT MEANING BETWEEN USERS
[0002] FIELD OF THE INVENTION
[0003] The present invention relates to a computer-implemented method for communicating inherent meaning between users, and, more specifically, to a computer-implemented method for communicating inherent meaning between two or more users having different language requirements while optimising computational resources and / or reducing latency.
[0004] BACKGROUND TO THE INVENTION
[0005] Language Models, particularly Large Language Models (LLMs), have significantly impacted the field of multilingual machine translation. They offer flexible, scalable, and highly effective solutions. Their application in machine translation demonstrates the power of deep learning in bridging linguistic divides and facilitating global communication. However, current systems do not optimise processing pipelines for real-time communication, which results in higher latency and resource consumption. In addition, current machine translation (MT) is plagued by cultural insensitivity, which occurs when algorithms, no matter how advanced, fail to recognize and respect cultural nuances (Akorbi, 2023). As a result, misinterpretations and unintentional perpetuation of stereotypes can occur. This limitation results in miscommunication, misunderstandings, and a lack of empathy, posing challenges in both personal and business contexts.
[0006] Traditional tools for translation, whether human-operated or technology-driven, frequently necessitate content simplification, especially when dealing with complex emotional or technical nuances. This often leads to a diluted exchange of ideas, cultural insensitivity, and inaccurate contextual interpretations (Harzing & Feely, 2007; da Costa, 2021). For example, context remains a significant challenge for machine learning algorithms, as recognizing context is an area where machine learning still faces difficulties. Current tools also lack adaptability to user-specific comprehension characteristics, which leads to semantic inaccuracies and communication breakdowns. As such, there exist significant differences in the accuracy and nuances captured across different currently available translation tools. There is an increasing need to improve real-time communication between two or more people, each using different language that is not easily understood by the other party, while optimising computational resources and reducing latency. This difference in language may arise due to certain characteristics of a person, including their native language, dialect and age to name but a few. It is against this background that the present invention has arisen.
[0007] SUMMARY OF THE INVENTION
[0008] According to the present invention there is provided a computer-implemented method for communicating inherent meaning between users, the method comprising: receiving, from a first user, an input communication comprising an input phrase in a source language, wherein the input phrase has an inherent meaning; identifying a second user profile comprising a plurality of characteristics of a second user; selecting, from a plurality of language models, an input phrase translating language model for translating the input phrase, wherein the input phrase translating language model is selected based on the input phrase and the second user profile; generating, using the input phrase translating language model, an output phrase having the same inherent meaning as the input phrase, wherein the output phrase is in a target language that is different from the source language, and wherein the target language is based, at least in part, on the plurality of characteristics of the second user; generating an output communication comprising the output phrase; and providing the output communication to the second user.
[0009] Selecting the input phrase translating language model based on the input phrase and the second user profile enables the processing power required to translate the input phrase to be optimised. For example, the plurality of language models may comprise a plurality of bespoke language models. Alternatively, or in addition, the plurality of language models may comprise a plurality of readily available / off-the-shelf language models, such as 'Chat GPT' and / or 'Gemini'. A plurality of the language models, whether bespoke and / or off-the-shelf, may be transformer-based language models. More specifically, a plurality of the language models may be transformer-based large language models.
[0010] Each language model, particularly the bespoke languages models, may be optimised for a specific type of input phrase and / or user profile. This enables the most efficient language model for a given input phrase and user profile to be selected. It also enables at least part of the method to be carried out using smaller, less performant hardware. As such, the present method enables real-time multilingual communication that dynamically adapts translation of the input phrase to the cognitive and linguistic profile of the second user while optimizing computational resources, thus resulting in tangible improvements in computer operation. The method may also reduce latency by selecting the most optimal language model for this purpose.
[0011] In this application, 'phrase' means one or more words that, in the context of the input communication, have meaning. As such, a phrase may be a single word, multiple words, a sentence, or multiple sentences. Moreover, in this application, 'language' means choice of word or words. As such, in the context of this application, a user's source and / or target language is dictated by a plurality of characteristics. Language is not defined solely by an officially recognised language for a given country.
[0012] The plurality of characteristics of a user, including the second user, may comprise one or more of: native language, vernacular language, and / or dialect, wherein 'native language' (also known as 'first language', 'native tongue', or 'mother tongue') refers to the language a user has been most commonly exposed to from birth or within a critical period; 'vernacular language' refers to the ordinary, informal, spoken form of a native language; and 'dialect' refers to a slight variation of a native language that is peculiar to a specific region or social group.
[0013] The plurality of characteristics may also comprise one or more of: ethnicity, cultural background, religion, age, socio-economic background, register, and situation, for example. However, any suitable user characteristic may be used. In some embodiments, the plurality of characteristics may comprise at least one of: user comprehension, cognitive load, communication modality, and a maximum time between the input communication and the output communication. The maximum time between the input communication and the output communication may be up to 0.1 seconds, 0.3 seconds, 0.5 seconds, or 1 seconds, for example. This enables synchronous communication between users. The maximum time between the input communication and the output communication may comprise a preferred processing speed. Consequently, the method may comprise: providing the output communication to the second user within 0.1, 0.3, 0.5, 1 or 2 seconds of receiving the input communication from the first user.
[0014] The claimed method, or at least a part thereof, may operate on a server. The server may be a central server. Alternatively, or in addition, the claimed method, or at least a part thereof, may operate on a distributed network. The distributed network may comprise the server and / or central server. The use of a distributed network enables the method to utilise edge computing. For example, there may be a plurality of servers and / or remote devices. The remote devices may comprise a first remote device belonging to the first user. The remote devices may comprise a second remote device belonging to the second user. These servers and / or remote devices may form the distributed network. The plurality of language models may be located on a variety of servers and / or remote devices within the network. Selection of the input phrase translating language model may be based on the location of the language model(s) within the distributed network. By selecting a language model local to the first and / or second user, processing power may be reduced.
[0015] In some embodiments, the input phrase translating language model may be selected additionally based on the required semantic accuracy of the resulting output communication; the processing speed required to translate the input communication; the input communication and / or the second user profile. One or more of these factors may be used to determine the processing power required to translate the input communication. Each of these factors may reduce computational power and / or overhead.
[0016] In some embodiments, the input phrase translating language model may be adjusted based on feedback from the second user. The feedback may be based on the output phrase and / or second user profile. The feedback may be provided in real-time. For example, the method may comprise: receiving, from the second user, feedback based on the output phrase and / or second user profile; and adjusting the input phrase translating language model based on the feedback. The adjustment may be carried out in real-time.
[0017] In some embodiments, the input phrase translating language model may be adjusted based on feedback from an external operative. The external operative may be a linguist. The feedback may be based on the input phrase, the output phrase, and / or the second user profile. The feedback may be provided in real-time. For example, the method may comprise providing the input phrase, output phrase, and second user profile to an external operative; and adjusting the input phrase translating language model based on the input phrase, output phrase, and second user profile. The adjustment may be carried out in real-time.
[0018] The second user profile may be generated by the second user. The second user profile may be generated using a remote device. The remote device may be remote from the server. The remote device may be a mobile device, such as a phone, tablet, or laptop. The remote device may form part of the distributed network.
[0019] The method may comprise generating a first user profile comprising a plurality of characteristics of the first user. The first user profile may be generated by the first user. The first user profile may be generated using a first remote device. As such, the second user's remote device may be a second remote device. The first remote device may be remote from the server. The first remote device may form part of the distributed network. Each remote device may comprise at least one language model. Each remote device may comprise a plurality of language models. This may include the input phrase translating language model. The language model(s) may be bespoke and / or off-the-shelf. As such, latency may be reduced through localised processing, where possible, with complex computations offloaded to a central server(s) only when necessary. Localised processing may comprise processing on a remote device, such as the first and / or second user's remote device.
[0020] The first and / or second user profile may be uploaded to a server and / or the distributed network. As such, there may be a first user profile belonging to the first user and a second user profile belonging to the second user. However, any number of user profiles may be present on the server and / or distributed network. Each user profile may comprise a plurality of characteristics of the corresponding user.
[0021] Each user profile may be continuously updated in real-time. More specifically, the plurality of user characteristics corresponding to each profile may be continuously updated. The update may be based on the input communication and / or output communication. In some embodiments, the update is based on a previous input communication and / or output communication. As such, each plurality of user characteristics may be a dynamic set of characteristics. For example, the method may comprise analysing the input communication and updating the first and / or second user profile and / or the plurality of characteristics of the first and / or second user in response thereto.
[0022] Any user may request a communication link between themselves and another user. A server may generate the communication link. For example, the first and / or second user may request a communication link between the first and second user profile. The server may generate a communication link between the first and second user profile.
[0023] The input communication may be generated by the first user. The first user may generate the input communication using the first remote device. The first remote device may also be a mobile device, such as a phone, tablet, or laptop. The input communication may be received at a server and / or a remote device, from the first user, via the communication link.
[0024] In some embodiments, the input communication comprises an inherent meaning. The inherent meaning of the input communication may comprise the inherent meaning of the input phrase. The inherent meaning of the input phrase may be based, at least in part, on the context of the input communication. The method may comprise: determining an inherent meaning of the input communication. Alternatively, or in addition, the method may comprise: determining an inherent meaning of the input phrase. The inherent meaning may be determined using the distributed network. Alternatively, the inherent meaning may be transferred from the source language to the target language.
[0025] The plurality of language models may comprise the input phrase translating language model. One of the plurality of language models may be used to generate the output phrase. Each language model may be pretrained on vast corpora that includes a wide variety of languages. This pretraining process involves learning from a massive amount of text data, enabling each model to understand input phrases and generate output phrases in multiple languages. However, each language model may be adapted to more efficiently understand an input phrase and / or generate an output phrase in a specific sub-set of languages, or specific source / target language pair (i.e., German adult to Spanish child; or American Man to Australian woman, for example). As such, the input phrase translating language model may be selected based on the input communication or, more specifically, the input phrase. Alternatively, or in addition, the input phrase translating language model may be selected based on the second user profile or, more specifically, the target language. This may reduce processing power, as the language model best adapted to translate an input phrase in a source large to an output phrase in a target language may be selected.
[0026] The distributed network may comprise a prompt. The distributed network may comprise a plurality of prompts. Each prompt may be configured to prevent a language model from containing hallucinations. A hallucination may occur when a language model, particularly a large language model, perceives patterns or objects that are non-existent or imperceptible to human observers, thus resulting in outputs that are nonsensical or altogether inaccurate. Alternatively, or in addition, one or more prompt may be configured to convey the inherent meaning from the input phrase to the output phrase.
[0027] The input communication may comprise a single input phrase. Alternatively, the input communication may comprise a plurality of input phrases. The input communication may comprise technical information. As such, the method may comprise: determining an inherent meaning of the technical information. The technical information may comprise an instruction. The technical information may relate to a user's profession. The technical information may be educational. Alternatively, or in addition, the technical information may comprise technical subject matter, such as engineering, science, or technology content. The input phrase may be a technical input phrase. For example, the input phrase may comprise the technical information. Consequently, the method may comprise: determining an inherent meaning of the technical input phrase. A server and / or a remote device may be used to identify the source language. The source language may be different from the target language of the user that generated the input communication.
[0028] The inherent meaning may be based, at least on part, on the context of the input communication. Alternatively, or in addition, inherent meaning may be based, at least on part, on at least one previous input communication. As such, the method may comprise: determining the inherent meaning based, at least on part, on at least one of context and a previous input communication. The previous input communication may have been provided by the first user. Alternatively, or in addition, the previous input communication may have been provided by the second user (or another user). This enables the overall context of a plurality of input communications to be used to determine the most accurate inherent meaning.
[0029] The output communication may be provided to the second user by the server and / or a remote device. In particular, the output communication may be provided to a remote device belonging to the second user.
[0030] There may be a lag between the input communication and the output communication. The lag may be up to 0.1 second, 0.5 seconds, 1 second, 2 seconds, or 5 seconds. In some embodiments, the lag may be more than 5 seconds. In other embodiments, the lag may be substantially unnoticeable to the users.
[0031] The method may comprise: receiving, from the second user, at least one characteristic of the second user. In other words, at least one characteristic of the second user may be provided by, or input by, the second user. This ensures that the output communication, particularly the output phrase, can accurately account for the user characteristics that have the greatest influence on the optimal target language. At least one received characteristic may be native language, or age, for example. However, any characteristic may be received. The at least one received characteristic may be requested. More specifically, the at least one received characteristic may be requested when the user's profile is created. However, the at least one received characteristic may be requested at any time. For example, the user may be requested to update at least one characteristic.
[0032] The method may comprise: determining, based on a previous input communication received from the second user, at least one characteristic of the second user. In other words, at least one characteristic of the second user may be determined by a server configured to carry out at least part of the claimed method. The previous input communication may be an input communication from a previous iteration of the method. This enables the target language to account for an increased number of user characteristics over time, which may help to improve the quality of the output communication. The determined characteristic may be vernacular language, or dialect, for example. However, any characteristic may be determined.
[0033] The method may comprise: receiving, from the first user, a desired attribute of the output phrase; and generating the output phrase based, at least in part, on the desired attribute. The desired attribute may be provided by, or input by, the first user. The attribute may form part of the first user profile. The desired attribute may be configured to alter the diction and / or register of the output phrase, wherein 'diction' is the choice and use of word or words; and 'register' is the variety of words used for a particular purpose or communicative situation. For example, a user having a naturally informal diction and / or register may set the desired attribute to a more formal diction and / or register. As such, the output phrase provided to the second user, in their target language, may have the same inherent meaning as the input phrase but be provided in a more formal manner. The same principle applies in other attributes, such as angry / calm, happy / sad, serious / funny etc.... Any suitable attribute may be provided.
[0034] The method may comprise: determining the inherent meaning of the input phrase using, at least in part, a manually generated dataset. A manually generated dataset may provide the most accurate inherent meaning for a specific phrase, or selection of phrases, which are not well, or easily, understood using a language model. The manually generated dataset may be located on the distributed network. For example, the dataset may be an add-on module that works in conjunction with the plurality of language models to override certain outputs. Alternatively, the manually generated dataset may be remote from the distributed network. For example, the dataset may be used offline to train the plurality of language models. Each language model may be updated and / or trained using the dataset to modify the language model itself.
[0035] The method may comprise: selecting the input phrase translating language model based on the input phrase, the second user profile, the source language and the target language. The input phrase translating language model may be a text translation tool. Alternatively, the input phrase translating language model may be a Large Language Model (LLM). Due to their deep learning architecture and self-attention mechanisms (the transformer architecture), LLMs can consider the overall inherent meaning of the input communication, allowing for more nuanced and accurate communications. This is a significant improvement over earlier statistical translation models, which often struggled with context and idiomatic expressions.
[0036] The input phrase translating language model may be selected from the plurality of language models based on the results of specific source / target language pair evaluations. These evaluations may be carried out manually, by experts, to determine whether the inherent meaning of an output phrase in a target language accurately reflects the inherent meaning of an input phrase in a source language. The language model deemed to communicate the most accurate inherent meaning may be selected to for use in subsequent identical source / target language pairs. However, these evaluations may be re-run continuously, or periodically, such that a different language model may be selected in the future for an identical source / target language pair. For example, these evaluations may be run daily, weekly, monthly, or yearly.
[0037] In some embodiments, the input phrase translating language model may be selected from the plurality of language models based on the source language; the target language; the second user profile; and / or the inherent meaning of the input phrase.
[0038] At least one language model may be located on a server. Alternatively, or in addition, at least one language model may be located externally from a server. For example, at least one language model may form part of the distributed network. As such, at least one language model may be in communication with a central server.
[0039] The method may comprise: determining the target language based, at least in part, on the plurality of characteristics of the second user; generating, using the input phrase translating language model, a processing phrase having the same inherent meaning as the input phrase, wherein the processing phrase is in a processing language that is different from the source language and the target language, and wherein the processing language is based, at least in part, on the source and target language; selecting, from the plurality of language models, a processing phrase translating language model for translating the processing phrase, wherein the processing phrase translating language model is selected based on the processing language and the target language; and generating the output phrase using the processing phrase translating language model, wherein the output phrase has the same inherent meaning as the processing phrase, and wherein the output phrase is in the target language. The processing phrase translating language model may be located on the distributed network. The use of an intermediate processing language may reduce computational complexity, processing power and / or latency.
[0040] Each of the processing phrase and the output phrase may be generated using a language model selected from the plurality of language models, based on the source and target language, as previously described. The same language model may be used for each language pair (i.e., source / processing and processing / target). Alternatively, a different language model may be used for each language pair. As such, in some embodiments, a plurality of language models may be used for a given source / target language pair. For example, a first language model may be used to convert the input phrase into the processing phrase. The processing language may be different from the source language and target language. The first language model may be selected based on the source language. More specifically, the first language model may be selected on the source language and a list of predetermined languages. The list of predetermined language may be a list of processing languages. Alternatively, or in addition, the list of predetermined language may be a list of common languages. The first language model may be the input phrase translating language model. A second language model may then be used to convert the processing phrase into the output phrase. The second language model may be selected based on the processing language and the target language. This enables a rare or unusual source language to be converted into a more common processing language and, subsequently, for the more common processing language to be converted into a rare or unusual target language. The second language model may be the processing phrase translating language model.
[0041] The method may comprise: determining the target language based, at least in part, on the plurality of characteristics of the second user; generating, using the input phrase translating language model, a first processing phrase having the same inherent meaning as the input phrase, wherein the first processing phrase is in a first processing language that is different from the source language and the target language, and wherein the first processing language is based, at least in part, on the source language; selecting, from the plurality of language models, a first processing phrase translating language model for translating the processing phrase, wherein the first processing phrase translating language model is selected based on the first processing language and the target language; generating, using the first processing phrase translating language model, a second processing phrase having the same inherent meaning as the first processing phrase, wherein the second processing phrase is in a second processing language that is different from the source language, the first processing language, and the target language, and wherein the second processing language is based, at least in part, on the first processing language and the target language; selecting, from the plurality of language models, a second processing phrase translating language model for translating the second processing phrase, wherein the second processing phrase translating language model is selected based on the second processing language and the target language; and generating the output phrase using the second processing phrase translating language model, wherein the output phrase has the same inherent meaning as the second processing phrase, and wherein the output phrase is in the target language.
[0042] As such, a plurality of processing languages may be used. For example, a first language model may be used to convert the input phrase into the first processing phrase. The first language model may be the input phrase translating language model. The first processing language may be different from the source language and target language. The first language model may be selected based on the source language. More specifically, the first language model may be selected on the source language and the list of predetermined languages.
[0043] A second language model may be used to convert the first processing phrase into the second processing phrase. The second language model may be the first processing phrase translating language model. The second processing language may be different from the source language, first processing language, and the target language. The second language model may be selected based on the first processing language and the target language. More specifically, the second language model may be selected on the first processing language, the target language, and the list of predetermined languages.
[0044] A third language model may then be used to convert the second processing phrase into the output phrase. The third language model may be a second processing phrase translating language model. The third language model may be selected based on the second processing language and the target language. This enables a first rare source language to be converted into a first common processing language and, subsequently, for the first common processing language to be converted into a second common processing language and, subsequently, for the second common processing language to be converted into a second rare or unusual target language. The conversion from a first common language into a second common language enables the most appropriate language model to be used for each rare / common language pair.
[0045] The first processing phrase may be generated using the first processing phrase translating language model. The second processing phrase may be generated using the second processing phrase translating language model. Each of the first and second processing phrase translating language models may be selected from the plurality of language models based on any of the previously mentioned criteria for selecting a language model. For example, the first processing phrase translating language model may be selected based on the input phrase and the second user profile. The second processing phrase translating language model may be selected based on the first processing phrase and the second user profile.
[0046] The first and second processing phrase translating language model may be selected, adjusted, and / or located based on the same inputs and / or criteria as the input phrase translating language model. In fact, all language models may be selected, adjusted, and / or located based on the same inputs and / or criteria as the input phrase translating language model. The input phrase may be a text-based input phrase or audio-based input phrase. Alternatively, or in addition, the output phrase may be a text-based output phrase or audio-based output phrase. The text-based input and / or output phrase may be a written message. The audio-based input and / or output phrase may be a pre-recorded audio clip or a live audio stream.
[0047] For example, the method may comprise: receiving, from the first user, an input communication comprising an audio-based input phrase; converting the audio-based input phrase into a text-based input phrase, wherein the text-based input phrase has an inherent meaning; generating a text-based output phrase having the same inherent meaning as the text-based input phrase, wherein the textbased output phrase is in the target language; and generating an output communication comprising the text-based output phrase.
[0048] The audio-based input phrase may be converted into a text-based input phrase by the distributed network or, in particular, a server within the distributed network. More specifically, the audio-based input phrase may be converted into a text-based input phrase using the audio-to-text capabilities of a language model. The language model may be a large language model. The language model may be integrated within the server. Alternatively, the language model may be remote from the server, but in communication therewith. The language model may be any of the previously mentioned language models.
[0049] In some embodiments, the audio-based input phrase may be converted into a text-based input phrase by a remote device. The remote device may be the first user's device. More specifically, the audiobased input phrase may be converted into a text-based input phrase using the audio-to-text capabilities of the remote device.
[0050] The method may comprise: continuously receiving, from the first user, an input communication comprising a plurality of audio-based input phrases; identifying an initial audio-based input phrase within the input communication; converting the initial audio-based input phrases into an initial textbased input phrase, wherein the initial text-based input phrase has an inherent meaning; generating an initial text-based output phrase having the same inherent meaning as the initial text-based input phrase, wherein the initial text-based output phrase is in the target language; generating an initial output communication comprising the initial text-based output phrase; and providing the initial output communication to the second user whilst the input communication is still being received.
[0051] In this context, 'continuously' means for the entire time taken to complete the aforementioned method steps. As such, an input communication comprising a plurality of audio-based input phrases may be broken down and communicated to the second user via a plurality of sequentially provided output communications. This enables longer audio-based input communication to be processed in segments and for the earlier segments to be communicated to the second user whilst the later segments are still being provided by the first user.
[0052] Each audio-based input phrase may be converted into a text-based input phrase as soon as it is identified. The subsequent method steps may be then carried out on each identified audio-based input phrase immediately, such that earlier audio-based input phrases are communicated to the second user whilst the later audio-based input phrases are still being generated by the first user.
[0053] The method may comprise: generating a text-based output phrase having the same inherent meaning as the input phrase, wherein the text-based output phrase is in the target language; converting the text-based output phrase into an audio-based output phrase; and generating an output communication comprising the audio-based output phrase. More specifically, the method may comprise generating a text-based output phrase having the same inherent meaning as a text-based input phrase.
[0054] The text-based output phrase may be converted into an audio-based output phrase by the distributed network or, in particular, a server within the distributed network. More specifically, the text-based output phrase may be converted into an audio-based output phrase using the text-to-audio capabilities of a language model. The language model may be a large language model. The language model may be integrated within the server. Alternatively, the language model may be remote from the server, but in communication therewith. Again, the language model may be any of the previously mentioned language models.
[0055] In some embodiments, the text-based output phrase may be converted into an audio-based output phrase by a remote device. The remote device may be the second user's device. More specifically, the text-based output phrase may be converted into an audio-based output phrase using the text-to-audio capabilities of the remote device.
[0056] Consequently, the claims method may comprise any combination of receiving a text-based input phrase and / or an audio-based input phrase and, subsequently, generating a text-based output phrase and / or an audio-based output phrase.
[0057] The input and / or output communication may further comprise video data. The video data may be a pre-recorded video clip or a live video stream. The video data may show sign language. In some embodiments, the input and / or output communication may comprise video data (e.g., a live stream or video clip); an audio-based input and / or output phrase (e.g., sound); and / or a text-based input and / or output phrase (e.g., sub-titles). Any text-based and / or audio-based output phrases within the output communication may be synchronized to the video data within the output communication.
[0058] The method may comprise: receiving, from the first user, an input communication comprising a foreground audio-based input phrase and background audio-based input noise; separating the foreground audio-based input phrase and background audio-based input noise; converting the foreground audio-based input phrase into a text-based input phrase, wherein the text-based input phrase has an inherent meaning; generating a text-based output phrase having the same inherent meaning as the text-based input phrase, wherein the text-based output phrase is in the target language; converting the text-based output phrase into a foreground audio-based output phrase; and generating an output communication comprising the foreground audio-based output phrase and the background audio-based input noise.
[0059] As such, the background noise within the input communication can be utilised within the output communication. This may be used to make the output communication more realistic and / or authentic.
[0060] Any means of generating audio may generate the background audio-based input noise. For example, the background audio-based input noise may comprise audio generated by another person or group of people (i.e., not the first user). Alternatively, or in addition, the background audio-based input noise may comprise audio generated by a vehicle, an animal, the weather and / or music, for example.
[0061] The foreground audio-based input phrase and background audio-based input noise may be separated using the distributed network or, in particular, a server within the distributed network. This may be achieved using readily available and / or known technologies and methods. For example, the foreground audio-based input phrase and background audio-based input noise may be separated using the audio-to-text capabilities of a language model. The language model may be any of the previously mentioned language models. For example, the language model may be a large language model. The language model may be integrated within the server. Alternatively, the language model may be remote from the server, but in communication therewith.
[0062] In some embodiments, the foreground audio-based input phrase and background audio-based input noise may be separated by a remote device. Again, this may be achieved using readily available and / or known technologies and methods. The remote device may be the first user's device. For example, the foreground audio-based input phrase and background audio-based input noise may be separated using the audio-to-text capabilities of the remote device. The method may comprise: receiving, from the first user, an input communication comprising a plurality of input phrases in the source language, wherein the plurality of input phrases each have an inherent meaning; generating at least one output phrase having the same inherent meaning as the plurality of input phrases, wherein the at least one output phrase is in the target language; and generating an output communication comprising the at least one output phrase.
[0063] The inherent meaning of the plurality of input phrases may be an overall inherent meaning of the plurality of input phrases. For example, the inherent meaning of the plurality of input phrases may be an inherent meaning of the input communication. Consequently, the output communication may better communicate the inherent meaning of a complex input communication. For example, the inherent meaning of an input communication may be considered as a whole, rather than focusing on an individual input phrase that may have a different inherent meaning when considered in the context of other input phrases.
[0064] In some embodiments, the plurality of input phrases may be in a plurality of different source languages. As such, the method may comprise: receiving, from the first user, an input communication comprising a plurality of input phrases in a plurality of different source languages, wherein each input phrase has an inherent meaning; generating an output phrase corresponding to each input phrase, wherein each output phrase has the same inherent meaning as the corresponding input phrase, wherein each output phrase is in the target language, and wherein the target language is different from each source language; and generating an output communication comprising each output phrase.
[0065] In other words, the claimed method can communicate inherent meaning between users even when an input communication comprises a plurality of different source languages. This enables the output communication to reflect the inherent meaning of a complex input communication more accurately. Alternatively, or in addition, the method may comprise: generating an output phrase having the same inherent meaning as the plurality of input phrases; and generating an output communication comprising the output phrase.
[0066] The method may comprise: identifying a third user profile comprising a plurality of characteristics of the third user; generating a second user output phrase having the same inherent meaning as the input phrase, wherein the second user output phrase is in a second user target language that is different from the source language, and wherein the second user target language is based, at least in part, on the plurality of characteristics of the second user; generating a third user output phrase having the same inherent meaning as the input phrase, wherein the third user output phrase is in a third user target language that is different from the source language and the second user target language, and wherein the third user target language is based, at least in part, on the plurality of characteristics of the third user; generating a second user output communication comprising the second user output phrase and generating a third user output communication comprising the third user output phrases; and providing the second user output communications to the second user and providing the third user output communication to the third user.
[0067] The method may comprise: selecting, from the plurality of language models, a plurality of input phrase translating language models for translating the input phrase, wherein the plurality of input phrase translating language models are based on the input phrase, the second user profile, and the third user profile; generating, using the plurality of input phrase translating language models, a plurality of output phrases having the same inherent meaning as the input phrase, wherein each output phrase is in a target language that is different from the source language, and wherein a first target language is based, at least in part, on the plurality of characteristics of the second user, and wherein a second target language is based, at least in part, on the plurality of characteristics of the third user. The plurality of output phrases may comprise the second and / or third user output phrases.
[0068] The second and third user output phrases may be generated simultaneously. Similarly, the second and third user output communications may be generated and / or provided to the second and third users, respectively, simultaneously. This enables three users to send and receive communications (i.e., within a group conversation) in a language that they understand, regardless of the language used by the other users. However, any number of users may be present with a group conversation. As such, a plurality of user profiles may be identified, and a plurality of output phrases may be generated and provided to the plurality of users in a plurality of different target languages.
[0069] According to the present invention, there is also provided a method for improving the authenticity of a machine generated communication, the method comprising: receiving, from the first user, an input communication comprising a foreground audio-based input phrase and background audio-based input noise; separating the foreground audio-based input phrase and background audio-based input noise; converting the foreground audio-based input phrase into a text-based input phrase, wherein the textbased input phrase has an inherent meaning; generating a text-based output phrase having the same inherent meaning as the text-based input phrase, wherein the text-based output phrase is in the target language; converting the text-based output phrase into a foreground audio-based output phrase; generating an output communication comprising the foreground audio-based output phrase and the background audio-based input noise; and providing the output communication to the second user. This method may further comprise any of the previously disclosed method steps. There is also provided a computer program for carrying out the previously disclosed method steps. There is also provided an apparatus comprising a memory storing computer readable instructions and at least one processor, which, when executing the computer readable instruction, is operable to control the apparatus to perform the previously disclosed method steps. The apparatus may be the distributed network. The apparatus may be a server. The apparatus may be a remote device.
[0070] The invention will now be further and more particularly described, by way of example only, with reference to the accompanying drawings.
[0071] FIGURES
[0072] Figure 1 shows the steps of a computer-implemented method for communicating inherent meaning between users; and
[0073] Figure 2 shows the computer-implemented method according to figure 1, further comprising a plurality of additional and optional steps.
[0074] DETAILED DESCRIPTION
[0075] Figure 1 shows the steps of a computer-implemented method 100 for communicating inherent meaning between users. In particular, the method comprises: receiving 110, from a first user, an input communication comprising an input phrase in a source language, wherein the input phrase has an inherent meaning; identifying 105 a second user profile comprising a plurality of characteristics of a second user; selecting 107, from a plurality of language models, an input phrase translating language model for translating the input phrase, wherein the input phrase translating language model is selected based on the input phrase and the second user profile; generating 130, using the input phrase translating language model, an output phrase having the same inherent meaning as the input phrase, wherein the output phrase is in a target language that is different from the source language, and wherein the target language is based, at least in part, on the plurality of characteristics of the second user; generating 140 an output communication comprising the output phrase; and providing 150 the output communication to the second user.
[0076] Figure 2 shows a flow diagram of the steps of the computer-implemented method 100 according to figure 1, further comprising a plurality of additional and optional steps. For example, in some embodiments, the method 100 comprises: determining 120 an inherent meaning of the input phrase. In some embodiments, the method 100 comprises: receiving 102, from the second user, at least one characteristic of the second user. Alternatively, or in addition, the method 100 may comprise: determining 104, based on a previous input communication received from the second user, at least one characteristic of the second user.
[0077] In some embodiments, the method comprises: receiving 101, from the first user, a desired attribute of the output phrase; and generating 130 the output phrase based, at least in part, on the desired attribute.
[0078] In some embodiments, the method comprises: determining 120 the inherent meaning of the input phrase using, at least in part, a manually generated dataset 121.
[0079] In some embodiments, the method comprises: generating 130 the output phrase using a language model 131 selected from a plurality of language models 131i.nbased on the source and target language. The language model 131 may be the input phrase translating language model selected in step 107.
[0080] In some embodiments, the method comprises: determining 122 the target language based, at least in part, on the plurality of characteristics of the second user; generating 124 a processing phrase having the same inherent meaning as the input phrase, wherein the processing phrase is in a processing language that is different from the source language and the target language, and wherein the processing language is based, at least in part, on the source and target language; and generating 130 the output phrase, wherein the output phrase has the same inherent meaning as the processing phrase, and wherein the output phrase is in the target language.
[0081] In some embodiments, the method comprises: determining 122 the target language based, at least in part, on the plurality of characteristics of the second user; generating 124 a first processing phrase having the same inherent meaning as the input phrase, wherein the first processing phrase is in a first processing language that is different from the source language and the target language, and wherein the first processing language is based, at least in part, on the source language; generating 126 a second processing phrase having the same inherent meaning as the first processing phrase, wherein the second processing phrase is in a second processing language that is different from the source language, the first processing language, and the target language, and wherein the second processing language is based, at least in part, on the first processing language and the target language; and generating 130 the output phrase, wherein the output phrase has the same inherent meaning as the second processing phrase, and wherein the output phrase is in the target language. In some embodiments, the method comprises: receiving 110, from the first user, an input communication comprising an audio-based input phrase; converting 116 the audio-based input phrase into a text-based input phrase, wherein the text-based input phrase has an inherent meaning; generating 130 a text-based output phrase having the same inherent meaning as the text-based input phrase, wherein the text-based output phrase is in the target language; and generating 140 an output communication comprising the text-based output phrase.
[0082] In some embodiments, the method comprises: continuously receiving 110, from the first user, an input communication comprising a plurality of audio-based input phrases; identifying 112 an initial audiobased input phrase within the input communication; converting 116 the initial audio-based input phrases into an initial text-based input phrase, wherein the initial text-based input phrase has an inherent meaning; generating 130 an initial text-based output phrase having the same inherent meaning as the initial text-based input phrase, wherein the initial text-based output phrase is in the target language; generating 140 an initial output communication comprising the initial text-based output phrase; and providing 150 the initial output communication to the second user whilst the input communication is still being received.
[0083] In some embodiments, the method comprises: generating 130 a text-based output phrase having the same inherent meaning as the input phrase, wherein the text-based output phrase is in the target language; converting 132 the text-based output phrase into an audio-based output phrase; and generating 140 an output communication comprising the audio-based output phrase.
[0084] In some embodiments, the method comprises: receiving 110, from the first user, an input communication comprising a foreground audio-based input phrase and background audio-based input noise; separating 114 the foreground audio-based input phrase and background audio-based input noise; converting 116 the foreground audio-based input phrase into a text-based input phrase, wherein the text-based input phrase has an inherent meaning; generating 130 a text-based output phrase having the same inherent meaning as the text-based input phrase, wherein the text-based output phrase is in the target language; converting 132 the text-based output phrase into a foreground audio-based output phrase; and generating 140 an output communication comprising the foreground audio-based output phrase and the background audio-based input noise.
[0085] In some embodiments, the method comprises: receiving 110, from the first user, an input communication comprising a plurality of input phrases in the source language, wherein the plurality of input phrases each have an inherent meaning; generating 130 at least one output phrase having the same inherent meaning as the plurality of input phrases, wherein the at least one output phrase is in the target language; and generating 140 an output communication comprising the at least one output phrase.
[0086] In some embodiments, the method comprises: receiving 110, from the first user, an input communication comprising a plurality of input phrases in a plurality of different source languages, wherein each input phrase has an inherent meaning; generating 130 an output phrase corresponding to each input phrase, wherein each output phrase has the same inherent meaning as the corresponding input phrase, wherein each output phrase is in the target language, and wherein the target language is different from each source language; and generating 140 an output communication comprising each output phrase.
[0087] In some embodiments, the method comprises: identifying 1052 a third user profile comprising a plurality of characteristics of the third user; generating 130 a second user output phrase having the same inherent meaning as the input phrase, wherein the second user output phrase is in a second user target language that is different from the source language, and wherein the second user target language is based, at least in part, on the plurality of characteristics of the second user; generating 1302 a third user output phrase having the same inherent meaning as the input phrase, wherein the third user output phrase is in a third user target language that is different from the source language and the second user target language, and wherein the third user target language is based, at least in part, on the plurality of characteristics of the third user; generating 140 a second user output communication comprising the second user output phrase and generating 1402 a third user output communication comprising the third user output phrases; and providing 150 the second user output communications to the second user and providing 1502 the third user output communication to the third user.
[0088] In some embodiments, the method comprises: identifying 105i. ,.na plurality of additional user profiles each comprising a plurality of characteristics of the corresponding user; generating 130i. ,.na plurality of user output phrases having the same inherent meaning as the input phrase, wherein each user output phrase is in a target language that is different from the source language, and wherein each target language is based, at least in part, on the plurality of characteristics of the corresponding user; generating 140i...na plurality of output communications, each comprising the user output phrase for a given user; and providing 150i...nthe plurality of user output communications to the corresponding users.
[0089] In the figures, 'line-arrows' have been used to link the method steps. Each 'line-arrow' comprises a pointed end and an opposing tail end. Each method step is connected to at least one other method step using a line-arrow. Some line-arrows overlap. A method step connected to the pointed end of a line-arrow must occur after the method step connected to the tail end of the corresponding linearrow. All other steps may occur in any order.
[0090] Various further aspects and embodiments of the present invention will be apparent to those skilled in the art in view of the present disclosure, "and / or" where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. For example, "A and / or B" is to be taken as specific disclosure of each of (i) A, (ii) B and (iii) A and B, just as if each is set out individually herein.
[0091] Unless context dictates otherwise, the descriptions and definitions of the features set out above are not limited to any particular aspect or embodiment of the invention and apply equally to all aspects and embodiments that are described. It will further be appreciated by those skilled in the art that although the invention has been described by way of example with reference to several embodiments, it is not limited to the disclosed embodiments and that alternative embodiments could be constructed without departing from the scope of the invention as defined in the appended claims.
Claims
CLAIMS1. A computer-implemented method for communicating inherent meaning between users, the method comprising: receiving, from a first user, an input communication comprising an input phrase in a source language, wherein the input phrase has an inherent meaning; identifying a second user profile comprising a plurality of characteristics of a second user; selecting, from a plurality of language models, an input phrase translating language model for translating the input phrase, wherein the input phrase translating language model is selected based on the input phrase and the second user profile; generating, using the input phrase translating language model, an output phrase having the same inherent meaning as the input phrase, wherein the output phrase is in a target language that is different from the source language, and wherein the target language is based, at least in part, on the plurality of characteristics of the second user; generating an output communication comprising the output phrase; and providing the output communication to the second user.
2. The method according to any preceding claim, comprising: receiving, from the second user, at least one characteristic of the second user.
3. The method according to any preceding claim, comprising: determining, based on a previous input communication received from the second user, at least one characteristic of the second user.
4. The method according to any preceding claim, comprising: receiving, from the first user, a desired attribute of the output phrase; and generating the output phrase based, at least in part, on the desired attribute.
5. The method according to any preceding claim, comprising: determining the inherent meaning of the input phrase using, at least in part, a manually generated dataset.
6. The method according to any preceding claim, comprising: selecting the input phrase translating language model based on the input phrase, the second user profile, the source language and the target language.
7. The method according to any preceding claim, comprising: determining the target language based, at least in part, on the plurality of characteristics of the second user; generating, using the input phrase translating language model, a processing phrase having the same inherent meaning as the input phrase, wherein the processing phrase is in a processing language that is different from the source language and the target language, and wherein the processing language is based, at least in part, on the source and target language; selecting, from the plurality of language models, a processing phrase translating language model for translating the processing phrase, wherein the processing phrase translating language model is selected based on the processing language and the target language; andgenerating the output phrase using the processing phrase translating language model, wherein the output phrase has the same inherent meaning as the processing phrase, and wherein the output phrase is in the target language.
8. The method according to any preceding claim, comprising: determining the target language based, at least in part, on the plurality of characteristics of the second user; generating, using the input phrase translating language model, a first processing phrase having the same inherent meaning as the input phrase, wherein the first processing phrase is in a first processing language that is different from the source language and the target language, and wherein the first processing language is based, at least in part, on the source language; selecting, from the plurality of language models, a first processing phrase translating language model for translating the processing phrase, wherein the first processing phrase translating language model is selected based on the first processing language and the target language; generating, using the first processing phrase translating language model, a second processing phrase having the same inherent meaning as the first processing phrase, wherein the second processing phrase is in a second processing language that is different from the source language, the first processing language, and the target language, and wherein the second processing language is based, at least in part, on the first processing language and the target language; selecting, from the plurality of language models, a second processing phrase translating language model for translating the second processing phrase, wherein the second processing phrase translating language model is selected based on the second processing language and the target language; and generating the output phrase using the second processing phrase translating language model, wherein the output phrase has the same inherent meaning as the second processing phrase, and wherein the output phrase is in the target language.
9. The method according to any preceding claim, comprising: receiving, from the first user, an input communication comprising an audio-based input phrase; converting the audio-based input phrase into a text-based input phrase, wherein the text-based input phrase has an inherent meaning; generating a text-based output phrase having the same inherent meaning as the text- based input phrase, wherein the text-based output phrase is in the target language; and generating an output communication comprising the text-based output phrase.
10. The method according to claim 9, comprising: continuously receiving, from the first user, an input communication comprising a plurality of audio-based input phrases; identifying an initial audio-based input phrase within the input communication; converting the initial audio-based input phrases into an initial text-based input phrase, wherein the initial text-based input phrase has an inherent meaning; generating an initial text-based output phrase having the same inherent meaning as the initial text-based input phrase, wherein the initial text-based output phrase is in the target language; generating an initial output communication comprising the initial text-based output phrase; and providing the initial output communication to the second user whilst the input communication is still being received.
11. The method according to any preceding claim, comprising:generating a text-based output phrase having the same inherent meaning as the input phrase, wherein the text-based output phrase is in the target language; converting the text-based output phrase into an audio-based output phrase; and generating an output communication comprising the audio-based output phrase.
12. The method according to any preceding claim, comprising: receiving, from the first user, an input communication comprising a foreground audiobased input phrase and background audio-based input noise; separating the foreground audio-based input phrase and background audio-based input noise; converting the foreground audio-based input phrase into a text-based input phrase, wherein the text-based input phrase has an inherent meaning; generating a text-based output phrase having the same inherent meaning as the text- based input phrase, wherein the text-based output phrase is in the target language; converting the text-based output phrase into a foreground audio-based output phrase; and generating an output communication comprising the foreground audio-based output phrase and the background audio-based input noise.
13. The method according to any preceding claim, comprising: receiving, from the first user, an input communication comprising a plurality of input phrases in the source language, wherein the plurality of input phrases each have an inherent meaning; generating at least one output phrase having the same inherent meaning as the plurality of input phrases, wherein the at least one output phrase is in the target language; andgenerating an output communication comprising the at least one output phrase.
14. The method according to any preceding claim, comprising: receiving, from the first user, an input communication comprising a plurality of input phrases in a plurality of different source languages, wherein each input phrase has an inherent meaning; generating an output phrase corresponding to each input phrase, wherein each output phrase has the same inherent meaning as the corresponding input phrase, wherein each output phrase is in the target language, and wherein the target language is different from each source language; and generating an output communication comprising each output phrase.
15. The method according to any preceding claim, comprising: identifying a third user profile comprising a plurality of characteristics of the third user; generating a second user output phrase having the same inherent meaning as the input phrase, wherein the second user output phrase is in a second user target language that is different from the source language, and wherein the second user target language is based, at least in part, on the plurality of characteristics of the second user; generating a third user output phrase having the same inherent meaning as the input phrase, wherein the third user output phrase is in a third user target language that is different from the source language and the second user target language, and wherein the third user target language is based, at least in part, on the plurality of characteristics of the third user; generating a second user output communication comprising the second user output phrase and generating a third user output communication comprising the third user output phrases; andproviding the second user output communications to the second user and providing the third user output communication to the third user.
Citation Information
Patent Citations
Facilitating end-to-end communications with automated assistants in multiple languages
CN113128239A
Systems and methods for multi-user multi-lingual communications
EP3005151A2
Automated systems and methods for providing bidirectional parallel language recognition and translation processing with machine speech production for two users simultaneously to enable gapless interactive conversational communication
US20190354592A1
System and method for direct speech translation system
US20200226327A1