User terminal and system capable of providing multiple interpretation services

The multi-interpretation system addresses the challenge of multi-party conversations by using individual AI modules on user terminals and a support server for simultaneous translation, enhancing accuracy and reducing costs.

WO2026059285A1PCT designated stage Publication Date: 2026-03-19KIM E E +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/014045
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-10
Filing Date
2025-09-10
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing conversational AI systems struggle with multi-party conversations due to difficulties in interpreting multiple voices simultaneously, leading to errors and high costs associated with hiring interpreters or relying on performance-dependent AI servers.

Method used

A multi-interpretation system utilizing individual AI modules on user terminals and an interpretation support server to facilitate simultaneous interpretation services between multiple parties, allowing for real-time translation and voice mimicry across different languages.

Benefits of technology

Enables efficient, cost-effective, and secure multi-party interpretation services by leveraging individual AI modules on user devices, reducing implementation and maintenance costs while maintaining high accuracy and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025014045_19032026_PF_FP_ABST
    Figure KR2025014045_19032026_PF_FP_ABST
Patent Text Reader

Abstract

A multi-interpretation system according to an embodiment disclosed in the present document comprises: a plurality of user terminals for providing interpretation services by using individual AI modules, respectively; and an interpretation support server for supporting the interpretation service between the plurality of user terminals. The interpretation support server may: share, with the plurality of user terminals, a plurality of used languages corresponding to the user terminals and information related to a first chat room when the plurality of user terminals participate in the first chat room; and when first conversation content is acquired from a speaker terminal among the plurality of user terminals, share the first conversation content with the remaining terminals except for the speaker terminal among the plurality of user terminals. The remaining terminals may translate the first conversation content from a first used language corresponding to the speaker terminal into used languages respectively corresponding to the remaining terminals through the individual AI modules, and output the translated first conversation content to each user.
Need to check novelty before this filing date? Find Prior Art

Description

User terminal and system capable of providing multi-interpretation services

[0001] The various embodiments disclosed in this document relate to automatic interpretation technology.

[0002] With the development of transportation and communication, communication between people who speak different languages ​​is becoming more frequent. To communicate with speakers of other languages, people are learning other languages ​​or relying on interpreters or automatic translation programs.

[0003] For example, automatic interpretation AI can already provide a significant level of interpretation capabilities. For instance, conversational AI based on Large Language Models (LM) can offer excellent natural language processing capabilities. Furthermore, LLM-based conversational AI can perform summarization and translation tasks of given content at a very high level. Consequently, people are accessing automatic interpretation AI on their smartphones to utilize the interpretation functions they desire.

[0004] However, since conversational AI is developed based on 1:1 services, it can be very difficult to simultaneously interpret multi-party conversations using conversational AI. Furthermore, when multiple voices are mixed during a multi-party conversation, many errors may occur during the process of a single conversational AI converting multiple voices into text.

[0005] As an example of a multi-party conversation interpretation method, multiple interpreters perform the interpretation, and the system receives and transmits the interpretation results. This method incurs the cost of hiring interpreters.

[0006] As another example of a multi-party conversation interpretation method, interpretation functions were provided using AI servers or machine learning servers equipped with translation capabilities. However, since this method is highly dependent on the performance of the interpretation server, the costs for its implementation, maintenance, and updates can be high. Consequently, even if an additional AI with superior interpretation capabilities is introduced, it may be difficult to apply the added AI to the interpretation server.

[0007] Various embodiments disclosed in this document can provide a user terminal and a multi-interpretation system capable of providing simultaneous interpretation services between multiple parties using different languages ​​by utilizing individual AI modules.

[0008] A multi-interpretation system according to an embodiment disclosed in this document comprises: a plurality of user terminals that provide interpretation services using individual AI modules; and an interpretation support server that supports the interpretation services between the plurality of user terminals. When the plurality of user terminals participate in a first chat room, the interpretation support server shares information related to the first chat room and a plurality of languages ​​corresponding to the user terminals with the plurality of user terminals. When the server obtains a first conversation content from a speaker terminal among the plurality of user terminals, it shares the first conversation content with the remaining terminals excluding the speaker terminal among the plurality of user terminals. The remaining terminals can translate the first conversation content from a first language corresponding to the speaker terminal to a language corresponding to each of the remaining terminals through their respective individual AI modules, and output the translated first conversation content to each user.

[0009] Additionally, a user terminal according to one embodiment disclosed in this document comprises an input module; an output module; a communication module capable of communicating with an interpretation support server; and a processor including a first app and an AI module that provide interpretation services. The first app sets a first language of a first user to be input into a first chat room, obtains information of another user who inputs another language participating in the first chat room from the interpretation support server, shares a first conversation content corresponding to the utterance of the first user obtained through the input module with respect to the first language through the interpretation support server to a participant terminal of the first chat room, and upon receiving a second conversation content spoken in the other language from the server, translates the second conversation content into the first language using the AI ​​module, and outputs the translated second conversation content through the output module.

[0010] According to the various embodiments disclosed in this document, simultaneous interpretation services between multiple parties using different languages ​​can be provided using individual AI modules. In addition, various effects that can be identified directly or indirectly through this document may be provided.

[0011] Figure 1 shows a configuration diagram of a multi-interpretation system according to one embodiment.

[0012] Figure 2 shows a configuration diagram of an interpretation support server according to one embodiment.

[0013] FIG. 3 shows a configuration diagram of a user terminal according to one embodiment.

[0014] FIG. 4 shows an example diagram of an interpretation operation according to one embodiment.

[0015] FIGS. 5 and FIGS. 6 show examples of dialogue content correction according to one embodiment.

[0016] FIG. 7 shows a flowchart of a multi-interpretation method according to one embodiment.

[0017] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.

[0018] Figure 1 shows a configuration diagram of a multi-interpretation system according to one embodiment.

[0019] Referring to FIG. 1, a multi-interpretation system (12) according to one embodiment may include an interpretation support server (100) and a plurality of user terminals (200). The plurality of user terminals (200) may include two or more user terminals using different languages. However, FIG. 1 describes an example in which the plurality of user terminals (200) include first to third user terminals (200).

[0020] Multiple user terminals (200) may be terminals of a user who wishes to use the interpretation service. Each user terminal (200) may be a various computing device such as a smartphone, smartpad, laptop, or personal computer.

[0021] Each user terminal (200) may have a first app for a multi-interpretation service installed. The first app may provide at least one of the following functions for the interpretation service: a chat room opening function, a function to join an opened chat room, a language setting function, an AI module to use, an input type (e.g., voice, text), a cleanup function setting, a correction function setting, and an intermediate language usage setting. The AI ​​module may be implemented by connecting to at least one AI model, for example, on-device AI, edge device AI, or cloud AI. The first app is executed by the processor of each user terminal (200). Therefore, in this document, at least some operations of the first app may be described with the user terminals (200) as the subject.

[0022] The interpretation support server (100) can provide (download) a first app for interpretation services to multiple user terminals (200) in response to requests from multiple user terminals (200). The interpretation support server (100) can provide multiple interpretation services by interacting with each user terminal (200) through the first app installed on each of the multiple user terminals (200).

[0023] The interpretation support server (100) can open a chat room (or create a multi-party chat group) upon the request of one of the multiple user terminals (200), for example, a first user terminal (200_1) (a first app installed on the first user terminal (200_1)). When the first chat room is opened, the interpretation support server (100) can store chat room information in a database, including a chat room identifier (e.g., chat room ID), a creator identifier (e.g., user ID), a participant identifier, the number of participants, and a participation link, as shown in Table 1 below.

[0024] [Table 1]

[0025]

[0026] The interpretation support server (100) can share chat room information (e.g., chat room identifier) ​​with the first user terminals (200). For example, the interpretation support server (100) can share the participation link of the first chat room with the first user terminal (200_1) who is the creator of the chat room.

[0027] When the first user terminal (200_1) obtains chat room information (e.g., chat room identifier, participation link) from the interpretation support server (100), it may share the chat room participation link with the second user terminal (200_2) and the third user terminal (200_3) to induce participation in the chat room. For example, the first user terminal (200_1) may provide the participation link of the first chat room, such as a barcode (or QR code), to the second and third user terminals (200_3). Subsequently, the second user terminal (200_2) and the third user terminal (200_3) may participate in the first chat room by taking a picture of the barcode with a camera and executing the participation link.

[0028] In an embodiment, the interpretation support server (100) may support setting a password for a participation link. A first user terminal (200_1) may set a password for a participation link through the interpretation support server (100), and second and third user terminals (200_2, 200_3) may participate in a first chat room by entering the password shared by the first user.

[0029] When the second and third user terminals (200_2, 200_3) join the first chat room, the interpretation support server (100) can register the second and third user terminals (200_2, 200_3) as participants and update the chat room information as shown in Table 2.

[0030] [Table 2]

[0031]

[0032] The first app of the first to third user terminals (200) can store the language of use of the first to third users set by the first to third users. In this embodiment, an example is described where the first user sets the language of use (first language of use) to Korean, the second user sets the language of use (second language of use) to Japanese, and the third user sets the language of use (third language of use) to English.

[0033] The interpretation support server (100) can acquire / generate first to third user information and store it in a database when first to third user terminals (200) open or participate in the first chat room. The first to third user information may each include first to third user identifiers, first to third terminal identifiers, and first to third languages ​​used (Korean / Japanese / English). Each of the above user information may be generated based on personal information such as the unique number of the user terminal, each user's mobile phone number, each user's email address, and each user account information.

[0034] When the interpretation support server (100) sets the first to third user information, it may share at least some of the first to third user information with the first to third user terminals (200) as participant information for the first chat room. For example, the interpretation support server (100) may share participant information including the first to third languages ​​used and the first to third user identifiers with the first to third user terminals (200). Specifically, if the first language used is Korean, the second language used is Japanese, and the third language used is English, the first participant information including the first to third languages ​​used and user identifiers may be shared as follows. The interpretation support server (100) may share the first to third user information with the first to third user terminals (200).

[0035] ▶ 1st Participant Information: KOR_User1, JAP_User2, USA_User3

[0036] Hereinafter, a specific example of a multi-interpretation service through a first chat room between the first to third user terminals (200) is described.

[0037] The first user (speaker) can input a voice utterance, "Hello. Nice to meet you," by running the first app of the first user terminal (200_1). Then, the first app of the first user terminal (200_1) can convert the voice utterance into a text utterance in Korean using an STT module. Afterwards, the first user terminal (200_1) can output the Korean utterances of the conversation participants of the first chat room and the first user as shown in the screen below.

[0038]

[0039] The first user terminal (200_1) can transmit the following first speaker conversation information to the interpretation support server (100). When the interpretation support server (100) receives the first speaker conversation information, it can share the received first speaker conversation information with the first user terminal (200_1) and the third user terminal (200_3).

[0040]

[0041] When the first app of the second user terminal (200_2) obtains the shared first speaker conversation information, it can translate the first conversation content from Korean into Japanese (second language). Referring to the example below, the first app of the second user terminal (200_2) inputs a command prompt to translate the first conversation content, which is in Korean, into Japanese into the individual AI module of the second user terminal (200_2), and can obtain the first conversation content translated into Japanese from the individual AI module.

[0042]

[0043] The first app of the second user terminal (200_2) can display the first conversation content translated into Japanese along with the original Korean text of the first conversation content.

[0044]

[0045] Similarly, the first app of the third user terminal (200_3) can obtain first speaker conversation information and translate the first conversation content from Korean into English (third language). As shown in the example below, the first app of the third user terminal (200_3) inputs a command prompt to translate the first conversation content, which is in Korean, into English into an individual AI module set in the third user terminal (200_3), and can obtain the first conversation content translated into English from the individual AI module.

[0046]

[0047] The first app of the third user terminal (200_3) can display the first conversation content translated into English as follows. The first app of the third user terminal (200_3) can further display the original Korean text of the first conversation content.

[0048]

[0049] The third user and the second user can greet each other simultaneously by voice through the third user terminal (200_3) and the second user terminal (200_2) as follows.

[0050] ▶ Third user: I'm very happy to talk with you guys.

[0051] ▶ Second User:

[0052] The second user terminal (200_2) and the third user terminal (200_3) can each receive the voice utterances of the second user and the third user, respectively, and convert them into text utterances through their respective STT modules.

[0053] The second user terminal (200_2) can transmit the following second speaker conversation information to the conversation participants through the interpretation support server (100).

[0054]

[0055] The third user terminal (200_3) can transmit the following third speaker conversation information to the conversation participants through the interpretation support server (100).

[0056]

[0057] When the first app of the first user terminal (200_1) simultaneously receives conversation information from the second and third speakers, it can translate the content of the second and third conversations using individual AI modules. For example, the first app of the first user terminal (200_1) separates the two speaker conversation information by speaker (second and third users) and language (Japanese, English) and inputs a command prompt to the AI ​​module, and as a result, can obtain the content of the second and third conversations translated into Korean (first language) from the AI ​​module.

[0058]

[0059] The first user terminal (200_1) can display the second and third conversation content translated into Korean in the first conversation room along with the original text of the second and third conversation content as follows.

[0060]

[0061] When the first app of the second user terminal (200_2) receives third speaker conversation information through the interpretation support server (100), it inputs a command prompt to an individual AI module asking to translate the third conversation content from English (the third language) into Japanese, and can receive the third conversation content translated into Japanese from the individual AI module.

[0062]

[0063] When the first app of the third user terminal (200_3) receives the second speaker conversation information through the interpretation support server (100), it inputs a command prompt to the individual AI module to translate the second conversation content from Japanese to English as follows, and can receive the second conversation content translated into English from the individual AI module.

[0064]

[0065] The first app of the second user terminal (200_2) can display the second conversation content spoken by the second user, the translated third conversation content, and the original text thereof in the first conversation room as follows.

[0066]

[0067] In addition, the first app of the third user terminal (200_3) can display the second conversation content translated into English, the original text, and the third conversation content spoken by the third user in the first conversation room as follows.

[0068]

[0069] According to various embodiments, a first app of user terminals (200) may generate user voice TTS information for generating each user's voice and provide the generated user voice TTS information to an interpretation support server (100). The interpretation support server (100) may share the user voice TTS information with other user terminals participating in the first chat room. The user voice TTS information may be used to mimic the user's voice, for example, as acoustic features extracted from each user's voice. To this end, the interpretation support server (100) may store first to third user information, which further includes the first to third user voice TTS information, in a database.

[0070] According to various embodiments, each user terminal (200) may output the first conversation content in relation to an object icon associated with each user identifier. The object icon may include, for example, a flag or symbol of the first language used. The object icon may further include other user identification information (e.g., name).

[0071] According to various embodiments, the first app of each user terminal (200) can translate the language designated by each user into an intermediate language (e.g., English) and share it. The intermediate language may be a widely used language such as English, Chinese, or Russian. Each user terminal (200) can translate and output the conversation content of the intermediate language shared by another user into the designated language. Accordingly, each user terminal (200) can provide a multi-interpretation service using a low-level AI module capable only of translation between two languages ​​(language of use ↔ intermediate language).

[0072] According to various embodiments, the creator (first user terminal (200_1)) can set a common topic in the created chat room. The common topic may be presented by the first app and set by the user. For example, the first app may present a list of frequently used common topics at the top of the chat room, such as 'Shopping', 'Restaurant', 'Travel', 'Movies / Art', 'Meetings', and 'Chat'. The user (e.g., the creator) may select any one of the presented common topics. In this case, the interpretation support server (100) may share the set common topic with the user terminals of the chat room participants. When the first to third user terminals (200) translate the shared conversation content through individual AI modules, they may add the common topic or related content to the instruction prompt. Alternatively, each of the first to third user terminals (200) can translate the shared conversation content by using a RAG (Retrieval Augmented Generation) module to additionally provide information inferable from a common topic to the AI ​​module. For example, if the first conversation room is about an 'Automotive Engineering Conference,' each of the first to third user terminals (200) can add a command prompt presenting a common topic such as 'The topic of conversation is automotive engineering,' or add a command prompt related to the common topic such as 'You are an expert with deep knowledge of automotive engineering.' Additionally or generally, the first to third user terminals (200) can provide reports, search results, etc. regarding automotive engineering to the AI ​​module through additional prompts or the RAG module. Accordingly, the first app according to one embodiment can interpret the conversation content more accurately.

[0073] According to various embodiments, the first app of each user terminal (200) may further provide a function to organize (e.g., summarize or divide) the input conversation content or to correct errors. Accordingly, in one embodiment, the accuracy of the multi-interpretation service may be improved.

[0074] In this way, the multi-interpretation system (12) according to one embodiment acquires speech content through the input means (e.g., microphone, keypad) of each user terminal (200), so that even if multiple users speak next to each other at the same time, the speech content is recognized at each user terminal (200) which has very strong directionality, so the likelihood of speech being mixed up or misrecognized is very low. Therefore, the multi-interpretation system (12) according to one embodiment can support each user in participating in the conversation without missing the conversation context by providing the content of each spoken word as translated text without omission, even in a situation where multiple users speak voice words at the same time.

[0075] In addition, the multi-interpretation system (12) according to one embodiment supports the diverse use of individual AIs of user terminals (200) and allows for easy replacement of AI modules according to user needs, thereby providing significant advantages in terms of implementation costs, system performance, and maintenance.

[0076] In addition, a multi-interpretation system (12) according to one embodiment can support multiple users who use different languages ​​to conveniently use an N:N interpretation service by utilizing the AI ​​modules of their terminals.

[0077] Figure 2 shows a configuration diagram of an interpretation support server according to one embodiment.

[0078] Referring to FIG. 2, an interpretation support server (100) according to one embodiment may include a communication module (110), a database (130), and a processor (150). In one embodiment, some components of the interpretation support server (100) may be omitted, or additional components may be included. Additionally, some of the components of the interpretation support server (100) may be combined to form a single entity, while performing the same functions as the components prior to the combination.

[0079] The communication module (110) can support the establishment of a communication channel or a wireless communication channel between the interpretation support server (100) and another device (e.g., user terminal (200)), and the performance of communication through the established communication channel. The communication channel may include, for example, at least one communication channel among LAN, FTTH, xDSL, Wibro, Wireless LAN, Wi-Fi, Bluetooth, Zigbee, WFD (Wi-Fi Direct), UWB (Ultrawideband), Infrared Data Association (IrDA), BLE (Bluetooth Low Energy), NFC (Near Field Communication), 3G, 4G, or 5G.

[0080] The database (130) may include various forms of volatile or non-volatile memory. For example, the database (130) may include storage devices such as RAM (random access memory), flash memory, or SSD (Solid State Drive). In one embodiment, the database (130) may be located inside or outside the processor (150), and the database (130) may be connected to the processor (150) through various known means. The database (130) may store various data used by at least one component of the interpretation support server (100) (e.g., processor (150)). The data may include, for example, input data or output data for software and related commands. For example, the database (130) may store at least one instruction and data for providing multiple interpretation services.

[0081] According to one embodiment, the database (130) may store data related to the first app. The first app may provide at least one of the following functions for an interpretation service: a function to open a chat room, a function to join an opened chat room, a function to set the language to use, an AI module to use, a type of input (e.g., voice, text), a setting for a cleanup function, a setting for a correction function, and a setting for using an intermediate language. The database (130) may store chat room information (e.g., chat room ID), information on participants of each chat room (e.g., participating user ID), and user information. The database (130) may further store each user's voice and object icons related to each user (e.g., icons and user names).

[0082] The processor (150) can control at least one other component (e.g., a hardware or software component) of the interpretation support server (100) and can perform various data processing or operations. The processor (150) may include, for example, at least one of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor, an application processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA), and may have multiple cores.

[0083] According to one embodiment, when a first chat room for a multi-interpretation service is opened, the processor (150) may store chat room information in a database (130). The chat room information may include a chat room identifier, chat room participant information (user identifiers and the number of participants in the chat room), and creator information. The chat room information may further include a participation link for the chat room.

[0084] In one embodiment, the chat room information may further include a common topic of the chat room. For example, if a common topic (e.g., automotive engineering conference) is set by a chat room creator terminal (e.g., 200_1) among a plurality of user terminals (200), the processor (150) may store the set common topic as chat room information.

[0085] According to one embodiment, the processor (150) can share chat room information with at least one user terminal (200) at the request of the creator. For example, the processor (150) can share a participation link of the first chat room with the creator's terminal and check information of other users who have joined the first chat room through the participation link.

[0086] The processor (150) can update the participant information of the shared first chat room in the chat room information. For example, the processor (150) can identify the user identifier and the language used by each user for all user terminals (200) that have joined the first chat room. The processor (150) can share the identified user identifier and language used with multiple user terminals (200) that have joined the first chat room. The user information may include, for example, a user identifier, a terminal identifier, and a language used.

[0087] According to one embodiment, the processor (150) can obtain a first conversation content from a speaker terminal among a plurality of user terminals (200). For example, the processor (150) can obtain first speaker conversation information including a first conversation room identifier, a first user identifier, the language used by the speaker (first language used), and a first conversation content as shown in Table 1.

[0088] According to one embodiment, the processor (150) can identify the remaining terminals among a plurality of user terminals (200), excluding the speaker terminal, based on a first user identifier and a first chat room identifier. The processor (150) can share the first speaker conversation information with the remaining terminals identified through the communication module (110). Subsequently, the remaining terminals can translate the first conversation content from the first language of use to the language of use set for each of the remaining terminals through individual AI modules. The remaining terminals can output the translated first conversation content as voice or text.

[0089] According to various embodiments, the processor (150) checks the language of use set by each user terminal (200) and, when sharing the conversation content of each user terminal with another user terminal, can provide an object icon associated with the language of use set by each user terminal. The object icon information can be stored in the database (130) as user information.

[0090] According to various embodiments, the user information may further include user voice TTS information for generating user voice. In this regard, the processor (150) may provide an interface for setting or inputting user voice to a user terminal (e.g., 200_1) participating in a chat room, and may acquire and analyze the user voice set / input through the interface to store user voice TTS information in a database. Processing the user voice to generate user voice TTS information may be supported by an external voice processing server (not shown) connected via a network. The external voice processing server may be included in the interpretation support server (100).

[0091] In one embodiment, the processor (150) can share a plurality of user voices corresponding to a plurality of user terminals (200) with the chat room participant terminals. Accordingly, each user terminal (200) can output a translation result of the conversation content by mimicking the voice of the shared speaker.

[0092] According to various embodiments, at least some operations of the interpretation support server (100) may be performed by any one of the plurality of user terminals (200). For example, the conversation content may be shared on the creator terminal instead of the interpretation support server (100). For example, a simple web server may be implemented on the creator terminal, and the chat room participants may connect to it. If the chat room participants are using a wireless LAN, communication occurs only on the local network, which has the effect of reducing latency.

[0093] In this way, the interpretation support server (100) according to one embodiment can support each user terminal in providing multi-party simultaneous interpretation services using individual processors and lightweight AI by opening and participating in a chat room for multi-interpretation services, managing chat room participants, and sharing conversation content between participants.

[0094] In addition, the interpretation support server (100) according to one embodiment manages only conversation sharing without storing or translating conversations, and the conversation sharing may also be configured to be handled by one of the user terminals, so the possibility of the conversation of the multi-interpretation process being exposed to a third party can be significantly reduced and thus high security can be provided.

[0095] FIG. 3 shows a configuration diagram of a user terminal according to one embodiment, and FIG. 4 shows an example diagram of an interpretation operation according to one embodiment.

[0096] Referring to FIG. 3, a user terminal (200) according to one embodiment may include an input module (210), an output module (220), a communication module (230), a memory (240), and a processor (250). In one embodiment, some components of the user terminal (200) may be omitted, or additional components may be included. Additionally, some of the components of the user terminal (200) may be combined to form a single entity, which can perform the same functions as the components prior to combination.

[0097] The input module (210) can receive user input using the user terminal (200). The input module (210) may include at least one input detection circuit, for example, a button, a touchscreen, or a microphone. According to one embodiment, the input module (210) can acquire voice or text input spoken by the user under the control of the processor (250).

[0098] The output module (220) can output at least one data of symbols, numbers, or characters visually or audibly under the control of the processor (250). The output module (220) may include at least one output device among, for example, a liquid crystal display, an OLED, a touchscreen display, or a speaker. The output module (220) can output conversation content shared in a chat room as voice, text, or icons under the control of the processor (250).

[0099] The communication module (230) can support the establishment of a communication channel or a wireless communication channel between a user terminal (200) and another device (e.g., an interpretation support server (100)), and the performance of communication through the established communication channel. The communication channel may include, for example, at least one communication channel among LAN, FTTH, xDSL, Wibro, Wireless LAN, Wi-Fi, Bluetooth, Zigbee, WFD (Wi-Fi Direct), UWB (Ultrawideband), Infrared Data Association (IrDA), BLE (Bluetooth Low Energy), NFC (Near Field Communication), 3G, 4G, or 5G.

[0100] Memory (240) may include various forms of volatile or non-volatile memory. For example, memory (240) may include read-only memory (ROM), flash memory, storage devices such as a Solid State Drive (SSD), and random access memory (RAM). In one embodiment, memory (240) may be located inside or outside the processor (250), and memory (240) may be connected to the processor (250) through various known means. Memory (240) may store various data used by at least one component of the user terminal (200) (e.g., processor (250)). The data may include, for example, input data or output data for software and related commands. For example, memory (240) may store at least one instruction and data for providing interpretation support. The memory (240) can store instructions and data for providing multi-interpretation services through the first app (251) and the AI ​​module (253).

[0101] The processor (250) can control at least one other component (e.g., a hardware or software component) of the user terminal (200) and can perform various data processing or operations. The processor (250) may include, for example, at least one of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor, an application processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA), and may have multiple cores. The processor (250) may include a first app (251) and an AI module (253). The processor (250) may further include a cleanup module (255) and a correction module (257). The first app (251), the AI ​​module (253), the cleanup module (255), and the correction module (257) may be executed by the processor (250) or may be software modules or hardware modules included in the processor (250). The operations of the first app (251), AI module (253), cleanup module (255), and correction module (257) are performed by the processor (250). Accordingly, in this document, at least some of the operations of the first app (251), AI module (253), cleanup module (255), and correction module (257) are described with the processor (250) as the subject. The AI ​​module (253) may be a default setting or an AI module set by the user. The AI ​​module (253) may be implemented by connecting to at least one AI model among on-device AI or cloud AI.

[0102] The processor (250) can acquire user utterance as text or voice through the input module (210). The processor (250) can share the acquired conversation content with the interpretation support server (100) through the communication module (230), or acquire the shared conversation content from the interpretation support server (100). The processor (250) can output the acquired conversation content through the output module (220).

[0103] According to one embodiment, the first app (251) can set a first language of use in a first chat room that the first user has opened or joined. For example, the processor (250) can open a first chat room, join a first chat room that has been opened, or set at least one of a language of use, a common topic, or a user voice in the first chat room through the first app (251).

[0104] According to one embodiment, the first app (251) can obtain information about other users who have joined the first chat room from the interpretation support server (100). The user information may include, for example, identifiers of all participants in the first chat room, user voices, and languages ​​used.

[0105] According to one embodiment, the first app (251) can obtain the voice or text of a first user who wishes to speak through the first chat room through an input module (210). The first app (251) can share the first conversation content (text speech) corresponding to the first user's speech with the participant terminal of the first chat room through an interpretation support server (100).

[0106] Referring to FIG. 4, when the first app (251) acquires a spoken voice, it can convert it into a text utterance through a speech-to-text (STT) module. The STT module can convert the voice into a text utterance based on a first language set by the first user. For example, the first app can share first speaker conversation information, including a speaker (user) identifier, a language used, conversation content, and a chat room identifier, with a first chat room participant terminal through an interpretation support server (100).

[0107] According to one embodiment, the first app (251) may receive a second conversation content spoken in a language other than the first language from the interpretation support server (100). For example, the first app (251) may obtain second speaker conversation information in the form of Table 3 above.

[0108] According to one embodiment, the first app (251) can translate the second conversation content from another language to the first language by using the second speaker conversation information AI module (253). For example, the first app (251) can input a command prompt that instructs the second conversation content to be translated from another language to the first language through the AI ​​module (253), and can obtain the second conversation content translated into the first language through the AI ​​module (253).

[0109] According to one embodiment, the first app (251) can output the translation result as voice through a speaker (output module (220)) or as text through a display (output module (220)). For example, the first app (251) can display the translated second conversation content in the first chat room in relation to the speaker information (e.g., object icon) of the conversation content corresponding to the translation result.

[0110] Referring to FIG. 4, the first app (251) can convert the second conversation content translated into the first language into speech through a TTS (text to speech) module. In this case, the first app (251) can output the translation result (translated second conversation content) as speech by simulating the speaker's voice using user voice TTS information obtained from the interpretation support server (100).

[0111] Additionally, the first app (251) can organize or correct the conversation content of the first user. This is explained below.

[0112] According to one embodiment, the first app (251) can use a cleanup module (255) to summarize or divide the conversation content of the first user. For example, the first app (251) can check whether the first conversation content exceeds a threshold length. If the first conversation content exceeds a threshold length, the cleanup module (255) can use an AI means (e.g., an AI model) to divide the first conversation content into threshold length units. Alternatively, the first app (251) can use the cleanup module (255) to summarize the first conversation content into short sentences (e.g., less than or equal to the threshold length).

[0113] Referring to FIG. 5, the organizing module (255) can input the language used for the first utterance (Korean), the content of the conversation, and an instruction prompt (e.g., result format (JSON), 'divide the conversation into short sentences' or 'summarize the conversation within 50 words') through an input prompt (510). As a result, the organizing module (255) can obtain the conversation content (520) organized by the AI ​​means. Accordingly, the first app (251) according to one embodiment can provide the user's utterance as several short sentences by cutting it into appropriate lengths when the user's utterance becomes excessively long, or share the summarized conversation content.

[0114] According to one embodiment, the first app (251) can correct errors in the conversation content based on the conversation content currently entered by the first user and the conversation history using a correction module (257).

[0115] Referring to FIG. 6, the correction module (257) can input an input prompt (610) that includes the currently entered conversation content, the previous conversation of the first conversation window, and an instruction prompt (e.g., 'Refer to the previous conversation to clarify the subject, object, predicate, etc., so as to fit the grammar and context'). As a result, the correction module (257) can obtain the conversation content (620) corrected into an accurate sentence by means of AI. Accordingly, the first app (251) according to one embodiment can prevent the occurrence of errors in the interpretation result caused by inaccurate expressions, such as the user's spoken pronunciation or abbreviated expressions.

[0116] In the above-described embodiment, the correction module (257) and the cleanup module (255) may be optionally used in the process of generating speaker conversation information.

[0117] According to one embodiment, if a common topic is set in the first chat room, the first app (251) may further utilize the common topic when translating the conversation content. For example, the first app (251) may further utilize the common topic when translating the first conversation content into an intermediate language or when translating the second conversation content shared in the first chat room through the interpretation support server (100). As another example, if the location of the first chat room is the venue of an automotive engineering academic conference, the chat room creator may set the common topic of the chat room to "Automotive Engineering Conference." In this case, the first app (251) may input a command prompt requesting a translation of the conversation content and the content of an interpretation service for the attendees of the automotive engineering conference into the AI ​​module (253). In this way, the first app (251) according to one embodiment can significantly improve the quality and accuracy of interpretation by explaining / instructing the purpose of use and the environment of use when translating each conversation content, thereby enabling the AI ​​module (253) to appropriately translate / interpret each conversation content according to the interpretation topic.

[0118] According to one embodiment, the first app (251) can translate and share the conversation content of each user terminal (200) into a widely used intermediate language, and can translate the conversation content of another user shared in the intermediate language into each user's language. For example, if an intermediate language is set in the first chat room, the first app (251) can translate the conversation content into the intermediate language before sharing it in the first chat room. Additionally, the first app (251) can share the conversation content translated into the intermediate language in the first chat room through the interpretation support server (100). The intermediate language may be a widely used language such as English, Chinese, or Russian. Accordingly, the first app (251) according to one embodiment can provide a multi-translation service even when direct translation between multiple languages ​​is not provided through the AI ​​module (253) of each user terminal (200), or when the size and performance of the AI ​​module (253) differ between the user terminals (200). For example, if the AI ​​module (253) is a super-large model AI such as ChatGPT, it can support translation between more than a hundred languages. On the other hand, if the AI ​​module (253) is a lightweight version of an open-source AI model such as llama, it can only support fewer than 10 languages. However, in one embodiment, translation between 'Korean ↔ Swahili', which is difficult to translate directly, can be easily translated in the manner of 'Korean ↔ English ↔ Swahili' using English as an intermediate language. Similarly, if direct translation between 'Korean ↔ Shanghai Chinese' is difficult, it can be translated into 'Korean ↔ Beijing Chinese ↔ Shanghai Chinese'.

[0119] In this way, the user terminal (200) according to one embodiment can use various individual AI modules (253) and can easily replace individual AI modules (253) according to the user's needs, thus providing significant advantages in terms of implementation costs, system performance, and maintenance.

[0120] Additionally, according to one embodiment, a user terminal (200) converts user conversation content entered as voice or text into text and transmits it to an interpretation support server (100), and as the interpretation support server (100) shares the conversation content with the participants of the conversation room along with their language, each user terminal (200) can support the translation into their own language so that they can listen to it as voice or view it as text.

[0121] FIG. 7 shows a flowchart of a multi-interpretation method according to one embodiment.

[0122] Referring to FIG. 7, in operation 710, the interpretation support server (100) can share participant information including the language used and user identifier with the first to third user terminals (200) that have joined the opened chat room.

[0123] In operation 720, when the first user terminal (200_1) receives the first user's utterance, it can generate a first conversation content corresponding to the input user utterance and display it on the display.

[0124] In operation 730, the first user terminal (200_1) can transmit the first conversation content to the interpretation support server (100).

[0125] In operation 740, the interpretation support server (100) can share the first conversation content with the second and third user terminals (200) that participated in the first conversation room.

[0126] In operation 750, the second user terminal (200_2) and the third user terminal (200_3) can translate the shared first conversation content through the language set by the second and third users, respectively, via the AI ​​module set by each.

[0127] In operation 760, the second user terminal (200_2) and the third user terminal (200_3) can output the translated first conversation content as voice or text.

[0128] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as “A or B,” “at least one of A and B,” “at least one of A or B,” “A, B or C,” “at least one of A, B and C,” and “at least one of A, B, or C” may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as “first,” “second,” or “first” or “second” may be used simply to distinguish a component from another corresponding component and do not limit the components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0129] As used herein, the term "module" may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0130] Various embodiments of the present document may be implemented as software (e.g., a program) comprising one or more instructions stored in a storage medium (e.g., memory (240) of FIG. 3) (e.g., internal memory or external memory) that can be read by a machine (e.g., a user terminal). For example, a processor (e.g., processor (250)) of a device (e.g., user terminal (200)) may call at least one of one or more instructions stored from a storage medium and execute it. This enables the device to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. A storage medium readable by the device may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily in the storage medium.

[0131] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0132] Components according to various embodiments of this document may be implemented in software or in hardware form, such as a digital signal processor (DSP), a field programmable gate array (FPGA), or an application-specific integrated circuit (ASIC), and may perform specific roles. The term "components" is not limited to software or hardware, and each component may be configured to reside in an addressable storage medium or configured to run one or more processors. As an example, components may include components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables.

[0133] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the components of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to the integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In a multi-interpretation system, Multiple user terminals providing interpretation services using individual AI modules; and It includes an interpretation support server that supports the interpretation service between the plurality of user terminals, and The above interpretation support server is, When the plurality of user terminals participate in the first chat room, information regarding the plurality of languages ​​used corresponding to the user terminals and the first chat room is shared with the plurality of user terminals, and When a first conversation content is obtained from a speaker terminal among the plurality of user terminals, the first conversation content is shared with the remaining terminals among the plurality of user terminals, excluding the speaker terminal. The remaining terminals mentioned above are, A multi-interpretation system that translates the first conversation content from a first language corresponding to the speaker terminal to a language corresponding to the remaining terminals through each of the individual AI modules, and outputs the translated first conversation content to each user terminal.

2. In claim 1, the interpretation support server is, A multi-interpretation system that acquires multiple user voice TTS information corresponding to the multiple user terminals and shares the multiple user voice TTS information with the multiple user terminals participating in the first chat room.

3. In claim 2, among the remaining terminals, the terminal with a voice interpretation function configured is, A multi-interpretation system that outputs the translated first conversation content as voice by simulating each user voice using user voice TTS information corresponding to the speaker terminal among the plurality of user voice TTS information.

4. In Claim 1, The above interpretation support server shares an object icon corresponding to the language used by each user terminal with the plurality of user terminals, and The above remaining terminals are a multi-interpretation system that outputs the above-mentioned translated first conversation content in relation to an object icon set by the speaker terminal.

5. In Claim 1, The interpretation support server shares a common topic of the first chat room set by at least one user terminal among the plurality of user terminals with the plurality of user terminals, and A multi-interpretation system in which each of the above user terminals translates the first conversation content by providing the above individual AI module with the above common topic, additional information inferred from the above common topic, or an instruction prompt related to the above common topic or the above additional information.

6. In claim 5, the interpretation support server is, A multi-interpretation system that provides a list of common topics with high usage frequency in a part area of ​​the first chat room and shares a common topic selected by at least one user terminal from the list of common topics.

7. In claim 1, the speaker terminal is, A multi-interpretation system that corrects errors in the first conversation content based on the first conversation content and the previous conversation history of the first conversation room, and shares the corrected first conversation content with the interpretation support server.

8. In claim 1, the speaker terminal is, A multi-interpretation system that, when the first conversation content exceeds a threshold length, generates the first conversation content organized by dividing or summarizing the first conversation content into units of the threshold length using an organization module.

9. In claim 1, where an intermediate language is set in the first chat room, The above speaker terminal translates the first conversation content into the intermediate language through the AI ​​module and shares the translated first conversation content, The above remaining terminals are a multi-interpretation system that translates the above-mentioned shared first conversation content from the intermediate language into the language of use set for each of the remaining terminals.

10. Input module; Output module; A communication module capable of communicating with an interpretation support server; and A processor including a first app and an AI module that provides interpretation services, wherein the first app, Setting the first language of the first user to be entered into the first chat room, and obtaining information of other users who enter other languages ​​participating in the first chat room from the interpretation support server, A first conversation content corresponding to the utterance of the first user obtained through the input module is shared with the participant terminal of the first chat room via the interpretation support server in relation to the first language used, and A user terminal that, upon receiving a second conversation content spoken in the other language from the interpretation support server, translates the second conversation content into the first language using the AI ​​module and outputs the translated second conversation content through the output module.

11. In claim 10, the first app is, A user terminal that, upon obtaining a voice utterance of the first user through the input module, converts the voice utterance into a text utterance through the AI ​​module and shares the first conversation content including the text utterance.

12. In claim 10, the first app is, A user terminal that corrects errors in the first conversation content based on the first conversation content and the previous conversation history of the first conversation room using a correction module.

13. In claim 10, the first app is, A user terminal that, when the above utterance exceeds a critical length, uses a sorting module to divide or summarize the above utterance into units of the critical length to generate the first conversation content.

14. In claim 10, the first app is, A user terminal that, if an intermediate language is set in the first chat room, translates the first conversation content into the intermediate language through the AI ​​module and shares the translated first conversation content.

15. In claim 10, the first app is, A user terminal that, if a common topic is set in the first chat room, provides the AI ​​module with the common topic, additional information inferred from the common topic, or an instruction prompt related to the common topic or the additional information to translate the second conversation content.

16. In claim 10, the first app is, A user terminal that transmits first user voice TTS information obtained by processing a first user voice acquired by the input module to the interpretation support server, thereby allowing the interpretation support server to share the first user voice TTS information with the first chat room.

17. In claim 10, the first app is, A user terminal that obtains second user voice TTS information corresponding to the second conversation content from the interpretation support server and outputs the translated second conversation content as a voice simulating the second user voice using the second user voice TTS information.

18. In claim 10, the first app is, A user terminal that corrects, organizes, or translates the second conversation content using the AI ​​module specified by the first user.

Citation Information

Patent Citations

  • A Cable For Preventing A Oil Vapor Explosion

    KR1020210090376A

  • Input sensing part and driving method thereof

    KR1020230133999A

  • Dried fruit tea composition and method for manufacturing thereof

    KR102109331B1

  • Context-aware and persona-based generative ai interpretation device and method for controlling the same

    KR102692549B1

  • Natural-language processing across multiple languages

    US20230096070A1