System

The system addresses the challenge of context-aware translation by using deep learning and user feedback to provide accurate translations that adapt to different communication scenarios.

JP2026017991APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024119052
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing translation systems fail to accurately understand the context and relationship between speakers, leading to misunderstandings and mistranslations, especially in diverse communication scenarios.

Method used

A system that utilizes deep learning-based models for speech recognition, natural language processing, and machine translation, combined with user feedback optimization, to adapt translations to the context and relationship, providing accurate and context-aware translations.

Benefits of technology

Enables smooth communication by generating appropriate translations that adapt to various situations, improving translation accuracy over time through continuous optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017991000001_ABST
    Figure 2026017991000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for acquiring an utterance of a user as voice data; means for converting the voice data into text data; means for analyzing a relationship between a context and a partner; means for translating the text data into a target language based on an analysis result; and means for notifying the user of a translation result.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, smooth communication between people who speak different languages ​​remains a major challenge. Misunderstandings and mistranslations can have serious consequences, especially in business situations. Even in everyday life, communication with family and friends can sometimes be difficult. To solve these problems, a system is needed that accurately understands the context and the relationship between the other person and provides appropriate translations in real time. [Means for solving the problem]

[0005] The present invention provides a system that includes a means for acquiring a user's speech as voice data, a means for converting the voice data into text data, a means for analyzing the context and the relationship with the other person, a means for translating the text data into a target language based on the analysis results, and a means for notifying the user of the translation results. A deep learning-based model is used for converting the voice data, and by collecting feedback from users and optimizing the translation algorithm, more accurate translations are achieved. This enables smooth communication in a variety of situations, from business to everyday conversation.

[0006] "User" refers to any person or entity who uses the system to speak and receive translation results.

[0007] "Utterance" refers to a form of information or communication that a user verbally conveys to a system.

[0008] "Voice data" refers to information that has been converted into a digital format from a user's speech.

[0009] "Text data" refers to data in a format in which voice data is converted into character information.

[0010] "Context" refers to the scene, situation, and background information in which the user's utterance is made.

[0011] "Relationship" refers to the social, professional, or personal relationship between the user and the target of the utterance.

[0012] "Analysis" refers to the process of understanding meaning and intent from audio and text data and identifying context and relationships.

[0013] "Target language" refers to the language into which the text is translated.

[0014] "Translation" refers to the act of converting text from one language into another.

[0015] A "deep learning-based model" is a form of artificial intelligence technology that uses multi-layer neural networks to learn complex patterns and perform speech recognition and translation.

[0016] "Feedback" refers to the evaluation and opinions provided by the user regarding the translation results.

[0017] An "algorithm" refers to a set of steps or a computational method for solving a particular problem.

[0018] "Optimization" refers to the adjustment or improvement of a system or algorithm to improve its performance.

[0019] "Notification" refers to the system's act of notifying the user of the translation results.

[0020] "System" refers to a set of components designed to analyze a user's speech and provide a translation result. [Brief explanation of the drawings]

[0021] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0022] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0023] First, the terms used in the following description will be explained.

[0024] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0025] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0026] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0027] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0029] [First embodiment]

[0030] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0031] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0032] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0033] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0034] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0036] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0037] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0038] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0039] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0040] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0041] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0042] This invention is a context-adaptive real-time translation system that uses speech recognition, natural language processing, and machine translation technologies. Specifically, the system acquires user utterances as voice data, converts the voice data into text data, analyzes the context and the relationship with the other party, and generates an appropriate translation based on the analysis results.

[0043] Program processing procedures and explanations

[0044] 1. User utterance transmission

[0045] The user speaks into a device (smartphone, tablet, PC, etc.), which captures the voice in real time and records it as audio data. The recorded audio data is compressed and sent to a server via the Internet.

[0046] 2. Audio data conversion

[0047] The server receives the voice data sent from the device and applies a speech recognition algorithm, which uses a deep learning-based model to convert the user's speech into text data with high accuracy. This process converts the voice data into text data and stores it in a temporary file or memory.

[0048] 3. Context and Relationship Analysis

[0049] The server analyzes the generated text data and estimates the context in which the user spoke and the relationship with the other person. Context analysis uses natural language processing technology to identify the scene by taking into account the content of the user's speech, time, location, and existing conversation history. The server also references a user profile database and past speech data to estimate the relationship between the user and the target person.

[0050] 4. Generating the Translation

[0051] The server uses a generative AI model to generate an appropriate translation based on information obtained from context and relationships. For example, it generates formal expressions for business situations and casual expressions for everyday conversations. The translation results are stored in temporary memory on the server.

[0052] 5. Gather feedback and optimize

[0053] The translation results are sent to the device and notified to the user. The user reviews the translation results and provides feedback if necessary. The device then sends this feedback to the server. The server analyzes the feedback and optimizes the parameters of the translation algorithm and generative AI model. This process improves future translation accuracy.

[0054] Specific examples

[0055] Example 1: Use in business situations

[0056] During a meeting, a user says, "How is the progress on this project?" The device captures the speech and sends it to the server. The server converts the speech into text and translates it, recognizing the formal context of the business setting. The resulting translation is "How is the progress on this project?" The translation is displayed on the device for the user to review.

[0057] Example 2: Use in everyday conversation

[0058] A user is with friends and says, "Where should we have dinner tonight?" The device captures the speech and sends it to the server. The server converts the speech to text and translates it, recognizing the casual context of everyday conversation. The resulting translation is "Where should we have dinner tonight?" The translation is displayed on the device for the user to confirm.

[0059] This invention combines speech recognition, natural language processing, and machine translation technologies to provide real-time translation suitable for a variety of situations. Furthermore, by optimizing the translation algorithm based on user feedback, it is possible to achieve more accurate translation.

[0060] The processing flow will be explained below.

[0061] Step 1:

[0062] The user speaks into the device, which uses a built-in microphone to capture the user's speech in real time, and this voice data is stored digitally in the device's temporary memory.

[0063] Step 2:

[0064] The device compresses the captured audio data and sends it over the internet to a server, where it is encrypted for security purposes.

[0065] Step 3:

[0066] The server receives the voice data sent from the terminal, decodes the received voice data, and prepares it for voice recognition processing.

[0067] Step 4:

[0068] The server applies a speech recognition algorithm to convert the audio data into text data, using a deep learning-based model to convert spoken content into text with high accuracy.

[0069] Step 5:

[0070] The server analyzes the generated text data and identifies contextual information and relationships with the other party. Contextual analysis uses natural language processing technology. This analysis categorizes the scene into categories such as business or everyday conversation.

[0071] Step 6:

[0072] The server references the user profile database and past conversation history to estimate the relationship between the user and the other party, and based on this, selects an appropriate communication style (formal, casual, etc.).

[0073] Step 7:

[0074] The server retrieves translation settings based on the analysis results and uses a generative AI model to translate the text data into the target language, adjusting the writing style and tone accordingly.

[0075] Step 8:

[0076] The server then sends the translation results to the device, which then displays them on the user interface. If necessary, speech synthesis technology can also be used to output the results as voice.

[0077] Step 9:

[0078] The user checks the translation results and provides feedback if necessary, which is sent from the device to the server.

[0079] Step 10:

[0080] The server analyzes the received feedback and updates and optimizes the parameters of the translation algorithm and generative AI model, thereby improving future translation accuracy.

[0081] In this way, this system translates user utterances in real time and provides appropriate translation results for a variety of situations. By optimizing the algorithm based on user feedback, it achieves high-accuracy translation over the long term.

[0082] Example 1

[0083] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0084] Conventional translation systems often perform simple translations without considering the context or the relationship with the target user, which can result in low translation accuracy. In addition, they lack the ability to optimize the system based on user feedback, making it difficult to continuously improve translation accuracy.

[0085] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0086] In this invention, the server includes means for acquiring user utterances as voice data, means for converting the voice data into text data, means for analyzing the relationship between the context and the target person, means for translating the text data into the target language based on the analysis results, means for notifying the user of the translation results, and means for collecting feedback from the user and optimizing the translation algorithm. This enables highly accurate translation that adapts to the context and relationship, and continuous system optimization is achieved by utilizing the feedback.

[0087] The "means for acquiring user speech as voice data" is a function for electronically capturing the voice uttered by the user and recording it as digital voice data.

[0088] The "means for converting voice data into text data" is a function for analyzing captured voice data and converting it into text data expressed as a corresponding character string.

[0089] The "means for analyzing the relationship between the context and the target person" is an analytical function for estimating the relationship between the context of an utterance and the target person of that utterance from the generated text data.

[0090] The "means for translating text data into a target language" refers to a translation function that converts the original text data into a different language based on the analyzed context and relationships.

[0091] The "means for notifying the user of the translation result" is a function for visually or audibly notifying the user of the generated translation result.

[0092] "Means for collecting user feedback and optimizing the translation algorithm" refers to a function that collects user evaluations and suggestions for corrections on translation results, and then adjusts the parameters and settings of the translation algorithm based on that information to improve accuracy.

[0093] A "deep learning-based model" is a type of algorithm that uses a multi-layer neural network to extract and transform data features, enabling advanced recognition and prediction.

[0094] A "generative AI model" is an algorithm that uses artificial intelligence technology to generate text or translate, and has the ability to generate appropriate output based on a prompt.

[0095] The present invention is a context-adaptive real-time translation system that uses speech recognition, natural language processing, and machine translation technologies. Specific embodiments for carrying out the present invention are described below.

[0096] First, the user speaks into the device (smartphone, tablet, PC, etc.). The device uses a built-in microphone to capture the user's voice in real time and record it as audio data. The recorded audio data is saved in a lossless compressed format such as FLAC and sent to a server via the Internet.

[0097] The server receives the voice data sent via the Internet, applies a speech recognition algorithm (e.g., Google Speech-to-Text API using a deep learning-based model) to analyze the voice data and convert it into text data, which is then saved in a temporary file or memory.

[0098] The server then analyzes the generated text data and estimates the context of the utterance and its relationship to the target person. This analysis uses natural language processing technology (e.g., the spaCy library). The server references a user profile database and a database of past utterances to identify the scene, taking into account the user's utterance content, time, location, and existing conversation history. It also references these databases to identify the relationship between the user and the target person.

[0099] The server generates an appropriate prompt based on the context and relationships, and generates the translation using a generative AI model (e.g., OpenAI GPT-3). The AI ​​model is fed with a prompt like this:

[0100] Input: "How is this project progressing?"

[0101] Context: Business meeting

[0102] Relationship: Superior-Subordinate

[0103] Produces the translation: "How is the progress on this project?"

[0104] The generated translation result is stored in temporary memory in the server and then sent to the terminal, where the user can check the translation result on the terminal screen or through voice output.

[0105] Finally, the user can review the translation results and provide feedback, such as ratings and suggestions for corrections. The device collects this feedback and sends it back to the server. The server analyzes the feedback and optimizes the parameters of the translation algorithm and generative AI model, thereby improving future translation accuracy.

[0106] As a concrete example, consider the case where a user says, "How is the progress on this project?" during a meeting. At this time, the device captures the audio and sends it to the server in FLAC format. The server converts it into text using the Google Speech-to-Text API and analyzes the context of the business meeting and the relationship between superiors and subordinates. Based on the analysis results, a prompt is sent to the GPT-3 model, which generates the translation result, "How is the progress on this project?". This translation result is sent to the device and confirmed by the user.

[0107] As another example, consider the case where a user is with friends and says, "Where should we have dinner tonight?" The device captures the audio and sends it to the server in FLAC format. The server converts it to text using the Google Speech-to-Text API and analyzes the casual context of everyday conversation and friendships. Based on the analysis results, the prompt is sent to the GPT-3 model, which generates the translation result, "Where should we have dinner tonight?" The translation result is then sent to the device for the user to confirm.

[0108] In this way, by combining speech recognition, natural language processing, and generative AI models, the present invention can provide highly accurate real-time translation that adapts to the user's context and relationships. Furthermore, by optimizing the translation algorithm based on user feedback, the system can be continuously improved.

[0109] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0110] Step 1: User utterance submission

[0111] explanation

[0112] The user speaks into a device (smartphone, tablet, PC, etc.), which then uses a built-in microphone to capture the voice in real time and record it as digital audio data.

[0113] input

[0114] User utterance: "How is this project going?"

[0115] output

[0116] FLAC format audio data

[0117] Specific actions

[0118] The user speaks to their smartphone, "How is this project progressing?" The device uses its built-in microphone to capture the audio and saves it as a FLAC audio file. The saved audio data is then sent to a server over the Internet.

[0119] Step 2: Convert the audio data

[0120] explanation

[0121] The server receives the voice data sent from the terminal and converts it into text data by applying a voice recognition algorithm.

[0122] input

[0123] FLAC format audio data

[0124] output

[0125] Text data: "How is this project progressing?"

[0126] Specific actions

[0127] The server receives FLAC audio data from the device. The server calls the Google Speech-to-Text API, analyzes the audio data, and converts it into text data such as "How is the progress on this project?" The converted text data is saved as a temporary file.

[0128] Step 3: Analyze context and relationships

[0129] explanation

[0130] The server analyzes the generated text data and estimates the context of the utterance and its relationship to the target person.

[0131] input

[0132] Text data: "How is this project progressing?"

[0133] output

[0134] Analysis results (context: business meeting, relationship: superior-subordinate)

[0135] Specific actions

[0136] The server analyzes the text data using the spaCy library, infers that the context of the utterance is a business meeting, and references a database of past utterances and a user profile database to determine that the user is a superior and the target person is a subordinate.

[0137] Step 4: Generate translations

[0138] explanation

[0139] The server generates prompt sentences based on context and relationships, and uses generative AI models to generate appropriate translations.

[0140] input

[0141] Analysis results (context: business meeting, relationship: superior-subordinate)

[0142] output

[0143] Translation result: "How is the progress on this project?"

[0144] Specific actions

[0145] The server inputs the following prompt into the generative AI model (OpenAI GPT-3):

[0146] Input: "How is this project progressing?"

[0147] Context: Business meeting

[0148] Relationship: Superior-Subordinate

[0149] Based on this prompt, the GPT-3 model generates the translation "How is the progress on this project?" This translation result is stored in temporary memory on the server.

[0150] Step 5: Notification of translation results

[0151] explanation

[0152] The server sends the generated translation result to the terminal and notifies the user.

[0153] input

[0154] Translation result: "How is the progress on this project?"

[0155] output

[0156] Terminal display: "How is the progress on this project?"

[0157] Specific actions

[0158] The server sends the generated translation to the device as an HTTPS response, and the user can check the translation result, "How is the progress on this project?", displayed on the device screen.

[0159] Step 6: Gather feedback and optimize

[0160] explanation

[0161] The user provides feedback on the translation result, and the terminal sends the feedback to the server, which analyzes the feedback and optimizes the translation algorithm.

[0162] input

[0163] User feedback: "The translation was accurate"

[0164] output

[0165] Optimized translation algorithm

[0166] Specific actions

[0167] The user reviews the translation result and provides feedback, rating it "the translation was accurate." The device then sends this feedback to the server, which analyzes the feedback and optimizes the translation algorithm by adjusting the model settings, resulting in higher accuracy in the next translation.

[0168] (Application example 1)

[0169] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0170] Food delivery services face the challenge of having delivery personnel communicate smoothly with foreign customers. In particular, the language barrier makes it difficult for delivery personnel to accurately understand order details and delivery instructions. This problem reduces delivery efficiency and reduces customer satisfaction.

[0171] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0172] In this invention, the server includes means for translating information and questions received by the delivery person from the customer in real time, means for acquiring the user's speech as voice data, means for converting the voice data into text data, means for translating the text data into the target language based on the analysis results, and means for notifying the user of the translation results. This enables the delivery person to communicate smoothly with foreign customers, improving delivery efficiency and customer satisfaction.

[0173] The "means for acquiring user speech as voice data" refers to a device or function for recording the user's speech as digital voice data.

[0174] The "means for converting voice data into text data" refers to an algorithm or software for analyzing acquired voice data and converting it into corresponding written text.

[0175] "Means for analyzing the context and the relationship with the other person" refers to natural language processing technologies and algorithms for understanding and analyzing the content of an utterance, its background, and the relationship between the speaker and the other person.

[0176] The "means for translating text data into a target language based on the analysis results" refers to a machine translation technology for generating text in a target language that is appropriately translated based on context and relationships.

[0177] "Means for notifying the user of the translation result" refers to an interface or function for displaying or audibly conveying the generated translation result to the user.

[0178] "A means for delivery personnel to translate information and questions received from customers in real time" refers to a system or algorithm that instantly translates the content of communication between delivery personnel and customers into other languages.

[0179] System Program

[0180] Hardware and Software Configuration

[0181] The system for implementing this invention includes a terminal operated by a user (e.g., a smartphone), a server connected via the Internet, and algorithms for performing speech recognition, natural language processing, and machine translation. The terminal has the function of capturing the user's voice and sending the voice data to the server. The server receives the voice data and performs the following main processes:

[0182] 1. Speech Recognition Using Deep Learning-Based Models

[0183] Use deep learning-based speech recognition algorithms (e.g., Google Speech-to-Text API) to convert voice data into text data with high accuracy.

[0184] 2. Natural language processing for context and relationship analysis

[0185] The converted text data is analyzed using natural language processing (e.g., the BERT model) to identify the context in which the utterance was made and the relationship between the user and the other party.

[0186] 3. Generating appropriate translations using machine translation

[0187] Uses a generative AI model (e.g. GPT-3) to generate an appropriate target language translation based on context and relationships.

[0188] 4. Feedback-driven optimization

[0189] Collect user feedback and optimize the parameters of the translation algorithm.

[0190] Specific example of operation procedure

[0191] Example 1: A delivery person communicating with a foreign customer

[0192] When the delivery person says into their smartphone, "Is this the correct address?", the smartphone captures the voice and sends the data to the server in real time.

[0193] The server receives the voice data and uses a deep learning model to convert it into text data such as "Is this address correct?"

[0194] The server uses natural language processing technology to identify that the spoken content is a question in a delivery situation and appropriately analyzes the communication with the customer.

[0195] Based on the analysis results, a generative AI model is used to generate the English translation "Is this the correct address?" and display it on the smartphone.

[0196] The delivery person checks the translation results, shows them to the customer, and provides feedback as needed to improve the system's accuracy.

[0197] Example prompt sentence:

[0198] "Please suggest a translation to ask the customer for confirmation during delivery."

[0199] "How do we understand our customers' needs and respond to them formally?"

[0200] "Please suggest a translation for a scenario in which a delivery person asks a customer where their pickup is located in a casual context."

[0201] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0202] Step 1:

[0203] User utterance transmission

[0204] The user (delivery person) speaks into a smartphone. The device (smartphone) captures the voice in real time and generates voice data. The voice data is compressed and sent to a server via the Internet. In this case, the input is the delivery person's speech, and the output is compressed voice data.

[0205] Step 2:

[0206] Audio data conversion

[0207] The server receives the voice data sent from the device. Then, it applies a deep learning-based speech recognition algorithm (e.g., Google Speech-to-Text API) to generate text data from the voice data. The input is compressed voice data, and the output is text data.

[0208] Step 3:

[0209] Context and relationship analysis

[0210] The server analyzes the generated text data. It uses a natural language processing model (e.g., BERT) to identify the context in which the user spoke and their relationship to the other party (customer). This process also references existing conversation history and user profile databases. The input is text data, and the output is information related to the context and relationships.

[0211] Step 4:

[0212] Generating Translations

[0213] The server uses a generative AI model (e.g., GPT-3) to generate an appropriate translation based on the information obtained from the context and relationships. For example, for business questions, it generates a translation using formal language. The input is information related to the context and relationships and text data, and the output is the text data translated into the target language.

[0214] Step 5:

[0215] Notification of translation results

[0216] The translation results are sent from the server to the device. The device displays the translation results to the user and notifies them by voice or text. The input is the translated text data, and the output is the translation results displayed on the device.

[0217] Step 6:

[0218] Feedback collection and optimization

[0219] The user checks the translation results and provides feedback if necessary. The device sends this feedback to the server, which analyzes the feedback and optimizes the parameters of the translation algorithm and generative AI model. The input is the user's feedback data, and the output is the optimized translation algorithm.

[0220] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0221] This invention is a context-adaptive real-time translation system that combines speech recognition, natural language processing, machine translation, and an emotion engine. The system acquires user utterances as voice data and converts the voice data into text data. In addition, the system analyzes the context, the relationship with the other person, and the user's emotions, and generates and provides an appropriate translation based on the analysis results.

[0222] Program processing procedures and explanations

[0223] 1. User utterance transmission

[0224] The user speaks into a device (smartphone, tablet, PC, etc.), which uses a built-in microphone to capture the user's speech in real time. This voice data is stored in a temporary memory in digital format, compressed, and then transmitted to a server via the Internet.

[0225] 2. Audio data conversion

[0226] The server analyzes the received voice data and applies a speech recognition algorithm, using a deep learning-based model to convert the voice data to text data with high accuracy, which is passed on to the next analysis step and stored in temporary memory.

[0227] 3. Context and Relationship Analysis

[0228] The server analyzes the generated text data and identifies the context of the user's speech and the relationship with the other party. Context analysis is performed using natural language processing technology, and a comprehensive assessment is made of the user's speech content, time, location, existing conversation history, etc. It also references a user profile database and past speech data to estimate the relationship between the user and the other party and select an appropriate communication style.

[0229] 4. Emotion Recognition Processing

[0230] The server uses an emotion engine to analyze the user's emotions from the voice data. The emotion engine analyzes the tone, pitch, and speed of the user's voice to identify the user's emotional state (e.g., joy, sadness, anger, etc.). In addition, the emotion engine also incorporates the user's facial expression recognition and biometric signal data (heart rate, skin potential, etc.) to perform more accurate emotion analysis.

[0231] 5. Generating the Translation

[0232] The server uses a generative AI model to translate the text data into the target language based on the results of contextual analysis and emotion recognition. Emotional information and writing style are also taken into consideration to generate a translation with appropriate expressions. The translated text data is stored in temporary memory.

[0233] 6. Gather feedback and optimize

[0234] The translation results are sent from the server to the device and notified to the user. The user can check the translation results on the device and provide feedback if necessary. The device then sends this feedback to the server. The server analyzes the feedback and optimizes the parameters of the translation algorithm and generative AI model to improve future translation accuracy.

[0235] Specific examples

[0236] Example 1: Use in business situations

[0237] During a meeting, a user says, "How is the progress on this project?" The device captures the speech and sends it to the server. The server converts the speech into text and recognizes the formal context of a business setting. The emotion engine also analyzes the user's tone to determine their seriousness. The result is translated into "How is the progress on this project?" in a formal and serious tone. The translation is then displayed on the device for the user to confirm.

[0238] Example 2: Use in everyday conversation

[0239] When a user is with friends, they say, "Where should we have dinner tonight?" The device captures the speech and sends it to the server. The server converts the speech to text and recognizes the casual context of everyday conversation. The emotion engine detects the joy in the user's voice. Based on this, the translation is performed, resulting in "Where should we have dinner tonight?" in a casual and joyful tone. The translation result is displayed on the device for the user to confirm.

[0240] The system of the present invention combines speech recognition, natural language processing, machine translation, and emotion recognition technologies to provide more natural and appropriate communication. Furthermore, it optimizes the algorithm based on user feedback to achieve high-accuracy translation over the long term.

[0241] The processing flow will be explained below.

[0242] Step 1:

[0243] The user speaks into the device, which uses a built-in microphone to capture the user's speech in real time, and this voice data is stored digitally in the device's temporary memory.

[0244] Step 2:

[0245] The device compresses the captured audio data and sends it to a server over the Internet, where it is encrypted for secure transmission.

[0246] Step 3:

[0247] The server receives the voice data sent from the terminal, decodes the received voice data, and prepares it for voice recognition processing.

[0248] Step 4:

[0249] The server applies a speech recognition algorithm to convert the audio data into text data. It uses a deep learning-based model to convert the speech into text with high accuracy. This text data is stored in the server's temporary memory.

[0250] Step 5:

[0251] The server then analyzes the generated text data using natural language processing technology to identify the context in which the utterance was made and the relationship with the other person. Context analysis takes into account the content of the utterance, time, location, and existing conversation history.

[0252] Step 6:

[0253] The server compares the user profile database and past conversation history to estimate the relationship between the user and the other party, and then determines the appropriate communication style, such as formal or casual.

[0254] Step 7:

[0255] The server uses an emotion engine to analyze the user's emotions from the voice data. The emotion engine analyzes the tone, pitch, and speed of the voice to identify the user's emotional state (e.g., joy, sadness, anger, etc.). It also uses facial expression recognition and biometric signal data as needed to improve the accuracy of the emotion detection.

[0256] Step 8:

[0257] Based on the results of contextual analysis and emotion recognition, the server uses a generative AI model to translate the text data into the target language, taking into account appropriate writing style and emotional expression. The generated translation data is stored in temporary memory.

[0258] Step 9:

[0259] The server sends the generated translation results to the terminal, which then displays the received translation results on the user interface. If necessary, speech synthesis technology can also be used to output the results as voice.

[0260] Step 10:

[0261] The user checks the translation results and provides feedback if necessary, which the device then sends to the server.

[0262] Step 11:

[0263] The server analyzes the received feedback and updates and optimizes the parameters of the translation algorithm and generative AI model, thereby improving future translation accuracy.

[0264] This allows the system of the present invention to translate user utterances in real time and provide appropriate translation results for various situations. Furthermore, by combining it with an emotion recognition engine, it achieves more natural translation that reflects the user's emotions. Optimization through feedback allows the system to be continuously improved, enabling highly accurate translation.

[0265] Example 2

[0266] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0267] Conventional translation systems struggle to accurately convert speech to text and to translate appropriately while taking into account the context and emotion of the translated text. They also lack effective means for collecting user feedback and improving translation accuracy. In particular, they are unable to accurately capture emotional nuances, making it difficult to achieve natural and appropriate communication.

[0268] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0269] In this invention, the server includes means for acquiring user speech as voice data, means for converting the voice data into text data, means for analyzing the context and the relationship with the other party, means for performing sentiment analysis and incorporating the results into the translation, and means for collecting user feedback and optimizing the translation algorithm. This not only improves the accuracy of voice data conversion, but also enables natural and appropriate translation that takes into account the relationship between the user and the other party, the context, and the emotions. Furthermore, translation accuracy can be continuously improved based on user feedback.

[0270] "User" refers to an individual who uses the system to make a speech.

[0271] A "terminal" is a device that a user uses to make a speech, and includes a smartphone, tablet, PC, etc.

[0272] "Voice Data" refers to information that digitally captures a user's speech.

[0273] "Text data" refers to character string information generated by analyzing voice data.

[0274] "Context" refers to the situation or background information obtained by comprehensively assessing the content of the user's utterance, time, location, existing conversation history, etc.

[0275] "Relationship" refers to the relationship between a user and a conversation partner, and is estimated by referring to past conversation history and a user profile database.

[0276] "Emotion analysis" refers to the process of analyzing a user's voice tone, pitch, rate, facial expression recognition, and biometric signal data to identify the user's emotional state.

[0277] "Translation" refers to the process of converting text data into a target language based on the analysis results.

[0278] "Feedback" refers to the evaluation and opinions of the translation results provided by the user.

[0279] A "generative AI model" refers to an algorithm that uses artificial intelligence technology to generate appropriate translations.

[0280] "Optimization" refers to the process of adjusting the system's parameters based on collected feedback to improve translation accuracy.

[0281] A "prompt sentence" is a portion of the text data input into a generative AI model, and refers to the phrase that serves as the basis for the model to generate a translation.

[0282] "Machine learning-based models" refer to trained algorithms or neural networks used to analyze audio data or generate text data.

[0283] "Natural language processing" refers to technology for analyzing the context and meaning of text data.

[0284] This invention is a context-adaptive real-time translation system that combines speech recognition, natural language processing, machine translation, and an emotion engine. Specifically, it uses the following hardware and software:

[0285] A user speaks into a device (such as a smartphone, tablet, or PC) and their speech is captured as audio data in real time. The device uses a built-in microphone and audio capture module to temporarily store this audio data in digital format (WAV format), compress it (e.g., to MP3 format), and transmit it to a server via the Internet.

[0286] The server analyzes the received voice data and converts it into text data using a speech recognition algorithm. Specifically, it uses a machine learning-based model (e.g., a model using an open-source deep learning library) to convert the voice data into text data. This text data is passed to the next analysis step and stored in temporary memory.

[0287] The server then analyzes the generated text data to identify the context of the user's speech and the relationship between the user and the other party. This context analysis uses natural language processing technology (e.g., BERT or GPT-based models). The server comprehensively determines the context by taking into account the content of the speech, time, location, existing conversation history, etc. It also estimates the relationship between the user and the other party by referencing a user profile database and past speech data.

[0288] Furthermore, the server uses an emotion engine to analyze the user's emotions from the voice data. The emotion engine analyzes the user's voice tone, pitch, speed, facial expression recognition, and bio-signal data (e.g., heart rate, skin potential) to identify the user's emotional state. A deep learning model is used for emotion analysis.

[0289] Based on the results of contextual analysis and emotion recognition, the server uses a generative AI model to translate the text data into the target language. Emotional information and writing style are also taken into consideration to generate a translation with appropriate expressions. For example, OpenAI's GPT-based model is used as the generative AI model. The generated translation data is stored in temporary memory.

[0290] The translation results are sent from the server to the device and notified to the user. The device is equipped with a notification function, allowing the user to check the translation results and provide feedback if necessary. Using the feedback function, users' ratings and opinions are entered into the device and sent to the server. The server analyzes this feedback and optimizes the parameters of the translation algorithm and generative AI model to improve future translation accuracy.

[0291] Specific use cases include the following:

[0292] Use in business situations

[0293] During a meeting, a user says, "How is the progress on this project?" The device captures the speech and sends it to the server. The server converts the speech into text and recognizes the formal context of a business setting. The emotion engine also analyzes the user's tone to determine seriousness. The result is translated into "How is the progress on this project?" in a formal and serious tone. The translation is then displayed on the device for the user to confirm.

[0294] Use in everyday conversation

[0295] When a user is with friends, they say, "Where should we have dinner tonight?" The device captures the speech and sends it to the server. The server converts the speech to text and recognizes the casual context of everyday conversation. The emotion engine detects the joy in the user's voice. Based on this, the translation is performed, resulting in "Where should we have dinner tonight?" in a casual and joyful tone. The translation result is displayed on the device for the user to confirm.

[0296] In this way, by using this system, more natural and appropriate translations are provided, enabling users to communicate smoothly. In addition, the feedback function allows for continuous system improvement, ensuring high translation accuracy over the long term.

[0297] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0298] Step 1:

[0299] The user speaks into the device. The device uses the built-in microphone to capture the user's speech in real time. The voice capture module is activated and captures the user's speech as digital audio data in WAV format. This audio data is stored in temporary memory.

[0300] Step 2:

[0301] The device compresses the captured audio data. Specifically, the audio data is converted to MP3 format and compressed. This compressed audio data is sent to a server via the Internet. The input is WAV format audio data, and the output is compressed MP3 format audio data.

[0302] Step 3:

[0303] The server analyzes the received audio data and applies a speech recognition algorithm. Specifically, the server uses a machine learning-based model (e.g., a model using a deep learning library) to convert the audio data into text data. The input is audio data in MP3 format, and the output is text data.

[0304] Step 4:

[0305] The server analyzes the generated text data and identifies the context and relationships with the other party. This analysis uses natural language processing technology (e.g., BERT or GPT-based models). The text data is input as a prompt, and the context and relationships are derived. The input is text data, and the output is context information and relationship information.

[0306] Step 5:

[0307] The server uses an emotion engine to analyze the user's emotions from the voice data. It identifies the user's emotional state by analyzing the tone, pitch, and rate of the user's voice, and also incorporating facial expression recognition and biosignal data (e.g., heart rate and skin potential). The input is text data and biosignal data, and the output is emotional state data.

[0308] Step 6:

[0309] The server uses a generative AI model to translate the text data into the target language based on the results of contextual analysis and emotion recognition. It generates an appropriate translation taking into account emotional information and writing style. The input is text data, context information, relationship information, and emotional state data, and the output is the translated text data.

[0310] Step 7:

[0311] The translation result is sent from the server to the terminal and notified to the user. The terminal uses a notification function to display the translation result to the user. The input is the translated text data, and the output is the translation result displayed to the user.

[0312] Step 8:

[0313] The user checks the translation results on the terminal and provides feedback if necessary. The terminal displays a feedback input form and sends the feedback entered by the user to the server. The input is the feedback data from the user, and the output is the feedback data sent to the server.

[0314] Step 9:

[0315] The server analyzes user feedback and optimizes the parameters of the translation algorithm and generative AI model, thereby improving future translation accuracy. The input is the feedback data, and the output is the optimized algorithm parameters.

[0316] (Application example 2)

[0317] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0318] In recent years, the use of self-driving vehicles has become more widespread, making appropriate communication with passengers crucial. However, current self-driving vehicle systems face challenges, such as language barriers and the inability to properly recognize and respond to emotions, which can reduce user satisfaction. Therefore, there is a need for a system that can analyze context and emotions in real time, translate them into the appropriate language, and even control in-vehicle functions.

[0319] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring a user's utterance as voice data, means for converting the voice data into text data, means for analyzing the context, the relationship with the other party, and the user's emotions, means for translating the text data into a target language based on the analysis results, means for notifying the user of the translation results, and means for controlling functions in the autonomous vehicle based on the user's utterance. This enables passengers to communicate naturally and appropriately in the autonomous vehicle and efficiently operate the vehicle's functions.

[0320] "User utterance" refers to words or instructions expressed by the user through speech.

[0321] "Voice data" refers to data that is a digital recording of a user's speech.

[0322] "Text data" is character information obtained by analyzing voice data.

[0323] "Context" refers to the linguistic background that is understood by taking into account the content of the user's speech, the surrounding circumstances, past speech history, and so on.

[0324] "Relationship with the other party" refers to the relationship between the user and the target person or system.

[0325] "Emotion" is a psychological state inferred based on the user's voice and other biometric signals.

[0326] "Target language" is the language to which the text is to be translated.

[0327] The "translation result" is the text resulting from converting the original text data into the target language.

[0328] An "autonomous vehicle" is a vehicle that is driven automatically without human intervention.

[0329] "Means for controlling functions" means methods or devices for operating or adjusting various devices or settings within an automated vehicle.

[0330] The system that realizes this application example includes the following components: Acquires user speech as voice data, converts the voice data into text data, and analyzes the context, the relationship with the other person, and the user's emotions. Based on the analysis results, translates the text data into the target language, notifies the user of the translation result, and controls functions within the autonomous vehicle.

[0331] Hardware and software used

[0332] 1. User utterance acquisition

[0333] Hardware: Self-driving vehicle system with microphone

[0334] Software: sounddevice library

[0335] The microphone inside the autonomous vehicle captures the user's speech in real time, temporarily stores it in the built-in memory as voice data, and then transmits it to the server. The user can interact with the system by speaking into the microphone.

[0336] 2. Audio data conversion

[0337] Hardware: Server or in-vehicle computer

[0338] Software: speech_recognition library

[0339] The server analyzes the received voice data and uses a deep learning-based model to convert the voice data into highly accurate text data, which is then passed on to the next analysis step.

[0340] 3. Context, Relational, and Sentiment Analysis

[0341] Hardware: Server or in-vehicle computer

[0342] Software: Natural language processing engine, emotion recognition module

[0343] The server analyzes the generated text data and analyzes the context of the user's speech, the relationship with the other person, and their emotions. It does this by using natural language processing technology and referring to existing conversation history and user profiles. It also uses an emotion analysis engine to identify the user's emotional state from the tone, pitch, and speed of their voice.

[0344] 4. Generating the Translation

[0345] Hardware: Server or in-vehicle computer

[0346] Software: translation engines (e.g., googletrans library), generative AI models

[0347] The server uses a generative AI model to translate the text data into the target language based on the results of contextual analysis and emotion recognition. An appropriate translation is generated that takes into account emotional information and writing style. The translated text data is stored in temporary memory.

[0348] 5. Notification of translation results

[0349] Hardware: Smartphones, tablets, dedicated devices

[0350] Software: Translation result display module

[0351] The translation results are sent from the server to the terminal and notified to the user, who can then check the translation results on the terminal screen.

[0352] 6. Control of functions inside autonomous vehicles

[0353] Hardware: Autonomous vehicle control systems

[0354] Software: Vehicle Control Module

[0355] The system controls in-vehicle functions (e.g., temperature adjustment, seat arrangement, destination change, etc.) based on the user's speech, allowing the user to efficiently operate in-vehicle functions using only speech.

[0356] Specific examples

[0357] Example 1:

[0358] When a passenger says "Adjust the air conditioning" in an autonomous vehicle, the system captures the speech and converts it into text. The server analyzes the context and sentiment of the text and translates it as "Adjust the air conditioning." The translation result is displayed on the user's device, and the air conditioning system in the vehicle is simultaneously adjusted.

[0359] Example prompt sentence:

[0360] Voice translation: Adjust the air conditioning

[0361] This will enable users to communicate naturally and smoothly within an autonomous vehicle and efficiently operate the vehicle's functions.

[0362] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0363] Step 1:

[0364] The user speaks. The user issues instructions or questions by voice into the microphone inside the autonomous vehicle. The input is voice data.

[0365] Step 2:

[0366] The device captures the user's speech and stores it in temporary memory as audio data. The input is audio data, and the output is a digital audio file. Specifically, it uses the sounddevice library to record audio and saves the data to a file.

[0367] Step 3:

[0368] The device compresses the stored audio data and transmits it to a server via the Internet. The input is a digital audio file, and the output is compressed audio data. Specifically, the audio data is compressed using a compression algorithm and then transmitted to the server.

[0369] Step 4:

[0370] The server analyzes the received voice data and converts it into text data by applying a speech recognition algorithm. The input is compressed voice data and the output is text data. Specifically, the server uses the speech_recognition library to convert voice to text data.

[0371] Step 5:

[0372] The server analyzes the generated text data to identify the context and the relationship with the other party. The input is the text data, and the output is the analyzed context and relationship information. The server uses a natural language processing engine to analyze the text and reference past conversation history and user profiles.

[0373] Step 6:

[0374] The server uses an emotion recognition engine to analyze the user's emotions. The input is voice data and text data, and the output is analyzed emotional information. Specifically, the server analyzes the tone, pitch, and speed of the voice to identify the user's emotional state.

[0375] Step 7:

[0376] The server translates the text data into the target language using a generative AI model based on the results of contextual analysis and emotion recognition. The input is the text data and the analysis results, and the output is the translated text data. The text is translated into the target language using the googletrans library.

[0377] Step 8:

[0378] The server sends the translation results to the terminal and notifies the user. The input is the translated text data, and the output is the translation results displayed to the user. The terminal displays the translation results on the screen so that the user can check them.

[0379] Step 9:

[0380] The device controls functions within the autonomous vehicle based on the user's speech. The input is translated text data and user instructions, and the output is changes to the vehicle's functions. Specifically, it sends signals to control the vehicle's air conditioning system, navigation system, etc.

[0381] Through the above processing steps, users can achieve natural and smooth communication within an autonomous vehicle and efficiently operate the vehicle's functions.

[0382] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0383] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0384] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0385] [Second embodiment]

[0386] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0387] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0388] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0389] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0390] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0391] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0392] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0393] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0394] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0395] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0396] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0397] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0398] This invention is a context-adaptive real-time translation system that uses speech recognition, natural language processing, and machine translation technologies. Specifically, the system acquires user utterances as voice data, converts the voice data into text data, analyzes the context and the relationship with the other party, and generates an appropriate translation based on the analysis results.

[0399] Program processing procedures and explanations

[0400] 1. User utterance transmission

[0401] The user speaks into a device (smartphone, tablet, PC, etc.), which captures the voice in real time and records it as audio data. The recorded audio data is compressed and sent to a server via the Internet.

[0402] 2. Audio data conversion

[0403] The server receives the voice data sent from the device and applies a speech recognition algorithm, which uses a deep learning-based model to convert the user's speech into text data with high accuracy. This process converts the voice data into text data and stores it in a temporary file or memory.

[0404] 3. Context and Relationship Analysis

[0405] The server analyzes the generated text data and estimates the context in which the user spoke and the relationship with the other person. Context analysis uses natural language processing technology to identify the scene by taking into account the content of the user's speech, time, location, and existing conversation history. The server also references a user profile database and past speech data to estimate the relationship between the user and the target person.

[0406] 4. Generating the Translation

[0407] The server uses a generative AI model to generate an appropriate translation based on information obtained from context and relationships. For example, it generates formal expressions for business situations and casual expressions for everyday conversations. The translation results are stored in temporary memory on the server.

[0408] 5. Gather feedback and optimize

[0409] The translation results are sent to the device and notified to the user. The user reviews the translation results and provides feedback if necessary. The device then sends this feedback to the server. The server analyzes the feedback and optimizes the parameters of the translation algorithm and generative AI model. This process improves future translation accuracy.

[0410] Specific examples

[0411] Example 1: Use in business situations

[0412] During a meeting, a user says, "How is the progress on this project?" The device captures the speech and sends it to the server. The server converts the speech into text and translates it, recognizing the formal context of the business setting. The resulting translation is "How is the progress on this project?" The translation is displayed on the device for the user to review.

[0413] Example 2: Use in everyday conversation

[0414] A user is with friends and says, "Where should we have dinner tonight?" The device captures the speech and sends it to the server. The server converts the speech to text and translates it, recognizing the casual context of everyday conversation. The resulting translation is "Where should we have dinner tonight?" The translation is displayed on the device for the user to confirm.

[0415] This invention combines speech recognition, natural language processing, and machine translation technologies to provide real-time translation suitable for a variety of situations. Furthermore, by optimizing the translation algorithm based on user feedback, it is possible to achieve more accurate translation.

[0416] The processing flow will be explained below.

[0417] Step 1:

[0418] The user speaks into the device, which uses a built-in microphone to capture the user's speech in real time, and this voice data is stored digitally in the device's temporary memory.

[0419] Step 2:

[0420] The device compresses the captured audio data and sends it over the internet to a server, where it is encrypted for security purposes.

[0421] Step 3:

[0422] The server receives the voice data sent from the terminal, decodes the received voice data, and prepares it for voice recognition processing.

[0423] Step 4:

[0424] The server applies a speech recognition algorithm to convert the audio data into text data, using a deep learning-based model to convert spoken content into text with high accuracy.

[0425] Step 5:

[0426] The server analyzes the generated text data and identifies contextual information and relationships with the other party. Contextual analysis uses natural language processing technology. This analysis categorizes the scene into categories such as business or everyday conversation.

[0427] Step 6:

[0428] The server references the user profile database and past conversation history to estimate the relationship between the user and the other party, and based on this, selects an appropriate communication style (formal, casual, etc.).

[0429] Step 7:

[0430] The server retrieves translation settings based on the analysis results and uses a generative AI model to translate the text data into the target language, adjusting the writing style and tone accordingly.

[0431] Step 8:

[0432] The server then sends the translation results to the device, which then displays them on the user interface. If necessary, speech synthesis technology can also be used to output the results as voice.

[0433] Step 9:

[0434] The user checks the translation results and provides feedback if necessary, which is sent from the device to the server.

[0435] Step 10:

[0436] The server analyzes the received feedback and updates and optimizes the parameters of the translation algorithm and generative AI model, thereby improving future translation accuracy.

[0437] In this way, this system translates user utterances in real time and provides appropriate translation results for a variety of situations. By optimizing the algorithm based on user feedback, it achieves high-accuracy translation over the long term.

[0438] Example 1

[0439] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0440] Conventional translation systems often perform simple translations without considering the context or the relationship with the target user, which can result in low translation accuracy. In addition, they lack the ability to optimize the system based on user feedback, making it difficult to continuously improve translation accuracy.

[0441] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0442] In this invention, the server includes means for acquiring user utterances as voice data, means for converting the voice data into text data, means for analyzing the relationship between the context and the target person, means for translating the text data into the target language based on the analysis results, means for notifying the user of the translation results, and means for collecting feedback from the user and optimizing the translation algorithm. This enables highly accurate translation that adapts to the context and relationship, and continuous system optimization is achieved by utilizing the feedback.

[0443] The "means for acquiring user speech as voice data" is a function for electronically capturing the voice uttered by the user and recording it as digital voice data.

[0444] The "means for converting voice data into text data" is a function for analyzing captured voice data and converting it into text data expressed as a corresponding character string.

[0445] The "means for analyzing the relationship between the context and the target person" is an analytical function for estimating the relationship between the context of an utterance and the target person of that utterance from the generated text data.

[0446] The "means for translating text data into a target language" refers to a translation function that converts the original text data into a different language based on the analyzed context and relationships.

[0447] The "means for notifying the user of the translation result" is a function for visually or audibly notifying the user of the generated translation result.

[0448] "Means for collecting user feedback and optimizing the translation algorithm" refers to a function that collects user evaluations and suggestions for corrections on translation results, and then adjusts the parameters and settings of the translation algorithm based on that information to improve accuracy.

[0449] A "deep learning-based model" is a type of algorithm that uses a multi-layer neural network to extract and transform data features, enabling advanced recognition and prediction.

[0450] A "generative AI model" is an algorithm that uses artificial intelligence technology to generate text or translate, and has the ability to generate appropriate output based on a prompt.

[0451] The present invention is a context-adaptive real-time translation system that uses speech recognition, natural language processing, and machine translation technologies. Specific embodiments for carrying out the present invention are described below.

[0452] First, the user speaks into the device (smartphone, tablet, PC, etc.). The device uses a built-in microphone to capture the user's voice in real time and record it as audio data. The recorded audio data is saved in a lossless compressed format such as FLAC and sent to a server via the Internet.

[0453] The server receives the voice data sent via the Internet, applies a speech recognition algorithm (e.g., Google Speech-to-Text API using a deep learning-based model) to analyze the voice data and convert it into text data, which is then saved in a temporary file or memory.

[0454] The server then analyzes the generated text data and estimates the context of the utterance and its relationship to the target person. This analysis uses natural language processing technology (e.g., the spaCy library). The server references a user profile database and a database of past utterances to identify the scene, taking into account the user's utterance content, time, location, and existing conversation history. It also references these databases to identify the relationship between the user and the target person.

[0455] The server generates an appropriate prompt based on the context and relationships, and generates the translation using a generative AI model (e.g., OpenAI GPT-3). The AI ​​model is fed with a prompt like this:

[0456] Input: "How is this project progressing?"

[0457] Context: Business meeting

[0458] Relationship: Superior-Subordinate

[0459] Produces the translation: "How is the progress on this project?"

[0460] The generated translation result is stored in temporary memory in the server and then sent to the terminal, where the user can check the translation result on the terminal screen or through voice output.

[0461] Finally, the user can review the translation results and provide feedback, such as ratings and suggestions for corrections. The device collects this feedback and sends it back to the server. The server analyzes the feedback and optimizes the parameters of the translation algorithm and generative AI model, thereby improving future translation accuracy.

[0462] As a concrete example, consider the case where a user says, "How is the progress on this project?" during a meeting. At this time, the device captures the audio and sends it to the server in FLAC format. The server converts it into text using the Google Speech-to-Text API and analyzes the context of the business meeting and the relationship between superiors and subordinates. Based on the analysis results, a prompt is sent to the GPT-3 model, which generates the translation result, "How is the progress on this project?". This translation result is sent to the device and confirmed by the user.

[0463] As another example, consider the case where a user is with friends and says, "Where should we have dinner tonight?" The device captures the audio and sends it to the server in FLAC format. The server converts it to text using the Google Speech-to-Text API and analyzes the casual context of everyday conversation and friendships. Based on the analysis results, the prompt is sent to the GPT-3 model, which generates the translation result, "Where should we have dinner tonight?" The translation result is then sent to the device for the user to confirm.

[0464] In this way, by combining speech recognition, natural language processing, and generative AI models, the present invention can provide highly accurate real-time translation that adapts to the user's context and relationships. Furthermore, by optimizing the translation algorithm based on user feedback, the system can be continuously improved.

[0465] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0466] Step 1: User utterance submission

[0467] explanation

[0468] The user speaks into a device (smartphone, tablet, PC, etc.), which then uses a built-in microphone to capture the voice in real time and record it as digital audio data.

[0469] input

[0470] User utterance: "How is this project going?"

[0471] output

[0472] FLAC format audio data

[0473] Specific actions

[0474] The user speaks to their smartphone, "How is this project progressing?" The device uses its built-in microphone to capture the audio and saves it as a FLAC audio file. The saved audio data is then sent to a server over the Internet.

[0475] Step 2: Convert the audio data

[0476] explanation

[0477] The server receives the voice data sent from the terminal and converts it into text data by applying a voice recognition algorithm.

[0478] input

[0479] FLAC format audio data

[0480] output

[0481] Text data: "How is this project progressing?"

[0482] Specific actions

[0483] The server receives FLAC audio data from the device. The server calls the Google Speech-to-Text API, analyzes the audio data, and converts it into text data such as "How is the progress on this project?" The converted text data is saved as a temporary file.

[0484] Step 3: Analyze context and relationships

[0485] explanation

[0486] The server analyzes the generated text data and estimates the context of the utterance and its relationship to the target person.

[0487] input

[0488] Text data: "How is this project progressing?"

[0489] output

[0490] Analysis results (context: business meeting, relationship: superior-subordinate)

[0491] Specific actions

[0492] The server analyzes the text data using the spaCy library, infers that the context of the utterance is a business meeting, and references a database of past utterances and a user profile database to determine that the user is a superior and the target person is a subordinate.

[0493] Step 4: Generate translations

[0494] explanation

[0495] The server generates prompt sentences based on context and relationships, and uses generative AI models to generate appropriate translations.

[0496] input

[0497] Analysis results (context: business meeting, relationship: superior-subordinate)

[0498] output

[0499] Translation result: "How is the progress on this project?"

[0500] Specific actions

[0501] The server inputs the following prompt into the generative AI model (OpenAI GPT-3):

[0502] Input: "How is this project progressing?"

[0503] Context: Business meeting

[0504] Relationship: Superior-Subordinate

[0505] Based on this prompt, the GPT-3 model generates the translation "How is the progress on this project?" This translation result is stored in temporary memory on the server.

[0506] Step 5: Notification of translation results

[0507] explanation

[0508] The server sends the generated translation result to the terminal and notifies the user.

[0509] input

[0510] Translation result: "How is the progress on this project?"

[0511] output

[0512] Terminal display: "How is the progress on this project?"

[0513] Specific actions

[0514] The server sends the generated translation to the device as an HTTPS response, and the user can check the translation result, "How is the progress on this project?", displayed on the device screen.

[0515] Step 6: Gather feedback and optimize

[0516] explanation

[0517] The user provides feedback on the translation result, and the terminal sends the feedback to the server, which analyzes the feedback and optimizes the translation algorithm.

[0518] input

[0519] User feedback: "The translation was accurate"

[0520] output

[0521] Optimized translation algorithm

[0522] Specific actions

[0523] The user reviews the translation result and provides feedback, rating it "the translation was accurate." The device then sends this feedback to the server, which analyzes the feedback and optimizes the translation algorithm by adjusting the model settings, resulting in higher accuracy in the next translation.

[0524] (Application example 1)

[0525] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0526] Food delivery services face the challenge of having delivery personnel communicate smoothly with foreign customers. In particular, the language barrier makes it difficult for delivery personnel to accurately understand order details and delivery instructions. This problem reduces delivery efficiency and reduces customer satisfaction.

[0527] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0528] In this invention, the server includes means for translating information and questions received by the delivery person from the customer in real time, means for acquiring the user's speech as voice data, means for converting the voice data into text data, means for translating the text data into the target language based on the analysis results, and means for notifying the user of the translation results. This enables the delivery person to communicate smoothly with foreign customers, improving delivery efficiency and customer satisfaction.

[0529] The "means for acquiring user speech as voice data" refers to a device or function for recording the user's speech as digital voice data.

[0530] The "means for converting voice data into text data" refers to an algorithm or software for analyzing acquired voice data and converting it into corresponding written text.

[0531] "Means for analyzing the context and the relationship with the other person" refers to natural language processing technologies and algorithms for understanding and analyzing the content of an utterance, its background, and the relationship between the speaker and the other person.

[0532] The "means for translating text data into a target language based on the analysis results" refers to a machine translation technology for generating text in a target language that is appropriately translated based on context and relationships.

[0533] "Means for notifying the user of the translation result" refers to an interface or function for displaying or audibly conveying the generated translation result to the user.

[0534] "A means for delivery personnel to translate information and questions received from customers in real time" refers to a system or algorithm that instantly translates the content of communication between delivery personnel and customers into other languages.

[0535] System Program

[0536] Hardware and Software Configuration

[0537] The system for implementing this invention includes a terminal operated by a user (e.g., a smartphone), a server connected via the Internet, and algorithms for performing speech recognition, natural language processing, and machine translation. The terminal has the function of capturing the user's voice and sending the voice data to the server. The server receives the voice data and performs the following main processes:

[0538] 1. Speech Recognition Using Deep Learning-Based Models

[0539] Use deep learning-based speech recognition algorithms (e.g., Google Speech-to-Text API) to convert voice data into text data with high accuracy.

[0540] 2. Natural language processing for context and relationship analysis

[0541] The converted text data is analyzed using natural language processing (e.g., the BERT model) to identify the context in which the utterance was made and the relationship between the user and the other party.

[0542] 3. Generating appropriate translations using machine translation

[0543] Uses a generative AI model (e.g. GPT-3) to generate an appropriate target language translation based on context and relationships.

[0544] 4. Feedback-driven optimization

[0545] Collect user feedback and optimize the parameters of the translation algorithm.

[0546] Specific example of operation procedure

[0547] Example 1: A delivery person communicating with a foreign customer

[0548] When the delivery person says into their smartphone, "Is this the correct address?", the smartphone captures the voice and sends the data to the server in real time.

[0549] The server receives the voice data and uses a deep learning model to convert it into text data such as "Is this address correct?"

[0550] The server uses natural language processing technology to identify that the spoken content is a question in a delivery situation and appropriately analyzes the communication with the customer.

[0551] Based on the analysis results, a generative AI model is used to generate the English translation "Is this the correct address?" and display it on the smartphone.

[0552] The delivery person checks the translation results, shows them to the customer, and provides feedback as needed to improve the system's accuracy.

[0553] Example prompt sentence:

[0554] "Please suggest a translation to ask the customer for confirmation during delivery."

[0555] "How do we understand our customers' needs and respond to them formally?"

[0556] "Please suggest a translation for a scenario in which a delivery person asks a customer where their pickup is located in a casual context."

[0557] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0558] Step 1:

[0559] User utterance transmission

[0560] The user (delivery person) speaks into a smartphone. The device (smartphone) captures the voice in real time and generates voice data. The voice data is compressed and sent to a server via the Internet. In this case, the input is the delivery person's speech, and the output is compressed voice data.

[0561] Step 2:

[0562] Audio data conversion

[0563] The server receives the voice data sent from the device. Then, it applies a deep learning-based speech recognition algorithm (e.g., Google Speech-to-Text API) to generate text data from the voice data. The input is compressed voice data, and the output is text data.

[0564] Step 3:

[0565] Context and relationship analysis

[0566] The server analyzes the generated text data. It uses a natural language processing model (e.g., BERT) to identify the context in which the user spoke and their relationship to the other party (customer). This process also references existing conversation history and user profile databases. The input is text data, and the output is information related to the context and relationships.

[0567] Step 4:

[0568] Generating Translations

[0569] The server uses a generative AI model (e.g., GPT-3) to generate an appropriate translation based on the information obtained from the context and relationships. For example, for business questions, it generates a translation using formal language. The input is information related to the context and relationships and text data, and the output is the text data translated into the target language.

[0570] Step 5:

[0571] Notification of translation results

[0572] The translation results are sent from the server to the device. The device displays the translation results to the user and notifies them by voice or text. The input is the translated text data, and the output is the translation results displayed on the device.

[0573] Step 6:

[0574] Feedback collection and optimization

[0575] The user checks the translation results and provides feedback if necessary. The device sends this feedback to the server, which analyzes the feedback and optimizes the parameters of the translation algorithm and generative AI model. The input is the user's feedback data, and the output is the optimized translation algorithm.

[0576] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0577] This invention is a context-adaptive real-time translation system that combines speech recognition, natural language processing, machine translation, and an emotion engine. The system acquires user utterances as voice data and converts the voice data into text data. In addition, the system analyzes the context, the relationship with the other person, and the user's emotions, and generates and provides an appropriate translation based on the analysis results.

[0578] Program processing procedures and explanations

[0579] 1. User utterance transmission

[0580] The user speaks into a device (smartphone, tablet, PC, etc.), which uses a built-in microphone to capture the user's speech in real time. This voice data is stored in a temporary memory in digital format, compressed, and then transmitted to a server via the Internet.

[0581] 2. Audio data conversion

[0582] The server analyzes the received voice data and applies a speech recognition algorithm, using a deep learning-based model to convert the voice data to text data with high accuracy, which is passed on to the next analysis step and stored in temporary memory.

[0583] 3. Context and Relationship Analysis

[0584] The server analyzes the generated text data and identifies the context of the user's speech and the relationship with the other party. Context analysis is performed using natural language processing technology, and a comprehensive assessment is made of the user's speech content, time, location, existing conversation history, etc. It also references a user profile database and past speech data to estimate the relationship between the user and the other party and select an appropriate communication style.

[0585] 4. Emotion Recognition Processing

[0586] The server uses an emotion engine to analyze the user's emotions from the voice data. The emotion engine analyzes the tone, pitch, and speed of the user's voice to identify the user's emotional state (e.g., joy, sadness, anger, etc.). In addition, the emotion engine also incorporates the user's facial expression recognition and biometric signal data (heart rate, skin potential, etc.) to perform more accurate emotion analysis.

[0587] 5. Generating the Translation

[0588] The server uses a generative AI model to translate the text data into the target language based on the results of contextual analysis and emotion recognition. Emotional information and writing style are also taken into consideration to generate a translation with appropriate expressions. The translated text data is stored in temporary memory.

[0589] 6. Gather feedback and optimize

[0590] The translation results are sent from the server to the device and notified to the user. The user can check the translation results on the device and provide feedback if necessary. The device then sends this feedback to the server. The server analyzes the feedback and optimizes the parameters of the translation algorithm and generative AI model to improve future translation accuracy.

[0591] Specific examples

[0592] Example 1: Use in business situations

[0593] During a meeting, a user says, "How is the progress on this project?" The device captures the speech and sends it to the server. The server converts the speech into text and recognizes the formal context of a business setting. The emotion engine also analyzes the user's tone to determine their seriousness. The result is translated into "How is the progress on this project?" in a formal and serious tone. The translation is then displayed on the device for the user to confirm.

[0594] Example 2: Use in everyday conversation

[0595] When a user is with friends, they say, "Where should we have dinner tonight?" The device captures the speech and sends it to the server. The server converts the speech to text and recognizes the casual context of everyday conversation. The emotion engine detects the joy in the user's voice. Based on this, the translation is performed, resulting in "Where should we have dinner tonight?" in a casual and joyful tone. The translation result is displayed on the device for the user to confirm.

[0596] The system of the present invention combines speech recognition, natural language processing, machine translation, and emotion recognition technologies to provide more natural and appropriate communication. Furthermore, it optimizes the algorithm based on user feedback to achieve high-accuracy translation over the long term.

[0597] The processing flow will be explained below.

[0598] Step 1:

[0599] The user speaks into the device, which uses a built-in microphone to capture the user's speech in real time, and this voice data is stored digitally in the device's temporary memory.

[0600] Step 2:

[0601] The device compresses the captured audio data and sends it to a server over the Internet, where it is encrypted for secure transmission.

[0602] Step 3:

[0603] The server receives the voice data sent from the terminal, decodes the received voice data, and prepares it for voice recognition processing.

[0604] Step 4:

[0605] The server applies a speech recognition algorithm to convert the audio data into text data. It uses a deep learning-based model to convert the speech into text with high accuracy. This text data is stored in the server's temporary memory.

[0606] Step 5:

[0607] The server then analyzes the generated text data using natural language processing technology to identify the context in which the utterance was made and the relationship with the other person. Context analysis takes into account the content of the utterance, time, location, and existing conversation history.

[0608] Step 6:

[0609] The server compares the user profile database and past conversation history to estimate the relationship between the user and the other party, and then determines the appropriate communication style, such as formal or casual.

[0610] Step 7:

[0611] The server uses an emotion engine to analyze the user's emotions from the voice data. The emotion engine analyzes the tone, pitch, and speed of the voice to identify the user's emotional state (e.g., joy, sadness, anger, etc.). It also uses facial expression recognition and biometric signal data as needed to improve the accuracy of the emotion detection.

[0612] Step 8:

[0613] Based on the results of contextual analysis and emotion recognition, the server uses a generative AI model to translate the text data into the target language, taking into account appropriate writing style and emotional expression. The generated translation data is stored in temporary memory.

[0614] Step 9:

[0615] The server sends the generated translation results to the terminal, which then displays the received translation results on the user interface. If necessary, speech synthesis technology can also be used to output the results as voice.

[0616] Step 10:

[0617] The user checks the translation results and provides feedback if necessary, which the device then sends to the server.

[0618] Step 11:

[0619] The server analyzes the received feedback and updates and optimizes the parameters of the translation algorithm and generative AI model, thereby improving future translation accuracy.

[0620] This allows the system of the present invention to translate user utterances in real time and provide appropriate translation results for various situations. Furthermore, by combining it with an emotion recognition engine, it achieves more natural translation that reflects the user's emotions. Optimization through feedback allows the system to be continuously improved, enabling highly accurate translation.

[0621] Example 2

[0622] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0623] Conventional translation systems struggle to accurately convert speech to text and to translate appropriately while taking into account the context and emotion of the translated text. They also lack effective means for collecting user feedback and improving translation accuracy. In particular, they are unable to accurately capture emotional nuances, making it difficult to achieve natural and appropriate communication.

[0624] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0625] In this invention, the server includes means for acquiring user speech as voice data, means for converting the voice data into text data, means for analyzing the context and the relationship with the other party, means for performing sentiment analysis and incorporating the results into the translation, and means for collecting user feedback and optimizing the translation algorithm. This not only improves the accuracy of voice data conversion, but also enables natural and appropriate translation that takes into account the relationship between the user and the other party, the context, and the emotions. Furthermore, translation accuracy can be continuously improved based on user feedback.

[0626] "User" refers to an individual who uses the system to make a speech.

[0627] A "terminal" is a device that a user uses to make a speech, and includes a smartphone, tablet, PC, etc.

[0628] "Voice Data" refers to information that digitally captures a user's speech.

[0629] "Text data" refers to character string information generated by analyzing voice data.

[0630] "Context" refers to the situation or background information obtained by comprehensively assessing the content of the user's utterance, time, location, existing conversation history, etc.

[0631] "Relationship" refers to the relationship between a user and a conversation partner, and is estimated by referring to past conversation history and a user profile database.

[0632] "Emotion analysis" refers to the process of analyzing a user's voice tone, pitch, rate, facial expression recognition, and biometric signal data to identify the user's emotional state.

[0633] "Translation" refers to the process of converting text data into a target language based on the analysis results.

[0634] "Feedback" refers to the evaluation and opinions of the translation results provided by the user.

[0635] A "generative AI model" refers to an algorithm that uses artificial intelligence technology to generate appropriate translations.

[0636] "Optimization" refers to the process of adjusting the system's parameters based on collected feedback to improve translation accuracy.

[0637] A "prompt sentence" is a portion of the text data input into a generative AI model, and refers to the phrase that serves as the basis for the model to generate a translation.

[0638] "Machine learning-based models" refer to trained algorithms or neural networks used to analyze audio data or generate text data.

[0639] "Natural language processing" refers to technology for analyzing the context and meaning of text data.

[0640] This invention is a context-adaptive real-time translation system that combines speech recognition, natural language processing, machine translation, and an emotion engine. Specifically, it uses the following hardware and software:

[0641] A user speaks into a device (such as a smartphone, tablet, or PC) and their speech is captured as audio data in real time. The device uses a built-in microphone and audio capture module to temporarily store this audio data in digital format (WAV format), compress it (e.g., to MP3 format), and transmit it to a server via the Internet.

[0642] The server analyzes the received voice data and converts it into text data using a speech recognition algorithm. Specifically, it uses a machine learning-based model (e.g., a model using an open-source deep learning library) to convert the voice data into text data. This text data is passed to the next analysis step and stored in temporary memory.

[0643] The server then analyzes the generated text data to identify the context of the user's speech and the relationship between the user and the other party. This context analysis uses natural language processing technology (e.g., BERT or GPT-based models). The server comprehensively determines the context by taking into account the content of the speech, time, location, existing conversation history, etc. It also estimates the relationship between the user and the other party by referencing a user profile database and past speech data.

[0644] Furthermore, the server uses an emotion engine to analyze the user's emotions from the voice data. The emotion engine analyzes the user's voice tone, pitch, speed, facial expression recognition, and bio-signal data (e.g., heart rate, skin potential) to identify the user's emotional state. A deep learning model is used for emotion analysis.

[0645] Based on the results of contextual analysis and emotion recognition, the server uses a generative AI model to translate the text data into the target language. Emotional information and writing style are also taken into consideration to generate a translation with appropriate expressions. For example, OpenAI's GPT-based model is used as the generative AI model. The generated translation data is stored in temporary memory.

[0646] The translation results are sent from the server to the device and notified to the user. The device is equipped with a notification function, allowing the user to check the translation results and provide feedback if necessary. Using the feedback function, users' ratings and opinions are entered into the device and sent to the server. The server analyzes this feedback and optimizes the parameters of the translation algorithm and generative AI model to improve future translation accuracy.

[0647] Specific use cases include the following:

[0648] Use in business situations

[0649] During a meeting, a user says, "How is the progress on this project?" The device captures the speech and sends it to the server. The server converts the speech into text and recognizes the formal context of a business setting. The emotion engine also analyzes the user's tone to determine seriousness. The result is translated into "How is the progress on this project?" in a formal and serious tone. The translation is then displayed on the device for the user to confirm.

[0650] Use in everyday conversation

[0651] When a user is with friends, they say, "Where should we have dinner tonight?" The device captures the speech and sends it to the server. The server converts the speech to text and recognizes the casual context of everyday conversation. The emotion engine detects the joy in the user's voice. Based on this, the translation is performed, resulting in "Where should we have dinner tonight?" in a casual and joyful tone. The translation result is displayed on the device for the user to confirm.

[0652] In this way, by using this system, more natural and appropriate translations are provided, enabling users to communicate smoothly. In addition, the feedback function allows for continuous system improvement, ensuring high translation accuracy over the long term.

[0653] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0654] Step 1:

[0655] The user speaks into the device. The device uses the built-in microphone to capture the user's speech in real time. The voice capture module is activated and captures the user's speech as digital audio data in WAV format. This audio data is stored in temporary memory.

[0656] Step 2:

[0657] The device compresses the captured audio data. Specifically, the audio data is converted to MP3 format and compressed. This compressed audio data is sent to a server via the Internet. The input is WAV format audio data, and the output is compressed MP3 format audio data.

[0658] Step 3:

[0659] The server analyzes the received audio data and applies a speech recognition algorithm. Specifically, the server uses a machine learning-based model (e.g., a model using a deep learning library) to convert the audio data into text data. The input is audio data in MP3 format, and the output is text data.

[0660] Step 4:

[0661] The server analyzes the generated text data and identifies the context and relationships with the other party. This analysis uses natural language processing technology (e.g., BERT or GPT-based models). The text data is input as a prompt, and the context and relationships are derived. The input is text data, and the output is context information and relationship information.

[0662] Step 5:

[0663] The server uses an emotion engine to analyze the user's emotions from the voice data. It identifies the user's emotional state by analyzing the tone, pitch, and rate of the user's voice, and also incorporating facial expression recognition and biosignal data (e.g., heart rate and skin potential). The input is text data and biosignal data, and the output is emotional state data.

[0664] Step 6:

[0665] The server uses a generative AI model to translate the text data into the target language based on the results of contextual analysis and emotion recognition. It generates an appropriate translation taking into account emotional information and writing style. The input is text data, context information, relationship information, and emotional state data, and the output is the translated text data.

[0666] Step 7:

[0667] The translation result is sent from the server to the terminal and notified to the user. The terminal uses a notification function to display the translation result to the user. The input is the translated text data, and the output is the translation result displayed to the user.

[0668] Step 8:

[0669] The user checks the translation results on the terminal and provides feedback if necessary. The terminal displays a feedback input form and sends the feedback entered by the user to the server. The input is the feedback data from the user, and the output is the feedback data sent to the server.

[0670] Step 9:

[0671] The server analyzes user feedback and optimizes the parameters of the translation algorithm and generative AI model, thereby improving future translation accuracy. The input is the feedback data, and the output is the optimized algorithm parameters.

[0672] (Application example 2)

[0673] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0674] In recent years, the use of self-driving vehicles has become more widespread, making appropriate communication with passengers crucial. However, current self-driving vehicle systems face challenges, such as language barriers and the inability to properly recognize and respond to emotions, which can reduce user satisfaction. Therefore, there is a need for a system that can analyze context and emotions in real time, translate them into the appropriate language, and even control in-vehicle functions.

[0675] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring a user's utterance as voice data, means for converting the voice data into text data, means for analyzing the context, the relationship with the other party, and the user's emotions, means for translating the text data into a target language based on the analysis results, means for notifying the user of the translation results, and means for controlling functions in the autonomous vehicle based on the user's utterance. This enables passengers to communicate naturally and appropriately in the autonomous vehicle and efficiently operate the vehicle's functions.

[0676] "User utterance" refers to words or instructions expressed by the user through speech.

[0677] "Voice data" refers to data that is a digital recording of a user's speech.

[0678] "Text data" is character information obtained by analyzing voice data.

[0679] "Context" refers to the linguistic background that is understood by taking into account the content of the user's speech, the surrounding circumstances, past speech history, and so on.

[0680] "Relationship with the other party" refers to the relationship between the user and the target person or system.

[0681] "Emotion" is a psychological state inferred based on the user's voice and other biometric signals.

[0682] "Target language" is the language to which the text is to be translated.

[0683] The "translation result" is the text resulting from converting the original text data into the target language.

[0684] An "autonomous vehicle" is a vehicle that is driven automatically without human intervention.

[0685] "Means for controlling functions" means methods or devices for operating or adjusting various devices or settings within an automated vehicle.

[0686] The system that realizes this application example includes the following components: Acquires user speech as voice data, converts the voice data into text data, and analyzes the context, the relationship with the other person, and the user's emotions. Based on the analysis results, translates the text data into the target language, notifies the user of the translation result, and controls functions within the autonomous vehicle.

[0687] Hardware and software used

[0688] 1. User utterance acquisition

[0689] Hardware: Self-driving vehicle system with microphone

[0690] Software: sounddevice library

[0691] The microphone inside the autonomous vehicle captures the user's speech in real time, temporarily stores it in the built-in memory as voice data, and then transmits it to the server. The user can interact with the system by speaking into the microphone.

[0692] 2. Audio data conversion

[0693] Hardware: Server or in-vehicle computer

[0694] Software: speech_recognition library

[0695] The server analyzes the received voice data and uses a deep learning-based model to convert the voice data into highly accurate text data, which is then passed on to the next analysis step.

[0696] 3. Context, Relational, and Sentiment Analysis

[0697] Hardware: Server or in-vehicle computer

[0698] Software: Natural language processing engine, emotion recognition module

[0699] The server analyzes the generated text data and analyzes the context of the user's speech, the relationship with the other person, and their emotions. It does this by using natural language processing technology and referring to existing conversation history and user profiles. It also uses an emotion analysis engine to identify the user's emotional state from the tone, pitch, and speed of their voice.

[0700] 4. Generating the Translation

[0701] Hardware: Server or in-vehicle computer

[0702] Software: translation engines (e.g., googletrans library), generative AI models

[0703] The server uses a generative AI model to translate the text data into the target language based on the results of contextual analysis and emotion recognition. An appropriate translation is generated that takes into account emotional information and writing style. The translated text data is stored in temporary memory.

[0704] 5. Notification of translation results

[0705] Hardware: Smartphones, tablets, dedicated devices

[0706] Software: Translation result display module

[0707] The translation results are sent from the server to the terminal and notified to the user, who can then check the translation results on the terminal screen.

[0708] 6. Control of functions inside autonomous vehicles

[0709] Hardware: Autonomous vehicle control systems

[0710] Software: Vehicle Control Module

[0711] The system controls in-vehicle functions (e.g., temperature adjustment, seat arrangement, destination change, etc.) based on the user's speech, allowing the user to efficiently operate in-vehicle functions using only speech.

[0712] Specific examples

[0713] Example 1:

[0714] When a passenger says "Adjust the air conditioning" in an autonomous vehicle, the system captures the speech and converts it into text. The server analyzes the context and sentiment of the text and translates it as "Adjust the air conditioning." The translation result is displayed on the user's device, and the air conditioning system in the vehicle is simultaneously adjusted.

[0715] Example prompt sentence:

[0716] Voice translation: Adjust the air conditioning

[0717] This will enable users to communicate naturally and smoothly within an autonomous vehicle and efficiently operate the vehicle's functions.

[0718] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0719] Step 1:

[0720] The user speaks. The user issues instructions or questions by voice into the microphone inside the autonomous vehicle. The input is voice data.

[0721] Step 2:

[0722] The device captures the user's speech and stores it in temporary memory as audio data. The input is audio data, and the output is a digital audio file. Specifically, it uses the sounddevice library to record audio and saves the data to a file.

[0723] Step 3:

[0724] The device compresses the stored audio data and transmits it to a server via the Internet. The input is a digital audio file, and the output is compressed audio data. Specifically, the audio data is compressed using a compression algorithm and then transmitted to the server.

[0725] Step 4:

[0726] The server analyzes the received voice data and converts it into text data by applying a speech recognition algorithm. The input is compressed voice data and the output is text data. Specifically, the server uses the speech_recognition library to convert voice to text data.

[0727] Step 5:

[0728] The server analyzes the generated text data to identify the context and the relationship with the other party. The input is the text data, and the output is the analyzed context and relationship information. The server uses a natural language processing engine to analyze the text and reference past conversation history and user profiles.

[0729] Step 6:

[0730] The server uses an emotion recognition engine to analyze the user's emotions. The input is voice data and text data, and the output is analyzed emotional information. Specifically, the server analyzes the tone, pitch, and speed of the voice to identify the user's emotional state.

[0731] Step 7:

[0732] The server translates the text data into the target language using a generative AI model based on the results of contextual analysis and emotion recognition. The input is the text data and the analysis results, and the output is the translated text data. The text is translated into the target language using the googletrans library.

[0733] Step 8:

[0734] The server sends the translation results to the terminal and notifies the user. The input is the translated text data, and the output is the translation results displayed to the user. The terminal displays the translation results on the screen so that the user can check them.

[0735] Step 9:

[0736] The device controls functions within the autonomous vehicle based on the user's speech. The input is translated text data and user instructions, and the output is changes to the vehicle's functions. Specifically, it sends signals to control the vehicle's air conditioning system, navigation system, etc.

[0737] Through the above processing steps, users can achieve natural and smooth communication within an autonomous vehicle and efficiently operate the vehicle's functions.

[0738] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0739] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0740] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0741] [Third embodiment]

[0742] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0743] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0744] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0745] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0746] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0747] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0748] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0749] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0750] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0751] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0752] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0753] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0754] This invention is a context-adaptive real-time translation system that uses speech recognition, natural language processing, and machine translation technologies. Specifically, the system acquires user utterances as voice data, converts the voice data into text data, analyzes the context and the relationship with the other party, and generates an appropriate translation based on the analysis results.

[0755] Program processing procedures and explanations

[0756] 1. User utterance transmission

[0757] The user speaks into a device (smartphone, tablet, PC, etc.), which captures the voice in real time and records it as audio data. The recorded audio data is compressed and sent to a server via the Internet.

[0758] 2. Audio data conversion

[0759] The server receives the voice data sent from the device and applies a speech recognition algorithm, which uses a deep learning-based model to convert the user's speech into text data with high accuracy. This process converts the voice data into text data and stores it in a temporary file or memory.

[0760] 3. Context and Relationship Analysis

[0761] The server analyzes the generated text data and estimates the context in which the user spoke and the relationship with the other person. Context analysis uses natural language processing technology to identify the scene by taking into account the content of the user's speech, time, location, and existing conversation history. The server also references a user profile database and past speech data to estimate the relationship between the user and the target person.

[0762] 4. Generating the Translation

[0763] The server uses a generative AI model to generate an appropriate translation based on information obtained from context and relationships. For example, it generates formal expressions for business situations and casual expressions for everyday conversations. The translation results are stored in temporary memory on the server.

[0764] 5. Gather feedback and optimize

[0765] The translation results are sent to the device and notified to the user. The user reviews the translation results and provides feedback if necessary. The device then sends this feedback to the server. The server analyzes the feedback and optimizes the parameters of the translation algorithm and generative AI model. This process improves future translation accuracy.

[0766] Specific examples

[0767] Example 1: Use in business situations

[0768] During a meeting, a user says, "How is the progress on this project?" The device captures the speech and sends it to the server. The server converts the speech into text and translates it, recognizing the formal context of the business setting. The resulting translation is "How is the progress on this project?" The translation is displayed on the device for the user to review.

[0769] Example 2: Use in everyday conversation

[0770] A user is with friends and says, "Where should we have dinner tonight?" The device captures the speech and sends it to the server. The server converts the speech to text and translates it, recognizing the casual context of everyday conversation. The resulting translation is "Where should we have dinner tonight?" The translation is displayed on the device for the user to confirm.

[0771] This invention combines speech recognition, natural language processing, and machine translation technologies to provide real-time translation suitable for a variety of situations. Furthermore, by optimizing the translation algorithm based on user feedback, it is possible to achieve more accurate translation.

[0772] The processing flow will be explained below.

[0773] Step 1:

[0774] The user speaks into the device, which uses a built-in microphone to capture the user's speech in real time, and this voice data is stored digitally in the device's temporary memory.

[0775] Step 2:

[0776] The device compresses the captured audio data and sends it over the internet to a server, where it is encrypted for security purposes.

[0777] Step 3:

[0778] The server receives the voice data sent from the terminal, decodes the received voice data, and prepares it for voice recognition processing.

[0779] Step 4:

[0780] The server applies a speech recognition algorithm to convert the audio data into text data, using a deep learning-based model to convert spoken content into text with high accuracy.

[0781] Step 5:

[0782] The server analyzes the generated text data and identifies contextual information and relationships with the other party. Contextual analysis uses natural language processing technology. This analysis categorizes the scene into categories such as business or everyday conversation.

[0783] Step 6:

[0784] The server references the user profile database and past conversation history to estimate the relationship between the user and the other party, and based on this, selects an appropriate communication style (formal, casual, etc.).

[0785] Step 7:

[0786] The server retrieves translation settings based on the analysis results and uses a generative AI model to translate the text data into the target language, adjusting the writing style and tone accordingly.

[0787] Step 8:

[0788] The server then sends the translation results to the device, which then displays them on the user interface. If necessary, speech synthesis technology can also be used to output the results as voice.

[0789] Step 9:

[0790] The user checks the translation results and provides feedback if necessary, which is sent from the device to the server.

[0791] Step 10:

[0792] The server analyzes the received feedback and updates and optimizes the parameters of the translation algorithm and generative AI model, thereby improving future translation accuracy.

[0793] In this way, this system translates user utterances in real time and provides appropriate translation results for a variety of situations. By optimizing the algorithm based on user feedback, it achieves high-accuracy translation over the long term.

[0794] Example 1

[0795] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0796] Conventional translation systems often perform simple translations without considering the context or the relationship with the target user, which can result in low translation accuracy. In addition, they lack the ability to optimize the system based on user feedback, making it difficult to continuously improve translation accuracy.

[0797] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0798] In this invention, the server includes means for acquiring user utterances as voice data, means for converting the voice data into text data, means for analyzing the relationship between the context and the target person, means for translating the text data into the target language based on the analysis results, means for notifying the user of the translation results, and means for collecting feedback from the user and optimizing the translation algorithm. This enables highly accurate translation that adapts to the context and relationship, and continuous system optimization is achieved by utilizing the feedback.

[0799] The "means for acquiring user speech as voice data" is a function for electronically capturing the voice uttered by the user and recording it as digital voice data.

[0800] The "means for converting voice data into text data" is a function for analyzing captured voice data and converting it into text data expressed as a corresponding character string.

[0801] The "means for analyzing the relationship between the context and the target person" is an analytical function for estimating the relationship between the context of an utterance and the target person of that utterance from the generated text data.

[0802] The "means for translating text data into a target language" refers to a translation function that converts the original text data into a different language based on the analyzed context and relationships.

[0803] The "means for notifying the user of the translation result" is a function for visually or audibly notifying the user of the generated translation result.

[0804] "Means for collecting user feedback and optimizing the translation algorithm" refers to a function that collects user evaluations and suggestions for corrections on translation results, and then adjusts the parameters and settings of the translation algorithm based on that information to improve accuracy.

[0805] A "deep learning-based model" is a type of algorithm that uses a multi-layer neural network to extract and transform data features, enabling advanced recognition and prediction.

[0806] A "generative AI model" is an algorithm that uses artificial intelligence technology to generate text or translate, and has the ability to generate appropriate output based on a prompt.

[0807] The present invention is a context-adaptive real-time translation system that uses speech recognition, natural language processing, and machine translation technologies. Specific embodiments for carrying out the present invention are described below.

[0808] First, the user speaks into the device (smartphone, tablet, PC, etc.). The device uses a built-in microphone to capture the user's voice in real time and record it as audio data. The recorded audio data is saved in a lossless compressed format such as FLAC and sent to a server via the Internet.

[0809] The server receives the voice data sent via the Internet, applies a speech recognition algorithm (e.g., Google Speech-to-Text API using a deep learning-based model) to analyze the voice data and convert it into text data, which is then saved in a temporary file or memory.

[0810] The server then analyzes the generated text data and estimates the context of the utterance and its relationship to the target person. This analysis uses natural language processing technology (e.g., the spaCy library). The server references a user profile database and a database of past utterances to identify the scene, taking into account the user's utterance content, time, location, and existing conversation history. It also references these databases to identify the relationship between the user and the target person.

[0811] The server generates an appropriate prompt based on the context and relationships, and generates the translation using a generative AI model (e.g., OpenAI GPT-3). The AI ​​model is fed with a prompt like this:

[0812] Input: "How is this project progressing?"

[0813] Context: Business meeting

[0814] Relationship: Superior-Subordinate

[0815] Produces the translation: "How is the progress on this project?"

[0816] The generated translation result is stored in temporary memory in the server and then sent to the terminal, where the user can check the translation result on the terminal screen or through voice output.

[0817] Finally, the user can review the translation results and provide feedback, such as ratings and suggestions for corrections. The device collects this feedback and sends it back to the server. The server analyzes the feedback and optimizes the parameters of the translation algorithm and generative AI model, thereby improving future translation accuracy.

[0818] As a concrete example, consider the case where a user says, "How is the progress on this project?" during a meeting. At this time, the device captures the audio and sends it to the server in FLAC format. The server converts it into text using the Google Speech-to-Text API and analyzes the context of the business meeting and the relationship between superiors and subordinates. Based on the analysis results, a prompt is sent to the GPT-3 model, which generates the translation result, "How is the progress on this project?". This translation result is sent to the device and confirmed by the user.

[0819] As another example, consider the case where a user is with friends and says, "Where should we have dinner tonight?" The device captures the audio and sends it to the server in FLAC format. The server converts it to text using the Google Speech-to-Text API and analyzes the casual context of everyday conversation and friendships. Based on the analysis results, the prompt is sent to the GPT-3 model, which generates the translation result, "Where should we have dinner tonight?" The translation result is then sent to the device for the user to confirm.

[0820] In this way, by combining speech recognition, natural language processing, and generative AI models, the present invention can provide highly accurate real-time translation that adapts to the user's context and relationships. Furthermore, by optimizing the translation algorithm based on user feedback, the system can be continuously improved.

[0821] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0822] Step 1: User utterance submission

[0823] explanation

[0824] The user speaks into a device (smartphone, tablet, PC, etc.), which then uses a built-in microphone to capture the voice in real time and record it as digital audio data.

[0825] input

[0826] User utterance: "How is this project going?"

[0827] output

[0828] FLAC format audio data

[0829] Specific actions

[0830] The user speaks to their smartphone, "How is this project progressing?" The device uses its built-in microphone to capture the audio and saves it as a FLAC audio file. The saved audio data is then sent to a server over the Internet.

[0831] Step 2: Convert the audio data

[0832] explanation

[0833] The server receives the voice data sent from the terminal and converts it into text data by applying a voice recognition algorithm.

[0834] input

[0835] FLAC format audio data

[0836] output

[0837] Text data: "How is this project progressing?"

[0838] Specific actions

[0839] The server receives FLAC audio data from the device. The server calls the Google Speech-to-Text API, analyzes the audio data, and converts it into text data such as "How is the progress on this project?" The converted text data is saved as a temporary file.

[0840] Step 3: Analyze context and relationships

[0841] explanation

[0842] The server analyzes the generated text data and estimates the context of the utterance and its relationship to the target person.

[0843] input

[0844] Text data: "How is this project progressing?"

[0845] output

[0846] Analysis results (context: business meeting, relationship: superior-subordinate)

[0847] Specific actions

[0848] The server analyzes the text data using the spaCy library, infers that the context of the utterance is a business meeting, and references a database of past utterances and a user profile database to determine that the user is a superior and the target person is a subordinate.

[0849] Step 4: Generate translations

[0850] explanation

[0851] The server generates prompt sentences based on context and relationships, and uses generative AI models to generate appropriate translations.

[0852] input

[0853] Analysis results (context: business meeting, relationship: superior-subordinate)

[0854] output

[0855] Translation result: "How is the progress on this project?"

[0856] Specific actions

[0857] The server inputs the following prompt into the generative AI model (OpenAI GPT-3):

[0858] Input: "How is this project progressing?"

[0859] Context: Business meeting

[0860] Relationship: Superior-Subordinate

[0861] Based on this prompt, the GPT-3 model generates the translation "How is the progress on this project?" This translation result is stored in temporary memory on the server.

[0862] Step 5: Notification of translation results

[0863] explanation

[0864] The server sends the generated translation result to the terminal and notifies the user.

[0865] input

[0866] Translation result: "How is the progress on this project?"

[0867] output

[0868] Terminal display: "How is the progress on this project?"

[0869] Specific actions

[0870] The server sends the generated translation to the device as an HTTPS response, and the user can check the translation result, "How is the progress on this project?", displayed on the device screen.

[0871] Step 6: Gather feedback and optimize

[0872] explanation

[0873] The user provides feedback on the translation result, and the terminal sends the feedback to the server, which analyzes the feedback and optimizes the translation algorithm.

[0874] input

[0875] User feedback: "The translation was accurate"

[0876] output

[0877] Optimized translation algorithm

[0878] Specific actions

[0879] The user reviews the translation result and provides feedback, rating it "the translation was accurate." The device then sends this feedback to the server, which analyzes the feedback and optimizes the translation algorithm by adjusting the model settings, resulting in higher accuracy in the next translation.

[0880] (Application example 1)

[0881] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0882] Food delivery services face the challenge of having delivery personnel communicate smoothly with foreign customers. In particular, the language barrier makes it difficult for delivery personnel to accurately understand order details and delivery instructions. This problem reduces delivery efficiency and reduces customer satisfaction.

[0883] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0884] In this invention, the server includes means for translating information and questions received by the delivery person from the customer in real time, means for acquiring the user's speech as voice data, means for converting the voice data into text data, means for translating the text data into the target language based on the analysis results, and means for notifying the user of the translation results. This enables the delivery person to communicate smoothly with foreign customers, improving delivery efficiency and customer satisfaction.

[0885] The "means for acquiring user speech as voice data" refers to a device or function for recording the user's speech as digital voice data.

[0886] The "means for converting voice data into text data" refers to an algorithm or software for analyzing acquired voice data and converting it into corresponding written text.

[0887] "Means for analyzing the context and the relationship with the other person" refers to natural language processing technologies and algorithms for understanding and analyzing the content of an utterance, its background, and the relationship between the speaker and the other person.

[0888] The "means for translating text data into a target language based on the analysis results" refers to a machine translation technology for generating text in a target language that is appropriately translated based on context and relationships.

[0889] "Means for notifying the user of the translation result" refers to an interface or function for displaying or audibly conveying the generated translation result to the user.

[0890] "A means for delivery personnel to translate information and questions received from customers in real time" refers to a system or algorithm that instantly translates the content of communication between delivery personnel and customers into other languages.

[0891] System Program

[0892] Hardware and Software Configuration

[0893] The system for implementing this invention includes a terminal operated by a user (e.g., a smartphone), a server connected via the Internet, and algorithms for performing speech recognition, natural language processing, and machine translation. The terminal has the function of capturing the user's voice and sending the voice data to the server. The server receives the voice data and performs the following main processes:

[0894] 1. Speech Recognition Using Deep Learning-Based Models

[0895] Use deep learning-based speech recognition algorithms (e.g., Google Speech-to-Text API) to convert voice data into text data with high accuracy.

[0896] 2. Natural language processing for context and relationship analysis

[0897] The converted text data is analyzed using natural language processing (e.g., the BERT model) to identify the context in which the utterance was made and the relationship between the user and the other party.

[0898] 3. Generating appropriate translations using machine translation

[0899] Uses a generative AI model (e.g. GPT-3) to generate an appropriate target language translation based on context and relationships.

[0900] 4. Feedback-driven optimization

[0901] Collect user feedback and optimize the parameters of the translation algorithm.

[0902] Specific example of operation procedure

[0903] Example 1: A delivery person communicating with a foreign customer

[0904] When the delivery person says into their smartphone, "Is this the correct address?", the smartphone captures the voice and sends the data to the server in real time.

[0905] The server receives the voice data and uses a deep learning model to convert it into text data such as "Is this address correct?"

[0906] The server uses natural language processing technology to identify that the spoken content is a question in a delivery situation and appropriately analyzes the communication with the customer.

[0907] Based on the analysis results, a generative AI model is used to generate the English translation "Is this the correct address?" and display it on the smartphone.

[0908] The delivery person checks the translation results, shows them to the customer, and provides feedback as needed to improve the system's accuracy.

[0909] Example prompt sentence:

[0910] "Please suggest a translation to ask the customer for confirmation during delivery."

[0911] "How do we understand our customers' needs and respond to them formally?"

[0912] "Please suggest a translation for a scenario in which a delivery person asks a customer where their pickup is located in a casual context."

[0913] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0914] Step 1:

[0915] User utterance transmission

[0916] The user (delivery person) speaks into a smartphone. The device (smartphone) captures the voice in real time and generates voice data. The voice data is compressed and sent to a server via the Internet. In this case, the input is the delivery person's speech, and the output is compressed voice data.

[0917] Step 2:

[0918] Audio data conversion

[0919] The server receives the voice data sent from the device. Then, it applies a deep learning-based speech recognition algorithm (e.g., Google Speech-to-Text API) to generate text data from the voice data. The input is compressed voice data, and the output is text data.

[0920] Step 3:

[0921] Context and relationship analysis

[0922] The server analyzes the generated text data. It uses a natural language processing model (e.g., BERT) to identify the context in which the user spoke and their relationship to the other party (customer). This process also references existing conversation history and user profile databases. The input is text data, and the output is information related to the context and relationships.

[0923] Step 4:

[0924] Generating Translations

[0925] The server uses a generative AI model (e.g., GPT-3) to generate an appropriate translation based on the information obtained from the context and relationships. For example, for business questions, it generates a translation using formal language. The input is information related to the context and relationships and text data, and the output is the text data translated into the target language.

[0926] Step 5:

[0927] Notification of translation results

[0928] The translation results are sent from the server to the device. The device displays the translation results to the user and notifies them by voice or text. The input is the translated text data, and the output is the translation results displayed on the device.

[0929] Step 6:

[0930] Feedback collection and optimization

[0931] The user checks the translation results and provides feedback if necessary. The device sends this feedback to the server, which analyzes the feedback and optimizes the parameters of the translation algorithm and generative AI model. The input is the user's feedback data, and the output is the optimized translation algorithm.

[0932] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0933] This invention is a context-adaptive real-time translation system that combines speech recognition, natural language processing, machine translation, and an emotion engine. The system acquires user utterances as voice data and converts the voice data into text data. In addition, the system analyzes the context, the relationship with the other person, and the user's emotions, and generates and provides an appropriate translation based on the analysis results.

[0934] Program processing procedures and explanations

[0935] 1. User utterance transmission

[0936] The user speaks into a device (smartphone, tablet, PC, etc.), which uses a built-in microphone to capture the user's speech in real time. This voice data is stored in a temporary memory in digital format, compressed, and then transmitted to a server via the Internet.

[0937] 2. Audio data conversion

[0938] The server analyzes the received voice data and applies a speech recognition algorithm, using a deep learning-based model to convert the voice data to text data with high accuracy, which is passed on to the next analysis step and stored in temporary memory.

[0939] 3. Context and Relationship Analysis

[0940] The server analyzes the generated text data and identifies the context of the user's speech and the relationship with the other party. Context analysis is performed using natural language processing technology, and a comprehensive assessment is made of the user's speech content, time, location, existing conversation history, etc. It also references a user profile database and past speech data to estimate the relationship between the user and the other party and select an appropriate communication style.

[0941] 4. Emotion Recognition Processing

[0942] The server uses an emotion engine to analyze the user's emotions from the voice data. The emotion engine analyzes the tone, pitch, and speed of the user's voice to identify the user's emotional state (e.g., joy, sadness, anger, etc.). In addition, the emotion engine also incorporates the user's facial expression recognition and biometric signal data (heart rate, skin potential, etc.) to perform more accurate emotion analysis.

[0943] 5. Generating the Translation

[0944] The server uses a generative AI model to translate the text data into the target language based on the results of contextual analysis and emotion recognition. Emotional information and writing style are also taken into consideration to generate a translation with appropriate expressions. The translated text data is stored in temporary memory.

[0945] 6. Gather feedback and optimize

[0946] The translation results are sent from the server to the device and notified to the user. The user can check the translation results on the device and provide feedback if necessary. The device then sends this feedback to the server. The server analyzes the feedback and optimizes the parameters of the translation algorithm and generative AI model to improve future translation accuracy.

[0947] Specific examples

[0948] Example 1: Use in business situations

[0949] During a meeting, a user says, "How is the progress on this project?" The device captures the speech and sends it to the server. The server converts the speech into text and recognizes the formal context of a business setting. The emotion engine also analyzes the user's tone to determine their seriousness. The result is translated into "How is the progress on this project?" in a formal and serious tone. The translation is then displayed on the device for the user to confirm.

[0950] Example 2: Use in everyday conversation

[0951] When a user is with friends, they say, "Where should we have dinner tonight?" The device captures the speech and sends it to the server. The server converts the speech to text and recognizes the casual context of everyday conversation. The emotion engine detects the joy in the user's voice. Based on this, the translation is performed, resulting in "Where should we have dinner tonight?" in a casual and joyful tone. The translation result is displayed on the device for the user to confirm.

[0952] The system of the present invention combines speech recognition, natural language processing, machine translation, and emotion recognition technologies to provide more natural and appropriate communication. Furthermore, it optimizes the algorithm based on user feedback to achieve high-accuracy translation over the long term.

[0953] The processing flow will be explained below.

[0954] Step 1:

[0955] The user speaks into the device, which uses a built-in microphone to capture the user's speech in real time, and this voice data is stored digitally in the device's temporary memory.

[0956] Step 2:

[0957] The device compresses the captured audio data and sends it to a server over the Internet, where it is encrypted for secure transmission.

[0958] Step 3:

[0959] The server receives the voice data sent from the terminal, decodes the received voice data, and prepares it for voice recognition processing.

[0960] Step 4:

[0961] The server applies a speech recognition algorithm to convert the audio data into text data. It uses a deep learning-based model to convert the speech into text with high accuracy. This text data is stored in the server's temporary memory.

[0962] Step 5:

[0963] The server then analyzes the generated text data using natural language processing technology to identify the context in which the utterance was made and the relationship with the other person. Context analysis takes into account the content of the utterance, time, location, and existing conversation history.

[0964] Step 6:

[0965] The server compares the user profile database and past conversation history to estimate the relationship between the user and the other party, and then determines the appropriate communication style, such as formal or casual.

[0966] Step 7:

[0967] The server uses an emotion engine to analyze the user's emotions from the voice data. The emotion engine analyzes the tone, pitch, and speed of the voice to identify the user's emotional state (e.g., joy, sadness, anger, etc.). It also uses facial expression recognition and biometric signal data as needed to improve the accuracy of the emotion detection.

[0968] Step 8:

[0969] Based on the results of contextual analysis and emotion recognition, the server uses a generative AI model to translate the text data into the target language, taking into account appropriate writing style and emotional expression. The generated translation data is stored in temporary memory.

[0970] Step 9:

[0971] The server sends the generated translation results to the terminal, which then displays the received translation results on the user interface. If necessary, speech synthesis technology can also be used to output the results as voice.

[0972] Step 10:

[0973] The user checks the translation results and provides feedback if necessary, which the device then sends to the server.

[0974] Step 11:

[0975] The server analyzes the received feedback and updates and optimizes the parameters of the translation algorithm and generative AI model, thereby improving future translation accuracy.

[0976] This allows the system of the present invention to translate user utterances in real time and provide appropriate translation results for various situations. Furthermore, by combining it with an emotion recognition engine, it achieves more natural translation that reflects the user's emotions. Optimization through feedback allows the system to be continuously improved, enabling highly accurate translation.

[0977] Example 2

[0978] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0979] Conventional translation systems struggle to accurately convert speech to text and to translate appropriately while taking into account the context and emotion of the translated text. They also lack effective means for collecting user feedback and improving translation accuracy. In particular, they are unable to accurately capture emotional nuances, making it difficult to achieve natural and appropriate communication.

[0980] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0981] In this invention, the server includes means for acquiring user speech as voice data, means for converting the voice data into text data, means for analyzing the context and the relationship with the other party, means for performing sentiment analysis and incorporating the results into the translation, and means for collecting user feedback and optimizing the translation algorithm. This not only improves the accuracy of voice data conversion, but also enables natural and appropriate translation that takes into account the relationship between the user and the other party, the context, and the emotions. Furthermore, translation accuracy can be continuously improved based on user feedback.

[0982] "User" refers to an individual who uses the system to make a speech.

[0983] A "terminal" is a device that a user uses to make a speech, and includes a smartphone, tablet, PC, etc.

[0984] "Voice Data" refers to information that digitally captures a user's speech.

[0985] "Text data" refers to character string information generated by analyzing voice data.

[0986] "Context" refers to the situation or background information obtained by comprehensively assessing the content of the user's utterance, time, location, existing conversation history, etc.

[0987] "Relationship" refers to the relationship between a user and a conversation partner, and is estimated by referring to past conversation history and a user profile database.

[0988] "Emotion analysis" refers to the process of analyzing a user's voice tone, pitch, rate, facial expression recognition, and biometric signal data to identify the user's emotional state.

[0989] "Translation" refers to the process of converting text data into a target language based on the analysis results.

[0990] "Feedback" refers to the evaluation and opinions of the translation results provided by the user.

[0991] A "generative AI model" refers to an algorithm that uses artificial intelligence technology to generate appropriate translations.

[0992] "Optimization" refers to the process of adjusting the system's parameters based on collected feedback to improve translation accuracy.

[0993] A "prompt sentence" is a portion of the text data input into a generative AI model, and refers to the phrase that serves as the basis for the model to generate a translation.

[0994] "Machine learning-based models" refer to trained algorithms or neural networks used to analyze audio data or generate text data.

[0995] "Natural language processing" refers to technology for analyzing the context and meaning of text data.

[0996] This invention is a context-adaptive real-time translation system that combines speech recognition, natural language processing, machine translation, and an emotion engine. Specifically, it uses the following hardware and software:

[0997] A user speaks into a device (such as a smartphone, tablet, or PC) and their speech is captured as audio data in real time. The device uses a built-in microphone and audio capture module to temporarily store this audio data in digital format (WAV format), compress it (e.g., to MP3 format), and transmit it to a server via the Internet.

[0998] The server analyzes the received voice data and converts it into text data using a speech recognition algorithm. Specifically, it uses a machine learning-based model (e.g., a model using an open-source deep learning library) to convert the voice data into text data. This text data is passed to the next analysis step and stored in temporary memory.

[0999] The server then analyzes the generated text data to identify the context of the user's speech and the relationship between the user and the other party. This context analysis uses natural language processing technology (e.g., BERT or GPT-based models). The server comprehensively determines the context by taking into account the content of the speech, time, location, existing conversation history, etc. It also estimates the relationship between the user and the other party by referencing a user profile database and past speech data.

[1000] Furthermore, the server uses an emotion engine to analyze the user's emotions from the voice data. The emotion engine analyzes the user's voice tone, pitch, speed, facial expression recognition, and bio-signal data (e.g., heart rate, skin potential) to identify the user's emotional state. A deep learning model is used for emotion analysis.

[1001] Based on the results of contextual analysis and emotion recognition, the server uses a generative AI model to translate the text data into the target language. Emotional information and writing style are also taken into consideration to generate a translation with appropriate expressions. For example, OpenAI's GPT-based model is used as the generative AI model. The generated translation data is stored in temporary memory.

[1002] The translation results are sent from the server to the device and notified to the user. The device is equipped with a notification function, allowing the user to check the translation results and provide feedback if necessary. Using the feedback function, users' ratings and opinions are entered into the device and sent to the server. The server analyzes this feedback and optimizes the parameters of the translation algorithm and generative AI model to improve future translation accuracy.

[1003] Specific use cases include the following:

[1004] Use in business situations

[1005] During a meeting, a user says, "How is the progress on this project?" The device captures the speech and sends it to the server. The server converts the speech into text and recognizes the formal context of a business setting. The emotion engine also analyzes the user's tone to determine seriousness. The result is translated into "How is the progress on this project?" in a formal and serious tone. The translation is then displayed on the device for the user to confirm.

[1006] Use in everyday conversation

[1007] When a user is with friends, they say, "Where should we have dinner tonight?" The device captures the speech and sends it to the server. The server converts the speech to text and recognizes the casual context of everyday conversation. The emotion engine detects the joy in the user's voice. Based on this, the translation is performed, resulting in "Where should we have dinner tonight?" in a casual and joyful tone. The translation result is displayed on the device for the user to confirm.

[1008] In this way, by using this system, more natural and appropriate translations are provided, enabling users to communicate smoothly. In addition, the feedback function allows for continuous system improvement, ensuring high translation accuracy over the long term.

[1009] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1010] Step 1:

[1011] The user speaks into the device. The device uses the built-in microphone to capture the user's speech in real time. The voice capture module is activated and captures the user's speech as digital audio data in WAV format. This audio data is stored in temporary memory.

[1012] Step 2:

[1013] The device compresses the captured audio data. Specifically, the audio data is converted to MP3 format and compressed. This compressed audio data is sent to a server via the Internet. The input is WAV format audio data, and the output is compressed MP3 format audio data.

[1014] Step 3:

[1015] The server analyzes the received audio data and applies a speech recognition algorithm. Specifically, the server uses a machine learning-based model (e.g., a model using a deep learning library) to convert the audio data into text data. The input is audio data in MP3 format, and the output is text data.

[1016] Step 4:

[1017] The server analyzes the generated text data and identifies the context and relationships with the other party. This analysis uses natural language processing technology (e.g., BERT or GPT-based models). The text data is input as a prompt, and the context and relationships are derived. The input is text data, and the output is context information and relationship information.

[1018] Step 5:

[1019] The server uses an emotion engine to analyze the user's emotions from the voice data. It identifies the user's emotional state by analyzing the tone, pitch, and rate of the user's voice, and also incorporating facial expression recognition and biosignal data (e.g., heart rate and skin potential). The input is text data and biosignal data, and the output is emotional state data.

[1020] Step 6:

[1021] The server uses a generative AI model to translate the text data into the target language based on the results of contextual analysis and emotion recognition. It generates an appropriate translation taking into account emotional information and writing style. The input is text data, context information, relationship information, and emotional state data, and the output is the translated text data.

[1022] Step 7:

[1023] The translation result is sent from the server to the terminal and notified to the user. The terminal uses a notification function to display the translation result to the user. The input is the translated text data, and the output is the translation result displayed to the user.

[1024] Step 8:

[1025] The user checks the translation results on the terminal and provides feedback if necessary. The terminal displays a feedback input form and sends the feedback entered by the user to the server. The input is the feedback data from the user, and the output is the feedback data sent to the server.

[1026] Step 9:

[1027] The server analyzes user feedback and optimizes the parameters of the translation algorithm and generative AI model, thereby improving future translation accuracy. The input is the feedback data, and the output is the optimized algorithm parameters.

[1028] (Application example 2)

[1029] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1030] In recent years, the use of self-driving vehicles has become more widespread, making appropriate communication with passengers crucial. However, current self-driving vehicle systems face challenges, such as language barriers and the inability to properly recognize and respond to emotions, which can reduce user satisfaction. Therefore, there is a need for a system that can analyze context and emotions in real time, translate them into the appropriate language, and even control in-vehicle functions.

[1031] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring a user's utterance as voice data, means for converting the voice data into text data, means for analyzing the context, the relationship with the other party, and the user's emotions, means for translating the text data into a target language based on the analysis results, means for notifying the user of the translation results, and means for controlling functions in the autonomous vehicle based on the user's utterance. This enables passengers to communicate naturally and appropriately in the autonomous vehicle and efficiently operate the vehicle's functions.

[1032] "User utterance" refers to words or instructions expressed by the user through speech.

[1033] "Voice data" refers to data that is a digital recording of a user's speech.

[1034] "Text data" is character information obtained by analyzing voice data.

[1035] "Context" refers to the linguistic background that is understood by taking into account the content of the user's speech, the surrounding circumstances, past speech history, and so on.

[1036] "Relationship with the other party" refers to the relationship between the user and the target person or system.

[1037] "Emotion" is a psychological state inferred based on the user's voice and other biometric signals.

[1038] "Target language" is the language to which the text is to be translated.

[1039] The "translation result" is the text resulting from converting the original text data into the target language.

[1040] An "autonomous vehicle" is a vehicle that is driven automatically without human intervention.

[1041] "Means for controlling functions" means methods or devices for operating or adjusting various devices or settings within an automated vehicle.

[1042] The system that realizes this application example includes the following components: Acquires user speech as voice data, converts the voice data into text data, and analyzes the context, the relationship with the other person, and the user's emotions. Based on the analysis results, translates the text data into the target language, notifies the user of the translation result, and controls functions within the autonomous vehicle.

[1043] Hardware and software used

[1044] 1. User utterance acquisition

[1045] Hardware: Self-driving vehicle system with microphone

[1046] Software: sounddevice library

[1047] The microphone inside the autonomous vehicle captures the user's speech in real time, temporarily stores it in the built-in memory as voice data, and then transmits it to the server. The user can interact with the system by speaking into the microphone.

[1048] 2. Audio data conversion

[1049] Hardware: Server or in-vehicle computer

[1050] Software: speech_recognition library

[1051] The server analyzes the received voice data and uses a deep learning-based model to convert the voice data into highly accurate text data, which is then passed on to the next analysis step.

[1052] 3. Context, Relational, and Sentiment Analysis

[1053] Hardware: Server or in-vehicle computer

[1054] Software: Natural language processing engine, emotion recognition module

[1055] The server analyzes the generated text data and analyzes the context of the user's speech, the relationship with the other person, and their emotions. It does this by using natural language processing technology and referring to existing conversation history and user profiles. It also uses an emotion analysis engine to identify the user's emotional state from the tone, pitch, and speed of their voice.

[1056] 4. Generating the Translation

[1057] Hardware: Server or in-vehicle computer

[1058] Software: translation engines (e.g., googletrans library), generative AI models

[1059] The server uses a generative AI model to translate the text data into the target language based on the results of contextual analysis and emotion recognition. An appropriate translation is generated that takes into account emotional information and writing style. The translated text data is stored in temporary memory.

[1060] 5. Notification of translation results

[1061] Hardware: Smartphones, tablets, dedicated devices

[1062] Software: Translation result display module

[1063] The translation results are sent from the server to the terminal and notified to the user, who can then check the translation results on the terminal screen.

[1064] 6. Control of functions inside autonomous vehicles

[1065] Hardware: Autonomous vehicle control systems

[1066] Software: Vehicle Control Module

[1067] The system controls in-vehicle functions (e.g., temperature adjustment, seat arrangement, destination change, etc.) based on the user's speech, allowing the user to efficiently operate in-vehicle functions using only speech.

[1068] Specific examples

[1069] Example 1:

[1070] When a passenger says "Adjust the air conditioning" in an autonomous vehicle, the system captures the speech and converts it into text. The server analyzes the context and sentiment of the text and translates it as "Adjust the air conditioning." The translation result is displayed on the user's device, and the air conditioning system in the vehicle is simultaneously adjusted.

[1071] Example prompt sentence:

[1072] Voice translation: Adjust the air conditioning

[1073] This will enable users to communicate naturally and smoothly within an autonomous vehicle and efficiently operate the vehicle's functions.

[1074] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1075] Step 1:

[1076] The user speaks. The user issues instructions or questions by voice into the microphone inside the autonomous vehicle. The input is voice data.

[1077] Step 2:

[1078] The device captures the user's speech and stores it in temporary memory as audio data. The input is audio data, and the output is a digital audio file. Specifically, it uses the sounddevice library to record audio and saves the data to a file.

[1079] Step 3:

[1080] The device compresses the stored audio data and transmits it to a server via the Internet. The input is a digital audio file, and the output is compressed audio data. Specifically, the audio data is compressed using a compression algorithm and then transmitted to the server.

[1081] Step 4:

[1082] The server analyzes the received voice data and converts it into text data by applying a speech recognition algorithm. The input is compressed voice data and the output is text data. Specifically, the server uses the speech_recognition library to convert voice to text data.

[1083] Step 5:

[1084] The server analyzes the generated text data to identify the context and the relationship with the other party. The input is the text data, and the output is the analyzed context and relationship information. The server uses a natural language processing engine to analyze the text and reference past conversation history and user profiles.

[1085] Step 6:

[1086] The server uses an emotion recognition engine to analyze the user's emotions. The input is voice data and text data, and the output is analyzed emotional information. Specifically, the server analyzes the tone, pitch, and speed of the voice to identify the user's emotional state.

[1087] Step 7:

[1088] The server translates the text data into the target language using a generative AI model based on the results of contextual analysis and emotion recognition. The input is the text data and the analysis results, and the output is the translated text data. The text is translated into the target language using the googletrans library.

[1089] Step 8:

[1090] The server sends the translation results to the terminal and notifies the user. The input is the translated text data, and the output is the translation results displayed to the user. The terminal displays the translation results on the screen so that the user can check them.

[1091] Step 9:

[1092] The device controls functions within the autonomous vehicle based on the user's speech. The input is translated text data and user instructions, and the output is changes to the vehicle's functions. Specifically, it sends signals to control the vehicle's air conditioning system, navigation system, etc.

[1093] Through the above processing steps, users can achieve natural and smooth communication within an autonomous vehicle and efficiently operate the vehicle's functions.

[1094] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1095] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1096] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1097] [Fourth embodiment]

[1098] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1099] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1100] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1101] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1102] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1103] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1104] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1105] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1106] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1107] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1108] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1109] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1110] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1111] This invention is a context-adaptive real-time translation system that uses speech recognition, natural language processing, and machine translation technologies. Specifically, the system acquires user utterances as voice data, converts the voice data into text data, analyzes the context and the relationship with the other party, and generates an appropriate translation based on the analysis results.

[1112] Program processing procedures and explanations

[1113] 1. User utterance transmission

[1114] The user speaks into a device (smartphone, tablet, PC, etc.), which captures the voice in real time and records it as audio data. The recorded audio data is compressed and sent to a server via the Internet.

[1115] 2. Audio data conversion

[1116] The server receives the voice data sent from the device and applies a speech recognition algorithm, which uses a deep learning-based model to convert the user's speech into text data with high accuracy. This process converts the voice data into text data and stores it in a temporary file or memory.

[1117] 3. Context and Relationship Analysis

[1118] The server analyzes the generated text data and estimates the context in which the user spoke and the relationship with the other person. Context analysis uses natural language processing technology to identify the scene by taking into account the content of the user's speech, time, location, and existing conversation history. The server also references a user profile database and past speech data to estimate the relationship between the user and the target person.

[1119] 4. Generating the Translation

[1120] The server uses a generative AI model to generate an appropriate translation based on information obtained from context and relationships. For example, it generates formal expressions for business situations and casual expressions for everyday conversations. The translation results are stored in temporary memory on the server.

[1121] 5. Gather feedback and optimize

[1122] The translation results are sent to the device and notified to the user. The user reviews the translation results and provides feedback if necessary. The device then sends this feedback to the server. The server analyzes the feedback and optimizes the parameters of the translation algorithm and generative AI model. This process improves future translation accuracy.

[1123] Specific examples

[1124] Example 1: Use in business situations

[1125] During a meeting, a user says, "How is the progress on this project?" The device captures the speech and sends it to the server. The server converts the speech into text and translates it, recognizing the formal context of the business setting. The resulting translation is "How is the progress on this project?" The translation is displayed on the device for the user to review.

[1126] Example 2: Use in everyday conversation

[1127] A user is with friends and says, "Where should we have dinner tonight?" The device captures the speech and sends it to the server. The server converts the speech to text and translates it, recognizing the casual context of everyday conversation. The resulting translation is "Where should we have dinner tonight?" The translation is displayed on the device for the user to confirm.

[1128] This invention combines speech recognition, natural language processing, and machine translation technologies to provide real-time translation suitable for a variety of situations. Furthermore, by optimizing the translation algorithm based on user feedback, it is possible to achieve more accurate translation.

[1129] The processing flow will be explained below.

[1130] Step 1:

[1131] The user speaks into the device, which uses a built-in microphone to capture the user's speech in real time, and this voice data is stored digitally in the device's temporary memory.

[1132] Step 2:

[1133] The device compresses the captured audio data and sends it over the internet to a server, where it is encrypted for security purposes.

[1134] Step 3:

[1135] The server receives the voice data sent from the terminal, decodes the received voice data, and prepares it for voice recognition processing.

[1136] Step 4:

[1137] The server applies a speech recognition algorithm to convert the audio data into text data, using a deep learning-based model to convert spoken content into text with high accuracy.

[1138] Step 5:

[1139] The server analyzes the generated text data and identifies contextual information and relationships with the other party. Contextual analysis uses natural language processing technology. This analysis categorizes the scene into categories such as business or everyday conversation.

[1140] Step 6:

[1141] The server references the user profile database and past conversation history to estimate the relationship between the user and the other party, and based on this, selects an appropriate communication style (formal, casual, etc.).

[1142] Step 7:

[1143] The server retrieves translation settings based on the analysis results and uses a generative AI model to translate the text data into the target language, adjusting the writing style and tone accordingly.

[1144] Step 8:

[1145] The server then sends the translation results to the device, which then displays them on the user interface. If necessary, speech synthesis technology can also be used to output the results as voice.

[1146] Step 9:

[1147] The user checks the translation results and provides feedback if necessary, which is sent from the device to the server.

[1148] Step 10:

[1149] The server analyzes the received feedback and updates and optimizes the parameters of the translation algorithm and generative AI model, thereby improving future translation accuracy.

[1150] In this way, this system translates user utterances in real time and provides appropriate translation results for a variety of situations. By optimizing the algorithm based on user feedback, it achieves high-accuracy translation over the long term.

[1151] Example 1

[1152] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1153] Conventional translation systems often perform simple translations without considering the context or the relationship with the target user, which can result in low translation accuracy. In addition, they lack the ability to optimize the system based on user feedback, making it difficult to continuously improve translation accuracy.

[1154] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1155] In this invention, the server includes means for acquiring user utterances as voice data, means for converting the voice data into text data, means for analyzing the relationship between the context and the target person, means for translating the text data into the target language based on the analysis results, means for notifying the user of the translation results, and means for collecting feedback from the user and optimizing the translation algorithm. This enables highly accurate translation that adapts to the context and relationship, and continuous system optimization is achieved by utilizing the feedback.

[1156] The "means for acquiring user speech as voice data" is a function for electronically capturing the voice uttered by the user and recording it as digital voice data.

[1157] The "means for converting voice data into text data" is a function for analyzing captured voice data and converting it into text data expressed as a corresponding character string.

[1158] The "means for analyzing the relationship between the context and the target person" is an analytical function for estimating the relationship between the context of an utterance and the target person of that utterance from the generated text data.

[1159] The "means for translating text data into a target language" refers to a translation function that converts the original text data into a different language based on the analyzed context and relationships.

[1160] The "means for notifying the user of the translation result" is a function for visually or audibly notifying the user of the generated translation result.

[1161] "Means for collecting user feedback and optimizing the translation algorithm" refers to a function that collects user evaluations and suggestions for corrections on translation results, and then adjusts the parameters and settings of the translation algorithm based on that information to improve accuracy.

[1162] A "deep learning-based model" is a type of algorithm that uses a multi-layer neural network to extract and transform data features, enabling advanced recognition and prediction.

[1163] A "generative AI model" is an algorithm that uses artificial intelligence technology to generate text or translate, and has the ability to generate appropriate output based on a prompt.

[1164] The present invention is a context-adaptive real-time translation system that uses speech recognition, natural language processing, and machine translation technologies. Specific embodiments for carrying out the present invention are described below.

[1165] First, the user speaks into the device (smartphone, tablet, PC, etc.). The device uses a built-in microphone to capture the user's voice in real time and record it as audio data. The recorded audio data is saved in a lossless compressed format such as FLAC and sent to a server via the Internet.

[1166] The server receives the voice data sent via the Internet, applies a speech recognition algorithm (e.g., Google Speech-to-Text API using a deep learning-based model) to analyze the voice data and convert it into text data, which is then saved in a temporary file or memory.

[1167] The server then analyzes the generated text data and estimates the context of the utterance and its relationship to the target person. This analysis uses natural language processing technology (e.g., the spaCy library). The server references a user profile database and a database of past utterances to identify the scene, taking into account the user's utterance content, time, location, and existing conversation history. It also references these databases to identify the relationship between the user and the target person.

[1168] The server generates an appropriate prompt based on the context and relationships, and generates the translation using a generative AI model (e.g., OpenAI GPT-3). The AI ​​model is fed with a prompt like this:

[1169] Input: "How is this project progressing?"

[1170] Context: Business meeting

[1171] Relationship: Superior-Subordinate

[1172] Produces the translation: "How is the progress on this project?"

[1173] The generated translation result is stored in temporary memory in the server and then sent to the terminal, where the user can check the translation result on the terminal screen or through voice output.

[1174] Finally, the user can review the translation results and provide feedback, such as ratings and suggestions for corrections. The device collects this feedback and sends it back to the server. The server analyzes the feedback and optimizes the parameters of the translation algorithm and generative AI model, thereby improving future translation accuracy.

[1175] As a concrete example, consider the case where a user says, "How is the progress on this project?" during a meeting. At this time, the device captures the audio and sends it to the server in FLAC format. The server converts it into text using the Google Speech-to-Text API and analyzes the context of the business meeting and the relationship between superiors and subordinates. Based on the analysis results, a prompt is sent to the GPT-3 model, which generates the translation result, "How is the progress on this project?". This translation result is sent to the device and confirmed by the user.

[1176] As another example, consider the case where a user is with friends and says, "Where should we have dinner tonight?" The device captures the audio and sends it to the server in FLAC format. The server converts it to text using the Google Speech-to-Text API and analyzes the casual context of everyday conversation and friendships. Based on the analysis results, the prompt is sent to the GPT-3 model, which generates the translation result, "Where should we have dinner tonight?" The translation result is then sent to the device for the user to confirm.

[1177] In this way, by combining speech recognition, natural language processing, and generative AI models, the present invention can provide highly accurate real-time translation that adapts to the user's context and relationships. Furthermore, by optimizing the translation algorithm based on user feedback, the system can be continuously improved.

[1178] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1179] Step 1: User utterance submission

[1180] explanation

[1181] The user speaks into a device (smartphone, tablet, PC, etc.), which then uses a built-in microphone to capture the voice in real time and record it as digital audio data.

[1182] input

[1183] User utterance: "How is this project going?"

[1184] output

[1185] FLAC format audio data

[1186] Specific actions

[1187] The user speaks to their smartphone, "How is this project progressing?" The device uses its built-in microphone to capture the audio and saves it as a FLAC audio file. The saved audio data is then sent to a server over the Internet.

[1188] Step 2: Convert the audio data

[1189] explanation

[1190] The server receives the voice data sent from the terminal and converts it into text data by applying a voice recognition algorithm.

[1191] input

[1192] FLAC format audio data

[1193] output

[1194] Text data: "How is this project progressing?"

[1195] Specific actions

[1196] The server receives FLAC audio data from the device. The server calls the Google Speech-to-Text API, analyzes the audio data, and converts it into text data such as "How is the progress on this project?" The converted text data is saved as a temporary file.

[1197] Step 3: Analyze context and relationships

[1198] explanation

[1199] The server analyzes the generated text data and estimates the context of the utterance and its relationship to the target person.

[1200] input

[1201] Text data: "How is this project progressing?"

[1202] output

[1203] Analysis results (context: business meeting, relationship: superior-subordinate)

[1204] Specific actions

[1205] The server analyzes the text data using the spaCy library, infers that the context of the utterance is a business meeting, and references a database of past utterances and a user profile database to determine that the user is a superior and the target person is a subordinate.

[1206] Step 4: Generate translations

[1207] explanation

[1208] The server generates prompt sentences based on context and relationships, and uses generative AI models to generate appropriate translations.

[1209] input

[1210] Analysis results (context: business meeting, relationship: superior-subordinate)

[1211] output

[1212] Translation result: "How is the progress on this project?"

[1213] Specific actions

[1214] The server inputs the following prompt into the generative AI model (OpenAI GPT-3):

[1215] Input: "How is this project progressing?"

[1216] Context: Business meeting

[1217] Relationship: Superior-Subordinate

[1218] Based on this prompt, the GPT-3 model generates the translation "How is the progress on this project?" This translation result is stored in temporary memory on the server.

[1219] Step 5: Notification of translation results

[1220] explanation

[1221] The server sends the generated translation result to the terminal and notifies the user.

[1222] input

[1223] Translation result: "How is the progress on this project?"

[1224] output

[1225] Terminal display: "How is the progress on this project?"

[1226] Specific actions

[1227] The server sends the generated translation to the device as an HTTPS response, and the user can check the translation result, "How is the progress on this project?", displayed on the device screen.

[1228] Step 6: Gather feedback and optimize

[1229] explanation

[1230] The user provides feedback on the translation result, and the terminal sends the feedback to the server, which analyzes the feedback and optimizes the translation algorithm.

[1231] input

[1232] User feedback: "The translation was accurate"

[1233] output

[1234] Optimized translation algorithm

[1235] Specific actions

[1236] The user reviews the translation result and provides feedback, rating it "the translation was accurate." The device then sends this feedback to the server, which analyzes the feedback and optimizes the translation algorithm by adjusting the model settings, resulting in higher accuracy in the next translation.

[1237] (Application example 1)

[1238] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1239] Food delivery services face the challenge of having delivery personnel communicate smoothly with foreign customers. In particular, the language barrier makes it difficult for delivery personnel to accurately understand order details and delivery instructions. This problem reduces delivery efficiency and reduces customer satisfaction.

[1240] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1241] In this invention, the server includes means for translating information and questions received by the delivery person from the customer in real time, means for acquiring the user's speech as voice data, means for converting the voice data into text data, means for translating the text data into the target language based on the analysis results, and means for notifying the user of the translation results. This enables the delivery person to communicate smoothly with foreign customers, improving delivery efficiency and customer satisfaction.

[1242] The "means for acquiring user speech as voice data" refers to a device or function for recording the user's speech as digital voice data.

[1243] The "means for converting voice data into text data" refers to an algorithm or software for analyzing acquired voice data and converting it into corresponding written text.

[1244] "Means for analyzing the context and the relationship with the other person" refers to natural language processing technologies and algorithms for understanding and analyzing the content of an utterance, its background, and the relationship between the speaker and the other person.

[1245] The "means for translating text data into a target language based on the analysis results" refers to a machine translation technology for generating text in a target language that is appropriately translated based on context and relationships.

[1246] "Means for notifying the user of the translation result" refers to an interface or function for displaying or audibly conveying the generated translation result to the user.

[1247] "A means for delivery personnel to translate information and questions received from customers in real time" refers to a system or algorithm that instantly translates the content of communication between delivery personnel and customers into other languages.

[1248] System Program

[1249] Hardware and Software Configuration

[1250] The system for implementing this invention includes a terminal operated by a user (e.g., a smartphone), a server connected via the Internet, and algorithms for performing speech recognition, natural language processing, and machine translation. The terminal has the function of capturing the user's voice and sending the voice data to the server. The server receives the voice data and performs the following main processes:

[1251] 1. Speech Recognition Using Deep Learning-Based Models

[1252] Use deep learning-based speech recognition algorithms (e.g., Google Speech-to-Text API) to convert voice data into text data with high accuracy.

[1253] 2. Natural language processing for context and relationship analysis

[1254] The converted text data is analyzed using natural language processing (e.g., the BERT model) to identify the context in which the utterance was made and the relationship between the user and the other party.

[1255] 3. Generating appropriate translations using machine translation

[1256] Uses a generative AI model (e.g. GPT-3) to generate an appropriate target language translation based on context and relationships.

[1257] 4. Feedback-driven optimization

[1258] Collect user feedback and optimize the parameters of the translation algorithm.

[1259] Specific example of operation procedure

[1260] Example 1: A delivery person communicating with a foreign customer

[1261] When the delivery person says into their smartphone, "Is this the correct address?", the smartphone captures the voice and sends the data to the server in real time.

[1262] The server receives the voice data and uses a deep learning model to convert it into text data such as "Is this address correct?"

[1263] The server uses natural language processing technology to identify that the spoken content is a question in a delivery situation and appropriately analyzes the communication with the customer.

[1264] Based on the analysis results, a generative AI model is used to generate the English translation "Is this the correct address?" and display it on the smartphone.

[1265] The delivery person checks the translation results, shows them to the customer, and provides feedback as needed to improve the system's accuracy.

[1266] Example prompt sentence:

[1267] "Please suggest a translation to ask the customer for confirmation during delivery."

[1268] "How do we understand our customers' needs and respond to them formally?"

[1269] "Please suggest a translation for a scenario in which a delivery person asks a customer where their pickup is located in a casual context."

[1270] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1271] Step 1:

[1272] User utterance transmission

[1273] The user (delivery person) speaks into a smartphone. The device (smartphone) captures the voice in real time and generates voice data. The voice data is compressed and sent to a server via the Internet. In this case, the input is the delivery person's speech, and the output is compressed voice data.

[1274] Step 2:

[1275] Audio data conversion

[1276] The server receives the voice data sent from the device. Then, it applies a deep learning-based speech recognition algorithm (e.g., Google Speech-to-Text API) to generate text data from the voice data. The input is compressed voice data, and the output is text data.

[1277] Step 3:

[1278] Context and relationship analysis

[1279] The server analyzes the generated text data. It uses a natural language processing model (e.g., BERT) to identify the context in which the user spoke and their relationship to the other party (customer). This process also references existing conversation history and user profile databases. The input is text data, and the output is information related to the context and relationships.

[1280] Step 4:

[1281] Generating Translations

[1282] The server uses a generative AI model (e.g., GPT-3) to generate an appropriate translation based on the information obtained from the context and relationships. For example, for business questions, it generates a translation using formal language. The input is information related to the context and relationships and text data, and the output is the text data translated into the target language.

[1283] Step 5:

[1284] Notification of translation results

[1285] The translation results are sent from the server to the device. The device displays the translation results to the user and notifies them by voice or text. The input is the translated text data, and the output is the translation results displayed on the device.

[1286] Step 6:

[1287] Feedback collection and optimization

[1288] The user checks the translation results and provides feedback if necessary. The device sends this feedback to the server, which analyzes the feedback and optimizes the parameters of the translation algorithm and generative AI model. The input is the user's feedback data, and the output is the optimized translation algorithm.

[1289] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1290] This invention is a context-adaptive real-time translation system that combines speech recognition, natural language processing, machine translation, and an emotion engine. The system acquires user utterances as voice data and converts the voice data into text data. In addition, the system analyzes the context, the relationship with the other person, and the user's emotions, and generates and provides an appropriate translation based on the analysis results.

[1291] Program processing procedures and explanations

[1292] 1. User utterance transmission

[1293] The user speaks into a device (smartphone, tablet, PC, etc.), which uses a built-in microphone to capture the user's speech in real time. This voice data is stored in a temporary memory in digital format, compressed, and then transmitted to a server via the Internet.

[1294] 2. Audio data conversion

[1295] The server analyzes the received voice data and applies a speech recognition algorithm, using a deep learning-based model to convert the voice data to text data with high accuracy, which is passed on to the next analysis step and stored in temporary memory.

[1296] 3. Context and Relationship Analysis

[1297] The server analyzes the generated text data and identifies the context of the user's speech and the relationship with the other party. Context analysis is performed using natural language processing technology, and a comprehensive assessment is made of the user's speech content, time, location, existing conversation history, etc. It also references a user profile database and past speech data to estimate the relationship between the user and the other party and select an appropriate communication style.

[1298] 4. Emotion Recognition Processing

[1299] The server uses an emotion engine to analyze the user's emotions from the voice data. The emotion engine analyzes the tone, pitch, and speed of the user's voice to identify the user's emotional state (e.g., joy, sadness, anger, etc.). In addition, the emotion engine also incorporates the user's facial expression recognition and biometric signal data (heart rate, skin potential, etc.) to perform more accurate emotion analysis.

[1300] 5. Generating the Translation

[1301] The server uses a generative AI model to translate the text data into the target language based on the results of contextual analysis and emotion recognition. Emotional information and writing style are also taken into consideration to generate a translation with appropriate expressions. The translated text data is stored in temporary memory.

[1302] 6. Gather feedback and optimize

[1303] The translation results are sent from the server to the device and notified to the user. The user can check the translation results on the device and provide feedback if necessary. The device then sends this feedback to the server. The server analyzes the feedback and optimizes the parameters of the translation algorithm and generative AI model to improve future translation accuracy.

[1304] Specific examples

[1305] Example 1: Use in business situations

[1306] During a meeting, a user says, "How is the progress on this project?" The device captures the speech and sends it to the server. The server converts the speech into text and recognizes the formal context of a business setting. The emotion engine also analyzes the user's tone to determine their seriousness. The result is translated into "How is the progress on this project?" in a formal and serious tone. The translation is then displayed on the device for the user to confirm.

[1307] Example 2: Use in everyday conversation

[1308] When a user is with friends, they say, "Where should we have dinner tonight?" The device captures the speech and sends it to the server. The server converts the speech to text and recognizes the casual context of everyday conversation. The emotion engine detects the joy in the user's voice. Based on this, the translation is performed, resulting in "Where should we have dinner tonight?" in a casual and joyful tone. The translation result is displayed on the device for the user to confirm.

[1309] The system of the present invention combines speech recognition, natural language processing, machine translation, and emotion recognition technologies to provide more natural and appropriate communication. Furthermore, it optimizes the algorithm based on user feedback to achieve high-accuracy translation over the long term.

[1310] The processing flow will be explained below.

[1311] Step 1:

[1312] The user speaks into the device, which uses a built-in microphone to capture the user's speech in real time, and this voice data is stored digitally in the device's temporary memory.

[1313] Step 2:

[1314] The device compresses the captured audio data and sends it to a server over the Internet, where it is encrypted for secure transmission.

[1315] Step 3:

[1316] The server receives the voice data sent from the terminal, decodes the received voice data, and prepares it for voice recognition processing.

[1317] Step 4:

[1318] The server applies a speech recognition algorithm to convert the audio data into text data. It uses a deep learning-based model to convert the speech into text with high accuracy. This text data is stored in the server's temporary memory.

[1319] Step 5:

[1320] The server then analyzes the generated text data using natural language processing technology to identify the context in which the utterance was made and the relationship with the other person. Context analysis takes into account the content of the utterance, time, location, and existing conversation history.

[1321] Step 6:

[1322] The server compares the user profile database and past conversation history to estimate the relationship between the user and the other party, and then determines the appropriate communication style, such as formal or casual.

[1323] Step 7:

[1324] The server uses an emotion engine to analyze the user's emotions from the voice data. The emotion engine analyzes the tone, pitch, and speed of the voice to identify the user's emotional state (e.g., joy, sadness, anger, etc.). It also uses facial expression recognition and biometric signal data as needed to improve the accuracy of the emotion detection.

[1325] Step 8:

[1326] Based on the results of contextual analysis and emotion recognition, the server uses a generative AI model to translate the text data into the target language, taking into account appropriate writing style and emotional expression. The generated translation data is stored in temporary memory.

[1327] Step 9:

[1328] The server sends the generated translation results to the terminal, which then displays the received translation results on the user interface. If necessary, speech synthesis technology can also be used to output the results as voice.

[1329] Step 10:

[1330] The user checks the translation results and provides feedback if necessary, which the device then sends to the server.

[1331] Step 11:

[1332] The server analyzes the received feedback and updates and optimizes the parameters of the translation algorithm and generative AI model, thereby improving future translation accuracy.

[1333] This allows the system of the present invention to translate user utterances in real time and provide appropriate translation results for various situations. Furthermore, by combining it with an emotion recognition engine, it achieves more natural translation that reflects the user's emotions. Optimization through feedback allows the system to be continuously improved, enabling highly accurate translation.

[1334] Example 2

[1335] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1336] Conventional translation systems struggle to accurately convert speech to text and to translate appropriately while taking into account the context and emotion of the translated text. They also lack effective means for collecting user feedback and improving translation accuracy. In particular, they are unable to accurately capture emotional nuances, making it difficult to achieve natural and appropriate communication.

[1337] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1338] In this invention, the server includes means for acquiring user speech as voice data, means for converting the voice data into text data, means for analyzing the context and the relationship with the other party, means for performing sentiment analysis and incorporating the results into the translation, and means for collecting user feedback and optimizing the translation algorithm. This not only improves the accuracy of voice data conversion, but also enables natural and appropriate translation that takes into account the relationship between the user and the other party, the context, and the emotions. Furthermore, translation accuracy can be continuously improved based on user feedback.

[1339] "User" refers to an individual who uses the system to make a speech.

[1340] A "terminal" is a device that a user uses to make a speech, and includes a smartphone, tablet, PC, etc.

[1341] "Voice Data" refers to information that digitally captures a user's speech.

[1342] "Text data" refers to character string information generated by analyzing voice data.

[1343] "Context" refers to the situation or background information obtained by comprehensively assessing the content of the user's utterance, time, location, existing conversation history, etc.

[1344] "Relationship" refers to the relationship between a user and a conversation partner, and is estimated by referring to past conversation history and a user profile database.

[1345] "Emotion analysis" refers to the process of analyzing a user's voice tone, pitch, rate, facial expression recognition, and biometric signal data to identify the user's emotional state.

[1346] "Translation" refers to the process of converting text data into a target language based on the analysis results.

[1347] "Feedback" refers to the evaluation and opinions of the translation results provided by the user.

[1348] A "generative AI model" refers to an algorithm that uses artificial intelligence technology to generate appropriate translations.

[1349] "Optimization" refers to the process of adjusting the system's parameters based on collected feedback to improve translation accuracy.

[1350] A "prompt sentence" is a portion of the text data input into a generative AI model, and refers to the phrase that serves as the basis for the model to generate a translation.

[1351] "Machine learning-based models" refer to trained algorithms or neural networks used to analyze audio data or generate text data.

[1352] "Natural language processing" refers to technology for analyzing the context and meaning of text data.

[1353] This invention is a context-adaptive real-time translation system that combines speech recognition, natural language processing, machine translation, and an emotion engine. Specifically, it uses the following hardware and software:

[1354] A user speaks into a device (such as a smartphone, tablet, or PC) and their speech is captured as audio data in real time. The device uses a built-in microphone and audio capture module to temporarily store this audio data in digital format (WAV format), compress it (e.g., to MP3 format), and transmit it to a server via the Internet.

[1355] The server analyzes the received voice data and converts it into text data using a speech recognition algorithm. Specifically, it uses a machine learning-based model (e.g., a model using an open-source deep learning library) to convert the voice data into text data. This text data is passed to the next analysis step and stored in temporary memory.

[1356] The server then analyzes the generated text data to identify the context of the user's speech and the relationship between the user and the other party. This context analysis uses natural language processing technology (e.g., BERT or GPT-based models). The server comprehensively determines the context by taking into account the content of the speech, time, location, existing conversation history, etc. It also estimates the relationship between the user and the other party by referencing a user profile database and past speech data.

[1357] Furthermore, the server uses an emotion engine to analyze the user's emotions from the voice data. The emotion engine analyzes the user's voice tone, pitch, speed, facial expression recognition, and bio-signal data (e.g., heart rate, skin potential) to identify the user's emotional state. A deep learning model is used for emotion analysis.

[1358] Based on the results of contextual analysis and emotion recognition, the server uses a generative AI model to translate the text data into the target language. Emotional information and writing style are also taken into consideration to generate a translation with appropriate expressions. For example, OpenAI's GPT-based model is used as the generative AI model. The generated translation data is stored in temporary memory.

[1359] The translation results are sent from the server to the device and notified to the user. The device is equipped with a notification function, allowing the user to check the translation results and provide feedback if necessary. Using the feedback function, users' ratings and opinions are entered into the device and sent to the server. The server analyzes this feedback and optimizes the parameters of the translation algorithm and generative AI model to improve future translation accuracy.

[1360] Specific use cases include the following:

[1361] Use in business situations

[1362] During a meeting, a user says, "How is the progress on this project?" The device captures the speech and sends it to the server. The server converts the speech into text and recognizes the formal context of a business setting. The emotion engine also analyzes the user's tone to determine seriousness. The result is translated into "How is the progress on this project?" in a formal and serious tone. The translation is then displayed on the device for the user to confirm.

[1363] Use in everyday conversation

[1364] When a user is with friends, they say, "Where should we have dinner tonight?" The device captures the speech and sends it to the server. The server converts the speech to text and recognizes the casual context of everyday conversation. The emotion engine detects the joy in the user's voice. Based on this, the translation is performed, resulting in "Where should we have dinner tonight?" in a casual and joyful tone. The translation result is displayed on the device for the user to confirm.

[1365] In this way, by using this system, more natural and appropriate translations are provided, enabling users to communicate smoothly. In addition, the feedback function allows for continuous system improvement, ensuring high translation accuracy over the long term.

[1366] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1367] Step 1:

[1368] The user speaks into the device. The device uses the built-in microphone to capture the user's speech in real time. The voice capture module is activated and captures the user's speech as digital audio data in WAV format. This audio data is stored in temporary memory.

[1369] Step 2:

[1370] The device compresses the captured audio data. Specifically, the audio data is converted to MP3 format and compressed. This compressed audio data is sent to a server via the Internet. The input is WAV format audio data, and the output is compressed MP3 format audio data.

[1371] Step 3:

[1372] The server analyzes the received audio data and applies a speech recognition algorithm. Specifically, the server uses a machine learning-based model (e.g., a model using a deep learning library) to convert the audio data into text data. The input is audio data in MP3 format, and the output is text data.

[1373] Step 4:

[1374] The server analyzes the generated text data and identifies the context and relationships with the other party. This analysis uses natural language processing technology (e.g., BERT or GPT-based models). The text data is input as a prompt, and the context and relationships are derived. The input is text data, and the output is context information and relationship information.

[1375] Step 5:

[1376] The server uses an emotion engine to analyze the user's emotions from the voice data. It identifies the user's emotional state by analyzing the tone, pitch, and rate of the user's voice, and also incorporating facial expression recognition and biosignal data (e.g., heart rate and skin potential). The input is text data and biosignal data, and the output is emotional state data.

[1377] Step 6:

[1378] The server uses a generative AI model to translate the text data into the target language based on the results of contextual analysis and emotion recognition. It generates an appropriate translation taking into account emotional information and writing style. The input is text data, context information, relationship information, and emotional state data, and the output is the translated text data.

[1379] Step 7:

[1380] The translation result is sent from the server to the terminal and notified to the user. The terminal uses a notification function to display the translation result to the user. The input is the translated text data, and the output is the translation result displayed to the user.

[1381] Step 8:

[1382] The user checks the translation results on the terminal and provides feedback if necessary. The terminal displays a feedback input form and sends the feedback entered by the user to the server. The input is the feedback data from the user, and the output is the feedback data sent to the server.

[1383] Step 9:

[1384] The server analyzes user feedback and optimizes the parameters of the translation algorithm and generative AI model, thereby improving future translation accuracy. The input is the feedback data, and the output is the optimized algorithm parameters.

[1385] (Application example 2)

[1386] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1387] In recent years, the use of self-driving vehicles has become more widespread, making appropriate communication with passengers crucial. However, current self-driving vehicle systems face challenges, such as language barriers and the inability to properly recognize and respond to emotions, which can reduce user satisfaction. Therefore, there is a need for a system that can analyze context and emotions in real time, translate them into the appropriate language, and even control in-vehicle functions.

[1388] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring a user's utterance as voice data, means for converting the voice data into text data, means for analyzing the context, the relationship with the other party, and the user's emotions, means for translating the text data into a target language based on the analysis results, means for notifying the user of the translation results, and means for controlling functions in the autonomous vehicle based on the user's utterance. This enables passengers to communicate naturally and appropriately in the autonomous vehicle and efficiently operate the vehicle's functions.

[1389] "User utterance" refers to words or instructions expressed by the user through speech.

[1390] "Voice data" refers to data that is a digital recording of a user's speech.

[1391] "Text data" is character information obtained by analyzing voice data.

[1392] "Context" refers to the linguistic background that is understood by taking into account the content of the user's speech, the surrounding circumstances, past speech history, and so on.

[1393] "Relationship with the other party" refers to the relationship between the user and the target person or system.

[1394] "Emotion" is a psychological state inferred based on the user's voice and other biometric signals.

[1395] "Target language" is the language to which the text is to be translated.

[1396] The "translation result" is the text resulting from converting the original text data into the target language.

[1397] An "autonomous vehicle" is a vehicle that is driven automatically without human intervention.

[1398] "Means for controlling functions" means methods or devices for operating or adjusting various devices or settings within an automated vehicle.

[1399] The system that realizes this application example includes the following components: Acquires user speech as voice data, converts the voice data into text data, and analyzes the context, the relationship with the other person, and the user's emotions. Based on the analysis results, translates the text data into the target language, notifies the user of the translation result, and controls functions within the autonomous vehicle.

[1400] Hardware and software used

[1401] 1. User utterance acquisition

[1402] Hardware: Self-driving vehicle system with microphone

[1403] Software: sounddevice library

[1404] The microphone inside the autonomous vehicle captures the user's speech in real time, temporarily stores it in the built-in memory as voice data, and then transmits it to the server. The user can interact with the system by speaking into the microphone.

[1405] 2. Audio data conversion

[1406] Hardware: Server or in-vehicle computer

[1407] Software: speech_recognition library

[1408] The server analyzes the received voice data and uses a deep learning-based model to convert the voice data into highly accurate text data, which is then passed on to the next analysis step.

[1409] 3. Context, Relational, and Sentiment Analysis

[1410] Hardware: Server or in-vehicle computer

[1411] Software: Natural language processing engine, emotion recognition module

[1412] The server analyzes the generated text data and analyzes the context of the user's speech, the relationship with the other person, and their emotions. It does this by using natural language processing technology and referring to existing conversation history and user profiles. It also uses an emotion analysis engine to identify the user's emotional state from the tone, pitch, and speed of their voice.

[1413] 4. Generating the Translation

[1414] Hardware: Server or in-vehicle computer

[1415] Software: translation engines (e.g., googletrans library), generative AI models

[1416] The server uses a generative AI model to translate the text data into the target language based on the results of contextual analysis and emotion recognition. An appropriate translation is generated that takes into account emotional information and writing style. The translated text data is stored in temporary memory.

[1417] 5. Notification of translation results

[1418] Hardware: Smartphones, tablets, dedicated devices

[1419] Software: Translation result display module

[1420] The translation results are sent from the server to the terminal and notified to the user, who can then check the translation results on the terminal screen.

[1421] 6. Control of functions inside autonomous vehicles

[1422] Hardware: Autonomous vehicle control systems

[1423] Software: Vehicle Control Module

[1424] The system controls in-vehicle functions (e.g., temperature adjustment, seat arrangement, destination change, etc.) based on the user's speech, allowing the user to efficiently operate in-vehicle functions using only speech.

[1425] Specific examples

[1426] Example 1:

[1427] When a passenger says "Adjust the air conditioning" in an autonomous vehicle, the system captures the speech and converts it into text. The server analyzes the context and sentiment of the text and translates it as "Adjust the air conditioning." The translation result is displayed on the user's device, and the air conditioning system in the vehicle is simultaneously adjusted.

[1428] Example prompt sentence:

[1429] Voice translation: Adjust the air conditioning

[1430] This will enable users to communicate naturally and smoothly within an autonomous vehicle and efficiently operate the vehicle's functions.

[1431] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1432] Step 1:

[1433] The user speaks. The user issues instructions or questions by voice into the microphone inside the autonomous vehicle. The input is voice data.

[1434] Step 2:

[1435] The device captures the user's speech and stores it in temporary memory as audio data. The input is audio data, and the output is a digital audio file. Specifically, it uses the sounddevice library to record audio and saves the data to a file.

[1436] Step 3:

[1437] The device compresses the stored audio data and transmits it to a server via the Internet. The input is a digital audio file, and the output is compressed audio data. Specifically, the audio data is compressed using a compression algorithm and then transmitted to the server.

[1438] Step 4:

[1439] The server analyzes the received voice data and converts it into text data by applying a speech recognition algorithm. The input is compressed voice data and the output is text data. Specifically, the server uses the speech_recognition library to convert voice to text data.

[1440] Step 5:

[1441] The server analyzes the generated text data to identify the context and the relationship with the other party. The input is the text data, and the output is the analyzed context and relationship information. The server uses a natural language processing engine to analyze the text and reference past conversation history and user profiles.

[1442] Step 6:

[1443] The server uses an emotion recognition engine to analyze the user's emotions. The input is voice data and text data, and the output is analyzed emotional information. Specifically, the server analyzes the tone, pitch, and speed of the voice to identify the user's emotional state.

[1444] Step 7:

[1445] The server translates the text data into the target language using a generative AI model based on the results of contextual analysis and emotion recognition. The input is the text data and the analysis results, and the output is the translated text data. The text is translated into the target language using the googletrans library.

[1446] Step 8:

[1447] The server sends the translation results to the terminal and notifies the user. The input is the translated text data, and the output is the translation results displayed to the user. The terminal displays the translation results on the screen so that the user can check them.

[1448] Step 9:

[1449] The device controls functions within the autonomous vehicle based on the user's speech. The input is translated text data and user instructions, and the output is changes to the vehicle's functions. Specifically, it sends signals to control the vehicle's air conditioning system, navigation system, etc.

[1450] Through the above processing steps, users can achieve natural and smooth communication within an autonomous vehicle and efficiently operate the vehicle's functions.

[1451] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1452] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1453] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1454] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1455] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1456] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1457] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1458] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1459] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1460] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1461] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1462] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1463] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1464] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1465] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1466] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1467] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1468] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1469] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1470] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1471] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1472] The following is further disclosed regarding the above embodiment.

[1473] (Claim 1)

[1474] means for acquiring user utterances as voice data;

[1475] means for converting voice data into text data;

[1476] A means of analyzing the context and the relationship with the other person,

[1477] means for translating the text data into a target language based on the analysis results;

[1478] means for notifying the user of the translation result;

[1479] A system including:

[1480] (Claim 2)

[1481] 10. The system of claim 1, further comprising means for using a deep learning-based model to convert the audio data.

[1482] (Claim 3)

[1483] 10. The system of claim 1, further comprising means for collecting user feedback and optimizing the translation algorithm.

[1484] "Example 1"

[1485] (Claim 1)

[1486] means for acquiring user utterances as voice data;

[1487] means for converting voice data into text data;

[1488] A means of analyzing the relationship between the context and the subject,

[1489] means for translating the text data into a target language based on the analysis results;

[1490] means for notifying the user of the translation result;

[1491] A means of collecting user feedback and optimizing the translation algorithm;

[1492] A system including:

[1493] (Claim 2)

[1494] 10. The system of claim 1, further comprising means for using a deep learning-based model to convert the audio data.

[1495] (Claim 3)

[1496] 10. The system of claim 1, further comprising means for generating a translation using a generative AI model.

[1497] "Application Example 1"

[1498] (Claim 1)

[1499] means for acquiring user utterances as voice data;

[1500] means for converting voice data into text data;

[1501] A means of analyzing the context and the relationship with the other person,

[1502] means for translating the text data into a target language based on the analysis results;

[1503] means for notifying the user of the translation result;

[1504] A means for delivery personnel to translate information and questions received from customers in real time,

[1505] A system including:

[1506] (Claim 2)

[1507] 10. The system of claim 1, further comprising means for using a deep learning-based model to convert the audio data.

[1508] (Claim 3)

[1509] 10. The system of claim 1, further comprising means for collecting user feedback and optimizing the translation algorithm.

[1510] "Example 2: Combining Emotion Engines"

[1511] (Claim 1)

[1512] means for acquiring user utterances as voice data;

[1513] means for converting voice data into text data;

[1514] A means of analyzing the context and the relationship with the other person,

[1515] means for translating the text data into a target language based on the analysis results;

[1516] means for notifying the user of the translation result;

[1517] A means of conducting sentiment analysis and reflecting the results in the translation,

[1518] A means of collecting user feedback to optimize the translation algorithm;

[1519] A system including:

[1520] (Claim 2)

[1521] 10. The system of claim 1, further comprising means for using a machine learning based model to convert the audio data.

[1522] (Claim 3)

[1523] 10. The system of claim 1, further comprising means for performing context analysis using natural language processing techniques.

[1524] (Claim 4)

[1525] 10. The system of claim 1, further comprising means for analyzing the user's voice tone, pitch, rate, facial expression recognition, and biometric signal data to identify emotional state.

[1526] (Claim 5)

[1527] The system of claim 1, further comprising means for translating text data using a generative AI model, taking into account emotional information and writing style.

[1528] (Claim 6)

[1529] 2. The system according to claim 1, further comprising means for estimating a relationship between a user and a partner by referring to a user profile database.

[1530] "Application example 2 when combining emotion engines"

[1531] (Claim 1)

[1532] means for acquiring user utterances as voice data;

[1533] means for converting voice data into text data;

[1534] A means for analyzing the context, the relationship with the other person, and the user's emotions;

[1535] means for translating the text data into a target language based on the analysis results;

[1536] means for notifying the user of the translation result;

[1537] means for controlling functions within the autonomous vehicle based on user utterances;

[1538] A system including:

[1539] (Claim 2)

[1540] 10. The system of claim 1, further comprising means for using a deep learning-based model to convert the audio data.

[1541] (Claim 3)

[1542] 10. The system of claim 1, further comprising means for collecting user feedback and optimizing the translation algorithm. [Explanation of symbols]

[1543] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for acquiring user utterances as voice data; means for converting voice data into text data; A means of analyzing the context and the relationship with the other person, means for translating the text data into a target language based on the analysis results; means for notifying the user of the translation result; A system including:

2. 10. The system of claim 1, further comprising means for using a deep learning based model to convert audio data.

3. 10. The system of claim 1, further comprising means for collecting user feedback and optimizing the translation algorithm.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A