system

JP2026085726APending Publication Date: 2026-05-25SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-11-13
Publication Date
2026-05-25

AI Technical Summary

Technical Problem

Conventional questionnaire systems are cumbersome and lack the ability to accurately analyze user emotions, leading to difficulty in eliciting genuine opinions and requiring significant time and effort in question design.

Method used

A system that uses voice input to convert user speech into text, analyzes emotions in real-time, and dynamically generates questions tailored to the user's emotional state, reducing effort and cost by presenting flexible questions.

Benefits of technology

Enables the collection of genuine user opinions efficiently by adapting survey flow to user emotions, improving product and service development through deeper insights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026085726000001_ABST
    Figure 2026085726000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Voice input method, A means of converting audio into text data, A means of analyzing user sentiment based on text data, A means for dynamically generating questions based on the results of user sentiment analysis, A means of formatting the generated questions into natural language, A means of presenting formatted questions to the user, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005] ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Conventional questionnaire systems adopt a fixed questioning method, which often makes respondents feel cumbersome and it is difficult to elicit the true opinions of users. Also, there is a problem that means for appropriately analyzing emotions during answering are lacking, so deeper insights cannot be obtained. Furthermore, it may sometimes take a lot of time and cost to design questionnaire questions.

Means for Solving the Problems

[0005] ​This invention reduces user effort by acquiring user voice data using voice input and converting it into text data. Furthermore, it analyzes the user's emotions in real time based on the text data and dynamically generates questions based on the analysis results. The generated questions are formatted into natural language and presented to the user. This allows the survey flow to flexibly change according to the user's emotions and needs, making it possible to elicit the user's true feelings. In addition, the effort and cost related to question design are reduced, allowing companies to collect customer opinions more efficiently.

[0006] "Voice input means" refers to a device or function for acquiring a user's voice as electronic data.

[0007] "Means for converting speech to text data" refers to a device or function that analyzes acquired speech and represents its content as a string of characters.

[0008] "Means for analyzing emotions" refers to a device or function that identifies a user's emotional state based on text data.

[0009] A "means for generating questions" refers to a device or function that automatically constructs questions appropriate to the situation based on the results of sentiment analysis.

[0010] "Means of shaping into natural language" refers to a device or function that converts a generated question into a natural language format that is easy for the user to understand.

[0011] "Means of presenting a question" refers to a device or function for visually or audibly presenting a formatted question to a user. [Brief explanation of the drawing]

[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0014] First, the language used in the following description will be explained.

[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0020] [First Embodiment]

[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0033] The system of the present invention is centered around a process that receives voice input from the user, analyzes its content, and dynamically generates and presents questions. The following describes the implementation of this system.

[0034] First, the user speaks into a survey terminal. This voice is captured by the voice input device built into the terminal. The captured voice data is converted into text data using a voice-to-text conversion mechanism within the terminal and sent to the server. Throughout this process, the user can spontaneously express their opinion without any special operation.

[0035] The server analyzes the received text data using a means of sentiment analysis. Sentiment analysis allows the server to determine the nuances of emotions from the user's statements. This helps identify emotions such as interest, dissatisfaction, and affection expressed by the user.

[0036] Based on the analysis results, the server automatically generates questions tailored to the user's emotions and past responses using a question generation mechanism. The generated questions are then formatted into natural language in a way that the user can easily understand.

[0037] The formatted questions are presented to the user via the terminal. For example, if the user answers, "I'm a little dissatisfied with recent products," the server can generate follow-up questions such as, "Specifically, what aspects did you find unsatisfactory?" to facilitate the conversation between the two parties.

[0038] In this way, this system enables the elicitation of genuine user opinions that are difficult to obtain through conventional, fixed survey methods, by presenting flexible questions based on user emotions. Through this process, companies can better understand user opinions and use them to improve their services and products.

[0039] The following describes the processing flow.

[0040] Step 1:

[0041] The user initiates voice input and speaks freely into the device. Questions related to the survey are presented to the user via the screen or speaker.

[0042] Step 2:

[0043] The device acquires the user's voice as audio data. This audio data is converted into text data in real time by a speech recognition function and immediately sent to the server.

[0044] Step 3:

[0045] The server uses natural language processing techniques to perform sentiment analysis on the received text data. This analysis identifies the emotional state embedded in the user's statements, revealing emotions such as dissatisfaction, satisfaction, and interest.

[0046] Step 4:

[0047] The server applies a question generation algorithm based on the analysis results to generate the next question appropriate for the situation. In this process, follow-up questions are designed to elicit further information, taking into account the user's emotions and statements.

[0048] Step 5:

[0049] The generated questions are formatted into natural language and rephrased to avoid misunderstandings. The formatted questions are then transferred to the terminal.

[0050] Step 6:

[0051] The terminal presents the received question to the user. The user is then asked to answer the question again via voice input. If necessary, steps 2 through 6 are repeated.

[0052] Step 7:

[0053] After collecting user responses, the server stores all text data and sentiment analysis results, and generates a report summarizing user opinions for company representatives.

[0054] This procedure allows the system to collect real-time feedback from users, which companies can then use to improve their products and services.

[0055] (Example 1)

[0056] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0057] Traditional survey systems primarily consist of fixed questions, making it difficult to engage in flexible dialogue based on users' emotions and needs. This resulted in a failure to elicit genuine user feedback, and the collected data was not being fully utilized for improving products and services.

[0058] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0059] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice into text data, and means for analyzing human emotions based on the text data. This enables the generation of dynamic dialogue that responds to the user's emotions.

[0060] A "device for acquiring voice input" is a device that accurately captures voice from a user and converts it into data that can be processed as a digital signal.

[0061] A "device for converting to text data" is a device that analyzes acquired audio data and converts its content into a corresponding text format.

[0062] A "device for analyzing human emotions" is a device that analyzes text data to detect the emotions and nuances contained in a statement.

[0063] A "dynamic dialogue generation device" is a device that has the function of automatically generating appropriate follow-up questions and dialogue content based on analyzed emotional data.

[0064] A "natural language formatting device" is a device that arranges generated dialogue into a natural and easily understandable format.

[0065] A "device for presenting to people" is a device for presenting formatted questions or dialogues to a user visually or audibly.

[0066] To implement this invention, an information processing system is used that includes a voice input device, a data conversion device, an emotion analysis device, a dialogue generation device, a natural language formatting device, and a presentation device. This system acquires voice input from the user, converts it into text data, and further analyzes it to generate dialogue that responds to the user's emotions, thereby enabling flexible dialogue.

[0067] Users express their opinions and feedback verbally using a dedicated terminal equipped with speech recognition software (e.g., a common speech recognition API). The terminal converts this audio into a digital signal, which is then converted into text data using the speech recognition API. At this stage, the data is sent to a server for further processing.

[0068] The server analyzes the received text data using sentiment analysis software (e.g., a general sentiment analysis model). This allows it to grasp the emotions contained in the utterances and identify the user's interests, frustrations, etc. Then, based on the analysis results, it uses a generative AI model (e.g., an AI language model) to generate dynamic dialogue. In this generation process, a highly accurate dialogue is generated using pre-configured prompt sentences. For example, a prompt sentence such as "Generate the following dialogue considering emotions based on the following statement: User statement: 'I feel this service needs improvement.'" could be used.

[0069] The generated dialogue is processed by a natural language formatting system and presented to the user via a terminal. The user can review this dialogue visually or audibly and provide further responses. Through this process, the system can respond flexibly to the user's emotions and gain a deeper understanding.

[0070] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0071] Step 1:

[0072] The user expresses their opinions and impressions verbally into a dedicated terminal. The terminal uses a voice input device to convert the voice into a digital signal. The input is the user's voice data, and the output is a digital signal. This signal is temporarily stored inside the terminal for later processing.

[0073] Step 2:

[0074] The terminal uses speech recognition software to convert digital speech signals into text data. The input is a digital signal, and the output is character data that reflects its content. The converted text data is ready to be sent to the server. During this process, noise filtering and frequency range analysis are performed to improve the accuracy of speech recognition.

[0075] Step 3:

[0076] The server analyzes received text data using sentiment analysis software. The input is text data, and the output is analyzed sentiment information. Specifically, the text data is classified into sentiment categories such as positive, negative, and neutral based on its syntax and keywords. This result is used as base data for dynamic dialogue generation.

[0077] Step 4:

[0078] The server uses a generative AI model to generate dynamic dialogue based on sentiment analysis results. This process utilizes specific prompts, such as "If the user expresses dissatisfaction, generate a dialogue asking for improvement." The input consists of sentiment information and prompts, while the output is a natural follow-up question. This question generation process also takes into account the user's past response history.

[0079] Step 5:

[0080] The server processes the generated dialogue through natural language formatting software to refine its structure before presenting it to the user via a terminal. The input is the generated follow-up question, and the output is a question in a language format easily understood by the user. The user then receives the dialogue and can prepare their next response. This process further enhances the naturalness of the dialogue.

[0081] (Application Example 1)

[0082] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0083] There is a need to provide companies with methods to gain deeper insights by collecting diverse consumer needs and feedback in real time and rapidly generating new questions through analysis. In this process, the challenge is to achieve more natural and efficient dialogue while reducing the burden on consumers.

[0084] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0085] In this invention, the server includes means for acquiring voice input, means for converting the voice into text data, and means for generating questions in response to sentiment analysis results. This makes it possible to collect consumer feedback in real time, analyze it on the spot, and quickly provide relevant questions.

[0086] "Voice input" is the process of converting human speech into digital signals using sound-sensing sensors, and acquiring the content in a format that the system can understand.

[0087] "Conversion to text data" is the process of visualizing audio, which has been acquired as a digital signal, as text, and making it a format that can be processed electronically.

[0088] "Emotional analysis" is a technique that analyzes the meaning and nuances contained in text data to identify the speaker's psychological state and emotions.

[0089] "Question generation" is the process of dynamically constructing questions to elicit new information based on the results of sentiment analysis, adapting to the situation.

[0090] "Formatting in natural language" is the process of adjusting generated questions and information into a natural language format so that users can easily understand them and do not feel any discomfort.

[0091] "To present" means to show formatted information to a user visually or audibly, thereby eliciting further responses from the user.

[0092] "Real-time processing" refers to the process of instantly processing input data without delay and quickly generating and providing results based on that data.

[0093] A "display device" is hardware that allows users to visually confirm information and data, and in this context, it includes smart glasses, etc.

[0094] The system for carrying out the present invention uses a display device such as smart glasses or a smartphone as a device that is easy for the user to use on a daily basis. The server works in cooperation with these devices to quickly process voice input data and present appropriate information to the user. The system includes the following elements and processes.

[0095] First, the user puts on smart glasses and provides their opinion about the product or service in voice. The smart glasses are equipped with a microphone to capture the voice, and this voice data is first captured within the glasses. The captured voice data is sent to a server and converted into text data using a speech recognition API such as Google® Speech-to-Text.

[0096] Subsequently, the text data is analyzed by sentiment analysis software such as IBM Watson® to identify the speaker's emotional state. Based on the information obtained from sentiment analysis, relevant questions are generated. The generated questions are then formatted using natural language processing techniques to ensure that the user can understand them without difficulty.

[0097] The formatted questions are displayed on the smart glasses' screen. At this point, processing takes place in real time, allowing users to provide immediate feedback and additional answers. The user's new responses are again captured by voice and incorporated into the new analysis and question generation process, enabling rapid interaction. For example, if a user says, "I don't really like using this product," a follow-up question such as, "What specific discomfort did you experience?" will quickly appear.

[0098] The generative AI model is prompted with the following message: "Based on the user's text and sentiment analysis results, generate the most appropriate follow-up question." This prompt helps to refine consumer feedback and facilitate more value-added conversations.

[0099] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0100] Step 1:

[0101] Users wearing smart glasses provide voice input to express their opinions about products and services. The input voice is captured by a microphone built into the glasses and converted into a digital signal as audio data.

[0102] Step 2:

[0103] Audio data is sent from the device to the server. The server calls the Google Speech-to-Text API to convert this audio data into text data. The input is audio data, and the output is text data in string format.

[0104] Step 3:

[0105] The server sends the obtained text data to IBM Watson's sentiment analysis software. The software analyzes the text and detects the emotional state. The input here is text data, and the output is metadata indicating the emotional state.

[0106] Step 4:

[0107] The server uses an AI model to generate relevant questions based on the sentiment analysis results. The generation process uses the prompt, "Based on the user's text and sentiment analysis results, please generate the most appropriate follow-up question." The output is a question in natural language.

[0108] Step 5:

[0109] The generated questions are formatted into a user-friendly format through a natural language processing algorithm. Here, the input is the generated question, and the output is the formatted question.

[0110] Step 6:

[0111] The formatted question is sent to the device and displayed on the smart glasses' screen. The user can visually confirm the displayed question and take the next action. The output is the question displayed on the screen.

[0112] Step 7:

[0113] The user enters their response again via voice through smart glasses. The new voice data is captured again by the microphone, and the process from the first step is repeated. In this step, the user provides the next opinion or response via voice.

[0114] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0115] The system of this invention uses voice input from the user to collect data through questionnaires and dialogues. This system incorporates an emotion recognition engine, enabling more accurate analysis of the user's emotions.

[0116] The user speaks freely into the device to answer a survey. This device has the ability to capture voice in real time and send the voice data to a server. The server instantly converts the voice data into text data. And here's where the emotion engine comes in. Based on the text data, the server uses the emotion engine to analyze the user's emotional state in detail. This emotion engine integrates natural language processing technology and machine learning algorithms, allowing it to recognize emotions from multiple perspectives based on the user's tone of voice and what they are saying.

[0117] The emotion engine adaptively customizes the content and format of the next questions to be presented based on the user's emotion recognition results. For example, if a user indicates a negative emotion, it is designed to generate detailed questions related to that emotion. Specifically, if a user indicates an emotion such as "I'm not satisfied with the recent service," the server will provide specific follow-up questions such as "Specifically, what aspects did not meet your expectations?"

[0118] The generated questions are formatted into natural-sounding sentences and presented to the user via voice or screen display through their device. The user can then respond verbally, and the process is repeated. This entire process allows companies to collect detailed user feedback in real time and use it to improve their products and services. This system enables data collection that accurately reflects user sentiment, making it an effective means of obtaining higher-quality feedback.

[0119] The following describes the processing flow.

[0120] Step 1:

[0121] The user speaks freely into the device to begin the survey. Questions are displayed on the device screen or presented to the user verbally.

[0122] Step 2:

[0123] The device captures the user's voice and imports it as audio data. This audio data is immediately sent to the server.

[0124] Step 3:

[0125] The server applies a speech recognition algorithm to convert the received audio data into text data. This conversion represents the user's speech as a string of characters.

[0126] Step 4:

[0127] The server analyzes the user's emotions using an emotion engine based on text data. This emotion engine uses natural language processing techniques and machine learning models to analyze the user's vocabulary and context, and to identify their emotional state (e.g., joy, dissatisfaction, surprise, etc.).

[0128] Step 5:

[0129] The server activates a question generation module based on the sentiment analysis results, dynamically creating the next question tailored to the user's current emotions. The content and format are specialized to take the user's emotions into consideration.

[0130] Step 6:

[0131] The generated questions are formatted into natural language and sent to the device. The formatting process adjusts them to a format that is easy for the user to understand and answer.

[0132] Step 7:

[0133] The device presents the user with a formatted question in either voice or text. The user can respond immediately by voice, which initiates the next step in the response process.

[0134] Step 8:

[0135] All collected data is stored and processed on servers and used for subsequent data analysis and report generation. Companies can use this data to improve their products and services.

[0136] (Example 2)

[0137] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0138] Conventional voice input systems have had the problem of difficulty in collecting data that accurately reflects the user's emotions, resulting in limited feedback. Furthermore, the inability to provide appropriate feedback and follow-up in real time, tailored to the user's emotions, has resulted in limitations on the quality and usability of the data.

[0139] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0140] In this invention, the server includes means for converting voice information into text information, means for identifying the user's emotions based on the text information, and means for dynamically creating queries according to the user's emotion identification results. This enables accurate understanding of the user's emotions, the generation of appropriate follow-up questions in real time that correspond to those emotions, and the collection of high-quality data.

[0141] "Auditory information" refers to the representation of human speech transmitted through sound in digital data format.

[0142] "Textual information" refers to text data recorded in a format that computers can understand.

[0143] A "user" is a person who uses this system, providing input and receiving feedback.

[0144] "Emotions" are elements that represent the psychological state and reactions of the user, and are analyzed from the voice and content.

[0145] "Identification" refers to the process of recognizing and classifying specific attributes or states based on data.

[0146] "Inquiry" refers to a linguistic expression that includes questions or points of clarification that the system should return to the user.

[0147] "Transformation" refers to the process of reorganizing information into other forms or styles to enable natural communication.

[0148] The system of the present invention consists of a terminal, a server, and software that links them together. The main purpose of this system is to analyze emotions in real time based on voice information from the user and to present follow-up questions appropriate to the user.

[0149] The device uses a microphone to acquire user voice information and sends it to the server as digital data. Specifically, it prepares the audio for conversion into text using speech recognition software (e.g., a speech-to-text API). This preparation involves acquiring the audio in a way that allows the user to respond naturally to surveys and conversational questions.

[0150] The server converts the received audio data into text information using a speech recognition engine (e.g., general-purpose speech recognition). Based on this text information, the server activates an emotion engine to identify the user's emotions. The emotion engine integrates natural language processing technology and machine learning algorithms to understand emotions from multiple perspectives, including the user's tone of voice and the content of their speech.

[0151] The server then uses the sentiment recognition results to dynamically generate appropriate queries through a generative AI model (e.g., a natural language generation tool). For example, if a user indicates that they are "not satisfied with recent services," the server might generate a question such as, "Specifically, what aspects did not meet your expectations?"

[0152] The generated query is transformed into natural language and presented to the user via the terminal. The user can respond verbally, and this voice response is also captured by the terminal. An example of a prompt might be, "Please tell us more details about the service that you were dissatisfied with."

[0153] This system allows companies to collect detailed user feedback in real time and use it to improve their products and services. It enables the collection of data that accurately reflects user sentiment and is an extremely effective means of obtaining high-quality feedback.

[0154] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0155] Step 1:

[0156] The device acquires voice information from the user. This input consists of what the user says during surveys or conversations. The device captures the voice through the microphone and formats it as digital audio data. Noise cancellation technology is used to improve the audio quality.

[0157] Step 2:

[0158] The terminal sends the captured audio data to the server. This data is transmitted over the network. The input the server receives is in a compressed digital audio format. The server immediately decompresses this data and prepares it to be sent to the speech recognition engine.

[0159] Step 3:

[0160] The server uses a speech recognition engine to convert speech data into text. The input is speech data, and the engine's process involves mapping acoustic patterns to phoneme representations. This results in a string of characters as output. Dictionaries and context models are used to improve speech recognition accuracy.

[0161] Step 4:

[0162] The server passes the generated text information to the emotion engine to identify the user's emotions. This input includes interpreted strings, and the engine performs emotion analysis using machine learning algorithms. The server outputs the user's emotional state with labels such as positive, negative, or neutral.

[0163] Step 5:

[0164] Based on the sentiment engine's identification results, the server dynamically creates new queries using a generative AI model. The input consists of sentiment labels and textual information, and the model performs natural language generation. In this step, follow-up questions tailored to the user's emotions are generated. For example, if a negative emotion is indicated, a question such as "Specifically, what aspects were you dissatisfied with?" might be output.

[0165] Step 6:

[0166] The server transforms the generated query into natural language and prepares it to be presented to the user via the terminal. The input is the generated query, and the server's processing is grammatical formatting using natural language processing. The terminal outputs this to the user in either speech synthesis or text display format.

[0167] Step 7:

[0168] The user responds to the presented inquiry using voice. This response is then captured again by the device as new voice input, and the process is repeated from step 1. This allows for precise data collection through repeated interactions.

[0169] (Application Example 2)

[0170] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0171] Existing voice input systems have struggled to accurately analyze user emotions and dynamically generate corresponding questions. This has resulted in a lack of effective feedback collection to improve the user experience. Furthermore, customizing responses based on emotions in real time is difficult, highlighting the need for customer support that enhances user satisfaction.

[0172] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0173] In this invention, the server includes means for converting speech into text data, means for analyzing the user's emotions based on the text data, and means for shaping questions generated using a generative AI model into natural language. This enables the generation of personalized questions tailored to the user's emotional state and the collection of high-quality feedback in real time.

[0174] "Voice input means" refers to devices or technologies for acquiring the voice spoken by a user as electronic data.

[0175] "Text data" refers to digital data that has been converted from speech into written information.

[0176] "Methods for analyzing user emotions" refer to technologies and algorithms for analyzing a user's emotional state based on text data.

[0177] "Methods for dynamically generating questions" refer to technologies and functions that appropriately create questions relevant to the situation based on the results of user sentiment analysis.

[0178] A "generative AI model" is an artificial intelligence model trained to perform adaptive text generation.

[0179] "Methods for formatting into natural language" refer to technologies and functions that convert generated text into a form that is natural and easy for users to understand.

[0180] This invention is a system that utilizes voice input from a user to perform emotion analysis and enables the generation of adaptive questions based on the results. Specifically, the process of this invention begins when a device such as a smartphone captures the voice and sends it to a server.

[0181] The server converts the audio data into text data using a speech recognition API (e.g., Google Speech-to-Text). Next, it analyzes the text data using a natural language processing library (e.g., Python's NLTK) and an emotion analysis model (e.g., BERT from the Transformers library) to comprehensively evaluate the user's emotional state.

[0182] Based on this sentiment analysis, a generative AI model (e.g., DialoGPT model) is used to generate the next question to be presented. In this step, the generated question is formatted in a way that is natural and easy for the user to understand. The server sends the generated question back to the smartphone as voice or text, presenting it to the user. The user's response is again captured via voice input, and the same process is repeated.

[0183] This configuration allows for the dynamic generation of follow-up questions based on sentiment analysis, such as "What specific impact did the delay have on you?", when a user expresses dissatisfaction, for example, with a delayed delivery of an ordered item, thereby improving the customer experience.

[0184] An example of a prompt might be something like, "Assume the user is dissatisfied, and ask him / her additional questions." This allows for highly accurate feedback collection and the generation of questions tailored to the individual needs of the user.

[0185] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0186] Step 1:

[0187] The terminal acquires the user's voice input via a microphone and captures it as digital audio data. This input audio data is then sent directly to the server.

[0188] Step 2:

[0189] The server uses a speech recognition API to convert received audio data into text data. This process analyzes the audio signal and converts it into a string using a language model. The output is text data representing the user's utterance.

[0190] Step 3:

[0191] The server analyzes text data using a natural language processing library and analyzes the user's emotions using the BERT sentiment analysis model. Based on the text data as input, it extracts specific elements of feelings and detects emotional states (e.g., joy, sadness, anger). The output is data indicating the emotional state.

[0192] Step 4:

[0193] The server uses an AI model to generate adaptive questions based on the emotional state obtained from the sentiment analysis. In this step, prompt sentences based on the sentiment data are input to the AI ​​model, and new questions are generated in text format. The output is the new question text.

[0194] Step 5:

[0195] The server formats the generated question into natural language that is easy for the user to understand and sends it to the terminal. Specifically, it adjusts the wording and grammar of the text to make it a natural response. The output is the formatted question.

[0196] Step 6:

[0197] The terminal presents the user with questions sent from the server. This presentation is done either as speech using a speech synthesis engine or as text displayed on the screen. The output is question information that the user hears or sees.

[0198] Step 7:

[0199] The user responds to the presented question again using voice, and the device acquires this audio. The input is the user's voice response data. This audio data is then processed again, returning to step 1.

[0200] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0201] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0202] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0203] [Second Embodiment]

[0204] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0205] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0206] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0207] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0208] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0209] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0210] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0211] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0212] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0213] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0214] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0215] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0216] The system of the present invention is centered around a process that receives voice input from the user, analyzes its content, and dynamically generates and presents questions. The following describes the implementation of this system.

[0217] First, the user speaks into a survey terminal. This voice is captured by the voice input device built into the terminal. The captured voice data is converted into text data using a voice-to-text conversion mechanism within the terminal and sent to the server. Throughout this process, the user can spontaneously express their opinion without any special operation.

[0218] The server analyzes the received text data using a means of sentiment analysis. Sentiment analysis allows the server to determine the nuances of emotions from the user's statements. This helps identify emotions such as interest, dissatisfaction, and affection expressed by the user.

[0219] Based on the analysis results, the server automatically generates questions tailored to the user's emotions and past responses using a question generation mechanism. The generated questions are then formatted into natural language in a way that the user can easily understand.

[0220] The formatted questions are presented to the user via the terminal. For example, if the user answers, "I'm a little dissatisfied with recent products," the server can generate follow-up questions such as, "Specifically, what aspects did you find unsatisfactory?" to facilitate the conversation between the two parties.

[0221] In this way, this system enables the elicitation of genuine user opinions that are difficult to obtain through conventional, fixed survey methods, by presenting flexible questions based on user emotions. Through this process, companies can better understand user opinions and use them to improve their services and products.

[0222] The following describes the processing flow.

[0223] Step 1:

[0224] The user initiates voice input and speaks freely into the device. Questions related to the survey are presented to the user via the screen or speaker.

[0225] Step 2:

[0226] The device acquires the user's voice as audio data. This audio data is converted into text data in real time by a speech recognition function and immediately sent to the server.

[0227] Step 3:

[0228] The server uses natural language processing techniques to perform sentiment analysis on the received text data. This analysis identifies the emotional state embedded in the user's statements, revealing emotions such as dissatisfaction, satisfaction, and interest.

[0229] Step 4:

[0230] The server applies a question generation algorithm based on the analysis results to generate the next question appropriate for the situation. In this process, follow-up questions are designed to elicit further information, taking into account the user's emotions and statements.

[0231] Step 5:

[0232] The generated questions are formatted into natural language and rephrased to avoid misunderstandings. The formatted questions are then transferred to the terminal.

[0233] Step 6:

[0234] The terminal presents the received question to the user. The user is then asked to answer the question again via voice input. If necessary, steps 2 through 6 are repeated.

[0235] Step 7:

[0236] After collecting user responses, the server stores all text data and sentiment analysis results, and generates a report summarizing user opinions for company representatives.

[0237] This procedure allows the system to collect real-time feedback from users, which companies can then use to improve their products and services.

[0238] (Example 1)

[0239] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0240] Traditional survey systems primarily consist of fixed questions, making it difficult to engage in flexible dialogue based on users' emotions and needs. This resulted in a failure to elicit genuine user feedback, and the collected data was not being fully utilized for improving products and services.

[0241] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0242] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice into text data, and means for analyzing human emotions based on the text data. This enables the generation of dynamic dialogue that responds to the user's emotions.

[0243] A "device for acquiring voice input" is a device that accurately captures voice from a user and converts it into data that can be processed as a digital signal.

[0244] A "device for converting to text data" is a device that analyzes acquired audio data and converts its content into a corresponding text format.

[0245] A "device for analyzing human emotions" is a device that analyzes text data to detect the emotions and nuances contained in a statement.

[0246] A "dynamic dialogue generation device" is a device that has the function of automatically generating appropriate follow-up questions and dialogue content based on analyzed emotional data.

[0247] A "natural language formatting device" is a device that arranges generated dialogue into a natural and easily understandable format.

[0248] A "device for presenting to people" is a device for presenting formatted questions or dialogues to a user visually or audibly.

[0249] To implement this invention, an information processing system is used that includes a voice input device, a data conversion device, an emotion analysis device, a dialogue generation device, a natural language formatting device, and a presentation device. This system acquires voice input from the user, converts it into text data, and further analyzes it to generate dialogue that responds to the user's emotions, thereby enabling flexible dialogue.

[0250] Users express their opinions and feedback verbally using a dedicated terminal equipped with speech recognition software (e.g., a common speech recognition API). The terminal converts this audio into a digital signal, which is then converted into text data using the speech recognition API. At this stage, the data is sent to a server for further processing.

[0251] The server analyzes the received text data using sentiment analysis software (e.g., a general sentiment analysis model). This allows it to grasp the emotions contained in the utterances and identify the user's interests, frustrations, etc. Then, based on the analysis results, it uses a generative AI model (e.g., an AI language model) to generate dynamic dialogue. In this generation process, a highly accurate dialogue is generated using pre-configured prompt sentences. For example, a prompt sentence such as "Generate the following dialogue considering emotions based on the following statement: User statement: 'I feel this service needs improvement.'" could be used.

[0252] The generated dialogue is processed by a natural language formatting system and presented to the user via a terminal. The user can review this dialogue visually or audibly and provide further responses. Through this process, the system can respond flexibly to the user's emotions and gain a deeper understanding.

[0253] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0254] Step 1:

[0255] The user expresses their opinions and impressions verbally into a dedicated terminal. The terminal uses a voice input device to convert the voice into a digital signal. The input is the user's voice data, and the output is a digital signal. This signal is temporarily stored inside the terminal for later processing.

[0256] Step 2:

[0257] The terminal uses speech recognition software to convert digital speech signals into text data. The input is a digital signal, and the output is character data that reflects its content. The converted text data is ready to be sent to the server. During this process, noise filtering and frequency range analysis are performed to improve the accuracy of speech recognition.

[0258] Step 3:

[0259] The server analyzes received text data using sentiment analysis software. The input is text data, and the output is analyzed sentiment information. Specifically, the text data is classified into sentiment categories such as positive, negative, and neutral based on its syntax and keywords. This result is used as base data for dynamic dialogue generation.

[0260] Step 4:

[0261] The server uses a generative AI model to generate dynamic dialogue based on sentiment analysis results. This process utilizes specific prompts, such as "If the user expresses dissatisfaction, generate a dialogue asking for improvement." The input consists of sentiment information and prompts, while the output is a natural follow-up question. This question generation process also takes into account the user's past response history.

[0262] Step 5:

[0263] The server processes the generated dialogue through natural language formatting software to refine its structure before presenting it to the user via a terminal. The input is the generated follow-up question, and the output is a question in a language format easily understood by the user. The user then receives the dialogue and can prepare their next response. This process further enhances the naturalness of the dialogue.

[0264] (Application Example 1)

[0265] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0266] There is a need to provide companies with methods to gain deeper insights by collecting diverse consumer needs and feedback in real time and rapidly generating new questions through analysis. In this process, the challenge is to achieve more natural and efficient dialogue while reducing the burden on consumers.

[0267] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0268] In this invention, the server includes means for acquiring voice input, means for converting the voice into text data, and means for generating questions in response to sentiment analysis results. This makes it possible to collect consumer feedback in real time, analyze it on the spot, and quickly provide relevant questions.

[0269] "Voice input" is the process of converting human speech into digital signals using sound-sensing sensors, and acquiring the content in a format that the system can understand.

[0270] "Conversion to text data" is the process of visualizing audio, which has been acquired as a digital signal, as text, and making it a format that can be processed electronically.

[0271] "Emotional analysis" is a technique that analyzes the meaning and nuances contained in text data to identify the speaker's psychological state and emotions.

[0272] "Question generation" is the process of dynamically constructing questions to elicit new information based on the results of sentiment analysis, adapting to the situation.

[0273] "Formatting in natural language" is the process of adjusting generated questions and information into a natural language format so that users can easily understand them and do not feel any discomfort.

[0274] "To present" means to show formatted information to a user visually or audibly, thereby eliciting further responses from the user.

[0275] "Real-time processing" refers to the process of instantly processing input data without delay and quickly generating and providing results based on that data.

[0276] A "display device" is hardware that allows users to visually confirm information and data, and in this context, it includes smart glasses, etc.

[0277] The system for carrying out the present invention uses a display device such as smart glasses or a smartphone as a device that is easy for the user to use on a daily basis. The server works in cooperation with these devices to quickly process voice input data and present appropriate information to the user. The system includes the following elements and processes.

[0278] First, the user puts on smart glasses and gives their opinion about the product or service in voice. The smart glasses are equipped with a microphone to capture the voice, and this voice data is first captured within the glasses. The captured voice data is sent to a server and converted into text data using a speech recognition API such as Google Speech-to-Text.

[0279] After that, the text data is analyzed by sentiment analysis software such as IBM Watson to identify the emotional state of the speaker. Based on the information obtained from the sentiment analysis, relevant questions are generated. The generated questions are formatted into a form that can be easily understood by the user using natural language processing technology.

[0280] The formatted questions are displayed on the display of the smart glasses. At this point, since real-time processing is performed, the user can immediately provide feedback or additional answers. The new answers from the user are again obtained as voice and incorporated into the new analysis and question generation process, enabling rapid interaction. For example, if the user states that "the usability of this product is not very good," a follow-up question such as "Specifically, what kind of discomfort did you feel?" will be quickly displayed.

[0281] For the generative AI model, a prompt sentence such as "Based on the user's text and the sentiment analysis results, please generate the next optimal follow-up question." is used. This prompt sentence helps to precisely understand the consumer feedback and facilitate more value-added conversations.

[0282] The flow of the specific process in Application Example 1 will be described using FIG.

[0283] Step 1:

[0284] A user wearing smart glasses inputs opinions about a product or service by voice. The input voice is captured by a microphone built into the glasses. As a result, it is converted into a digital signal as voice data.

[0285] Step 2:

[0286] Audio data is sent from the terminal to the server. The server calls the Google Speech-to-Text API to convert this audio data into text data. The input is audio data, and the output is text data in string format.

[0287] Step 3:

[0288] The server sends the obtained text data to the sentiment analysis software of IBM Watson. The software analyzes the text and detects the emotional state. Here, the input is text data, and the output is metadata indicating the emotional state.

[0289] Step 4:

[0290] Based on the results of the sentiment analysis, the server uses a question generation AI model to create relevant questions. In the generation process, a prompt sentence "Please generate the next optimal follow-up question based on the user's text and the sentiment analysis results." is used. The output is a question sentence in natural language.

[0291] Step 5:

[0292] The generated questions are formatted into a form that is easy for the user to understand through natural language processing algorithms. Here, the input is the generated question sentence, and the output is the formatted question sentence. [[ID=Z5]]

[0293] Step 6: [[ID=Z9]]

[0294] The formatted question sentence is sent to the terminal and presented on the display of the smart glasses. The user can visually confirm the displayed question and take the next action. The output is the question sentence displayed on the display. <000-0931> Step 7:

[0296] The user enters their response again via voice through smart glasses. The new voice data is captured again by the microphone, and the process from the first step is repeated. In this step, the user provides the next opinion or response via voice.

[0297] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0298] The system of this invention uses voice input from the user to collect data through questionnaires and dialogues. This system incorporates an emotion recognition engine, enabling more accurate analysis of the user's emotions.

[0299] The user speaks freely into the device to answer a survey. This device has the ability to capture voice in real time and send the voice data to a server. The server instantly converts the voice data into text data. And here's where the emotion engine comes in. Based on the text data, the server uses the emotion engine to analyze the user's emotional state in detail. This emotion engine integrates natural language processing technology and machine learning algorithms, allowing it to recognize emotions from multiple perspectives based on the user's tone of voice and what they are saying.

[0300] The emotion engine adaptively customizes the content and format of the next questions to be presented based on the user's emotion recognition results. For example, if a user indicates a negative emotion, it is designed to generate detailed questions related to that emotion. Specifically, if a user indicates an emotion such as "I'm not satisfied with the recent service," the server will provide specific follow-up questions such as "Specifically, what aspects did not meet your expectations?"

[0301] The generated questions are formatted into natural sentences and presented to the user via the terminal in the form of voice or on-screen display. The user can then respond verbally accordingly, and the process is repeated. Through this series of processes, the company can collect detailed opinions from the user in real time and utilize them for the improvement of products and services. This system is capable of collecting data that accurately reflects the user's emotions and is effective as a means of obtaining higher-quality feedback.

[0302] The following describes the process flow.

[0303] Step 1:

[0304] The user freely speaks towards the terminal to start the questionnaire. The question is displayed on the terminal screen or presented to the user via voice.

[0305] Step 2:

[0306] The terminal captures the user's voice and incorporates it as voice data. This voice data is immediately transmitted to the server.

[0307] Step 3:

[0308] The server applies a speech recognition algorithm to convert the received voice data into text data. Through this conversion, the content of the user's speech is expressed as a character string.

[0309] Step 4:

[0310] The server analyzes the user's emotions using an emotion engine based on the text data. This emotion engine uses natural language processing technology and a machine learning model to analyze the user's vocabulary and context and identify the emotional state (e.g., joy, dissatisfaction, surprise, etc.).

[0311] Step 5:

[0312] The server activates a question generation module based on the sentiment analysis results, dynamically creating the next question tailored to the user's current emotions. The content and format are specialized to take the user's emotions into consideration.

[0313] Step 6:

[0314] The generated questions are formatted into natural language and sent to the device. The formatting process adjusts them to a format that is easy for the user to understand and answer.

[0315] Step 7:

[0316] The device presents the user with a formatted question in either voice or text. The user can respond immediately by voice, which initiates the next step in the response process.

[0317] Step 8:

[0318] All collected data is stored and processed on servers and used for subsequent data analysis and report generation. Companies can use this data to improve their products and services.

[0319] (Example 2)

[0320] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0321] Conventional voice input systems have had the problem of difficulty in collecting data that accurately reflects the user's emotions, resulting in limited feedback. Furthermore, the inability to provide appropriate feedback and follow-up in real time, tailored to the user's emotions, has resulted in limitations on the quality and usability of the data.

[0322] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0323] In this invention, the server includes means for converting voice information into text information, means for identifying the user's emotions based on the text information, and means for dynamically creating queries according to the user's emotion identification results. This enables accurate understanding of the user's emotions, the generation of appropriate follow-up questions in real time that correspond to those emotions, and the collection of high-quality data.

[0324] "Auditory information" refers to the representation of human speech transmitted through sound in digital data format.

[0325] "Textual information" refers to text data recorded in a format that computers can understand.

[0326] A "user" is a person who uses this system, providing input and receiving feedback.

[0327] "Emotions" are elements that represent the psychological state and reactions of the user, and are analyzed from the voice and content.

[0328] "Identification" refers to the process of recognizing and classifying specific attributes or states based on data.

[0329] "Inquiry" refers to a linguistic expression that includes questions or points of clarification that the system should return to the user.

[0330] "Transformation" refers to the process of reorganizing information into other forms or styles to enable natural communication.

[0331] The system of the present invention consists of a terminal, a server, and software that links them together. The main purpose of this system is to analyze emotions in real time based on voice information from the user and to present follow-up questions appropriate to the user.

[0332] The device uses a microphone to acquire user voice information and sends it to the server as digital data. Specifically, it prepares the audio for conversion into text using speech recognition software (e.g., a speech-to-text API). This preparation involves acquiring the audio in a way that allows the user to respond naturally to surveys and conversational questions.

[0333] The server converts the received audio data into text information using a speech recognition engine (e.g., general-purpose speech recognition). Based on this text information, the server activates an emotion engine to identify the user's emotions. The emotion engine integrates natural language processing technology and machine learning algorithms to understand emotions from multiple perspectives, including the user's tone of voice and the content of their speech.

[0334] The server then uses the sentiment recognition results to dynamically generate appropriate queries through a generative AI model (e.g., a natural language generation tool). For example, if a user indicates that they are "not satisfied with recent services," the server might generate a question such as, "Specifically, what aspects did not meet your expectations?"

[0335] The generated query is transformed into natural language and presented to the user via the terminal. The user can respond verbally, and this voice response is also captured by the terminal. An example of a prompt might be, "Please tell us more details about the service that you were dissatisfied with."

[0336] This system allows companies to collect detailed user feedback in real time and use it to improve their products and services. It enables the collection of data that accurately reflects user sentiment and is an extremely effective means of obtaining high-quality feedback.

[0337] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0338] Step 1:

[0339] The device acquires voice information from the user. This input consists of what the user says during surveys or conversations. The device captures the voice through the microphone and formats it as digital audio data. Noise cancellation technology is used to improve the audio quality.

[0340] Step 2:

[0341] The terminal sends the captured audio data to the server. This data is transmitted over the network. The input the server receives is in a compressed digital audio format. The server immediately decompresses this data and prepares it to be sent to the speech recognition engine.

[0342] Step 3:

[0343] The server uses a speech recognition engine to convert speech data into text. The input is speech data, and the engine's process involves mapping acoustic patterns to phoneme representations. This results in a string of characters as output. Dictionaries and context models are used to improve speech recognition accuracy.

[0344] Step 4:

[0345] The server passes the generated text information to the emotion engine to identify the user's emotions. This input includes interpreted strings, and the engine performs emotion analysis using machine learning algorithms. The server outputs the user's emotional state with labels such as positive, negative, or neutral.

[0346] Step 5:

[0347] Based on the sentiment engine's identification results, the server dynamically creates new queries using a generative AI model. The input consists of sentiment labels and textual information, and the model performs natural language generation. In this step, follow-up questions tailored to the user's emotions are generated. For example, if a negative emotion is indicated, a question such as "Specifically, what aspects were you dissatisfied with?" might be output.

[0348] Step 6:

[0349] The server transforms the generated query into natural language and prepares it to be presented to the user via the terminal. The input is the generated query, and the server's processing is grammatical formatting using natural language processing. The terminal outputs this to the user in either speech synthesis or text display format.

[0350] Step 7:

[0351] The user responds to the presented inquiry using voice. This response is then captured again by the device as new voice input, and the process is repeated from step 1. This allows for precise data collection through repeated interactions.

[0352] (Application Example 2)

[0353] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0354] Existing voice input systems have struggled to accurately analyze user emotions and dynamically generate corresponding questions. This has resulted in a lack of effective feedback collection to improve the user experience. Furthermore, customizing responses based on emotions in real time is difficult, highlighting the need for customer support that enhances user satisfaction.

[0355] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0356] In this invention, the server includes means for converting speech into text data, means for analyzing the user's emotions based on the text data, and means for shaping questions generated using a generative AI model into natural language. This enables the generation of personalized questions tailored to the user's emotional state and the collection of high-quality feedback in real time.

[0357] "Voice input means" refers to devices or technologies for acquiring the voice spoken by a user as electronic data.

[0358] "Text data" refers to digital data that has been converted from speech into written information.

[0359] "Methods for analyzing user emotions" refer to technologies and algorithms for analyzing a user's emotional state based on text data.

[0360] "Methods for dynamically generating questions" refer to technologies and functions that appropriately create questions relevant to the situation based on the results of user sentiment analysis.

[0361] A "generative AI model" is an artificial intelligence model trained to perform adaptive text generation.

[0362] "Methods for formatting into natural language" refer to technologies and functions that convert generated text into a form that is natural and easy for users to understand.

[0363] This invention is a system that utilizes voice input from a user to perform emotion analysis and enables the generation of adaptive questions based on the results. Specifically, the process of this invention begins when a device such as a smartphone captures the voice and sends it to a server.

[0364] The server converts the audio data into text data using a speech recognition API (e.g., Google Speech-to-Text). Next, it analyzes the text data using a natural language processing library (e.g., Python's NLTK) and an emotion analysis model (e.g., BERT from the Transformers library) to comprehensively evaluate the user's emotional state.

[0365] Based on this sentiment analysis, a generative AI model (e.g., DialoGPT model) is used to generate the next question to be presented. In this step, the generated question is formatted in a way that is natural and easy for the user to understand. The server sends the generated question back to the smartphone as voice or text, presenting it to the user. The user's response is again captured via voice input, and the same process is repeated.

[0366] This configuration allows for the dynamic generation of follow-up questions based on sentiment analysis, such as "What specific impact did the delay have on you?", when a user expresses dissatisfaction, for example, with a delayed delivery of an ordered item, thereby improving the customer experience.

[0367] An example of a prompt might be something like, "Assume the user is dissatisfied, and ask him / her additional questions." This allows for highly accurate feedback collection and the generation of questions tailored to the individual needs of the user.

[0368] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0369] Step 1:

[0370] The terminal acquires the user's voice input via a microphone and captures it as digital audio data. This input audio data is then sent directly to the server.

[0371] Step 2:

[0372] The server uses a speech recognition API to convert received audio data into text data. This process analyzes the audio signal and converts it into a string using a language model. The output is text data representing the user's utterance.

[0373] Step 3:

[0374] The server analyzes text data using a natural language processing library and analyzes the user's emotions using the BERT sentiment analysis model. Based on the text data as input, it extracts specific elements of feelings and detects emotional states (e.g., joy, sadness, anger). The output is data indicating the emotional state.

[0375] Step 4:

[0376] The server uses an AI model to generate adaptive questions based on the emotional state obtained from the sentiment analysis. In this step, prompt sentences based on the sentiment data are input to the AI ​​model, and new questions are generated in text format. The output is the new question text.

[0377] Step 5:

[0378] The server formats the generated question into natural language that is easy for the user to understand and sends it to the terminal. Specifically, it adjusts the wording and grammar of the text to make it a natural response. The output is the formatted question.

[0379] Step 6:

[0380] The terminal presents the user with questions sent from the server. This presentation is done either as speech using a speech synthesis engine or as text displayed on the screen. The output is question information that the user hears or sees.

[0381] Step 7:

[0382] The user responds to the presented question again using voice, and the device acquires this audio. The input is the user's voice response data. This audio data is then processed again, returning to step 1.

[0383] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0384] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0385] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0386] [Third Embodiment]

[0387] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0388] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0389] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0390] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0391] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0392] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0393] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0394] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0395] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0396] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0397] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0398] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0399] The system of the present invention is centered around a process that receives voice input from the user, analyzes its content, and dynamically generates and presents questions. The following describes the implementation of this system.

[0400] First, the user speaks into a survey terminal. This voice is captured by the voice input device built into the terminal. The captured voice data is converted into text data using a voice-to-text conversion mechanism within the terminal and sent to the server. Throughout this process, the user can spontaneously express their opinion without any special operation.

[0401] The server analyzes the received text data using a means of sentiment analysis. Sentiment analysis allows the server to determine the nuances of emotions from the user's statements. This helps identify emotions such as interest, dissatisfaction, and affection expressed by the user.

[0402] Based on the analysis results, the server automatically generates questions tailored to the user's emotions and past responses using a question generation mechanism. The generated questions are then formatted into natural language in a way that the user can easily understand.

[0403] The formatted questions are presented to the user via the terminal. For example, if the user answers, "I'm a little dissatisfied with recent products," the server can generate follow-up questions such as, "Specifically, what aspects did you find unsatisfactory?" to facilitate the conversation between the two parties.

[0404] In this way, this system enables the elicitation of genuine user opinions that are difficult to obtain through conventional, fixed survey methods, by presenting flexible questions based on user emotions. Through this process, companies can better understand user opinions and use them to improve their services and products.

[0405] The following describes the processing flow.

[0406] Step 1:

[0407] The user initiates voice input and speaks freely into the device. Questions related to the survey are presented to the user via the screen or speaker.

[0408] Step 2:

[0409] The device acquires the user's voice as audio data. This audio data is converted into text data in real time by a speech recognition function and immediately sent to the server.

[0410] Step 3:

[0411] The server uses natural language processing techniques to perform sentiment analysis on the received text data. This analysis identifies the emotional state embedded in the user's statements, revealing emotions such as dissatisfaction, satisfaction, and interest.

[0412] Step 4:

[0413] The server applies a question generation algorithm based on the analysis results to generate the next question appropriate for the situation. In this process, follow-up questions are designed to elicit further information, taking into account the user's emotions and statements.

[0414] Step 5:

[0415] The generated questions are formatted into natural language and rephrased to avoid misunderstandings. The formatted questions are then transferred to the terminal.

[0416] Step 6:

[0417] The terminal presents the received question to the user. The user is then asked to answer the question again via voice input. If necessary, steps 2 through 6 are repeated.

[0418] Step 7:

[0419] After collecting user responses, the server stores all text data and sentiment analysis results, and generates a report summarizing user opinions for company representatives.

[0420] This procedure allows the system to collect real-time feedback from users, which companies can then use to improve their products and services.

[0421] (Example 1)

[0422] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0423] Traditional survey systems primarily consist of fixed questions, making it difficult to engage in flexible dialogue based on users' emotions and needs. This resulted in a failure to elicit genuine user feedback, and the collected data was not being fully utilized for improving products and services.

[0424] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0425] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice into text data, and means for analyzing human emotions based on the text data. This enables the generation of dynamic dialogue that responds to the user's emotions.

[0426] A "device for acquiring voice input" is a device that accurately captures voice from a user and converts it into data that can be processed as a digital signal.

[0427] A "device for converting to text data" is a device that analyzes acquired audio data and converts its content into a corresponding text format.

[0428] A "device for analyzing human emotions" is a device that analyzes text data to detect the emotions and nuances contained in a statement.

[0429] A "dynamic dialogue generation device" is a device that has the function of automatically generating appropriate follow-up questions and dialogue content based on analyzed emotional data.

[0430] A "natural language formatting device" is a device that arranges generated dialogue into a natural and easily understandable format.

[0431] A "device for presenting to people" is a device for presenting formatted questions or dialogues to a user visually or audibly.

[0432] To implement this invention, an information processing system is used that includes a voice input device, a data conversion device, an emotion analysis device, a dialogue generation device, a natural language formatting device, and a presentation device. This system acquires voice input from the user, converts it into text data, and further analyzes it to generate dialogue that responds to the user's emotions, thereby enabling flexible dialogue.

[0433] Users express their opinions and feedback verbally using a dedicated terminal equipped with speech recognition software (e.g., a common speech recognition API). The terminal converts this audio into a digital signal, which is then converted into text data using the speech recognition API. At this stage, the data is sent to a server for further processing.

[0434] The server analyzes the received text data using sentiment analysis software (e.g., a general sentiment analysis model). This allows it to grasp the emotions contained in the utterances and identify the user's interests, frustrations, etc. Then, based on the analysis results, it uses a generative AI model (e.g., an AI language model) to generate dynamic dialogue. In this generation process, a highly accurate dialogue is generated using pre-configured prompt sentences. For example, a prompt sentence such as "Generate the following dialogue considering emotions based on the following statement: User statement: 'I feel this service needs improvement.'" could be used.

[0435] The generated dialogue is processed by a natural language formatting system and presented to the user via a terminal. The user can review this dialogue visually or audibly and provide further responses. Through this process, the system can respond flexibly to the user's emotions and gain a deeper understanding.

[0436] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0437] Step 1:

[0438] The user expresses their opinions and impressions verbally into a dedicated terminal. The terminal uses a voice input device to convert the voice into a digital signal. The input is the user's voice data, and the output is a digital signal. This signal is temporarily stored inside the terminal for later processing.

[0439] Step 2:

[0440] The terminal uses speech recognition software to convert digital speech signals into text data. The input is a digital signal, and the output is character data that reflects its content. The converted text data is ready to be sent to the server. During this process, noise filtering and frequency range analysis are performed to improve the accuracy of speech recognition.

[0441] Step 3:

[0442] The server analyzes received text data using sentiment analysis software. The input is text data, and the output is analyzed sentiment information. Specifically, the text data is classified into sentiment categories such as positive, negative, and neutral based on its syntax and keywords. This result is used as base data for dynamic dialogue generation.

[0443] Step 4:

[0444] The server uses a generative AI model to generate dynamic dialogue based on sentiment analysis results. This process utilizes specific prompts, such as "If the user expresses dissatisfaction, generate a dialogue asking for improvement." The input consists of sentiment information and prompts, while the output is a natural follow-up question. This question generation process also takes into account the user's past response history.

[0445] Step 5:

[0446] The server processes the generated dialogue through natural language formatting software to refine its structure before presenting it to the user via a terminal. The input is the generated follow-up question, and the output is a question in a language format easily understood by the user. The user then receives the dialogue and can prepare their next response. This process further enhances the naturalness of the dialogue.

[0447] (Application Example 1)

[0448] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0449] There is a need to provide companies with methods to gain deeper insights by collecting diverse consumer needs and feedback in real time and rapidly generating new questions through analysis. In this process, the challenge is to achieve more natural and efficient dialogue while reducing the burden on consumers.

[0450] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0451] In this invention, the server includes means for acquiring voice input, means for converting the voice into text data, and means for generating questions in response to sentiment analysis results. This makes it possible to collect consumer feedback in real time, analyze it on the spot, and quickly provide relevant questions.

[0452] "Voice input" is the process of converting human speech into digital signals using sound-sensing sensors, and acquiring the content in a format that the system can understand.

[0453] "Conversion to text data" is the process of visualizing audio, which has been acquired as a digital signal, as text, and making it a format that can be processed electronically.

[0454] "Emotional analysis" is a technique that analyzes the meaning and nuances contained in text data to identify the speaker's psychological state and emotions.

[0455] "Question generation" is the process of dynamically constructing questions to elicit new information based on the results of sentiment analysis, adapting to the situation.

[0456] "Formatting in natural language" is the process of adjusting generated questions and information into a natural language format so that users can easily understand them and do not feel any discomfort.

[0457] "To present" means to show formatted information to a user visually or audibly, thereby eliciting further responses from the user.

[0458] "Real-time processing" refers to the process of instantly processing input data without delay and quickly generating and providing results based on that data.

[0459] A "display device" is hardware that allows users to visually confirm information and data, and in this context, it includes smart glasses, etc.

[0460] The system for carrying out the present invention uses a display device such as smart glasses or a smartphone as a device that is easy for the user to use on a daily basis. The server works in cooperation with these devices to quickly process voice input data and present appropriate information to the user. The system includes the following elements and processes.

[0461] First, the user puts on smart glasses and gives their opinion about the product or service in voice. The smart glasses are equipped with a microphone to capture the voice, and this voice data is first captured within the glasses. The captured voice data is sent to a server and converted into text data using a speech recognition API such as Google Speech-to-Text.

[0462] Subsequently, the text data is analyzed by sentiment analysis software such as IBM Watson to identify the speaker's emotional state. Based on the information obtained from sentiment analysis, relevant questions are generated. The generated questions are then formatted using natural language processing techniques to ensure that the user can understand them without difficulty.

[0463] The formatted questions are displayed on the smart glasses' screen. At this point, processing takes place in real time, allowing users to provide immediate feedback and additional answers. The user's new responses are again captured by voice and incorporated into the new analysis and question generation process, enabling rapid interaction. For example, if a user says, "I don't really like using this product," a follow-up question such as, "What specific discomfort did you experience?" will quickly appear.

[0464] The generative AI model is prompted with the following message: "Based on the user's text and sentiment analysis results, generate the most appropriate follow-up question." This prompt helps to refine consumer feedback and facilitate more value-added conversations.

[0465] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0466] Step 1:

[0467] Users wearing smart glasses provide voice input to express their opinions about products and services. The input voice is captured by a microphone built into the glasses and converted into a digital signal as audio data.

[0468] Step 2:

[0469] Audio data is sent from the device to the server. The server calls the Google Speech-to-Text API to convert this audio data into text data. The input is audio data, and the output is text data in string format.

[0470] Step 3:

[0471] The server sends the obtained text data to IBM Watson's sentiment analysis software. The software analyzes the text and detects the emotional state. The input here is text data, and the output is metadata indicating the emotional state.

[0472] Step 4:

[0473] The server uses an AI model to generate relevant questions based on the sentiment analysis results. The generation process uses the prompt, "Based on the user's text and sentiment analysis results, please generate the most appropriate follow-up question." The output is a question in natural language.

[0474] Step 5:

[0475] The generated questions are formatted into a user-friendly format through a natural language processing algorithm. Here, the input is the generated question, and the output is the formatted question.

[0476] Step 6:

[0477] The formatted question is sent to the device and displayed on the smart glasses' screen. The user can visually confirm the displayed question and take the next action. The output is the question displayed on the screen.

[0478] Step 7:

[0479] The user enters their response again via voice through smart glasses. The new voice data is captured again by the microphone, and the process from the first step is repeated. In this step, the user provides the next opinion or response via voice.

[0480] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0481] The system of this invention uses voice input from the user to collect data through questionnaires and dialogues. This system incorporates an emotion recognition engine, enabling more accurate analysis of the user's emotions.

[0482] The user speaks freely into the device to answer a survey. This device has the ability to capture voice in real time and send the voice data to a server. The server instantly converts the voice data into text data. And here's where the emotion engine comes in. Based on the text data, the server uses the emotion engine to analyze the user's emotional state in detail. This emotion engine integrates natural language processing technology and machine learning algorithms, allowing it to recognize emotions from multiple perspectives based on the user's tone of voice and what they are saying.

[0483] The emotion engine adaptively customizes the content and format of the next questions to be presented based on the user's emotion recognition results. For example, if a user indicates a negative emotion, it is designed to generate detailed questions related to that emotion. Specifically, if a user indicates an emotion such as "I'm not satisfied with the recent service," the server will provide specific follow-up questions such as "Specifically, what aspects did not meet your expectations?"

[0484] The generated questions are formatted into natural-sounding sentences and presented to the user via voice or screen display through their device. The user can then respond verbally, and the process is repeated. This entire process allows companies to collect detailed user feedback in real time and use it to improve their products and services. This system enables data collection that accurately reflects user sentiment, making it an effective means of obtaining higher-quality feedback.

[0485] The following describes the processing flow.

[0486] Step 1:

[0487] The user speaks freely into the device to begin the survey. Questions are displayed on the device screen or presented to the user verbally.

[0488] Step 2:

[0489] The device captures the user's voice and imports it as audio data. This audio data is immediately sent to the server.

[0490] Step 3:

[0491] The server applies a speech recognition algorithm to convert the received audio data into text data. This conversion represents the user's speech as a string of characters.

[0492] Step 4:

[0493] The server analyzes the user's emotions using an emotion engine based on text data. This emotion engine uses natural language processing techniques and machine learning models to analyze the user's vocabulary and context, and to identify their emotional state (e.g., joy, dissatisfaction, surprise, etc.).

[0494] Step 5:

[0495] The server activates a question generation module based on the sentiment analysis results, dynamically creating the next question tailored to the user's current emotions. The content and format are specialized to take the user's emotions into consideration.

[0496] Step 6:

[0497] The generated questions are formatted into natural language and sent to the device. The formatting process adjusts them to a format that is easy for the user to understand and answer.

[0498] Step 7:

[0499] The device presents the user with a formatted question in either voice or text. The user can respond immediately by voice, which initiates the next step in the response process.

[0500] Step 8:

[0501] All collected data is stored and processed on servers and used for subsequent data analysis and report generation. Companies can use this data to improve their products and services.

[0502] (Example 2)

[0503] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0504] Conventional voice input systems have had the problem of difficulty in collecting data that accurately reflects the user's emotions, resulting in limited feedback. Furthermore, the inability to provide appropriate feedback and follow-up in real time, tailored to the user's emotions, has resulted in limitations on the quality and usability of the data.

[0505] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0506] In this invention, the server includes means for converting voice information into text information, means for identifying the user's emotions based on the text information, and means for dynamically creating queries according to the user's emotion identification results. This enables accurate understanding of the user's emotions, the generation of appropriate follow-up questions in real time that correspond to those emotions, and the collection of high-quality data.

[0507] "Auditory information" refers to the representation of human speech transmitted through sound in digital data format.

[0508] "Textual information" refers to text data recorded in a format that computers can understand.

[0509] A "user" is a person who uses this system, providing input and receiving feedback.

[0510] "Emotions" are elements that represent the psychological state and reactions of the user, and are analyzed from the voice and content.

[0511] "Identification" refers to the process of recognizing and classifying specific attributes or states based on data.

[0512] "Inquiry" refers to a linguistic expression that includes questions or points of clarification that the system should return to the user.

[0513] "Transformation" refers to the process of reorganizing information into other forms or styles to enable natural communication.

[0514] The system of the present invention consists of a terminal, a server, and software that links them together. The main purpose of this system is to analyze emotions in real time based on voice information from the user and to present follow-up questions appropriate to the user.

[0515] The device uses a microphone to acquire user voice information and sends it to the server as digital data. Specifically, it prepares the audio for conversion into text using speech recognition software (e.g., a speech-to-text API). This preparation involves acquiring the audio in a way that allows the user to respond naturally to surveys and conversational questions.

[0516] The server converts the received audio data into text information using a speech recognition engine (e.g., general-purpose speech recognition). Based on this text information, the server activates an emotion engine to identify the user's emotions. The emotion engine integrates natural language processing technology and machine learning algorithms to understand emotions from multiple perspectives, including the user's tone of voice and the content of their speech.

[0517] The server then uses the sentiment recognition results to dynamically generate appropriate queries through a generative AI model (e.g., a natural language generation tool). For example, if a user indicates that they are "not satisfied with recent services," the server might generate a question such as, "Specifically, what aspects did not meet your expectations?"

[0518] The generated query is transformed into natural language and presented to the user via the terminal. The user can respond verbally, and this voice response is also captured by the terminal. An example of a prompt might be, "Please tell us more details about the service that you were dissatisfied with."

[0519] This system allows companies to collect detailed user feedback in real time and use it to improve their products and services. It enables the collection of data that accurately reflects user sentiment and is an extremely effective means of obtaining high-quality feedback.

[0520] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0521] Step 1:

[0522] The device acquires voice information from the user. This input consists of what the user says during surveys or conversations. The device captures the voice through the microphone and formats it as digital audio data. Noise cancellation technology is used to improve the audio quality.

[0523] Step 2:

[0524] The terminal sends the captured audio data to the server. This data is transmitted over the network. The input the server receives is in a compressed digital audio format. The server immediately decompresses this data and prepares it to be sent to the speech recognition engine.

[0525] Step 3:

[0526] The server uses a speech recognition engine to convert speech data into text. The input is speech data, and the engine's process involves mapping acoustic patterns to phoneme representations. This results in a string of characters as output. Dictionaries and context models are used to improve speech recognition accuracy.

[0527] Step 4:

[0528] The server passes the generated text information to the emotion engine to identify the user's emotions. This input includes interpreted strings, and the engine performs emotion analysis using machine learning algorithms. The server outputs the user's emotional state with labels such as positive, negative, or neutral.

[0529] Step 5:

[0530] Based on the sentiment engine's identification results, the server dynamically creates new queries using a generative AI model. The input consists of sentiment labels and textual information, and the model performs natural language generation. In this step, follow-up questions tailored to the user's emotions are generated. For example, if a negative emotion is indicated, a question such as "Specifically, what aspects were you dissatisfied with?" might be output.

[0531] Step 6:

[0532] The server transforms the generated query into natural language and prepares it to be presented to the user via the terminal. The input is the generated query, and the server's processing is grammatical formatting using natural language processing. The terminal outputs this to the user in either speech synthesis or text display format.

[0533] Step 7:

[0534] The user responds to the presented inquiry using voice. This response is then captured again by the device as new voice input, and the process is repeated from step 1. This allows for precise data collection through repeated interactions.

[0535] (Application Example 2)

[0536] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0537] Existing voice input systems have struggled to accurately analyze user emotions and dynamically generate corresponding questions. This has resulted in a lack of effective feedback collection to improve the user experience. Furthermore, customizing responses based on emotions in real time is difficult, highlighting the need for customer support that enhances user satisfaction.

[0538] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0539] In this invention, the server includes means for converting speech into text data, means for analyzing the user's emotions based on the text data, and means for shaping questions generated using a generative AI model into natural language. This enables the generation of personalized questions tailored to the user's emotional state and the collection of high-quality feedback in real time.

[0540] "Voice input means" refers to devices or technologies for acquiring the voice spoken by a user as electronic data.

[0541] "Text data" refers to digital data that has been converted from speech into written information.

[0542] "Methods for analyzing user emotions" refer to technologies and algorithms for analyzing a user's emotional state based on text data.

[0543] "Methods for dynamically generating questions" refer to technologies and functions that appropriately create questions relevant to the situation based on the results of user sentiment analysis.

[0544] A "generative AI model" is an artificial intelligence model trained to perform adaptive text generation.

[0545] "Methods for formatting into natural language" refer to technologies and functions that convert generated text into a form that is natural and easy for users to understand.

[0546] This invention is a system that utilizes voice input from a user to perform emotion analysis and enables the generation of adaptive questions based on the results. Specifically, the process of this invention begins when a device such as a smartphone captures the voice and sends it to a server.

[0547] The server converts the audio data into text data using a speech recognition API (e.g., Google Speech-to-Text). Next, it analyzes the text data using a natural language processing library (e.g., Python's NLTK) and an emotion analysis model (e.g., BERT from the Transformers library) to comprehensively evaluate the user's emotional state.

[0548] Based on this sentiment analysis, a generative AI model (e.g., DialoGPT model) is used to generate the next question to be presented. In this step, the generated question is formatted in a way that is natural and easy for the user to understand. The server sends the generated question back to the smartphone as voice or text, presenting it to the user. The user's response is again captured via voice input, and the same process is repeated.

[0549] This configuration allows for the dynamic generation of follow-up questions based on sentiment analysis, such as "What specific impact did the delay have on you?", when a user expresses dissatisfaction, for example, with a delayed delivery of an ordered item, thereby improving the customer experience.

[0550] An example of a prompt might be something like, "Assume the user is dissatisfied, and ask him / her additional questions." This allows for highly accurate feedback collection and the generation of questions tailored to the individual needs of the user.

[0551] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0552] Step 1:

[0553] The terminal acquires the user's voice input via a microphone and captures it as digital audio data. This input audio data is then sent directly to the server.

[0554] Step 2:

[0555] The server uses a speech recognition API to convert received audio data into text data. This process analyzes the audio signal and converts it into a string using a language model. The output is text data representing the user's utterance.

[0556] Step 3:

[0557] The server analyzes text data using a natural language processing library and analyzes the user's emotions using the BERT sentiment analysis model. Based on the text data as input, it extracts specific elements of feelings and detects emotional states (e.g., joy, sadness, anger). The output is data indicating the emotional state.

[0558] Step 4:

[0559] The server uses an AI model to generate adaptive questions based on the emotional state obtained from the sentiment analysis. In this step, prompt sentences based on the sentiment data are input to the AI ​​model, and new questions are generated in text format. The output is the new question text.

[0560] Step 5:

[0561] The server formats the generated question into natural language that is easy for the user to understand and sends it to the terminal. Specifically, it adjusts the wording and grammar of the text to make it a natural response. The output is the formatted question.

[0562] Step 6:

[0563] The terminal presents the user with questions sent from the server. This presentation is done either as speech using a speech synthesis engine or as text displayed on the screen. The output is question information that the user hears or sees.

[0564] Step 7:

[0565] The user responds to the presented question again using voice, and the device acquires this audio. The input is the user's voice response data. This audio data is then processed again, returning to step 1.

[0566] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0567] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0568] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0569] [Fourth Embodiment]

[0570] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0571] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0572] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0573] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0574] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0575] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0576] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0577] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0578] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0579] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0580] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0581] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0582] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0583] The system of the present invention is centered around a process that receives voice input from the user, analyzes its content, and dynamically generates and presents questions. The following describes the implementation of this system.

[0584] First, the user speaks into a survey terminal. This voice is captured by the voice input device built into the terminal. The captured voice data is converted into text data using a voice-to-text conversion mechanism within the terminal and sent to the server. Throughout this process, the user can spontaneously express their opinion without any special operation.

[0585] The server analyzes the received text data using a means of sentiment analysis. Sentiment analysis allows the server to determine the nuances of emotions from the user's statements. This helps identify emotions such as interest, dissatisfaction, and affection expressed by the user.

[0586] Based on the analysis results, the server automatically generates questions tailored to the user's emotions and past responses using a question generation mechanism. The generated questions are then formatted into natural language in a way that the user can easily understand.

[0587] The formatted questions are presented to the user via the terminal. For example, if the user answers, "I'm a little dissatisfied with recent products," the server can generate follow-up questions such as, "Specifically, what aspects did you find unsatisfactory?" to facilitate the conversation between the two parties.

[0588] In this way, this system enables the elicitation of genuine user opinions that are difficult to obtain through conventional, fixed survey methods, by presenting flexible questions based on user emotions. Through this process, companies can better understand user opinions and use them to improve their services and products.

[0589] The following describes the processing flow.

[0590] Step 1:

[0591] The user initiates voice input and speaks freely into the device. Questions related to the survey are presented to the user via the screen or speaker.

[0592] Step 2:

[0593] The device acquires the user's voice as audio data. This audio data is converted into text data in real time by a speech recognition function and immediately sent to the server.

[0594] Step 3:

[0595] The server uses natural language processing techniques to perform sentiment analysis on the received text data. This analysis identifies the emotional state embedded in the user's statements, revealing emotions such as dissatisfaction, satisfaction, and interest.

[0596] Step 4:

[0597] The server applies a question generation algorithm based on the analysis results to generate the next question appropriate for the situation. In this process, follow-up questions are designed to elicit further information, taking into account the user's emotions and statements.

[0598] Step 5:

[0599] The generated questions are formatted into natural language and rephrased to avoid misunderstandings. The formatted questions are then transferred to the terminal.

[0600] Step 6:

[0601] The terminal presents the received question to the user. The user is then asked to answer the question again via voice input. If necessary, steps 2 through 6 are repeated.

[0602] Step 7:

[0603] After collecting user responses, the server stores all text data and sentiment analysis results, and generates a report summarizing user opinions for company representatives.

[0604] This procedure allows the system to collect real-time feedback from users, which companies can then use to improve their products and services.

[0605] (Example 1)

[0606] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0607] Traditional survey systems primarily consist of fixed questions, making it difficult to engage in flexible dialogue based on users' emotions and needs. This resulted in a failure to elicit genuine user feedback, and the collected data was not being fully utilized for improving products and services.

[0608] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0609] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice into text data, and means for analyzing human emotions based on the text data. This enables the generation of dynamic dialogue that responds to the user's emotions.

[0610] A "device for acquiring voice input" is a device that accurately captures voice from a user and converts it into data that can be processed as a digital signal.

[0611] A "device for converting to text data" is a device that analyzes acquired audio data and converts its content into a corresponding text format.

[0612] A "device for analyzing human emotions" is a device that analyzes text data to detect the emotions and nuances contained in a statement.

[0613] A "dynamic dialogue generation device" is a device that has the function of automatically generating appropriate follow-up questions and dialogue content based on analyzed emotional data.

[0614] A "natural language formatting device" is a device that arranges generated dialogue into a natural and easily understandable format.

[0615] A "device for presenting to people" is a device for presenting formatted questions or dialogues to a user visually or audibly.

[0616] To implement this invention, an information processing system is used that includes a voice input device, a data conversion device, an emotion analysis device, a dialogue generation device, a natural language formatting device, and a presentation device. This system acquires voice input from the user, converts it into text data, and further analyzes it to generate dialogue that responds to the user's emotions, thereby enabling flexible dialogue.

[0617] Users express their opinions and feedback verbally using a dedicated terminal equipped with speech recognition software (e.g., a common speech recognition API). The terminal converts this audio into a digital signal, which is then converted into text data using the speech recognition API. At this stage, the data is sent to a server for further processing.

[0618] The server analyzes the received text data using sentiment analysis software (e.g., a general sentiment analysis model). This allows it to grasp the emotions contained in the utterances and identify the user's interests, frustrations, etc. Then, based on the analysis results, it uses a generative AI model (e.g., an AI language model) to generate dynamic dialogue. In this generation process, a highly accurate dialogue is generated using pre-configured prompt sentences. For example, a prompt sentence such as "Generate the following dialogue considering emotions based on the following statement: User statement: 'I feel this service needs improvement.'" could be used.

[0619] The generated dialogue is processed by a natural language formatting system and presented to the user via a terminal. The user can review this dialogue visually or audibly and provide further responses. Through this process, the system can respond flexibly to the user's emotions and gain a deeper understanding.

[0620] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0621] Step 1:

[0622] The user expresses their opinions and impressions verbally into a dedicated terminal. The terminal uses a voice input device to convert the voice into a digital signal. The input is the user's voice data, and the output is a digital signal. This signal is temporarily stored inside the terminal for later processing.

[0623] Step 2:

[0624] The terminal uses speech recognition software to convert digital speech signals into text data. The input is a digital signal, and the output is character data that reflects its content. The converted text data is ready to be sent to the server. During this process, noise filtering and frequency range analysis are performed to improve the accuracy of speech recognition.

[0625] Step 3:

[0626] The server analyzes received text data using sentiment analysis software. The input is text data, and the output is analyzed sentiment information. Specifically, the text data is classified into sentiment categories such as positive, negative, and neutral based on its syntax and keywords. This result is used as base data for dynamic dialogue generation.

[0627] Step 4:

[0628] The server uses a generative AI model to generate dynamic dialogue based on sentiment analysis results. This process utilizes specific prompts, such as "If the user expresses dissatisfaction, generate a dialogue asking for improvement." The input consists of sentiment information and prompts, while the output is a natural follow-up question. This question generation process also takes into account the user's past response history.

[0629] Step 5:

[0630] The server processes the generated dialogue through natural language formatting software to refine its structure before presenting it to the user via a terminal. The input is the generated follow-up question, and the output is a question in a language format easily understood by the user. The user then receives the dialogue and can prepare their next response. This process further enhances the naturalness of the dialogue.

[0631] (Application Example 1)

[0632] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0633] There is a need to provide companies with methods to gain deeper insights by collecting diverse consumer needs and feedback in real time and rapidly generating new questions through analysis. In this process, the challenge is to achieve more natural and efficient dialogue while reducing the burden on consumers.

[0634] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0635] In this invention, the server includes means for acquiring voice input, means for converting the voice into text data, and means for generating questions in response to sentiment analysis results. This makes it possible to collect consumer feedback in real time, analyze it on the spot, and quickly provide relevant questions.

[0636] "Voice input" is the process of converting human speech into digital signals using sound-sensing sensors, and acquiring the content in a format that the system can understand.

[0637] "Conversion to text data" is the process of visualizing audio, which has been acquired as a digital signal, as text, and making it a format that can be processed electronically.

[0638] "Emotional analysis" is a technique that analyzes the meaning and nuances contained in text data to identify the speaker's psychological state and emotions.

[0639] "Question generation" is the process of dynamically constructing questions to elicit new information based on the results of sentiment analysis, adapting to the situation.

[0640] "Formatting in natural language" is the process of adjusting generated questions and information into a natural language format so that users can easily understand them and do not feel any discomfort.

[0641] "To present" means to show formatted information to a user visually or audibly, thereby eliciting further responses from the user.

[0642] "Real-time processing" refers to the process of instantly processing input data without delay and quickly generating and providing results based on that data.

[0643] A "display device" is hardware that allows users to visually confirm information and data, and in this context, it includes smart glasses, etc.

[0644] The system for carrying out the present invention uses a display device such as smart glasses or a smartphone as a device that is easy for the user to use on a daily basis. The server works in cooperation with these devices to quickly process voice input data and present appropriate information to the user. The system includes the following elements and processes.

[0645] First, the user puts on smart glasses and gives their opinion about the product or service in voice. The smart glasses are equipped with a microphone to capture the voice, and this voice data is first captured within the glasses. The captured voice data is sent to a server and converted into text data using a speech recognition API such as Google Speech-to-Text.

[0646] Subsequently, the text data is analyzed by sentiment analysis software such as IBM Watson to identify the speaker's emotional state. Based on the information obtained from sentiment analysis, relevant questions are generated. The generated questions are then formatted using natural language processing techniques to ensure that the user can understand them without difficulty.

[0647] The formatted questions are displayed on the smart glasses' screen. At this point, processing takes place in real time, allowing users to provide immediate feedback and additional answers. The user's new responses are again captured by voice and incorporated into the new analysis and question generation process, enabling rapid interaction. For example, if a user says, "I don't really like using this product," a follow-up question such as, "What specific discomfort did you experience?" will quickly appear.

[0648] The generative AI model is prompted with the following message: "Based on the user's text and sentiment analysis results, generate the most appropriate follow-up question." This prompt helps to refine consumer feedback and facilitate more value-added conversations.

[0649] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0650] Step 1:

[0651] Users wearing smart glasses provide voice input to express their opinions about products and services. The input voice is captured by a microphone built into the glasses and converted into a digital signal as audio data.

[0652] Step 2:

[0653] Audio data is sent from the device to the server. The server calls the Google Speech-to-Text API to convert this audio data into text data. The input is audio data, and the output is text data in string format.

[0654] Step 3:

[0655] The server sends the obtained text data to IBM Watson's sentiment analysis software. The software analyzes the text and detects the emotional state. The input here is text data, and the output is metadata indicating the emotional state.

[0656] Step 4:

[0657] The server uses an AI model to generate relevant questions based on the sentiment analysis results. The generation process uses the prompt, "Based on the user's text and sentiment analysis results, please generate the most appropriate follow-up question." The output is a question in natural language.

[0658] Step 5:

[0659] The generated questions are formatted into a user-friendly format through a natural language processing algorithm. Here, the input is the generated question, and the output is the formatted question.

[0660] Step 6:

[0661] The formatted question is sent to the device and displayed on the smart glasses' screen. The user can visually confirm the displayed question and take the next action. The output is the question displayed on the screen.

[0662] Step 7:

[0663] The user enters their response again via voice through smart glasses. The new voice data is captured again by the microphone, and the process from the first step is repeated. In this step, the user provides the next opinion or response via voice.

[0664] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0665] The system of this invention uses voice input from the user to collect data through questionnaires and dialogues. This system incorporates an emotion recognition engine, enabling more accurate analysis of the user's emotions.

[0666] The user speaks freely into the device to answer a survey. This device has the ability to capture voice in real time and send the voice data to a server. The server instantly converts the voice data into text data. And here's where the emotion engine comes in. Based on the text data, the server uses the emotion engine to analyze the user's emotional state in detail. This emotion engine integrates natural language processing technology and machine learning algorithms, allowing it to recognize emotions from multiple perspectives based on the user's tone of voice and what they are saying.

[0667] The emotion engine adaptively customizes the content and format of the next questions to be presented based on the user's emotion recognition results. For example, if a user indicates a negative emotion, it is designed to generate detailed questions related to that emotion. Specifically, if a user indicates an emotion such as "I'm not satisfied with the recent service," the server will provide specific follow-up questions such as "Specifically, what aspects did not meet your expectations?"

[0668] The generated questions are formatted into natural-sounding sentences and presented to the user via voice or screen display through their device. The user can then respond verbally, and the process is repeated. This entire process allows companies to collect detailed user feedback in real time and use it to improve their products and services. This system enables data collection that accurately reflects user sentiment, making it an effective means of obtaining higher-quality feedback.

[0669] The following describes the processing flow.

[0670] Step 1:

[0671] The user speaks freely into the device to begin the survey. Questions are displayed on the device screen or presented to the user verbally.

[0672] Step 2:

[0673] The device captures the user's voice and imports it as audio data. This audio data is immediately sent to the server.

[0674] Step 3:

[0675] The server applies a speech recognition algorithm to convert the received audio data into text data. This conversion represents the user's speech as a string of characters.

[0676] Step 4:

[0677] The server analyzes the user's emotions using an emotion engine based on text data. This emotion engine uses natural language processing techniques and machine learning models to analyze the user's vocabulary and context, and to identify their emotional state (e.g., joy, dissatisfaction, surprise, etc.).

[0678] Step 5:

[0679] The server activates a question generation module based on the sentiment analysis results, dynamically creating the next question tailored to the user's current emotions. The content and format are specialized to take the user's emotions into consideration.

[0680] Step 6:

[0681] The generated questions are formatted into natural language and sent to the device. The formatting process adjusts them to a format that is easy for the user to understand and answer.

[0682] Step 7:

[0683] The device presents the user with a formatted question in either voice or text. The user can respond immediately by voice, which initiates the next step in the response process.

[0684] Step 8:

[0685] All collected data is stored and processed on servers and used for subsequent data analysis and report generation. Companies can use this data to improve their products and services.

[0686] (Example 2)

[0687] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0688] Conventional voice input systems have had the problem of difficulty in collecting data that accurately reflects the user's emotions, resulting in limited feedback. Furthermore, the inability to provide appropriate feedback and follow-up in real time, tailored to the user's emotions, has resulted in limitations on the quality and usability of the data.

[0689] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0690] In this invention, the server includes means for converting voice information into text information, means for identifying the user's emotions based on the text information, and means for dynamically creating queries according to the user's emotion identification results. This enables accurate understanding of the user's emotions, the generation of appropriate follow-up questions in real time that correspond to those emotions, and the collection of high-quality data.

[0691] "Auditory information" refers to the representation of human speech transmitted through sound in digital data format.

[0692] "Textual information" refers to text data recorded in a format that computers can understand.

[0693] A "user" is a person who uses this system, providing input and receiving feedback.

[0694] "Emotions" are elements that represent the psychological state and reactions of the user, and are analyzed from the voice and content.

[0695] "Identification" refers to the process of recognizing and classifying specific attributes or states based on data.

[0696] "Inquiry" refers to a linguistic expression that includes questions or points of clarification that the system should return to the user.

[0697] "Transformation" refers to the process of reorganizing information into other forms or styles to enable natural communication.

[0698] The system of the present invention consists of a terminal, a server, and software that links them together. The main purpose of this system is to analyze emotions in real time based on voice information from the user and to present follow-up questions appropriate to the user.

[0699] The device uses a microphone to acquire user voice information and sends it to the server as digital data. Specifically, it prepares the audio for conversion into text using speech recognition software (e.g., a speech-to-text API). This preparation involves acquiring the audio in a way that allows the user to respond naturally to surveys and conversational questions.

[0700] The server converts the received audio data into text information using a speech recognition engine (e.g., general-purpose speech recognition). Based on this text information, the server activates an emotion engine to identify the user's emotions. The emotion engine integrates natural language processing technology and machine learning algorithms to understand emotions from multiple perspectives, including the user's tone of voice and the content of their speech.

[0701] The server then uses the sentiment recognition results to dynamically generate appropriate queries through a generative AI model (e.g., a natural language generation tool). For example, if a user indicates that they are "not satisfied with recent services," the server might generate a question such as, "Specifically, what aspects did not meet your expectations?"

[0702] The generated query is transformed into natural language and presented to the user via the terminal. The user can respond verbally, and this voice response is also captured by the terminal. An example of a prompt might be, "Please tell us more details about the service that you were dissatisfied with."

[0703] This system allows companies to collect detailed user feedback in real time and use it to improve their products and services. It enables the collection of data that accurately reflects user sentiment and is an extremely effective means of obtaining high-quality feedback.

[0704] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0705] Step 1:

[0706] The device acquires voice information from the user. This input consists of what the user says during surveys or conversations. The device captures the voice through the microphone and formats it as digital audio data. Noise cancellation technology is used to improve the audio quality.

[0707] Step 2:

[0708] The terminal sends the captured audio data to the server. This data is transmitted over the network. The input the server receives is in a compressed digital audio format. The server immediately decompresses this data and prepares it to be sent to the speech recognition engine.

[0709] Step 3:

[0710] The server uses a speech recognition engine to convert speech data into text. The input is speech data, and the engine's process involves mapping acoustic patterns to phoneme representations. This results in a string of characters as output. Dictionaries and context models are used to improve speech recognition accuracy.

[0711] Step 4:

[0712] The server passes the generated text information to the emotion engine to identify the user's emotions. This input includes interpreted strings, and the engine performs emotion analysis using machine learning algorithms. The server outputs the user's emotional state with labels such as positive, negative, or neutral.

[0713] Step 5:

[0714] Based on the sentiment engine's identification results, the server dynamically creates new queries using a generative AI model. The input consists of sentiment labels and textual information, and the model performs natural language generation. In this step, follow-up questions tailored to the user's emotions are generated. For example, if a negative emotion is indicated, a question such as "Specifically, what aspects were you dissatisfied with?" might be output.

[0715] Step 6:

[0716] The server transforms the generated query into natural language and prepares it to be presented to the user via the terminal. The input is the generated query, and the server's processing is grammatical formatting using natural language processing. The terminal outputs this to the user in either speech synthesis or text display format.

[0717] Step 7:

[0718] The user responds to the presented inquiry using voice. This response is then captured again by the device as new voice input, and the process is repeated from step 1. This allows for precise data collection through repeated interactions.

[0719] (Application Example 2)

[0720] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0721] Existing voice input systems have struggled to accurately analyze user emotions and dynamically generate corresponding questions. This has resulted in a lack of effective feedback collection to improve the user experience. Furthermore, customizing responses based on emotions in real time is difficult, highlighting the need for customer support that enhances user satisfaction.

[0722] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0723] In this invention, the server includes means for converting speech into text data, means for analyzing the user's emotions based on the text data, and means for shaping questions generated using a generative AI model into natural language. This enables the generation of personalized questions tailored to the user's emotional state and the collection of high-quality feedback in real time.

[0724] "Voice input means" refers to devices or technologies for acquiring the voice spoken by a user as electronic data.

[0725] "Text data" refers to digital data that has been converted from speech into written information.

[0726] "Methods for analyzing user emotions" refer to technologies and algorithms for analyzing a user's emotional state based on text data.

[0727] "Methods for dynamically generating questions" refer to technologies and functions that appropriately create questions relevant to the situation based on the results of user sentiment analysis.

[0728] A "generative AI model" is an artificial intelligence model trained to perform adaptive text generation.

[0729] "Methods for formatting into natural language" refer to technologies and functions that convert generated text into a form that is natural and easy for users to understand.

[0730] This invention is a system that utilizes voice input from a user to perform emotion analysis and enables the generation of adaptive questions based on the results. Specifically, the process of this invention begins when a device such as a smartphone captures the voice and sends it to a server.

[0731] The server converts the audio data into text data using a speech recognition API (e.g., Google Speech-to-Text). Next, it analyzes the text data using a natural language processing library (e.g., Python's NLTK) and an emotion analysis model (e.g., BERT from the Transformers library) to comprehensively evaluate the user's emotional state.

[0732] Based on this sentiment analysis, a generative AI model (e.g., DialoGPT model) is used to generate the next question to be presented. In this step, the generated question is formatted in a way that is natural and easy for the user to understand. The server sends the generated question back to the smartphone as voice or text, presenting it to the user. The user's response is again captured via voice input, and the same process is repeated.

[0733] This configuration allows for the dynamic generation of follow-up questions based on sentiment analysis, such as "What specific impact did the delay have on you?", when a user expresses dissatisfaction, for example, with a delayed delivery of an ordered item, thereby improving the customer experience.

[0734] An example of a prompt might be something like, "Assume the user is dissatisfied, and ask him / her additional questions." This allows for highly accurate feedback collection and the generation of questions tailored to the individual needs of the user.

[0735] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0736] Step 1:

[0737] The terminal acquires the user's voice input via a microphone and captures it as digital audio data. This input audio data is then sent directly to the server.

[0738] Step 2:

[0739] The server uses a speech recognition API to convert received audio data into text data. This process analyzes the audio signal and converts it into a string using a language model. The output is text data representing the user's utterance.

[0740] Step 3:

[0741] The server analyzes text data using a natural language processing library and analyzes the user's emotions using the BERT sentiment analysis model. Based on the text data as input, it extracts specific elements of feelings and detects emotional states (e.g., joy, sadness, anger). The output is data indicating the emotional state.

[0742] Step 4:

[0743] The server uses an AI model to generate adaptive questions based on the emotional state obtained from the sentiment analysis. In this step, prompt sentences based on the sentiment data are input to the AI ​​model, and new questions are generated in text format. The output is the new question text.

[0744] Step 5:

[0745] The server formats the generated question into natural language that is easy for the user to understand and sends it to the terminal. Specifically, it adjusts the wording and grammar of the text to make it a natural response. The output is the formatted question.

[0746] Step 6:

[0747] The terminal presents the user with questions sent from the server. This presentation is done either as speech using a speech synthesis engine or as text displayed on the screen. The output is question information that the user hears or sees.

[0748] Step 7:

[0749] The user responds to the presented question again using voice, and the device acquires this audio. The input is the user's voice response data. This audio data is then processed again, returning to step 1.

[0750] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0751] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0752] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0753] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0754] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0755] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0756] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0757] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0758] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0759] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0760] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0761] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0762] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0763] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0764] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0765] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0766] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0767] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0768] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0769] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0770] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0771] The following is further disclosed regarding the embodiments described above.

[0772] (Claim 1)

[0773] Voice input method,

[0774] A means of converting audio into text data,

[0775] A means of analyzing user sentiment based on text data,

[0776] A means for dynamically generating questions based on the results of user sentiment analysis,

[0777] A means of formatting the generated questions into natural language,

[0778] A means of presenting formatted questions to the user,

[0779] A system that includes this.

[0780] (Claim 2)

[0781] The system according to claim 1, comprising means for processing voice input in real time.

[0782] (Claim 3)

[0783] The system according to claim 1, characterized in that the user's response to the generated question is again obtained by voice input means.

[0784] "Example 1"

[0785] (Claim 1)

[0786] A device for acquiring voice input,

[0787] A device that converts acquired audio into text data,

[0788] A device that analyzes human emotions based on text data,

[0789] A device that dynamically generates dialogue based on the results of emotion analysis,

[0790] A device that formats the generated dialogue into natural language,

[0791] A device that presents a restructured dialogue to a person,

[0792] An information processing system that includes this.

[0793] (Claim 2)

[0794] The information processing system according to claim 1, comprising a device for instantly processing voice input.

[0795] (Claim 3)

[0796] The information processing system according to claim 1, characterized in that it acquires a person's response to a generated dialogue again using a voice input device.

[0797] "Application Example 1"

[0798] (Claim 1)

[0799] A means of acquiring voice input,

[0800] A means of converting audio into text data,

[0801] A means of analyzing emotions based on text data,

[0802] A means of generating questions based on the results of sentiment analysis,

[0803] A method for formatting the generated questions into natural language,

[0804] A means of presenting a formatted question,

[0805] Means for displaying on a terminal device,

[0806] A system that includes this.

[0807] (Claim 2)

[0808] The system according to claim 1, comprising a display device for processing audio data in real time and enabling an immediate response to a presented question.

[0809] (Claim 3)

[0810] The system according to claim 1, which obtains answers to the generated questions again using a voice input means and proposes new questions based on the analysis results.

[0811] "Example 2 of combining an emotion engine"

[0812] (Claim 1)

[0813] Means for acquiring audio information,

[0814] A means of converting audio information into text information,

[0815] A means of identifying a user's emotions based on textual information,

[0816] A means of dynamically creating queries based on the results of user sentiment recognition,

[0817] A means of transforming the generated query into natural language,

[0818] A means of presenting a modified query to the user,

[0819] A system that includes this.

[0820] (Claim 2)

[0821] The system according to claim 1, comprising means for processing audio information in real time.

[0822] (Claim 3)

[0823] The system according to claim 1, characterized in that the user's response to the created inquiry is again obtained using voice information acquisition means.

[0824] "Application example 2 when combining with an emotional engine"

[0825] (Claim 1)

[0826] Voice input method,

[0827] A means of converting audio into text data,

[0828] A means of analyzing user sentiment based on text data,

[0829] A means for dynamically generating questions based on the results of user sentiment analysis,

[0830] A method for formatting questions generated using a generative AI model into natural language,

[0831] A means of presenting formatted questions to users,

[0832] A system that includes this.

[0833] (Claim 2)

[0834] The system according to claim 1, comprising means for processing voice input in real time and individually customizing responses according to the user's emotional state.

[0835] (Claim 3)

[0836] The system according to claim 1, further comprising means for obtaining the user's response to a generated question again using voice input means, and for generating questions continuously based on that response. [Explanation of symbols]

[0837] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Voice input method, A means of converting audio into text data, A means of analyzing user sentiment based on text data, A means for dynamically generating questions based on the results of user sentiment analysis, A means of formatting the generated questions into natural language, A means of presenting formatted questions to the user, A system that includes this.

2. The system according to claim 1, comprising means for processing voice input in real time.

3. The system according to claim 1, characterized in that the user's response to the generated question is again obtained by voice input means.