system

A voice-activated system generates and distributes questionnaires to wireless devices, enabling efficient and rapid opinion collection from family members, addressing the challenge of modern communication difficulties by allowing quick decision-making and personalized surveys.

JP2026103441APending Publication Date: 2026-06-24SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-12
Publication Date
2026-06-24

AI Technical Summary

Technical Problem

Modern communication among family members is often difficult due to differences in schedules and lifestyles, and existing methods lack immediacy and convenience, especially when quick decisions are needed, requiring a system to efficiently aggregate opinions without direct interaction.

Method used

A system that recognizes voice instructions, generates questionnaires, distributes them to wireless devices, and collects responses, supporting multiple-choice and comment input, allowing users to gather opinions quickly and make decisions based on the results.

Benefits of technology

Enables efficient and rapid collection of opinions from family members, facilitating quick decision-making without physical contact or manual input, and supports personalized surveys based on emotional analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026103441000001_ABST
    Figure 2026103441000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of receiving voice instructions from a user and converting them into text using speech recognition technology, A means for generating an information gathering task based on the converted text and distributing it to an information sharing device, A means of collecting input from an information sharing device and analyzing the data through statistical processing, Means for providing analysis results to users, An information processing system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern times, communication among family members often becomes difficult due to differences in individual schedules and lifestyles. In particular, when it is necessary to collect opinions quickly for making important decisions, there is a need for a means to efficiently aggregate opinions while saving direct interaction and effort. However, many existing means lack immediacy and convenience, and are not practical especially when on the move or performing other tasks. Therefore, there is a need for a system that allows users to easily and quickly collect opinions from family members and share the response results.

Means for Solving the Problems

[0005] This invention provides a system that recognizes a user's voice instructions and generates a questionnaire based on those instructions. This system can distribute the generated questionnaire to multiple wireless devices and collect responses from each device. The results are then quickly notified to the user. Furthermore, the system supports cases where voice instructions are given via earphones, and the questionnaire format supports not only multiple-choice options but also comment input. This allows the user to instantly gather opinions from family members and make decisions based on the results.

[0006] "User" refers to any person who operates this system and sends and receives questionnaires.

[0007] "Voice commands" refer to operational instructions given by the user through earphones.

[0008] "Recognition" refers to the process of converting voice commands into digital data and understanding its content.

[0009] A "survey" refers to a digital questionnaire that includes specific questions and multiple answer choices or comment fields.

[0010] "Wireless device" refers to a handheld device, especially wireless communication-enabled devices such as smartphones and tablets.

[0011] "Distribution" refers to the process by which a server sends a survey to a wireless device.

[0012] "Data aggregation" refers to the statistical processing of responses transmitted from wireless devices and the summarization of the results.

[0013] "Notification" refers to the process by which the server sends a message to the user informing them of the aggregated results.

[0014] "Earphones" refers to audio devices worn in the ear by users, primarily for receiving voice commands. [Brief explanation of the drawing]

[0015] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

MODE FOR CARRYING OUT THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes. <00001​In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention relates to a system that starts operating when a user issues a voice command via earphones. The user provides the voice command to the AI ​​agent using the earphone microphone. This voice command is sent directly to the server. Upon receiving the command, the server converts it into text using speech recognition technology and generates a questionnaire based on it.

[0037] The generated questionnaires are stored on the server in a state where they can be retrieved via random access. The server then distributes the questionnaire information to each device of family members and related parties. When a device (smartphone or tablet) receives the questionnaire, it displays it on the screen and prompts the user to answer.

[0038] When a user enters their response on their device, the response is sent to the server. The server statistically processes each collected response and compiles the results. The compiled results are then reviewed by the user again and notified to the original user who requested the survey.

[0039] For example, consider a scenario where a user conducts a survey asking "What would you like to eat for dinner tonight?" The user gives a voice command, the server transcribes it into text, creates the survey, and distributes it. Each family member answers the survey on their own device, and the information is collected by the server. The server then compiles the results, such as 3 votes for curry and 2 votes for pasta, and notifies the user of the results.

[0040] This system enables the collection and rapid use of opinions from family and friends without physical contact or manual input. Thus, the goal of this invention is to facilitate efficient information gathering and communication.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The user gives voice commands to the AI ​​agent through earphones. During this process, they verbally specify the content of the survey.

[0044] Step 2:

[0045] The server receives voice commands and converts them into text data using speech recognition technology. Based on this text data, it generates the survey questions.

[0046] Step 3:

[0047] The server saves the generated questionnaire to a database and prepares the questionnaire data, including the necessary information (questions and answer choices).

[0048] Step 4:

[0049] The server distributes the questionnaire to each device of the target family member or related party. At this time, the server identifies the recipient based on the device information.

[0050] Step 5:

[0051] The terminal receives survey data distributed from the server and displays it on the screen in a survey format. During this process, notifications are also sent to the user.

[0052] Step 6:

[0053] Users respond by operating the survey screen on their device, selecting options, and entering comments.

[0054] Step 7:

[0055] The terminal sends the user's response to the server. The transmitted data is then aggregated again on the server.

[0056] Step 8:

[0057] The server aggregates all received responses and generates statistical data. Based on these aggregated results, it forms an overview of the responses.

[0058] Step 9:

[0059] The server then notifies the user again of the compiled results. The user checks the notification on their smartphone and views the survey results.

[0060] (Example 1)

[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0062] In today's information society, it is crucial to efficiently collect information and communicate regardless of time or place. In particular, quickly gathering opinions from family or groups is not easy and is a challenge many people face. Traditional methods require manual tabulation and physical contact, which is time-consuming. Furthermore, the difficulty in sharing information across multiple devices limits their use in situations requiring rapid decision-making.

[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0064] In this invention, the server includes means for acquiring voice instructions from a user and converting them into digital information, means for generating a survey based on the converted digital information, and means for transmitting the generated survey to multiple communication devices. This enables users to quickly and efficiently generate questionnaires, distribute them to multiple locations, and aggregate the results using only voice.

[0065] A "user" is a person or group that issues voice commands to the system and generates and receives the results of the survey.

[0066] "Voice instructions" refer to the form of instructions or requests that users give to a system through hearing devices.

[0067] "Digital information" refers to electronic text data converted based on voice instructions, and forms the basis for generating questionnaires.

[0068] "Survey" refers to a form of questionnaire or inquiry generated based on the user's voice instructions.

[0069] "Communication equipment" refers to devices that receive generated surveys and allow users to input their responses, such as smartphones and tablets.

[0070] "Response" refers to the data provided by users via communication devices in response to a survey.

[0071] A "server" is a central computing system that performs a series of processes, from receiving voice commands to generating survey results, organizing responses, and communicating results.

[0072] This invention provides a system that allows users to easily generate questionnaires and streamline information gathering. Users use an earphone microphone (a type of hearing device) to issue voice commands to a server. The server uses speech recognition technology to convert the received voice commands into digital information. A general-purpose speech recognition API can be used as the speech recognition software for this process.

[0073] The server generates a survey based on the converted digital information and sends it to a communication device. This device, such as a smartphone or tablet, receives the survey, displays it to the user, and allows them to enter a response. After the user enters their response via the communication device, these responses are sent back to the server for further processing.

[0074] As a concrete example, we can consider a scenario where a user gives a voice command on the topic of "What do you want to eat for dinner tonight?". The server converts the voice into text data, "What do you want to eat for dinner tonight?", and generates a survey with options such as curry and pasta. This survey is distributed to the family's communication devices, and each member responds individually. The server then compiles these results and notifies the user of the number of votes for each option.

[0075] An example of a prompt statement generated using a generative AI model is as follows:

[0076] Prompt example:

[0077] "The user has initiated a voice survey about dinner. Please generate the survey based on the voice instructions and distribute it to the relevant parties' communication devices."

[0078] This system enables efficient investigations and supports rapid decision-making through voice commands alone.

[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0080] Step 1:

[0081] The user provides voice instructions using an audio device. The user's input is specific, such as "What do you want to eat for dinner tonight?" This voice data is captured as an acoustic signal. This includes the user pressing the record button using the microphone function of their earphones and transmitting the voice instruction.

[0082] Step 2:

[0083] The server converts received audio data into text using speech recognition technology. The input is an audio signal, and the output is character data. The server applies a speech recognition algorithm, analyzes the audio waveform, and converts it into a string of characters. A speech recognition API is used in this process.

[0084] Step 3:

[0085] The server generates a survey based on the converted character data. The input is text data, and the output is data in the survey format. The server calculates answer choices to match the questions inferred from the text content and organizes the survey format.

[0086] Step 4:

[0087] The server transmits the generated survey data to communication devices. The input is data in the survey format, and the output is data distributed to each communication device. The server uses network functionality to broadcast the data to multiple terminals.

[0088] Step 5:

[0089] The terminal displays the received survey to the user and allows them to enter a response. The input is the survey data, and the output is the user's response. The terminal displays the survey content on the screen and enables the user to tap options or enter text.

[0090] Step 6:

[0091] Users input their responses to a survey via a terminal. Input consists of user selections from multiple-choice options or free text, while output is response data. Users select options on the terminal screen, confirm their input, and press the response button.

[0092] Step 7:

[0093] The terminal sends the collected responses to the server. The input is the response data entered on the terminal, and the output is the data sent to the server. The terminal aggregates the responses via the internet connection and sends them directly to the server.

[0094] Step 8:

[0095] The server organizes the received responses and generates aggregated results. The input is the response data from each user, and the output is the aggregated data. The server performs statistical processing, calculates results such as the number of votes for each option, and stores them in the database.

[0096] Step 9:

[0097] The server communicates the compiled and summarized results to the user. The input is the aggregated data, and the output is notification data for the user. The server delivers the results to the user via push notifications or email.

[0098] (Application Example 1)

[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0100] In modern communication, efficient information sharing is essential. However, traditional technologies require users to manually collect and share information, which is time-consuming and cumbersome. This is particularly problematic in family decision-making, where smoothly gathering and reflecting opinions is difficult.

[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0102] In this invention, the server includes means for receiving voice instructions from a user and converting them into text using speech recognition technology; means for generating information gathering tasks based on the converted text and distributing them to an information sharing device; and means for collecting input from the information sharing device and analyzing the data through statistical processing. This enables users to efficiently collect and quickly share information using only voice instructions.

[0103] A "user" is an entity that uses the system, and is a person or group that initiates information gathering tasks through voice commands.

[0104] "Voice commands" are instructions issued by users through audio equipment and are signals that are converted into text using speech recognition technology.

[0105] "Speech recognition technology" is a technology that converts speech signals into text data, and it is an important process for understanding and processing user instructions.

[0106] An "information gathering task" is an activity generated by the system based on transcribed voice instructions, and includes components such as questionnaires and questions to collect specific information.

[0107] An "information sharing device" is a digital device that can receive collected information and allows for the input of responses.

[0108] "Statistical processing" is a method of analyzing collected data to derive certain trends and results, and it is the process of providing aggregated results to users.

[0109] "Audio equipment" refers to devices used when voice instructions are given, such as earphones and speakers.

[0110] The server's role in this system is to receive voice commands transmitted by users via audio equipment and convert that voice into text data using speech recognition technology. For this speech recognition, the `speech_recognition` software library is used as an example. The server then utilizes a generative AI model to create information gathering tasks based on the transcribed voice commands. These tasks are stored in a randomly accessible format and distributed to the information sharing device.

[0111] Information sharing devices (such as smartphones and tablets) display received information gathering tasks on their screens and accept input from users. This input information is then sent back to the server, which performs statistical processing on the input data. Specifically, it processes the data using numerical analysis libraries such as statistics and derives aggregated results. These results are ultimately notified to the user and used to facilitate rapid decision-making.

[0112] As a concrete example, the server receives a voice command from a user saying, "Please conduct a survey on movies we want to see this weekend," and generates a survey on related movies. Then, it collects the family's choices through an information sharing device, determines the most popular movie through statistical processing, and communicates the results to the family.

[0113] An example of a prompt might be: "Use a speech recognition API to convert speech instructions into text, create a questionnaire using a natural language generation model, and collect and summarize the responses."

[0114] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0115] Step 1:

[0116] The user inputs voice commands using audio equipment. The audio equipment captures the user's voice commands as audio signals and sends these signals to the server. The audio signals serve as input and provide the data necessary for subsequent processing.

[0117] Step 2:

[0118] When the server receives an audio signal, it uses speech recognition technology to convert the signal into text data. Here, a speech recognition software library (e.g., speech_recognition) is used to output the audio signal as text. This conversion prepares the text data for subsequent information processing.

[0119] Step 3:

[0120] The server uses the acquired text data and a generative AI model to automatically generate information gathering tasks. The text data is used as input, and survey content is output based on prompts (e.g., "Please conduct a survey about movies you'd like to see this weekend"). This process generates surveys tailored to the user's needs.

[0121] Step 4:

[0122] The generated information gathering tasks are distributed from the server to the information sharing device. The distributed information gathering tasks are displayed on the terminal screen and become ready to receive responses from the user. The information sharing device constructs a user interface based on the data received from the server.

[0123] Step 5:

[0124] Users enter their answers to a survey via an information sharing device. The terminal collects user input in real time and sends the data back to the server. The entered response data becomes the raw material used for statistical processing.

[0125] Step 6:

[0126] The server collects responses submitted from terminals and performs statistical processing. Here, numerical analysis libraries such as Python's `statistics` library are used to analyze multiple response data and calculate aggregated results. These results serve as information for the user's final decision-making.

[0127] Step 7:

[0128] The server notifies the user of the aggregated results obtained through statistical processing. These results represent the user's final information and can be used, for example, as a reference when choosing a movie with family members. This output supports the decision-making process and improves user satisfaction.

[0129] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0130] In this invention, the user gives voice instructions through earphones. The voice spoken by the user is transmitted to a server. In addition to normal voice instructions, this system incorporates an emotion engine to extract emotional information from the voice. This emotion engine analyzes the emotional state from the user's voice tone, speed, and emphasized words. As a result, for example, if the user is speaking cheerfully, the questionnaire is adjusted to present casual options, while if the user is nervous, it presents more cautious options.

[0131] The server performs sentiment analysis simultaneously with speech-to-text conversion, and generates a questionnaire based on this analysis. This questionnaire includes questions and answer choices, and is personalized according to the user's emotions. The generated questionnaire is then delivered wirelessly to the devices of family members and related users.

[0132] The device displays the survey received from the server on its screen and notifies the user. The device displays questions and options that match the recipient's emotions, making the respondent feel more comfortable than usual. Once the user answers the survey on the device, the data is sent back to the server, and all responses are tallied.

[0133] The server statistically processes the aggregated results and summarizes them. The aggregated results are notified to the user, allowing them to easily review the information. For example, if a user asks, "What type of leisure activity would you like to enjoy on your next holiday?", the server detects positive emotions and provides a variety of cheerful leisure options.

[0134] This system enables a more personalized survey process that takes user emotions into account, resulting in more appropriate and contextually relevant information gathering.

[0135] The following describes the processing flow.

[0136] Step 1:

[0137] The user gives voice commands to the AI ​​agent through earphones. These commands can include the survey topic and questions they want to ask.

[0138] Step 2:

[0139] The server receives voice commands from the user and converts them into text data via a speech recognition system. Simultaneously, an emotion engine analyzes emotional information from the user's voice.

[0140] Step 3:

[0141] The server generates a survey tailored to the user's purpose based on recognized text data and sentiment information. It customizes the tone of questions and the content of answer choices based on sentiment information.

[0142] Step 4:

[0143] The server distributes the generated questionnaire to the devices of the target family members and related parties using wireless communication. This distribution takes place in real time.

[0144] Step 5:

[0145] The terminal receives the survey delivered from the server, displays it on the screen, and simultaneously notifies the user of the survey's arrival. The notification includes the survey's subject and response deadline.

[0146] Step 6:

[0147] Users respond to a survey displayed on their device. They can select options and enter comments as needed.

[0148] Step 7:

[0149] The terminal sends the user's entered answers to the server. The response data is compiled on the server, and statistical information for each item is calculated.

[0150] Step 8:

[0151] The server analyzes the aggregated results, compiles the overall findings, and then notifies the original user. In this process, the presentation of the results may be adjusted based on the user's emotional state.

[0152] Step 9:

[0153] Users receive notifications and can review the aggregated results. These results can then be used as material for future decision-making and discussions.

[0154] (Example 2)

[0155] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0156] Conventional voice-based survey systems lacked personalization that took into account the user's emotional state, resulting in difficulties in collecting information relevant to the user's context. Therefore, there is a need for a system that is more user-friendly and allows users to answer surveys more appropriately.

[0157] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0158] In this invention, the server includes means for recognizing the user's voice instructions, means for analyzing the emotional state from the recognized voice instructions, and means for generating a questionnaire based on the analysis results. This makes it possible to generate a personalized questionnaire that responds to the user's emotions.

[0159] A "user" is an individual or group that uses the system to give voice instructions and receives the results.

[0160] "Means for recognizing voice instructions" refers to technology or devices that capture a user's speech as voice data and convert it into text in order to understand its content.

[0161] "Means for analyzing emotional state" refers to technologies or devices that infer and analyze a user's emotions from the content and tone of voice instructions, speaking speed, emphasized words, etc.

[0162] "Means for generating questionnaires" refers to technologies or devices that dynamically create questions and answer choices in response to analyzed emotional states and present them to users.

[0163] A "communication device" is an electronic device that communicates with a server to send and receive data and notifies users of information.

[0164] "Means for aggregating responses" refers to a technology or device that statistically compiles and analyzes the responses to questionnaires submitted by users.

[0165] This system starts operating when the user gives voice commands through earphones. The earphones connect to a device (such as a smartphone or tablet) via Bluetooth and transmit the acquired voice data to a server.

[0166] The server processes the audio data using speech recognition software and an emotion analysis engine. Specifically, it converts speech to text using Google® Cloud Speech-to-Text, and then extracts the user's emotions from the text using emotion analysis tools such as IBM Watson® Tone Analyzer.

[0167] The interpreted sentiment information is used by a generative AI model to generate questionnaires. This generative AI model uses an open-source natural language processing library, and for example, if a user expresses a positive emotion, it dynamically creates a questionnaire that includes casual and diverse questions tailored to that emotion.

[0168] The generated questionnaire is delivered from the server to the relevant terminal via wireless communication. The terminal immediately displays the received questionnaire on its screen and notifies the user with audio and pop-up notifications. The user then answers the questionnaire, and the response is resent from the terminal to the server. This transmitted data is statistically analyzed through aggregation, and the results are fed back to the user from the server.

[0169] As a concrete example, consider conducting a survey about holiday plans. An example of a prompt question might be, "What type of leisure activity would you like to enjoy on your next holiday?" If positive emotions are detected, the survey will then provide a variety of bright leisure options, such as nature experiences, sports, and cultural activities. This allows users to answer the question in a more natural way.

[0170] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0171] Step 1:

[0172] The user gives voice commands through earphones. The voice is picked up by the earphone's microphone and transmitted as digital data to the device via Bluetooth. At this stage, the input is the user's voice, and the output is digital voice data.

[0173] Step 2:

[0174] The device transmits digital audio data to the server using Wi-Fi or mobile communication. The input here is the digital audio data acquired by the device, and the output is the data packets sent to the server via the network.

[0175] Step 3:

[0176] The server uses speech recognition software to convert audio data into text. This process analyzes the audio waveform and converts its content into character data. The input for this step is digital audio data, and the output is the converted text data.

[0177] Step 4:

[0178] The server uses an emotion analysis engine to extract emotional states from text data. Specifically, it analyzes the sentence structure, the presence or absence of exclamation marks, and specific words and phrases in the text to measure the user's emotions. In this step, the input is text data, and the output is the emotion analysis result.

[0179] Step 5:

[0180] The server uses a generative AI model to generate questionnaires based on sentiment analysis results. The AI ​​model, for example, utilizes open-source natural language processing libraries to design questions and answer choices that respond to emotions. The input to this process is the sentiment analysis results, and the output is questionnaire data containing questions and answer choices.

[0181] Step 6:

[0182] The server transmits survey data to the terminal via wireless communication. Wireless communication includes data encryption to ensure secure information transfer. The input is the survey data, and the output is the data packets sent to the terminal.

[0183] Step 7:

[0184] The terminal displays the survey on its screen interface and notifies the user. Here, the questions and answer choices are displayed on the terminal's screen, and the user is notified with a notification sound and banner. The input is the received survey data, and the output is the displayed visual interface.

[0185] Step 8:

[0186] The user answers a questionnaire and inputs their responses into the device. They select options using a touch panel or voice input and confirm their answers on the device. The input is the user's selection, and the output is the response data.

[0187] Step 9:

[0188] The terminal sends the response data to the server. Communication is conducted using network standards, ensuring accurate data transmission. The input is the response data from the terminal, and the output is a data structure for aggregation that is sent to the server.

[0189] Step 10:

[0190] The server aggregates the received response data and performs statistical processing. It uses a database system to efficiently consolidate responses and generate analysis results. The input is a large amount of response data, and the output is an aggregated statistical report.

[0191] (Application Example 2)

[0192] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0193] Conventional information gathering systems using listening devices failed to consider the emotional state of users, making it difficult to provide responses and suggestions tailored to individual users. This resulted in insufficient improvements in convenience and user satisfaction, and potentially impacted the accuracy of the information.

[0194] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0195] In this invention, the server includes a device for receiving a user's voice signal, means for analyzing the characteristics of the received voice signal to identify emotional information, and means for generating information collection documents based on the identified emotional information. This enables personalized information collection that responds to the user's emotions.

[0196] "User voice signal" refers to the electrical signal obtained by converting the user's voice into an electrical signal.

[0197] A "receiving device" is a machine or system that takes in an audio signal and uses it for subsequent processing.

[0198] "Means for analyzing the characteristics of an audio signal to identify emotional information" refers to a process or technique that analyzes the tone, speed, and emphasized words of speech and evaluates the emotional state from the results.

[0199] "Means for generating information gathering documents based on emotional information" refers to the process of creating information gathering documents according to identified emotional states.

[0200] A "communication device" is a technology or device used to exchange information or data with other systems or devices.

[0201] "Means of aggregation" refers to techniques or methods for collecting multiple responses and processing them statistically.

[0202] "Means of providing aggregated results to users" refers to technologies or methods for presenting processed aggregated data to users in an easily understandable manner.

[0203] The system for implementing this invention consists of a consumer robot installed in the home and an external server.

[0204] The server receives voice signals that the user makes to the robot. These signals are captured and digitized through a microphone built into the robot. Next, the server uses an emotion analysis engine to analyze the received voice signals. For example, IBM Watson or Google Cloud's voice analysis software may be used. Through this analysis, the server identifies the user's emotional state based on the tone, speed, and emphasized words of the voice.

[0205] Next, the server generates a personalized information gathering document based on the identified emotional information. For example, if the user indicates a cheerful mood, the document will include suggestions for relaxing activities. This information gathering document is sent to the robot and guided to the user via voice. Software such as Google Text-to-Speech is used for speech synthesis.

[0206] This robot can play relaxing music or suggest simple games through conversations with users in the home. For example, if a user says, "I'm a little tired today," the robot can play soothing music and offer a warm drink.

[0207] The following is an example of a prompt message:

[0208] User's voice: "I'm a little tired today."

[0209] Based on the results of the emotion recognition engine, please recognize the emotion of fatigue and suggest the following reactions:

[0210] "Thank you for your hard work. I'll play some relaxing music."

[0211] "Shall I make you some warm tea?"

[0212] In this way, this system allows users to receive personalized support tailored to their emotional state, making their daily lives more comfortable.

[0213] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0214] Step 1:

[0215] The user speaks to the robot. A microphone built into the robot converts this audio signal from analog to digital and sends it to the server. The input is the user's voice, and the output is digitized audio data. The server receives this data.

[0216] Step 2:

[0217] The server passes the received audio data to the emotion analysis engine. This process analyzes factors such as voice tone, speed, and emphasized words. The input is digitized audio data, and the output is information indicating the user's emotional state. Specifically, the server uses speech analysis software to identify the user's emotions.

[0218] Step 3:

[0219] Based on the identified sentiment information, the server generates personalized information gathering documents. The input is the analyzed sentiment information, and the output is an information gathering document tailored to that sentiment. For example, if a cheerful sentiment is detected, activity suggestions will be included.

[0220] Step 4:

[0221] The generated information gathering documents are sent from the server to the robot. The input is an information gathering document based on emotions, and the output is data that is transferred to the robot. The robot receives this data.

[0222] Step 5:

[0223] The robot uses the received data to provide voice guidance to the user. It uses speech synthesis software to generate reactions and interact with the user. The input is information collected from a server, and the output is voice guidance to the user. For example, if the user is tired, the robot might suggest playing relaxing music.

[0224] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0225] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0226] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0227] [Second Embodiment]

[0228] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0229] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0230] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0231] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0232] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0233] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0234] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0235] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0236] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0237] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0238] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0239] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0240] This invention relates to a system that starts operating when a user issues a voice command via earphones. The user provides the voice command to the AI ​​agent using the earphone microphone. This voice command is sent directly to the server. Upon receiving the command, the server converts it into text using speech recognition technology and generates a questionnaire based on it.

[0241] The generated questionnaires are stored on the server in a state where they can be retrieved via random access. The server then distributes the questionnaire information to each device of family members and related parties. When a device (smartphone or tablet) receives the questionnaire, it displays it on the screen and prompts the user to answer.

[0242] When a user enters their response on their device, the response is sent to the server. The server statistically processes each collected response and compiles the results. The compiled results are then reviewed by the user again and notified to the original user who requested the survey.

[0243] For example, consider a scenario where a user conducts a survey asking "What would you like to eat for dinner tonight?" The user gives a voice command, the server transcribes it into text, creates the survey, and distributes it. Each family member answers the survey on their own device, and the information is collected by the server. The server then compiles the results, such as 3 votes for curry and 2 votes for pasta, and notifies the user of the results.

[0244] This system enables the collection and rapid use of opinions from family and friends without physical contact or manual input. Thus, the goal of this invention is to facilitate efficient information gathering and communication.

[0245] The following describes the processing flow.

[0246] Step 1:

[0247] The user gives voice commands to the AI ​​agent through earphones. During this process, they verbally specify the content of the survey.

[0248] Step 2:

[0249] The server receives voice commands and converts them into text data using speech recognition technology. Based on this text data, it generates the survey questions.

[0250] Step 3:

[0251] The server saves the generated questionnaire to a database and prepares the questionnaire data, including the necessary information (questions and answer choices).

[0252] Step 4:

[0253] The server distributes the questionnaire to each device of the target family member or related party. At this time, the server identifies the recipient based on the device information.

[0254] Step 5:

[0255] The terminal receives survey data distributed from the server and displays it on the screen in a survey format. During this process, notifications are also sent to the user.

[0256] Step 6:

[0257] Users respond by operating the survey screen on their device, selecting options, and entering comments.

[0258] Step 7:

[0259] The terminal sends the user's response to the server. The transmitted data is then aggregated again on the server.

[0260] Step 8:

[0261] The server aggregates all received responses and generates statistical data. Based on these aggregated results, it forms an overview of the responses.

[0262] Step 9:

[0263] The server then notifies the user again of the compiled results. The user checks the notification on their smartphone and views the survey results.

[0264] (Example 1)

[0265] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0266] In today's information society, it is crucial to efficiently collect information and communicate regardless of time or place. In particular, quickly gathering opinions from family or groups is not easy and is a challenge many people face. Traditional methods require manual tabulation and physical contact, which is time-consuming. Furthermore, the difficulty in sharing information across multiple devices limits their use in situations requiring rapid decision-making.

[0267] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0268] In this invention, the server includes means for acquiring voice instructions from a user and converting them into digital information, means for generating a survey based on the converted digital information, and means for transmitting the generated survey to multiple communication devices. This enables users to quickly and efficiently generate questionnaires, distribute them to multiple locations, and aggregate the results using only voice.

[0269] A "user" is a person or group that issues voice commands to the system and generates and receives the results of the survey.

[0270] "Voice instructions" refer to the form of instructions or requests that users give to a system through hearing devices.

[0271] "Digital information" refers to electronic text data converted based on voice instructions, and forms the basis for generating questionnaires.

[0272] "Survey" refers to a form of questionnaire or inquiry generated based on the user's voice instructions.

[0273] "Communication equipment" refers to devices that receive generated surveys and allow users to input their responses, such as smartphones and tablets.

[0274] "Response" refers to the data provided by users via communication devices in response to a survey.

[0275] A "server" is a central computing system that performs a series of processes, from receiving voice commands to generating survey results, organizing responses, and communicating results.

[0276] This invention provides a system that allows users to easily generate questionnaires and streamline information gathering. Users use an earphone microphone (a type of hearing device) to issue voice commands to a server. The server uses speech recognition technology to convert the received voice commands into digital information. A general-purpose speech recognition API can be used as the speech recognition software for this process.

[0277] The server generates a survey based on the converted digital information and sends it to a communication device. This device, such as a smartphone or tablet, receives the survey, displays it to the user, and allows them to enter a response. After the user enters their response via the communication device, these responses are sent back to the server for further processing.

[0278] As a concrete example, we can consider a scenario where a user gives a voice command on the topic of "What do you want to eat for dinner tonight?". The server converts the voice into text data, "What do you want to eat for dinner tonight?", and generates a survey with options such as curry and pasta. This survey is distributed to the family's communication devices, and each member responds individually. The server then compiles these results and notifies the user of the number of votes for each option.

[0279] An example of a prompt statement generated using a generative AI model is as follows:

[0280] Prompt example:

[0281] "The user has initiated a voice survey about dinner. Please generate the survey based on the voice instructions and distribute it to the relevant parties' communication devices."

[0282] This system realizes a function that can streamline the investigation with only voice instructions and support rapid decision-making.

[0283] The flow of the specific process in Example 1 will be described with reference to FIG. 11.

[0284] Step 1:

[0285] The user gives voice instructions using an auditory device. The user's input is a specific voice such as "What would I like to eat for dinner today?" This voice data is acquired as an acoustic signal. This includes the operation where the user presses the recording button using the microphone function of the earphone and transmits a voice instruction.

[0286] Step 2:

[0287] The server converts the received voice data into text using voice recognition technology. The input is an acoustic signal, and the output is character data. The server applies a voice recognition algorithm, analyzes the voice waveform, and performs an operation of converting it into a character string. At this time, a voice recognition API is utilized.

[0288] Step 3:

[0289] The server generates an investigation based on the converted character data. The input is text data, and the output is data in the investigation format. The server calculates options according to the question content inferred from the text content and performs an operation of organizing the form of the investigation. ​​​​​​​​​​​​​​The terminal displays the received survey to the user and allows them to enter a response. The input is the survey data, and the output is the user's response. The terminal displays the survey content on the screen and enables the user to tap options or enter text.

[0294] Step 6:

[0295] Users input their responses to a survey via a terminal. Input consists of user selections from multiple-choice options or free text, while output is response data. Users select options on the terminal screen, confirm their input, and press the response button.

[0296] Step 7:

[0297] The terminal sends the collected responses to the server. The input is the response data entered on the terminal, and the output is the data sent to the server. The terminal aggregates the responses via the internet connection and sends them directly to the server.

[0298] Step 8:

[0299] The server organizes the received responses and generates aggregated results. The input is the response data from each user, and the output is the aggregated data. The server performs statistical processing, calculates results such as the number of votes for each option, and stores them in the database.

[0300] Step 9:

[0301] The server communicates the compiled and summarized results to the user. The input is the aggregated data, and the output is notification data for the user. The server delivers the results to the user via push notifications or email.

[0302] (Application Example 1)

[0303] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0304] In modern communication, it is required to efficiently perform information sharing. However, in the conventional technology, users need to manually collect and share information, which has the problem of taking time and effort. In particular, in the decision-making within a family, there is a problem that it is difficult to smoothly aggregate and reflect opinions.

[0305] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0306] In this invention, the server includes means for receiving a user's voice instruction, converting it into text using voice recognition technology, generating an information collection task based on the converted text, and distributing it to an information sharing device, and means for collecting inputs at the information sharing device and analyzing data through statistical processing. Thereby, the user can efficiently collect information and quickly share it only with a voice instruction.

[0307] A "user" is a subject that uses the system, such as a person or a group that starts an information collection task through a voice instruction.

[0308] A "voice instruction" is a command issued by a user through an acoustic device and is a signal that is converted into text by voice recognition technology.

[0309] "Voice recognition technology" is a technology that converts a voice signal into text data and is an important process for understanding and processing a user's instruction.

[0310] An "information collection task" is an activity generated by the system based on a texturized voice instruction and is a component including a questionnaire or a question format for collecting specific information.

[0311] An "information sharing device" is a terminal that can receive the collected information and is a digital device that enables input of answers.

[0312] "Statistical processing" is a method of analyzing collected data to derive certain trends and results, and it is the process of providing aggregated results to users.

[0313] "Audio equipment" refers to devices used when voice instructions are given, such as earphones and speakers.

[0314] The server's role in this system is to receive voice commands transmitted by users via audio equipment and convert that voice into text data using speech recognition technology. For this speech recognition, the `speech_recognition` software library is used as an example. The server then utilizes a generative AI model to create information gathering tasks based on the transcribed voice commands. These tasks are stored in a randomly accessible format and distributed to the information sharing device.

[0315] Information sharing devices (such as smartphones and tablets) display received information gathering tasks on their screens and accept input from users. This input information is then sent back to the server, which performs statistical processing on the input data. Specifically, it processes the data using numerical analysis libraries such as statistics and derives aggregated results. These results are ultimately notified to the user and used to facilitate rapid decision-making.

[0316] As a concrete example, the server receives a voice command from a user saying, "Please conduct a survey on movies we want to see this weekend," and generates a survey on related movies. Then, it collects the family's choices through an information sharing device, determines the most popular movie through statistical processing, and communicates the results to the family.

[0317] An example of a prompt might be: "Use a speech recognition API to convert speech instructions into text, create a questionnaire using a natural language generation model, and collect and summarize the responses."

[0318] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0319] Step 1:

[0320] The user inputs voice commands using audio equipment. The audio equipment captures the user's voice commands as audio signals and sends these signals to the server. The audio signals serve as input and provide the data necessary for subsequent processing.

[0321] Step 2:

[0322] When the server receives an audio signal, it uses speech recognition technology to convert the signal into text data. Here, a speech recognition software library (e.g., speech_recognition) is used to output the audio signal as text. This conversion prepares the text data for subsequent information processing.

[0323] Step 3:

[0324] The server uses the acquired text data and a generative AI model to automatically generate information gathering tasks. The text data is used as input, and survey content is output based on prompts (e.g., "Please conduct a survey about movies you'd like to see this weekend"). This process generates surveys tailored to the user's needs.

[0325] Step 4:

[0326] The generated information gathering tasks are distributed from the server to the information sharing device. The distributed information gathering tasks are displayed on the terminal screen and become ready to receive responses from the user. The information sharing device constructs a user interface based on the data received from the server.

[0327] Step 5:

[0328] Users enter their answers to a survey via an information sharing device. The terminal collects user input in real time and sends the data back to the server. The entered response data becomes the raw material used for statistical processing.

[0329] Step 6:

[0330] The server collects responses submitted from terminals and performs statistical processing. Here, numerical analysis libraries such as Python's `statistics` library are used to analyze multiple response data and calculate aggregated results. These results serve as information for the user's final decision-making.

[0331] Step 7:

[0332] The server notifies the user of the aggregated results obtained through statistical processing. These results represent the user's final information and can be used, for example, as a reference when choosing a movie with family members. This output supports the decision-making process and improves user satisfaction.

[0333] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0334] In this invention, the user gives voice instructions through earphones. The voice spoken by the user is transmitted to a server. In addition to normal voice instructions, this system incorporates an emotion engine to extract emotional information from the voice. This emotion engine analyzes the emotional state from the user's voice tone, speed, and emphasized words. As a result, for example, if the user is speaking cheerfully, the questionnaire is adjusted to present casual options, while if the user is nervous, it presents more cautious options.

[0335] The server performs sentiment analysis simultaneously with speech-to-text conversion, and generates a questionnaire based on this analysis. This questionnaire includes questions and answer choices, and is personalized according to the user's emotions. The generated questionnaire is then delivered wirelessly to the devices of family members and related users.

[0336] The device displays the survey received from the server on its screen and notifies the user. The device displays questions and options that match the recipient's emotions, making the respondent feel more comfortable than usual. Once the user answers the survey on the device, the data is sent back to the server, and all responses are tallied.

[0337] The server statistically processes the aggregated results and summarizes them. The aggregated results are notified to the user, allowing them to easily review the information. For example, if a user asks, "What type of leisure activity would you like to enjoy on your next holiday?", the server detects positive emotions and provides a variety of cheerful leisure options.

[0338] This system enables a more personalized survey process that takes user emotions into account, resulting in more appropriate and contextually relevant information gathering.

[0339] The following describes the processing flow.

[0340] Step 1:

[0341] The user gives voice commands to the AI ​​agent through earphones. These commands can include the survey topic and questions they want to ask.

[0342] Step 2:

[0343] The server receives voice commands from the user and converts them into text data via a speech recognition system. Simultaneously, an emotion engine analyzes emotional information from the user's voice.

[0344] Step 3:

[0345] The server generates a survey tailored to the user's purpose based on recognized text data and sentiment information. It customizes the tone of questions and the content of answer choices based on sentiment information.

[0346] Step 4:

[0347] The server distributes the generated questionnaire to the devices of the target family members and related parties using wireless communication. This distribution takes place in real time.

[0348] Step 5:

[0349] The terminal receives the survey delivered from the server, displays it on the screen, and simultaneously notifies the user of the survey's arrival. The notification includes the survey's subject and response deadline.

[0350] Step 6:

[0351] Users respond to a survey displayed on their device. They can select options and enter comments as needed.

[0352] Step 7:

[0353] The terminal sends the user's entered answers to the server. The response data is compiled on the server, and statistical information for each item is calculated.

[0354] Step 8:

[0355] The server analyzes the aggregated results, compiles the overall findings, and then notifies the original user. In this process, the presentation of the results may be adjusted based on the user's emotional state.

[0356] Step 9:

[0357] Users receive notifications and can review the aggregated results. These results can then be used as material for future decision-making and discussions.

[0358] (Example 2)

[0359] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0360] Conventional voice-based survey systems lacked personalization that took into account the user's emotional state, resulting in difficulties in collecting information relevant to the user's context. Therefore, there is a need for a system that is more user-friendly and allows users to answer surveys more appropriately.

[0361] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0362] In this invention, the server includes means for recognizing the user's voice instructions, means for analyzing the emotional state from the recognized voice instructions, and means for generating a questionnaire based on the analysis results. This makes it possible to generate a personalized questionnaire that responds to the user's emotions.

[0363] A "user" is an individual or group that uses the system to give voice instructions and receives the results.

[0364] "Means for recognizing voice instructions" refers to technology or devices that capture a user's speech as voice data and convert it into text in order to understand its content.

[0365] "Means for analyzing emotional state" refers to technologies or devices that infer and analyze a user's emotions from the content and tone of voice instructions, speaking speed, emphasized words, etc.

[0366] "Means for generating questionnaires" refers to technologies or devices that dynamically create questions and answer choices in response to analyzed emotional states and present them to users.

[0367] A "communication device" is an electronic device that communicates with a server to send and receive data and notifies users of information.

[0368] "Means for aggregating responses" refers to a technology or device that statistically compiles and analyzes the responses to questionnaires submitted by users.

[0369] This system starts operating when the user gives voice commands through earphones. The earphones connect to a device (such as a smartphone or tablet) via Bluetooth and transmit the acquired voice data to a server.

[0370] The server processes the audio data using speech recognition software and an emotion analysis engine. Specifically, it converts speech to text using Google Cloud Speech-to-Text, and then extracts the user's emotions from the text using emotion analysis tools such as IBM Watson Tone Analyzer.

[0371] The interpreted sentiment information is used by a generative AI model to generate questionnaires. This generative AI model uses an open-source natural language processing library, and for example, if a user expresses a positive emotion, it dynamically creates a questionnaire that includes casual and diverse questions tailored to that emotion.

[0372] The generated questionnaire is delivered from the server to the relevant terminal via wireless communication. The terminal immediately displays the received questionnaire on its screen and notifies the user with audio and pop-up notifications. The user then answers the questionnaire, and the response is resent from the terminal to the server. This transmitted data is statistically analyzed through aggregation, and the results are fed back to the user from the server.

[0373] As a concrete example, consider conducting a survey about holiday plans. An example of a prompt question might be, "What type of leisure activity would you like to enjoy on your next holiday?" If positive emotions are detected, the survey will then provide a variety of bright leisure options, such as nature experiences, sports, and cultural activities. This allows users to answer the question in a more natural way.

[0374] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0375] Step 1:

[0376] The user gives voice commands through earphones. The voice is picked up by the earphone's microphone and transmitted as digital data to the device via Bluetooth. At this stage, the input is the user's voice, and the output is digital voice data.

[0377] Step 2:

[0378] The device transmits digital audio data to the server using Wi-Fi or mobile communication. The input here is the digital audio data acquired by the device, and the output is the data packets sent to the server via the network.

[0379] Step 3:

[0380] The server uses speech recognition software to convert audio data into text. This process analyzes the audio waveform and converts its content into character data. The input for this step is digital audio data, and the output is the converted text data.

[0381] Step 4:

[0382] The server uses an emotion analysis engine to extract emotional states from text data. Specifically, it analyzes the sentence structure, the presence or absence of exclamation marks, and specific words and phrases in the text to measure the user's emotions. In this step, the input is text data, and the output is the emotion analysis result.

[0383] Step 5:

[0384] The server uses a generative AI model to generate questionnaires based on sentiment analysis results. The AI ​​model, for example, utilizes open-source natural language processing libraries to design questions and answer choices that respond to emotions. The input to this process is the sentiment analysis results, and the output is questionnaire data containing questions and answer choices.

[0385] Step 6:

[0386] The server transmits survey data to the terminal via wireless communication. Wireless communication includes data encryption to ensure secure information transfer. The input is the survey data, and the output is the data packets sent to the terminal.

[0387] Step 7:

[0388] The terminal displays the survey on its screen interface and notifies the user. Here, the questions and answer choices are displayed on the terminal's screen, and the user is notified with a notification sound and banner. The input is the received survey data, and the output is the displayed visual interface.

[0389] Step 8:

[0390] The user answers a questionnaire and inputs their responses into the device. They select options using a touch panel or voice input and confirm their answers on the device. The input is the user's selection, and the output is the response data.

[0391] Step 9:

[0392] The terminal sends the response data to the server. Communication is conducted using network standards, ensuring accurate data transmission. The input is the response data from the terminal, and the output is a data structure for aggregation that is sent to the server.

[0393] Step 10:

[0394] The server aggregates the received response data and performs statistical processing. It uses a database system to efficiently consolidate responses and generate analysis results. The input is a large amount of response data, and the output is an aggregated statistical report.

[0395] (Application Example 2)

[0396] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0397] Conventional information gathering systems using listening devices failed to consider the emotional state of users, making it difficult to provide responses and suggestions tailored to individual users. This resulted in insufficient improvements in convenience and user satisfaction, and potentially impacted the accuracy of the information.

[0398] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0399] In this invention, the server includes a device for receiving a user's voice signal, means for analyzing the characteristics of the received voice signal to identify emotional information, and means for generating information collection documents based on the identified emotional information. This enables personalized information collection that responds to the user's emotions.

[0400] "User voice signal" refers to the electrical signal obtained by converting the user's voice into an electrical signal.

[0401] A "receiving device" is a machine or system that takes in an audio signal and uses it for subsequent processing.

[0402] "Means for analyzing the characteristics of an audio signal to identify emotional information" refers to a process or technique that analyzes the tone, speed, and emphasized words of speech and evaluates the emotional state from the results.

[0403] "Means for generating information gathering documents based on emotional information" refers to the process of creating information gathering documents according to identified emotional states.

[0404] A "communication device" is a technology or device used to exchange information or data with other systems or devices.

[0405] "Means of aggregation" refers to techniques or methods for collecting multiple responses and processing them statistically.

[0406] "Means of providing aggregated results to users" refers to technologies or methods for presenting processed aggregated data to users in an easily understandable manner.

[0407] The system for implementing this invention consists of a consumer robot installed in the home and an external server.

[0408] The server receives voice signals that the user makes to the robot. These signals are captured and digitized through a microphone built into the robot. Next, the server uses an emotion analysis engine to analyze the received voice signals. For example, IBM Watson or Google Cloud's voice analysis software may be used. Through this analysis, the server identifies the user's emotional state based on the tone, speed, and emphasized words of the voice.

[0409] Next, the server generates a personalized information gathering document based on the identified emotional information. For example, if the user indicates a cheerful mood, the document will include suggestions for relaxing activities. This information gathering document is sent to the robot and guided to the user via voice. Software such as Google Text-to-Speech is used for speech synthesis.

[0410] This robot can play relaxing music or suggest simple games through conversations with users in the home. For example, if a user says, "I'm a little tired today," the robot can play soothing music and offer a warm drink.

[0411] The following is an example of a prompt message:

[0412] User's voice: "I'm a little tired today."

[0413] Based on the results of the emotion recognition engine, please recognize the emotion of fatigue and suggest the following reactions:

[0414] "Thank you for your hard work. I'll play some relaxing music."

[0415] "Shall I make you some warm tea?"

[0416] In this way, this system allows users to receive personalized support tailored to their emotional state, making their daily lives more comfortable.

[0417] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0418] Step 1:

[0419] The user speaks to the robot. A microphone built into the robot converts this audio signal from analog to digital and sends it to the server. The input is the user's voice, and the output is digitized audio data. The server receives this data.

[0420] Step 2:

[0421] The server passes the received audio data to the emotion analysis engine. This process analyzes factors such as voice tone, speed, and emphasized words. The input is digitized audio data, and the output is information indicating the user's emotional state. Specifically, the server uses speech analysis software to identify the user's emotions.

[0422] Step 3:

[0423] Based on the identified sentiment information, the server generates personalized information gathering documents. The input is the analyzed sentiment information, and the output is an information gathering document tailored to that sentiment. For example, if a cheerful sentiment is detected, activity suggestions will be included.

[0424] Step 4:

[0425] The generated information gathering documents are sent from the server to the robot. The input is an information gathering document based on emotions, and the output is data that is transferred to the robot. The robot receives this data.

[0426] Step 5:

[0427] The robot uses the received data to provide voice guidance to the user. It uses speech synthesis software to generate reactions and interact with the user. The input is information collected from a server, and the output is voice guidance to the user. For example, if the user is tired, the robot might suggest playing relaxing music.

[0428] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0429] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0430] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0431] [Third Embodiment]

[0432] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0433] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0434] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0435] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0436] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0437] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0438] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0439] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0440] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0441] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0442] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0443] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0444] This invention relates to a system that starts operating when a user issues a voice command via earphones. The user provides the voice command to the AI ​​agent using the earphone microphone. This voice command is sent directly to the server. Upon receiving the command, the server converts it into text using speech recognition technology and generates a questionnaire based on it.

[0445] The generated questionnaires are stored on the server in a state where they can be retrieved via random access. The server then distributes the questionnaire information to each device of family members and related parties. When a device (smartphone or tablet) receives the questionnaire, it displays it on the screen and prompts the user to answer.

[0446] When a user enters their response on their device, the response is sent to the server. The server statistically processes each collected response and compiles the results. The compiled results are then reviewed by the user again and notified to the original user who requested the survey.

[0447] For example, consider a scenario where a user conducts a survey asking "What would you like to eat for dinner tonight?" The user gives a voice command, the server transcribes it into text, creates the survey, and distributes it. Each family member answers the survey on their own device, and the information is collected by the server. The server then compiles the results, such as 3 votes for curry and 2 votes for pasta, and notifies the user of the results.

[0448] This system enables the collection and rapid use of opinions from family and friends without physical contact or manual input. Thus, the goal of this invention is to facilitate efficient information gathering and communication.

[0449] The following describes the processing flow.

[0450] Step 1:

[0451] The user gives voice commands to the AI ​​agent through earphones. During this process, they verbally specify the content of the survey.

[0452] Step 2:

[0453] The server receives voice commands and converts them into text data using speech recognition technology. Based on this text data, it generates the survey questions.

[0454] Step 3:

[0455] The server saves the generated questionnaire to a database and prepares the questionnaire data, including the necessary information (questions and answer choices).

[0456] Step 4:

[0457] The server distributes the questionnaire to each device of the target family member or related party. At this time, the server identifies the recipient based on the device information.

[0458] Step 5:

[0459] The terminal receives survey data distributed from the server and displays it on the screen in a survey format. During this process, notifications are also sent to the user.

[0460] Step 6:

[0461] Users respond by operating the survey screen on their device, selecting options, and entering comments.

[0462] Step 7:

[0463] The terminal sends the user's response to the server. The transmitted data is then aggregated again on the server.

[0464] Step 8:

[0465] The server aggregates all received responses and generates statistical data. Based on these aggregated results, it forms an overview of the responses.

[0466] Step 9:

[0467] The server then notifies the user again of the compiled results. The user checks the notification on their smartphone and views the survey results.

[0468] (Example 1)

[0469] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0470] In today's information society, it is crucial to efficiently collect information and communicate regardless of time or place. In particular, quickly gathering opinions from family or groups is not easy and is a challenge many people face. Traditional methods require manual tabulation and physical contact, which is time-consuming. Furthermore, the difficulty in sharing information across multiple devices limits their use in situations requiring rapid decision-making.

[0471] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0472] In this invention, the server includes means for acquiring voice instructions from a user and converting them into digital information, means for generating a survey based on the converted digital information, and means for transmitting the generated survey to multiple communication devices. This enables users to quickly and efficiently generate questionnaires, distribute them to multiple locations, and aggregate the results using only voice.

[0473] A "user" is a person or group that issues voice commands to the system and generates and receives the results of the survey.

[0474] "Voice instructions" refer to the form of instructions or requests that users give to a system through hearing devices.

[0475] "Digital information" refers to electronic text data converted based on voice instructions, and forms the basis for generating questionnaires.

[0476] "Survey" refers to a form of questionnaire or inquiry generated based on the user's voice instructions.

[0477] "Communication equipment" refers to devices that receive generated surveys and allow users to input their responses, such as smartphones and tablets.

[0478] "Response" refers to the data provided by users via communication devices in response to a survey.

[0479] A "server" is a central computing system that performs a series of processes, from receiving voice commands to generating survey results, organizing responses, and communicating results.

[0480] This invention provides a system that allows users to easily generate questionnaires and streamline information gathering. Users use an earphone microphone (a type of hearing device) to issue voice commands to a server. The server uses speech recognition technology to convert the received voice commands into digital information. A general-purpose speech recognition API can be used as the speech recognition software for this process.

[0481] The server generates a survey based on the converted digital information and sends it to a communication device. This device, such as a smartphone or tablet, receives the survey, displays it to the user, and allows them to enter a response. After the user enters their response via the communication device, these responses are sent back to the server for further processing.

[0482] As a concrete example, we can consider a scenario where a user gives a voice command on the topic of "What do you want to eat for dinner tonight?". The server converts the voice into text data, "What do you want to eat for dinner tonight?", and generates a survey with options such as curry and pasta. This survey is distributed to the family's communication devices, and each member responds individually. The server then compiles these results and notifies the user of the number of votes for each option.

[0483] An example of a prompt statement generated using a generative AI model is as follows:

[0484] Prompt example:

[0485] "The user has initiated a voice survey about dinner. Please generate the survey based on the voice instructions and distribute it to the relevant parties' communication devices."

[0486] This system enables efficient investigations and supports rapid decision-making through voice commands alone.

[0487] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0488] Step 1:

[0489] The user provides voice instructions using an audio device. The user's input is specific, such as "What do you want to eat for dinner tonight?" This voice data is captured as an acoustic signal. This includes the user pressing the record button using the microphone function of their earphones and transmitting the voice instruction.

[0490] Step 2:

[0491] The server converts received audio data into text using speech recognition technology. The input is an audio signal, and the output is character data. The server applies a speech recognition algorithm, analyzes the audio waveform, and converts it into a string of characters. A speech recognition API is used in this process.

[0492] Step 3:

[0493] The server generates a survey based on the converted character data. The input is text data, and the output is data in the survey format. The server calculates answer choices based on the questions inferred from the text content and organizes the survey format.

[0494] Step 4:

[0495] The server transmits the generated survey data to communication devices. The input is data in the survey format, and the output is data distributed to each communication device. The server uses network functionality to broadcast the data to multiple terminals.

[0496] Step 5:

[0497] The terminal displays the received survey to the user and allows them to enter a response. The input is the survey data, and the output is the user's response. The terminal displays the survey content on the screen and enables the user to tap options or enter text.

[0498] Step 6:

[0499] Users input their responses to a survey via a terminal. Input consists of user selections from multiple-choice options or free text, while output is response data. Users select options on the terminal screen, confirm their input, and press the response button.

[0500] Step 7:

[0501] The terminal sends the collected responses to the server. The input is the response data entered on the terminal, and the output is the data sent to the server. The terminal aggregates the responses via the internet connection and sends them directly to the server.

[0502] Step 8:

[0503] The server organizes the received responses and generates aggregated results. The input is the response data from each user, and the output is the aggregated data. The server performs statistical processing, calculates results such as the number of votes for each option, and stores them in the database.

[0504] Step 9:

[0505] The server communicates the compiled and summarized results to the user. The input is the aggregated data, and the output is notification data for the user. The server delivers the results to the user via push notifications or email.

[0506] (Application Example 1)

[0507] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0508] In modern communication, efficient information sharing is essential. However, traditional technologies require users to manually collect and share information, which is time-consuming and cumbersome. This is particularly problematic in family decision-making, where smoothly gathering and reflecting opinions is difficult.

[0509] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0510] In this invention, the server includes means for receiving voice instructions from a user and converting them into text using speech recognition technology; means for generating information gathering tasks based on the converted text and distributing them to an information sharing device; and means for collecting input from the information sharing device and analyzing the data through statistical processing. This enables users to efficiently collect and quickly share information using only voice instructions.

[0511] A "user" is an entity that uses the system, and is a person or group that initiates information gathering tasks through voice commands.

[0512] "Voice commands" are instructions issued by users through audio equipment and are signals that are converted into text using speech recognition technology.

[0513] "Speech recognition technology" is a technology that converts speech signals into text data, and it is an important process for understanding and processing user instructions.

[0514] An "information gathering task" is an activity generated by the system based on transcribed voice instructions, and includes components such as questionnaires and questions to collect specific information.

[0515] An "information sharing device" is a digital device that can receive collected information and allows for the input of responses.

[0516] "Statistical processing" is a method of analyzing collected data to derive certain trends and results, and it is the process of providing aggregated results to users.

[0517] "Audio equipment" refers to devices used when voice instructions are given, such as earphones and speakers.

[0518] The server's role in this system is to receive voice commands transmitted by users via audio equipment and convert that voice into text data using speech recognition technology. For this speech recognition, the `speech_recognition` software library is used as an example. The server then utilizes a generative AI model to create information gathering tasks based on the transcribed voice commands. These tasks are stored in a randomly accessible format and distributed to the information sharing device.

[0519] Information sharing devices (such as smartphones and tablets) display received information gathering tasks on their screens and accept input from users. This input information is then sent back to the server, which performs statistical processing on the input data. Specifically, it processes the data using numerical analysis libraries such as statistics and derives aggregated results. These results are ultimately notified to the user and used to facilitate rapid decision-making.

[0520] As a concrete example, the server receives a voice command from a user saying, "Please conduct a survey on movies we want to see this weekend," and generates a survey on related movies. Then, it collects the family's choices through an information sharing device, determines the most popular movie through statistical processing, and communicates the results to the family.

[0521] An example of a prompt might be: "Use a speech recognition API to convert speech instructions into text, create a questionnaire using a natural language generation model, and collect and summarize the responses."

[0522] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0523] Step 1:

[0524] The user inputs voice commands using audio equipment. The audio equipment captures the user's voice commands as audio signals and sends these signals to the server. The audio signals serve as input and provide the data necessary for subsequent processing.

[0525] Step 2:

[0526] When the server receives an audio signal, it uses speech recognition technology to convert the signal into text data. Here, a speech recognition software library (e.g., speech_recognition) is used to output the audio signal as text. This conversion prepares the text data for subsequent information processing.

[0527] Step 3:

[0528] The server uses the acquired text data and a generative AI model to automatically generate information gathering tasks. The text data is used as input, and survey content is output based on prompts (e.g., "Please conduct a survey about movies you'd like to see this weekend"). This process generates surveys tailored to the user's needs.

[0529] Step 4:

[0530] The generated information gathering tasks are distributed from the server to the information sharing device. The distributed information gathering tasks are displayed on the terminal screen and become ready to receive responses from the user. The information sharing device constructs a user interface based on the data received from the server.

[0531] Step 5:

[0532] Users enter their answers to a questionnaire via an information sharing device. The terminal collects the user's input in real time and sends the data back to the server. The entered response data becomes the raw material used for statistical processing.

[0533] Step 6:

[0534] The server collects responses submitted from terminals and performs statistical processing. Here, numerical analysis libraries such as Python's `statistics` library are used to analyze multiple response data and calculate aggregated results. These results serve as information for the user's final decision-making.

[0535] Step 7:

[0536] The server notifies the user of the aggregated results obtained through statistical processing. These results represent the user's final information and can be used, for example, as a reference when choosing a movie with family members. This output supports the decision-making process and improves user satisfaction.

[0537] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0538] In this invention, the user gives voice instructions through earphones. The voice spoken by the user is transmitted to a server. In addition to normal voice instructions, this system incorporates an emotion engine to extract emotional information from the voice. This emotion engine analyzes the emotional state from the user's voice tone, speed, and emphasized words. As a result, for example, if the user is speaking cheerfully, the questionnaire is adjusted to present casual options, while if the user is nervous, it presents more cautious options.

[0539] The server performs sentiment analysis simultaneously with speech-to-text conversion, and generates a questionnaire based on this analysis. This questionnaire includes questions and answer choices, and is personalized according to the user's emotions. The generated questionnaire is then delivered wirelessly to the devices of family members and related users.

[0540] The device displays the survey received from the server on its screen and notifies the user. The device displays questions and options that match the recipient's emotions, making the respondent feel more comfortable than usual. Once the user answers the survey on the device, the data is sent back to the server, and all responses are tallied.

[0541] The server statistically processes the aggregated results and summarizes them. The aggregated results are notified to the user, allowing them to easily review the information. For example, if a user asks, "What type of leisure activity would you like to enjoy on your next holiday?", the server detects positive emotions and provides a variety of cheerful leisure options.

[0542] This system enables a more personalized survey process that takes user emotions into account, resulting in more appropriate and contextually relevant information gathering.

[0543] The following describes the processing flow.

[0544] Step 1:

[0545] The user gives voice commands to the AI ​​agent through earphones. These commands can include the survey topic and questions they want to ask.

[0546] Step 2:

[0547] The server receives voice commands from the user and converts them into text data via a speech recognition system. Simultaneously, an emotion engine analyzes emotional information from the user's voice.

[0548] Step 3:

[0549] The server generates a survey tailored to the user's purpose based on recognized text data and sentiment information. It customizes the tone of questions and the content of answer choices based on sentiment information.

[0550] Step 4:

[0551] The server distributes the generated questionnaire to the devices of the target family members and related parties using wireless communication. This distribution takes place in real time.

[0552] Step 5:

[0553] The terminal receives the survey delivered from the server, displays it on the screen, and simultaneously notifies the user of the survey's arrival. The notification includes the survey's subject and response deadline.

[0554] Step 6:

[0555] Users respond to a survey displayed on their device. They can select options and enter comments as needed.

[0556] Step 7:

[0557] The terminal sends the user's entered answers to the server. The response data is compiled on the server, and statistical information for each item is calculated.

[0558] Step 8:

[0559] The server analyzes the aggregated results, compiles the overall findings, and then notifies the original user. In this process, the presentation of the results may be adjusted based on the user's emotional state.

[0560] Step 9:

[0561] Users receive notifications and can review the aggregated results. These results can then be used as material for future decision-making and discussions.

[0562] (Example 2)

[0563] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0564] Conventional voice-based survey systems lacked personalization that took into account the user's emotional state, resulting in difficulties in collecting information relevant to the user's context. Therefore, there is a need for a system that is more user-friendly and allows users to answer surveys more appropriately.

[0565] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0566] In this invention, the server includes means for recognizing the user's voice instructions, means for analyzing the emotional state from the recognized voice instructions, and means for generating a questionnaire based on the analysis results. This makes it possible to generate a personalized questionnaire that responds to the user's emotions.

[0567] A "user" is an individual or group that uses the system to give voice instructions and receives the results.

[0568] "Means for recognizing voice instructions" refers to technology or devices that capture a user's speech as voice data and convert it into text in order to understand its content.

[0569] "Means for analyzing emotional state" refers to technologies or devices that infer and analyze a user's emotions from the content and tone of voice instructions, speaking speed, emphasized words, etc.

[0570] "Means for generating questionnaires" refers to technologies or devices that dynamically create questions and answer choices in response to analyzed emotional states and present them to users.

[0571] A "communication device" is an electronic device that communicates with a server to send and receive data and notifies users of information.

[0572] "Means for aggregating responses" refers to a technology or device that statistically compiles and analyzes the responses to questionnaires submitted by users.

[0573] This system starts operating when the user gives voice commands through earphones. The earphones connect to a device (such as a smartphone or tablet) via Bluetooth and transmit the acquired voice data to a server.

[0574] The server processes the audio data using speech recognition software and an emotion analysis engine. Specifically, it converts speech to text using Google Cloud Speech-to-Text, and then extracts the user's emotions from the text using emotion analysis tools such as IBM Watson Tone Analyzer.

[0575] The interpreted sentiment information is used by a generative AI model to generate questionnaires. This generative AI model uses an open-source natural language processing library, and for example, if a user expresses a positive emotion, it dynamically creates a questionnaire that includes casual and diverse questions tailored to that emotion.

[0576] The generated questionnaire is delivered from the server to the relevant terminal via wireless communication. The terminal immediately displays the received questionnaire on its screen and notifies the user with audio and pop-up notifications. The user then answers the questionnaire, and the response is resent from the terminal to the server. This transmitted data is statistically analyzed through aggregation, and the results are fed back to the user from the server.

[0577] As a concrete example, consider conducting a survey about holiday plans. An example of a prompt question might be, "What type of leisure activity would you like to enjoy on your next holiday?" If positive emotions are detected, the survey will then provide a variety of bright leisure options, such as nature experiences, sports, and cultural activities. This allows users to answer the question in a more natural way.

[0578] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0579] Step 1:

[0580] The user gives voice commands through earphones. The voice is picked up by the earphone's microphone and transmitted as digital data to the device via Bluetooth. At this stage, the input is the user's voice, and the output is digital voice data.

[0581] Step 2:

[0582] The device transmits digital audio data to the server using Wi-Fi or mobile communication. The input here is the digital audio data acquired by the device, and the output is the data packets sent to the server via the network.

[0583] Step 3:

[0584] The server uses speech recognition software to convert audio data into text. This process analyzes the audio waveform and converts its content into character data. The input for this step is digital audio data, and the output is the converted text data.

[0585] Step 4:

[0586] The server uses an emotion analysis engine to extract emotional states from text data. Specifically, it analyzes the sentence structure, the presence or absence of exclamation marks, and specific words and phrases in the text to measure the user's emotions. In this step, the input is text data, and the output is the emotion analysis result.

[0587] Step 5:

[0588] The server uses a generative AI model to generate questionnaires based on sentiment analysis results. The AI ​​model, for example, utilizes open-source natural language processing libraries to design questions and answer choices that respond to emotions. The input to this process is the sentiment analysis results, and the output is questionnaire data containing questions and answer choices.

[0589] Step 6:

[0590] The server transmits survey data to the terminal via wireless communication. Wireless communication includes data encryption to ensure secure information transfer. The input is the survey data, and the output is the data packets sent to the terminal.

[0591] Step 7:

[0592] The terminal displays the survey on its screen interface and notifies the user. Here, the questions and answer choices are displayed on the terminal's screen, and the user is notified with a notification sound and banner. The input is the received survey data, and the output is the displayed visual interface.

[0593] Step 8:

[0594] The user answers a questionnaire and inputs their responses into the device. They select options using a touch panel or voice input and confirm their answers on the device. The input is the user's selection, and the output is the response data.

[0595] Step 9:

[0596] The terminal sends the response data to the server. Communication is conducted using network standards, ensuring accurate data transmission. The input is the response data from the terminal, and the output is a data structure for aggregation that is sent to the server.

[0597] Step 10:

[0598] The server aggregates the received response data and performs statistical processing. It uses a database system to efficiently consolidate responses and generate analysis results. The input is a large amount of response data, and the output is an aggregated statistical report.

[0599] (Application Example 2)

[0600] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0601] Conventional information gathering systems using listening devices failed to consider the emotional state of users, making it difficult to provide responses and suggestions tailored to individual users. This resulted in insufficient improvements in convenience and user satisfaction, and potentially impacted the accuracy of the information.

[0602] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0603] In this invention, the server includes a device for receiving a user's voice signal, means for analyzing the characteristics of the received voice signal to identify emotional information, and means for generating information collection documents based on the identified emotional information. This enables personalized information collection that responds to the user's emotions.

[0604] "User voice signal" refers to the electrical signal obtained by converting the user's voice into an electrical signal.

[0605] A "receiving device" is a machine or system that takes in an audio signal and uses it for subsequent processing.

[0606] "Means for analyzing the characteristics of an audio signal to identify emotional information" refers to a process or technique that analyzes the tone, speed, and emphasized words of speech and evaluates the emotional state from the results.

[0607] "Means for generating information gathering documents based on emotional information" refers to the process of creating information gathering documents according to identified emotional states.

[0608] A "communication device" is a technology or device used to exchange information or data with other systems or devices.

[0609] "Means of aggregation" refers to techniques or methods for collecting multiple responses and processing them statistically.

[0610] "Means of providing aggregated results to users" refers to technologies or methods for presenting processed aggregated data to users in an easily understandable manner.

[0611] The system for implementing this invention consists of a consumer robot installed in the home and an external server.

[0612] The server receives voice signals that the user makes to the robot. These signals are captured and digitized through a microphone built into the robot. Next, the server uses an emotion analysis engine to analyze the received voice signals. For example, IBM Watson or Google Cloud's voice analysis software may be used. Through this analysis, the server identifies the user's emotional state based on the tone, speed, and emphasized words of the voice.

[0613] Next, the server generates a personalized information gathering document based on the identified emotional information. For example, if the user indicates a cheerful mood, the document will include suggestions for relaxing activities. This information gathering document is sent to the robot and guided to the user via voice. Software such as Google Text-to-Speech is used for speech synthesis.

[0614] This robot can play relaxing music or suggest simple games through conversations with users in the home. For example, if a user says, "I'm a little tired today," the robot can play soothing music and offer a warm drink.

[0615] The following is an example of a prompt message:

[0616] User's voice: "I'm a little tired today."

[0617] Based on the results of the emotion recognition engine, please recognize the emotion of fatigue and suggest the following reactions:

[0618] "Thank you for your hard work. I'll play some relaxing music."

[0619] "Shall I make you some warm tea?"

[0620] In this way, this system allows users to receive personalized support tailored to their emotional state, making their daily lives more comfortable.

[0621] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0622] Step 1:

[0623] The user speaks to the robot. A microphone built into the robot converts this audio signal from analog to digital and sends it to the server. The input is the user's voice, and the output is digitized audio data. The server receives this data.

[0624] Step 2:

[0625] The server passes the received audio data to the emotion analysis engine. This process analyzes factors such as voice tone, speed, and emphasized words. The input is digitized audio data, and the output is information indicating the user's emotional state. Specifically, the server uses speech analysis software to identify the user's emotions.

[0626] Step 3:

[0627] Based on the identified sentiment information, the server generates personalized information gathering documents. The input is the analyzed sentiment information, and the output is an information gathering document tailored to that sentiment. For example, if a cheerful sentiment is detected, activity suggestions will be included.

[0628] Step 4:

[0629] The generated information gathering documents are sent from the server to the robot. The input is an information gathering document based on emotions, and the output is data that is transferred to the robot. The robot receives this data.

[0630] Step 5:

[0631] The robot uses the received data to provide voice guidance to the user. It uses speech synthesis software to generate reactions and interact with the user. The input is information collected from a server, and the output is voice guidance to the user. For example, if the user is tired, the robot might suggest playing relaxing music.

[0632] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0633] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0634] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0635] [Fourth Embodiment]

[0636] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0637] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0638] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0639] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0640] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0641] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0642] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0643] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0644] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0645] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0646] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0647] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0648] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0649] This invention relates to a system that starts operating when a user issues a voice command via earphones. The user provides the voice command to the AI ​​agent using the earphone microphone. This voice command is sent directly to the server. Upon receiving the command, the server converts it into text using speech recognition technology and generates a questionnaire based on it.

[0650] The generated questionnaires are stored on the server in a state where they can be retrieved via random access. The server then distributes the questionnaire information to each device of family members and related parties. When a device (smartphone or tablet) receives the questionnaire, it displays it on the screen and prompts the user to answer.

[0651] When a user enters their response on their device, the response is sent to the server. The server statistically processes each collected response and compiles the results. The compiled results are then reviewed by the user again and notified to the original user who requested the survey.

[0652] For example, consider a scenario where a user conducts a survey asking "What would you like to eat for dinner tonight?" The user gives a voice command, the server transcribes it into text, creates the survey, and distributes it. Each family member answers the survey on their own device, and the information is collected by the server. The server then compiles the results, such as 3 votes for curry and 2 votes for pasta, and notifies the user of the results.

[0653] This system enables the collection and rapid use of opinions from family and friends without physical contact or manual input. Thus, the goal of this invention is to facilitate efficient information gathering and communication.

[0654] The following describes the processing flow.

[0655] Step 1:

[0656] The user gives voice commands to the AI ​​agent through earphones. During this process, they verbally specify the content of the survey.

[0657] Step 2:

[0658] The server receives voice commands and converts them into text data using speech recognition technology. Based on this text data, it generates the survey questions.

[0659] Step 3:

[0660] The server saves the generated questionnaire to a database and prepares the questionnaire data, including the necessary information (questions and answer choices).

[0661] Step 4:

[0662] The server distributes the questionnaire to each device of the target family member or related party. At this time, the server identifies the recipient based on the device information.

[0663] Step 5:

[0664] The terminal receives survey data distributed from the server and displays it on the screen in a survey format. During this process, notifications are also sent to the user.

[0665] Step 6:

[0666] Users respond by operating the survey screen on their device, selecting options, and entering comments.

[0667] Step 7:

[0668] The terminal sends the user's response to the server. The transmitted data is then aggregated again on the server.

[0669] Step 8:

[0670] The server aggregates all received responses and generates statistical data. Based on these aggregated results, it forms an overview of the responses.

[0671] Step 9:

[0672] The server then notifies the user again of the compiled results. The user checks the notification on their smartphone and views the survey results.

[0673] (Example 1)

[0674] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0675] In today's information society, it is crucial to efficiently collect information and communicate regardless of time or place. In particular, quickly gathering opinions from family or groups is not easy and is a challenge many people face. Traditional methods require manual tabulation and physical contact, which is time-consuming. Furthermore, the difficulty in sharing information across multiple devices limits their use in situations requiring rapid decision-making.

[0676] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0677] In this invention, the server includes means for acquiring voice instructions from a user and converting them into digital information, means for generating a survey based on the converted digital information, and means for transmitting the generated survey to multiple communication devices. This enables users to quickly and efficiently generate questionnaires, distribute them to multiple locations, and aggregate the results using only voice.

[0678] A "user" is a person or group that issues voice commands to the system and generates and receives the results of the survey.

[0679] "Voice instructions" refer to the form of instructions or requests that users give to a system through hearing devices.

[0680] "Digital information" refers to electronic text data converted based on voice instructions, and forms the basis for generating questionnaires.

[0681] "Survey" refers to a form of questionnaire or inquiry generated based on the user's voice instructions.

[0682] "Communication equipment" refers to devices that receive generated surveys and allow users to input their responses, such as smartphones and tablets.

[0683] "Response" refers to the data provided by users via communication devices in response to a survey.

[0684] A "server" is a central computing system that performs a series of processes, from receiving voice commands to generating survey results, organizing responses, and communicating results.

[0685] This invention provides a system that allows users to easily generate questionnaires and streamline information gathering. Users use an earphone microphone (a type of hearing device) to issue voice commands to a server. The server uses speech recognition technology to convert the received voice commands into digital information. A general-purpose speech recognition API can be used as the speech recognition software for this process.

[0686] The server generates a survey based on the converted digital information and sends it to a communication device. This device, such as a smartphone or tablet, receives the survey, displays it to the user, and allows them to enter a response. After the user enters their response via the communication device, these responses are sent back to the server for further processing.

[0687] As a concrete example, we can consider a scenario where a user gives a voice command on the topic of "What do you want to eat for dinner tonight?". The server converts the voice into text data, "What do you want to eat for dinner tonight?", and generates a survey with options such as curry and pasta. This survey is distributed to the family's communication devices, and each member responds individually. The server then compiles these results and notifies the user of the number of votes for each option.

[0688] An example of a prompt statement generated using a generative AI model is as follows:

[0689] Prompt example:

[0690] "The user has initiated a voice survey about dinner. Please generate the survey based on the voice instructions and distribute it to the relevant parties' communication devices."

[0691] This system enables efficient investigations and supports rapid decision-making through voice commands alone.

[0692] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0693] Step 1:

[0694] The user provides voice instructions using an audio device. The user's input is specific, such as "What do you want to eat for dinner tonight?" This voice data is captured as an acoustic signal. This includes the user pressing the record button using the microphone function of their earphones and transmitting the voice instruction.

[0695] Step 2:

[0696] The server converts received audio data into text using speech recognition technology. The input is an audio signal, and the output is character data. The server applies a speech recognition algorithm, analyzes the audio waveform, and converts it into a string of characters. A speech recognition API is used in this process.

[0697] Step 3:

[0698] The server generates a survey based on the converted character data. The input is text data, and the output is data in the survey format. The server calculates answer choices based on the questions inferred from the text content and organizes the survey format.

[0699] Step 4:

[0700] The server transmits the generated survey data to communication devices. The input is data in the survey format, and the output is data distributed to each communication device. The server uses network functionality to broadcast the data to multiple terminals.

[0701] Step 5:

[0702] The terminal displays the received survey to the user and allows them to enter a response. The input is the survey data, and the output is the user's response. The terminal displays the survey content on the screen and enables the user to tap options or enter text.

[0703] Step 6:

[0704] Users input their responses to a survey via a terminal. Input consists of user selections from multiple-choice options or free text, while output is response data. Users select options on the terminal screen, confirm their input, and press the response button.

[0705] Step 7:

[0706] The terminal sends the collected responses to the server. The input is the response data entered on the terminal, and the output is the data sent to the server. The terminal aggregates the responses via the internet connection and sends them directly to the server.

[0707] Step 8:

[0708] The server organizes the received responses and generates aggregated results. The input is the response data from each user, and the output is the aggregated data. The server performs statistical processing, calculates results such as the number of votes for each option, and stores them in the database.

[0709] Step 9:

[0710] The server communicates the compiled and summarized results to the user. The input is the aggregated data, and the output is notification data for the user. The server delivers the results to the user via push notifications or email.

[0711] (Application Example 1)

[0712] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0713] In modern communication, efficient information sharing is essential. However, traditional technologies require users to manually collect and share information, which is time-consuming and cumbersome. This is particularly problematic in family decision-making, where smoothly gathering and reflecting opinions is difficult.

[0714] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0715] In this invention, the server includes means for receiving voice instructions from a user and converting them into text using speech recognition technology; means for generating information gathering tasks based on the converted text and distributing them to an information sharing device; and means for collecting input from the information sharing device and analyzing the data through statistical processing. This enables users to efficiently collect and quickly share information using only voice instructions.

[0716] A "user" is an entity that uses the system, and is a person or group that initiates information gathering tasks through voice commands.

[0717] "Voice commands" are instructions issued by users through audio equipment and are signals that are converted into text using speech recognition technology.

[0718] "Speech recognition technology" is a technology that converts speech signals into text data, and it is an important process for understanding and processing user instructions.

[0719] An "information gathering task" is an activity generated by the system based on transcribed voice instructions, and includes components such as questionnaires and questions to collect specific information.

[0720] An "information sharing device" is a digital device that can receive collected information and allows for the input of responses.

[0721] "Statistical processing" is a method of analyzing collected data to derive certain trends and results, and it is the process of providing aggregated results to users.

[0722] "Audio equipment" refers to devices used when voice instructions are given, such as earphones and speakers.

[0723] The server's role in this system is to receive voice commands transmitted by users via audio equipment and convert that voice into text data using speech recognition technology. For this speech recognition, the `speech_recognition` software library is used as an example. The server then utilizes a generative AI model to create information gathering tasks based on the transcribed voice commands. These tasks are stored in a randomly accessible format and distributed to the information sharing device.

[0724] Information sharing devices (such as smartphones and tablets) display received information gathering tasks on their screens and accept input from users. This input information is then sent back to the server, which performs statistical processing on the input data. Specifically, it processes the data using numerical analysis libraries such as statistics and derives aggregated results. These results are ultimately notified to the user and used to facilitate rapid decision-making.

[0725] As a concrete example, the server receives a voice command from a user saying, "Please conduct a survey on movies we want to see this weekend," and generates a survey on related movies. Then, it collects the family's choices through an information sharing device, determines the most popular movie through statistical processing, and communicates the results to the family.

[0726] An example of a prompt might be: "Use a speech recognition API to convert speech instructions into text, create a questionnaire using a natural language generation model, and collect and summarize the responses."

[0727] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0728] Step 1:

[0729] The user inputs voice commands using audio equipment. The audio equipment captures the user's voice commands as audio signals and sends these signals to the server. The audio signals serve as input and provide the data necessary for subsequent processing.

[0730] Step 2:

[0731] When the server receives an audio signal, it uses speech recognition technology to convert the signal into text data. Here, a speech recognition software library (e.g., speech_recognition) is used to output the audio signal as text. This conversion prepares the text data for subsequent information processing.

[0732] Step 3:

[0733] The server uses the acquired text data and a generative AI model to automatically generate information gathering tasks. The text data is used as input, and survey content is output based on prompts (e.g., "Please conduct a survey about movies you'd like to see this weekend"). This process generates surveys tailored to the user's needs.

[0734] Step 4:

[0735] The generated information gathering tasks are distributed from the server to the information sharing device. The distributed information gathering tasks are displayed on the terminal screen and become ready to receive responses from the user. The information sharing device constructs a user interface based on the data received from the server.

[0736] Step 5:

[0737] Users enter their answers to a questionnaire via an information sharing device. The terminal collects the user's input in real time and sends the data back to the server. The entered response data becomes the raw material used for statistical processing.

[0738] Step 6:

[0739] The server collects responses submitted from terminals and performs statistical processing. Here, numerical analysis libraries such as Python's `statistics` library are used to analyze multiple response data and calculate aggregated results. These results serve as information for the user's final decision-making.

[0740] Step 7:

[0741] The server notifies the user of the aggregated results obtained through statistical processing. These results represent the user's final information and can be used, for example, as a reference when choosing a movie with family members. This output supports the decision-making process and improves user satisfaction.

[0742] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0743] In this invention, the user gives voice instructions through earphones. The voice spoken by the user is transmitted to a server. In addition to normal voice instructions, this system incorporates an emotion engine to extract emotional information from the voice. This emotion engine analyzes the emotional state from the user's voice tone, speed, and emphasized words. As a result, for example, if the user is speaking cheerfully, the questionnaire is adjusted to present casual options, while if the user is nervous, it presents more cautious options.

[0744] The server performs sentiment analysis simultaneously with speech-to-text conversion, and generates a questionnaire based on this analysis. This questionnaire includes questions and answer choices, and is personalized according to the user's emotions. The generated questionnaire is then delivered wirelessly to the devices of family members and related users.

[0745] The device displays the survey received from the server on its screen and notifies the user. The device displays questions and options that match the recipient's emotions, making the respondent feel more comfortable than usual. Once the user answers the survey on the device, the data is sent back to the server, and all responses are tallied.

[0746] The server statistically processes the aggregated results and summarizes them. The aggregated results are notified to the user, allowing them to easily review the information. For example, if a user asks, "What type of leisure activity would you like to enjoy on your next holiday?", the server detects positive emotions and provides a variety of cheerful leisure options.

[0747] This system enables a more personalized survey process that takes user emotions into account, resulting in more appropriate and contextually relevant information gathering.

[0748] The following describes the processing flow.

[0749] Step 1:

[0750] The user gives voice commands to the AI ​​agent through earphones. These commands can include the survey topic and questions they want to ask.

[0751] Step 2:

[0752] The server receives voice commands from the user and converts them into text data via a speech recognition system. Simultaneously, an emotion engine analyzes emotional information from the user's voice.

[0753] Step 3:

[0754] The server generates a survey tailored to the user's purpose based on recognized text data and sentiment information. It customizes the tone of questions and the content of answer choices based on sentiment information.

[0755] Step 4:

[0756] The server distributes the generated questionnaire to the devices of the target family members and related parties using wireless communication. This distribution takes place in real time.

[0757] Step 5:

[0758] The terminal receives the survey delivered from the server, displays it on the screen, and simultaneously notifies the user of the survey's arrival. The notification includes the survey's subject and response deadline.

[0759] Step 6:

[0760] Users respond to a survey displayed on their device. They can select options and enter comments as needed.

[0761] Step 7:

[0762] The terminal sends the user's entered answers to the server. The response data is compiled on the server, and statistical information for each item is calculated.

[0763] Step 8:

[0764] The server analyzes the aggregated results, compiles the overall findings, and then notifies the original user. In this process, the presentation of the results may be adjusted based on the user's emotional state.

[0765] Step 9:

[0766] Users receive notifications and can review the aggregated results. These results can then be used as material for future decision-making and discussions.

[0767] (Example 2)

[0768] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0769] Conventional voice-based survey systems lacked personalization that took into account the user's emotional state, resulting in difficulties in collecting information relevant to the user's context. Therefore, there is a need for a system that is more user-friendly and allows users to answer surveys more appropriately.

[0770] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0771] In this invention, the server includes means for recognizing the user's voice instructions, means for analyzing the emotional state from the recognized voice instructions, and means for generating a questionnaire based on the analysis results. This makes it possible to generate a personalized questionnaire that responds to the user's emotions.

[0772] A "user" is an individual or group that uses the system to give voice instructions and receives the results.

[0773] "Means for recognizing voice instructions" refers to technology or devices that capture a user's speech as voice data and convert it into text in order to understand its content.

[0774] "Means for analyzing emotional state" refers to technologies or devices that infer and analyze a user's emotions from the content and tone of voice instructions, speaking speed, emphasized words, etc.

[0775] "Means for generating questionnaires" refers to technologies or devices that dynamically create questions and answer choices in response to analyzed emotional states and present them to users.

[0776] A "communication device" is an electronic device that communicates with a server to send and receive data and notifies users of information.

[0777] "Means for aggregating responses" refers to a technology or device that statistically compiles and analyzes the responses to questionnaires submitted by users.

[0778] This system starts operating when the user gives voice commands through earphones. The earphones connect to a device (such as a smartphone or tablet) via Bluetooth and transmit the acquired voice data to a server.

[0779] The server processes the audio data using speech recognition software and an emotion analysis engine. Specifically, it converts speech to text using Google Cloud Speech-to-Text, and then extracts the user's emotions from the text using emotion analysis tools such as IBM Watson Tone Analyzer.

[0780] The interpreted sentiment information is used by a generative AI model to generate questionnaires. This generative AI model uses an open-source natural language processing library, and for example, if a user expresses a positive emotion, it dynamically creates a questionnaire that includes casual and diverse questions tailored to that emotion.

[0781] The generated questionnaire is delivered from the server to the relevant terminal via wireless communication. The terminal immediately displays the received questionnaire on its screen and notifies the user with audio and pop-up notifications. The user then answers the questionnaire, and the response is resent from the terminal to the server. This transmitted data is statistically analyzed through aggregation, and the results are fed back to the user from the server.

[0782] As a concrete example, consider conducting a survey about holiday plans. An example of a prompt question might be, "What type of leisure activity would you like to enjoy on your next holiday?" If positive emotions are detected, the survey will then provide a variety of bright leisure options, such as nature experiences, sports, and cultural activities. This allows users to answer the question in a more natural way.

[0783] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0784] Step 1:

[0785] The user gives voice commands through earphones. The voice is picked up by the earphone's microphone and transmitted as digital data to the device via Bluetooth. At this stage, the input is the user's voice, and the output is digital voice data.

[0786] Step 2:

[0787] The device transmits digital audio data to the server using Wi-Fi or mobile communication. The input here is the digital audio data acquired by the device, and the output is the data packets sent to the server via the network.

[0788] Step 3:

[0789] The server uses speech recognition software to convert audio data into text. This process analyzes the audio waveform and converts its content into character data. The input for this step is digital audio data, and the output is the converted text data.

[0790] Step 4:

[0791] The server uses an emotion analysis engine to extract emotional states from text data. Specifically, it analyzes the sentence structure, the presence or absence of exclamation marks, and specific words and phrases in the text to measure the user's emotions. In this step, the input is text data, and the output is the emotion analysis result.

[0792] Step 5:

[0793] The server uses a generative AI model to generate questionnaires based on sentiment analysis results. The AI ​​model, for example, utilizes open-source natural language processing libraries to design questions and answer choices that respond to emotions. The input to this process is the sentiment analysis results, and the output is questionnaire data containing questions and answer choices.

[0794] Step 6:

[0795] The server transmits survey data to the terminal via wireless communication. Wireless communication includes data encryption to ensure secure information transfer. The input is the survey data, and the output is the data packets sent to the terminal.

[0796] Step 7:

[0797] The terminal displays the survey on its screen interface and notifies the user. Here, the questions and answer choices are displayed on the terminal's screen, and the user is notified with a notification sound and banner. The input is the received survey data, and the output is the displayed visual interface.

[0798] Step 8:

[0799] The user answers a questionnaire and inputs their responses into the device. They select options using a touch panel or voice input and confirm their answers on the device. The input is the user's selection, and the output is the response data.

[0800] Step 9:

[0801] The terminal sends the response data to the server. Communication is conducted using network standards, ensuring accurate data transmission. The input is the response data from the terminal, and the output is a data structure for aggregation that is sent to the server.

[0802] Step 10:

[0803] The server aggregates the received response data and performs statistical processing. It uses a database system to efficiently consolidate responses and generate analysis results. The input is a large amount of response data, and the output is an aggregated statistical report.

[0804] (Application Example 2)

[0805] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0806] Conventional information gathering systems using listening devices failed to consider the emotional state of users, making it difficult to provide responses and suggestions tailored to individual users. This resulted in insufficient improvements in convenience and user satisfaction, and potentially impacted the accuracy of the information.

[0807] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0808] In this invention, the server includes a device for receiving a user's voice signal, means for analyzing the characteristics of the received voice signal to identify emotional information, and means for generating information collection documents based on the identified emotional information. This enables personalized information collection that responds to the user's emotions.

[0809] "User voice signal" refers to the electrical signal obtained by converting the user's voice into an electrical signal.

[0810] A "receiving device" is a machine or system that takes in an audio signal and uses it for subsequent processing.

[0811] "Means for analyzing the characteristics of an audio signal to identify emotional information" refers to a process or technique that analyzes the tone, speed, and emphasized words of speech and evaluates the emotional state from the results.

[0812] "Means for generating information gathering documents based on emotional information" refers to the process of creating information gathering documents according to identified emotional states.

[0813] A "communication device" is a technology or device used to exchange information or data with other systems or devices.

[0814] "Means of aggregation" refers to techniques or methods for collecting multiple responses and processing them statistically.

[0815] "Means of providing aggregated results to users" refers to technologies or methods for presenting processed aggregated data to users in an easily understandable manner.

[0816] The system for implementing this invention consists of a consumer robot installed in the home and an external server.

[0817] The server receives voice signals that the user makes to the robot. These signals are captured and digitized through a microphone built into the robot. Next, the server uses an emotion analysis engine to analyze the received voice signals. For example, IBM Watson or Google Cloud's voice analysis software may be used. Through this analysis, the server identifies the user's emotional state based on the tone, speed, and emphasized words of the voice.

[0818] Next, the server generates a personalized information gathering document based on the identified emotional information. For example, if the user indicates a cheerful mood, the document will include suggestions for relaxing activities. This information gathering document is sent to the robot and guided to the user via voice. Software such as Google Text-to-Speech is used for speech synthesis.

[0819] This robot can play relaxing music or suggest simple games through conversations with users in the home. For example, if a user says, "I'm a little tired today," the robot can play soothing music and offer a warm drink.

[0820] The following is an example of a prompt message:

[0821] User's voice: "I'm a little tired today."

[0822] Based on the results of the emotion recognition engine, please recognize the emotion of fatigue and suggest the following reactions:

[0823] "Thank you for your hard work. I'll play some relaxing music."

[0824] "Shall I make you some warm tea?"

[0825] In this way, this system allows users to receive personalized support tailored to their emotional state, making their daily lives more comfortable.

[0826] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0827] Step 1:

[0828] The user speaks to the robot. A microphone built into the robot converts this audio signal from analog to digital and sends it to the server. The input is the user's voice, and the output is digitized audio data. The server receives this data.

[0829] Step 2:

[0830] The server passes the received audio data to the emotion analysis engine. This process analyzes factors such as voice tone, speed, and emphasized words. The input is digitized audio data, and the output is information indicating the user's emotional state. Specifically, the server uses speech analysis software to identify the user's emotions.

[0831] Step 3:

[0832] Based on the identified sentiment information, the server generates personalized information gathering documents. The input is the analyzed sentiment information, and the output is an information gathering document tailored to that sentiment. For example, if a cheerful sentiment is detected, activity suggestions will be included.

[0833] Step 4:

[0834] The generated information gathering documents are sent from the server to the robot. The input is an information gathering document based on emotions, and the output is data that is transferred to the robot. The robot receives this data.

[0835] Step 5:

[0836] The robot uses the received data to provide voice guidance to the user. It uses speech synthesis software to generate reactions and interact with the user. The input is information collected from a server, and the output is voice guidance to the user. For example, if the user is tired, the robot might suggest playing relaxing music.

[0837] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0838] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0839] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0840] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0841] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0842] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0843] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0844] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0845] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0846] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0847] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0848] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0849] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0850] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0851] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0852] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0853] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0854] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0855] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0856] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0857] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0858] The following is further disclosed regarding the embodiments described above.

[0859] (Claim 1)

[0860] A means of recognizing the user's voice instructions,

[0861] A means for generating a questionnaire based on recognized voice instructions,

[0862] A means for distributing the generated questionnaire to multiple wireless devices,

[0863] A means of collecting responses from wireless devices,

[0864] A means of notifying users of the aggregated results,

[0865] A system that includes this.

[0866] (Claim 2)

[0867] The system according to claim 1, in which recognition and distribution are performed when voice instructions are given through earphones.

[0868] (Claim 3)

[0869] The system according to claim 1, wherein the questionnaire supports not only multiple-choice format but also comment input.

[0870] "Example 1"

[0871] (Claim 1)

[0872] A means of acquiring user voice commands and converting them into digital information,

[0873] Means for generating a survey based on converted digital information,

[0874] A means for transmitting the generated survey to multiple communication devices,

[0875] A means of organizing responses from communication devices,

[0876] A means of communicating the organized results to the user,

[0877] A system that includes this.

[0878] (Claim 2)

[0879] The system according to claim 1, wherein voice commands are acquired and transmitted when given through an auditory device.

[0880] (Claim 3)

[0881] The system according to claim 1, wherein the survey supports free input in addition to multiple-choice format.

[0882] "Application Example 1"

[0883] (Claim 1)

[0884] A means of receiving voice instructions from a user and converting them into text using speech recognition technology,

[0885] A means for generating an information gathering task based on the converted text and distributing it to an information sharing device,

[0886] A means of collecting input from an information sharing device and analyzing the data through statistical processing,

[0887] Means for providing analysis results to users,

[0888] An information processing system that includes this.

[0889] (Claim 2)

[0890] The information processing system according to claim 1, which delivers recognition and information processing tasks when voice instructions are provided through an audio device.

[0891] (Claim 3)

[0892] The information processing system according to claim 1, wherein the information gathering task supports both a multiple-choice format and a free-form input format.

[0893] "Example 2 of combining an emotion engine"

[0894] (Claim 1)

[0895] A means of recognizing the user's voice instructions,

[0896] A means of analyzing emotional states from recognized voice commands,

[0897] A means for generating a questionnaire based on the analysis results,

[0898] A means for distributing the generated questionnaire to multiple communication devices,

[0899] A means of collecting responses from communication devices,

[0900] A means of notifying users of the aggregated results,

[0901] A system that includes this.

[0902] (Claim 2)

[0903] The system according to claim 1, wherein recognition and emotion analysis are performed when voice commands are given through a headset.

[0904] (Claim 3)

[0905] The system according to claim 1, wherein the questionnaire supports not only multiple-choice format but also open-ended format.

[0906] "Application example 2 when combining with an emotional engine"

[0907] (Claim 1)

[0908] A device that receives the user's voice signal,

[0909] A means of identifying emotional information by analyzing the characteristics of the received audio signal,

[0910] A means for generating information gathering documents based on identified emotional information,

[0911] A means for transmitting the generated information collection document to a communication device,

[0912] A means for aggregating responses from communication devices,

[0913] Means of providing aggregated results to users,

[0914] A system that includes this.

[0915] (Claim 2)

[0916] The system according to claim 1, wherein analysis and transmission are performed when an audio signal is input through a listening device.

[0917] (Claim 3)

[0918] The system according to claim 1, wherein the information gathering document supports not only multiple-choice format but also free-text format. [Explanation of Symbols]

[0919] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of receiving voice instructions from a user and converting them into text using speech recognition technology, A means for generating an information gathering task based on the converted text and distributing it to an information sharing device, A means of collecting input from an information sharing device and analyzing the data through statistical processing, Means for providing analysis results to users, An information processing system that includes this.

2. The information processing system according to claim 1, which delivers recognition and information processing tasks when voice instructions are provided through an audio device.

3. The information processing system according to claim 1, wherein the information gathering task supports both a multiple-choice format and a free-form input format.