system
A system that converts voice data to text and generates relevant questions addresses the issue of missed confirmations in meetings, enhancing meeting quality and efficiency by ensuring all important points are addressed.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-12-12
- Publication Date
- 2026-06-24
Smart Images

Figure 2026103444000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In a communication place such as a meeting, participants may not be able to ask appropriate questions promptly, resulting in important confirmations being missed. Such problems may cause a decline in the quality of the meeting and overlooking of important matters, and may have an adverse impact on work efficiency. To solve this problem, there is a demand for realizing a system that grasps the conversation content in real time and generates and presents relevant questions.
Means for Solving the Problems
[0005] The present invention provides a system including voice collection means for collecting voice data, conversion means for converting the voice data into text, analysis and generation means for analyzing the text and generating related questions, and presentation means for presenting the generated questions to the user. This makes it possible to quickly ask appropriate questions without missing important confirmations during a conversation, thereby improving the quality of meetings and work efficiency.
[0006] "Voice collection means" refers to functions or devices that record participants' voices in real time during meetings or discussions and supply that voice data to a system.
[0007] The "conversion means" refers to a function that uses speech recognition technology to convert collected audio data into text data in preparation for subsequent processing.
[0008] "Analysis and generation means" refers to functions that use machine learning algorithms and natural language processing techniques to analyze text data and automatically generate relevant questions based on the context of the conversation.
[0009] A "presentation means" is an interface or device that presents generated questions to the user visually or audibly, helping the user to smoothly confirm conversations and supplement discussions. [Brief explanation of the drawing]
[0010] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5]This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0011] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0012] First, let's explain the terminology used in the following explanation.
[0013] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0014] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0015] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0016] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0017] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0018] [First Embodiment]
[0019] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0020] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0021] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0022] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0023] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0024] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0025] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0026] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0027] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0028] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0029] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0030] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0031] To implement this invention, a system is constructed in which three main entities—a terminal, a server, and a user—work in cooperation. First, the terminal uses a microphone at the meeting to collect participants' voices in real time. On the terminal, the voice collection means plays the role of converting the voice data into text. In this conversion, the terminal uses speech recognition technology to accurately transcribe what it hears into text.
[0032] Next, the terminal sends the converted text to the server. On the server, an analysis and generation system receives the text data, analyzes it using natural language processing algorithms, and extracts important keywords and topics. Based on these analysis results, the server uses machine learning techniques to generate relevant questions that are appropriate to the context of the conversation.
[0033] The generated questions are sent from the server to the terminal and presented to the user visually or audibly through a presentation method. For example, if the meeting is about the development of a new product, specific questions such as "Where is the target market?" or "What are the challenges during development?" are generated and presented. By using these questions as a reference during the meeting, users can ensure that all important points of the discussion are covered and that no details are overlooked.
[0034] In this way, the system of the present invention supports important confirmations in meetings and enables the efficient promotion of discussions. Furthermore, by continuously learning and improving the server-side AI model using user feedback, it becomes possible to generate even more accurate questions. This system is expected to improve the quality of meetings and the efficiency of work.
[0035] The following describes the processing flow.
[0036] Step 1:
[0037] The device activates the microphone at the start of the meeting and collects participants' voices in real time. The voice data is temporarily stored in a buffer to prepare for the next processing.
[0038] Step 2:
[0039] The device uses speech recognition technology to convert collected audio data into text data. This text is grammatically formatted, and the meeting content is stored in an easily understandable format.
[0040] Step 3:
[0041] The terminal sends formatted text data to the server. The transmission takes place over the network and is done in streaming format rather than batch processing to enable real-time data processing.
[0042] Step 4:
[0043] The server analyzes the received text data using analysis and generation tools. Natural language processing (NLP) is used to extract keywords and classify topics. This analysis identifies key points in the conversation.
[0044] Step 5:
[0045] Based on the analysis results, the server uses machine learning algorithms to generate relevant questions. This process takes into account past conversation patterns and contextual data to create questions that are useful and appropriate for the user.
[0046] Step 6:
[0047] The server sends the generated question to the terminal. The sent question is displayed on the terminal's user interface and presented to the user visually or audibly.
[0048] Step 7:
[0049] Users review the presented questions and make adjustments based on important confirmations and discussions as the meeting progresses. If necessary, users can improve the quality of the meeting by adding or refining questions.
[0050] Step 8:
[0051] The server collects user feedback and uses it to refine and train the AI model. This feedback helps improve the accuracy of question generation for the next meeting.
[0052] (Example 1)
[0053] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0054] In meetings and discussions, participants often miss important points, making it difficult to efficiently and effectively grasp the meeting content. Furthermore, there is a need to generate relevant questions to support meeting progress and improve the quality of meetings.
[0055] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0056] In this invention, the server includes an information gathering means for acquiring audio information, a conversion means for converting audio information into text information, an analysis and generation means for analyzing text information and generating related queries, and a learning means for collecting feedback from users and improving the analysis and generation means. This makes it possible to efficiently and effectively support the progress of meetings without missing important information during meetings, and to improve the quality of meetings.
[0057] "Audio information" refers to the waveforms of sounds emitted by participants during meetings and discussions, captured as digital data.
[0058] "Information gathering means" refers to devices and software installed to acquire audio information.
[0059] "Textual information" refers to digital data obtained by converting audio information into text or written characters.
[0060] "Conversion means" refers to a technology or device for converting audio information into text information, and utilizes speech recognition technology.
[0061] "Analysis and generation means" refers to technologies and processes for analyzing textual information and generating related queries based on that analysis.
[0062] "Users" refer to individuals or organizations that use the system and are responsible for receiving inquiries and facilitating the meeting.
[0063] "Presentation means" refers to methods or devices for informing the user of a generated inquiry, and includes visual or auditory methods.
[0064] "Feedback" refers to user reactions and evaluations of a system, and is information used to improve the system.
[0065] "Learning methods" refer to techniques and methods for continuously improving analysis and generation methods using feedback.
[0066] To implement this invention, three entities—a terminal, a server, and a user—must work together in coordination. In particular, this system is designed to support important discussions in meetings and improve the quality of those meetings.
[0067] Collection and conversion of audio information
[0068] The device uses a microphone during the meeting to collect audio information in real time. This information is stored as digital data and converted using speech recognition technology. Specifically, it uses speech recognition services such as Google® Cloud Speech-to-Text API and Microsoft® Azure® Speech Service. Through this process, the audio information is instantly converted into text.
[0069] Text information analysis and question generation
[0070] The terminal sends text information to the server, where analysis begins. The server can use libraries such as NLTK and SpaCy for natural language processing analysis. This analysis extracts important keywords and topics from the text information. Then, using generative AI models like GPT and BERT, contextual questions are generated based on the text information. These generated questions are highly relevant and designed to further deepen the discussion.
[0071] Question formulation and improvement
[0072] The generated questions are sent from the server to the terminal and presented to the user via the terminal's display or speech synthesis tool. This presentation allows the user to receive the questions visually or audibly, facilitating smoother meeting progress. Furthermore, users provide feedback after the meeting. This feedback information is stored on the server and used to train the generating AI model. This improves the accuracy of question generation, making support in future meetings more effective.
[0073] For example, in a meeting about a new product, participants' statements are recognized by speech recognition, and keywords such as "target market" and "competitor analysis" are extracted. Based on this, the server generates specific questions such as "How will we differentiate ourselves from competitors?" and presents them to the user via their terminal.
[0074] An example of a prompt message would be: "Listen to the discussion about new product development in the meeting and generate relevant questions, specifically about the target market and the challenges under development."
[0075] As described above, this system can support the efficient and effective conduct of meetings, thereby improving work efficiency and the quality of discussions.
[0076] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0077] Step 1:
[0078] Collection of audio information
[0079] The terminal uses a microphone during the meeting to collect participants' voice information in real time. In this step, the voice waveform is input to the terminal as digital data. Voice collection software runs on the terminal to properly record the voice waveform. This digital data is used in a later conversion step.
[0080] Step 2:
[0081] Converting speech to text information
[0082] The terminal begins processing the collected audio information into text information. It processes the audio digital data obtained in step 1 as input. Specifically, it utilizes speech recognition technology, leveraging services such as the Google Cloud Speech-to-Text API and Microsoft Azure Speech Service. This outputs the audio waveform as text information in the form of words and sentences. This text information is then used in the next analysis step.
[0083] Step 3:
[0084] Text information transfer and analysis
[0085] The terminal securely transfers the converted character information to the server. The input is the character information generated in step 2. The server receives this character information and analyzes it using natural language processing algorithms such as Python's NLTK or SpaCy. The purpose of the analysis is to extract important keywords and topics from the text data. The output of this analysis is useful in the subsequent question generation step.
[0086] Step 4:
[0087] Question generation
[0088] Based on the analysis results from step 3, the server generates questions using a generative AI model. This process utilizes generative AI models such as GPT and BERT. The input consists of keywords and contextual information extracted through the analysis. The output is a list of highly relevant questions, which are then used to further the meeting.
[0089] Step 5:
[0090] Question presentation
[0091] The server sends the generated questions to the terminal. The terminal presents this list of questions to the user via a display or speech synthesis tool. The input is the list of questions generated in step 4. By receiving these questions, the user can effectively conduct the meeting. The output is the visual or auditory presentation of questions that the user uses.
[0092] Step 6:
[0093] Gathering feedback and improving the model
[0094] After the meeting, users provide feedback to the system. The server receives this feedback as input and incorporates it into the training data for the generative AI model. This enables more accurate question generation in the next meeting. The output is the improved generative AI model.
[0095] (Application Example 1)
[0096] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0097] There is a need to improve the efficiency of customer service in stores and to provide information more quickly and accurately. However, currently, there is insufficient support to accurately respond to individual customer questions, which places a heavy burden on staff and limits the improvement of customer satisfaction. In addition, the automatic generation of relevant information based on the content of the conversation is often not performed properly, hindering the smooth flow of conversation.
[0098] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0099] In this invention, the server includes an acquisition means for acquiring acoustic data, a conversion means for converting acoustic data into text information, and an analysis and generation means for analyzing the text information and generating related information. This enables store staff to instantly acquire accurate and relevant information during interactions with customers, allowing for efficient customer service.
[0100] "Audio data" refers to a data format that digitizes audio or sound information.
[0101] "Means of acquisition" refers to the devices and methods used to collect acoustic data.
[0102] "Conversion means" refers to a technical process for converting acoustic data into textual information.
[0103] "Textual information" refers to information expressed in text format.
[0104] "Analysis and generation means" refers to a method of analyzing textual information and automatically generating related information based on that analysis.
[0105] "Presentation means" refers to a function for presenting generated information to the user visually or audibly.
[0106] "User" refers to a person or group that uses the means of presentation.
[0107] "Recommended information" refers to additional information that is presumed to be useful to the user.
[0108] "Dialogue data" refers to the content of conversations recorded in past audio or text.
[0109] "Learning methods" refer to machine learning techniques used to analyze dialogue data and optimize relevant information.
[0110] "Immediate" means that an action is performed instantly, without delay.
[0111] To implement this invention, a system for acquiring acoustic data is first constructed. The terminal acquires the acoustic data and converts it into text information using a conversion means. The "speech_recognition" library, which performs speech recognition, is suitable as the software to be used here. The converted text information is sent to the server.
[0112] The server analyzes the received text information using parsing and generation tools and generates relevant information. Natural language processing technology using "transformers" is employed in this process. This makes it possible to instantly generate relevant recommendation information based on customer interaction. The information obtained through analysis is sent back to the terminal and displayed to the user by a presentation tool. Based on this information, the user can smoothly proceed with the interaction with the customer.
[0113] This system allows store staff, for example, when asked by a customer, "What are the features of the new product?", to instantly obtain relevant features and recommended uses, enabling appropriate communication. As described above, the system of the present invention can efficiently handle the process from voice data acquisition to analysis and information presentation.
[0114] The generative AI model generates relevant information based on conversations in various contexts. For example, a prompt might read, "Generate a revised question to quickly obtain detailed information about the products or services offered. Please also provide information that can help improve the question."
[0115] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0116] Step 1:
[0117] The device uses its built-in microphone to capture audio during conversations with customers. It receives an acoustic signal as input and records it as digital audio data. Specifically, the device launches a speech recognition application in the background and performs audio capture.
[0118] Step 2:
[0119] The device converts acquired audio data into text. It uses audio data as input and generates text information through a conversion mechanism. In this process, a speech recognition library is used to convert speech to text. Specifically, it analyzes the audio waveform based on a language model and outputs the corresponding words as characters.
[0120] Step 3:
[0121] The server analyzes text data sent from the terminal. It receives text data as input and uses analysis and generation tools to extract important keywords and related information. Here, natural language processing techniques are used to analyze the text content and generate highly relevant information. Specifically, it extracts frequently occurring words from the received text while understanding the context and participating in the generation of recommended information.
[0122] Step 4:
[0123] The server generates relevant information based on the analysis results and sends it to the terminal. It generates recommended information as output and sends it back to the terminal via a presentation mechanism. Specifically, the process involves structuring the generated questions and information and conveying them to the terminal using a communication protocol.
[0124] Step 5:
[0125] The terminal presents relevant information received from the server to the user visually or audibly. It utilizes the received recommendation information as input and presents it via screen display or audio output. Specifically, the received information is delivered to the user in an easy-to-understand manner via a GUI or voice assistant.
[0126] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0127] For this invention to be implemented, the interaction between the terminal, server, and user is crucial. First, the terminal incorporates a voice collection means and an emotion engine. The voice collection means collects participants' voices in real time during a meeting and converts them into text. The emotion engine analyzes the user's emotions from this voice and text data. Specifically, it detects changes in voice tone and speaking style to infer the user's emotional state, such as whether they are tense or relaxed.
[0128] Next, the terminal sends the converted text data and sentiment analysis results to the server. The server analyzes the text and sentiment data using analysis and generation tools and generates questions that take into account the user's current emotional state. In this process, to support emotion-based communication, for example, if the user is feeling anxious, questions that encourage relaxation can be generated. Furthermore, the sentiment analysis results can be compared with past data to enable more personalized questions.
[0129] The generated questions are sent from the server to the terminal and presented to the user through a presentation mechanism. This allows the user to review the questions provided during the meeting and obtain optimal guidance that takes their emotional state into consideration. For example, if the emotional engine detects increased tension during a meeting to discuss the design of a new product, a question such as, "What kind of environment do you think is needed to think about this point more relaxed?" might be provided, enabling a discussion that is also supported emotionally.
[0130] As a result, the system of the present invention supports important discussions while considering the emotional balance in meetings, and improves the efficiency and quality of communication. In this way, emotionally considerate conversations are realized, providing a more productive environment.
[0131] The following describes the processing flow.
[0132] Step 1:
[0133] The device activates its audio collection system at the start of the meeting, collecting the voices of meeting participants in real time. Simultaneously, an emotion engine analyzes changes in voice tone and speaking style to infer the emotional state of the participants.
[0134] Step 2:
[0135] The device converts the collected audio into text. Using speech recognition technology, grammatically correct text is generated from the audio data. During this process, the emotion engine continuously monitors the audio data and updates the emotional state.
[0136] Step 3:
[0137] The terminal sends the converted text and sentiment analysis results to the server. The transmission is done in real time, and the system is designed to allow data processing in line with the progress of the meeting.
[0138] Step 4:
[0139] The server analyzes the received text and sentiment data. Analysis and generation tools analyze the text content and extract important topics. In addition, sentiment data is used to design questions that are appropriate to the user's current emotions.
[0140] Step 5:
[0141] The server uses a machine learning model to generate adaptive questions by comparing them with past data. For example, the generated questions might be designed to reassure a user who is in an unstable state.
[0142] Step 6:
[0143] The server sends the generated questions to the terminal and presents them to the user visually or audibly through a presentation mechanism. The user reviews these questions and responds with consideration for emotions as the meeting discussion progresses.
[0144] Step 7:
[0145] Users utilize the presented questions to continue the discussion appropriately within the meeting. Furthermore, they can modify the conversation or add new questions based on their own emotional state if necessary.
[0146] Step 8:
[0147] The server collects user feedback after the meeting and uses it to update and improve the AI model. This feedback is used to improve the accuracy of question generation in the next meeting.
[0148] (Example 2)
[0149] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0150] Traditional meeting systems simply convert audio data into text, lacking communication that takes into account the emotional state of the speaker. As a result, there was insufficient support to reduce tension and stress among meeting participants, leading to problems with the efficiency and quality of communication.
[0151] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0152] In this invention, the server includes means for converting voice data into language data, analysis means for analyzing the language data and voice data and inferring the user's emotional state, and generation means for generating questions to support conversation based on the user's emotional state. This enables communication support tailored to the user's emotional state.
[0153] "Audio data" refers to information used to record or transmit audio.
[0154] A "device" is a component of hardware or software designed to perform a specific function or purpose.
[0155] "Linguistic data" refers to information expressed in natural language, which is usually stored or transmitted in text format.
[0156] An "analysis device" is a device used to examine data and extract the meaning and information behind it.
[0157] "User" refers to an individual or group that operates a system or device and benefits from it.
[0158] "Emotional state" refers to information that indicates the speaker's psychological state or mood, and is usually inferred from the tone of voice and facial expressions.
[0159] A "generation device" is a device that has the function of creating new information or content based on input data.
[0160] A "machine learning system" is a computer system that uses data to automatically learn and execute algorithms that optimize a specific task.
[0161] "Instantly" refers to processing in near real-time, minimizing delays.
[0162] This invention is a system for improving the efficiency and quality of communication during meetings. Specifically, it is a system that collects speaker voice data in real time, analyzes that data to understand the speaker's emotional state, and generates customized questions based on that emotion.
[0163] The device is equipped with a microphone and speaker to collect participants' voices in real time. Once this voice data is collected, it is converted into text data using speech recognition software on the device. Generally, commercially available speech recognition APIs are used for speech recognition.
[0164] Next, the terminal performs sentiment analysis using the converted text data. Here, a sentiment analysis engine is used to analyze changes in voice tone and speaking rhythm. This engine utilizes generally known sentiment analysis tools.
[0165] The analyzed data is sent from the terminal to the server. The server inputs this data into a generative AI model to generate questions that take into account the user's emotional state. A generative AI model is a program that analyzes data and evolutionarily generates the optimal response. For example, machine learning frameworks such as TENSORFLOW® are often used.
[0166] The questions generated on the server are sent back to the terminal and displayed to the user by a presentation device. The presentation device includes displays and visual media to clearly display the questions visually and aid in user comprehension.
[0167] As a concrete example, when discussing a new project in a meeting, if the emotion engine detects tension among participants, it will generate a question such as, "What kind of environment do you think is needed to help us think about this more relaxed?" This allows for a conversation that encourages participants to relax, leading to a more lively discussion. An example of a prompt would be, "The user's emotional state is tense. Please generate appropriate questions to support a conversation that takes this situation into consideration." This specific content would be used as input to the generating AI model.
[0168] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0169] Step 1:
[0170] The device collects participants' voices in real time. The input is raw audio data acquired using a microphone. Specifically, during a meeting, the device captures each participant's speech, resulting in digital audio data.
[0171] Step 2:
[0172] The device converts the collected audio data into text data using speech recognition software. The input is the audio data obtained in step 1, and the output is a string of text. In this process, a speech recognition engine is used to convert the content of each participant's speech into text information.
[0173] Step 3:
[0174] The device inputs the converted text and audio data into an emotion analysis engine to analyze the user's emotional state. The input is the text data obtained in step 2 and the original audio data, and the output is information representing the user's emotional state. The emotion analysis engine infers emotions such as tension and relaxation based on audio attributes such as voice tone and speed.
[0175] Step 4:
[0176] The terminal sends the analyzed emotion data and text data to the server. The input is the emotion and text information obtained in step 3, and the output is the communication data that sends them to the server. The data is encrypted and securely relayed to the server.
[0177] Step 5:
[0178] The server uses a generative AI model based on the received data to generate questions that respond to the user's emotions. The input is the emotion and text information sent to the server in step 4, and the output is the text data of the generated questions. Specifically, if the emotional state is "tension," the server will generate questions that promote relaxation.
[0179] Step 6:
[0180] The server sends the generated question to the terminal, which then presents it to the user. The input is the question data generated in step 5, and the output is the information presented to the user. The presentation is done via a display, allowing the user to directly visually confirm the question.
[0181] Step 7:
[0182] The user reviews the presented questions and decides on their next action based on them. The input is the question displayed on the device, and the output is the user's response or answer. Based on the information provided, the user can engage in more effective discussions.
[0183] (Application Example 2)
[0184] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0185] Maintaining a harmonious atmosphere during meetings and discussions at home and in the office presents challenges, particularly in considering the emotional states of participants. There is a need for means to support smooth communication when tension or conflict arises.
[0186] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0187] In this invention, the server includes an information gathering means for collecting voice information, a conversion means for converting the voice information into text data, and an analysis and generation means for analyzing the text data and voice characteristics to generate emotion-based dialogue. This enables the provision of optimal dialogue in real time that responds to the emotions of the participants, easing the atmosphere of the conversation and facilitating smoother communication.
[0188] "Auditory information" refers to the data of sounds emitted during conversations and dialogues, which are recorded as audio signals.
[0189] "Information gathering means" refers to a device or process that has the function of acquiring voice information and transmitting it to a system.
[0190] "Text data" refers to data in text format that has been converted from audio information, and is treated as text information.
[0191] "Conversion means" refers to a process or device that analyzes audio information and converts it into text data.
[0192] "Speech characteristics" refer to distinctive elements contained in speech information, such as pitch, tone, speed, and intonation.
[0193] "Analysis and generation means" refers to a process or device that analyzes the user's emotions based on text data and voice characteristics and generates appropriate dialogue.
[0194] An "information user" is someone who receives the analyzed and generated information, or someone who utilizes that information.
[0195] An "automated learning method" is a process or device that uses machine learning techniques to improve the performance of a system based on empirically obtained data.
[0196] "Immediately" refers to an action that takes place in real time without any delay.
[0197] The system implementing this invention includes voice collection means, conversion to text data means, analysis and generation means, information presentation means, and automatic learning means. First, voice information of participants is collected in real time by terminals placed in homes or offices. The hardware used includes a high-sensitivity microphone and a voice recognition device. The voice collection means collects voice information and transmits it to a server.
[0198] Next, the server uses a conversion mechanism to convert the received audio information into text data. Speech recognition software (such as the Google Speech Recognition API) is utilized here. The resulting text data is then analyzed by an analysis and generation mechanism to determine the user's emotional state. The elements analyzed include speech characteristics such as voice tone and pitch. Emotion analysis software (such as IBM's Tone Analyzer or similar systems) is used in this process.
[0199] Based on the analyzed emotional state, the server generates appropriate questions and suggestions. This uses a generative AI model to create dialogue that takes into account the user's current emotional state. The generated dialogue is presented to the information user via the terminal. Presentation methods include displays and voice assistants.
[0200] For example, if analysis indicates that one participant is emotionally tense during a family discussion, a prompt such as, "Do you have any ideas for how we can relax and continue the discussion?" might be generated. This helps to alleviate tension and promote a more harmonious discussion.
[0201] Examples of prompts to input into a generative AI model are as follows:
[0202] "Audio data: "Hello. I'd like to discuss recent projects." Text data: "Hello. I'd like to discuss recent projects." Based on this, and analyzing the user's emotions, please consider what questions would be appropriate. Please suggest three questions that would help the user relax."
[0203] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0204] Step 1:
[0205] The device uses a high-sensitivity microphone to collect participants' voice information in real time during meetings and discussions. The input is raw voice data, and the output is a voice sample. This voice sample is then converted into a format that can be processed by speech recognition software.
[0206] Step 2:
[0207] The terminal sends the collected audio information to the server. The server uses speech recognition software to convert the audio information into text data. The input is an audio sample, and the output is text data. Through this conversion process, the conversation content is obtained as text.
[0208] Step 3:
[0209] The server analyzes the user's emotional state using analysis and generation methods based on acquired text data and speech characteristics (pitch, tone, etc.). The input is text data and speech characteristics, and the output is the result of the emotional analysis. Emotional analysis software is used to estimate the user's emotional state (e.g., relaxed, tense).
[0210] Step 4:
[0211] The server uses a generative AI model to generate appropriate dialogue based on the results of sentiment analysis. In this generation process, the sentiment analysis results are used as input, and questions and responses to the user are generated as output. The generated questions are tailored to the user's emotional state.
[0212] Step 5:
[0213] The server sends the generated dialogue to the terminal, which then presents it to the information user using a presentation method. The input is the generated question, and the output is a visual or auditory presentation to the information user. This allows the user to receive information that is sensitive to their emotions.
[0214] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0215] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0216] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0217] [Second Embodiment]
[0218] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0219] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0220] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0221] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0222] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0223] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0224] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0225] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0226] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0227] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0228] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0229] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0230] To implement this invention, a system is constructed in which three main entities—a terminal, a server, and a user—work in cooperation. First, the terminal uses a microphone at the meeting to collect participants' voices in real time. On the terminal, the voice collection means plays the role of converting the voice data into text. In this conversion, the terminal uses speech recognition technology to accurately transcribe what it hears into text.
[0231] Next, the terminal sends the converted text to the server. On the server, an analysis and generation system receives the text data, analyzes it using natural language processing algorithms, and extracts important keywords and topics. Based on these analysis results, the server uses machine learning techniques to generate relevant questions that are appropriate to the context of the conversation.
[0232] The generated questions are sent from the server to the terminal and presented to the user visually or audibly through a presentation method. For example, if the meeting is about the development of a new product, specific questions such as "Where is the target market?" or "What are the challenges during development?" are generated and presented. By using these questions as a reference during the meeting, users can ensure that all important points of the discussion are covered and that no details are overlooked.
[0233] In this way, the system of the present invention supports important confirmations in meetings and enables the efficient promotion of discussions. Furthermore, by continuously learning and improving the server-side AI model using user feedback, it becomes possible to generate even more accurate questions. This system is expected to improve the quality of meetings and the efficiency of work.
[0234] The following describes the processing flow.
[0235] Step 1:
[0236] The device activates the microphone at the start of the meeting and collects participants' voices in real time. The voice data is temporarily stored in a buffer to prepare for the next processing.
[0237] Step 2:
[0238] The device uses speech recognition technology to convert collected audio data into text data. This text is grammatically formatted, and the meeting content is stored in an easily understandable format.
[0239] Step 3:
[0240] The terminal sends formatted text data to the server. The transmission takes place over the network and is done in streaming format rather than batch processing to enable real-time data processing.
[0241] Step 4:
[0242] The server analyzes the received text data using analysis and generation tools. Natural language processing (NLP) is used to extract keywords and classify topics. This analysis identifies key points in the conversation.
[0243] Step 5:
[0244] Based on the analysis results, the server uses machine learning algorithms to generate relevant questions. This process takes into account past conversation patterns and contextual data to create questions that are useful and appropriate for the user.
[0245] Step 6:
[0246] The server sends the generated question to the terminal. The sent question is displayed on the terminal's user interface and presented to the user visually or audibly.
[0247] Step 7:
[0248] Users review the presented questions and make adjustments based on important confirmations and discussions as the meeting progresses. If necessary, users can improve the quality of the meeting by adding or refining questions.
[0249] Step 8:
[0250] The server collects user feedback and uses it to refine and train the AI model. This feedback helps improve the accuracy of question generation for the next meeting.
[0251] (Example 1)
[0252] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0253] In meetings and discussions, participants often miss important points, making it difficult to efficiently and effectively grasp the meeting content. Furthermore, there is a need to generate relevant questions to support meeting progress and improve the quality of meetings.
[0254] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0255] In this invention, the server includes an information gathering means for acquiring audio information, a conversion means for converting audio information into text information, an analysis and generation means for analyzing text information and generating related queries, and a learning means for collecting feedback from users and improving the analysis and generation means. This makes it possible to efficiently and effectively support the progress of meetings without missing important information during meetings, and to improve the quality of meetings.
[0256] "Audio information" refers to the waveforms of sounds emitted by participants during meetings and discussions, captured as digital data.
[0257] "Information gathering means" refers to devices and software installed to acquire audio information.
[0258] "Textual information" refers to digital data obtained by converting audio information into text or written characters.
[0259] "Conversion means" refers to a technology or device for converting audio information into text information, and utilizes speech recognition technology.
[0260] "Analysis and generation means" refers to technologies and processes for analyzing textual information and generating related queries based on that analysis.
[0261] "Users" refer to individuals or organizations that use the system and are responsible for receiving inquiries and facilitating the meeting.
[0262] "Presentation means" refers to methods or devices for informing the user of a generated inquiry, and includes visual or auditory methods.
[0263] "Feedback" refers to user reactions and evaluations of a system, and is information used to improve the system.
[0264] "Learning methods" refer to techniques and methods for continuously improving analysis and generation methods using feedback.
[0265] To implement this invention, three entities—a terminal, a server, and a user—must work together in coordination. In particular, this system is designed to support important discussions in meetings and improve the quality of those meetings.
[0266] Collection and conversion of audio information
[0267] The device uses a microphone during the meeting to collect audio information in real time. This information is stored as digital data and converted using speech recognition technology. Specifically, it uses speech recognition services such as the Google Cloud Speech-to-Text API and Microsoft Azure Speech Service. Through this process, the audio information is instantly converted into text.
[0268] Text information analysis and question generation
[0269] The terminal sends text information to the server, where analysis begins. The server can use libraries such as NLTK and SpaCy for natural language processing analysis. This analysis extracts important keywords and topics from the text information. Then, using generative AI models like GPT and BERT, contextual questions are generated based on the text information. These generated questions are highly relevant and designed to further deepen the discussion.
[0270] Question formulation and improvement
[0271] The generated questions are sent from the server to the terminal and presented to the user via the terminal's display or speech synthesis tool. This presentation allows the user to receive the questions visually or audibly, facilitating smoother meeting progress. Furthermore, users provide feedback after the meeting. This feedback information is stored on the server and used to train the generating AI model. This improves the accuracy of question generation, making support in future meetings more effective.
[0272] For example, in a meeting about a new product, participants' statements are recognized by speech recognition, and keywords such as "target market" and "competitor analysis" are extracted. Based on this, the server generates specific questions such as "How will we differentiate ourselves from competitors?" and presents them to the user via their terminal.
[0273] An example of a prompt message would be: "Listen to the discussion about new product development in the meeting and generate relevant questions, specifically about the target market and the challenges under development."
[0274] As described above, this system can support the efficient and effective conduct of meetings, thereby improving work efficiency and the quality of discussions.
[0275] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0276] Step 1:
[0277] Collection of Voice Information
[0278] The terminal uses a microphone at the meeting place to collect the voice information of participants in real time. In this step, the voice waveform is input into the terminal as digital data. The voice collection software on the terminal side operates to appropriately record the voice waveform for the purpose of subsequent conversion steps. This digital data is used in the subsequent conversion steps.
[0279] Step 2:
[0280] Conversion of Voice to Character Information
[0281] The terminal starts the process of converting the collected voice information into character information. As input, it processes the voice digital data obtained in Step 1. Specifically, it uses voice recognition technology and utilizes Google Cloud Speech-to-Text API, Microsoft Azure Speech Service, etc. As a result, the voice waveform is output as character information in the form of words and sentences. This character information is used in the next analysis step.
[0282] Step 3:
[0283] Transfer and Analysis of Character Information
[0284] The terminal securely transfers the converted character information to the server. As input, it is the character information generated in Step 2. The server receives this character information and analyzes it using natural language processing algorithms such as Python's NLTK and SpaCy. The purpose of the analysis is to extract important keywords and topics from the text data. The output of this analysis is useful in the subsequent question generation step.
[0285] Step 4:
[0286] Generation of Questions
[0287] Based on the analysis results from step 3, the server generates questions using a generative AI model. This process utilizes generative AI models such as GPT and BERT. The input consists of keywords and contextual information extracted through the analysis. The output is a list of highly relevant questions, which are then used to further the meeting.
[0288] Step 5:
[0289] Question presentation
[0290] The server sends the generated questions to the terminal. The terminal presents this list of questions to the user via a display or speech synthesis tool. The input is the list of questions generated in step 4. By receiving these questions, the user can effectively conduct the meeting. The output is the visual or auditory presentation of questions that the user uses.
[0291] Step 6:
[0292] Gathering feedback and improving the model
[0293] After the meeting, users provide feedback to the system. The server receives this feedback as input and incorporates it into the training data for the generative AI model. This enables more accurate question generation in the next meeting. The output is the improved generative AI model.
[0294] (Application Example 1)
[0295] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0296] There is a need to improve the efficiency of customer service in stores and to provide information more quickly and accurately. However, currently, there is insufficient support to accurately respond to individual customer questions, which places a heavy burden on staff and limits the improvement of customer satisfaction. In addition, the automatic generation of relevant information based on the content of the conversation is often not performed properly, hindering the smooth flow of conversation.
[0297] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0298] In this invention, the server includes an acquisition means for acquiring acoustic data, a conversion means for converting acoustic data into text information, and an analysis and generation means for analyzing the text information and generating related information. This enables store staff to instantly acquire accurate and relevant information during interactions with customers, allowing for efficient customer service.
[0299] "Audio data" refers to a data format that digitizes audio or sound information.
[0300] "Means of acquisition" refers to the devices and methods used to collect acoustic data.
[0301] "Conversion means" refers to a technical process for converting acoustic data into textual information.
[0302] "Textual information" refers to information expressed in text format.
[0303] "Analysis and generation means" refers to a method of analyzing textual information and automatically generating related information based on that analysis.
[0304] "Presentation means" refers to a function for presenting generated information to the user visually or audibly.
[0305] "User" refers to a person or group that uses the means of presentation.
[0306] "Recommended information" refers to additional information that is presumed to be useful to the user.
[0307] "Dialogue data" refers to the content of conversations recorded in past voice or text.
[0308] "Learning means" refers to machine learning techniques for analyzing dialogue data and optimizing related information.
[0309] "Instantaneous" means that the operation is performed immediately without delay.
[0310] To implement this invention, first, a system for acquiring acoustic data is constructed. The terminal acquires acoustic data and converts it into character information using conversion means. As the software to be used here, the "speech_recognition" library for performing speech recognition is suitable. The converted character information is transmitted to the server.
[0311] The server analyzes the received character information by analysis / generation means and generates related information. In this process, natural language processing technology using "transformers" is used. As a result, it is possible to immediately generate relevant recommended information based on the dialogue with the customer. The information obtained by the analysis is transmitted to the terminal again and displayed to the user by presentation means. The user can smoothly proceed with the dialogue with the customer based on this information.
[0312] With this system, when a store staff is asked, for example, "What are the features of the new product?" by a customer, they can immediately obtain relevant features and recommended usage methods and conduct appropriate communication. As described above, the system of the present invention can efficiently process the process from the acquisition of acoustic data to analysis and information presentation.
[0313] The generative AI model generates relevant information based on conversations in various contexts. For example, a prompt might read, "Generate a revised question to quickly obtain detailed information about the products or services offered. Please also provide information that can help improve the question."
[0314] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0315] Step 1:
[0316] The device uses its built-in microphone to capture audio during conversations with customers. It receives an acoustic signal as input and records it as digital audio data. Specifically, the device launches a speech recognition application in the background and performs audio capture.
[0317] Step 2:
[0318] The device converts acquired audio data into text. It uses audio data as input and generates text information through a conversion mechanism. In this process, a speech recognition library is used to convert speech to text. Specifically, it analyzes the audio waveform based on a language model and outputs the corresponding words as characters.
[0319] Step 3:
[0320] The server analyzes text data sent from the terminal. It receives text data as input and uses analysis and generation tools to extract important keywords and related information. Here, natural language processing techniques are used to analyze the text content and generate highly relevant information. Specifically, it extracts frequently occurring words from the received text while understanding the context and participating in the generation of recommended information.
[0321] Step 4:
[0322] The server generates relevant information based on the analysis results and sends it to the terminal. It generates recommended information as output and sends it back to the terminal via a presentation mechanism. Specifically, the process involves structuring the generated questions and information and conveying them to the terminal using a communication protocol.
[0323] Step 5:
[0324] The terminal presents relevant information received from the server to the user visually or audibly. It utilizes the received recommendation information as input and presents it via screen display or audio output. Specifically, the received information is delivered to the user in an easy-to-understand manner via a GUI or voice assistant.
[0325] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0326] For this invention to be implemented, the interaction between the terminal, server, and user is crucial. First, the terminal incorporates a voice collection means and an emotion engine. The voice collection means collects participants' voices in real time during a meeting and converts them into text. The emotion engine analyzes the user's emotions from this voice and text data. Specifically, it detects changes in voice tone and speaking style to infer the user's emotional state, such as whether they are tense or relaxed.
[0327] Next, the terminal sends the converted text data and sentiment analysis results to the server. The server analyzes the text and sentiment data using analysis and generation tools and generates questions that take into account the user's current emotional state. In this process, to support emotion-based communication, for example, if the user is feeling anxious, questions that encourage relaxation can be generated. Furthermore, the sentiment analysis results can be compared with past data to enable more personalized questions.
[0328] The generated questions are sent from the server to the terminal and presented to the user through a presentation mechanism. This allows the user to review the questions provided during the meeting and obtain optimal guidance that takes their emotional state into consideration. For example, if the emotional engine detects increased tension during a meeting to discuss the design of a new product, a question such as, "What kind of environment do you think is needed to think about this point more relaxed?" might be provided, enabling a discussion that is also supported emotionally.
[0329] As a result, the system of the present invention supports important discussions while considering the emotional balance in meetings, and improves the efficiency and quality of communication. In this way, emotionally considerate conversations are realized, providing a more productive environment.
[0330] The following describes the processing flow.
[0331] Step 1:
[0332] The device activates its audio collection system at the start of the meeting, collecting the voices of meeting participants in real time. Simultaneously, an emotion engine analyzes changes in voice tone and speaking style to infer the emotional state of the participants.
[0333] Step 2:
[0334] The device converts the collected audio into text. Using speech recognition technology, grammatically correct text is generated from the audio data. During this process, the emotion engine continuously monitors the audio data and updates the emotional state.
[0335] Step 3:
[0336] The terminal sends the converted text and sentiment analysis results to the server. The transmission is done in real time, and the system is designed to allow data processing in line with the progress of the meeting.
[0337] Step 4:
[0338] The server analyzes the received text and sentiment data. Analysis and generation tools analyze the text content and extract important topics. In addition, sentiment data is used to design questions that are appropriate to the user's current emotions.
[0339] Step 5:
[0340] The server uses a machine learning model to generate adaptive questions by comparing them with past data. For example, the generated questions might be designed to reassure a user who is in an unstable state.
[0341] Step 6:
[0342] The server sends the generated questions to the terminal and presents them to the user visually or audibly through a presentation mechanism. The user reviews these questions and responds with consideration for emotions as the meeting discussion progresses.
[0343] Step 7:
[0344] Users utilize the presented questions to continue the discussion appropriately within the meeting. Furthermore, they can modify the conversation or add new questions based on their own emotional state if necessary.
[0345] Step 8:
[0346] The server collects user feedback after the meeting and uses it to update and improve the AI model. This feedback is used to improve the accuracy of question generation in the next meeting.
[0347] (Example 2)
[0348] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0349] Traditional meeting systems simply convert audio data into text, lacking communication that takes into account the emotional state of the speaker. As a result, there was insufficient support to reduce tension and stress among meeting participants, leading to problems with the efficiency and quality of communication.
[0350] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0351] In this invention, the server includes means for converting voice data into language data, analysis means for analyzing the language data and voice data and inferring the user's emotional state, and generation means for generating questions to support conversation based on the user's emotional state. This enables communication support tailored to the user's emotional state.
[0352] "Audio data" refers to information used to record or transmit audio.
[0353] A "device" is a component of hardware or software designed to perform a specific function or purpose.
[0354] "Linguistic data" refers to information expressed in natural language, which is usually stored or transmitted in text format.
[0355] An "analysis device" is a device used to examine data and extract the meaning and information behind it.
[0356] "User" refers to an individual or group that operates a system or device and benefits from it.
[0357] "Emotional state" refers to information that indicates the speaker's psychological state or mood, and is usually inferred from the tone of voice and facial expressions.
[0358] A "generation device" is a device that has the function of creating new information or content based on input data.
[0359] A "machine learning system" is a computer system that uses data to automatically learn and execute algorithms that optimize a specific task.
[0360] "Instantly" refers to processing in near real-time, minimizing delays.
[0361] This invention is a system for improving the efficiency and quality of communication during meetings. Specifically, it is a system that collects speaker voice data in real time, analyzes that data to understand the speaker's emotional state, and generates customized questions based on that emotion.
[0362] The device is equipped with a microphone and speaker to collect participants' voices in real time. Once this voice data is collected, it is converted into text data using speech recognition software on the device. Generally, commercially available speech recognition APIs are used for speech recognition.
[0363] Next, the terminal performs sentiment analysis using the converted text data. Here, a sentiment analysis engine is used to analyze changes in voice tone and speaking rhythm. This engine utilizes generally known sentiment analysis tools.
[0364] The analyzed data is sent from the terminal to the server. The server inputs this data into a generative AI model to generate questions that take into account the user's emotional state. A generative AI model is a program that analyzes data and evolutionarily generates the optimal response. For example, machine learning frameworks such as TensorFlow are often used.
[0365] The questions generated on the server are sent back to the terminal and displayed to the user by a presentation device. The presentation device includes displays and visual media to clearly display the questions visually and aid in user comprehension.
[0366] As a concrete example, when discussing a new project in a meeting, if the emotion engine detects tension among participants, it will generate a question such as, "What kind of environment do you think is needed to help us think about this more relaxed?" This allows for a conversation that encourages participants to relax, leading to a more lively discussion. An example of a prompt would be, "The user's emotional state is tense. Please generate appropriate questions to support a conversation that takes this situation into consideration." This specific content would be used as input to the generating AI model.
[0367] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0368] Step 1:
[0369] The device collects participants' voices in real time. The input is raw audio data acquired using a microphone. Specifically, during a meeting, the device captures each participant's speech, resulting in digital audio data.
[0370] Step 2:
[0371] The device converts the collected audio data into text data using speech recognition software. The input is the audio data obtained in step 1, and the output is a string of text. In this process, a speech recognition engine is used to convert the content of each participant's speech into text information.
[0372] Step 3:
[0373] The device inputs the converted text and audio data into an emotion analysis engine to analyze the user's emotional state. The input is the text data obtained in step 2 and the original audio data, and the output is information representing the user's emotional state. The emotion analysis engine infers emotions such as tension and relaxation based on audio attributes such as voice tone and speed.
[0374] Step 4:
[0375] The terminal sends the analyzed emotion data and text data to the server. The input is the emotion and text information obtained in step 3, and the output is the communication data that sends them to the server. The data is encrypted and securely relayed to the server.
[0376] Step 5:
[0377] The server uses a generative AI model based on the received data to generate questions that respond to the user's emotions. The input is the emotion and text information sent to the server in step 4, and the output is the text data of the generated questions. Specifically, if the emotional state is "tension," the server will generate questions that promote relaxation.
[0378] Step 6:
[0379] The server sends the generated question to the terminal, which then presents it to the user. The input is the question data generated in step 5, and the output is the information presented to the user. The presentation is done via a display, allowing the user to directly visually confirm the question.
[0380] Step 7:
[0381] The user reviews the presented questions and decides on their next action based on them. The input is the question displayed on the device, and the output is the user's response or answer. Based on the information provided, the user can engage in more effective discussions.
[0382] (Application Example 2)
[0383] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0384] Maintaining a harmonious atmosphere during meetings and discussions at home and in the office presents challenges, particularly in considering the emotional states of participants. There is a need for means to support smooth communication when tension or conflict arises.
[0385] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0386] In this invention, the server includes an information gathering means for collecting voice information, a conversion means for converting the voice information into text data, and an analysis and generation means for analyzing the text data and voice characteristics to generate emotion-based dialogue. This enables the provision of optimal dialogue in real time that responds to the emotions of the participants, easing the atmosphere of the conversation and facilitating smoother communication.
[0387] "Auditory information" refers to the data of sounds emitted during conversations and dialogues, which are recorded as audio signals.
[0388] "Information gathering means" refers to a device or process that has the function of acquiring voice information and transmitting it to a system.
[0389] "Text data" refers to data in text format that has been converted from audio information, and is treated as text information.
[0390] "Conversion means" refers to a process or device that analyzes audio information and converts it into text data.
[0391] "Speech characteristics" refer to distinctive elements contained in speech information, such as pitch, tone, speed, and intonation.
[0392] "Analysis and generation means" refers to a process or device that analyzes the user's emotions based on text data and voice characteristics and generates appropriate dialogue.
[0393] An "information user" is someone who receives the analyzed and generated information, or someone who utilizes that information.
[0394] An "automated learning method" is a process or device that uses machine learning techniques to improve the performance of a system based on empirically obtained data.
[0395] "Immediately" refers to an action that takes place in real time without any delay.
[0396] The system implementing this invention includes voice collection means, conversion to text data means, analysis and generation means, information presentation means, and automatic learning means. First, voice information of participants is collected in real time by terminals placed in homes or offices. The hardware used includes a high-sensitivity microphone and a voice recognition device. The voice collection means collects voice information and transmits it to a server.
[0397] Next, the server uses a conversion mechanism to convert the received audio information into text data. Speech recognition software (such as the Google Speech Recognition API) is utilized here. The resulting text data is then analyzed by an analysis and generation mechanism to determine the user's emotional state. The elements analyzed include speech characteristics such as voice tone and pitch. Emotion analysis software (such as IBM's Tone Analyzer or similar systems) is used in this process.
[0398] Based on the analyzed emotional state, the server generates appropriate questions and suggestions. This uses a generative AI model to create dialogue that takes into account the user's current emotional state. The generated dialogue is presented to the information user via the terminal. Presentation methods include displays and voice assistants.
[0399] For example, if analysis indicates that one participant is emotionally tense during a family discussion, a prompt such as, "Do you have any ideas for how we can relax and continue the discussion?" might be generated. This helps to alleviate tension and promote a more harmonious discussion.
[0400] Examples of prompts to input into a generative AI model are as follows:
[0401] "Audio data: "Hello. I'd like to discuss recent projects." Text data: "Hello. I'd like to discuss recent projects." Based on this, and analyzing the user's emotions, please consider what questions would be appropriate. Please suggest three questions that would help the user relax."
[0402] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0403] Step 1:
[0404] The device uses a high-sensitivity microphone to collect participants' voice information in real time during meetings and discussions. The input is raw voice data, and the output is a voice sample. This voice sample is then converted into a format that can be processed by speech recognition software.
[0405] Step 2:
[0406] The terminal sends the collected audio information to the server. The server uses speech recognition software to convert the audio information into text data. The input is an audio sample, and the output is text data. Through this conversion process, the conversation content is obtained as text.
[0407] Step 3:
[0408] The server analyzes the user's emotional state using analysis and generation methods based on acquired text data and speech characteristics (pitch, tone, etc.). The input is text data and speech characteristics, and the output is the result of the emotional analysis. Emotional analysis software is used to estimate the user's emotional state (e.g., relaxed, tense).
[0409] Step 4:
[0410] The server uses a generative AI model to generate appropriate dialogue based on the results of sentiment analysis. In this generation process, the sentiment analysis results are used as input, and questions and responses to the user are generated as output. The generated questions are tailored to the user's emotional state.
[0411] Step 5:
[0412] The server sends the generated dialogue to the terminal, which then presents it to the information user using a presentation method. The input is the generated question, and the output is a visual or auditory presentation to the information user. This allows the user to receive information that is sensitive to their emotions.
[0413] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0414] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0415] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0416] [Third Embodiment]
[0417] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0418] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0419] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0420] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0421] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0422] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0423] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0424] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0425] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0426] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0427] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0428] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0429] To implement this invention, a system is constructed in which three main entities—a terminal, a server, and a user—work in cooperation. First, the terminal uses a microphone at the meeting to collect participants' voices in real time. On the terminal, the voice collection means plays the role of converting the voice data into text. In this conversion, the terminal uses speech recognition technology to accurately transcribe what it hears into text.
[0430] Next, the terminal sends the converted text to the server. On the server, an analysis and generation system receives the text data, analyzes it using natural language processing algorithms, and extracts important keywords and topics. Based on these analysis results, the server uses machine learning techniques to generate relevant questions that are appropriate to the context of the conversation.
[0431] The generated questions are sent from the server to the terminal and presented to the user visually or audibly through a presentation method. For example, if the meeting is about the development of a new product, specific questions such as "Where is the target market?" or "What are the challenges during development?" are generated and presented. By using these questions as a reference during the meeting, users can ensure that all important points of the discussion are covered and that no details are overlooked.
[0432] In this way, the system of the present invention supports important confirmations in meetings and enables the efficient promotion of discussions. Furthermore, by continuously learning and improving the server-side AI model using user feedback, it becomes possible to generate even more accurate questions. This system is expected to improve the quality of meetings and the efficiency of work.
[0433] The following describes the processing flow.
[0434] Step 1:
[0435] The device activates the microphone at the start of the meeting and collects participants' voices in real time. The voice data is temporarily stored in a buffer to prepare for the next processing.
[0436] Step 2:
[0437] The device uses speech recognition technology to convert collected audio data into text data. This text is grammatically formatted, and the meeting content is stored in an easily understandable format.
[0438] Step 3:
[0439] The terminal sends formatted text data to the server. The transmission takes place over the network and is done in streaming format rather than batch processing to enable real-time data processing.
[0440] Step 4:
[0441] The server analyzes the received text data using analysis and generation tools. Natural language processing (NLP) is used to extract keywords and classify topics. This analysis identifies key points in the conversation.
[0442] Step 5:
[0443] Based on the analysis results, the server uses machine learning algorithms to generate relevant questions. This process takes into account past conversation patterns and contextual data to create questions that are useful and appropriate for the user.
[0444] Step 6:
[0445] The server sends the generated question to the terminal. The sent question is displayed on the terminal's user interface and presented to the user visually or audibly.
[0446] Step 7:
[0447] Users review the presented questions and make adjustments based on important confirmations and discussions as the meeting progresses. If necessary, users can improve the quality of the meeting by adding or refining questions.
[0448] Step 8:
[0449] The server collects user feedback and uses it to refine and train the AI model. This feedback helps improve the accuracy of question generation for the next meeting.
[0450] (Example 1)
[0451] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0452] In meetings and discussions, participants often miss important points, making it difficult to efficiently and effectively grasp the meeting content. Furthermore, there is a need to generate relevant questions to support meeting progress and improve the quality of meetings.
[0453] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0454] In this invention, the server includes an information gathering means for acquiring audio information, a conversion means for converting audio information into text information, an analysis and generation means for analyzing text information and generating related queries, and a learning means for collecting feedback from users and improving the analysis and generation means. This makes it possible to efficiently and effectively support the progress of meetings without missing important information during meetings, and to improve the quality of meetings.
[0455] "Audio information" refers to the waveforms of sounds emitted by participants during meetings and discussions, captured as digital data.
[0456] "Information gathering means" refers to devices and software installed to acquire audio information.
[0457] "Textual information" refers to digital data obtained by converting audio information into text or written characters.
[0458] "Conversion means" refers to a technology or device for converting audio information into text information, and utilizes speech recognition technology.
[0459] "Analysis and generation means" refers to technologies and processes for analyzing textual information and generating related queries based on that analysis.
[0460] "Users" refer to individuals or organizations that use the system and are responsible for receiving inquiries and facilitating the meeting.
[0461] "Presentation means" refers to methods or devices for informing the user of a generated inquiry, and includes visual or auditory methods.
[0462] "Feedback" refers to user reactions and evaluations of a system, and is information used to improve the system.
[0463] "Learning methods" refer to techniques and methods for continuously improving analysis and generation methods using feedback.
[0464] To implement this invention, three entities—a terminal, a server, and a user—must work together in coordination. In particular, this system is designed to support important discussions in meetings and improve the quality of those meetings.
[0465] Collection and conversion of audio information
[0466] The device uses a microphone during the meeting to collect audio information in real time. This information is stored as digital data and converted using speech recognition technology. Specifically, it uses speech recognition services such as the Google Cloud Speech-to-Text API and Microsoft Azure Speech Service. Through this process, the audio information is instantly converted into text.
[0467] Text information analysis and question generation
[0468] The terminal sends text information to the server, where analysis begins. The server can use libraries such as NLTK and SpaCy for natural language processing analysis. This analysis extracts important keywords and topics from the text information. Then, using generative AI models like GPT and BERT, contextual questions are generated based on the text information. These generated questions are highly relevant and designed to further deepen the discussion.
[0469] Question formulation and improvement
[0470] The generated questions are sent from the server to the terminal and presented to the user via the terminal's display or speech synthesis tool. This presentation allows the user to receive the questions visually or audibly, facilitating smoother meeting progress. Furthermore, users provide feedback after the meeting. This feedback information is stored on the server and used to train the generating AI model. This improves the accuracy of question generation, making support in future meetings more effective.
[0471] For example, in a meeting about a new product, participants' statements are recognized by speech recognition, and keywords such as "target market" and "competitor analysis" are extracted. Based on this, the server generates specific questions such as "How will we differentiate ourselves from competitors?" and presents them to the user via their terminal.
[0472] An example of a prompt message would be: "Listen to the discussion about new product development in the meeting and generate relevant questions, specifically about the target market and the challenges under development."
[0473] As described above, this system can support the efficient and effective conduct of meetings, thereby improving work efficiency and the quality of discussions.
[0474] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0475] Step 1:
[0476] Collection of audio information
[0477] The terminal uses a microphone during the meeting to collect participants' voice information in real time. In this step, the voice waveform is input to the terminal as digital data. Voice collection software runs on the terminal to properly record the voice waveform. This digital data is used in a later conversion step.
[0478] Step 2:
[0479] Converting speech to text information
[0480] The terminal begins processing the collected audio information into text information. It processes the audio digital data obtained in step 1 as input. Specifically, it utilizes speech recognition technology, leveraging services such as the Google Cloud Speech-to-Text API and Microsoft Azure Speech Service. This outputs the audio waveform as text information in the form of words and sentences. This text information is then used in the next analysis step.
[0481] Step 3:
[0482] Text information transfer and analysis
[0483] The terminal securely transfers the converted character information to the server. The input is the character information generated in step 2. The server receives this character information and analyzes it using natural language processing algorithms such as Python's NLTK or SpaCy. The purpose of the analysis is to extract important keywords and topics from the text data. The output of this analysis is useful in the subsequent question generation step.
[0484] Step 4:
[0485] Question generation
[0486] Based on the analysis results from step 3, the server generates questions using a generative AI model. This process utilizes generative AI models such as GPT and BERT. The input consists of keywords and contextual information extracted through the analysis. The output is a list of highly relevant questions, which are then used to further the meeting.
[0487] Step 5:
[0488] Question presentation
[0489] The server sends the generated questions to the terminal. The terminal presents this list of questions to the user via a display or speech synthesis tool. The input is the list of questions generated in step 4. By receiving these questions, the user can effectively conduct the meeting. The output is the visual or auditory presentation of questions that the user uses.
[0490] Step 6:
[0491] Gathering feedback and improving the model
[0492] After the meeting, users provide feedback to the system. The server receives this feedback as input and incorporates it into the training data for the generative AI model. This enables more accurate question generation in the next meeting. The output is the improved generative AI model.
[0493] (Application Example 1)
[0494] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0495] There is a need to improve the efficiency of customer service in stores and to provide information more quickly and accurately. However, currently, there is insufficient support to accurately respond to individual customer questions, which places a heavy burden on staff and limits the improvement of customer satisfaction. In addition, the automatic generation of relevant information based on the content of the conversation is often not performed properly, hindering the smooth flow of conversation.
[0496] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0497] In this invention, the server includes an acquisition means for acquiring acoustic data, a conversion means for converting acoustic data into text information, and an analysis and generation means for analyzing the text information and generating related information. This enables store staff to instantly acquire accurate and relevant information during interactions with customers, allowing for efficient customer service.
[0498] "Audio data" refers to a data format that digitizes audio or sound information.
[0499] "Means of acquisition" refers to the devices and methods used to collect acoustic data.
[0500] "Conversion means" refers to a technical process for converting acoustic data into textual information.
[0501] "Textual information" refers to information expressed in text format.
[0502] "Analysis and generation means" refers to a method of analyzing textual information and automatically generating related information based on that analysis.
[0503] "Presentation means" refers to a function for presenting generated information to the user visually or audibly.
[0504] "User" refers to a person or group that uses the means of presentation.
[0505] "Recommended information" refers to additional information that is presumed to be useful to the user.
[0506] "Dialogue data" refers to the content of conversations recorded in past audio or text.
[0507] "Learning methods" refer to machine learning techniques used to analyze dialogue data and optimize relevant information.
[0508] "Immediate" means that an action is performed instantly, without delay.
[0509] To implement this invention, a system for acquiring acoustic data is first constructed. The terminal acquires the acoustic data and converts it into text information using a conversion means. The "speech_recognition" library, which performs speech recognition, is suitable as the software to be used here. The converted text information is sent to the server.
[0510] The server analyzes the received text information using parsing and generation tools and generates relevant information. Natural language processing technology using "transformers" is employed in this process. This makes it possible to instantly generate relevant recommendation information based on customer interaction. The information obtained through analysis is sent back to the terminal and displayed to the user by a presentation tool. Based on this information, the user can smoothly proceed with the interaction with the customer.
[0511] This system allows store staff, for example, when asked by a customer, "What are the features of the new product?", to instantly obtain relevant features and recommended uses, enabling appropriate communication. As described above, the system of the present invention can efficiently handle the process from voice data acquisition to analysis and information presentation.
[0512] The generative AI model generates relevant information based on conversations in various contexts. For example, a prompt might read, "Generate a revised question to quickly obtain detailed information about the products or services offered. Please also provide information that can help improve the question."
[0513] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0514] Step 1:
[0515] The device uses its built-in microphone to capture audio during conversations with customers. It receives an acoustic signal as input and records it as digital audio data. Specifically, the device launches a speech recognition application in the background and performs audio capture.
[0516] Step 2:
[0517] The device converts acquired audio data into text. It uses audio data as input and generates text information through a conversion mechanism. In this process, a speech recognition library is used to convert speech to text. Specifically, it analyzes the audio waveform based on a language model and outputs the corresponding words as characters.
[0518] Step 3:
[0519] The server analyzes text data sent from the terminal. It receives text data as input and uses analysis and generation tools to extract important keywords and related information. Here, natural language processing techniques are used to analyze the text content and generate highly relevant information. Specifically, it extracts frequently occurring words from the received text while understanding the context and participating in the generation of recommended information.
[0520] Step 4:
[0521] The server generates relevant information based on the analysis results and sends it to the terminal. It generates recommended information as output and sends it back to the terminal via a presentation mechanism. Specifically, the process involves structuring the generated questions and information and conveying them to the terminal using a communication protocol.
[0522] Step 5:
[0523] The terminal presents relevant information received from the server to the user visually or audibly. It utilizes the received recommendation information as input and presents it via screen display or audio output. Specifically, the received information is delivered to the user in an easy-to-understand manner via a GUI or voice assistant.
[0524] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0525] For this invention to be implemented, the interaction between the terminal, server, and user is crucial. First, the terminal incorporates a voice collection means and an emotion engine. The voice collection means collects participants' voices in real time during a meeting and converts them into text. The emotion engine analyzes the user's emotions from this voice and text data. Specifically, it detects changes in voice tone and speaking style to infer the user's emotional state, such as whether they are tense or relaxed.
[0526] Next, the terminal sends the converted text data and sentiment analysis results to the server. The server analyzes the text and sentiment data using analysis and generation tools and generates questions that take into account the user's current emotional state. In this process, to support emotion-based communication, for example, if the user is feeling anxious, questions that encourage relaxation can be generated. Furthermore, the sentiment analysis results can be compared with past data to enable more personalized questions.
[0527] The generated questions are sent from the server to the terminal and presented to the user through a presentation mechanism. This allows the user to review the questions provided during the meeting and obtain optimal guidance that takes their emotional state into consideration. For example, if the emotional engine detects increased tension during a meeting to discuss the design of a new product, a question such as, "What kind of environment do you think is needed to think about this point more relaxed?" might be provided, enabling a discussion that is also supported emotionally.
[0528] As a result, the system of the present invention supports important discussions while considering the emotional balance in meetings, and improves the efficiency and quality of communication. In this way, emotionally considerate conversations are realized, providing a more productive environment.
[0529] The following describes the processing flow.
[0530] Step 1:
[0531] The device activates its audio collection system at the start of the meeting, collecting the voices of meeting participants in real time. Simultaneously, an emotion engine analyzes changes in voice tone and speaking style to infer the emotional state of the participants.
[0532] Step 2:
[0533] The device converts the collected audio into text. Using speech recognition technology, grammatically correct text is generated from the audio data. During this process, the emotion engine continuously monitors the audio data and updates the emotional state.
[0534] Step 3:
[0535] The terminal sends the converted text and sentiment analysis results to the server. The transmission is done in real time, and the system is designed to allow data processing in line with the progress of the meeting.
[0536] Step 4:
[0537] The server analyzes the received text and sentiment data. Analysis and generation tools analyze the text content and extract important topics. In addition, sentiment data is used to design questions that are appropriate to the user's current emotions.
[0538] Step 5:
[0539] The server uses a machine learning model to generate adaptive questions by comparing them with past data. For example, the generated questions might be designed to reassure a user who is in an unstable state.
[0540] Step 6:
[0541] The server sends the generated questions to the terminal and presents them to the user visually or audibly through a presentation mechanism. The user reviews these questions and responds with consideration for emotions as the meeting discussion progresses.
[0542] Step 7:
[0543] Users utilize the presented questions to continue the discussion appropriately within the meeting. Furthermore, they can modify the conversation or add new questions based on their own emotional state if necessary.
[0544] Step 8:
[0545] The server collects user feedback after the meeting and uses it to update and improve the AI model. This feedback is used to improve the accuracy of question generation in the next meeting.
[0546] (Example 2)
[0547] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0548] Traditional meeting systems simply convert audio data into text, lacking communication that takes into account the emotional state of the speaker. As a result, there was insufficient support to reduce tension and stress among meeting participants, leading to problems with the efficiency and quality of communication.
[0549] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0550] In this invention, the server includes means for converting voice data into language data, analysis means for analyzing the language data and voice data and inferring the user's emotional state, and generation means for generating questions to support conversation based on the user's emotional state. This enables communication support tailored to the user's emotional state.
[0551] "Audio data" refers to information used to record or transmit audio.
[0552] A "device" is a component of hardware or software designed to perform a specific function or purpose.
[0553] "Linguistic data" refers to information expressed in natural language, which is usually stored or transmitted in text format.
[0554] An "analysis device" is a device used to examine data and extract the meaning and information behind it.
[0555] "User" refers to an individual or group that operates a system or device and benefits from it.
[0556] "Emotional state" refers to information that indicates the speaker's psychological state or mood, and is usually inferred from the tone of voice and facial expressions.
[0557] A "generation device" is a device that has the function of creating new information or content based on input data.
[0558] A "machine learning system" is a computer system that uses data to automatically learn and execute algorithms that optimize a specific task.
[0559] "Instantly" refers to processing in near real-time, minimizing delays.
[0560] This invention is a system for improving the efficiency and quality of communication during meetings. Specifically, it is a system that collects speaker voice data in real time, analyzes that data to understand the speaker's emotional state, and generates customized questions based on that emotion.
[0561] The device is equipped with a microphone and speaker to collect participants' voices in real time. Once this voice data is collected, it is converted into text data using speech recognition software on the device. Generally, commercially available speech recognition APIs are used for speech recognition.
[0562] Next, the terminal performs sentiment analysis using the converted text data. Here, a sentiment analysis engine is used to analyze changes in voice tone and speaking rhythm. This engine utilizes generally known sentiment analysis tools.
[0563] The analyzed data is sent from the terminal to the server. The server inputs this data into a generative AI model to generate questions that take into account the user's emotional state. A generative AI model is a program that analyzes data and evolutionarily generates the optimal response. For example, machine learning frameworks such as TensorFlow are often used.
[0564] The questions generated on the server are sent back to the terminal and displayed to the user by a presentation device. The presentation device includes displays and visual media to clearly display the questions visually and aid in user comprehension.
[0565] As a concrete example, when discussing a new project in a meeting, if the emotion engine detects tension among participants, it will generate a question such as, "What kind of environment do you think is needed to help us think about this more relaxed?" This allows for a conversation that encourages participants to relax, leading to a more lively discussion. An example of a prompt would be, "The user's emotional state is tense. Please generate appropriate questions to support a conversation that takes this situation into consideration." This specific content would be used as input to the generating AI model.
[0566] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0567] Step 1:
[0568] The device collects participants' voices in real time. The input is raw audio data acquired using a microphone. Specifically, during a meeting, the device captures each participant's speech, resulting in digital audio data.
[0569] Step 2:
[0570] The device converts the collected audio data into text data using speech recognition software. The input is the audio data obtained in step 1, and the output is a string of text. In this process, a speech recognition engine is used to convert the content of each participant's speech into text information.
[0571] Step 3:
[0572] The device inputs the converted text and audio data into an emotion analysis engine to analyze the user's emotional state. The input is the text data obtained in step 2 and the original audio data, and the output is information representing the user's emotional state. The emotion analysis engine infers emotions such as tension and relaxation based on audio attributes such as voice tone and speed.
[0573] Step 4:
[0574] The terminal sends the analyzed emotion data and text data to the server. The input is the emotion and text information obtained in step 3, and the output is the communication data that sends them to the server. The data is encrypted and securely relayed to the server.
[0575] Step 5:
[0576] The server uses a generative AI model based on the received data to generate questions that respond to the user's emotions. The input is the emotion and text information sent to the server in step 4, and the output is the text data of the generated questions. Specifically, if the emotional state is "tension," the server will generate questions that promote relaxation.
[0577] Step 6:
[0578] The server sends the generated question to the terminal, which then presents it to the user. The input is the question data generated in step 5, and the output is the information presented to the user. The presentation is done via a display, allowing the user to directly visually confirm the question.
[0579] Step 7:
[0580] The user reviews the presented questions and decides on their next action based on them. The input is the question displayed on the device, and the output is the user's response or answer. Based on the information provided, the user can engage in more effective discussions.
[0581] (Application Example 2)
[0582] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0583] Maintaining a harmonious atmosphere during meetings and discussions at home and in the office presents challenges, particularly in considering the emotional states of participants. There is a need for means to support smooth communication when tension or conflict arises.
[0584] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0585] In this invention, the server includes an information gathering means for collecting voice information, a conversion means for converting the voice information into text data, and an analysis and generation means for analyzing the text data and voice characteristics to generate emotion-based dialogue. This enables the provision of optimal dialogue in real time that responds to the emotions of the participants, easing the atmosphere of the conversation and facilitating smoother communication.
[0586] "Auditory information" refers to the data of sounds emitted during conversations and dialogues, which are recorded as audio signals.
[0587] "Information gathering means" refers to a device or process that has the function of acquiring voice information and transmitting it to a system.
[0588] "Text data" refers to data in text format that has been converted from audio information, and is treated as text information.
[0589] "Conversion means" refers to a process or device that analyzes audio information and converts it into text data.
[0590] "Speech characteristics" refer to distinctive elements contained in speech information, such as pitch, tone, speed, and intonation.
[0591] "Analysis and generation means" refers to a process or device that analyzes the user's emotions based on text data and voice characteristics and generates appropriate dialogue.
[0592] An "information user" is someone who receives the analyzed and generated information, or someone who utilizes that information.
[0593] An "automated learning method" is a process or device that uses machine learning techniques to improve the performance of a system based on empirically obtained data.
[0594] "Immediately" refers to an action that takes place in real time without any delay.
[0595] The system implementing this invention includes voice collection means, conversion to text data means, analysis and generation means, information presentation means, and automatic learning means. First, voice information of participants is collected in real time by terminals placed in homes or offices. The hardware used includes a high-sensitivity microphone and a voice recognition device. The voice collection means collects voice information and transmits it to a server.
[0596] Next, the server uses a conversion mechanism to convert the received audio information into text data. Speech recognition software (such as the Google Speech Recognition API) is utilized here. The resulting text data is then analyzed by an analysis and generation mechanism to determine the user's emotional state. The elements analyzed include speech characteristics such as voice tone and pitch. Emotion analysis software (such as IBM's Tone Analyzer or similar systems) is used in this process.
[0597] Based on the analyzed emotional state, the server generates appropriate questions and suggestions. This uses a generative AI model to create dialogue that takes into account the user's current emotional state. The generated dialogue is presented to the information user via the terminal. Presentation methods include displays and voice assistants.
[0598] For example, if analysis indicates that one participant is emotionally tense during a family discussion, a prompt such as, "Do you have any ideas for how we can relax and continue the discussion?" might be generated. This helps to alleviate tension and promote a more harmonious discussion.
[0599] Examples of prompts to input into a generative AI model are as follows:
[0600] "Audio data: "Hello. I'd like to discuss recent projects." Text data: "Hello. I'd like to discuss recent projects." Based on this, and analyzing the user's emotions, please consider what questions would be appropriate. Please suggest three questions that would help the user relax."
[0601] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0602] Step 1:
[0603] The device uses a high-sensitivity microphone to collect participants' voice information in real time during meetings and discussions. The input is raw voice data, and the output is a voice sample. This voice sample is then converted into a format that can be processed by speech recognition software.
[0604] Step 2:
[0605] The terminal sends the collected audio information to the server. The server uses speech recognition software to convert the audio information into text data. The input is an audio sample, and the output is text data. Through this conversion process, the conversation content is obtained as text.
[0606] Step 3:
[0607] The server analyzes the user's emotional state using analysis and generation methods based on acquired text data and speech characteristics (pitch, tone, etc.). The input is text data and speech characteristics, and the output is the result of the emotional analysis. Emotional analysis software is used to estimate the user's emotional state (e.g., relaxed, tense).
[0608] Step 4:
[0609] The server uses a generative AI model to generate appropriate dialogue based on the results of sentiment analysis. In this generation process, the sentiment analysis results are used as input, and questions and responses to the user are generated as output. The generated questions are tailored to the user's emotional state.
[0610] Step 5:
[0611] The server sends the generated dialogue to the terminal, which then presents it to the information user using a presentation method. The input is the generated question, and the output is a visual or auditory presentation to the information user. This allows the user to receive information that is sensitive to their emotions.
[0612] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0613] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0614] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0615] [Fourth Embodiment]
[0616] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0617] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0618] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0619] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0620] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0621] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0622] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0623] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0624] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0625] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0626] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0627] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0628] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0629] To implement this invention, a system is constructed in which three main entities—a terminal, a server, and a user—work in cooperation. First, the terminal uses a microphone at the meeting to collect participants' voices in real time. On the terminal, the voice collection means plays the role of converting the voice data into text. In this conversion, the terminal uses speech recognition technology to accurately transcribe what it hears into text.
[0630] Next, the terminal sends the converted text to the server. On the server, an analysis and generation system receives the text data, analyzes it using natural language processing algorithms, and extracts important keywords and topics. Based on these analysis results, the server uses machine learning techniques to generate relevant questions that are appropriate to the context of the conversation.
[0631] The generated questions are sent from the server to the terminal and presented to the user visually or audibly through a presentation method. For example, if the meeting is about the development of a new product, specific questions such as "Where is the target market?" or "What are the challenges during development?" are generated and presented. By using these questions as a reference during the meeting, users can ensure that all important points of the discussion are covered and that no details are overlooked.
[0632] In this way, the system of the present invention supports important confirmations in meetings and enables the efficient promotion of discussions. Furthermore, by continuously learning and improving the server-side AI model using user feedback, it becomes possible to generate even more accurate questions. This system is expected to improve the quality of meetings and the efficiency of work.
[0633] The following describes the processing flow.
[0634] Step 1:
[0635] The device activates the microphone at the start of the meeting and collects participants' voices in real time. The voice data is temporarily stored in a buffer to prepare for the next processing.
[0636] Step 2:
[0637] The device uses speech recognition technology to convert collected audio data into text data. This text is grammatically formatted, and the meeting content is stored in an easily understandable format.
[0638] Step 3:
[0639] The terminal sends formatted text data to the server. The transmission takes place over the network and is done in streaming format rather than batch processing to enable real-time data processing.
[0640] Step 4:
[0641] The server analyzes the received text data using analysis and generation tools. Natural language processing (NLP) is used to extract keywords and classify topics. This analysis identifies key points in the conversation.
[0642] Step 5:
[0643] Based on the analysis results, the server uses machine learning algorithms to generate relevant questions. This process takes into account past conversation patterns and contextual data to create questions that are useful and appropriate for the user.
[0644] Step 6:
[0645] The server sends the generated question to the terminal. The sent question is displayed on the terminal's user interface and presented to the user visually or audibly.
[0646] Step 7:
[0647] Users review the presented questions and make adjustments based on important confirmations and discussions as the meeting progresses. If necessary, users can improve the quality of the meeting by adding or refining questions.
[0648] Step 8:
[0649] The server collects user feedback and uses it to refine and train the AI model. This feedback helps improve the accuracy of question generation for the next meeting.
[0650] (Example 1)
[0651] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0652] In meetings and discussions, participants often miss important points, making it difficult to efficiently and effectively grasp the meeting content. Furthermore, there is a need to generate relevant questions to support meeting progress and improve the quality of meetings.
[0653] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0654] In this invention, the server includes an information gathering means for acquiring audio information, a conversion means for converting audio information into text information, an analysis and generation means for analyzing text information and generating related queries, and a learning means for collecting feedback from users and improving the analysis and generation means. This makes it possible to efficiently and effectively support the progress of meetings without missing important information during meetings, and to improve the quality of meetings.
[0655] "Audio information" refers to the waveforms of sounds emitted by participants during meetings and discussions, captured as digital data.
[0656] "Information gathering means" refers to devices and software installed to acquire audio information.
[0657] "Textual information" refers to digital data obtained by converting audio information into text or written characters.
[0658] "Conversion means" refers to a technology or device for converting audio information into text information, and utilizes speech recognition technology.
[0659] "Analysis and generation means" refers to technologies and processes for analyzing textual information and generating related queries based on that analysis.
[0660] "Users" refer to individuals or organizations that use the system and are responsible for receiving inquiries and facilitating the meeting.
[0661] "Presentation means" refers to methods or devices for informing the user of a generated inquiry, and includes visual or auditory methods.
[0662] "Feedback" refers to user reactions and evaluations of a system, and is information used to improve the system.
[0663] "Learning methods" refer to techniques and methods for continuously improving analysis and generation methods using feedback.
[0664] To implement this invention, three entities—a terminal, a server, and a user—must work together in coordination. In particular, this system is designed to support important discussions in meetings and improve the quality of those meetings.
[0665] Collection and conversion of audio information
[0666] The device uses a microphone during the meeting to collect audio information in real time. This information is stored as digital data and converted using speech recognition technology. Specifically, it uses speech recognition services such as the Google Cloud Speech-to-Text API and Microsoft Azure Speech Service. Through this process, the audio information is instantly converted into text.
[0667] Text information analysis and question generation
[0668] The terminal sends text information to the server, where analysis begins. The server can use libraries such as NLTK and SpaCy for natural language processing analysis. This analysis extracts important keywords and topics from the text information. Then, using generative AI models like GPT and BERT, contextual questions are generated based on the text information. These generated questions are highly relevant and designed to further deepen the discussion.
[0669] Question formulation and improvement
[0670] The generated questions are sent from the server to the terminal and presented to the user via the terminal's display or speech synthesis tool. This presentation allows the user to receive the questions visually or audibly, facilitating smoother meeting progress. Furthermore, users provide feedback after the meeting. This feedback information is stored on the server and used to train the generating AI model. This improves the accuracy of question generation, making support in future meetings more effective.
[0671] For example, in a meeting about a new product, participants' statements are recognized by speech recognition, and keywords such as "target market" and "competitor analysis" are extracted. Based on this, the server generates specific questions such as "How will we differentiate ourselves from competitors?" and presents them to the user via their terminal.
[0672] An example of a prompt message would be: "Listen to the discussion about new product development in the meeting and generate relevant questions, specifically about the target market and the challenges under development."
[0673] As described above, this system can support the efficient and effective conduct of meetings, thereby improving work efficiency and the quality of discussions.
[0674] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0675] Step 1:
[0676] Collection of audio information
[0677] The terminal uses a microphone during the meeting to collect participants' voice information in real time. In this step, the voice waveform is input to the terminal as digital data. Voice collection software runs on the terminal to properly record the voice waveform. This digital data is used in a later conversion step.
[0678] Step 2:
[0679] Converting speech to text information
[0680] The terminal begins processing the collected audio information into text information. It processes the audio digital data obtained in step 1 as input. Specifically, it utilizes speech recognition technology, leveraging services such as the Google Cloud Speech-to-Text API and Microsoft Azure Speech Service. This outputs the audio waveform as text information in the form of words and sentences. This text information is then used in the next analysis step.
[0681] Step 3:
[0682] Text information transfer and analysis
[0683] The terminal securely transfers the converted character information to the server. The input is the character information generated in step 2. The server receives this character information and analyzes it using natural language processing algorithms such as Python's NLTK or SpaCy. The purpose of the analysis is to extract important keywords and topics from the text data. The output of this analysis is useful in the subsequent question generation step.
[0684] Step 4:
[0685] Question generation
[0686] Based on the analysis results from step 3, the server generates questions using a generative AI model. This process utilizes generative AI models such as GPT and BERT. The input consists of keywords and contextual information extracted through the analysis. The output is a list of highly relevant questions, which are then used to further the meeting.
[0687] Step 5:
[0688] Question presentation
[0689] The server sends the generated questions to the terminal. The terminal presents this list of questions to the user via a display or speech synthesis tool. The input is the list of questions generated in step 4. By receiving these questions, the user can effectively conduct the meeting. The output is the visual or auditory presentation of questions that the user uses.
[0690] Step 6:
[0691] Gathering feedback and improving the model
[0692] After the meeting, users provide feedback to the system. The server receives this feedback as input and incorporates it into the training data for the generative AI model. This enables more accurate question generation in the next meeting. The output is the improved generative AI model.
[0693] (Application Example 1)
[0694] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0695] There is a need to improve the efficiency of customer service in stores and to provide information more quickly and accurately. However, currently, there is insufficient support to accurately respond to individual customer questions, which places a heavy burden on staff and limits the improvement of customer satisfaction. In addition, the automatic generation of relevant information based on the content of the conversation is often not performed properly, hindering the smooth flow of conversation.
[0696] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0697] In this invention, the server includes an acquisition means for acquiring acoustic data, a conversion means for converting acoustic data into text information, and an analysis and generation means for analyzing the text information and generating related information. This enables store staff to instantly acquire accurate and relevant information during interactions with customers, allowing for efficient customer service.
[0698] "Audio data" refers to a data format that digitizes audio or sound information.
[0699] "Means of acquisition" refers to the devices and methods used to collect acoustic data.
[0700] "Conversion means" refers to a technical process for converting acoustic data into textual information.
[0701] "Textual information" refers to information expressed in text format.
[0702] "Analysis and generation means" refers to a method of analyzing textual information and automatically generating related information based on that analysis.
[0703] "Presentation means" refers to a function for presenting generated information to the user visually or audibly.
[0704] "User" refers to a person or group that uses the means of presentation.
[0705] "Recommended information" refers to additional information that is presumed to be useful to the user.
[0706] "Dialogue data" refers to the content of conversations recorded in past audio or text.
[0707] "Learning methods" refer to machine learning techniques used to analyze dialogue data and optimize relevant information.
[0708] "Immediate" means that an action is performed instantly, without delay.
[0709] To implement this invention, a system for acquiring acoustic data is first constructed. The terminal acquires the acoustic data and converts it into text information using a conversion means. The "speech_recognition" library, which performs speech recognition, is suitable as the software to be used here. The converted text information is sent to the server.
[0710] The server analyzes the received text information using parsing and generation tools and generates relevant information. Natural language processing technology using "transformers" is employed in this process. This makes it possible to instantly generate relevant recommendation information based on customer interaction. The information obtained through analysis is sent back to the terminal and displayed to the user by a presentation tool. Based on this information, the user can smoothly proceed with the interaction with the customer.
[0711] This system allows store staff, for example, when asked by a customer, "What are the features of the new product?", to instantly obtain relevant features and recommended uses, enabling appropriate communication. As described above, the system of the present invention can efficiently handle the process from voice data acquisition to analysis and information presentation.
[0712] The generative AI model generates relevant information based on conversations in various contexts. For example, a prompt might read, "Generate a revised question to quickly obtain detailed information about the products or services offered. Please also provide information that can help improve the question."
[0713] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0714] Step 1:
[0715] The device uses its built-in microphone to capture audio during conversations with customers. It receives an acoustic signal as input and records it as digital audio data. Specifically, the device launches a speech recognition application in the background and performs audio capture.
[0716] Step 2:
[0717] The device converts acquired audio data into text. It uses audio data as input and generates text information through a conversion mechanism. In this process, a speech recognition library is used to convert speech to text. Specifically, it analyzes the audio waveform based on a language model and outputs the corresponding words as characters.
[0718] Step 3:
[0719] The server analyzes text data sent from the terminal. It receives text data as input and uses analysis and generation tools to extract important keywords and related information. Here, natural language processing techniques are used to analyze the text content and generate highly relevant information. Specifically, it extracts frequently occurring words from the received text while understanding the context and participating in the generation of recommended information.
[0720] Step 4:
[0721] The server generates relevant information based on the analysis results and sends it to the terminal. It generates recommended information as output and sends it back to the terminal via a presentation mechanism. Specifically, the process involves structuring the generated questions and information and conveying them to the terminal using a communication protocol.
[0722] Step 5:
[0723] The terminal presents relevant information received from the server to the user visually or audibly. It utilizes the received recommendation information as input and presents it via screen display or audio output. Specifically, the received information is delivered to the user in an easy-to-understand manner via a GUI or voice assistant.
[0724] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0725] For this invention to be implemented, the interaction between the terminal, server, and user is crucial. First, the terminal incorporates a voice collection means and an emotion engine. The voice collection means collects participants' voices in real time during a meeting and converts them into text. The emotion engine analyzes the user's emotions from this voice and text data. Specifically, it detects changes in voice tone and speaking style to infer the user's emotional state, such as whether they are tense or relaxed.
[0726] Next, the terminal sends the converted text data and sentiment analysis results to the server. The server analyzes the text and sentiment data using analysis and generation tools and generates questions that take into account the user's current emotional state. In this process, to support emotion-based communication, for example, if the user is feeling anxious, questions that encourage relaxation can be generated. Furthermore, the sentiment analysis results can be compared with past data to enable more personalized questions.
[0727] The generated questions are sent from the server to the terminal and presented to the user through a presentation mechanism. This allows the user to review the questions provided during the meeting and obtain optimal guidance that takes their emotional state into consideration. For example, if the emotional engine detects increased tension during a meeting to discuss the design of a new product, a question such as, "What kind of environment do you think is needed to think about this point more relaxed?" might be provided, enabling a discussion that is also supported emotionally.
[0728] As a result, the system of the present invention supports important discussions while considering the emotional balance in meetings, and improves the efficiency and quality of communication. In this way, emotionally considerate conversations are realized, providing a more productive environment.
[0729] The following describes the processing flow.
[0730] Step 1:
[0731] The device activates its audio collection system at the start of the meeting, collecting the voices of meeting participants in real time. Simultaneously, an emotion engine analyzes changes in voice tone and speaking style to infer the emotional state of the participants.
[0732] Step 2:
[0733] The device converts the collected audio into text. Using speech recognition technology, grammatically correct text is generated from the audio data. During this process, the emotion engine continuously monitors the audio data and updates the emotional state.
[0734] Step 3:
[0735] The terminal sends the converted text and sentiment analysis results to the server. The transmission is done in real time, and the system is designed to allow data processing in line with the progress of the meeting.
[0736] Step 4:
[0737] The server analyzes the received text and sentiment data. Analysis and generation tools analyze the text content and extract important topics. In addition, sentiment data is used to design questions that are appropriate to the user's current emotions.
[0738] Step 5:
[0739] The server uses a machine learning model to generate adaptive questions by comparing them with past data. For example, the generated questions might be designed to reassure a user who is in an unstable state.
[0740] Step 6:
[0741] The server sends the generated questions to the terminal and presents them to the user visually or audibly through a presentation mechanism. The user reviews these questions and responds with consideration for emotions as the meeting discussion progresses.
[0742] Step 7:
[0743] Users utilize the presented questions to continue the discussion appropriately within the meeting. Furthermore, they can modify the conversation or add new questions based on their own emotional state if necessary.
[0744] Step 8:
[0745] The server collects user feedback after the meeting and uses it to update and improve the AI model. This feedback is used to improve the accuracy of question generation in the next meeting.
[0746] (Example 2)
[0747] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0748] Traditional meeting systems simply convert audio data into text, lacking communication that takes into account the emotional state of the speaker. As a result, there was insufficient support to reduce tension and stress among meeting participants, leading to problems with the efficiency and quality of communication.
[0749] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0750] In this invention, the server includes means for converting voice data into language data, analysis means for analyzing the language data and voice data and inferring the user's emotional state, and generation means for generating questions to support conversation based on the user's emotional state. This enables communication support tailored to the user's emotional state.
[0751] "Audio data" refers to information used to record or transmit audio.
[0752] A "device" is a component of hardware or software designed to perform a specific function or purpose.
[0753] "Linguistic data" refers to information expressed in natural language, which is usually stored or transmitted in text format.
[0754] An "analysis device" is a device used to examine data and extract the meaning and information behind it.
[0755] "User" refers to an individual or group that operates a system or device and benefits from it.
[0756] "Emotional state" refers to information that indicates the speaker's psychological state or mood, and is usually inferred from the tone of voice and facial expressions.
[0757] A "generation device" is a device that has the function of creating new information or content based on input data.
[0758] A "machine learning system" is a computer system that uses data to automatically learn and execute algorithms that optimize a specific task.
[0759] "Instantly" refers to processing in near real-time, minimizing delays.
[0760] This invention is a system for improving the efficiency and quality of communication during meetings. Specifically, it is a system that collects speaker voice data in real time, analyzes that data to understand the speaker's emotional state, and generates customized questions based on that emotion.
[0761] The device is equipped with a microphone and speaker to collect participants' voices in real time. Once this voice data is collected, it is converted into text data using speech recognition software on the device. Generally, commercially available speech recognition APIs are used for speech recognition.
[0762] Next, the terminal performs sentiment analysis using the converted text data. Here, a sentiment analysis engine is used to analyze changes in voice tone and speaking rhythm. This engine utilizes generally known sentiment analysis tools.
[0763] The analyzed data is sent from the terminal to the server. The server inputs this data into a generative AI model to generate questions that take into account the user's emotional state. A generative AI model is a program that analyzes data and evolutionarily generates the optimal response. For example, machine learning frameworks such as TensorFlow are often used.
[0764] The questions generated on the server are sent back to the terminal and displayed to the user by a presentation device. The presentation device includes displays and visual media to clearly display the questions visually and aid in user comprehension.
[0765] As a concrete example, when discussing a new project in a meeting, if the emotion engine detects tension among participants, it will generate a question such as, "What kind of environment do you think is needed to help us think about this more relaxed?" This allows for a conversation that encourages participants to relax, leading to a more lively discussion. An example of a prompt would be, "The user's emotional state is tense. Please generate appropriate questions to support a conversation that takes this situation into consideration." This specific content would be used as input to the generating AI model.
[0766] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0767] Step 1:
[0768] The device collects participants' voices in real time. The input is raw audio data acquired using a microphone. Specifically, during a meeting, the device captures each participant's speech, resulting in digital audio data.
[0769] Step 2:
[0770] The device converts the collected audio data into text data using speech recognition software. The input is the audio data obtained in step 1, and the output is a string of text. In this process, a speech recognition engine is used to convert the content of each participant's speech into text information.
[0771] Step 3:
[0772] The device inputs the converted text and audio data into an emotion analysis engine to analyze the user's emotional state. The input is the text data obtained in step 2 and the original audio data, and the output is information representing the user's emotional state. The emotion analysis engine infers emotions such as tension and relaxation based on audio attributes such as voice tone and speed.
[0773] Step 4:
[0774] The terminal sends the analyzed emotion data and text data to the server. The input is the emotion and text information obtained in step 3, and the output is the communication data that sends them to the server. The data is encrypted and securely relayed to the server.
[0775] Step 5:
[0776] The server uses a generative AI model based on the received data to generate questions that respond to the user's emotions. The input is the emotion and text information sent to the server in step 4, and the output is the text data of the generated questions. Specifically, if the emotional state is "tension," the server will generate questions that promote relaxation.
[0777] Step 6:
[0778] The server sends the generated question to the terminal, which then presents it to the user. The input is the question data generated in step 5, and the output is the information presented to the user. The presentation is done via a display, allowing the user to directly visually confirm the question.
[0779] Step 7:
[0780] The user reviews the presented questions and decides on their next action based on them. The input is the question displayed on the device, and the output is the user's response or answer. Based on the information provided, the user can engage in more effective discussions.
[0781] (Application Example 2)
[0782] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0783] Maintaining a harmonious atmosphere during meetings and discussions at home and in the office presents challenges, particularly in considering the emotional states of participants. There is a need for means to support smooth communication when tension or conflict arises.
[0784] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0785] In this invention, the server includes an information gathering means for collecting voice information, a conversion means for converting the voice information into text data, and an analysis and generation means for analyzing the text data and voice characteristics to generate emotion-based dialogue. This enables the provision of optimal dialogue in real time that responds to the emotions of the participants, easing the atmosphere of the conversation and facilitating smoother communication.
[0786] "Auditory information" refers to the data of sounds emitted during conversations and dialogues, which are recorded as audio signals.
[0787] "Information gathering means" refers to a device or process that has the function of acquiring voice information and transmitting it to a system.
[0788] "Text data" refers to data in text format that has been converted from audio information, and is treated as text information.
[0789] "Conversion means" refers to a process or device that analyzes audio information and converts it into text data.
[0790] "Speech characteristics" refer to distinctive elements contained in speech information, such as pitch, tone, speed, and intonation.
[0791] "Analysis and generation means" refers to a process or device that analyzes the user's emotions based on text data and voice characteristics and generates appropriate dialogue.
[0792] An "information user" is someone who receives the analyzed and generated information, or someone who utilizes that information.
[0793] An "automated learning method" is a process or device that uses machine learning techniques to improve the performance of a system based on empirically obtained data.
[0794] "Immediately" refers to an action that takes place in real time without any delay.
[0795] The system implementing this invention includes voice collection means, conversion to text data means, analysis and generation means, information presentation means, and automatic learning means. First, voice information of participants is collected in real time by terminals placed in homes or offices. The hardware used includes a high-sensitivity microphone and a voice recognition device. The voice collection means collects voice information and transmits it to a server.
[0796] Next, the server uses a conversion mechanism to convert the received audio information into text data. Speech recognition software (such as the Google Speech Recognition API) is utilized here. The resulting text data is then analyzed by an analysis and generation mechanism to determine the user's emotional state. The elements analyzed include speech characteristics such as voice tone and pitch. Emotion analysis software (such as IBM's Tone Analyzer or similar systems) is used in this process.
[0797] Based on the analyzed emotional state, the server generates appropriate questions and suggestions. This uses a generative AI model to create dialogue that takes into account the user's current emotional state. The generated dialogue is presented to the information user via the terminal. Presentation methods include displays and voice assistants.
[0798] For example, if analysis indicates that one participant is emotionally tense during a family discussion, a prompt such as, "Do you have any ideas for how we can relax and continue the discussion?" might be generated. This helps to alleviate tension and promote a more harmonious discussion.
[0799] Examples of prompts to input into a generative AI model are as follows:
[0800] "Audio data: "Hello. I'd like to discuss recent projects." Text data: "Hello. I'd like to discuss recent projects." Based on this, and analyzing the user's emotions, please consider what questions would be appropriate. Please suggest three questions that would help the user relax."
[0801] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0802] Step 1:
[0803] The device uses a high-sensitivity microphone to collect participants' voice information in real time during meetings and discussions. The input is raw voice data, and the output is a voice sample. This voice sample is then converted into a format that can be processed by speech recognition software.
[0804] Step 2:
[0805] The terminal sends the collected audio information to the server. The server uses speech recognition software to convert the audio information into text data. The input is an audio sample, and the output is text data. Through this conversion process, the conversation content is obtained as text.
[0806] Step 3:
[0807] The server analyzes the user's emotional state using analysis and generation methods based on acquired text data and speech characteristics (pitch, tone, etc.). The input is text data and speech characteristics, and the output is the result of the emotional analysis. Emotional analysis software is used to estimate the user's emotional state (e.g., relaxed, tense).
[0808] Step 4:
[0809] The server uses a generative AI model to generate appropriate dialogue based on the results of sentiment analysis. In this generation process, the sentiment analysis results are used as input, and questions and responses to the user are generated as output. The generated questions are tailored to the user's emotional state.
[0810] Step 5:
[0811] The server sends the generated dialogue to the terminal, which then presents it to the information user using a presentation method. The input is the generated question, and the output is a visual or auditory presentation to the information user. This allows the user to receive information that is sensitive to their emotions.
[0812] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0813] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0814] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0815] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0816] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0817] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0818] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0819] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0820] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0821] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0822] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0823] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0824] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0825] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0826] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0827] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0828] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0829] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0830] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0831] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0832] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0833] The following is further disclosed regarding the embodiments described above.
[0834] (Claim 1)
[0835] A voice collection method for collecting voice data,
[0836] A conversion means for converting the aforementioned audio data into text,
[0837] An analysis and generation means for analyzing the aforementioned text and generating related questions,
[0838] A presentation means for presenting the generated question to the user,
[0839] A system that includes this.
[0840] (Claim 2)
[0841] The system according to claim 1, wherein the analysis and generation means further includes machine learning means for optimizing questions using past conversation data.
[0842] (Claim 3)
[0843] The system according to claim 1, wherein the voice collection means collects voice data in real time.
[0844] "Example 1"
[0845] (Claim 1)
[0846] Information gathering means for acquiring audio information,
[0847] A conversion means for converting the aforementioned audio information into text information,
[0848] An analysis and generation means for analyzing the aforementioned textual information and generating related queries,
[0849] A presentation means for presenting the generated inquiry to the user,
[0850] A learning means for collecting user feedback and improving the analysis and generation means,
[0851] A system that includes this.
[0852] (Claim 2)
[0853] The system according to claim 1, further comprising machine learning means for optimizing inquiries using past dialogue information as analysis and generation means.
[0854] (Claim 3)
[0855] The system according to claim 1, wherein the voice information collection means acquires voice information in real time.
[0856] "Application Example 1"
[0857] (Claim 1)
[0858] Acquisition method for acquiring acoustic data,
[0859] A conversion means for converting the aforementioned acoustic data into textual information,
[0860] An analysis and generation means for analyzing the aforementioned textual information and generating related information,
[0861] A presentation means for presenting the generated information to the user,
[0862] The aforementioned presentation means includes means for presenting recommended information as related information,
[0863] A system that includes this.
[0864] (Claim 2)
[0865] The system according to claim 1, wherein the analysis and generation means further includes a learning means that optimizes information using prior dialogue data.
[0866] (Claim 3)
[0867] The system according to claim 1, wherein the acquisition means acquires acoustic data immediately.
[0868] "Example 2 of combining an emotion engine"
[0869] (Claim 1)
[0870] A device for collecting audio data,
[0871] A device for converting the aforementioned audio data into language data,
[0872] An analysis device that analyzes the aforementioned language data and voice data to infer the user's emotional state,
[0873] A device that generates questions to support conversation based on the emotional state of the user,
[0874] A device for transmitting the generated question,
[0875] A system that includes this.
[0876] (Claim 2)
[0877] The system according to claim 1, further comprising a machine learning device that uses past conversation information to personalize questions.
[0878] (Claim 3)
[0879] The system according to claim 1, wherein the device for collecting the audio data collects the audio data immediately.
[0880] "Application example 2 when combining with an emotional engine"
[0881] (Claim 1)
[0882] Information gathering means for collecting audio information,
[0883] A conversion means for converting the aforementioned audio information into text data,
[0884] An analysis and generation means for analyzing the aforementioned text data and voice characteristics and generating emotion-based dialogue,
[0885] A presentation means for presenting the generated dialogue to the information user,
[0886] A system that includes this.
[0887] (Claim 2)
[0888] The system according to claim 1, wherein the analysis and generation means further comprises an automatic learning means that optimizes dialogue using empirical information.
[0889] (Claim 3)
[0890] The system according to claim 1, wherein the information gathering means collects voice information in real time. [Explanation of Symbols]
[0891] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Acquisition method for acquiring acoustic data, A conversion means for converting the aforementioned acoustic data into textual information, An analysis and generation means for analyzing the aforementioned textual information and generating related information, A presentation means for presenting the generated information to the user, The aforementioned presentation means includes means for presenting recommended information as related information, A system that includes this.
2. The system according to claim 1, wherein the analysis and generation means further includes a learning means that optimizes information using prior dialogue data.
3. The system according to claim 1, wherein the acquisition means acquires acoustic data immediately.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A