system

A system for real-time transcription and analysis of meeting audio data, with feedback loops, addresses inefficiencies in information management during meetings, enabling rapid information acquisition and improved operational efficiency.

JP2026069168APending Publication Date: 2026-04-23SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-11
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

In modern business environments, there is a need for quick information acquisition, accurate recording, and efficient sharing during meetings, but existing systems are inefficient, making it difficult to retrieve necessary materials and reuse information, leading to reduced work efficiency.

Method used

A system that includes real-time audio data transcription, analysis of user statements, information retrieval from a database, and automatic summary generation, with feedback loops to improve accuracy and efficiency.

Benefits of technology

Enables rapid acquisition of necessary information during meetings, facilitates information reuse, and improves operational efficiency by providing timely and accurate summaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069168000001_ABST
    Figure 2026069168000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A method for acquiring audio data and performing real-time transcription, A means of analyzing necessary information from the user's statements and searching for related information in a database, A means to present the acquired information to the user and to automatically generate a summary of the meeting content, A means of improving the accuracy of the system based on user feedback, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In a modern business environment, it is required to quickly obtain important information during a meeting, accurately record and share it. However, in many situations, time is spent on recording conversations and creating materials, resulting in a decline in work efficiency. In addition, there are problems that it is difficult to quickly retrieve necessary materials and past cases, and it is difficult to reuse information. Furthermore, the lack of information support for efficiently conducting meetings has an impact.

Means for Solving the Problems

[0005] To solve this problem, the present invention provides a system that includes means for acquiring audio data and transcribing it in real time, and means for analyzing necessary information from the user's speech and searching for related information from a database. Furthermore, it includes means for presenting the acquired information to the user and automatically generating a summary of the meeting content, allowing for system accuracy improvement based on user feedback. This enables rapid acquisition of necessary information during meetings and efficient information support. Moreover, it facilitates information reuse and improves operational efficiency.

[0006] "Audio data" refers to sound wave information acquired through conversation or vocalization, recorded in digital or analog format.

[0007] "Real-time" refers to the property of processing and reactions occurring in sync with real time, minimizing delays.

[0008] "Transcription" is the process of analyzing audio data to generate corresponding textual information and then writing it down.

[0009] "User statements" refer to linguistic information expressed by users during meetings or conversations, and are acquired as audio or text information.

[0010] "Analysis" is the process of understanding acquired data and organizing, judging, and extracting the information necessary for a specific purpose.

[0011] "Relevant information" refers to information deemed highly necessary based on the user's requests and context, and is retrieved from the database.

[0012] A "database" is a system designed to efficiently store, search, and manage large amounts of data.

[0013] "Presentation" means displaying acquired information in a way that is easy to see or understand, and is the act of providing information to the user.

[0014] "Summarizing" is the process of extracting the main points and essence from a vast amount of information and condensing them into a short form.

[0015] "Feedback" refers to the information used to improve a system based on the reactions and opinions received from users.

[0016] "Accuracy" refers to the degree to which information processing or system operation is accurate in relation to requirements and expectations. [Brief explanation of the drawing]

[0017] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Mode for Carrying Out the Invention

[0018] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a processor with a reference numeral (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0021] In the following embodiments, a RAM (Random Access Memory) with a reference numeral is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0025] [First Embodiment]

[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0038] This invention provides a means to streamline information management in meetings by constructing a system that transcribes conversations in real time, analyzes the necessary information, and presents it appropriately. This system is implemented as follows.

[0039] First, the terminal acquires the audio data of the meeting. Here, a high-sensitivity microphone is used to pick up the conversation, and the audio signal is converted into digital data. The converted audio data is then sent to a server via the network.

[0040] Next, the server analyzes the received audio data and performs real-time transcription using a speech recognition engine. This converts the audio data into text data, which is then temporarily stored in a database. To improve speech recognition accuracy, specialized terms and proper nouns can be pre-registered as a dictionary.

[0041] The server then analyzes the transcribed data to identify necessary information from the user's statements. For example, if a user is looking for similar cases during a meeting, the server searches past conversation history and case databases based on those keywords.

[0042] Once the information retrieval is complete, the server retrieves the relevant information and processes it into a user-friendly format. This information is sent to the terminal and presented to the user in real time. This allows the user to quickly obtain the information needed during the meeting. The server also has a function to automatically generate a meeting summary, and can output it as a report after the meeting based on preset templates as needed.

[0043] Finally, after the meeting concludes, users review the content via their terminals and provide feedback, offering data to improve the system's accuracy. This feedback is then used by the server to inform future processing, providing more accurate and useful information.

[0044] As a concrete example, in product development meetings, when searching for past solutions to technical problems, the system can be used to support decision-making by quickly searching a database of past projects and presenting relevant technical documents in real time. Such functionality allows users to conduct meetings more efficiently and make decisions more quickly.

[0045] The following describes the processing flow.

[0046] Step 1:

[0047] The terminal acquires audio data from the meeting. It collects sound using a high-sensitivity microphone and converts the audio signal into a digital format. The converted digital audio data is then transmitted to the server via the network.

[0048] Step 2:

[0049] The server inputs the received audio data into the speech recognition engine and performs transcription in real time. The recognized text data is temporarily stored in a database. At this stage, a specialized terminology dictionary can be used to improve the accuracy of speech recognition.

[0050] Step 3:

[0051] The server analyzes the transcribed text data and infers the information the user is looking for. For example, it extracts important keywords based on the user's statements and identifies their intent.

[0052] Step 4:

[0053] The server searches for relevant information from past databases and related resources based on identified keywords and intentions. It performs query searches to quickly retrieve similar cases and necessary documents.

[0054] Step 5:

[0055] The server organizes the search results and processes the information into a format that is easy for the user to understand. The processed information is then sent to the terminal and displayed to the user on the screen.

[0056] Step 6:

[0057] Users review the information displayed on their devices and utilize it as needed during the meeting. If a user has specific inquiries or feedback, they send that information to the server via their device.

[0058] Step 7:

[0059] The server automatically generates a meeting summary and outputs it in report format after the meeting, if necessary. It uses a pre-configured template to organize the information.

[0060] Step 8:

[0061] The server receives feedback data from users and uses it as training data to improve the accuracy of future processing. This improves the system's algorithms and enhances the accuracy of the information it provides.

[0062] (Example 1)

[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0064] In today's business environment, meetings are frequently held as a forum for important decision-making, but it is often difficult to properly record what is said during meetings and to quickly obtain the necessary information. Furthermore, summarizing meetings and organizing information afterward takes a lot of time and effort, which reduces work efficiency. In addition, obtaining and presenting relevant information in real time during meetings contributes to quick and accurate decision-making, but traditional systems have the problem of only being able to utilize a limited amount of information.

[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0066] In this invention, the server includes means for acquiring audio information and converting it into a digital signal using a highly sensitive receiver; means for analyzing the received audio information and converting it into text information using a speech recognition function; means for analyzing the transcribed text information, identifying necessary information from the content of speech, and searching for information sources including past meeting history; means for processing the search results and presenting them in a format that is easy for the user to understand, as well as automatically generating a summary of the meeting content; and means for collecting user feedback and using it as data to improve the system's performance. As a result, necessary information can be acquired in a timely and accurate manner during meetings, and summaries can be created efficiently after meetings, thereby improving the overall efficiency of operations.

[0067] "Audio information" refers to data that digitally represents speech and sounds from settings such as meetings.

[0068] A "receiver" is a device used to pick up sound, and its high sensitivity allows it to clearly detect conversations.

[0069] A "digital signal" is a format in which analog audio information is converted into binary data that can be processed by a computer.

[0070] "Speech recognition" refers to the technology that analyzes speech and converts it into text, and is used as part of natural language processing.

[0071] "Text information" refers to a written representation converted from audio data, and is a document format that can be processed electronically.

[0072] "Past meeting history" refers to data that records the content of meetings held in the past, and is used as the basis for information retrieval.

[0073] An "information source" is a collection of data used when searching or referencing, and it forms the basis for obtaining the necessary information.

[0074] "Search results" refer to data obtained from information sources based on specified criteria, and the content presented to the user.

[0075] "Processing" refers to the process of formatting and refining search results and information to make them easier for users to understand.

[0076] A "summary" is a short document that compiles the main points of a meeting and decisions, and it concisely conveys the outcome of the meeting.

[0077] "Automatic generation" generally refers to the process of constructing information using computer programs without human intervention.

[0078] "Opinions" refer to feedback that specifically outlines areas for improvement or suggestions that users have felt regarding the system's functions and results.

[0079] "Performance" is a general term encompassing the functions, processing speed, accuracy, and utilization efficiency of a system, and it is an indicator that should be strived to improve.

[0080] This system includes a series of steps for processing audio data to streamline information management in meetings. A specific embodiment is described below.

[0081] First, the terminal uses a highly sensitive receiver, specifically a standard microphone, to acquire audio information from the meeting. This audio information is received as an analog signal and converted into a digital signal. A computer or a dedicated encoding device is used for this process. The digitized audio information is then transmitted to a server via the network.

[0082] The server analyzes the received digital audio data and converts it into text information in real time using speech recognition technology. Specifically, this could involve utilizing an API equipped with a speech recognition engine. After conversion to text information, this data is temporarily stored in a database.

[0083] Next, the server analyzes the text information and identifies the necessary information from the content of the statements. Natural language processing technology is used to identify the information, and a search engine is used to search past meeting history and related information sources. These search results are processed into a format that is easy for the user to understand and presented to the user through the terminal.

[0084] Furthermore, the server has a function to automatically generate meeting summaries, and can improve system performance based on user feedback. This feedback is reflected in subsequent processing, resulting in further improvements in accuracy. Specifically, the system operates based on the prompt message, "We plan to discuss technical solutions at the next meeting. Please search for relevant past cases." This allows users to quickly obtain relevant information and improve the efficiency of meetings.

[0085] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0086] Step 1:

[0087] The terminal uses a highly sensitive microphone installed in the conference room to acquire real-time audio information from the meeting. This input is an analog audio signal. The audio signal is converted to a digital signal and transmitted to the server via the network as PCM digital data. Digitization makes the audio format usable by a computer.

[0088] Step 2:

[0089] The server receives digital audio information as input and converts it into text information using a speech recognition engine. The speech recognition engine analyzes the audio signal using a specific algorithm and transcribes it into words and sentences. The resulting output is transcribed text data, which is then temporarily stored in a database.

[0090] Step 3:

[0091] The server analyzes the transcribed text data. The input is the text data obtained in step 2. Using natural language processing, important keywords and phrases are extracted from the user's speech. This extracted information serves as foundational data for searching past meeting histories and related documents. The output is a list of the identified information.

[0092] Step 4:

[0093] The server uses the keywords obtained in the previous step to search for information sources within the database. The input is the identified keywords, and the search engine quickly extracts past meeting materials and related information. The output is a set of information that the user needs, and this set is then prepared for further presentation to the user.

[0094] Step 5:

[0095] The server processes search results into a format that is easy for the user to understand. The input is a set of information obtained from the search engine, which is then formatted appropriately, including summarizing and highlighting. The output is the formatted information, which is then sent to the terminal. This allows the user to efficiently obtain the information they need.

[0096] Step 6:

[0097] The server automatically generates a summary after the meeting concludes. The input is data collected during the meeting; a generation AI model extracts key topics and points, which are then compiled into a report format based on a specified template. The output is a summary report of the entire meeting.

[0098] Step 7:

[0099] Users provide feedback through their devices. Input consists of comments and suggestions regarding the meeting content and system output, which are sent to the server and used to improve speech recognition accuracy and information retrieval accuracy in future sessions. Output consists of insights that contribute to improving system performance.

[0100] (Application Example 1)

[0101] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0102] In typical brick-and-mortar stores, it is difficult to respond quickly to customer inquiries and immediately present products that customers are interested in. Therefore, there is a need to improve customer service. Furthermore, store staff have difficulty memorizing vast amounts of product information, resulting in delays in providing appropriate information.

[0103] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0104] In this invention, the server includes means for acquiring acoustic information and performing real-time text conversion, a computer device for analyzing the user's statements and immediately presenting product information, and means for providing information so that the user can confirm the information using a visual device. This makes it possible to provide product information immediately based on the content of the customer's statements.

[0105] "Acoustic information" refers to audible data such as voices and sound effects, which are acquired via input devices such as microphones.

[0106] "Real-time text conversion" refers to the process of instantly converting acquired audio information into text information, indicating that time delays are minimized.

[0107] "User" refers to an individual or organization that uses a system or service, and is primarily an entity that interacts with the system through voice.

[0108] "Analyzing the content of a user's statements" refers to analyzing the linguistic and semantic elements contained in the user's statements, with the aim of extracting necessary information within the system.

[0109] A "computer device that instantly displays product information" refers to a device that has the function of quickly acquiring, processing, and displaying product-related information according to the user's needs.

[0110] "Visual devices" are devices used to visually present textual information and images to users, and include displays such as smart glasses.

[0111] "Means of providing information" refers to a system function that outputs the information a user needs in a format that the user can view, after a series of processing steps.

[0112] As a concrete example of the present invention, a system for improving customer service in a physical store is implemented. This system includes a high-sensitivity microphone, smart glasses, a server, and a database management system.

[0113] First, the smart glasses, which serve as the device, have a built-in high-sensitivity microphone that constantly acquires ambient acoustic information. The acquired audio data is immediately digitized and sent to a server via the network. The server uses the Google® Cloud Speech-to-Text API to convert the audio data into text in real time.

[0114] Next, the text data of the spoken content is analyzed using a generative AI model, and relevant product information is searched from the database. At this time, the server extracts keywords from the spoken content, and the computer immediately processes and provides product information based on those keywords.

[0115] The information is displayed on the smart glasses' screen, allowing store staff (users) to visually confirm it and appropriately guide customers with product information. For example, if a customer says, "I want a camera that's perfect for travel," the server searches for relevant camera information and displays it on the smart glasses. The user can then use this information to immediately explain the most suitable product to the customer.

[0116] In embodiments of the present invention, an example of a prompt sentence based on the generated AI model is, "Analyze the voice of a customer who spoke about a camera suitable for travel, and generate a prompt sentence that suggests relevant information." This prompt enables the system to provide information that immediately responds to the customer's needs.

[0117] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0118] Step 1:

[0119] The smart glasses, which act as the device, acquire acoustic information using a high-sensitivity microphone. This allows surrounding conversations and sounds to be collected as digital data. The input is an analog audio signal, and the output is audio data in digital format.

[0120] Step 2:

[0121] The terminal transmits the acquired digital audio data to the server via the network. Here, the audio data is converted into packets and transmitted to the server. The input is digital audio data, and the output is the data received by the server.

[0122] Step 3:

[0123] The server uses the Google Cloud Speech-to-Text API to convert incoming audio data into text in real time. The speech recognition engine analyzes phonemes and generates text data. The input is audio data sent to the server, and the output is data in text format.

[0124] Step 4:

[0125] The server uses a generative AI model to extract keywords from text data and analyze the content of the speech. Here, it identifies important words and phrases within the text data. The input is text data, and the output is a list of extracted keywords.

[0126] Step 5:

[0127] The server uses a database management system to search for product information based on extracted keywords. Relevant product information is retrieved instantly via database queries. The input is a list of keywords, and the output is the corresponding product information.

[0128] Step 6:

[0129] The server converts the acquired product information into a format that users can visually verify and transmits it to the terminal via a computer. Data optimized for display is generated. The input is product information, and the output is a displayable data format.

[0130] Step 7:

[0131] The smart glasses on the device display product information transmitted from the server. Store staff, who are the users, can check the information in real time and use it to assist customers. Input is in a displayable data format, and output is information confirmation by the user's eyes.

[0132] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0133] This invention provides a system that utilizes audio data from meetings to recognize participants' emotions and optimize information delivery. This system consists of real-time transcription of audio data, information retrieval from a database, information presentation, automatic summarization of meeting content, and an emotion recognition engine.

[0134] First, the terminal acquires audio data during the meeting and sends this audio data to the server. The server uses a speech recognition engine to transcribe the audio into text and passes this text data to an emotion engine, which analyzes the emotional state of the participants based on their statements. The emotion engine uses elements such as tone, speed, and language patterns to determine the speaker's emotions (e.g., excitement, calmness, anger).

[0135] Once an emotion is recognized, the server optimizes the information provided to support the meeting based on the acquired emotion information. For example, if the system detects that the user is stressed, it will simplify the information it suggests or visually highlight meeting materials to make the information easier for the user to understand.

[0136] Furthermore, the server creates a summary of the meeting content at the end of the meeting, reflecting the user's emotional state. This ensures that feedback and important points are appropriately highlighted, and the report is provided in a way that aligns with the user's emotional state.

[0137] For example, in a kick-off meeting for a new project, the emotion engine may detect that the presenter is feeling anxious about new technical challenges. Based on this emotional information, the system immediately searches for relevant technical data and past success stories, presenting them instantly to alleviate anxiety and help the meeting proceed smoothly.

[0138] Thus, the present invention supports improved meeting environments and efficient decision-making processes through an information provision system that combines emotion recognition.

[0139] The following describes the processing flow.

[0140] Step 1:

[0141] The terminal acquires audio data from the meeting. The audio data is collected using a high-sensitivity microphone, and the digital audio signal is transmitted to the server via the network.

[0142] Step 2:

[0143] The server passes the received audio data to the speech recognition engine, which performs transcription in real time. The generated text data is then supplied to the sentiment engine and the database search engine.

[0144] Step 3:

[0145] The server uses an emotion engine to analyze the speaker's emotions from text data. It uses voice tone, speaking speed, and linguistic features to determine the speaker's emotional state and records the emotional information.

[0146] Step 4:

[0147] Based on the sentiment analysis results, the server searches the database for relevant information to provide the user with appropriate information. If necessary, it also collects similar past cases and related materials.

[0148] Step 5:

[0149] The server uses emotional information to optimize information display. For example, if a user is confused, it simplifies the information presented or changes the interface to make it more visually understandable.

[0150] Step 6:

[0151] The device receives optimized information from the server and displays it to the user. The way the information is presented is tailored to the user's current emotional state and is adjusted for easy understanding.

[0152] Step 7:

[0153] Users review the information presented via their device and provide feedback as needed. This feedback is sent from the device to the server to be used for future system improvements.

[0154] Step 8:

[0155] At the end of the meeting, the server creates a summary that reflects sentiment information and outputs it as an automatically generated meeting report. This allows users to review key points from the meeting based on sentiment data.

[0156] (Example 2)

[0157] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 will be referred to as the "terminal".

[0158] The challenge lies in efficiently utilizing audio data from meetings to understand participants' emotional states, thereby optimizing meeting progress and providing an environment where participants can quickly and accurately grasp information. Furthermore, it is crucial to appropriately summarize the meeting content and provide it as a useful resource for subsequent reviews and decision-making.

[0159] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0160] In this invention, the server includes means for acquiring audio signals and performing data conversion in real time, means for analyzing emotional states from the content and characteristics of speech and retrieving relevant information from an external storage device, and means for optimizing information presentation based on the analyzed emotional information and providing it to the user, as well as automatically generating a summary of the meeting content. This makes it possible to improve the efficiency of meetings and the quality of decision-making by analyzing participants' emotions in real time during meetings and providing appropriate information.

[0161] An "audio signal" is an electrical representation of audio information, treating sound vibrations as signals that change along the time axis.

[0162] "Data conversion" is the process of converting audio signals into other formats, such as text data, in real time.

[0163] "Content and characteristics of speech" refers to the meaning of words within the audio data and the characteristics that indicate the speaker's emotional nuances.

[0164] "Analyzing emotional states" is the process of estimating a speaker's emotions from audio or text data and identifying that state.

[0165] "External storage device" refers to a storage medium used to store data for a long period of time, and includes, but is not limited to, databases.

[0166] "Optimizing information presentation" refers to displaying and providing information in the most appropriate format based on the user's situation and emotions.

[0167] "Users" refers to anyone who uses the system and receives information, including meeting participants and general individuals who receive information.

[0168] A "meeting summary" is a compilation of the key points and decisions discussed at a meeting, reorganized into a concise format for later use or reference.

[0169] This invention is a system that optimizes the progress of a meeting and streamlines information presentation by acquiring audio signals during a meeting, converting those signals into data in real time, and analyzing the emotional state of the speakers.

[0170] The terminal uses a microphone device to acquire the user's audio signal during the meeting. This audio signal is transmitted to the server via the network.

[0171] The server uses speech recognition software (e.g., a general speech recognition engine) to convert the received audio signal into text format. The converted text data is sent to an emotion analysis engine (e.g., a natural language processing engine) to analyze the participant's emotions based on the content and characteristics of their speech.

[0172] Based on the analyzed emotional information, the server retrieves the most relevant information from external storage devices and provides it to the user, tailored to the meeting content. Specifically, if the emotional state is one of anxiety, the server presents relevant technical documents and past success stories to improve meeting efficiency.

[0173] Users can view information provided in real time and provide feedback. This feedback is aggregated on the server as data to improve the accuracy of the system's analysis.

[0174] A concrete example would be a meeting to launch a new project, where technical challenges are addressed. If the presenter expresses concerns about these challenges, the system can instantly provide documentation of past cases and solutions, reducing anxiety and ensuring the meeting proceeds smoothly.

[0175] Examples of prompts to input into a generative AI model include: "Design a system that recognizes the emotions of meeting participants in real time and presents the most relevant information based on those emotions. For example, if a participant is feeling anxious about a technical issue, present them with relevant data."

[0176] This system improves the meeting environment and significantly enhances the efficiency of the decision-making process.

[0177] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0178] Step 1:

[0179] The terminal uses a microphone to acquire audio signals in real time during a meeting. The acquired audio signals are then transmitted directly to the server via the network. Its input is raw audio data, and its output is digital audio data that can be transferred to the server.

[0180] Step 2:

[0181] The server converts the received audio signal into text data using a speech recognition engine (e.g., general speech recognition software). In this process, the input is digital audio data, and the output is transcribed text data. Specifically, the process involves analyzing the audio waveform along the time axis and converting phonemes into text strings.

[0182] Step 3:

[0183] The transcribed text data is sent to an emotion analysis engine on the server and used for analysis to determine the speaker's emotional state. The input is text data, and the output is data tagged with the speaker's emotional state (e.g., relief, excitement, doubt). In this step, natural language processing techniques are applied to analyze language patterns and extract keywords.

[0184] Step 4:

[0185] Based on the analyzed emotional state, the server initiates a process to retrieve appropriate information from external storage and provide efficient information presentation. The input is emotional state-tagged data, and the output is a list of relevant information matching the emotion. Specifically, it executes database queries to list examples and literature.

[0186] Step 5:

[0187] During a meeting, relevant information is displayed in real time on the user's device (terminal). The input is a list of relevant information, and the output provides the user with visually organized information. The system utilizes a graphical user interface to visually highlight and categorize the information.

[0188] Step 6:

[0189] The server automatically generates a summary of the meeting content based on data including changes in emotional states obtained after the meeting. The input is all meeting data and emotional information, and the output is a concise meeting report summarizing the key points. Specific operations include selecting high-priority topics and analyzing emotional changes over time.

[0190] (Application Example 2)

[0191] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0192] In modern meetings and customer service, accurately understanding the emotions of participants and customers and responding appropriately is essential. However, traditional systems have struggled to analyze emotional states in real time from tone of voice and content of speech, and to provide optimal information. Furthermore, the lack of emotion-based suggestions has made it difficult to improve the quality of customer service. These challenges need to be addressed.

[0193] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0194] In this invention, the server includes means for acquiring audio information and transcribing it in real time, means for analyzing the emotional state of participants from their statements, and means for retrieving and presenting relevant information from a database according to the emotional state. This enables the provision of appropriate information according to the emotional state of participants and improves the quality of customer service.

[0195] "Audio information" refers to data on the content of speech uttered by participants and customers, collected during meetings and interactions with customers.

[0196] "Real-time" refers to technology that processes and analyzes events as they occur.

[0197] "Transcription" is a technology that converts audio information into text format.

[0198] "Participants" are individuals involved in a meeting or discussion.

[0199] "Statements" are words expressed orally during meetings or discussions.

[0200] "Emotional state" refers to the psychological state and emotions of participants or customers.

[0201] "Analysis" is the process of analyzing data and information to understand their meaning and patterns.

[0202] "Relevant information" refers to data and knowledge that are presented appropriately in response to the statements and emotional states of participants and customers.

[0203] A "database" is a collection of information organized in a systematic way that allows for efficient storage, searching, and use of information.

[0204] A "proposal" is to recommend appropriate policies or actions based on the situation.

[0205] "Quality improvement" means enhancing the value and effectiveness of the services and information provided.

[0206] This invention is implemented as a system to improve customer service in physical stores. The system consists of a program for effectively utilizing various data in meetings and interactions with customers in stores. The embodiments thereof are described below.

[0207] The server first collects customer voices from smart devices in the store and obtains audio information. The acquired audio information is transcribed in real time using the Google Cloud Speech-to-Text API. This transcription result is sent to Amazon Comprehend for real-time sentiment analysis, where the customer's emotional state is analyzed. Based on the analyzed emotional state, the server searches the database for relevant information and selects the appropriate information.

[0208] The smart glasses, acting as the terminal, visually display the customer's emotional state and related information transmitted from the server in real time. This allows the store staff, acting as users, to provide service tailored to the customer's emotional state. For example, if the glasses detect that the customer is confused, information carefully explaining how to set up and use the product will be displayed on them.

[0209] As a concrete example, suppose emotion analysis detects that customer A is looking for a specific product but cannot find it and appears confused. In this case, the smart glasses would display a specific response plan, such as "Suggest to guide customer A to the location of the product on the shelf."

[0210] Furthermore, the server also includes a function to ensure that the information presented is based on the customer's emotional state. This is expected to improve the quality of customer service and increase customer satisfaction.

[0211] Example of a prompt

[0212] Convert speech to text in real time and analyze customer emotions. If a customer is having trouble, immediately suggest helpful information and display it on smart glasses.

[0213] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0214] Step 1:

[0215] The server acquires customer voices as audio data from smart devices acting as terminals. The input is real-time audio information collected from smart devices within the store. The output is audio data used in the speech recognition process.

[0216] Step 2:

[0217] The server transcribes the acquired audio data in real time using the Google Cloud Speech-to-Text API. The input is the audio data obtained in step 1, and the output is the transcribed text. This process prepares the spoken content for analysis as digital data.

[0218] Step 3:

[0219] The server passes text data to Amazon Comprehend, which analyzes the spoken content to identify the customer's emotional state. The input is the transcribed text from step 2, and the output is information about the emotional state. This procedure allows the server to understand the customer's psychological state in real time.

[0220] Step 4:

[0221] The server searches the database for appropriate relevant information based on the analyzed emotional state. The input is the emotional state information obtained in step 3, and the output is relevant information that matches the emotional state. A prompt sentence is generated through a generative AI model, and the optimal information is selected.

[0222] Step 5:

[0223] The smart glasses, acting as the terminal, visually display relevant information sent from the server to the customer. The input is the relevant information selected in step 4, and the output is information presented visually in a user-friendly format. This enables the user to respond quickly and appropriately to the customer.

[0224] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0225] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0226] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0227] [Second Embodiment]

[0228] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0229] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0230] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0231] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0232] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0233] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0234] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0235] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0236] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0237] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0238] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0239] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0240] This invention provides a means to streamline information management in meetings by constructing a system that transcribes conversations in real time, analyzes the necessary information, and presents it appropriately. This system is implemented as follows.

[0241] First, the terminal acquires the audio data of the meeting. Here, a high-sensitivity microphone is used to pick up the conversation, and the audio signal is converted into digital data. The converted audio data is then sent to a server via the network.

[0242] Next, the server analyzes the received audio data and performs real-time transcription using a speech recognition engine. This converts the audio data into text data, which is then temporarily stored in a database. To improve speech recognition accuracy, specialized terms and proper nouns can be pre-registered as a dictionary.

[0243] The server then analyzes the transcribed data to identify necessary information from the user's statements. For example, if a user is looking for similar cases during a meeting, the server searches past conversation history and case databases based on those keywords.

[0244] Once the information retrieval is complete, the server retrieves the relevant information and processes it into a user-friendly format. This information is sent to the terminal and presented to the user in real time. This allows the user to quickly obtain the information needed during the meeting. The server also has a function to automatically generate a meeting summary, and can output it as a report after the meeting based on preset templates as needed.

[0245] Finally, after the meeting concludes, users review the content via their terminals and provide feedback, offering data to improve the system's accuracy. This feedback is then used by the server to inform future processing, providing more accurate and useful information.

[0246] As a concrete example, in product development meetings, when searching for past solutions to technical problems, the system can be used to support decision-making by quickly searching a database of past projects and presenting relevant technical documents in real time. Such functionality allows users to conduct meetings more efficiently and make decisions more quickly.

[0247] The following describes the processing flow.

[0248] Step 1:

[0249] The terminal acquires audio data from the meeting. It collects sound using a high-sensitivity microphone and converts the audio signal into a digital format. The converted digital audio data is then transmitted to the server via the network.

[0250] Step 2:

[0251] The server inputs the received audio data into the speech recognition engine and performs transcription in real time. The recognized text data is temporarily stored in a database. At this stage, a specialized terminology dictionary can be used to improve the accuracy of speech recognition.

[0252] Step 3:

[0253] The server analyzes the transcribed text data and infers the information the user is looking for. For example, it extracts important keywords based on the user's statements and identifies their intent.

[0254] Step 4:

[0255] The server searches for relevant information from past databases and related resources based on identified keywords and intentions. It performs query searches to quickly retrieve similar cases and necessary documents.

[0256] Step 5:

[0257] The server organizes the search results and processes the information into a format that is easy for the user to understand. The processed information is then sent to the terminal and displayed to the user on the screen.

[0258] Step 6:

[0259] Users review the information displayed on their devices and utilize it as needed during the meeting. If a user has specific inquiries or feedback, they send that information to the server via their device.

[0260] Step 7:

[0261] The server automatically generates a meeting summary and outputs it in report format after the meeting, if necessary. It uses a pre-configured template to organize the information.

[0262] Step 8:

[0263] The server receives feedback data from users and uses it as training data to improve the accuracy of future processing. This improves the system's algorithms and enhances the accuracy of the information it provides.

[0264] (Example 1)

[0265] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0266] In today's business environment, meetings are frequently held as a forum for important decision-making, but it is often difficult to properly record what is said during meetings and to quickly obtain the necessary information. Furthermore, summarizing meetings and organizing information afterward takes a lot of time and effort, which reduces work efficiency. In addition, obtaining and presenting relevant information in real time during meetings contributes to quick and accurate decision-making, but traditional systems have the problem of only being able to utilize a limited amount of information.

[0267] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0268] In this invention, the server includes means for acquiring audio information and converting it into a digital signal using a highly sensitive receiver; means for analyzing the received audio information and converting it into text information using a speech recognition function; means for analyzing the transcribed text information, identifying necessary information from the content of speech, and searching for information sources including past meeting history; means for processing the search results and presenting them in a format that is easy for the user to understand, as well as automatically generating a summary of the meeting content; and means for collecting user feedback and using it as data to improve the system's performance. As a result, necessary information can be acquired in a timely and accurate manner during meetings, and summaries can be created efficiently after meetings, thereby improving the overall efficiency of operations.

[0269] "Audio information" refers to data that digitally represents speech and sounds from settings such as meetings.

[0270] A "receiver" is a device used to pick up sound, and its high sensitivity allows it to clearly detect conversations.

[0271] A "digital signal" is a format in which analog audio information is converted into binary data that can be processed by a computer.

[0272] "Speech recognition" refers to the technology that analyzes speech and converts it into text, and is used as part of natural language processing.

[0273] "Text information" refers to a written representation converted from audio data, and is a document format that can be processed electronically.

[0274] "Past meeting history" refers to data that records the content of meetings held in the past, and is used as the basis for information retrieval.

[0275] An "information source" is a collection of data used when searching or referencing, and it forms the basis for obtaining the necessary information.

[0276] "Search results" refer to the data obtained from information sources based on specified criteria and indicate the content presented to the user.

[0277] "Processing" is the process of organizing the format and content to make search results and information more understandable to the user.

[0278] "Summary" is a short document that summarizes the main statements and decisions of a meeting and concisely conveys the results of the meeting.

[0279] "Automatic generation" generally refers to the process of constructing information without human intervention using computer programs.

[0280] "Opinion" is specific feedback that indicates the areas for improvement and suggestions felt by the user regarding the functions and results of the system.

[0281] "Performance" is a general term for the functions, processing speed, accuracy, and utilization efficiency of a system and is an indicator that should be aimed for improvement.

[0282] This system includes a series of steps for voice data processing to streamline information management in meetings. The specific implementation forms will be described below.

[0283] First, the terminal uses a sensitive receiver, specifically a general microphone, to obtain the voice information of the meeting. This voice information is received as an analog signal and converted into a digital signal. A computer or a dedicated encoding device is used in this process. The digitized voice information is transmitted to the server through the network.

[0284] The server analyzes the received digital voice data and converts it into text information in real time using the voice recognition function. As specific software, it is conceivable to utilize an API equipped with a voice recognition engine. After being converted into text information, this data is temporarily stored in the database.

[0285] Next, the server analyzes the text information and identifies the necessary information from the speech content. Natural language processing technology is utilized for information identification, and a search engine is used when searching the past meeting history and related information sources. This search result is processed into a user-friendly format and presented to the user through the terminal.

[0286] Furthermore, the server has a function of automatically generating the summary of the meeting, and the performance of the system can be improved based on the user's feedback. These feedbacks are reflected in the subsequent processing, and as a result, further improvement in accuracy is achieved. As a specific example, the system is implemented based on the prompt sentence "It is planned to discuss technical solutions in the next meeting. Please search for relevant past cases." Thereby, the user can quickly obtain relevant information and improve the efficiency of the meeting.

[0287] The flow of the specific process in Example 1 will be described using FIG. 11.

[0288] Step 1:

[0289] The terminal uses a high-sensitivity microphone installed in the meeting room to obtain the audio information of the meeting in real time. This input is an analog audio signal. The audio signal is converted into a digital signal and transmitted to the server via the network as digital data in PCM format. By digitization, the audio becomes a format that can be processed by a computer.

[0290] Step 2:

[0291] Using the received digital audio information as input, the server converts it into text information using a speech recognition engine. The speech recognition engine analyzes the audio signal using a specific algorithm and converts it into text as phrases and sentences. The output of this result is the transcribed text data, which is further temporarily stored in the database.

[0292] Step 3:

[0293] The server analyzes the transcribed text data. The input is the text data obtained in step 2. Using natural language processing, important keywords and phrases are extracted from the user's speech. This extracted information serves as foundational data for searching past meeting histories and related documents. The output is a list of the identified information.

[0294] Step 4:

[0295] The server uses the keywords obtained in the previous step to search for information sources within the database. The input is the identified keywords, and the search engine quickly extracts past meeting materials and related information. The output is a set of information that the user needs, and this set is then prepared for further presentation to the user.

[0296] Step 5:

[0297] The server processes search results into a format that is easy for the user to understand. The input is a set of information obtained from the search engine, which is then formatted appropriately, including summarizing and highlighting. The output is the formatted information, which is then sent to the terminal. This allows the user to efficiently obtain the information they need.

[0298] Step 6:

[0299] The server automatically generates a summary after the meeting concludes. The input is data collected during the meeting; a generation AI model extracts key topics and points, which are then compiled into a report format based on a specified template. The output is a summary report of the entire meeting.

[0300] Step 7:

[0301] The user provides feedback through the terminal. The input is impressions and suggestions regarding the content of the meeting and the output of the system, which are sent to the server and become data for improving the speech recognition accuracy and information search accuracy from the next time onwards. The output is insights that lead to performance improvement of the system.

[0302] (Application Example 1)

[0303] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0304] In a general physical store, it is difficult to quickly respond to customers' questions and immediately present products that they are interested in. Therefore, improvement of customer service is required. Also, there is a problem that it is difficult for store staff to remember a huge amount of product information and it takes time to provide appropriate information.

[0305] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0306] In this invention, the server includes means for acquiring acoustic information and performing real-time character conversion, a computer device for analyzing the user's speech and immediately presenting product information, and means for providing information so that the user can confirm the information with a visual device. Thereby, it becomes possible to immediately provide product information based on the content of the customer's speech.

[0307] "Acoustic information" refers to audible data such as speech and sound effects, and is acquired via an input device such as a microphone.

[0308] "Real-time character conversion" refers to the process of immediately converting the acquired acoustic information into character information, indicating that the time delay is minimized.

[0309] "User" refers to an individual or organization that uses a system or service, and is primarily an entity that interacts with the system through voice.

[0310] "Analyzing the content of a user's statements" refers to analyzing the linguistic and semantic elements contained in the user's statements, with the aim of extracting necessary information within the system.

[0311] A "computer device that instantly displays product information" refers to a device that has the function of quickly acquiring, processing, and displaying product-related information according to the user's needs.

[0312] "Visual devices" are devices used to visually present textual information and images to users, and include displays such as smart glasses.

[0313] "Means of providing information" refers to a system function that outputs the information a user needs in a format that the user can view, after a series of processing steps.

[0314] As a concrete example of the present invention, a system for improving customer service in a physical store is implemented. This system includes a high-sensitivity microphone, smart glasses, a server, and a database management system.

[0315] First, the smart glasses, which serve as the device, have a built-in high-sensitivity microphone that constantly acquires ambient acoustic information. The acquired audio data is immediately digitized and sent to a server via the network. The server uses the Google Cloud Speech-to-Text API to convert the audio data into text in real time.

[0316] Next, the text data of the spoken content is analyzed using a generative AI model, and relevant product information is searched from the database. At this time, the server extracts keywords from the spoken content, and the computer immediately processes and provides product information based on those keywords.

[0317] The information is displayed on the smart glasses' screen, allowing store staff (users) to visually confirm it and appropriately guide customers with product information. For example, if a customer says, "I want a camera that's perfect for travel," the server searches for relevant camera information and displays it on the smart glasses. The user can then use this information to immediately explain the most suitable product to the customer.

[0318] In embodiments of the present invention, an example of a prompt sentence based on the generated AI model is, "Analyze the voice of a customer who spoke about a camera suitable for travel, and generate a prompt sentence that suggests relevant information." This prompt enables the system to provide information that immediately responds to the customer's needs.

[0319] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0320] Step 1:

[0321] The smart glasses, which act as the device, acquire acoustic information using a high-sensitivity microphone. This allows surrounding conversations and sounds to be collected as digital data. The input is an analog audio signal, and the output is audio data in digital format.

[0322] Step 2:

[0323] The terminal transmits the acquired digital audio data to the server via the network. Here, the audio data is converted into packets and transmitted to the server. The input is digital audio data, and the output is the data received by the server.

[0324] Step 3:

[0325] The server uses the Google Cloud Speech-to-Text API to convert incoming audio data into text in real time. The speech recognition engine analyzes phonemes and generates text data. The input is audio data sent to the server, and the output is data in text format.

[0326] Step 4:

[0327] The server uses a generative AI model to extract keywords from text data and analyze the content of the speech. Here, it identifies important words and phrases within the text data. The input is text data, and the output is a list of extracted keywords.

[0328] Step 5:

[0329] The server uses a database management system to search for product information based on extracted keywords. Relevant product information is retrieved instantly via database queries. The input is a list of keywords, and the output is the corresponding product information.

[0330] Step 6:

[0331] The server converts the acquired product information into a format that users can visually verify and transmits it to the terminal via a computer. Data optimized for display is generated. The input is product information, and the output is a displayable data format.

[0332] Step 7:

[0333] The smart glasses on the device display product information transmitted from the server. Store staff, who are the users, can check the information in real time and use it to assist customers. Input is in a displayable data format, and output is information confirmation by the user's eyes.

[0334] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0335] This invention provides a system that utilizes audio data from meetings to recognize participants' emotions and optimize information delivery. This system consists of real-time transcription of audio data, information retrieval from a database, information presentation, automatic summarization of meeting content, and an emotion recognition engine.

[0336] First, the terminal acquires audio data during the meeting and sends this audio data to the server. The server uses a speech recognition engine to transcribe the audio into text and passes this text data to an emotion engine, which analyzes the emotional state of the participants based on their statements. The emotion engine uses elements such as tone, speed, and language patterns to determine the speaker's emotions (e.g., excitement, calmness, anger).

[0337] Once an emotion is recognized, the server optimizes the information provided to support the meeting based on the acquired emotion information. For example, if the system detects that the user is stressed, it will simplify the information it suggests or visually highlight meeting materials to make the information easier for the user to understand.

[0338] Furthermore, the server creates a summary of the meeting content at the end of the meeting, reflecting the user's emotional state. This ensures that feedback and important points are appropriately highlighted, and the report is provided in a way that aligns with the user's emotional state.

[0339] For example, in a kick-off meeting for a new project, the emotion engine may detect that the presenter is feeling anxious about new technical challenges. Based on this emotional information, the system immediately searches for relevant technical data and past success stories, presenting them instantly to alleviate anxiety and help the meeting proceed smoothly.

[0340] Thus, the present invention supports improved meeting environments and efficient decision-making processes through an information provision system that combines emotion recognition.

[0341] The following describes the processing flow.

[0342] Step 1:

[0343] The terminal acquires audio data from the meeting. The audio data is collected using a high-sensitivity microphone, and the digital audio signal is transmitted to the server via the network.

[0344] Step 2:

[0345] The server passes the received audio data to the speech recognition engine, which performs transcription in real time. The generated text data is then supplied to the sentiment engine and the database search engine.

[0346] Step 3:

[0347] The server uses an emotion engine to analyze the speaker's emotions from text data. It uses voice tone, speaking speed, and linguistic features to determine the speaker's emotional state and records the emotional information.

[0348] Step 4:

[0349] Based on the sentiment analysis results, the server searches the database for relevant information to provide the user with appropriate information. If necessary, it also collects similar past cases and related materials.

[0350] Step 5:

[0351] The server uses emotional information to optimize information display. For example, if a user is confused, it simplifies the information presented or changes the interface to make it more visually understandable.

[0352] Step 6:

[0353] The device receives optimized information from the server and displays it to the user. The way the information is presented is tailored to the user's current emotional state and is adjusted for easy understanding.

[0354] Step 7:

[0355] Users review the information presented via their device and provide feedback as needed. This feedback is sent from the device to the server to be used for future system improvements.

[0356] Step 8:

[0357] At the end of the meeting, the server creates a summary that reflects sentiment information and outputs it as an automatically generated meeting report. This allows users to review key points from the meeting based on sentiment data.

[0358] (Example 2)

[0359] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0360] The challenge lies in efficiently utilizing audio data from meetings to understand participants' emotional states, thereby optimizing meeting progress and providing an environment where participants can quickly and accurately grasp information. Furthermore, it is crucial to appropriately summarize the meeting content and provide it as a useful resource for subsequent reviews and decision-making.

[0361] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0362] In this invention, the server includes means for acquiring audio signals and performing data conversion in real time, means for analyzing emotional states from the content and characteristics of speech and retrieving relevant information from an external storage device, and means for optimizing information presentation based on the analyzed emotional information and providing it to the user, as well as automatically generating a summary of the meeting content. This makes it possible to improve the efficiency of meetings and the quality of decision-making by analyzing participants' emotions in real time during meetings and providing appropriate information.

[0363] An "audio signal" is an electrical representation of audio information, treating sound vibrations as signals that change along the time axis.

[0364] "Data conversion" is the process of converting audio signals into other formats, such as text data, in real time.

[0365] "Content and characteristics of speech" refers to the meaning of words within the audio data and the characteristics that indicate the speaker's emotional nuances.

[0366] "Analyzing emotional states" is the process of estimating a speaker's emotions from audio or text data and identifying that state.

[0367] "External storage device" refers to a storage medium used to store data for a long period of time, and includes, but is not limited to, databases.

[0368] "Optimizing information presentation" refers to displaying and providing information in the most appropriate format based on the user's situation and emotions.

[0369] "Users" refers to anyone who uses the system and receives information, including meeting participants and general individuals who receive information.

[0370] A "meeting summary" is a compilation of the key points and decisions discussed at a meeting, reorganized into a concise format for later use or reference.

[0371] This invention is a system that optimizes the progress of a meeting and streamlines information presentation by acquiring audio signals during a meeting, converting those signals into data in real time, and analyzing the emotional state of the speakers.

[0372] The terminal uses a microphone device to acquire the user's audio signal during the meeting. This audio signal is transmitted to the server via the network.

[0373] The server uses speech recognition software (e.g., a general speech recognition engine) to convert the received audio signal into text format. The converted text data is sent to an emotion analysis engine (e.g., a natural language processing engine) to analyze the participant's emotions based on the content and characteristics of their speech.

[0374] Based on the analyzed emotional information, the server retrieves the most relevant information from external storage devices and provides it to the user, tailored to the meeting content. Specifically, if the emotional state is one of anxiety, the server presents relevant technical documents and past success stories to improve meeting efficiency.

[0375] Users can view information provided in real time and provide feedback. This feedback is aggregated on the server as data to improve the accuracy of the system's analysis.

[0376] A concrete example would be a meeting to launch a new project, where technical challenges are addressed. If the presenter expresses concerns about these challenges, the system can instantly provide documentation of past cases and solutions, reducing anxiety and ensuring the meeting proceeds smoothly.

[0377] Examples of prompts to input into a generative AI model include: "Design a system that recognizes the emotions of meeting participants in real time and presents the most relevant information based on those emotions. For example, if a participant is feeling anxious about a technical issue, present them with relevant data."

[0378] This system improves the meeting environment and significantly enhances the efficiency of the decision-making process.

[0379] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0380] Step 1:

[0381] The terminal uses a microphone to acquire audio signals in real time during a meeting. The acquired audio signals are then transmitted directly to the server via the network. Its input is raw audio data, and its output is digital audio data that can be transferred to the server.

[0382] Step 2:

[0383] The server converts the received audio signal into text data using a speech recognition engine (e.g., general speech recognition software). In this process, the input is digital audio data, and the output is transcribed text data. Specifically, the process involves analyzing the audio waveform along the time axis and converting phonemes into text strings.

[0384] Step 3:

[0385] The transcribed text data is sent to an emotion analysis engine on the server and used for analysis to determine the speaker's emotional state. The input is text data, and the output is data tagged with the speaker's emotional state (e.g., relief, excitement, doubt). In this step, natural language processing techniques are applied to analyze language patterns and extract keywords.

[0386] Step 4:

[0387] Based on the analyzed emotional state, the server initiates a process to retrieve appropriate information from external storage and provide efficient information presentation. The input is emotional state-tagged data, and the output is a list of relevant information matching the emotion. Specifically, it executes database queries to list examples and literature.

[0388] Step 5:

[0389] During a meeting, relevant information is displayed in real time on the user's device (terminal). The input is a list of relevant information, and the output provides the user with visually organized information. The system utilizes a graphical user interface to visually highlight and categorize the information.

[0390] Step 6:

[0391] The server automatically generates a summary of the meeting content based on data including changes in emotional states obtained after the meeting. The input is all meeting data and emotional information, and the output is a concise meeting report summarizing the key points. Specific operations include selecting high-priority topics and analyzing emotional changes over time.

[0392] (Application Example 2)

[0393] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0394] In modern meetings and customer service, accurately understanding the emotions of participants and customers and responding appropriately is essential. However, traditional systems have struggled to analyze emotional states in real time from tone of voice and content of speech, and to provide optimal information. Furthermore, the lack of emotion-based suggestions has made it difficult to improve the quality of customer service. These challenges need to be addressed.

[0395] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0396] In this invention, the server includes means for acquiring audio information and transcribing it in real time, means for analyzing the emotional state of participants from their statements, and means for retrieving and presenting relevant information from a database according to the emotional state. This enables the provision of appropriate information according to the emotional state of participants and improves the quality of customer service.

[0397] "Audio information" refers to data on the content of speech uttered by participants and customers, collected during meetings and interactions with customers.

[0398] "Real-time" refers to technology that processes and analyzes events as they occur.

[0399] "Transcription" is a technology that converts audio information into text format.

[0400] "Participants" are individuals involved in a meeting or discussion.

[0401] "Statements" are words expressed orally during meetings or discussions.

[0402] "Emotional state" refers to the psychological state and emotions of participants or customers.

[0403] "Analysis" is the process of analyzing data and information to understand their meaning and patterns.

[0404] "Relevant information" refers to data and knowledge that are presented appropriately in response to the statements and emotional states of participants and customers.

[0405] A "database" is a collection of information organized in a systematic way that allows for efficient storage, searching, and use of information.

[0406] A "proposal" is to recommend appropriate policies or actions based on the situation.

[0407] "Quality improvement" means enhancing the value and effectiveness of the services and information provided.

[0408] This invention is implemented as a system to improve customer service in physical stores. The system consists of a program for effectively utilizing various data in meetings and interactions with customers in stores. The embodiments thereof are described below.

[0409] The server first collects customer voices from smart devices in the store and obtains audio information. The acquired audio information is transcribed in real time using the Google Cloud Speech-to-Text API. This transcription result is sent to Amazon Comprehend for real-time sentiment analysis, where the customer's emotional state is analyzed. Based on the analyzed emotional state, the server searches the database for relevant information and selects the appropriate information.

[0410] The smart glasses, acting as the terminal, visually display the customer's emotional state and related information transmitted from the server in real time. This allows the store staff, acting as users, to provide service tailored to the customer's emotional state. For example, if the glasses detect that the customer is confused, information carefully explaining how to set up and use the product will be displayed on them.

[0411] As a concrete example, suppose emotion analysis detects that customer A is looking for a specific product but cannot find it and appears confused. In this case, the smart glasses would display a specific response plan, such as "Suggest to guide customer A to the location of the product on the shelf."

[0412] Furthermore, the server also includes a function to ensure that the information presented is based on the customer's emotional state. This is expected to improve the quality of customer service and increase customer satisfaction.

[0413] Example of a prompt

[0414] Convert speech to text in real time and analyze customer emotions. If a customer is having trouble, immediately suggest helpful information and display it on smart glasses.

[0415] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0416] Step 1:

[0417] The server acquires customer voices as audio data from smart devices acting as terminals. The input is real-time audio information collected from smart devices within the store. The output is audio data used in the speech recognition process.

[0418] Step 2:

[0419] The server transcribes the acquired audio data in real time using the Google Cloud Speech-to-Text API. The input is the audio data obtained in step 1, and the output is the transcribed text. This process prepares the spoken content for analysis as digital data.

[0420] Step 3:

[0421] The server passes text data to Amazon Comprehend, which analyzes the spoken content to identify the customer's emotional state. The input is the transcribed text from step 2, and the output is information about the emotional state. This procedure allows the server to understand the customer's psychological state in real time.

[0422] Step 4:

[0423] The server searches the database for appropriate relevant information based on the analyzed emotional state. The input is the emotional state information obtained in step 3, and the output is relevant information that matches the emotional state. A prompt sentence is generated through a generative AI model, and the optimal information is selected.

[0424] Step 5:

[0425] The smart glasses, acting as the terminal, visually display relevant information sent from the server to the customer. The input is the relevant information selected in step 4, and the output is information presented visually in a user-friendly format. This enables the user to respond quickly and appropriately to the customer.

[0426] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0427] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0428] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0429] [Third Embodiment]

[0430] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0431] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0432] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0433] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0434] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0435] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0436] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0437] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0438] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0439] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0440] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0441] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0442] This invention provides a means to streamline information management in meetings by constructing a system that transcribes conversations in real time, analyzes the necessary information, and presents it appropriately. This system is implemented as follows.

[0443] First, the terminal acquires the audio data of the meeting. Here, a high-sensitivity microphone is used to pick up the conversation, and the audio signal is converted into digital data. The converted audio data is then sent to a server via the network.

[0444] Next, the server analyzes the received audio data and performs real-time transcription using a speech recognition engine. This converts the audio data into text data, which is then temporarily stored in a database. To improve speech recognition accuracy, specialized terms and proper nouns can be pre-registered as a dictionary.

[0445] The server then analyzes the transcribed data to identify necessary information from the user's statements. For example, if a user is looking for similar cases during a meeting, the server searches past conversation history and case databases based on those keywords.

[0446] Once the information retrieval is complete, the server retrieves the relevant information and processes it into a user-friendly format. This information is sent to the terminal and presented to the user in real time. This allows the user to quickly obtain the information needed during the meeting. The server also has a function to automatically generate a meeting summary, and can output it as a report after the meeting based on preset templates as needed.

[0447] Finally, after the meeting concludes, users review the content via their terminals and provide feedback, offering data to improve the system's accuracy. This feedback is then used by the server to inform future processing, providing more accurate and useful information.

[0448] As a concrete example, in product development meetings, when searching for past solutions to technical problems, the system can be used to support decision-making by quickly searching a database of past projects and presenting relevant technical documents in real time. Such functionality allows users to conduct meetings more efficiently and make decisions more quickly.

[0449] The following describes the processing flow.

[0450] Step 1:

[0451] The terminal acquires audio data from the meeting. It collects sound using a high-sensitivity microphone and converts the audio signal into a digital format. The converted digital audio data is then transmitted to the server via the network.

[0452] Step 2:

[0453] The server inputs the received audio data into the speech recognition engine and performs transcription in real time. The recognized text data is temporarily stored in a database. At this stage, a specialized terminology dictionary can be used to improve the accuracy of speech recognition.

[0454] Step 3:

[0455] The server analyzes the transcribed text data and infers the information the user is looking for. For example, it extracts important keywords based on the user's statements and identifies their intent.

[0456] Step 4:

[0457] The server searches for relevant information from past databases and related resources based on identified keywords and intentions. It performs query searches to quickly retrieve similar cases and necessary documents.

[0458] Step 5:

[0459] The server organizes the search results and processes the information into a format that is easy for the user to understand. The processed information is then sent to the terminal and displayed to the user on the screen.

[0460] Step 6:

[0461] Users review the information displayed on their devices and utilize it as needed during the meeting. If a user has specific inquiries or feedback, they send that information to the server via their device.

[0462] Step 7:

[0463] The server automatically generates a meeting summary and outputs it in report format after the meeting, if necessary. It uses a pre-configured template to organize the information.

[0464] Step 8:

[0465] The server receives feedback data from users and uses it as training data to improve the accuracy of future processing. This improves the system's algorithms and enhances the accuracy of the information it provides.

[0466] (Example 1)

[0467] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0468] In today's business environment, meetings are frequently held as a forum for important decision-making, but it is often difficult to properly record what is said during meetings and to quickly obtain the necessary information. Furthermore, summarizing meetings and organizing information afterward takes a lot of time and effort, which reduces work efficiency. In addition, obtaining and presenting relevant information in real time during meetings contributes to quick and accurate decision-making, but traditional systems have the problem of only being able to utilize a limited amount of information.

[0469] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0470] In this invention, the server includes means for acquiring audio information and converting it into a digital signal using a highly sensitive receiver; means for analyzing the received audio information and converting it into text information using a speech recognition function; means for analyzing the transcribed text information, identifying necessary information from the content of speech, and searching for information sources including past meeting history; means for processing the search results and presenting them in a format that is easy for the user to understand, as well as automatically generating a summary of the meeting content; and means for collecting user feedback and using it as data to improve the system's performance. As a result, necessary information can be acquired in a timely and accurate manner during meetings, and summaries can be created efficiently after meetings, thereby improving the overall efficiency of operations.

[0471] "Audio information" refers to data that digitally represents speech and sounds from settings such as meetings.

[0472] A "receiver" is a device used to pick up sound, and its high sensitivity allows it to clearly detect conversations.

[0473] A "digital signal" is a format in which analog audio information is converted into binary data that can be processed by a computer.

[0474] "Speech recognition" refers to the technology that analyzes speech and converts it into text, and is used as part of natural language processing.

[0475] "Text information" refers to a written representation converted from audio data, and is a document format that can be processed electronically.

[0476] "Past meeting history" refers to data that records the content of meetings held in the past, and is used as the basis for information retrieval.

[0477] An "information source" is a collection of data used when searching or referencing, and it forms the basis for obtaining the necessary information.

[0478] "Search results" refer to data obtained from information sources based on specified criteria, and the content presented to the user.

[0479] "Processing" refers to the process of formatting and refining search results and information to make them easier for users to understand.

[0480] A "summary" is a short document that compiles the main points of a meeting and decisions, and it concisely conveys the outcome of the meeting.

[0481] "Automatic generation" generally refers to the process of constructing information using computer programs without human intervention.

[0482] "Opinions" refer to feedback that specifically outlines areas for improvement or suggestions that users have felt regarding the system's functions and results.

[0483] "Performance" is a general term encompassing the functions, processing speed, accuracy, and utilization efficiency of a system, and it is an indicator that should be strived to improve.

[0484] This system includes a series of steps for processing audio data to streamline information management in meetings. A specific embodiment is described below.

[0485] First, the terminal uses a highly sensitive receiver, specifically a standard microphone, to acquire audio information from the meeting. This audio information is received as an analog signal and converted into a digital signal. A computer or a dedicated encoding device is used for this process. The digitized audio information is then transmitted to a server via the network.

[0486] The server analyzes the received digital audio data and converts it into text information in real time using speech recognition technology. Specifically, this could involve utilizing an API equipped with a speech recognition engine. After conversion to text information, this data is temporarily stored in a database.

[0487] Next, the server analyzes the text information and identifies the necessary information from the content of the statements. Natural language processing technology is used to identify the information, and a search engine is used to search past meeting history and related information sources. These search results are processed into a format that is easy for the user to understand and presented to the user through the terminal.

[0488] Furthermore, the server has a function to automatically generate meeting summaries, and can improve system performance based on user feedback. This feedback is reflected in subsequent processing, resulting in further improvements in accuracy. Specifically, the system operates based on the prompt message, "We plan to discuss technical solutions at the next meeting. Please search for relevant past cases." This allows users to quickly obtain relevant information and improve the efficiency of meetings.

[0489] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0490] Step 1:

[0491] The terminal uses a highly sensitive microphone installed in the conference room to acquire real-time audio information from the meeting. This input is an analog audio signal. The audio signal is converted to a digital signal and transmitted to the server via the network as PCM digital data. Digitization makes the audio format usable by a computer.

[0492] Step 2:

[0493] The server receives digital audio information as input and converts it into text information using a speech recognition engine. The speech recognition engine analyzes the audio signal using a specific algorithm and transcribes it into words and sentences. The resulting output is transcribed text data, which is then temporarily stored in a database.

[0494] Step 3:

[0495] The server analyzes the transcribed text data. The input is the text data obtained in step 2. Using natural language processing, important keywords and phrases are extracted from the user's speech. This extracted information serves as foundational data for searching past meeting histories and related documents. The output is a list of the identified information.

[0496] Step 4:

[0497] The server uses the keywords obtained in the previous step to search for information sources within the database. The input is the identified keywords, and the search engine quickly extracts past meeting materials and related information. The output is a set of information that the user needs, and this set is then prepared for further presentation to the user.

[0498] Step 5:

[0499] The server processes search results into a format that is easy for the user to understand. The input is a set of information obtained from the search engine, which is then formatted appropriately, including summarizing and highlighting. The output is the formatted information, which is then sent to the terminal. This allows the user to efficiently obtain the information they need.

[0500] Step 6:

[0501] The server automatically generates a summary after the meeting concludes. The input is data collected during the meeting; a generation AI model extracts key topics and points, which are then compiled into a report format based on a specified template. The output is a summary report of the entire meeting.

[0502] Step 7:

[0503] Users provide feedback through their devices. Input consists of comments and suggestions regarding the meeting content and system output, which are sent to the server and used to improve speech recognition accuracy and information retrieval accuracy in future sessions. Output consists of insights that contribute to improving system performance.

[0504] (Application Example 1)

[0505] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0506] In typical brick-and-mortar stores, it is difficult to respond quickly to customer inquiries and immediately present products that customers are interested in. Therefore, there is a need to improve customer service. Furthermore, store staff have difficulty memorizing vast amounts of product information, resulting in delays in providing appropriate information.

[0507] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0508] In this invention, the server includes means for acquiring acoustic information and performing real-time text conversion, a computer device for analyzing the user's statements and immediately presenting product information, and means for providing information so that the user can confirm the information using a visual device. This makes it possible to provide product information immediately based on the content of the customer's statements.

[0509] "Acoustic information" refers to audible data such as voices and sound effects, which are acquired via input devices such as microphones.

[0510] "Real-time text conversion" refers to the process of instantly converting acquired audio information into text information, indicating that time delays are minimized.

[0511] "User" refers to an individual or organization that uses a system or service, and is primarily an entity that interacts with the system through voice.

[0512] "Analyzing the content of a user's statements" refers to analyzing the linguistic and semantic elements contained in the user's statements, with the aim of extracting necessary information within the system.

[0513] A "computer device that instantly displays product information" refers to a device that has the function of quickly acquiring, processing, and displaying product-related information according to the user's needs.

[0514] "Visual devices" are devices used to visually present textual information and images to users, and include displays such as smart glasses.

[0515] "Means of providing information" refers to a system function that outputs the information a user needs in a format that the user can view, after a series of processing steps.

[0516] As a concrete example of the present invention, a system for improving customer service in a physical store is implemented. This system includes a high-sensitivity microphone, smart glasses, a server, and a database management system.

[0517] First, the smart glasses, which serve as the device, have a built-in high-sensitivity microphone that constantly acquires ambient acoustic information. The acquired audio data is immediately digitized and sent to a server via the network. The server uses the Google Cloud Speech-to-Text API to convert the audio data into text in real time.

[0518] Next, the text data of the spoken content is analyzed using a generative AI model, and relevant product information is searched from the database. At this time, the server extracts keywords from the spoken content, and the computer immediately processes and provides product information based on those keywords.

[0519] The information is displayed on the smart glasses' screen, allowing store staff (users) to visually confirm it and appropriately guide customers with product information. For example, if a customer says, "I want a camera that's perfect for travel," the server searches for relevant camera information and displays it on the smart glasses. The user can then use this information to immediately explain the most suitable product to the customer.

[0520] In embodiments of the present invention, an example of a prompt sentence based on the generated AI model is, "Analyze the voice of a customer who spoke about a camera suitable for travel, and generate a prompt sentence that suggests relevant information." This prompt enables the system to provide information that immediately responds to the customer's needs.

[0521] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0522] Step 1:

[0523] The smart glasses, which act as the device, acquire acoustic information using a high-sensitivity microphone. This allows surrounding conversations and sounds to be collected as digital data. The input is an analog audio signal, and the output is audio data in digital format.

[0524] Step 2:

[0525] The terminal transmits the acquired digital audio data to the server via the network. Here, the audio data is converted into packets and transmitted to the server. The input is digital audio data, and the output is the data received by the server.

[0526] Step 3:

[0527] The server uses the Google Cloud Speech-to-Text API to convert incoming audio data into text in real time. The speech recognition engine analyzes phonemes and generates text data. The input is audio data sent to the server, and the output is data in text format.

[0528] Step 4:

[0529] The server uses a generative AI model to extract keywords from text data and analyze the content of the speech. Here, it identifies important words and phrases within the text data. The input is text data, and the output is a list of extracted keywords.

[0530] Step 5:

[0531] The server uses a database management system to search for product information based on extracted keywords. Relevant product information is retrieved instantly via database queries. The input is a list of keywords, and the output is the corresponding product information.

[0532] Step 6:

[0533] The server converts the acquired product information into a format that users can visually verify and transmits it to the terminal via a computer. Data optimized for display is generated. The input is product information, and the output is a displayable data format.

[0534] Step 7:

[0535] The smart glasses on the device display product information transmitted from the server. Store staff, who are the users, can check the information in real time and use it to assist customers. Input is in a displayable data format, and output is information confirmation by the user's eyes.

[0536] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0537] This invention provides a system that utilizes audio data from meetings to recognize participants' emotions and optimize information delivery. This system consists of real-time transcription of audio data, information retrieval from a database, information presentation, automatic summarization of meeting content, and an emotion recognition engine.

[0538] First, the terminal acquires audio data during the meeting and sends this audio data to the server. The server uses a speech recognition engine to transcribe the audio into text and passes this text data to an emotion engine, which analyzes the emotional state of the participants based on their statements. The emotion engine uses elements such as tone, speed, and language patterns to determine the speaker's emotions (e.g., excitement, calmness, anger).

[0539] Once an emotion is recognized, the server optimizes the information provided to support the meeting based on the acquired emotion information. For example, if the system detects that the user is stressed, it will simplify the information it suggests or visually highlight meeting materials to make the information easier for the user to understand.

[0540] Furthermore, the server creates a summary of the meeting content at the end of the meeting, reflecting the user's emotional state. This ensures that feedback and important points are appropriately highlighted, and the report is provided in a way that aligns with the user's emotional state.

[0541] For example, in a kick-off meeting for a new project, the emotion engine may detect that the presenter is feeling anxious about new technical challenges. Based on this emotional information, the system immediately searches for relevant technical data and past success stories, presenting them instantly to alleviate anxiety and help the meeting proceed smoothly.

[0542] Thus, the present invention supports improved meeting environments and efficient decision-making processes through an information provision system that combines emotion recognition.

[0543] The following describes the processing flow.

[0544] Step 1:

[0545] The terminal acquires audio data from the meeting. The audio data is collected using a high-sensitivity microphone, and the digital audio signal is transmitted to the server via the network.

[0546] Step 2:

[0547] The server passes the received audio data to the speech recognition engine, which performs transcription in real time. The generated text data is then supplied to the sentiment engine and the database search engine.

[0548] Step 3:

[0549] The server uses an emotion engine to analyze the speaker's emotions from text data. It uses voice tone, speaking speed, and linguistic features to determine the speaker's emotional state and records the emotional information.

[0550] Step 4:

[0551] Based on the sentiment analysis results, the server searches the database for relevant information to provide the user with appropriate information. If necessary, it also collects similar past cases and related materials.

[0552] Step 5:

[0553] The server uses emotional information to optimize information display. For example, if a user is confused, it simplifies the information presented or changes the interface to make it more visually understandable.

[0554] Step 6:

[0555] The device receives optimized information from the server and displays it to the user. The way the information is presented is tailored to the user's current emotional state and is adjusted for easy understanding.

[0556] Step 7:

[0557] Users review the information presented via their device and provide feedback as needed. This feedback is sent from the device to the server to be used for future system improvements.

[0558] Step 8:

[0559] At the end of the meeting, the server creates a summary that reflects sentiment information and outputs it as an automatically generated meeting report. This allows users to review key points from the meeting based on sentiment data.

[0560] (Example 2)

[0561] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0562] The challenge lies in efficiently utilizing audio data from meetings to understand participants' emotional states, thereby optimizing meeting progress and providing an environment where participants can quickly and accurately grasp information. Furthermore, it is crucial to appropriately summarize the meeting content and provide it as a useful resource for subsequent reviews and decision-making.

[0563] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0564] In this invention, the server includes means for acquiring audio signals and performing data conversion in real time, means for analyzing emotional states from the content and characteristics of speech and retrieving relevant information from an external storage device, and means for optimizing information presentation based on the analyzed emotional information and providing it to the user, as well as automatically generating a summary of the meeting content. This makes it possible to improve the efficiency of meetings and the quality of decision-making by analyzing participants' emotions in real time during meetings and providing appropriate information.

[0565] An "audio signal" is an electrical representation of audio information, treating sound vibrations as signals that change along the time axis.

[0566] "Data conversion" is the process of converting audio signals into other formats, such as text data, in real time.

[0567] "Content and characteristics of speech" refers to the meaning of words within the audio data and the characteristics that indicate the speaker's emotional nuances.

[0568] "Analyzing emotional states" is the process of estimating a speaker's emotions from audio or text data and identifying that state.

[0569] "External storage device" refers to a storage medium used to store data for a long period of time, and includes, but is not limited to, databases.

[0570] "Optimizing information presentation" refers to displaying and providing information in the most appropriate format based on the user's situation and emotions.

[0571] "Users" refers to anyone who uses the system and receives information, including meeting participants and general individuals who receive information.

[0572] A "meeting summary" is a compilation of the key points and decisions discussed at a meeting, reorganized into a concise format for later use or reference.

[0573] This invention is a system that optimizes the progress of a meeting and streamlines information presentation by acquiring audio signals during a meeting, converting those signals into data in real time, and analyzing the emotional state of the speakers.

[0574] The terminal uses a microphone device to acquire the user's audio signal during the meeting. This audio signal is transmitted to the server via the network.

[0575] The server uses speech recognition software (e.g., a general speech recognition engine) to convert the received audio signal into text format. The converted text data is sent to an emotion analysis engine (e.g., a natural language processing engine) to analyze the participant's emotions based on the content and characteristics of their speech.

[0576] Based on the analyzed emotional information, the server retrieves the most relevant information from external storage devices and provides it to the user, tailored to the meeting content. Specifically, if the emotional state is one of anxiety, the server presents relevant technical documents and past success stories to improve meeting efficiency.

[0577] Users can view information provided in real time and provide feedback. This feedback is aggregated on the server as data to improve the accuracy of the system's analysis.

[0578] A concrete example would be a meeting to launch a new project, where technical challenges are addressed. If the presenter expresses concerns about these challenges, the system can instantly provide documentation of past cases and solutions, reducing anxiety and ensuring the meeting proceeds smoothly.

[0579] Examples of prompts to input into a generative AI model include: "Design a system that recognizes the emotions of meeting participants in real time and presents the most relevant information based on those emotions. For example, if a participant is feeling anxious about a technical issue, present them with relevant data."

[0580] This system improves the meeting environment and significantly enhances the efficiency of the decision-making process.

[0581] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0582] Step 1:

[0583] The terminal uses a microphone to acquire audio signals in real time during a meeting. The acquired audio signals are then transmitted directly to the server via the network. Its input is raw audio data, and its output is digital audio data that can be transferred to the server.

[0584] Step 2:

[0585] The server converts the received audio signal into text data using a speech recognition engine (e.g., general speech recognition software). In this process, the input is digital audio data, and the output is transcribed text data. Specifically, the process involves analyzing the audio waveform along the time axis and converting phonemes into text strings.

[0586] Step 3:

[0587] The transcribed text data is sent to an emotion analysis engine on the server and used for analysis to determine the speaker's emotional state. The input is text data, and the output is data tagged with the speaker's emotional state (e.g., relief, excitement, doubt). In this step, natural language processing techniques are applied to analyze language patterns and extract keywords.

[0588] Step 4:

[0589] Based on the analyzed emotional state, the server initiates a process to retrieve appropriate information from external storage and provide efficient information presentation. The input is emotional state-tagged data, and the output is a list of relevant information matching the emotion. Specifically, it executes database queries to list examples and literature.

[0590] Step 5:

[0591] During a meeting, relevant information is displayed in real time on the user's device (terminal). The input is a list of relevant information, and the output provides the user with visually organized information. The system utilizes a graphical user interface to visually highlight and categorize the information.

[0592] Step 6:

[0593] The server automatically generates a summary of the meeting content based on data including changes in emotional states obtained after the meeting. The input is all meeting data and emotional information, and the output is a concise meeting report summarizing the key points. Specific operations include selecting high-priority topics and analyzing emotional changes over time.

[0594] (Application Example 2)

[0595] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0596] In modern meetings and customer service, accurately understanding the emotions of participants and customers and responding appropriately is essential. However, traditional systems have struggled to analyze emotional states in real time from tone of voice and content of speech, and to provide optimal information. Furthermore, the lack of emotion-based suggestions has made it difficult to improve the quality of customer service. These challenges need to be addressed.

[0597] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0598] In this invention, the server includes means for acquiring audio information and transcribing it in real time, means for analyzing the emotional state of participants from their statements, and means for retrieving and presenting relevant information from a database according to the emotional state. This enables the provision of appropriate information according to the emotional state of participants and improves the quality of customer service.

[0599] "Audio information" refers to data on the content of speech uttered by participants and customers, collected during meetings and interactions with customers.

[0600] "Real-time" refers to technology that processes and analyzes events as they occur.

[0601] "Transcription" is a technology that converts audio information into text format.

[0602] "Participants" are individuals involved in a meeting or discussion.

[0603] "Statements" are words expressed orally during meetings or discussions.

[0604] "Emotional state" refers to the psychological state and emotions of participants or customers.

[0605] "Analysis" is the process of analyzing data and information to understand their meaning and patterns.

[0606] "Relevant information" refers to data and knowledge that are presented appropriately in response to the statements and emotional states of participants and customers.

[0607] A "database" is a collection of information organized in a systematic way that allows for efficient storage, searching, and use of information.

[0608] A "proposal" is to recommend appropriate policies or actions based on the situation.

[0609] "Quality improvement" means enhancing the value and effectiveness of the services and information provided.

[0610] This invention is implemented as a system to improve customer service in physical stores. The system consists of a program for effectively utilizing various data in meetings and interactions with customers in stores. The embodiments thereof are described below.

[0611] The server first collects customer voices from smart devices in the store and obtains audio information. The acquired audio information is transcribed in real time using the Google Cloud Speech-to-Text API. This transcription result is sent to Amazon Comprehend for real-time sentiment analysis, where the customer's emotional state is analyzed. Based on the analyzed emotional state, the server searches the database for relevant information and selects the appropriate information.

[0612] The smart glasses, acting as the terminal, visually display the customer's emotional state and related information transmitted from the server in real time. This allows the store staff, acting as users, to provide service tailored to the customer's emotional state. For example, if the glasses detect that the customer is confused, information carefully explaining how to set up and use the product will be displayed on them.

[0613] As a concrete example, suppose emotion analysis detects that customer A is looking for a specific product but cannot find it and appears confused. In this case, the smart glasses would display a specific response plan, such as "Suggest to guide customer A to the location of the product on the shelf."

[0614] Furthermore, the server also includes a function to ensure that the information presented is based on the customer's emotional state. This is expected to improve the quality of customer service and increase customer satisfaction.

[0615] Example of a prompt

[0616] Convert speech to text in real time and analyze customer emotions. If a customer is having trouble, immediately suggest helpful information and display it on smart glasses.

[0617] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0618] Step 1:

[0619] The server acquires customer voices as audio data from smart devices acting as terminals. The input is real-time audio information collected from smart devices within the store. The output is audio data used in the speech recognition process.

[0620] Step 2:

[0621] The server transcribes the acquired audio data in real time using the Google Cloud Speech-to-Text API. The input is the audio data obtained in step 1, and the output is the transcribed text. This process prepares the spoken content for analysis as digital data.

[0622] Step 3:

[0623] The server passes text data to Amazon Comprehend, which analyzes the spoken content to identify the customer's emotional state. The input is the transcribed text from step 2, and the output is information about the emotional state. This procedure allows the server to understand the customer's psychological state in real time.

[0624] Step 4:

[0625] The server searches the database for appropriate relevant information based on the analyzed emotional state. The input is the emotional state information obtained in step 3, and the output is relevant information that matches the emotional state. A prompt sentence is generated through a generative AI model, and the optimal information is selected.

[0626] Step 5:

[0627] The smart glasses, acting as the terminal, visually display relevant information sent from the server to the customer. The input is the relevant information selected in step 4, and the output is information presented visually in a user-friendly format. This enables the user to respond quickly and appropriately to the customer.

[0628] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0629] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0630] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0631] [Fourth Embodiment]

[0632] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0633] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0634] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0635] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0636] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0637] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0638] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0639] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0640] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0641] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0642] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0643] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0644] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0645] This invention provides a means to streamline information management in meetings by constructing a system that transcribes conversations in real time, analyzes the necessary information, and presents it appropriately. This system is implemented as follows.

[0646] First, the terminal acquires the audio data of the meeting. Here, a high-sensitivity microphone is used to pick up the conversation, and the audio signal is converted into digital data. The converted audio data is then sent to a server via the network.

[0647] Next, the server analyzes the received audio data and performs real-time transcription using a speech recognition engine. This converts the audio data into text data, which is then temporarily stored in a database. To improve speech recognition accuracy, specialized terms and proper nouns can be pre-registered as a dictionary.

[0648] The server then analyzes the transcribed data to identify necessary information from the user's statements. For example, if a user is looking for similar cases during a meeting, the server searches past conversation history and case databases based on those keywords.

[0649] Once the information retrieval is complete, the server retrieves the relevant information and processes it into a user-friendly format. This information is sent to the terminal and presented to the user in real time. This allows the user to quickly obtain the information needed during the meeting. The server also has a function to automatically generate a meeting summary, and can output it as a report after the meeting based on preset templates as needed.

[0650] Finally, after the meeting concludes, users review the content via their terminals and provide feedback, offering data to improve the system's accuracy. This feedback is then used by the server to inform future processing, providing more accurate and useful information.

[0651] As a concrete example, in product development meetings, when searching for past solutions to technical problems, the system can be used to support decision-making by quickly searching a database of past projects and presenting relevant technical documents in real time. Such functionality allows users to conduct meetings more efficiently and make decisions more quickly.

[0652] The following describes the processing flow.

[0653] Step 1:

[0654] The terminal acquires audio data from the meeting. It collects sound using a high-sensitivity microphone and converts the audio signal into a digital format. The converted digital audio data is then transmitted to the server via the network.

[0655] Step 2:

[0656] The server inputs the received audio data into the speech recognition engine and performs transcription in real time. The recognized text data is temporarily stored in a database. At this stage, a specialized terminology dictionary can be used to improve the accuracy of speech recognition.

[0657] Step 3:

[0658] The server analyzes the transcribed text data and infers the information the user is looking for. For example, it extracts important keywords based on the user's statements and identifies their intent.

[0659] Step 4:

[0660] The server searches for relevant information from past databases and related resources based on identified keywords and intentions. It performs query searches to quickly retrieve similar cases and necessary documents.

[0661] Step 5:

[0662] The server organizes the search results and processes the information into a format that is easy for the user to understand. The processed information is then sent to the terminal and displayed to the user on the screen.

[0663] Step 6:

[0664] Users review the information displayed on their devices and utilize it as needed during the meeting. If a user has specific inquiries or feedback, they send that information to the server via their device.

[0665] Step 7:

[0666] The server automatically generates a meeting summary and outputs it in report format after the meeting, if necessary. It uses a pre-configured template to organize the information.

[0667] Step 8:

[0668] The server receives feedback data from users and uses it as training data to improve the accuracy of future processing. This improves the system's algorithms and enhances the accuracy of the information it provides.

[0669] (Example 1)

[0670] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0671] In today's business environment, meetings are frequently held as a forum for important decision-making, but it is often difficult to properly record what is said during meetings and to quickly obtain the necessary information. Furthermore, summarizing meetings and organizing information afterward takes a lot of time and effort, which reduces work efficiency. In addition, obtaining and presenting relevant information in real time during meetings contributes to quick and accurate decision-making, but traditional systems have the problem of only being able to utilize a limited amount of information.

[0672] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0673] In this invention, the server includes means for acquiring audio information and converting it into a digital signal using a highly sensitive receiver; means for analyzing the received audio information and converting it into text information using a speech recognition function; means for analyzing the transcribed text information, identifying necessary information from the content of speech, and searching for information sources including past meeting history; means for processing the search results and presenting them in a format that is easy for the user to understand, as well as automatically generating a summary of the meeting content; and means for collecting user feedback and using it as data to improve the system's performance. As a result, necessary information can be acquired in a timely and accurate manner during meetings, and summaries can be created efficiently after meetings, thereby improving the overall efficiency of operations.

[0674] "Audio information" refers to data that digitally represents speech and sounds from settings such as meetings.

[0675] A "receiver" is a device used to pick up sound, and its high sensitivity allows it to clearly detect conversations.

[0676] A "digital signal" is a format in which analog audio information is converted into binary data that can be processed by a computer.

[0677] "Speech recognition" refers to the technology that analyzes speech and converts it into text, and is used as part of natural language processing.

[0678] "Text information" refers to a written representation converted from audio data, and is a document format that can be processed electronically.

[0679] "Past meeting history" refers to data that records the content of meetings held in the past, and is used as the basis for information retrieval.

[0680] An "information source" is a collection of data used when searching or referencing, and it forms the basis for obtaining the necessary information.

[0681] "Search results" refer to data obtained from information sources based on specified criteria, and the content presented to the user.

[0682] "Processing" refers to the process of formatting and refining search results and information to make them easier for users to understand.

[0683] A "summary" is a short document that compiles the main points of a meeting and decisions, and it concisely conveys the outcome of the meeting.

[0684] "Automatic generation" generally refers to the process of constructing information using computer programs without human intervention.

[0685] "Opinions" refer to feedback that specifically outlines areas for improvement or suggestions that users have felt regarding the system's functions and results.

[0686] "Performance" is a general term encompassing the functions, processing speed, accuracy, and utilization efficiency of a system, and it is an indicator that should be strived to improve.

[0687] This system includes a series of steps for processing audio data to streamline information management in meetings. A specific embodiment is described below.

[0688] First, the terminal uses a highly sensitive receiver, specifically a standard microphone, to acquire audio information from the meeting. This audio information is received as an analog signal and converted into a digital signal. A computer or a dedicated encoding device is used for this process. The digitized audio information is then transmitted to a server via the network.

[0689] The server analyzes the received digital audio data and converts it into text information in real time using speech recognition technology. Specifically, this could involve utilizing an API equipped with a speech recognition engine. After conversion to text information, this data is temporarily stored in a database.

[0690] Next, the server analyzes the text information and identifies the necessary information from the content of the statements. Natural language processing technology is used to identify the information, and a search engine is used to search past meeting history and related information sources. These search results are processed into a format that is easy for the user to understand and presented to the user through the terminal.

[0691] Furthermore, the server has a function to automatically generate meeting summaries, and can improve system performance based on user feedback. This feedback is reflected in subsequent processing, resulting in further improvements in accuracy. Specifically, the system operates based on the prompt message, "We plan to discuss technical solutions at the next meeting. Please search for relevant past cases." This allows users to quickly obtain relevant information and improve the efficiency of meetings.

[0692] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0693] Step 1:

[0694] The terminal uses a highly sensitive microphone installed in the conference room to acquire real-time audio information from the meeting. This input is an analog audio signal. The audio signal is converted to a digital signal and transmitted to the server via the network as PCM digital data. Digitization makes the audio format usable by a computer.

[0695] Step 2:

[0696] The server receives digital audio information as input and converts it into text information using a speech recognition engine. The speech recognition engine analyzes the audio signal using a specific algorithm and transcribes it into words and sentences. The resulting output is transcribed text data, which is then temporarily stored in a database.

[0697] Step 3:

[0698] The server analyzes the transcribed text data. The input is the text data obtained in step 2. Using natural language processing, important keywords and phrases are extracted from the user's speech. This extracted information serves as foundational data for searching past meeting histories and related documents. The output is a list of the identified information.

[0699] Step 4:

[0700] The server uses the keywords obtained in the previous step to search for information sources within the database. The input is the identified keywords, and the search engine quickly extracts past meeting materials and related information. The output is a set of information that the user needs, and this set is then prepared for further presentation to the user.

[0701] Step 5:

[0702] The server processes search results into a format that is easy for the user to understand. The input is a set of information obtained from the search engine, which is then formatted appropriately, including summarizing and highlighting. The output is the formatted information, which is then sent to the terminal. This allows the user to efficiently obtain the information they need.

[0703] Step 6:

[0704] The server automatically generates a summary after the meeting concludes. The input is data collected during the meeting; a generation AI model extracts key topics and points, which are then compiled into a report format based on a specified template. The output is a summary report of the entire meeting.

[0705] Step 7:

[0706] Users provide feedback through their devices. Input consists of comments and suggestions regarding the meeting content and system output, which are sent to the server and used to improve speech recognition accuracy and information retrieval accuracy in future sessions. Output consists of insights that contribute to improving system performance.

[0707] (Application Example 1)

[0708] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0709] In typical brick-and-mortar stores, it is difficult to respond quickly to customer inquiries and immediately present products that customers are interested in. Therefore, there is a need to improve customer service. Furthermore, store staff have difficulty memorizing vast amounts of product information, resulting in delays in providing appropriate information.

[0710] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0711] In this invention, the server includes means for acquiring acoustic information and performing real-time text conversion, a computer device for analyzing the user's statements and immediately presenting product information, and means for providing information so that the user can confirm the information using a visual device. This makes it possible to provide product information immediately based on the content of the customer's statements.

[0712] "Acoustic information" refers to audible data such as voices and sound effects, which are acquired via input devices such as microphones.

[0713] "Real-time text conversion" refers to the process of instantly converting acquired audio information into text information, indicating that time delays are minimized.

[0714] "User" refers to an individual or organization that uses a system or service, and is primarily an entity that interacts with the system through voice.

[0715] "Analyzing the content of a user's statements" refers to analyzing the linguistic and semantic elements contained in the user's statements, with the aim of extracting necessary information within the system.

[0716] A "computer device that instantly displays product information" refers to a device that has the function of quickly acquiring, processing, and displaying product-related information according to the user's needs.

[0717] "Visual devices" are devices used to visually present textual information and images to users, and include displays such as smart glasses.

[0718] "Means of providing information" refers to a system function that outputs the information a user needs in a format that the user can view, after a series of processing steps.

[0719] As a concrete example of the present invention, a system for improving customer service in a physical store is implemented. This system includes a high-sensitivity microphone, smart glasses, a server, and a database management system.

[0720] First, the smart glasses, which serve as the device, have a built-in high-sensitivity microphone that constantly acquires ambient acoustic information. The acquired audio data is immediately digitized and sent to a server via the network. The server uses the Google Cloud Speech-to-Text API to convert the audio data into text in real time.

[0721] Next, the text data of the spoken content is analyzed using a generative AI model, and relevant product information is searched from the database. At this time, the server extracts keywords from the spoken content, and the computer immediately processes and provides product information based on those keywords.

[0722] The information is displayed on the smart glasses' screen, allowing store staff (users) to visually confirm it and appropriately guide customers with product information. For example, if a customer says, "I want a camera that's perfect for travel," the server searches for relevant camera information and displays it on the smart glasses. The user can then use this information to immediately explain the most suitable product to the customer.

[0723] In embodiments of the present invention, an example of a prompt sentence based on the generated AI model is, "Analyze the voice of a customer who spoke about a camera suitable for travel, and generate a prompt sentence that suggests relevant information." This prompt enables the system to provide information that immediately responds to the customer's needs.

[0724] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0725] Step 1:

[0726] The smart glasses, which act as the device, acquire acoustic information using a high-sensitivity microphone. This allows surrounding conversations and sounds to be collected as digital data. The input is an analog audio signal, and the output is audio data in digital format.

[0727] Step 2:

[0728] The terminal transmits the acquired digital audio data to the server via the network. Here, the audio data is converted into packets and transmitted to the server. The input is digital audio data, and the output is the data received by the server.

[0729] Step 3:

[0730] The server uses the Google Cloud Speech-to-Text API to convert incoming audio data into text in real time. The speech recognition engine analyzes phonemes and generates text data. The input is audio data sent to the server, and the output is data in text format.

[0731] Step 4:

[0732] The server uses a generative AI model to extract keywords from text data and analyze the content of the speech. Here, it identifies important words and phrases within the text data. The input is text data, and the output is a list of extracted keywords.

[0733] Step 5:

[0734] The server uses a database management system to search for product information based on extracted keywords. Relevant product information is retrieved instantly via database queries. The input is a list of keywords, and the output is the corresponding product information.

[0735] Step 6:

[0736] The server converts the acquired product information into a format that users can visually verify and transmits it to the terminal via a computer. Data optimized for display is generated. The input is product information, and the output is a displayable data format.

[0737] Step 7:

[0738] The smart glasses on the device display product information transmitted from the server. Store staff, who are the users, can check the information in real time and use it to assist customers. Input is in a displayable data format, and output is information confirmation by the user's eyes.

[0739] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0740] This invention provides a system that utilizes audio data from meetings to recognize participants' emotions and optimize information delivery. This system consists of real-time transcription of audio data, information retrieval from a database, information presentation, automatic summarization of meeting content, and an emotion recognition engine.

[0741] First, the terminal acquires audio data during the meeting and sends this audio data to the server. The server uses a speech recognition engine to transcribe the audio into text and passes this text data to an emotion engine, which analyzes the emotional state of the participants based on their statements. The emotion engine uses elements such as tone, speed, and language patterns to determine the speaker's emotions (e.g., excitement, calmness, anger).

[0742] Once an emotion is recognized, the server optimizes the information provided to support the meeting based on the acquired emotion information. For example, if the system detects that the user is stressed, it will simplify the information it suggests or visually highlight meeting materials to make the information easier for the user to understand.

[0743] Furthermore, the server creates a summary of the meeting content at the end of the meeting, reflecting the user's emotional state. This ensures that feedback and important points are appropriately highlighted, and the report is provided in a way that aligns with the user's emotional state.

[0744] For example, in a kick-off meeting for a new project, the emotion engine may detect that the presenter is feeling anxious about new technical challenges. Based on this emotional information, the system immediately searches for relevant technical data and past success stories, presenting them instantly to alleviate anxiety and help the meeting proceed smoothly.

[0745] Thus, the present invention supports improved meeting environments and efficient decision-making processes through an information provision system that combines emotion recognition.

[0746] The following describes the processing flow.

[0747] Step 1:

[0748] The terminal acquires audio data from the meeting. The audio data is collected using a high-sensitivity microphone, and the digital audio signal is transmitted to the server via the network.

[0749] Step 2:

[0750] The server passes the received audio data to the speech recognition engine, which performs transcription in real time. The generated text data is then supplied to the sentiment engine and the database search engine.

[0751] Step 3:

[0752] The server uses an emotion engine to analyze the speaker's emotions from text data. It uses voice tone, speaking speed, and linguistic features to determine the speaker's emotional state and records the emotional information.

[0753] Step 4:

[0754] Based on the sentiment analysis results, the server searches the database for relevant information to provide the user with appropriate information. If necessary, it also collects similar past cases and related materials.

[0755] Step 5:

[0756] The server uses emotional information to optimize information display. For example, if a user is confused, it simplifies the information presented or changes the interface to make it more visually understandable.

[0757] Step 6:

[0758] The device receives optimized information from the server and displays it to the user. The way the information is presented is tailored to the user's current emotional state and is adjusted for easy understanding.

[0759] Step 7:

[0760] Users review the information presented via their device and provide feedback as needed. This feedback is sent from the device to the server to be used for future system improvements.

[0761] Step 8:

[0762] At the end of the meeting, the server creates a summary that reflects sentiment information and outputs it as an automatically generated meeting report. This allows users to review key points from the meeting based on sentiment data.

[0763] (Example 2)

[0764] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0765] The challenge lies in efficiently utilizing audio data from meetings to understand participants' emotional states, thereby optimizing meeting progress and providing an environment where participants can quickly and accurately grasp information. Furthermore, it is crucial to appropriately summarize the meeting content and provide it as a useful resource for subsequent reviews and decision-making.

[0766] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0767] In this invention, the server includes means for acquiring audio signals and performing data conversion in real time, means for analyzing emotional states from the content and characteristics of speech and retrieving relevant information from an external storage device, and means for optimizing information presentation based on the analyzed emotional information and providing it to the user, as well as automatically generating a summary of the meeting content. This makes it possible to improve the efficiency of meetings and the quality of decision-making by analyzing participants' emotions in real time during meetings and providing appropriate information.

[0768] An "audio signal" is an electrical representation of audio information, treating sound vibrations as signals that change along the time axis.

[0769] "Data conversion" is the process of converting audio signals into other formats, such as text data, in real time.

[0770] "Content and characteristics of speech" refers to the meaning of words within the audio data and the characteristics that indicate the speaker's emotional nuances.

[0771] "Analyzing emotional states" is the process of estimating a speaker's emotions from audio or text data and identifying that state.

[0772] "External storage device" refers to a storage medium used to store data for a long period of time, and includes, but is not limited to, databases.

[0773] "Optimizing information presentation" refers to displaying and providing information in the most appropriate format based on the user's situation and emotions.

[0774] "Users" refers to anyone who uses the system and receives information, including meeting participants and general individuals who receive information.

[0775] A "meeting summary" is a compilation of the key points and decisions discussed at a meeting, reorganized into a concise format for later use or reference.

[0776] This invention is a system that optimizes the progress of a meeting and streamlines information presentation by acquiring audio signals during a meeting, converting those signals into data in real time, and analyzing the emotional state of the speakers.

[0777] The terminal uses a microphone device to acquire the user's audio signal during the meeting. This audio signal is transmitted to the server via the network.

[0778] The server uses speech recognition software (e.g., a general speech recognition engine) to convert the received audio signal into text format. The converted text data is sent to an emotion analysis engine (e.g., a natural language processing engine) to analyze the participant's emotions based on the content and characteristics of their speech.

[0779] Based on the analyzed emotional information, the server retrieves the most relevant information from external storage devices and provides it to the user, tailored to the meeting content. Specifically, if the emotional state is one of anxiety, the server presents relevant technical documents and past success stories to improve meeting efficiency.

[0780] Users can view information provided in real time and provide feedback. This feedback is aggregated on the server as data to improve the accuracy of the system's analysis.

[0781] A concrete example would be a meeting to launch a new project, where technical challenges are addressed. If the presenter expresses concerns about these challenges, the system can instantly provide documentation of past cases and solutions, reducing anxiety and ensuring the meeting proceeds smoothly.

[0782] Examples of prompts to input into a generative AI model include: "Design a system that recognizes the emotions of meeting participants in real time and presents the most relevant information based on those emotions. For example, if a participant is feeling anxious about a technical issue, present them with relevant data."

[0783] This system improves the meeting environment and significantly enhances the efficiency of the decision-making process.

[0784] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0785] Step 1:

[0786] The terminal uses a microphone to acquire audio signals in real time during a meeting. The acquired audio signals are then transmitted directly to the server via the network. Its input is raw audio data, and its output is digital audio data that can be transferred to the server.

[0787] Step 2:

[0788] The server converts the received audio signal into text data using a speech recognition engine (e.g., general speech recognition software). In this process, the input is digital audio data, and the output is transcribed text data. Specifically, the process involves analyzing the audio waveform along the time axis and converting phonemes into text strings.

[0789] Step 3:

[0790] The transcribed text data is sent to an emotion analysis engine on the server and used for analysis to determine the speaker's emotional state. The input is text data, and the output is data tagged with the speaker's emotional state (e.g., relief, excitement, doubt). In this step, natural language processing techniques are applied to analyze language patterns and extract keywords.

[0791] Step 4:

[0792] Based on the analyzed emotional state, the server initiates a process to retrieve appropriate information from external storage and provide efficient information presentation. The input is emotional state-tagged data, and the output is a list of relevant information matching the emotion. Specifically, it executes database queries to list examples and literature.

[0793] Step 5:

[0794] During a meeting, relevant information is displayed in real time on the user's device (terminal). The input is a list of relevant information, and the output provides the user with visually organized information. The system utilizes a graphical user interface to visually highlight and categorize the information.

[0795] Step 6:

[0796] The server automatically generates a summary of the meeting content based on data including changes in emotional states obtained after the meeting. The input is all meeting data and emotional information, and the output is a concise meeting report summarizing the key points. Specific operations include selecting high-priority topics and analyzing emotional changes over time.

[0797] (Application Example 2)

[0798] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0799] In modern meetings and customer service, accurately understanding the emotions of participants and customers and responding appropriately is essential. However, traditional systems have struggled to analyze emotional states in real time from tone of voice and content of speech, and to provide optimal information. Furthermore, the lack of emotion-based suggestions has made it difficult to improve the quality of customer service. These challenges need to be addressed.

[0800] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0801] In this invention, the server includes means for acquiring audio information and transcribing it in real time, means for analyzing the emotional state of participants from their statements, and means for retrieving and presenting relevant information from a database according to the emotional state. This enables the provision of appropriate information according to the emotional state of participants and improves the quality of customer service.

[0802] "Audio information" refers to data on the content of speech uttered by participants and customers, collected during meetings and interactions with customers.

[0803] "Real-time" refers to technology that processes and analyzes events as they occur.

[0804] "Transcription" is a technology that converts audio information into text format.

[0805] "Participants" are individuals involved in a meeting or discussion.

[0806] "Statements" are words expressed orally during meetings or discussions.

[0807] "Emotional state" refers to the psychological state and emotions of participants or customers.

[0808] "Analysis" is the process of analyzing data and information to understand their meaning and patterns.

[0809] "Relevant information" refers to data and knowledge that are presented appropriately in response to the statements and emotional states of participants and customers.

[0810] A "database" is a collection of information organized in a systematic way that allows for efficient storage, searching, and use of information.

[0811] A "proposal" is to recommend appropriate policies or actions based on the situation.

[0812] "Quality improvement" means enhancing the value and effectiveness of the services and information provided.

[0813] This invention is implemented as a system to improve customer service in physical stores. The system consists of a program for effectively utilizing various data in meetings and interactions with customers in stores. The embodiments thereof are described below.

[0814] The server first collects customer voices from smart devices in the store and obtains audio information. The acquired audio information is transcribed in real time using the Google Cloud Speech-to-Text API. This transcription result is sent to Amazon Comprehend for real-time sentiment analysis, where the customer's emotional state is analyzed. Based on the analyzed emotional state, the server searches the database for relevant information and selects the appropriate information.

[0815] The smart glasses, acting as the terminal, visually display the customer's emotional state and related information transmitted from the server in real time. This allows the store staff, acting as users, to provide service tailored to the customer's emotional state. For example, if the glasses detect that the customer is confused, information carefully explaining how to set up and use the product will be displayed on them.

[0816] As a concrete example, suppose emotion analysis detects that customer A is looking for a specific product but cannot find it and appears confused. In this case, the smart glasses would display a specific response plan, such as "Suggest to guide customer A to the location of the product on the shelf."

[0817] Furthermore, the server also includes a function to ensure that the information presented is based on the customer's emotional state. This is expected to improve the quality of customer service and increase customer satisfaction.

[0818] Example of a prompt

[0819] Convert speech to text in real time and analyze customer emotions. If a customer is having trouble, immediately suggest helpful information and display it on smart glasses.

[0820] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0821] Step 1:

[0822] The server acquires customer voices as audio data from smart devices acting as terminals. The input is real-time audio information collected from smart devices within the store. The output is audio data used in the speech recognition process.

[0823] Step 2:

[0824] The server transcribes the acquired audio data in real time using the Google Cloud Speech-to-Text API. The input is the audio data obtained in step 1, and the output is the transcribed text. This process prepares the spoken content for analysis as digital data.

[0825] Step 3:

[0826] The server passes text data to Amazon Comprehend, which analyzes the spoken content to identify the customer's emotional state. The input is the transcribed text from step 2, and the output is information about the emotional state. This procedure allows the server to understand the customer's psychological state in real time.

[0827] Step 4:

[0828] The server searches the database for appropriate relevant information based on the analyzed emotional state. The input is the emotional state information obtained in step 3, and the output is relevant information that matches the emotional state. A prompt sentence is generated through a generative AI model, and the optimal information is selected.

[0829] Step 5:

[0830] The smart glasses, acting as the terminal, visually display relevant information sent from the server to the customer. The input is the relevant information selected in step 4, and the output is information presented visually in a user-friendly format. This enables the user to respond quickly and appropriately to the customer.

[0831] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0832] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0833] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0834] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0835] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. In the upper and lower directions of the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. Also, the upper side of the concentric circles is where "pleasant" emotions are located, and the lower side is where "unpleasant" emotions are located. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0836] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0837] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0838] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0839] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0840] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0841] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0842] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0843] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0844] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0845] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0846] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0847] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0848] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0849] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0850] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0851] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0852] The following is further disclosed regarding the embodiments described above.

[0853] (Claim 1)

[0854] A method for acquiring audio data and performing real-time transcription,

[0855] A means of analyzing necessary information from the user's statements and searching for related information in a database,

[0856] A means to present the acquired information to the user and to automatically generate a summary of the meeting content,

[0857] A means of improving the accuracy of the system based on user feedback,

[0858] A system that includes this.

[0859] (Claim 2)

[0860] The system according to claim 1, comprising means for collecting the aforementioned transcription results and automatically documenting important agenda items and decisions based on templates.

[0861] (Claim 3)

[0862] The system according to claim 1, further comprising output means for presenting the acquired information in a manner that can be viewed by a user.

[0863] "Example 1"

[0864] (Claim 1)

[0865] A means for acquiring audio information and converting it into a digital signal using a highly sensitive receiver,

[0866] A means for analyzing received audio information and converting it into text information using a speech recognition function,

[0867] A means of analyzing transcribed text information, identifying necessary information from the content of speeches, and searching for information sources including past meeting histories,

[0868] A means to process search results and present them in a format that is easy for users to understand, as well as to automatically generate a summary of the meeting content.

[0869] A means of collecting user feedback and using it as data to improve system performance,

[0870] A system that includes this.

[0871] (Claim 2)

[0872] The system according to claim 1, comprising means for saving the transcription results and automatically compiling important agenda items and decisions based on an existing format.

[0873] (Claim 3)

[0874] The system according to claim 1, further comprising a function for presenting the processed information using an output device.

[0875] "Application Example 1"

[0876] (Claim 1)

[0877] A means of acquiring acoustic information and performing real-time text conversion,

[0878] A means of analyzing the necessary knowledge from the user's statements and searching for related knowledge from storage devices,

[0879] A means to present acquired knowledge to the user and to automatically generate a summary of the dialogue content,

[0880] A means of improving the accuracy of the system based on user feedback,

[0881] A computer device that analyzes user statements and instantly displays product information,

[0882] A means of providing information so that users can confirm the information using visual devices,

[0883] A system that includes this.

[0884] (Claim 2)

[0885] The system according to claim 1, comprising means for collecting the character conversion results and automatically documenting important topics and decisions based on a template.

[0886] (Claim 3)

[0887] The system according to claim 1, further comprising output means for displaying the acquired knowledge in a manner accessible to users.

[0888] "Example 2 of combining an emotion engine"

[0889] (Claim 1)

[0890] A means of acquiring audio signals and performing data conversion in real time,

[0891] A means of analyzing emotional states from the content and characteristics of statements and retrieving related information from an external memory device,

[0892] A means to optimize information presentation based on analyzed sentiment information and provide it to users, as well as to automatically generate a summary of the meeting content,

[0893] A means of obtaining user responses to improve the accuracy of system analysis,

[0894] ...

[0895] A system that includes this.

[0896] (Claim 2)

[0897] The system according to claim 1, comprising means for collecting the conversion results and automatically documenting important topics and decisions based on templates.

[0898] (Claim 3)

[0899] The system according to claim 1, further comprising display means for providing the acquired information in a manner that is visible to the user.

[0900] "Application example 2 of combining emotional engines"

[0901] (Claim 1)

[0902] A method for acquiring audio information and transcribing it in real time,

[0903] A method for analyzing participants' emotional states from their statements,

[0904] A means of retrieving and presenting relevant information from a database according to the emotional state,

[0905] A means to ensure that the information presented is based on the user's emotional state,

[0906] A means to automatically generate a summary of meeting content and reflect the emotional trends of users,

[0907] A means of improving the accuracy of the system based on user feedback,

[0908] A system that includes this.

[0909] (Claim 2)

[0910] The system according to claim 1, comprising means for collecting the transcription results, automatically documenting important topics and decisions based on templates, and visually highlighting and displaying them according to the user's emotional state.

[0911] (Claim 3)

[0912] The system according to claim 1, further comprising output means for presenting the acquired information in a manner accessible to the user, and display means for proposing customer service methods based on emotion analysis results. [Explanation of Symbols]

[0913] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A method for acquiring audio data and performing real-time transcription, A means of analyzing necessary information from the user's statements and searching for related information in a database, A means to present the acquired information to the user and to automatically generate a summary of the meeting content, A means of improving the accuracy of the system based on user feedback, A system that includes this.

2. The system according to claim 1, comprising means for collecting the aforementioned transcription results and automatically documenting important agenda items and decisions based on a template.

3. The system according to claim 1, further comprising output means for presenting the acquired information in a manner that can be viewed by a user.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A