system

The system addresses long meetings by converting audio to text, performing natural language processing, and generating responses, enhancing meeting efficiency and productivity through real-time participation and summary automation.

JP2026069154APending Publication Date: 2026-04-23SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-11
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Long meetings reduce productivity, and existing systems fail to efficiently automate information sharing, decision-making, and meeting summarization, leading to delays in document preparation and inefficient use of working hours.

Method used

A system that converts audio streams to text, performs natural language processing, generates responses, and summarizes conversations in real-time, using user feedback to improve accuracy and automate meeting participation.

Benefits of technology

Enhances meeting efficiency by allowing real-time participation and summary generation, reducing the burden on participants and improving productivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069154000001_ABST
    Figure 2026069154000001_ABST
Patent Text Reader

Abstract

Provide a system. 【Solution means】 Means for storing conversation information in a database, Means for converting an audio stream into character information, Means for performing natural language processing using the character information, Means for retrieving relevant information from the database and generating a response, Means for converting the generated response into speech and outputting it, Means for performing video generation as needed, Means for summarizing the content of the conversation, Means for sending the summary information to the user, A system including the above.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In a modern business environment, long meetings often compress working hours and lead to a decline in productivity. Also, if information sharing and decision-making in meetings are not carried out quickly, the efficiency of corporate activities may also decline. Furthermore, in order to grasp the content of a meeting that one could not attend, it is necessary to follow up on the key points, which takes a lot of time. As a result, there may be a delay in document preparation and the preparation for the next meeting. Therefore, there is a need for a technology that automates and improves the efficiency of meetings and provides an easy-to-participate environment.

Means for Solving the Problems

[0005] This invention provides a system that ensures real-time performance by storing conversation information in a database and converting audio streams into text information. This system uses the text information for natural language processing, retrieving relevant information from the database and automatically generating responses. Furthermore, the generated responses are converted into audio, and video is generated as needed, allowing the system to participate in meetings on behalf of the user. It also summarizes the conversation content and sends the summary to the user, enabling efficient understanding of meetings missed. In addition, it utilizes user feedback to improve the system's accuracy and provides a function to automatically generate materials, supporting preparation for future meetings. This improves the time efficiency of meetings and reduces the burden on participants.

[0006] "Conversational information" refers to audio and related text data, and includes all information generated during the process of meetings and communication.

[0007] A "database" is a system or platform for storing information in a structured format, making it accessible and manipulated quickly and effectively.

[0008] An "audio stream" is a continuous flow of data that transmits real-time or recorded audio signals in digital format.

[0009] "Textual information" refers to information obtained by converting audio data into text, and is digital data expressed in the form of strings or documents.

[0010] "Natural language processing" is a technology that uses computers to understand, interpret, and generate human language, and it is a process that mechanically performs the analysis of text data and extracts information.

[0011] "Relevant information" refers to data and knowledge that are appropriate to the current context or topic and necessary for the requested answers or actions.

[0012] "Generating a response" refers to the process of creating a message as an appropriate reply or reaction based on the information received.

[0013] "Converting to speech" is the process of converting text data into a human voice using speech synthesis technology and conveying it audibly.

[0014] "Image generation" is the process of creating virtual visual content using computer graphics and AI technology and communicating it visually.

[0015] "Summary" refers to extracting the main points from detailed information or text and expressing them in a concise and clear format.

[0016] "Feedback" refers to opinions and evaluations provided about the performance and functionality of a system, and is information used for improvement.

[0017] "Creating documents" is the process of organizing and structuring information to generate documents, reports, presentations, and other materials tailored to a specific purpose. [Brief explanation of the drawing]

[0018] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0019] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0020] First, the language used in the following description will be explained.

[0021] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0022] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0023] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0024] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0026] [First Embodiment]

[0027] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0028] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0031] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0034] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0038] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0039] This invention is implemented as a system that processes information in real time by storing conversation information in a database and converting audio streams into text information, in order to improve the efficiency and automation of meeting participation.

[0040] Server Functions

[0041] The server stores user-registered materials and past conversation history in a database. This data is used to provide necessary information during the meeting. When the meeting begins, the server uses speech recognition technology to convert the conversation into text in real time and puts it into a natural language processing process. The server analyzes the generated text, searches the database for relevant information, and generates an appropriate response. The generated response is converted into speech using speech synthesis technology and sent to the terminal in real time. After the meeting ends, the server automatically summarizes the conversation and sends it to the user so that they can quickly grasp the main points.

[0042] Device functions

[0043] The terminal plays synthesized speech transmitted from the server through its speaker, participating in the meeting on behalf of the user. When necessary, it also generates virtual images using video generation technology and distributes them to other participants. This allows for visual and auditory interaction even without the user's physical presence.

[0044] User actions

[0045] Users register their meeting attendance requests in the system while simultaneously managing their own schedules. After the meeting, they send necessary feedback to the server based on the summary they receive, contributing to the system's accuracy improvement. They can also request the creation of materials for future meetings. For example, even if a user is unable to attend a sales meeting, they can review the summary of the negotiation report and decisions, and send a request to the server to create additional materials if necessary.

[0046] Therefore, the present invention adopts a form that significantly improves the efficiency of information transmission by automating meetings through the cooperation of a server, terminal, and user. This system allows users to make effective use of their time and dramatically improve the productivity of their business activities.

[0047] The following describes the processing flow.

[0048] Step 1:

[0049] When the server detects the start of a meeting, it activates a speech recognition AI to collect the audio stream transmitted from the terminal and convert it into text in real time. The converted text is stored as a temporary file for smooth natural language processing.

[0050] Step 2:

[0051] When the server acquires textual information, it inputs it into a large-scale language model to analyze its content and understand the context. As the conversation progresses, the server searches for relevant information in the database and automatically generates the next necessary response. This response is intended to allow the server to participate in the meeting on behalf of the user.

[0052] Step 3:

[0053] The server converts the generated response into audio data using speech synthesis technology. It then sends the converted audio data to the terminal, conveying the specific response content to the meeting participants via the speaker.

[0054] Step 4:

[0055] The device activates a video generation AI as needed to virtually create the user's avatar and video. The generated video is updated in real time and distributed as visual information to other meeting participants. This video generation provides users in remote locations with a sense of presence as if they were actually there.

[0056] Step 5:

[0057] After the meeting ends, the server analyzes the recorded conversation and participants' comments to extract the main points and conclusions of the discussion. Using this information, the server creates a concise and easy-to-understand summary and sends it to the user via email or application.

[0058] Step 6:

[0059] Users review the summary provided by the server and provide feedback as needed. This feedback helps improve the system. Furthermore, they can request the creation of materials for the next meeting. This request is sent to the server and processed.

[0060] This processing flow allows users to efficiently participate in meetings and grasp their key points. The system provides information and automation to support business activities, contributing to increased user productivity.

[0061] (Example 1)

[0062] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0063] In today's business environment, efficient participation in meetings is crucial, but it can be difficult to attend due to busy work schedules. Furthermore, there is a need to effectively manage and utilize the information provided during meetings. Existing systems lack the functionality to completely replace meeting attendance or the ability to quickly summarize and create materials after meetings, thus limiting improvements in work efficiency.

[0064] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0065] In this invention, the server includes means for storing conversation data in an information storage device, a device for converting audio signals into text data, and a device for performing language processing using the text data. This makes it possible to convert information during a meeting into text data in real time and to immediately search for and provide relevant information. Furthermore, by automatically generating materials using the generated summary data, users can efficiently prepare for the next meeting.

[0066] "Conversation data" refers to information that shows the content of words recorded during voice communication.

[0067] An "information storage device" is hardware or software used to systematically store and manage large amounts of data.

[0068] An "audio signal" is waveform data in electrical or digital form generated by sound.

[0069] "Text data" refers to information that represents sounds, symbols, and other elements in both text and digital formats.

[0070] "Language processing" is the technology used to analyze, understand, and generate natural language using computers.

[0071] "Relevant data" refers to information that is deemed useful or necessary in relation to a particular data point or question.

[0072] A "response" is the processing result or answer returned in response to an input.

[0073] "Audio data" refers to information that represents sound as electrical signals or in digital format.

[0074] "Image data generation" is the process of creating visual representations using computer graphics.

[0075] "Summary data" refers to information that has been shortened by extracting the most important content from the original information.

[0076] "User" refers to a person or organization that uses a system or service.

[0077] "Event scheduling" is the process of planning and tracking the schedules of events and meetings for individuals or organizations.

[0078] This invention constitutes a system that streamlines and automates meeting participation. The invention mainly consists of three elements: a server, a terminal, and a user.

[0079] The server has a database as an information storage device, where it stores conversation data and materials registered in advance by users. The server processes audio signals and converts them into text data using speech recognition technology. Here, a common example of speech recognition software is a speech processing API. The text data is analyzed using a natural language processing library to quickly retrieve relevant data and generate a response. This response is converted into audio data using speech synthesis technology and sent to the terminal.

[0080] The terminal plays audio data transmitted by the server through its speaker and participates in the meeting on behalf of the user. It also generates video using image data generation technology when requested and provides it to other participants. The terminal transmits information collected through voice input to the server in real time.

[0081] Users register their attendance at meetings using an event scheduling application. After the meeting, users can review the key points based on the provided summary data and contribute to improving the system's accuracy by sending feedback to the server. Users can also request document generation from a system utilizing a generative AI model by entering prompt messages. A specific example of such a prompt message is, "Please prepare a proposal for the next technical meeting."

[0082] In this way, servers, terminals, and users cooperate, enabling users to efficiently obtain information and participate in meetings on their behalf.

[0083] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0084] Step 1:

[0085] The server stores meeting materials and past conversation data registered by users in advance in an information storage device. It receives documents and audio files provided by users as input and systematically stores them in a database. This process makes it easy to retrieve information needed in subsequent processes.

[0086] Step 2:

[0087] The device acquires audio signals in real time during meetings via a microphone. It receives physical audio as input and transmits it to the server as a digital audio signal. Noise cancellation technology is implemented to ensure high-quality audio signals.

[0088] Step 3:

[0089] The server uses speech recognition technology to convert the received audio signal into text data. Specifically, it uses a speech recognition API to perform the conversion from speech to text. This allows the conversation content to be stored as text, enabling subsequent natural language processing.

[0090] Step 4:

[0091] The server analyzes the obtained character data using a natural language processing library. Using the character data as input, it extracts linguistic context and important keywords. This lays the foundation for the server to quickly retrieve relevant information from the database.

[0092] Step 5:

[0093] The server searches the information storage device based on the analysis results to retrieve relevant data and generates a response that meets the user's needs. A generative AI model is used here, and the generated response is meaningful to the user. This output is then converted into speech in the next step.

[0094] Step 6:

[0095] The server converts the generated response into audio data using speech synthesis technology. A speech synthesis API is used for this process, and the generated audio is sent to the terminal. The terminal plays the received audio data through its speaker to communicate with the meeting participants.

[0096] Step 7:

[0097] After the meeting ends, the server automatically summarizes the conversation. Using a generative AI model, it extracts key points from the overall conversation as input. This summary is sent to the user via email or other means, allowing them to quickly review the main points of the meeting.

[0098] (Application Example 1)

[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0100] In today's commercial environment, it can be difficult for customers to efficiently obtain product information and receive smooth support when making purchasing decisions within virtual stores. Furthermore, the inability to quickly understand individual customer needs and provide product suggestions and follow-up based on those needs makes improving customer satisfaction and the purchasing experience a challenge.

[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0102] In this invention, the server includes means for storing conversation information in a storage device, means for converting audio signals into text information, and means for performing natural language processing using the text information. This makes it possible to analyze customer interactions in real time and present product information based on individual needs.

[0103] A "storage device" is a medium for saving conversational information and related information.

[0104] An "audio signal" is sound data acquired by an input device such as a microphone.

[0105] "Textual information" refers to text data obtained by converting audio signals.

[0106] "Natural language processing" is a technology that analyzes the meaning of text data and processes the information.

[0107] "Response" refers to the content of the reply to the customer generated by natural language processing.

[0108] An "image" is visual data that is generated as needed and provided as visual information.

[0109] "Summary information" refers to data that briefly summarizes the content of a conversation or piece of information.

[0110] "Users" refer to customers or users who utilize virtual stores or systems.

[0111] "Purchasing behavior" refers to a series of actions taken by a user to select a product and decide to purchase it.

[0112] "Opinions" refer to feedback information provided by users.

[0113] "Materials" refer to documents or data provided to support users in their preparation and decision-making.

[0114] The server receives an audio signal and converts it into text using the Google® Cloud Speech-to-Text API. Furthermore, it analyzes this text using an OpenAI® language model and performs natural language processing. Based on the analyzed content, it searches a database in the storage device for relevant information and generates an appropriate response. This response is then converted back into speech using Google Cloud Text-to-Speech and sent to the terminal.

[0115] The terminal plays back the received response audio and provides it to the user. It also generates and displays images as needed to supplement visual information. Through this, users can smoothly carry out purchasing actions within the virtual store.

[0116] Users interact with an AI assistant in a virtual store and receive product information. For example, if a user is looking for sneakers, they can ask, "Which sneakers do you recommend?" and the AI ​​assistant will suggest products based on past purchase history and trend data. This system allows users to obtain information efficiently and improves the shopping experience.

[0117] An example of a prompt might be, "Based on this customer's past purchase history and trend data, please provide recommendations for appropriate sneakers." This prompt prompts the generative AI model to process information to provide the customer with the most suitable product information.

[0118] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0119] Step 1:

[0120] The server receives audio signals from the terminal. The input is audio data spoken by the user through the microphone. The audio signals are converted into a digital format and sent to the server.

[0121] Step 2:

[0122] The server converts the received audio signal into text information using the Google Cloud Speech-to-Text API. The input is a digital audio signal, and the output is text data in string format. This conversion is performed using speech recognition technology.

[0123] Step 3:

[0124] The server analyzes the generated text information using natural language processing techniques. The input is text data in string format, and the output is the analyzed meaning and related information. OpenAI's language model is used to understand the text content and extract contextually relevant information.

[0125] Step 4:

[0126] The server searches for relevant information from the database in the storage device based on the analyzed information. The input is keywords and contextual information obtained through natural language processing, and the output is the optimal response information. The database search takes into account the user's past history and trend information.

[0127] Step 5:

[0128] The server converts the generated response into speech using Google Cloud Text-to-Speech. The input is text data of the response information, and the output is synthesized speech data. This speech data is provided to the user in a natural and easy-to-understand format.

[0129] Step 6:

[0130] The terminal plays synthesized speech data received from the server. The input is speech data, and the output is sound played from the speaker. Through this, users can obtain both visual and auditory information.

[0131] Step 7:

[0132] The device generates relevant images as needed and displays them to the user. Input is an instruction to generate an image based on response information, and output is the visual information displayed on the screen. This facilitates the user's purchasing behavior.

[0133] Step 8:

[0134] Users make purchasing decisions based on the provided information and images, and send feedback to the server. Input consists of user ratings and comments, and output is stored as feedback information. This allows the system to improve its accuracy for future interactions.

[0135] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0136] This invention aims to automate and improve the efficiency of meeting systems, and is implemented in a form that includes a function to recognize user emotions and optimize responses.

[0137] Server Functions

[0138] The server stores user-uploaded materials and past conversation history in a database. It receives meeting audio streams and uses speech recognition AI to convert them into text in real time. Furthermore, the server uses an emotion engine in its natural language processing to recognize user emotions and generate contextually appropriate responses. This enables not only the exchange of information but also communication that is sensitive to the user's emotional state.

[0139] The generated responses are converted into synthesized speech data and transmitted to the terminal in real time. Furthermore, after the meeting ends, the server analyzes the statements and materials, creates a summary including changes in emotion, and provides it to the user. In particular, the user's emotion data is stored in a database as relevant information and used for future meetings.

[0140] Device functions

[0141] The device plays synthesized speech transmitted from the server through its speaker. It also uses video generation technology to create emotionally appropriate facial expressions as needed, broadcasting a virtual video of the user to other participants. This video reflects the user's current emotional state, creating a more natural participation experience.

[0142] User actions

[0143] Users can use data provided by the emotion engine to understand the impact of their own emotions on the progress of a meeting. For example, if an emotional instability occurs during a meeting, the emotion engine can detect this, and the server can generate a calming response to help the meeting proceed smoothly. Users can also review the summary provided after the meeting and, if necessary, instruct the system to prepare for the next meeting, requesting the creation of relevant materials.

[0144] Thus, the present invention provides a system in which servers, terminals, and users cooperate with each other to realize emotionally resonant information transmission and efficient meetings. The overall system aims to improve the productivity and satisfaction of meeting participants.

[0145] The following describes the processing flow.

[0146] Step 1:

[0147] As soon as the meeting starts, the server activates the speech recognition AI, receiving the audio stream transmitted from the terminal in real time and converting it into text. This text data serves as the foundational data for storing the conversation content in a database.

[0148] Step 2:

[0149] The server inputs textual information into a natural language processing system and activates an emotion engine to analyze the user's emotional state. Based on the emotion data, the server understands the context of the conversation and designs an optimal response that takes the user's emotions into consideration. This response takes into account the user's emotions and the meeting situation.

[0150] Step 3:

[0151] The server sends the generated response to a speech synthesis engine, where it is converted into speech data. This converted speech data is then sent to the terminal to be spoken on behalf of the user during the meeting. This speech is responsible for interacting with other meeting participants.

[0152] Step 4:

[0153] The device utilizes video generation technology as needed to create avatar expressions and postures that match the user's emotions. The generated video is updated in real time and displayed to other participants, contributing to the meeting as visual information.

[0154] Step 5:

[0155] After the meeting ends, the server analyzes the entire conversation and creates a summary that reflects the key points of the discussion and the users' emotional changes. This summary is organized with emotional data and provided to the users as useful data for future meetings and follow-ups.

[0156] Step 6:

[0157] Users review summaries and sentiment data sent from the server and provide feedback as needed to help improve the system. They can also request the server to create materials for the next meeting, facilitating efficient meeting participation. This request is processed automatically, and the materials are created in the specified format.

[0158] This processing flow allows the system to provide users with an emotionally responsive and sophisticated meeting experience while simultaneously supporting improved productivity and satisfaction.

[0159] (Example 2)

[0160] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0161] Traditional meeting systems merely transmit information without considering user emotions, limiting their ability to improve participant satisfaction and productivity. Furthermore, they lacked efficient support for post-meeting summarization and preparation for future meetings.

[0162] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0163] In this invention, the server includes means for storing conversation information in a storage device, means for converting audio signals into text information, and means for performing natural language analysis using the text information. This enables richer communication by analyzing the user's emotions and reflecting them in the response.

[0164] "Conversational information" refers to data on statements and content exchanged during meetings and dialogues, which is stored in a data storage device.

[0165] "Audio signal" refers to audio data in analog or digital format collected by a microphone or similar device.

[0166] "Textual information" refers to data in text format that is converted from an audio signal using speech recognition technology.

[0167] "Natural language processing" is the process of assigning meaning to and structuring textual information using algorithms such as generative AI models.

[0168] An "emotion engine" is an algorithm or software that analyzes and extracts emotions from a user's statements and actions.

[0169] A "response" is a system message generated based on the user's statements and emotions.

[0170] "Means of converting to speech and outputting" refers to the process of converting text-based response information into speech data using speech synthesis technology and providing it through speakers or other means.

[0171] "Video generation" is the process of generating videos and animations that visually represent a user's emotions and situation based on audio and text information.

[0172] "Summary information" refers to information that has been shortened and organized to include key points and changes in emotion from a meeting or conversation.

[0173] An "information terminal" is a device used by users to receive and view conversational information and summary information.

[0174] The system of this invention aims to improve meeting efficiency and enable emotionally resonant communication by having the server, terminals, and users work together in a coordinated manner.

[0175] The server supports meeting preparation by storing meeting materials and past conversation information uploaded by users in advance in a storage device. The server receives audio signals in real time and converts them into text information using a speech recognition module. Specifically, a speech recognition API can be used. The recognized text information is interpreted by a natural language processing engine, and an emotion engine is used to recognize the user's emotions. The server generates an appropriate response from the analyzed information and outputs it as speech using a speech synthesis module. This outputted speech is sent to the terminal and communicated to the user.

[0176] The device plays synthesized speech transmitted from the server through its speaker. Additionally, it uses video generation technology as needed to deliver virtual images reflecting the user's emotions to other participants. This feature helps to visually communicate the user's current emotional state to other participants, facilitating more natural communication.

[0177] Users can conduct meetings while receiving responses from the server in real time. They can also utilize emotional data analyzed by the emotion engine to understand their own emotional state and use that information to guide the meeting. After the meeting, they can review a summary provided by the server and use it to prepare for the next meeting.

[0178] A specific example of a prompt would be, "Analyze emotional changes during the meeting and generate feedback to help users relax." This allows users to reduce stress during meetings and communicate more effectively.

[0179] As described above, this system improves participant productivity and satisfaction by managing meeting information, analyzing emotions and generating responses, and distributing video based on the results.

[0180] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0181] Step 1:

[0182] The server stores meeting materials and past conversation information uploaded in advance by users in a storage device. It receives various meeting-related data provided by users as input and saves it in a structured database, making it quickly accessible when needed. As output, it generates a state that facilitates searching based on specific topics and past discussions.

[0183] Step 2:

[0184] The server receives audio signals in real time during the meeting. It receives digital audio data obtained through the microphone as input and converts it into text information using a speech recognition engine. Specifically, it utilizes a speech recognition API to output the spoken words in text format.

[0185] Step 3:

[0186] The server performs natural language processing using textual information. It receives textual information generated by speech recognition as input, and this information is analyzed by a natural language processing engine. The output provides contextual and semantic information for each utterance in a structured data format.

[0187] Step 4:

[0188] The server uses the results of natural language processing to run an emotion engine and recognize the user's emotions. It takes parsed text information as input and extracts the emotional state based on its context. The output explicitly indicates the user's emotional state by assigning tags and scores corresponding to those emotions.

[0189] Step 5:

[0190] The server generates appropriate responses based on emotion recognition. It uses the output data of an emotion engine as input and a generative AI model to generate contextually appropriate responses. The output is a text-based response that aligns with the user's emotions.

[0191] Step 6:

[0192] The server converts the generated response into synthesized speech and sends it to the terminal. It receives text responses from a generation AI model as input and converts them into speech data using a speech synthesis engine. As output, it creates a playable audio file and transfers it to the terminal in real time.

[0193] Step 7:

[0194] The terminal plays synthesized speech sent from the server through its speaker. It receives synthesized speech data as input and plays the speech via an audio output module. The output is the actual speech that the user and other participants can hear.

[0195] Step 8:

[0196] The device will use video generation technology as needed to deliver videos that express the user's emotions. It will receive emotion-tagged data as input and utilize video generation software. As output, it will generate animations or video content reflecting those emotions, which will then be distributed to other meeting participants.

[0197] Step 9:

[0198] After the meeting ends, the server analyzes each statement and document to create a summary that includes changes in emotions. As input, it integrates all textual information and emotional state data collected during the meeting, and generates a summary using an analysis algorithm. As output, a summary formatted for easy user understanding is sent to the terminal.

[0199] (Application Example 2)

[0200] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0201] In autonomous vehicles, accurately understanding passengers' emotional states is difficult, making it challenging to provide appropriate responses and support. This can result in a compromised passenger experience and prevent the vehicle from fully achieving its safety and comfort levels. Therefore, a system is needed that analyzes passengers' emotional states and provides optimal information and responses based on that analysis.

[0202] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0203] In this invention, the server includes a function for storing conversation information, a function for converting audio signals into text, a function for performing natural language understanding, a function for retrieving relevant information from the information storage unit and generating a response, a function for converting the generated response into audio data and outputting it, a function for visual representation as needed, a function for summarizing the content of the conversation, a function for transmitting the summarized information to the user, a function for analyzing the emotional state of the passengers, and a function for optimizing information presentation based on the emotional state. This makes it possible to grasp the emotional state of the passengers in real time and provide optimal information and responses accordingly.

[0204] "Conversational information" refers to audio and text-based information exchanged between users, which is stored in a database and used for analysis.

[0205] An "audio signal" is an electrical representation of sound acquired by a microphone or other audio equipment, and is used as input data for conversion into text.

[0206] "Natural language understanding" is a technology that allows computers to analyze human language and understand its meaning, enabling the generation of appropriate responses and information retrieval.

[0207] The "information storage unit" is a memory area that stores past conversation data and related information, providing the data necessary for response generation and information retrieval.

[0208] "Visual expression" refers to videos and images generated to visually convey passengers' emotional states and information, and is used to improve the user experience.

[0209] "Emotional state" refers to the psychological state of passengers and represents information analyzed from audio and video data.

[0210] "Information presentation" refers to the information and responses provided to the user, and by presenting them at the appropriate time and in the appropriate format, it improves the user experience.

[0211] In this invention, the server plays a central role in the information processing system installed in the autonomous vehicle. The server receives passenger voice signals and converts them into text information in real time using speech recognition AI. Speech recognition technologies such as Google Speech-to-Text can be used for this process. The converted text is analyzed using natural language processing tools, and the passenger's emotional state is understood by an emotion recognition engine such as IBM Watson® Tone Analyzer.

[0212] Based on the analyzed data, the server generates optimal information and responses tailored to the passenger's emotions. To generate responses in natural language, a generative AI model can be used to select contextually appropriate phrases. These generated responses are converted into speech using text-to-speech technology such as Amazon Polly and output through the vehicle's speakers.

[0213] The terminal can display visuals based on the passenger's emotional state. Specifically, if a passenger is feeling stressed, it can display relaxing scenes to create a sense of security. Furthermore, the terminal can store conversational information and emotional data collected during the journey, which can then be referenced in future information provision.

[0214] For example, if the emotion engine detects that a passenger appears tired, the server can suggest, "You seem a little tired. Shall I play some soothing music?" In this way, appropriate communication can be provided according to the passenger's situation.

[0215] Examples of prompt statements are as follows:

[0216] "Perform a sentiment analysis on this text: {text}. Suggest appropriate responses based on the passenger's emotions."

[0217] This allows the system to provide information efficiently and in a user-centric manner, in response to changes in passengers' emotions.

[0218] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0219] Step 1:

[0220] The server receives audio signals from microphones placed inside the vehicle. The audio data acquired by the microphones is used as input. The data processing performed here involves converting analog audio signals into digital data.

[0221] Step 2:

[0222] The server converts the received audio data into text using speech recognition AI (e.g., Google Speech-to-Text). The input is the digital audio data acquired in step 1, and the output is text data. The speech recognition algorithm analyzes words and generates the corresponding text.

[0223] Step 3:

[0224] The server processes the generated text data using a natural language processing tool (e.g., NLTK) to analyze its content. Here, the converted text is used as input for syntactic and semantic analysis. This results in detailed, context-based text information as output.

[0225] Step 4:

[0226] The server passes the analyzed text to an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to detect the passenger's emotional state. The input is the text information from step 3, and the output is data about the recognized emotional state. The data calculation performed here involves assigning emotional weights to each word and phrase to calculate an overall emotion score.

[0227] Step 5:

[0228] The server generates the optimal response based on emotion data using an AI model. The input is emotion state data, and the output is the text of the generated response. In this process, emotion data is input to the model as a prompt sentence, and the generated response is obtained.

[0229] Step 6:

[0230] The server converts the generated response text into speech using speech synthesis technology (e.g., Amazon Polly) and sends it to the terminal. The input is the response text, and the output is audio data. A speech synthesis algorithm analyzes the text and generates an audio waveform.

[0231] Step 7:

[0232] The terminal outputs the received audio data through its speaker. It also creates and displays visual representations on the screen based on the passenger's emotional state. The input consists of the audio data and emotional data transmitted in step 6, and the output consists of audio from the speaker and visuals displayed on the screen.

[0233] Step 8:

[0234] Users monitor their emotional state based on the information and voice responses provided. This allows them to enjoy a comfortable travel experience in the environment the system offers.

[0235] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0236] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0237] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0238] [Second Embodiment]

[0239] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0240] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0241] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0242] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0243] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0244] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0245] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0246] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0247] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0248] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0249] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0250] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0251] This invention is implemented as a system that processes information in real time by storing conversation information in a database and converting audio streams into text information, in order to improve the efficiency and automation of meeting participation.

[0252] Server Functions

[0253] The server stores user-registered materials and past conversation history in a database. This data is used to provide necessary information during the meeting. When the meeting begins, the server uses speech recognition technology to convert the conversation into text in real time and puts it into a natural language processing process. The server analyzes the generated text, searches the database for relevant information, and generates an appropriate response. The generated response is converted into speech using speech synthesis technology and sent to the terminal in real time. After the meeting ends, the server automatically summarizes the conversation and sends it to the user so that they can quickly grasp the main points.

[0254] Device functions

[0255] The terminal plays synthesized speech transmitted from the server through its speaker, participating in the meeting on behalf of the user. When necessary, it also generates virtual images using video generation technology and distributes them to other participants. This allows for visual and auditory interaction even without the user's physical presence.

[0256] User actions

[0257] Users register their meeting attendance requests in the system while simultaneously managing their own schedules. After the meeting, they send necessary feedback to the server based on the summary they receive, contributing to the system's accuracy improvement. They can also request the creation of materials for future meetings. For example, even if a user is unable to attend a sales meeting, they can review the summary of the negotiation report and decisions, and send a request to the server to create additional materials if necessary.

[0258] Therefore, the present invention adopts a form that significantly improves the efficiency of information transmission by automating meetings through the cooperation of a server, terminal, and user. This system allows users to make effective use of their time and dramatically improve the productivity of their business activities.

[0259] The following describes the processing flow.

[0260] Step 1:

[0261] When the server detects the start of a meeting, it activates a speech recognition AI to collect the audio stream transmitted from the terminal and convert it into text in real time. The converted text is stored as a temporary file for smooth natural language processing.

[0262] Step 2:

[0263] When the server acquires textual information, it inputs it into a large-scale language model to analyze its content and understand the context. As the conversation progresses, the server searches for relevant information in the database and automatically generates the next necessary response. This response is intended to allow the server to participate in the meeting on behalf of the user.

[0264] Step 3:

[0265] The server converts the generated response into audio data using speech synthesis technology. It then sends the converted audio data to the terminal, conveying the specific response content to the meeting participants via the speaker.

[0266] Step 4:

[0267] The device activates a video generation AI as needed to virtually create the user's avatar and video. The generated video is updated in real time and distributed as visual information to other meeting participants. This video generation provides users in remote locations with a sense of presence as if they were actually there.

[0268] Step 5:

[0269] After the meeting ends, the server analyzes the recorded conversation and participants' comments to extract the main points and conclusions of the discussion. Using this information, the server creates a concise and easy-to-understand summary and sends it to the user via email or application.

[0270] Step 6:

[0271] Users review the summary provided by the server and provide feedback as needed. This feedback helps improve the system. Furthermore, they can request the creation of materials for the next meeting. This request is sent to the server and processed.

[0272] This processing flow allows users to efficiently participate in meetings and grasp their key points. The system provides information and automation to support business activities, contributing to increased user productivity.

[0273] (Example 1)

[0274] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0275] In today's business environment, efficient participation in meetings is crucial, but it can be difficult to attend due to busy work schedules. Furthermore, there is a need to effectively manage and utilize the information provided during meetings. Existing systems lack the functionality to completely replace meeting attendance or the ability to quickly summarize and create materials after meetings, thus limiting improvements in work efficiency.

[0276] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0277] In this invention, the server includes means for storing conversation data in an information storage device, a device for converting audio signals into text data, and a device for performing language processing using the text data. This makes it possible to convert information during a meeting into text data in real time and to immediately search for and provide relevant information. Furthermore, by automatically generating materials using the generated summary data, users can efficiently prepare for the next meeting.

[0278] "Conversation data" refers to information indicating the content of words recorded during voice communication.

[0279] "Information storage device" refers to hardware or software for systematically storing and managing large amounts of data.

[0280] "Voice signal" refers to waveform data in electrical or digital form generated by voice.

[0281] "Character data" refers to information representing voices, symbols, etc. in characters and digital form.

[0282] "Language processing" refers to technology for analyzing, understanding, and generating natural language by computer.

[0283] "Related data" refers to information judged to be useful or necessary for certain data or questions.

[0284] "Response" refers to the processing result or answer returned for an input.

[0285] "Voice data" refers to information representing sound in electrical signal or digital form.

[0286] "Image data generation" refers to the process of creating visual representations using computer graphics.

[0287] "Summary data" refers to information obtained by extracting and shortening important content from the original information.

[0288] "User" refers to a person or organization using a system or service.

[0289] "Event schedule management" refers to the process of planning and tracking the schedules of personal or organizational events and meetings.

[0290] This invention constitutes a system that streamlines and automates meeting participation. The invention mainly consists of three elements: a server, a terminal, and a user.

[0291] The server has a database as an information storage device, where it stores conversation data and materials registered in advance by users. The server processes audio signals and converts them into text data using speech recognition technology. Here, a common example of speech recognition software is a speech processing API. The text data is analyzed using a natural language processing library to quickly retrieve relevant data and generate a response. This response is converted into audio data using speech synthesis technology and sent to the terminal.

[0292] The terminal plays audio data transmitted by the server through its speaker and participates in the meeting on behalf of the user. It also generates video using image data generation technology when requested and provides it to other participants. The terminal transmits information collected through voice input to the server in real time.

[0293] Users register their attendance at meetings using an event scheduling application. After the meeting, users can review the key points based on the provided summary data and contribute to improving the system's accuracy by sending feedback to the server. Users can also request document generation from a system utilizing a generative AI model by entering prompt messages. A specific example of such a prompt message is, "Please prepare a proposal for the next technical meeting."

[0294] In this way, servers, terminals, and users cooperate, enabling users to efficiently obtain information and participate in meetings on their behalf.

[0295] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0296] Step 1:

[0297] The server stores meeting materials and past conversation data registered by users in advance in an information storage device. It receives documents and audio files provided by users as input and systematically stores them in a database. This process makes it easy to retrieve information needed in subsequent processes.

[0298] Step 2:

[0299] The device acquires audio signals in real time during meetings via a microphone. It receives physical audio as input and transmits it to the server as a digital audio signal. Noise cancellation technology is implemented to ensure high-quality audio signals.

[0300] Step 3:

[0301] The server uses speech recognition technology to convert the received audio signal into text data. Specifically, it uses a speech recognition API to perform the conversion from speech to text. This allows the conversation content to be stored as text, enabling subsequent natural language processing.

[0302] Step 4:

[0303] The server analyzes the obtained character data using a natural language processing library. Using the character data as input, it extracts linguistic context and important keywords. This lays the foundation for the server to quickly retrieve relevant information from the database.

[0304] Step 5:

[0305] The server searches the information storage device based on the analysis results to retrieve relevant data and generates a response that meets the user's needs. A generative AI model is used here, and the generated response is meaningful to the user. This output is then converted into speech in the next step.

[0306] Step 6:

[0307] The server converts the generated response into voice data using text-to-speech technology. A text-to-speech API is utilized in this process, and the generated voice is transmitted to the terminal. The terminal plays the received voice data from the speaker to convey it to the participants in the meeting.

[0308] Step 7:

[0309] After the meeting ends, the server automatically summarizes the content of the conversation. By leveraging a generative AI model and using the overall content of the conversation as input, it extracts important points. This summary is sent to the user via email or the like, enabling the user to quickly review the key points of the meeting.

[0310] (Application Example 1)

[0311] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0312] In a modern business environment, it can be difficult for customers to efficiently obtain product information and receive smooth support when making purchasing decisions within a virtual store. Additionally, since it is not possible to quickly understand individual customer needs and make product proposals and follow-ups based on them, improving customer satisfaction and the purchasing experience has become an issue.

[0313] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0314] In this invention, the server includes means for storing conversation information in a storage device, means for converting voice signals into character information, and means for performing natural language processing using the character information. Thereby, it becomes possible to analyze the conversation with the customer in real time and present product information based on individual needs.

[0315] The "storage device" is a medium for storing conversation information and related information.

[0316] An "audio signal" is sound data acquired by an input device such as a microphone.

[0317] "Textual information" refers to text data obtained by converting audio signals.

[0318] "Natural language processing" is a technology that analyzes the meaning of text data and processes the information.

[0319] "Response" refers to the content of the reply to the customer generated by natural language processing.

[0320] An "image" is visual data that is generated as needed and provided as visual information.

[0321] "Summary information" refers to data that briefly summarizes the content of a conversation or piece of information.

[0322] "Users" refer to customers or users who utilize virtual stores or systems.

[0323] "Purchasing behavior" refers to a series of actions taken by a user to select a product and decide to purchase it.

[0324] "Opinions" refer to feedback information provided by users.

[0325] "Materials" refer to documents or data provided to support users in their preparation and decision-making.

[0326] The server receives an audio signal and converts it into text using the Google Cloud Speech-to-Text API. Furthermore, it analyzes this text using an OpenAI language model and performs natural language processing. Based on the analyzed content, it searches for relevant information in a database within the storage device and generates an appropriate response. This response is then converted back into speech using Google Cloud Text-to-Speech and sent to the terminal.

[0327] The terminal plays back the received response audio and provides it to the user. It also generates and displays images as needed to supplement visual information. Through this, users can smoothly carry out purchasing actions within the virtual store.

[0328] Users interact with an AI assistant in a virtual store and receive product information. For example, if a user is looking for sneakers, they can ask, "Which sneakers do you recommend?" and the AI ​​assistant will suggest products based on past purchase history and trend data. This system allows users to obtain information efficiently and improves the shopping experience.

[0329] An example of a prompt might be, "Based on this customer's past purchase history and trend data, please provide recommendations for appropriate sneakers." This prompt prompts the generative AI model to process information to provide the customer with the most suitable product information.

[0330] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0331] Step 1:

[0332] The server receives audio signals from the terminal. The input is audio data spoken by the user through the microphone. The audio signals are converted into a digital format and sent to the server.

[0333] Step 2:

[0334] The server converts the received audio signal into text information using the Google Cloud Speech-to-Text API. The input is a digital audio signal, and the output is text data in string format. This conversion is performed using speech recognition technology.

[0335] Step 3:

[0336] The server analyzes the generated text information using natural language processing techniques. The input is text data in string format, and the output is the analyzed meaning and related information. OpenAI's language model is used to understand the text content and extract contextually relevant information.

[0337] Step 4:

[0338] The server searches for relevant information from the database in the storage device based on the analyzed information. The input is keywords and contextual information obtained through natural language processing, and the output is the optimal response information. The database search takes into account the user's past history and trend information.

[0339] Step 5:

[0340] The server converts the generated response into speech using Google Cloud Text-to-Speech. The input is text data of the response information, and the output is synthesized speech data. This speech data is provided to the user in a natural and easy-to-understand format.

[0341] Step 6:

[0342] The terminal plays synthesized speech data received from the server. The input is speech data, and the output is sound played from the speaker. Through this, users can obtain both visual and auditory information.

[0343] Step 7:

[0344] The device generates relevant images as needed and displays them to the user. Input is an instruction to generate an image based on response information, and output is the visual information displayed on the screen. This facilitates the user's purchasing behavior.

[0345] Step 8:

[0346] Users make purchasing decisions based on the provided information and images, and send feedback to the server. Input consists of user ratings and comments, and output is stored as feedback information. This allows the system to improve its accuracy for future interactions.

[0347] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0348] This invention aims to automate and improve the efficiency of meeting systems, and is implemented in a form that includes a function to recognize user emotions and optimize responses.

[0349] Server Functions

[0350] The server stores user-uploaded materials and past conversation history in a database. It receives meeting audio streams and uses speech recognition AI to convert them into text in real time. Furthermore, the server uses an emotion engine in its natural language processing to recognize user emotions and generate contextually appropriate responses. This enables not only the exchange of information but also communication that is sensitive to the user's emotional state.

[0351] The generated responses are converted into synthesized speech data and transmitted to the terminal in real time. Furthermore, after the meeting ends, the server analyzes the statements and materials, creates a summary including changes in emotion, and provides it to the user. In particular, the user's emotion data is stored in a database as relevant information and used for future meetings.

[0352] Device functions

[0353] The device plays synthesized speech transmitted from the server through its speaker. It also uses video generation technology to create emotionally appropriate facial expressions as needed, broadcasting a virtual video of the user to other participants. This video reflects the user's current emotional state, creating a more natural participation experience.

[0354] User actions

[0355] Users can use data provided by the emotion engine to understand the impact of their own emotions on the progress of a meeting. For example, if an emotional instability occurs during a meeting, the emotion engine can detect this, and the server can generate a calming response to help the meeting proceed smoothly. Users can also review the summary provided after the meeting and, if necessary, instruct the system to prepare for the next meeting, requesting the creation of relevant materials.

[0356] Thus, the present invention provides a system in which servers, terminals, and users cooperate with each other to realize emotionally resonant information transmission and efficient meetings. The overall system aims to improve the productivity and satisfaction of meeting participants.

[0357] The following describes the processing flow.

[0358] Step 1:

[0359] As soon as the meeting starts, the server activates the speech recognition AI, receiving the audio stream transmitted from the terminal in real time and converting it into text. This text data serves as the foundational data for storing the conversation content in a database.

[0360] Step 2:

[0361] The server inputs textual information into a natural language processing system and activates an emotion engine to analyze the user's emotional state. Based on the emotion data, the server understands the context of the conversation and designs an optimal response that takes the user's emotions into consideration. This response takes into account the user's emotions and the meeting situation.

[0362] Step 3:

[0363] The server sends the generated response to a speech synthesis engine, where it is converted into speech data. This converted speech data is then sent to the terminal to be spoken on behalf of the user during the meeting. This speech is responsible for interacting with other meeting participants.

[0364] Step 4:

[0365] The device utilizes video generation technology as needed to create avatar expressions and postures that match the user's emotions. The generated video is updated in real time and displayed to other participants, contributing to the meeting as visual information.

[0366] Step 5:

[0367] After the meeting ends, the server analyzes the entire conversation and creates a summary that reflects the key points of the discussion and the users' emotional changes. This summary is organized with emotional data and provided to the users as useful data for future meetings and follow-ups.

[0368] Step 6:

[0369] Users review summaries and sentiment data sent from the server and provide feedback as needed to help improve the system. They can also request the server to create materials for the next meeting, facilitating efficient meeting participation. This request is processed automatically, and the materials are created in the specified format.

[0370] This processing flow allows the system to provide users with an emotionally responsive and sophisticated meeting experience while simultaneously supporting improved productivity and satisfaction.

[0371] (Example 2)

[0372] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0373] Traditional meeting systems merely transmit information without considering user emotions, limiting their ability to improve participant satisfaction and productivity. Furthermore, they lacked efficient support for post-meeting summarization and preparation for future meetings.

[0374] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0375] In this invention, the server includes means for storing conversation information in a storage device, means for converting audio signals into text information, and means for performing natural language analysis using the text information. This enables richer communication by analyzing the user's emotions and reflecting them in the response.

[0376] "Conversational information" refers to data on statements and content exchanged during meetings and dialogues, which is stored in a data storage device.

[0377] "Audio signal" refers to audio data in analog or digital format collected by a microphone or similar device.

[0378] "Textual information" refers to data in text format that is converted from an audio signal using speech recognition technology.

[0379] "Natural language processing" is the process of assigning meaning to and structuring textual information using algorithms such as generative AI models.

[0380] An "emotion engine" is an algorithm or software that analyzes and extracts emotions from a user's statements and actions.

[0381] A "response" is a system message generated based on the user's statements and emotions.

[0382] "Means of converting to speech and outputting" refers to the process of converting text-based response information into speech data using speech synthesis technology and providing it through speakers or other means.

[0383] "Video generation" is the process of generating videos and animations that visually represent a user's emotions and situation based on audio and text information.

[0384] "Summary information" refers to information that has been shortened and organized to include key points and changes in emotion from a meeting or conversation.

[0385] An "information terminal" is a device used by users to receive and view conversational information and summary information.

[0386] The system of this invention aims to improve meeting efficiency and enable emotionally resonant communication by having the server, terminals, and users work together in a coordinated manner.

[0387] The server supports meeting preparation by storing meeting materials and past conversation information uploaded by users in advance in a storage device. The server receives audio signals in real time and converts them into text information using a speech recognition module. Specifically, a speech recognition API can be used. The recognized text information is interpreted by a natural language processing engine, and an emotion engine is used to recognize the user's emotions. The server generates an appropriate response from the analyzed information and outputs it as speech using a speech synthesis module. This outputted speech is sent to the terminal and communicated to the user.

[0388] The device plays synthesized speech transmitted from the server through its speaker. Additionally, it uses video generation technology as needed to deliver virtual images reflecting the user's emotions to other participants. This feature helps to visually communicate the user's current emotional state to other participants, facilitating more natural communication.

[0389] Users can conduct meetings while receiving responses from the server in real time. They can also utilize emotional data analyzed by the emotion engine to understand their own emotional state and use that information to guide the meeting. After the meeting, they can review a summary provided by the server and use it to prepare for the next meeting.

[0390] A specific example of a prompt would be, "Analyze emotional changes during the meeting and generate feedback to help users relax." This allows users to reduce stress during meetings and communicate more effectively.

[0391] As described above, this system improves participant productivity and satisfaction by managing meeting information, analyzing emotions and generating responses, and distributing video based on the results.

[0392] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0393] Step 1:

[0394] The server stores meeting materials and past conversation information uploaded in advance by users in a storage device. It receives various meeting-related data provided by users as input and saves it in a structured database, making it quickly accessible when needed. As output, it generates a state that facilitates searching based on specific topics and past discussions.

[0395] Step 2:

[0396] The server receives audio signals in real time during the meeting. It receives digital audio data obtained through the microphone as input and converts it into text information using a speech recognition engine. Specifically, it utilizes a speech recognition API to output the spoken words in text format.

[0397] Step 3:

[0398] The server performs natural language processing using textual information. It receives textual information generated by speech recognition as input, and this information is analyzed by a natural language processing engine. The output provides contextual and semantic information for each utterance in a structured data format.

[0399] Step 4:

[0400] The server uses the results of natural language processing to run an emotion engine and recognize the user's emotions. It takes parsed text information as input and extracts the emotional state based on its context. The output explicitly indicates the user's emotional state by assigning tags and scores corresponding to those emotions.

[0401] Step 5:

[0402] The server generates appropriate responses based on emotion recognition. It uses the output data of an emotion engine as input and a generative AI model to generate contextually appropriate responses. The output is a text-based response that aligns with the user's emotions.

[0403] Step 6:

[0404] The server converts the generated response into synthesized speech and sends it to the terminal. It receives text responses from a generation AI model as input and converts them into speech data using a speech synthesis engine. As output, it creates a playable audio file and transfers it to the terminal in real time.

[0405] Step 7:

[0406] The terminal plays synthesized speech sent from the server through its speaker. It receives synthesized speech data as input and plays the speech via an audio output module. The output is the actual speech that the user and other participants can hear.

[0407] Step 8:

[0408] The device will use video generation technology as needed to deliver videos that express the user's emotions. It will receive emotion-tagged data as input and utilize video generation software. As output, it will generate animations or video content reflecting those emotions, which will then be distributed to other meeting participants.

[0409] Step 9:

[0410] After the meeting ends, the server analyzes each statement and document to create a summary that includes changes in emotions. As input, it integrates all textual information and emotional state data collected during the meeting, and generates a summary using an analysis algorithm. As output, a summary formatted for easy user understanding is sent to the terminal.

[0411] (Application Example 2)

[0412] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0413] In autonomous vehicles, accurately understanding passengers' emotional states is difficult, making it challenging to provide appropriate responses and support. This can result in a compromised passenger experience and prevent the vehicle from fully achieving its safety and comfort levels. Therefore, a system is needed that analyzes passengers' emotional states and provides optimal information and responses based on that analysis.

[0414] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0415] In this invention, the server includes a function for storing conversation information, a function for converting audio signals into text, a function for performing natural language understanding, a function for retrieving relevant information from the information storage unit and generating a response, a function for converting the generated response into audio data and outputting it, a function for visual representation as needed, a function for summarizing the content of the conversation, a function for transmitting the summarized information to the user, a function for analyzing the emotional state of the passengers, and a function for optimizing information presentation based on the emotional state. This makes it possible to grasp the emotional state of the passengers in real time and provide optimal information and responses accordingly.

[0416] "Conversational information" refers to audio and text-based information exchanged between users, which is stored in a database and used for analysis.

[0417] An "audio signal" is an electrical representation of sound acquired by a microphone or other audio equipment, and is used as input data for conversion into text.

[0418] "Natural language understanding" is a technology that allows computers to analyze human language and understand its meaning, enabling the generation of appropriate responses and information retrieval.

[0419] The "information storage unit" is a memory area that stores past conversation data and related information, providing the data necessary for response generation and information retrieval.

[0420] "Visual expression" refers to videos and images generated to visually convey passengers' emotional states and information, and is used to improve the user experience.

[0421] "Emotional state" refers to the psychological state of passengers and represents information analyzed from audio and video data.

[0422] "Information presentation" refers to the information and responses provided to the user, and by presenting them at the appropriate time and in the appropriate format, it improves the user experience.

[0423] In this invention, the server plays a central role in the information processing system installed in the autonomous vehicle. The server receives passenger voice signals and converts them into text information in real time using speech recognition AI. Speech recognition technologies such as Google Speech-to-Text can be used for this process. The converted text is analyzed using natural language processing tools, and the passenger's emotional state is understood by an emotion recognition engine such as IBM Watson Tone Analyzer.

[0424] Based on the analyzed data, the server generates optimal information and responses tailored to the passenger's emotions. To generate responses in natural language, a generative AI model can be used to select contextually appropriate phrases. These generated responses are converted into speech using text-to-speech technology such as Amazon Polly and output through the vehicle's speakers.

[0425] The terminal can display visuals based on the passenger's emotional state. Specifically, if a passenger is feeling stressed, it can display relaxing scenes to create a sense of security. Furthermore, the terminal can store conversational information and emotional data collected during the journey, which can then be referenced in future information provision.

[0426] For example, if the emotion engine detects that a passenger appears tired, the server can suggest, "You seem a little tired. Shall I play some soothing music?" In this way, appropriate communication can be provided according to the passenger's situation.

[0427] Examples of prompt statements are as follows:

[0428] "Perform a sentiment analysis on this text: {text}. Suggest appropriate responses based on the passenger's emotions."

[0429] This allows the system to provide information efficiently and in a user-centric manner, in response to changes in passengers' emotions.

[0430] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0431] Step 1:

[0432] The server receives audio signals from microphones placed inside the vehicle. The audio data acquired by the microphones is used as input. The data processing performed here involves converting analog audio signals into digital data.

[0433] Step 2:

[0434] The server converts the received audio data into text using speech recognition AI (e.g., Google Speech-to-Text). The input is the digital audio data acquired in step 1, and the output is text data. The speech recognition algorithm analyzes words and generates the corresponding text.

[0435] Step 3:

[0436] The server processes the generated text data using a natural language processing tool (e.g., NLTK) to analyze its content. Here, the converted text is used as input for syntactic and semantic analysis. This results in detailed, context-based text information as output.

[0437] Step 4:

[0438] The server passes the analyzed text to an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to detect the passenger's emotional state. The input is the text information from step 3, and the output is data about the recognized emotional state. The data calculation performed here involves assigning emotional weights to each word and phrase to calculate an overall emotion score.

[0439] Step 5:

[0440] The server generates the optimal response based on emotion data using an AI model. The input is emotion state data, and the output is the text of the generated response. In this process, emotion data is input to the model as a prompt sentence, and the generated response is obtained.

[0441] Step 6:

[0442] The server converts the generated response text into speech using speech synthesis technology (e.g., Amazon Polly) and sends it to the terminal. The input is the response text, and the output is audio data. A speech synthesis algorithm analyzes the text and generates an audio waveform.

[0443] Step 7:

[0444] The terminal outputs the received audio data through its speaker. It also creates and displays visual representations on the screen based on the passenger's emotional state. The input consists of the audio data and emotional data transmitted in step 6, and the output consists of audio from the speaker and visuals displayed on the screen.

[0445] Step 8:

[0446] Users monitor their emotional state based on the information and voice responses provided. This allows them to enjoy a comfortable travel experience in the environment the system offers.

[0447] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0448] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0449] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0450] [Third Embodiment]

[0451] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0452] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0453] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0454] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0455] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0456] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0457] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0458] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0459] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0460] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0461] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0462] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0463] This invention is implemented as a system that processes information in real time by storing conversation information in a database and converting audio streams into text information, in order to improve the efficiency and automation of meeting participation.

[0464] Server Functions

[0465] The server stores user-registered materials and past conversation history in a database. This data is used to provide necessary information during the meeting. When the meeting begins, the server uses speech recognition technology to convert the conversation into text in real time and puts it into a natural language processing process. The server analyzes the generated text, searches the database for relevant information, and generates an appropriate response. The generated response is converted into speech using speech synthesis technology and sent to the terminal in real time. After the meeting ends, the server automatically summarizes the conversation and sends it to the user so that they can quickly grasp the main points.

[0466] Device functions

[0467] The terminal plays synthesized speech transmitted from the server through its speaker, participating in the meeting on behalf of the user. When necessary, it also generates virtual images using video generation technology and distributes them to other participants. This allows for visual and auditory interaction even without the user's physical presence.

[0468] User actions

[0469] Users register their meeting attendance requests in the system while simultaneously managing their own schedules. After the meeting, they send necessary feedback to the server based on the summary they receive, contributing to the system's accuracy improvement. They can also request the creation of materials for future meetings. For example, even if a user is unable to attend a sales meeting, they can review the summary of the negotiation report and decisions, and send a request to the server to create additional materials if necessary.

[0470] Therefore, the present invention adopts a form that significantly improves the efficiency of information transmission by automating meetings through the cooperation of a server, terminal, and user. This system allows users to make effective use of their time and dramatically improve the productivity of their business activities.

[0471] The following describes the processing flow.

[0472] Step 1:

[0473] When the server detects the start of a meeting, it activates a speech recognition AI to collect the audio stream transmitted from the terminal and convert it into text in real time. The converted text is stored as a temporary file for smooth natural language processing.

[0474] Step 2:

[0475] When the server acquires textual information, it inputs it into a large-scale language model to analyze its content and understand the context. As the conversation progresses, the server searches for relevant information in the database and automatically generates the next necessary response. This response is intended to allow the server to participate in the meeting on behalf of the user.

[0476] Step 3:

[0477] The server converts the generated response into audio data using speech synthesis technology. It then sends the converted audio data to the terminal, conveying the specific response content to the meeting participants via the speaker.

[0478] Step 4:

[0479] The device activates a video generation AI as needed to virtually create the user's avatar and video. The generated video is updated in real time and distributed as visual information to other meeting participants. This video generation provides users in remote locations with a sense of presence as if they were actually there.

[0480] Step 5:

[0481] After the meeting ends, the server analyzes the recorded conversation and participants' comments to extract the main points and conclusions of the discussion. Using this information, the server creates a concise and easy-to-understand summary and sends it to the user via email or application.

[0482] Step 6:

[0483] Users review the summary provided by the server and provide feedback as needed. This feedback helps improve the system. Furthermore, they can request the creation of materials for the next meeting. This request is sent to the server and processed.

[0484] This processing flow allows users to efficiently participate in meetings and grasp their key points. The system provides information and automation to support business activities, contributing to increased user productivity.

[0485] (Example 1)

[0486] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0487] In today's business environment, efficient participation in meetings is crucial, but it can be difficult to attend due to busy work schedules. Furthermore, there is a need to effectively manage and utilize the information provided during meetings. Existing systems lack the functionality to completely replace meeting attendance or the ability to quickly summarize and create materials after meetings, thus limiting improvements in work efficiency.

[0488] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0489] In this invention, the server includes means for storing conversation data in an information storage device, a device for converting audio signals into text data, and a device for performing language processing using the text data. This makes it possible to convert information during a meeting into text data in real time and to immediately search for and provide relevant information. Furthermore, by automatically generating materials using the generated summary data, users can efficiently prepare for the next meeting.

[0490] "Conversation data" refers to information that shows the content of words recorded during voice communication.

[0491] An "information storage device" is hardware or software used to systematically store and manage large amounts of data.

[0492] An "audio signal" is waveform data in electrical or digital form generated by sound.

[0493] "Text data" refers to information that represents sounds, symbols, and other elements in both text and digital formats.

[0494] "Language processing" is the technology used to analyze, understand, and generate natural language using computers.

[0495] "Relevant data" refers to information that is deemed useful or necessary in relation to a particular data point or question.

[0496] A "response" is the processing result or answer returned in response to an input.

[0497] "Audio data" refers to information that represents sound as electrical signals or in digital format.

[0498] "Image data generation" is the process of creating visual representations using computer graphics.

[0499] "Summary data" refers to information that has been shortened by extracting the most important content from the original information.

[0500] "User" refers to a person or organization that uses a system or service.

[0501] "Event scheduling" is the process of planning and tracking the schedules of events and meetings for individuals or organizations.

[0502] This invention constitutes a system that streamlines and automates meeting participation. The invention mainly consists of three elements: a server, a terminal, and a user.

[0503] The server has a database as an information storage device, where it stores conversation data and materials registered in advance by users. The server processes audio signals and converts them into text data using speech recognition technology. Here, a common example of speech recognition software is a speech processing API. The text data is analyzed using a natural language processing library to quickly retrieve relevant data and generate a response. This response is converted into audio data using speech synthesis technology and sent to the terminal.

[0504] The terminal plays audio data transmitted by the server through its speaker and participates in the meeting on behalf of the user. It also generates video using image data generation technology when requested and provides it to other participants. The terminal transmits information collected through voice input to the server in real time.

[0505] Users register their attendance at meetings using an event scheduling application. After the meeting, users can review the key points based on the provided summary data and contribute to improving the system's accuracy by sending feedback to the server. Users can also request document generation from a system utilizing a generative AI model by entering prompt messages. A specific example of such a prompt message is, "Please prepare a proposal for the next technical meeting."

[0506] In this way, servers, terminals, and users cooperate, enabling users to efficiently obtain information and participate in meetings on their behalf.

[0507] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0508] Step 1:

[0509] The server stores meeting materials and past conversation data registered by users in advance in an information storage device. It receives documents and audio files provided by users as input and systematically stores them in a database. This process makes it easy to retrieve information needed in subsequent processes.

[0510] Step 2:

[0511] The device acquires audio signals in real time during meetings via a microphone. It receives physical audio as input and transmits it to the server as a digital audio signal. Noise cancellation technology is implemented to ensure high-quality audio signals.

[0512] Step 3:

[0513] The server uses speech recognition technology to convert the received audio signal into text data. Specifically, it uses a speech recognition API to perform the conversion from speech to text. This allows the conversation content to be stored as text, enabling subsequent natural language processing.

[0514] Step 4:

[0515] The server analyzes the obtained character data using a natural language processing library. Using the character data as input, it extracts linguistic context and important keywords. This lays the foundation for the server to quickly retrieve relevant information from the database.

[0516] Step 5:

[0517] The server searches the information storage device based on the analysis results to retrieve relevant data and generates a response that meets the user's needs. A generative AI model is used here, and the generated response is meaningful to the user. This output is then converted into speech in the next step.

[0518] Step 6:

[0519] The server converts the generated response into audio data using speech synthesis technology. A speech synthesis API is used for this process, and the generated audio is sent to the terminal. The terminal plays the received audio data through its speaker to communicate with the meeting participants.

[0520] Step 7:

[0521] After the meeting ends, the server automatically summarizes the conversation. Using a generative AI model, it extracts key points from the overall conversation as input. This summary is sent to the user via email or other means, allowing them to quickly review the main points of the meeting.

[0522] (Application Example 1)

[0523] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0524] In today's commercial environment, it can be difficult for customers to efficiently obtain product information and receive smooth support when making purchasing decisions within virtual stores. Furthermore, the inability to quickly understand individual customer needs and provide product suggestions and follow-up based on those needs makes improving customer satisfaction and the purchasing experience a challenge.

[0525] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0526] In this invention, the server includes means for storing conversation information in a storage device, means for converting audio signals into text information, and means for performing natural language processing using the text information. This makes it possible to analyze customer interactions in real time and present product information based on individual needs.

[0527] A "storage device" is a medium for saving conversational information and related information.

[0528] An "audio signal" is sound data acquired by an input device such as a microphone.

[0529] "Textual information" refers to text data obtained by converting audio signals.

[0530] "Natural language processing" is a technology that analyzes the meaning of text data and processes the information.

[0531] "Response" refers to the content of the reply to the customer generated by natural language processing.

[0532] An "image" is visual data that is generated as needed and provided as visual information.

[0533] "Summary information" refers to data that briefly summarizes the content of a conversation or piece of information.

[0534] "Users" refer to customers or users who utilize virtual stores or systems.

[0535] "Purchasing behavior" refers to a series of actions taken by a user to select a product and decide to purchase it.

[0536] "Opinions" refer to feedback information provided by users.

[0537] "Materials" refer to documents or data provided to support users in their preparation and decision-making.

[0538] The server receives an audio signal and converts it into text using the Google Cloud Speech-to-Text API. Furthermore, it analyzes this text using an OpenAI language model and performs natural language processing. Based on the analyzed content, it searches for relevant information in a database within the storage device and generates an appropriate response. This response is then converted back into speech using Google Cloud Text-to-Speech and sent to the terminal.

[0539] The terminal plays back the received response audio and provides it to the user. It also generates and displays images as needed to supplement visual information. Through this, users can smoothly carry out purchasing actions within the virtual store.

[0540] Users interact with an AI assistant in a virtual store and receive product information. For example, if a user is looking for sneakers, they can ask, "Which sneakers do you recommend?" and the AI ​​assistant will suggest products based on past purchase history and trend data. This system allows users to obtain information efficiently and improves the shopping experience.

[0541] An example of a prompt might be, "Based on this customer's past purchase history and trend data, please provide recommendations for appropriate sneakers." This prompt prompts the generative AI model to process information to provide the customer with the most suitable product information.

[0542] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0543] Step 1:

[0544] The server receives audio signals from the terminal. The input is audio data spoken by the user through the microphone. The audio signals are converted into a digital format and sent to the server.

[0545] Step 2:

[0546] The server converts the received audio signal into text information using the Google Cloud Speech-to-Text API. The input is a digital audio signal, and the output is text data in string format. This conversion is performed using speech recognition technology.

[0547] Step 3:

[0548] The server analyzes the generated text information using natural language processing techniques. The input is text data in string format, and the output is the analyzed meaning and related information. OpenAI's language model is used to understand the text content and extract contextually relevant information.

[0549] Step 4:

[0550] The server searches for relevant information from the database in the storage device based on the analyzed information. The input is keywords and contextual information obtained through natural language processing, and the output is the optimal response information. The database search takes into account the user's past history and trend information.

[0551] Step 5:

[0552] The server converts the generated response into speech using Google Cloud Text-to-Speech. The input is text data of the response information, and the output is synthesized speech data. This speech data is provided to the user in a natural and easy-to-understand format.

[0553] Step 6:

[0554] The terminal plays synthesized speech data received from the server. The input is speech data, and the output is sound played from the speaker. Through this, users can obtain both visual and auditory information.

[0555] Step 7:

[0556] The device generates relevant images as needed and displays them to the user. Input is an instruction to generate an image based on response information, and output is the visual information displayed on the screen. This facilitates the user's purchasing behavior.

[0557] Step 8:

[0558] Users make purchasing decisions based on the provided information and images, and send feedback to the server. Input consists of user ratings and comments, and output is stored as feedback information. This allows the system to improve its accuracy for future interactions.

[0559] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0560] This invention aims to automate and improve the efficiency of meeting systems, and is implemented in a form that includes a function to recognize user emotions and optimize responses.

[0561] Server Functions

[0562] The server stores user-uploaded materials and past conversation history in a database. It receives meeting audio streams and uses speech recognition AI to convert them into text in real time. Furthermore, the server uses an emotion engine in its natural language processing to recognize user emotions and generate contextually appropriate responses. This enables not only the exchange of information but also communication that is sensitive to the user's emotional state.

[0563] The generated responses are converted into synthesized speech data and transmitted to the terminal in real time. Furthermore, after the meeting ends, the server analyzes the statements and materials, creates a summary including changes in emotion, and provides it to the user. In particular, the user's emotion data is stored in a database as relevant information and used for future meetings.

[0564] Device functions

[0565] The device plays synthesized speech transmitted from the server through its speaker. It also uses video generation technology to create emotionally appropriate facial expressions as needed, broadcasting a virtual video of the user to other participants. This video reflects the user's current emotional state, creating a more natural participation experience.

[0566] User actions

[0567] Users can use data provided by the emotion engine to understand the impact of their own emotions on the progress of a meeting. For example, if an emotional instability occurs during a meeting, the emotion engine can detect this, and the server can generate a calming response to help the meeting proceed smoothly. Users can also review the summary provided after the meeting and, if necessary, instruct the system to prepare for the next meeting, requesting the creation of relevant materials.

[0568] Thus, the present invention provides a system in which servers, terminals, and users cooperate with each other to realize emotionally resonant information transmission and efficient meetings. The overall system aims to improve the productivity and satisfaction of meeting participants.

[0569] The following describes the processing flow.

[0570] Step 1:

[0571] As soon as the meeting starts, the server activates the speech recognition AI, receiving the audio stream transmitted from the terminal in real time and converting it into text. This text data serves as the foundational data for storing the conversation content in a database.

[0572] Step 2:

[0573] The server inputs textual information into a natural language processing system and activates an emotion engine to analyze the user's emotional state. Based on the emotion data, the server understands the context of the conversation and designs an optimal response that takes the user's emotions into consideration. This response takes into account the user's emotions and the meeting situation.

[0574] Step 3:

[0575] The server sends the generated response to a speech synthesis engine, where it is converted into speech data. This converted speech data is then sent to the terminal to be spoken on behalf of the user during the meeting. This speech is responsible for interacting with other meeting participants.

[0576] Step 4:

[0577] The device utilizes video generation technology as needed to create avatar expressions and postures that match the user's emotions. The generated video is updated in real time and displayed to other participants, contributing to the meeting as visual information.

[0578] Step 5:

[0579] After the meeting ends, the server analyzes the entire conversation and creates a summary that reflects the key points of the discussion and the users' emotional changes. This summary is organized with emotional data and provided to the users as useful data for future meetings and follow-ups.

[0580] Step 6:

[0581] Users review summaries and sentiment data sent from the server and provide feedback as needed to help improve the system. They can also request the server to create materials for the next meeting, facilitating efficient meeting participation. This request is processed automatically, and the materials are created in the specified format.

[0582] This processing flow allows the system to provide users with an emotionally responsive and sophisticated meeting experience while simultaneously supporting improved productivity and satisfaction.

[0583] (Example 2)

[0584] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0585] Traditional meeting systems merely transmit information without considering user emotions, limiting their ability to improve participant satisfaction and productivity. Furthermore, they lacked efficient support for post-meeting summarization and preparation for future meetings.

[0586] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0587] In this invention, the server includes means for storing conversation information in a storage device, means for converting audio signals into text information, and means for performing natural language analysis using the text information. This enables richer communication by analyzing the user's emotions and reflecting them in the response.

[0588] "Conversational information" refers to data on statements and content exchanged during meetings and dialogues, which is stored in a data storage device.

[0589] "Audio signal" refers to audio data in analog or digital format collected by a microphone or similar device.

[0590] "Textual information" refers to data in text format that is converted from an audio signal using speech recognition technology.

[0591] "Natural language processing" is the process of assigning meaning to and structuring textual information using algorithms such as generative AI models.

[0592] An "emotion engine" is an algorithm or software that analyzes and extracts emotions from a user's statements and actions.

[0593] A "response" is a system message generated based on the user's statements and emotions.

[0594] "Means of converting to speech and outputting" refers to the process of converting text-based response information into speech data using speech synthesis technology and providing it through speakers or other means.

[0595] "Video generation" is the process of generating videos and animations that visually represent a user's emotions and situation based on audio and text information.

[0596] "Summary information" refers to information that has been shortened and organized to include key points and changes in emotion from a meeting or conversation.

[0597] An "information terminal" is a device used by users to receive and view conversational information and summary information.

[0598] The system of this invention aims to improve meeting efficiency and enable emotionally resonant communication by having the server, terminals, and users work together in a coordinated manner.

[0599] The server supports meeting preparation by storing meeting materials and past conversation information uploaded by users in advance in a storage device. The server receives audio signals in real time and converts them into text information using a speech recognition module. Specifically, a speech recognition API can be used. The recognized text information is interpreted by a natural language processing engine, and an emotion engine is used to recognize the user's emotions. The server generates an appropriate response from the analyzed information and outputs it as speech using a speech synthesis module. This outputted speech is sent to the terminal and communicated to the user.

[0600] The device plays synthesized speech transmitted from the server through its speaker. Additionally, it uses video generation technology as needed to deliver virtual images reflecting the user's emotions to other participants. This feature helps to visually communicate the user's current emotional state to other participants, facilitating more natural communication.

[0601] Users can conduct meetings while receiving responses from the server in real time. They can also utilize emotional data analyzed by the emotion engine to understand their own emotional state and use that information to guide the meeting. After the meeting, they can review a summary provided by the server and use it to prepare for the next meeting.

[0602] A specific example of a prompt would be, "Analyze emotional changes during the meeting and generate feedback to help users relax." This allows users to reduce stress during meetings and communicate more effectively.

[0603] As described above, this system improves participant productivity and satisfaction by managing meeting information, analyzing emotions and generating responses, and distributing video based on the results.

[0604] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0605] Step 1:

[0606] The server stores meeting materials and past conversation information uploaded in advance by users in a storage device. It receives various meeting-related data provided by users as input and saves it in a structured database, making it quickly accessible when needed. As output, it generates a state that facilitates searching based on specific topics and past discussions.

[0607] Step 2:

[0608] The server receives audio signals in real time during the meeting. It receives digital audio data obtained through the microphone as input and converts it into text information using a speech recognition engine. Specifically, it utilizes a speech recognition API to output the spoken words in text format.

[0609] Step 3:

[0610] The server performs natural language processing using textual information. It receives textual information generated by speech recognition as input, and this information is analyzed by a natural language processing engine. The output provides contextual and semantic information for each utterance in a structured data format.

[0611] Step 4:

[0612] The server uses the results of natural language processing to run an emotion engine and recognize the user's emotions. It takes parsed text information as input and extracts the emotional state based on its context. The output explicitly indicates the user's emotional state by assigning tags and scores corresponding to those emotions.

[0613] Step 5:

[0614] The server generates appropriate responses based on emotion recognition. It uses the output data of an emotion engine as input and a generative AI model to generate contextually appropriate responses. The output is a text-based response that aligns with the user's emotions.

[0615] Step 6:

[0616] The server converts the generated response into synthesized speech and sends it to the terminal. It receives text responses from a generation AI model as input and converts them into speech data using a speech synthesis engine. As output, it creates a playable audio file and transfers it to the terminal in real time.

[0617] Step 7:

[0618] The terminal plays synthesized speech sent from the server through its speaker. It receives synthesized speech data as input and plays the speech via an audio output module. The output is the actual speech that the user and other participants can hear.

[0619] Step 8:

[0620] The device will use video generation technology as needed to deliver videos that express the user's emotions. It will receive emotion-tagged data as input and utilize video generation software. As output, it will generate animations or video content reflecting those emotions, which will then be distributed to other meeting participants.

[0621] Step 9:

[0622] After the meeting ends, the server analyzes each statement and document to create a summary that includes changes in emotions. As input, it integrates all textual information and emotional state data collected during the meeting, and generates a summary using an analysis algorithm. As output, a summary formatted for easy user understanding is sent to the terminal.

[0623] (Application Example 2)

[0624] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0625] In autonomous vehicles, accurately understanding passengers' emotional states is difficult, making it challenging to provide appropriate responses and support. This can result in a compromised passenger experience and prevent the vehicle from fully achieving its safety and comfort levels. Therefore, a system is needed that analyzes passengers' emotional states and provides optimal information and responses based on that analysis.

[0626] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0627] In this invention, the server includes a function for storing conversation information, a function for converting audio signals into text, a function for performing natural language understanding, a function for retrieving relevant information from the information storage unit and generating a response, a function for converting the generated response into audio data and outputting it, a function for visual representation as needed, a function for summarizing the content of the conversation, a function for transmitting the summarized information to the user, a function for analyzing the emotional state of the passengers, and a function for optimizing information presentation based on the emotional state. This makes it possible to grasp the emotional state of the passengers in real time and provide optimal information and responses accordingly.

[0628] "Conversational information" refers to audio and text-based information exchanged between users, which is stored in a database and used for analysis.

[0629] An "audio signal" is an electrical representation of sound acquired by a microphone or other audio equipment, and is used as input data for conversion into text.

[0630] "Natural language understanding" is a technology that allows computers to analyze human language and understand its meaning, enabling the generation of appropriate responses and information retrieval.

[0631] The "information storage unit" is a memory area that stores past conversation data and related information, providing the data necessary for response generation and information retrieval.

[0632] "Visual expression" refers to videos and images generated to visually convey passengers' emotional states and information, and is used to improve the user experience.

[0633] "Emotional state" refers to the psychological state of passengers and represents information analyzed from audio and video data.

[0634] "Information presentation" refers to the information and responses provided to the user, and by presenting them at the appropriate time and in the appropriate format, it improves the user experience.

[0635] In this invention, the server plays a central role in the information processing system installed in the autonomous vehicle. The server receives passenger voice signals and converts them into text information in real time using speech recognition AI. Speech recognition technologies such as Google Speech-to-Text can be used for this process. The converted text is analyzed using natural language processing tools, and the passenger's emotional state is understood by an emotion recognition engine such as IBM Watson Tone Analyzer.

[0636] Based on the analyzed data, the server generates optimal information and responses tailored to the passenger's emotions. To generate responses in natural language, a generative AI model can be used to select contextually appropriate phrases. These generated responses are converted into speech using text-to-speech technology such as Amazon Polly and output through the vehicle's speakers.

[0637] The terminal can display visuals based on the passenger's emotional state. Specifically, if a passenger is feeling stressed, it can display relaxing scenes to create a sense of security. Furthermore, the terminal can store conversational information and emotional data collected during the journey, which can then be referenced in future information provision.

[0638] For example, if the emotion engine detects that a passenger appears tired, the server can suggest, "You seem a little tired. Shall I play some soothing music?" In this way, appropriate communication can be provided according to the passenger's situation.

[0639] Examples of prompt statements are as follows:

[0640] "Perform a sentiment analysis on this text: {text}. Suggest appropriate responses based on the passenger's emotions."

[0641] This allows the system to provide information efficiently and in a user-centric manner, in response to changes in passengers' emotions.

[0642] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0643] Step 1:

[0644] The server receives audio signals from microphones placed inside the vehicle. The audio data acquired by the microphones is used as input. The data processing performed here involves converting analog audio signals into digital data.

[0645] Step 2:

[0646] The server converts the received audio data into text using speech recognition AI (e.g., Google Speech-to-Text). The input is the digital audio data acquired in step 1, and the output is text data. The speech recognition algorithm analyzes words and generates the corresponding text.

[0647] Step 3:

[0648] The server processes the generated text data using a natural language processing tool (e.g., NLTK) to analyze its content. Here, the converted text is used as input for syntactic and semantic analysis. This results in detailed, context-based text information as output.

[0649] Step 4:

[0650] The server passes the analyzed text to an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to detect the passenger's emotional state. The input is the text information from step 3, and the output is data about the recognized emotional state. The data calculation performed here involves assigning emotional weights to each word and phrase to calculate an overall emotion score.

[0651] Step 5:

[0652] The server generates the optimal response based on emotion data using an AI model. The input is emotion state data, and the output is the text of the generated response. In this process, emotion data is input to the model as a prompt sentence, and the generated response is obtained.

[0653] Step 6:

[0654] The server converts the generated response text into speech using speech synthesis technology (e.g., Amazon Polly) and sends it to the terminal. The input is the response text, and the output is audio data. A speech synthesis algorithm analyzes the text and generates an audio waveform.

[0655] Step 7:

[0656] The terminal outputs the received audio data through its speaker. It also creates and displays visual representations on the screen based on the passenger's emotional state. The input consists of the audio data and emotional data transmitted in step 6, and the output consists of audio from the speaker and visuals displayed on the screen.

[0657] Step 8:

[0658] Users monitor their emotional state based on the information and voice responses provided. This allows them to enjoy a comfortable travel experience in the environment the system provides.

[0659] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0660] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0661] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0662] [Fourth Embodiment]

[0663] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0664] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0665] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0666] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0667] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0668] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0669] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0670] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0671] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0672] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0673] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0674] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0675] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0676] This invention is implemented as a system that processes information in real time by storing conversation information in a database and converting audio streams into text information, in order to improve the efficiency and automation of meeting participation.

[0677] Server Functions

[0678] The server stores user-registered materials and past conversation history in a database. This data is used to provide necessary information during the meeting. When the meeting begins, the server uses speech recognition technology to convert the conversation into text in real time and puts it into a natural language processing process. The server analyzes the generated text, searches the database for relevant information, and generates an appropriate response. The generated response is converted into speech using speech synthesis technology and sent to the terminal in real time. After the meeting ends, the server automatically summarizes the conversation and sends it to the user so that they can quickly grasp the main points.

[0679] Device functions

[0680] The terminal plays synthesized speech transmitted from the server through its speaker, participating in the meeting on behalf of the user. When necessary, it also generates virtual images using video generation technology and distributes them to other participants. This allows for visual and auditory interaction even without the user's physical presence.

[0681] User actions

[0682] Users register their meeting attendance requests in the system while simultaneously managing their own schedules. After the meeting, they send necessary feedback to the server based on the summary they receive, contributing to the system's accuracy improvement. They can also request the creation of materials for future meetings. For example, even if a user is unable to attend a sales meeting, they can review the summary of the negotiation report and decisions, and send a request to the server to create additional materials if necessary.

[0683] Therefore, the present invention adopts a form that significantly improves the efficiency of information transmission by automating meetings through the cooperation of a server, terminal, and user. This system allows users to make effective use of their time and dramatically improve the productivity of their business activities.

[0684] The following describes the processing flow.

[0685] Step 1:

[0686] When the server detects the start of a meeting, it activates a speech recognition AI to collect the audio stream transmitted from the terminal and convert it into text in real time. The converted text is stored as a temporary file for smooth natural language processing.

[0687] Step 2:

[0688] When the server acquires textual information, it inputs it into a large-scale language model to analyze its content and understand the context. As the conversation progresses, the server searches for relevant information in the database and automatically generates the next necessary response. This response is intended to allow the server to participate in the meeting on behalf of the user.

[0689] Step 3:

[0690] The server converts the generated response into audio data using speech synthesis technology. It then sends the converted audio data to the terminal, conveying the specific response content to the meeting participants via the speaker.

[0691] Step 4:

[0692] The device activates a video generation AI as needed to virtually create the user's avatar and video. The generated video is updated in real time and distributed as visual information to other meeting participants. This video generation provides users in remote locations with a sense of presence as if they were actually there.

[0693] Step 5:

[0694] After the meeting ends, the server analyzes the recorded conversation and participants' comments to extract the main points and conclusions of the discussion. Using this information, the server creates a concise and easy-to-understand summary and sends it to the user via email or application.

[0695] Step 6:

[0696] Users review the summary provided by the server and provide feedback as needed. This feedback helps improve the system. Furthermore, they can request the creation of materials for the next meeting. This request is sent to the server and processed.

[0697] This processing flow allows users to efficiently participate in meetings and grasp their key points. The system provides information and automation to support business activities, contributing to increased user productivity.

[0698] (Example 1)

[0699] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0700] In today's business environment, efficient participation in meetings is crucial, but it can be difficult to attend due to busy work schedules. Furthermore, there is a need to effectively manage and utilize the information provided during meetings. Existing systems lack the functionality to completely replace meeting attendance or the ability to quickly summarize and create materials after meetings, thus limiting improvements in work efficiency.

[0701] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0702] In this invention, the server includes means for storing conversation data in an information storage device, a device for converting audio signals into text data, and a device for performing language processing using the text data. This makes it possible to convert information during a meeting into text data in real time and to immediately search for and provide relevant information. Furthermore, by automatically generating materials using the generated summary data, users can efficiently prepare for the next meeting.

[0703] "Conversation data" refers to information that shows the content of words recorded during voice communication.

[0704] An "information storage device" is hardware or software used to systematically store and manage large amounts of data.

[0705] An "audio signal" is waveform data in electrical or digital form generated by sound.

[0706] "Text data" refers to information that represents sounds, symbols, and other elements in both text and digital formats.

[0707] "Language processing" is the technology used to analyze, understand, and generate natural language using computers.

[0708] "Relevant data" refers to information that is deemed useful or necessary in relation to a particular data point or question.

[0709] A "response" is the processing result or answer returned in response to an input.

[0710] "Audio data" refers to information that represents sound as electrical signals or in digital format.

[0711] "Image data generation" is the process of creating visual representations using computer graphics.

[0712] "Summary data" refers to information that has been shortened by extracting the most important content from the original information.

[0713] "User" refers to a person or organization that uses a system or service.

[0714] "Event scheduling" is the process of planning and tracking the schedules of events and meetings for individuals or organizations.

[0715] This invention constitutes a system that streamlines and automates meeting participation. The invention mainly consists of three elements: a server, a terminal, and a user.

[0716] The server has a database as an information storage device, where it stores conversation data and materials registered in advance by users. The server processes audio signals and converts them into text data using speech recognition technology. Here, a common example of speech recognition software is a speech processing API. The text data is analyzed using a natural language processing library to quickly retrieve relevant data and generate a response. This response is converted into audio data using speech synthesis technology and sent to the terminal.

[0717] The terminal plays audio data transmitted by the server through its speaker and participates in the meeting on behalf of the user. It also generates video using image data generation technology when requested and provides it to other participants. The terminal transmits information collected through voice input to the server in real time.

[0718] Users register their attendance at meetings using an event scheduling application. After the meeting, users can review the key points based on the provided summary data and contribute to improving the system's accuracy by sending feedback to the server. Users can also request document generation from a system utilizing a generative AI model by entering prompt messages. A specific example of such a prompt message is, "Please prepare a proposal for the next technical meeting."

[0719] In this way, servers, terminals, and users cooperate, enabling users to efficiently obtain information and participate in meetings on their behalf.

[0720] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0721] Step 1:

[0722] The server stores meeting materials and past conversation data registered by users in advance in an information storage device. It receives documents and audio files provided by users as input and systematically stores them in a database. This process makes it easy to retrieve information needed in subsequent processes.

[0723] Step 2:

[0724] The device acquires audio signals in real time during meetings via a microphone. It receives physical audio as input and transmits it to the server as a digital audio signal. Noise cancellation technology is implemented to ensure high-quality audio signals.

[0725] Step 3:

[0726] The server uses speech recognition technology to convert the received audio signal into text data. Specifically, it uses a speech recognition API to perform the conversion from speech to text. This allows the conversation content to be stored as text, enabling subsequent natural language processing.

[0727] Step 4:

[0728] The server analyzes the obtained character data using a natural language processing library. Using the character data as input, it extracts linguistic context and important keywords. This lays the foundation for the server to quickly retrieve relevant information from the database.

[0729] Step 5:

[0730] The server searches the information storage device based on the analysis results to retrieve relevant data and generates a response that meets the user's needs. A generative AI model is used here, and the generated response is meaningful to the user. This output is then converted into speech in the next step.

[0731] Step 6:

[0732] The server converts the generated response into audio data using speech synthesis technology. A speech synthesis API is used for this process, and the generated audio is sent to the terminal. The terminal plays the received audio data through its speaker to communicate with the meeting participants.

[0733] Step 7:

[0734] After the meeting ends, the server automatically summarizes the conversation. Using a generative AI model, it extracts key points from the overall conversation as input. This summary is sent to the user via email or other means, allowing them to quickly review the main points of the meeting.

[0735] (Application Example 1)

[0736] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0737] In today's commercial environment, it can be difficult for customers to efficiently obtain product information and receive smooth support when making purchasing decisions within virtual stores. Furthermore, the inability to quickly understand individual customer needs and provide product suggestions and follow-up based on those needs makes improving customer satisfaction and the purchasing experience a challenge.

[0738] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0739] In this invention, the server includes means for storing conversation information in a storage device, means for converting audio signals into text information, and means for performing natural language processing using the text information. This makes it possible to analyze customer interactions in real time and present product information based on individual needs.

[0740] A "storage device" is a medium for saving conversational information and related information.

[0741] An "audio signal" is sound data acquired by an input device such as a microphone.

[0742] "Textual information" refers to text data obtained by converting audio signals.

[0743] "Natural language processing" is a technology that analyzes the meaning of text data and processes the information.

[0744] "Response" refers to the content of the reply to the customer generated by natural language processing.

[0745] An "image" is visual data that is generated as needed and provided as visual information.

[0746] "Summary information" refers to data that briefly summarizes the content of a conversation or piece of information.

[0747] "Users" refer to customers or users who utilize virtual stores or systems.

[0748] "Purchasing behavior" refers to a series of actions taken by a user to select a product and decide to purchase it.

[0749] "Opinions" refer to feedback information provided by users.

[0750] "Materials" refer to documents or data provided to support users in their preparation and decision-making.

[0751] The server receives an audio signal and converts it into text using the Google Cloud Speech-to-Text API. Furthermore, it analyzes this text using an OpenAI language model and performs natural language processing. Based on the analyzed content, it searches for relevant information in a database within the storage device and generates an appropriate response. This response is then converted back into speech using Google Cloud Text-to-Speech and sent to the terminal.

[0752] The terminal plays back the received response audio and provides it to the user. It also generates and displays images as needed to supplement visual information. Through this, users can smoothly carry out purchasing actions within the virtual store.

[0753] Users interact with an AI assistant in a virtual store and receive product information. For example, if a user is looking for sneakers, they can ask, "Which sneakers do you recommend?" and the AI ​​assistant will suggest products based on past purchase history and trend data. This system allows users to obtain information efficiently and improves the shopping experience.

[0754] An example of a prompt might be, "Based on this customer's past purchase history and trend data, please provide recommendations for appropriate sneakers." This prompt prompts the generative AI model to process information to provide the customer with the most suitable product information.

[0755] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0756] Step 1:

[0757] The server receives audio signals from the terminal. The input is audio data spoken by the user through the microphone. The audio signals are converted into a digital format and sent to the server.

[0758] Step 2:

[0759] The server converts the received audio signal into text information using the Google Cloud Speech-to-Text API. The input is a digital audio signal, and the output is text data in string format. This conversion is performed using speech recognition technology.

[0760] Step 3:

[0761] The server analyzes the generated text information using natural language processing techniques. The input is text data in string format, and the output is the analyzed meaning and related information. OpenAI's language model is used to understand the text content and extract contextually relevant information.

[0762] Step 4:

[0763] The server searches for relevant information from the database in the storage device based on the analyzed information. The input is keywords and contextual information obtained through natural language processing, and the output is the optimal response information. The database search takes into account the user's past history and trend information.

[0764] Step 5:

[0765] The server converts the generated response into speech using Google Cloud Text-to-Speech. The input is text data of the response information, and the output is synthesized speech data. This speech data is provided to the user in a natural and easy-to-understand format.

[0766] Step 6:

[0767] The terminal plays synthesized speech data received from the server. The input is speech data, and the output is sound played from the speaker. Through this, users can obtain both visual and auditory information.

[0768] Step 7:

[0769] The device generates relevant images as needed and displays them to the user. Input is an instruction to generate an image based on response information, and output is the visual information displayed on the screen. This facilitates the user's purchasing behavior.

[0770] Step 8:

[0771] Users make purchasing decisions based on the provided information and images, and send feedback to the server. Input consists of user ratings and comments, and output is stored as feedback information. This allows the system to improve its accuracy for future interactions.

[0772] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0773] This invention aims to automate and improve the efficiency of meeting systems, and is implemented in a form that includes a function to recognize user emotions and optimize responses.

[0774] Server Functions

[0775] The server stores user-uploaded materials and past conversation history in a database. It receives meeting audio streams and uses speech recognition AI to convert them into text in real time. Furthermore, the server uses an emotion engine in its natural language processing to recognize user emotions and generate contextually appropriate responses. This enables not only the exchange of information but also communication that is sensitive to the user's emotional state.

[0776] The generated responses are converted into synthesized speech data and transmitted to the terminal in real time. Furthermore, after the meeting ends, the server analyzes the statements and materials, creates a summary including changes in emotion, and provides it to the user. In particular, the user's emotion data is stored in a database as relevant information and used for future meetings.

[0777] Device functions

[0778] The device plays synthesized speech transmitted from the server through its speaker. It also uses video generation technology to create emotionally appropriate facial expressions as needed, broadcasting a virtual video of the user to other participants. This video reflects the user's current emotional state, creating a more natural participation experience.

[0779] User actions

[0780] Users can use data provided by the emotion engine to understand the impact of their own emotions on the progress of a meeting. For example, if an emotional instability occurs during a meeting, the emotion engine can detect this, and the server can generate a calming response to help the meeting proceed smoothly. Users can also review the summary provided after the meeting and, if necessary, instruct the system to prepare for the next meeting, requesting the creation of relevant materials.

[0781] Thus, the present invention provides a system in which servers, terminals, and users cooperate with each other to realize emotionally resonant information transmission and efficient meetings. The overall system aims to improve the productivity and satisfaction of meeting participants.

[0782] The following describes the processing flow.

[0783] Step 1:

[0784] As soon as the meeting starts, the server activates the speech recognition AI, receiving the audio stream transmitted from the terminal in real time and converting it into text. This text serves as the basic data for storing the conversation content in a database.

[0785] Step 2:

[0786] The server inputs textual information into a natural language processing system and activates an emotion engine to analyze the user's emotional state. Based on the emotion data, the server understands the context of the conversation and designs an optimal response that takes the user's emotions into consideration. This response takes into account the user's emotions and the meeting situation.

[0787] Step 3:

[0788] The server sends the generated response to a speech synthesis engine, where it is converted into speech data. This converted speech data is then sent to the terminal to be spoken on behalf of the user during the meeting. This speech is responsible for interacting with other meeting participants.

[0789] Step 4:

[0790] The device utilizes video generation technology as needed to create avatar expressions and postures that match the user's emotions. The generated video is updated in real time and displayed to other participants, contributing to the meeting as visual information.

[0791] Step 5:

[0792] After the meeting ends, the server analyzes the entire conversation and creates a summary that reflects the key points of the discussion and the users' emotional changes. This summary is organized with emotional data and provided to the users as useful data for future meetings and follow-ups.

[0793] Step 6:

[0794] Users review summaries and sentiment data sent from the server and provide feedback as needed to help improve the system. They can also request the server to create materials for the next meeting, facilitating efficient meeting participation. This request is processed automatically, and the materials are created in the specified format.

[0795] This processing flow allows the system to provide users with an emotionally responsive and sophisticated meeting experience while simultaneously supporting improved productivity and satisfaction.

[0796] (Example 2)

[0797] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0798] Traditional meeting systems merely transmit information without considering user emotions, limiting their ability to improve participant satisfaction and productivity. Furthermore, they lacked efficient support for post-meeting summarization and preparation for future meetings.

[0799] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0800] In this invention, the server includes means for storing conversation information in a storage device, means for converting audio signals into text information, and means for performing natural language analysis using the text information. This enables richer communication by analyzing the user's emotions and reflecting them in the response.

[0801] "Conversational information" refers to data on statements and content exchanged during meetings and dialogues, which is stored in a data storage device.

[0802] "Audio signal" refers to audio data in analog or digital format collected by a microphone or similar device.

[0803] "Textual information" refers to data in text format that is converted from an audio signal using speech recognition technology.

[0804] "Natural language processing" is the process of assigning meaning to and structuring textual information using algorithms such as generative AI models.

[0805] An "emotion engine" is an algorithm or software that analyzes and extracts emotions from a user's statements and actions.

[0806] A "response" is a system message generated based on the user's statements and emotions.

[0807] "Means of converting to speech and outputting" refers to the process of converting text-based response information into speech data using speech synthesis technology and providing it through speakers or other means.

[0808] "Video generation" is the process of generating videos and animations that visually represent a user's emotions and situation based on audio and text information.

[0809] "Summary information" refers to information that has been shortened and organized to include key points and changes in emotion from a meeting or conversation.

[0810] An "information terminal" is a device used by users to receive and view conversational information and summary information.

[0811] The system of this invention aims to improve meeting efficiency and enable emotionally resonant communication by having the server, terminals, and users work together in a coordinated manner.

[0812] The server supports meeting preparation by storing meeting materials and past conversation information uploaded by users in advance in a storage device. The server receives audio signals in real time and converts them into text information using a speech recognition module. Specifically, a speech recognition API can be used. The recognized text information is interpreted by a natural language processing engine, and an emotion engine is used to recognize the user's emotions. The server generates an appropriate response from the analyzed information and outputs it as speech using a speech synthesis module. This outputted speech is sent to the terminal and communicated to the user.

[0813] The device plays synthesized speech transmitted from the server through its speaker. Additionally, it uses video generation technology as needed to deliver virtual images reflecting the user's emotions to other participants. This feature helps to visually communicate the user's current emotional state to other participants, facilitating more natural communication.

[0814] Users can conduct meetings while receiving responses from the server in real time. They can also utilize emotional data analyzed by the emotion engine to understand their own emotional state and use that information to guide the meeting. After the meeting, they can review a summary provided by the server and use it to prepare for the next meeting.

[0815] A specific example of a prompt would be, "Analyze emotional changes during the meeting and generate feedback to help users relax." This allows users to reduce stress during meetings and communicate more effectively.

[0816] As described above, this system improves participant productivity and satisfaction by managing meeting information, analyzing emotions and generating responses, and distributing video based on the results.

[0817] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0818] Step 1:

[0819] The server stores meeting materials and past conversation information uploaded in advance by users in a storage device. It receives various meeting-related data provided by users as input and saves it in a structured database, making it quickly accessible when needed. As output, it generates a state that facilitates searching based on specific topics and past discussions.

[0820] Step 2:

[0821] The server receives audio signals in real time during the meeting. It receives digital audio data obtained through the microphone as input and converts it into text information using a speech recognition engine. Specifically, it utilizes a speech recognition API to output the spoken words in text format.

[0822] Step 3:

[0823] The server performs natural language processing using textual information. It receives textual information generated by speech recognition as input, and this information is analyzed by a natural language processing engine. The output provides contextual and semantic information for each utterance in a structured data format.

[0824] Step 4:

[0825] The server uses the results of natural language processing to run an emotion engine and recognize the user's emotions. It takes parsed text information as input and extracts the emotional state based on its context. The output explicitly indicates the user's emotional state by assigning tags and scores corresponding to those emotions.

[0826] Step 5:

[0827] The server generates appropriate responses based on emotion recognition. It uses the output data of an emotion engine as input and a generative AI model to generate contextually appropriate responses. The output is a text-based response that aligns with the user's emotions.

[0828] Step 6:

[0829] The server converts the generated response into synthesized speech and sends it to the terminal. It receives text responses from a generation AI model as input and converts them into speech data using a speech synthesis engine. As output, it creates a playable audio file and transfers it to the terminal in real time.

[0830] Step 7:

[0831] The terminal plays synthesized speech sent from the server through its speaker. It receives synthesized speech data as input and plays the speech via an audio output module. The output is the actual speech that the user and other participants can hear.

[0832] Step 8:

[0833] The device will use video generation technology as needed to deliver videos that express the user's emotions. It will receive emotion-tagged data as input and utilize video generation software. As output, it will generate animations or video content reflecting those emotions, which will then be distributed to other meeting participants.

[0834] Step 9:

[0835] After the meeting ends, the server analyzes each statement and document to create a summary that includes changes in emotions. As input, it integrates all textual information and emotional state data collected during the meeting, and generates a summary using an analysis algorithm. As output, a summary formatted for easy user understanding is sent to the terminal.

[0836] (Application Example 2)

[0837] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0838] In autonomous vehicles, accurately understanding passengers' emotional states is difficult, making it challenging to provide appropriate responses and support. This can result in a compromised passenger experience and prevent the vehicle from fully achieving its safety and comfort levels. Therefore, a system is needed that analyzes passengers' emotional states and provides optimal information and responses based on that analysis.

[0839] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0840] In this invention, the server includes a function for storing conversation information, a function for converting audio signals into text, a function for performing natural language understanding, a function for retrieving relevant information from the information storage unit and generating a response, a function for converting the generated response into audio data and outputting it, a function for visual representation as needed, a function for summarizing the content of the conversation, a function for transmitting the summarized information to the user, a function for analyzing the emotional state of the passengers, and a function for optimizing information presentation based on the emotional state. This makes it possible to grasp the emotional state of the passengers in real time and provide optimal information and responses accordingly.

[0841] "Conversational information" refers to audio and text-based information exchanged between users, which is stored in a database and used for analysis.

[0842] An "audio signal" is an electrical representation of sound acquired by a microphone or other audio equipment, and is used as input data for conversion into text.

[0843] "Natural language understanding" is a technology that allows computers to analyze human language and understand its meaning, enabling the generation of appropriate responses and information retrieval.

[0844] The "information storage unit" is a memory area that stores past conversation data and related information, providing the data necessary for response generation and information retrieval.

[0845] "Visual expression" refers to videos and images generated to visually convey passengers' emotional states and information, and is used to improve the user experience.

[0846] "Emotional state" refers to the psychological state of passengers and represents information analyzed from audio and video data.

[0847] "Information presentation" refers to the information and responses provided to the user, and by presenting them at the appropriate time and in the appropriate format, it improves the user experience.

[0848] In this invention, the server plays a central role in the information processing system installed in the autonomous vehicle. The server receives passenger voice signals and converts them into text information in real time using speech recognition AI. Speech recognition technologies such as Google Speech-to-Text can be used for this process. The converted text is analyzed using natural language processing tools, and the passenger's emotional state is understood by an emotion recognition engine such as IBM Watson Tone Analyzer.

[0849] Based on the analyzed data, the server generates optimal information and responses tailored to the passenger's emotions. To generate responses in natural language, a generative AI model can be used to select contextually appropriate phrases. These generated responses are converted into speech using text-to-speech technology such as Amazon Polly and output through the vehicle's speakers.

[0850] The terminal can display visuals based on the passenger's emotional state. Specifically, if a passenger is feeling stressed, it can display relaxing scenes to create a sense of security. Furthermore, the terminal can store conversational information and emotional data collected during the journey, which can then be referenced in future information provision.

[0851] For example, if the emotion engine detects that a passenger appears tired, the server can suggest, "You seem a little tired. Shall I play some soothing music?" In this way, appropriate communication can be provided according to the passenger's situation.

[0852] Examples of prompt statements are as follows:

[0853] "Perform a sentiment analysis on this text: {text}. Suggest appropriate responses based on the passenger's emotions."

[0854] This allows the system to provide information efficiently and in a user-centric manner, in response to changes in passengers' emotions.

[0855] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0856] Step 1:

[0857] The server receives audio signals from microphones placed inside the vehicle. The audio data acquired by the microphones is used as input. The data processing performed here involves converting analog audio signals into digital data.

[0858] Step 2:

[0859] The server converts the received audio data into text using speech recognition AI (e.g., Google Speech-to-Text). The input is the digital audio data acquired in step 1, and the output is text data. The speech recognition algorithm analyzes words and generates the corresponding text.

[0860] Step 3:

[0861] The server processes the generated text data using a natural language processing tool (e.g., NLTK) to analyze its content. Here, the converted text is used as input for syntactic and semantic analysis. This results in detailed, context-based text information as output.

[0862] Step 4:

[0863] The server passes the analyzed text to an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to detect the passenger's emotional state. The input is the text information from step 3, and the output is data about the recognized emotional state. The data calculation performed here involves assigning emotional weights to each word and phrase to calculate an overall emotion score.

[0864] Step 5:

[0865] The server generates the optimal response based on emotion data using an AI model. The input is emotion state data, and the output is the text of the generated response. In this process, emotion data is input to the model as a prompt sentence, and the generated response is obtained.

[0866] Step 6:

[0867] The server converts the generated response text into speech using speech synthesis technology (e.g., Amazon Polly) and sends it to the terminal. The input is the response text, and the output is audio data. A speech synthesis algorithm analyzes the text and generates an audio waveform.

[0868] Step 7:

[0869] The terminal outputs the received audio data through its speaker. It also creates and displays visual representations on the screen based on the passenger's emotional state. The input consists of the audio data and emotional data transmitted in step 6, and the output consists of audio from the speaker and visuals displayed on the screen.

[0870] Step 8:

[0871] Users monitor their emotional state based on the information and voice responses provided. This allows them to enjoy a comfortable travel experience in the environment the system offers.

[0872] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0873] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0874] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0875] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0876] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0877] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0878] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0879] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0880] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0881] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0882] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0883] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0884] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0885] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0886] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0887] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0888] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0889] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0890] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0891] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0892] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0893] The following is further disclosed regarding the embodiments described above.

[0894] (Claim 1)

[0895] A means of storing conversation information in a database,

[0896] A means of converting an audio stream into text information,

[0897] A means of performing natural language processing using textual information,

[0898] A means for retrieving relevant information from a database and generating a response,

[0899] A means for converting the generated response into speech and outputting it,

[0900] A means of generating video as needed,

[0901] Means of summarizing the content of a conversation,

[0902] A means of sending summary information to the user,

[0903] A system that includes this.

[0904] (Claim 2)

[0905] The system according to claim 1, further comprising means for receiving user feedback to improve the accuracy of the system.

[0906] (Claim 3)

[0907] The system according to claim 1, comprising means for automatically creating materials using the generated summaries to assist in preparing for the next meeting.

[0908] "Example 1"

[0909] (Claim 1)

[0910] Means for storing conversation data in an information storage device,

[0911] A device that converts audio signals into text data,

[0912] A device that performs language processing using character data,

[0913] A device that retrieves relevant data from an information storage device and generates a response,

[0914] A means for converting the generated response into audio data and transmitting it,

[0915] Means for generating image data as needed,

[0916] A means of summarizing the content of conversation data,

[0917] A means of sending summary data to users,

[0918] A means of managing users' event schedules and registering instructions for meeting participation,

[0919] A system that includes this.

[0920] (Claim 2)

[0921] The system according to claim 1, further comprising means for receiving user evaluation data and improving the accuracy of the system.

[0922] (Claim 3)

[0923] The system according to claim 1, comprising means for automatically generating materials using generated summary data and assisting in the preparation of the next meeting.

[0924] "Application Example 1"

[0925] (Claim 1)

[0926] A means for storing conversation information in a storage device,

[0927] A means of converting audio signals into text information,

[0928] A means of performing natural language processing using textual information,

[0929] A means for retrieving relevant information from a storage device and generating a response,

[0930] A means for converting the generated response into speech and outputting it,

[0931] Means for generating images as needed,

[0932] Means of summarizing the content of a conversation,

[0933] A means of sending summary information to the user,

[0934] A means of interacting with users in real time and supporting their purchasing behavior,

[0935] A system that includes this.

[0936] (Claim 2)

[0937] The system according to claim 1, further comprising means for receiving user feedback and improving the accuracy of the system.

[0938] (Claim 3)

[0939] The system according to claim 1, comprising means for automatically creating materials using the generated summaries and assisting in preparation for the next time.

[0940] "Example 2 of combining an emotion engine"

[0941] (Claim 1)

[0942] A means for storing conversation information in a storage device,

[0943] A means of converting audio signals into text information,

[0944] A means of performing natural language analysis using textual information,

[0945] A means of recognizing emotions and generating responses from the analysis results,

[0946] A means for converting the generated response into speech and outputting it,

[0947] A means of generating video as needed,

[0948] A means of summarizing the content of a conversation and including emotional changes in that content,

[0949] A means of transmitting summary information to an information terminal,

[0950] A system that includes this.

[0951] (Claim 2)

[0952] The system according to claim 1, comprising means for analyzing a user's emotions using an emotion engine.

[0953] (Claim 3)

[0954] The system according to claim 1, comprising means for automatically creating materials using the generated summaries and assisting in the preparation of the next meeting.

[0955] "Application example 2 when combining with an emotional engine"

[0956] (Claim 1)

[0957] A function to store conversation information,

[0958] A function that converts audio signals to text,

[0959] Features for natural language comprehension,

[0960] A function that retrieves relevant information from the information storage unit and generates a response,

[0961] A function that converts the generated response into audio data and outputs it,

[0962] A function to perform visual expression as needed,

[0963] A function to summarize the content of the conversation,

[0964] A function to send summary information to the user,

[0965] A function to analyze the emotional state of passengers,

[0966] A function that optimizes information presentation based on emotional state,

[0967] An information processing system that includes this.

[0968] (Claim 2)

[0969] The information processing system according to claim 1, comprising a function to acquire user responses and improve the system's performance.

[0970] (Claim 3)

[0971] The information processing system according to claim 1, comprising a function to automatically create informational materials using the generated summaries and to assist in preparing for the next information exchange. [Explanation of Symbols]

[0972] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of storing conversation information in a database, A means of converting an audio stream into text information, A means of performing natural language processing using textual information, A means for retrieving relevant information from a database and generating a response, A means for converting the generated response into speech and outputting it, A means of generating video as needed, Means of summarizing the content of a conversation, A means of sending summary information to the user, A system that includes this.

2. The system according to claim 1, further comprising means for receiving user feedback and improving the accuracy of the system.

3. The system according to claim 1, comprising means for automatically creating materials using the generated summary and assisting in the preparation of the next meeting.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A