system

The system addresses inefficiencies in knowledge management by collecting and training data for voice-based responses, enhancing data management and user interaction for immediate access and effective knowledge sharing.

JP2026064806APending Publication Date: 2026-04-14SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Conventional knowledge management systems face challenges in efficiently collecting, managing, and searching for in-company information, require frequent updates, and lack interactive voice synthesis for immediate and easy access, making effective knowledge sharing and personnel training difficult.

Method used

A system that collects knowledge data from databases, trains machine learning models to generate answers, converts them into audio using speech synthesis software, and provides audio data to users, enabling efficient and intuitive knowledge management and talent development.

Benefits of technology

Enables flexible data manipulation, centralized management, and immediate access to knowledge through voice responses, improving user experience and knowledge sharing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026064806000001_ABST
    Figure 2026064806000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of collecting knowledge data from a database, A method for feeding collected knowledge data into a machine learning model and performing training, A means for generating answers to user questions using a trained model, A means of converting the generated response into speech data using speech synthesis software, Means of providing audio data to users, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In conventional knowledge management systems, it has been difficult to efficiently collect, manage, and search for in - company information and acquired data, and frequent updates have been required. Also, conventional personnel training methods require individual support, making effective knowledge sharing and response difficult. Furthermore, the lack of an interactive response system using voice synthesis has hindered immediacy and ease of access. The object of the present invention is to solve these problems and provide a new tool that enables efficient knowledge management and interactive personnel training.

Means for Solving the Problems

[0005] This invention provides a system that includes means for collecting knowledge data from a database, feeding it to a machine learning model for training, and generating answers to user questions. Furthermore, it includes means for converting the generated answers into audio data using speech synthesis software and providing the audio data to the user. This enables efficient and intuitive knowledge management and talent development. In addition, by collecting knowledge data through APIs and data sources, and managing the knowledge data stored in the database in JSON format, flexible data manipulation and centralized management are achieved.

[0006] A "database" is a digital repository used to efficiently and quickly search, store, edit, and delete structured data.

[0007] "Knowledge data" refers to a collection of information and knowledge gathered and stored within a company or organization, and includes data related to specific tasks or projects.

[0008] A "machine learning model" is a set of algorithms that are trained using data to learn patterns and rules without explicit human programming, and to perform predictions and classifications.

[0009] "Training" is the process of providing training data to a machine learning model so that the model can effectively perform a specific task.

[0010] A "user" refers to an end-user who utilizes a system or software application, and in this case, it is a person who obtains information from knowledge data.

[0011] A "question" is a query that a user submits to the system to retrieve specific information from knowledge data.

[0012] An "answer" is the information or knowledge that a system provides based on a user's question using generative AI.

[0013] "Speech synthesis software" refers to software or an API used to convert text data into speech data, and is a technology that provides speech information to users.

[0014] "Audio data" refers to digital audio files generated by speech synthesis software that users can listen to.

[0015] "API" stands for Application Programming Interface, which is a set of definitions and protocols for different software systems to communicate with each other.

[0016] "JSON format" is an abbreviation for JavaScript (registered trademark) Object Notation, and is a standard format for structuring data in a human-readable text format.

[0017] A "system" is an integrated whole in which multiple components or subsystems work together to perform a specific function or task. [Brief explanation of the drawing]

[0018] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0019] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0022] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0023] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0024] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0026] [First Embodiment]

[0027] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0028] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0031] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0034] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0038] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0039] The system of the present invention is implemented according to the following procedure.

[0040] First, knowledge data is collected. During this collection process, the server periodically retrieves necessary information from the company's internal knowledge database and external data sources. For example, knowledge data can be retrieved via an API, converted into a common format (e.g., JSON), and stored in the database. This ensures data consistency and availability.

[0041] Next, the process of generating answers to user questions is executed using a trained generative AI model. The server trains the machine learning model using knowledge data collected in advance as training data. Through this training process, the model learns patterns and rules in the knowledge data and gains the ability to generate appropriate answers to user questions.

[0042] When a user enters a question, the device sends that question to a server. The server inputs the question into a generative AI and generates an appropriate answer. The generated answer is sent to speech synthesis software, where the text data is converted into speech data. This is done, for example, using the Google® Text-to-Speech API.

[0043] The generated audio data is sent from the server to the terminal, which then plays it back to provide the user with an answer. This entire process allows the user to quickly and intuitively obtain useful information from the knowledge data.

[0044] Furthermore, the system saves usage history as logs and performs effectiveness analysis. Success stories are selected and documented as specific examples. These success stories are packaged as proposal materials for implementation to other companies. This packaged material is used as a guideline when other companies implement similar systems.

[0045] As a concrete example, consider a scenario where a project manager at a certain company asks about the procedure for launching a new project. The user enters the question, "How do I create a new project?" into the terminal. The terminal sends this query to the server, which then sends it to a generative AI model. The model generates the appropriate procedure from its trained knowledge data, and the server sends this answer to speech synthesis software, which generates audio data. The terminal plays the audio data received from the server, providing the project manager with the appropriate creation procedure in audio format.

[0046] In this way, the system of the present invention realizes efficient and intuitive knowledge management and information provision to users.

[0047] The following describes the processing flow.

[0048] Step 1:

[0049] Server: Accesses internal knowledge databases and external data sources, and collects knowledge data using APIs. Specifically, it retrieves data using HTTP requests and converts the data into JSON format.

[0050] Step 2:

[0051] Server: Stores the collected knowledge data in a database. For example, it uses a NoSQL database such as MongoDB to store data in JSON format.

[0052] Step 3:

[0053] Server: Initializes generative AI models for machine learning. Sets up the training environment and retrieves knowledge data from the database in batches.

[0054] Step 4:

[0055] Server: Trains a generative AI model using the acquired knowledge data. This process uses pairs of questions and their corresponding answers.

[0056] Step 5:

[0057] User: Enter your question from your device and click the submit button. For example, enter "How do I create a new project?"

[0058] Step 6:

[0059] Terminal: Sends the entered question to the server. Specifically, it passes the question data to the server using an HTTP POST request.

[0060] Step 7:

[0061] Server: Inputs questions received from users into a generative AI model and generates the optimal answer.

[0062] Step 8:

[0063] Server: Retrieves the generated response and sends it to the text-to-speech software. For example, it sends a request to convert text into speech data using the Google Text-to-Speech API.

[0064] Step 9:

[0065] Server: Receives the speech data returned from the speech synthesis software and caches or temporarily stores it as needed.

[0066] Step 10:

[0067] Server: Sends audio data to the terminal. Specifically, it returns the audio data to the terminal as an HTTP response.

[0068] Step 11:

[0069] Terminal: Plays back the received audio data and provides the user with an audio response. The user can listen to the audio and obtain the answer to the question.

[0070] Step 12:

[0071] Server: Use history is saved as logs, and effectiveness analysis is performed. Specifically, log data stored in the database is analyzed, and successful cases are identified.

[0072] Step 13:

[0073] Server: Document and package success stories to create proposal materials for implementation at other companies. Specifically, compile them into presentation materials and digital booklets.

[0074] Through this series of processing steps, the system of the present invention achieves efficient and intuitive knowledge management and information provision to users.

[0075] (Example 1)

[0076] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0077] Traditional knowledge database systems had problems with quickly obtaining the information users needed. Furthermore, text-based responses alone were difficult to understand intuitively, potentially degrading the user experience. In addition, the lack of usage history analysis and effective feedback made system improvement difficult.

[0078] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0079] In this invention, the server includes means for collecting knowledge data from a database, means for feeding the collected knowledge data to a machine learning model and training it, means for receiving questions entered by the user through a terminal and sending them to the server, means for generating answers to the user's questions using the trained generative AI model, means for converting the generated answers into audio data using speech synthesis software, means for sending the generated audio data to the terminal and providing the audio data to the user, and means for saving the system usage history as a log and performing effectiveness analysis. This enables the user to efficiently and intuitively obtain useful information from the knowledge data.

[0080] A "database" is a computer system that systematically manages a collection of data and allows for efficient searching and updating.

[0081] "Knowledge data" refers to data that systematically organizes the knowledge and information held by a company or organization.

[0082] "Means of collection" refers to methods and technologies for gathering and acquiring specific information and storing it in a database or similar system.

[0083] A "machine learning model" is an algorithm or mathematical model that learns patterns from data and uses them to make predictions and classifications about unseen data.

[0084] "Training methods" refer to the methods and techniques used to execute the process by which a machine learning model learns from knowledge data.

[0085] A "generative AI model" is a trained model that uses artificial intelligence technology to generate appropriate outputs for specific inputs.

[0086] A "user question" is a question that a user enters through their device to find information or to solve a problem they want to resolve.

[0087] A "trained generative AI model" is an AI model that has been trained using knowledge data through a machine learning algorithm.

[0088] "Means for generating answers" refers to methods and technologies for generating appropriate information in response to a user's question and creating an answer.

[0089] "Speech synthesis software" is software that takes text data as input and generates speech that sounds like a human voice.

[0090] "Means of converting to audio data" refers to methods and technologies for converting generated text-formatted responses into audio data.

[0091] "Means of providing audio data to users" refers to methods and technologies for presenting and allowing users to listen to the generated audio data.

[0092] "Usage history" refers to a record of how the system was used, including logs of user questions and responses.

[0093] "Effectiveness analysis" is the process of analyzing system performance and user satisfaction based on collected usage history to identify areas for improvement.

[0094] The system of this invention mainly consists of three components: a server, a terminal, and a user. The roles of each component and specific operating procedures are described below.

[0095] First, the server collects knowledge data from databases. This collection is done via APIs from internal corporate knowledge databases or external data sources. The acquired data is converted to JSON format and stored in a database such as MongoDB. The collected knowledge data is then trained using machine learning models (e.g., BERT or GPT-3®). Python libraries such as TENSORFLOW® and PyTorch are used for this training process. The trained generative AI model is then used to generate answers to user questions.

[0096] The user enters a question through the device's interface (e.g., a web browser). For example, they might enter a question like, "How do I create a new project?" into a text box. The device converts the user's question into JSON format and sends it to the server using an HTTP POST request. The server decodes the received question and inputs it into a generating AI model. The AI ​​model generates an appropriate answer based on a comprehensive knowledge database.

[0097] The generated responses are converted into audio data using text-to-speech software (e.g., Google Text-to-Speech API). The text-to-speech software takes text data as input and generates an audio file (e.g., MP3 format). The generated audio data is sent from the server to the terminal, which plays it using its built-in audio player to provide the user with the response.

[0098] Furthermore, the server stores system usage history as logs and performs effectiveness analysis. For example, log data is stored and analyzed using Elasticsearch®. This allows for evaluation of system performance and user satisfaction, and identification of areas for improvement. Success stories are compiled into regular reports and used as proposal materials for implementation to other companies.

[0099] As a concrete example, let's consider a scenario where a project manager at a certain company asks about the procedure for launching a new project. In this case, the user enters the question "How do I create a new project?" into the terminal. The terminal sends this query to the server, which then sends it to a generative AI model. The model generates the appropriate procedure from its trained knowledge data, and the server sends that answer to speech synthesis software to generate audio data. The terminal plays the audio data received from the server, providing the project manager with the appropriate creation procedure in audio format.

[0100] Other examples of prompt statements include:

[0101] "Please tell me about best practices for project management."

[0102] "What are the key points when drafting contracts with clients?"

[0103] In this way, the system of the present invention realizes efficient and intuitive knowledge management and information provision to users.

[0104] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0105] Step 1: Collecting Knowledge Data

[0106] The server collects knowledge data from databases and external data sources. For example, the server uses HTTP requests to retrieve knowledge data within the company. The retrieved data is converted to JSON format and stored in a database such as MongoDB. The input is raw data retrieved from APIs and data sources, and the output is knowledge data in JSON format.

[0107] Step 2: Training the machine learning model

[0108] The server trains a generative AI model using the collected knowledge data. Specifically, it trains the model using Python's TensorFlow or PyTorch libraries. Data preprocessing (cleaning, normalization, etc.) is also performed during training. The input is knowledge data in JSON format and an initial machine learning model, and the output is a trained generative AI model.

[0109] Step 3: Enter user questions

[0110] The user enters the question through the terminal's interface. For example, they might type "How do I create a new project?" into a text box in a web browser. The input is the text-based question entered by the user, and the output is the question data in JSON format generated by the terminal.

[0111] Step 4: Sending the question to the server

[0112] The terminal converts the user's input into JSON format and sends it to the server using an HTTP POST request. The input is the JSON formatted question data generated by the terminal, and the output is the question request sent to the server.

[0113] Step 5: Generating answers using a generative AI model

[0114] The server decodes the received question and inputs it into a generative AI model. The AI ​​model generates an appropriate answer to the question based on the trained knowledge data. The input consists of the question data sent to the server and the trained generative AI model, and the output is the generated answer in text format.

[0115] Step 6: Speech synthesis processing

[0116] The server sends the generated text-based response to speech synthesis software, which converts it into audio data. Specifically, it calls the Google Text-to-Speech API to generate an audio file (e.g., in MP3 format). The input is the generated text-based response, and the output is the generated audio data.

[0117] Step 7: Providing the response to the user

[0118] The server sends the generated audio data to the terminal as an HTTP response. The terminal receives this and plays it using its built-in audio player. The input is the audio data sent from the server, and the output is the audio response presented to the user.

[0119] Step 8: Log usage history and analyze its effectiveness.

[0120] The server stores system usage history as logs and performs periodic effectiveness analyses. For example, it uses tools such as Elasticsearch to analyze log data and evaluate system performance and user satisfaction. The input is the usage history logs, and the output is analysis results and improvement suggestions.

[0121] (Application Example 1)

[0122] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0123] Traditional knowledge management systems typically rely on text-based interfaces for users to obtain answers to their questions, making intuitive and rapid information provision difficult. Furthermore, there is a lack of tools that allow store staff to instantly access knowledge data and provide accurate answers during customer service in physical stores. Therefore, improving the efficiency of customer service in physical stores and enhancing customer satisfaction are key challenges.

[0124] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0125] In this invention, the server includes means for collecting knowledge data from a database, means for feeding the collected knowledge data to a machine learning model and training it, means for generating answers to user questions using the trained model, means for converting the generated answers into voice data using speech synthesis software, means for providing the voice data while the user is wearing smart glasses, and means for running the speech synthesis software on the smart glasses when providing the voice data. This enables store clerks in physical stores to quickly and accurately provide answers to customer questions through smart glasses, thereby improving the efficiency of customer service and increasing customer satisfaction.

[0126] A "database" is an information system used to centrally store and manage knowledge data collected from both inside and outside a company.

[0127] "Knowledge data" refers to knowledge and information collected from data sources both inside and outside a company, and is used in question-answering systems.

[0128] A "machine learning model" is an algorithm that is trained on collected knowledge data to generate appropriate answers to user questions.

[0129] "Training" is the process of inputting collected knowledge data into a machine learning model and allowing that model to learn.

[0130] "Generated answers" refer to the information that a machine learning model outputs in response to a user's question.

[0131] "Speech synthesis software" is a program that converts text data into speech data.

[0132] "Smart glasses" are wearable devices that users wear and that provide information via displays and audio.

[0133] A "user" is a person who uses this system to ask questions and receive answers.

[0134] A "server" is a central computer that collects, processes, and stores data, and is responsible for providing various services.

[0135] As an embodiment of this invention, a system is constructed using the following configuration and procedure. The system mainly consists of a server, a terminal (smart glasses), a knowledge database, a machine learning model, and speech synthesis software.

[0136] First, the server periodically collects knowledge data from data sources both inside and outside the company. During this collection process, knowledge data is retrieved via APIs, converted into a common format (e.g., JSON), and stored in a knowledge database. This ensures data consistency and availability.

[0137] Next, the server uses the collected knowledge data to train a machine learning model. Through this training process, the machine learning model learns patterns and rules in the knowledge data and gains the ability to generate appropriate answers to user questions.

[0138] When a user wears smart glasses and enters a question, the device sends the question to a server. The server inputs the question into a generative AI model and generates an appropriate answer. The generated answer is sent to speech synthesis software, where the text data is converted into speech data. Specifically, the Google Text-to-Speech API can be used.

[0139] The generated audio data is sent from the server to the smart glasses, which then play it back to provide the user with an answer. This entire process allows the user to quickly and intuitively obtain useful information from the knowledge data.

[0140] Furthermore, the system saves usage history as logs and performs effectiveness analysis. Success stories are selected and documented as specific examples. These success stories are packaged as proposal materials for implementation in other companies. This packaged material is used as a guideline when other companies implement similar systems.

[0141] As a concrete example, consider a scenario where a store clerk in a physical store is wearing smart glasses while performing their duties. If a customer asks, "How is the sizing of these shoes?", an application built into the clerk's smart glasses recognizes the question and transmits it to a server. A generative AI model generates a response such as, "These shoes run a little smaller than usual, so we recommend going up one size," and provides this as voice data to the clerk's smart glasses. The clerk can then provide the customer with an appropriate answer through this voice data.

[0142] Examples of prompt statements include:

[0143] "Knowledge Data API Endpoint: https: / / example.com / api / knowledge"

[0144] "AI Model API Endpoint: https: / / example.com / api / ai_model"

[0145] Question: How is the sizing of these shoes?

[0146] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0147] Step 1:

[0148] The server periodically collects knowledge data. During this process, it retrieves necessary information from internal and external data sources via APIs, converts this data into a common format (such as JSON), and stores it in the knowledge database. The input is raw data obtained from data sources, and the output is knowledge data converted into the common format.

[0149] Step 2:

[0150] The server feeds the collected knowledge data into a machine learning model for training. Here, the input is the knowledge data stored in the knowledge database, and the output is the trained machine learning model. Each data point is input into the model, and an iterative training process is performed to teach the model patterns and rules in the knowledge data.

[0151] Step 3:

[0152] When a user wears smart glasses and inputs a question from a customer, the terminal (smart glasses) recognizes the question and sends it to the server. Here, the input is the question in the form of voice or text, and the output is the transfer of the question data to the server. This includes the specific operation of using the smart glasses' microphone and voice recognition function to convert the question into text.

[0153] Step 4:

[0154] The server inputs the received question into a generative AI model and generates an appropriate answer. Here, the input is the question text sent by the user, and the output is the generated answer text. The generative AI model executes an algorithm that generates the best possible answer to the question based on its pre-trained knowledge.

[0155] Step 5:

[0156] The server sends the generated response text to text-to-speech software, which converts it into audio data. Here, the input is the generated response text, and the output is the audio data. Specifically, the text is converted into an audio file using APIs such as the Google Text-to-Speech API.

[0157] Step 6:

[0158] The server sends the generated audio data to the smart glasses, and the device plays it back and provides it to the user. Here, the input is the audio data, and the output is the audio played back through the smart glasses. This includes the specific operation of playing the audio using the smart glasses' speaker.

[0159] Step 7:

[0160] The system saves the entire usage history as a log and performs effectiveness analysis. Here, the input is the usage history data, and the output is the analysis results. Through log saving and analysis, success stories are selected and implementation proposal materials are created for other companies.

[0161] This series of processing steps enables store staff in physical stores to provide quick and accurate answers to customer questions through smart glasses, thereby improving the efficiency of customer service and increasing customer satisfaction.

[0162] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0163] The system of the present invention includes means for collecting knowledge data from a database, feeding it to a machine learning model for training, and generating answers to user questions. It also includes means for converting the generated answers into audio data using speech synthesis software and providing the audio data to the user. By combining this system with an emotion engine, it is possible to recognize the user's emotions and provide optimal answers based on them.

[0164] First, let me explain how this system works. The server periodically collects knowledge data from the company's internal knowledge database and external data sources. For example, it retrieves data via an API, converts it to a common format (e.g., JSON), and stores it in the database.

[0165] Next, a process is executed to generate answers to user questions using a trained generative AI model. The server trains the machine learning model using pre-collected knowledge data. Through this training process, the model learns patterns and rules in the knowledge data and gains the ability to generate appropriate answers to user questions.

[0166] When a user enters a question, the device sends it to the server. The server inputs the question into a generative AI model and generates the optimal answer. At this time, the emotion engine analyzes the user's input and recognizes their emotions. The recognized emotion information is fed back to the generative AI model, which then generates the optimal answer.

[0167] The generated responses are sent to text-to-speech software, where the text data is converted into speech data. This is done, for example, using the Google Text-to-Speech API. The text-to-speech software adjusts the speech tone based on the emotion information obtained from the emotion engine and generates speech corresponding to each emotion.

[0168] The generated audio data is sent from the server to the terminal, which then plays it back to provide the user with an answer. This entire process allows the user to quickly and intuitively obtain useful information from knowledge data, while also receiving emotionally sensitive responses.

[0169] As a concrete example, consider a scenario where a customer support representative at a certain company asks, "How should I handle this complaint?" The user enters this question into a terminal and clicks the send button. The terminal sends this query to a server, which then sends it to a generative AI model. The model generates an appropriate response method from its trained knowledge data, and the server sends this response to speech synthesis software to generate audio data. In this process, an emotion engine recognizes the user's emotions, such as urgency and dissatisfaction, and generates a response in a corresponding voice tone. The terminal plays the audio data received from the server, providing the customer support representative with an appropriate response method in audio.

[0170] In this way, the system of the present invention realizes efficient and intuitive knowledge management and information provision to users, as well as enabling responses that take into account the user's feelings.

[0171] The following describes the processing flow.

[0172] Step 1:

[0173] Server: Accesses internal knowledge databases and external data sources, and collects knowledge data using APIs. For example, it retrieves data using HTTP requests and converts the data into JSON format.

[0174] Step 2:

[0175] Server: Stores the collected knowledge data in a database. For example, it uses a NoSQL database such as MongoDB to store data in JSON format.

[0176] Step 3:

[0177] Server: Initializes generative AI models for machine learning. Sets up the training environment and retrieves knowledge data from the database in batches.

[0178] Step 4:

[0179] Server: Trains a generative AI model using the acquired knowledge data. This process uses pairs of questions and their corresponding answers.

[0180] Step 5:

[0181] User: Enter your question from your device and click the submit button. For example, enter "How do I create a new project?"

[0182] Step 6:

[0183] Terminal: Sends the entered question to the server. Specifically, it passes the question data to the server using an HTTP POST request.

[0184] Step 7:

[0185] Server: Inputs questions received from users into a generative AI model to generate the optimal answer. During this process, the emotion engine analyzes the user's input and recognizes their emotions. For example, natural language processing techniques are used to extract the user's emotions from the text.

[0186] Step 8:

[0187] Server: Based on recognized emotion information, a generative AI model generates responses with a tone and content appropriate to that emotion. For example, if the user is expressing dissatisfaction, it will generate a more helpful and empathetic response.

[0188] Step 9:

[0189] Server: Retrieves the generated response and sends it to the text-to-speech software. For example, it sends a request to convert text into speech data using the Google Text-to-Speech API.

[0190] Step 10:

[0191] Server: Receives the voice data returned from the speech synthesis software and caches or temporarily stores it as needed. At this time, the speech synthesis software adjusts the voice tone based on emotional information.

[0192] Step 11:

[0193] Server: Sends audio data to the terminal. Specifically, it returns the audio data to the terminal as an HTTP response.

[0194] Step 12:

[0195] Terminal: Plays back received audio data and provides the user with an audio response. The user can listen to the audio and obtain the answer to their question. For example, the voice tone may be friendly and empathetic.

[0196] Step 13:

[0197] Server: Use history is saved as logs, and effectiveness analysis is performed. Specifically, log data stored in the database is analyzed, and successful cases are identified.

[0198] Step 14:

[0199] Server: Document and package success stories to create proposal materials for implementation at other companies. Specifically, compile them into presentation materials and digital booklets.

[0200] This series of processing steps enables the system of the present invention not only to achieve efficient and intuitive knowledge management and information provision to users, but also to respond in a way that takes users' emotions into consideration.

[0201] (Example 2)

[0202] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0203] Conventional knowledge management systems have struggled to provide quick and appropriate answers to user questions. Furthermore, they have been unable to provide answers that take user emotions into consideration, making it difficult to improve user satisfaction. This invention aims to solve these problems and provide a superior user experience.

[0204] In Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for collecting knowledge data from a database, means for feeding the collected knowledge data to a machine learning model and training it, means for generating answers to user questions using the trained model, means for converting the generated answers into voice data using speech synthesis software, means for analyzing the user's emotions, means for generating the optimal answer based on the analyzed emotional information, and means for providing the voice data to the user. This enables quick and appropriate answers to user questions, and further improves user satisfaction by providing answers that take the user's emotions into consideration.

[0205] A "database" is a software system for systematically organizing, storing, searching, updating, and managing information.

[0206] "Knowledge data" refers to data that companies and organizations have accumulated and stored in digital format, encompassing experience, knowledge, and information.

[0207] A "machine learning model" is a collection of algorithms that learn patterns and rules based on data and use them to make predictions and classifications on new data.

[0208] "Training" is the process of using knowledge data to teach a machine learning model patterns and rules.

[0209] A "generative AI model" is a model trained using machine learning that has the ability to generate answers to user questions.

[0210] "Speech synthesis software" is software that converts text data into speech data.

[0211] An "emotion engine" is a technology that analyzes user emotions from their input and actions and generates information based on those emotions.

[0212] An "API" is an interface for exchanging functions and data between different software programs.

[0213] A "data source" refers to an external system or service that provides data.

[0214] The "JSON format" is a text-based data format for representing data in a lightweight and highly readable format.

[0215] The system of the present invention includes means for collecting knowledge data from a database, feeding it to a machine learning model for training, and generating answers to user questions. It also includes means for converting the generated answers into audio data using speech synthesis software and providing the audio data to the user. Furthermore, this system incorporates an emotion engine that can recognize the user's emotions and provide the most appropriate answers based on those emotions.

[0216] First, the server collects knowledge data from the company's internal knowledge database and external data sources. Specifically, it retrieves data from external data sources via APIs, converts it to JSON format, and stores it in the database. Next, the server uses the collected knowledge data to train a generative AI model. This training process uses machine learning frameworks such as TensorFlow.

[0217] When a user enters a question, the device sends it to the server. The server inputs the received question into a generative AI model to generate the best possible answer. In this process, an emotion engine analyzes the user's input and recognizes their emotions. The recognized emotion information is fed back to the generative AI model, which then generates the best possible answer.

[0218] The generated response is sent to text-to-speech software (e.g., Google Text-to-Speech API), where the text data is converted into audio data. During this speech synthesis process, emotional information obtained from an emotion engine is used to adjust the voice tone. The generated audio data is then sent from the server to the device, which plays it back to provide the response to the user.

[0219] As a concrete example, consider a scenario where a customer support representative at a certain company asks, "How do I handle complaints?" The user enters this question into a device and clicks the send button. The device sends the query to a server, which passes it to a generative AI model. The generative AI model generates an appropriate answer from its trained knowledge data, and the server converts that answer into audio data using the Google Text-to-Speech API. In this process, an emotion engine recognizes the user's emotions, such as urgency and frustration, and generates an answer with an appropriate tone of voice. The device plays the audio data received from the server and provides the answer to the customer support representative.

[0220] Example of a prompt:

[0221] "Create a customer support system that generates responses to user inquiries about how to handle complaints. Please explain the process for outputting responses in an appropriate tone of voice, taking into account the user's emotional state and potential anxiety."

[0222] In this way, the system of the present invention realizes efficient and intuitive knowledge management and information provision to users, and further enables responses that take into account the user's feelings.

[0223] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0224] Step 1: Data Collection

[0225] The server collects knowledge data from the company's internal knowledge database and external data sources. Specifically, the server uses APIs to retrieve data from external data sources. This data is converted to JSON format and stored in the database. The input is raw data retrieved from the API, and the output is knowledge data converted to JSON format.

[0226] Step 2: Model Training

[0227] The server uses the collected knowledge data to train a generative AI model. Specifically, the server loads knowledge data from a database and performs data preprocessing. For training, it uses a machine learning framework such as TensorFlow, and optimizes the model by repeating multiple epochs. The input is knowledge data in JSON format, and the output is a trained generative AI model.

[0228] Step 3: Receiving questions from users

[0229] The user enters a question via the device. For example, they might enter "How should I handle complaints?" through a web browser's input form. When the user presses the submit button, the device sends the question to the server as an HTTP POST request. The input is the user's question text, and the output is the question data in HTTP request format.

[0230] Step 4: Analyzing the Question

[0231] The server inputs the received question into a generative AI model and generates an answer. Specifically, the server first parses the question text and inputs it into the generative AI model. During this process, the emotion engine analyzes the user's input and recognizes their emotions. The input is question data in HTTP request format, and the output is answer text with added emotion information.

[0232] Step 5: Generating the answer

[0233] The generative AI model generates the optimal answer based on the input question and sentiment information. The generated answer is temporarily stored in the server's memory. The input is question data including sentiment information, and the output is the generated answer text.

[0234] Step 6: Generate audio data

[0235] The server sends the generated response to speech synthesis software, which converts it into speech data. Specifically, the server uses the Google Text-to-Speech API to convert the text data into speech data. In this process, emotional information obtained from the emotion engine is reflected in the tone of voice. The input is the response text, and the output is emotion-adjusted speech data.

[0236] Step 7: Providing audio data

[0237] The generated audio data is sent from the server to the terminal, which then plays it back to provide a response to the user. Specifically, the terminal plays the audio data received as an HTTP response using the Audio API of the web browser. The input is the audio data sent from the server, and the output is the played audio.

[0238] (Application Example 2)

[0239] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0240] Traditional knowledge management systems often suffer from usability and poor information comprehension because they provide information without considering the user's emotions. Furthermore, systems for factory workers, in particular, require immediate responses in emergencies and high-stress situations, necessitating emotionally sensitive support. Therefore, there is a growing need for systems that recognize users' emotions in real time and provide optimal responses accordingly.

[0241] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0242] In this invention, the server includes means for collecting knowledge data from a database, means for feeding the collected knowledge data to a machine learning model and training it, means for generating answers to user questions using the trained model, means for converting the generated answers into voice data using speech synthesis software, means for providing the voice data to the user, means including an emotion engine that recognizes the user's emotions, and means for adjusting the voice tone based on emotion information from the emotion engine. This enables the rapid and appropriate provision of information that takes the user's emotions into consideration.

[0243] A "database" is an information storage system for systematically storing and managing knowledge data.

[0244] "Knowledge data" refers to a collection of information and knowledge used in a knowledge management system.

[0245] A "machine learning model" is an algorithm that learns from large amounts of data and uses that learning to make predictions and classifications for new inputs.

[0246] "Training methods" refer to the process of providing data to a machine learning model to allow it to learn and improve its performance.

[0247] "Generated answers" refer to information that a machine learning model generates based on the user's questions.

[0248] "Speech synthesis software" is a program that converts text data into speech data.

[0249] An "emotion engine" is a technology that recognizes a user's emotions in real time.

[0250] "Voice tone" refers to the characteristics of sound in audio data that express emotions and intentions.

[0251] A "server" is a computer system that collects and processes knowledge data and provides services to users.

[0252] A "user" is an individual or organization that uses a knowledge management system to retrieve information.

[0253] The system for carrying out this invention consists of multiple components. The main components include a database, a machine learning model, an emotion engine, speech synthesis software, and a server. The functions and roles of these components are described below.

[0254] The server periodically collects knowledge data from internal and external knowledge databases. The collected data is converted into a common format (e.g., JSON) and stored in the database. This knowledge data stored in the database serves as the basis for generating answers to user questions.

[0255] The collected knowledge data is used to train a machine learning model. The training process extracts patterns and rules from the large amount of knowledge data, enabling the model to generate appropriate answers to user questions. Because this training process requires significant computing resources, cloud computing environments are often used.

[0256] When a user enters a question, the device sends it to the server. The server inputs the question into a generative AI model to generate the best answer. At this time, the emotion engine analyzes the user's input and recognizes their emotions. The emotion engine uses a transformer model (e.g., DistilBERT) to determine the user's emotions. The recognized emotion information is fed back to the generative AI model, which generates an answer that takes the user's emotions into consideration.

[0257] The generated responses are sent to text-to-speech software, where the text data is converted into audio data. At this stage, the voice tone is adjusted based on emotional information obtained from the emotion engine, generating audio corresponding to each emotion. Google Cloud Text-to-Speech API is often used for this text-to-speech process. The generated audio data is sent from the server to the device, which then plays it back to provide the user with the response.

[0258] As a concrete example, consider a scenario where a factory worker asks, "What should I do if a machine stops?" The user inputs this question via a smartphone or smart glasses and sends it. The device sends this query to a server, which then sends it to a generative AI model. The model generates an appropriate response from its trained knowledge data, and the server sends this response to speech synthesis software to generate audio data. In this process, an emotion engine recognizes emotions such as stress and urgency from the user's input, and generates a response in a corresponding voice tone. The device plays the audio data received from the server, providing the factory worker with an appropriate response via voice.

[0259] The system of this invention realizes efficient and intuitive knowledge management and information delivery that takes user emotions into consideration. Users can receive timely and appropriate information and take appropriate action, especially in highly urgent situations.

[0260] Example of a prompt:

[0261] User question: "What should I do if the machine stops working?"

[0262] Generated response: "Try restarting the machine, and if it still doesn't work, contact maintenance."

[0263] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0264] Step 1:

[0265] The server collects knowledge data from internal and external knowledge databases. Specifically, it retrieves data via APIs, converts it to JSON format, and stores it in the database.

[0266] Input: Knowledge data

[0267] Data processing: Retrieve data using the API and convert it to JSON format.

[0268] Output: Knowledge data stored in the database

[0269] Step 2:

[0270] The server feeds the collected knowledge data into a machine learning model for training. During the training process, the model learns patterns and rules in the knowledge data, improving its accuracy.

[0271] Input: Knowledge data stored in the database

[0272] Data processing: Learning using machine learning algorithms

[0273] Output: Trained machine learning model

[0274] Step 3:

[0275] The user enters the question. Specifically, they enter the question in text format via a smartphone or smart glasses, and the device sends it to the server.

[0276] Input: User's question (text format)

[0277] Data processing: None

[0278] Output: User questions sent to the server

[0279] Step 4:

[0280] The server inputs the questions received from the user into a generating AI model to produce the optimal answer. In this process, the emotion engine analyzes the user's input and recognizes their emotions.

[0281] Input: User's question

[0282] Data operation: Sentiment analysis by the sentiment engine, answer generation by the generative AI model

[0283] Output: Generated answer and user sentiment information

[0284] Step 5:

[0285] Based on the sentiment information obtained from the sentiment engine, to adjust the voice tone of the generated answer, the server uses voice synthesis software to convert the text data into voice data.

[0286] Input: Generated answer, user sentiment information

[0287] Data operation: Voice synthesis (conversion from text to voice)

[0288] Output: Voice data with adjusted voice tone

[0289] Step 6:

[0290] The server transmits the generated voice data to the terminal. The terminal plays the received voice data to provide an answer to the user.

[0291] Input: Voice data

[0292] Data processing: None

[0293] Output: Voice data provided to the user

[0294] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires voice indicating user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0295] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0296] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0297] [Second Embodiment]

[0298] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0299] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0300] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0301] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0302] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0303] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0304] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0305] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0306] The specific processing program 56 is an example of the "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by operating as the specific processing unit 290 according to the specific processing program 56 executed by the processor 28 on the RAM 30.

[0307] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the specific processing unit 290.

[0308] In the smart glasses 214, the processor 46 performs the reception and output processing. The storage 50 stores a reception and output program 60. The processor 46 reads the reception and output program 60 from the storage 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is realized by operating as the control unit 46A according to the reception and output program 60 executed by the processor 46 on the RAM 48.

[0309] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".

[0310] The system of the present invention is implemented in the following procedure.

[0311] First, the collection of knowledge data is performed. In this collection process, the server periodically acquires necessary information from the in-company knowledge database and external data sources. For example, the knowledge data is acquired via an API, converted into a common format (e.g., JSON), and stored in the database. Thereby, the consistency and availability of the data are ensured.

[0312] Next, the process of generating answers to user questions is executed using a trained generative AI model. The server trains the machine learning model using knowledge data collected in advance as training data. Through this training process, the model learns patterns and rules in the knowledge data and gains the ability to generate appropriate answers to user questions.

[0313] When a user enters a question, the device sends it to a server. The server inputs the question into a generative AI and generates an appropriate answer. The generated answer is sent to speech synthesis software, where the text data is converted into speech data. This is done, for example, using the Google Text-to-Speech API.

[0314] The generated audio data is sent from the server to the terminal, which then plays it back to provide the user with an answer. This entire process allows the user to quickly and intuitively obtain useful information from the knowledge data.

[0315] Furthermore, the system saves usage history as logs and performs effectiveness analysis. Success stories are selected and documented as specific examples. These success stories are packaged as proposal materials for implementation to other companies. This packaged material is used as a guideline when other companies implement similar systems.

[0316] As a concrete example, consider a scenario where a project manager at a certain company asks about the procedure for launching a new project. The user enters the question, "How do I create a new project?" into the terminal. The terminal sends this query to the server, which then sends it to a generative AI model. The model generates the appropriate procedure from its trained knowledge data, and the server sends this answer to speech synthesis software, which generates audio data. The terminal plays the audio data received from the server, providing the project manager with the appropriate creation procedure in audio format.

[0317] In this way, the system of the present invention realizes efficient and intuitive knowledge management and information provision to users.

[0318] The following describes the processing flow.

[0319] Step 1:

[0320] Server: Accesses internal knowledge databases and external data sources, and collects knowledge data using APIs. Specifically, it retrieves data using HTTP requests and converts the data into JSON format.

[0321] Step 2:

[0322] Server: Stores the collected knowledge data in a database. For example, it uses a NoSQL database such as MongoDB to store data in JSON format.

[0323] Step 3:

[0324] Server: Initializes generative AI models for machine learning. Sets up the training environment and retrieves knowledge data from the database in batches.

[0325] Step 4:

[0326] Server: Trains a generative AI model using the acquired knowledge data. This process uses pairs of questions and their corresponding answers.

[0327] Step 5:

[0328] User: Enter your question from your device and click the submit button. For example, enter "How do I create a new project?"

[0329] Step 6:

[0330] Terminal: Sends the entered question to the server. Specifically, it passes the question data to the server using an HTTP POST request.

[0331] Step 7:

[0332] Server: Inputs questions received from users into a generative AI model and generates the optimal answer.

[0333] Step 8:

[0334] Server: Retrieves the generated response and sends it to the text-to-speech software. For example, it sends a request to convert text into speech data using the Google Text-to-Speech API.

[0335] Step 9:

[0336] Server: Receives the speech data returned from the speech synthesis software and caches or temporarily stores it as needed.

[0337] Step 10:

[0338] Server: Sends audio data to the terminal. Specifically, it returns the audio data to the terminal as an HTTP response.

[0339] Step 11:

[0340] Terminal: Plays back the received audio data and provides the user with an audio response. The user can listen to the audio and obtain the answer to the question.

[0341] Step 12:

[0342] Server: Use history is saved as logs, and effectiveness analysis is performed. Specifically, log data stored in the database is analyzed, and successful cases are identified.

[0343] Step 13:

[0344] Server: Document and package success stories to create proposal materials for implementation at other companies. Specifically, compile them into presentation materials and digital booklets.

[0345] Through this series of processing steps, the system of the present invention achieves efficient and intuitive knowledge management and information provision to users.

[0346] (Example 1)

[0347] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0348] Traditional knowledge database systems had problems with quickly obtaining the information users needed. Furthermore, text-based responses alone were difficult to understand intuitively, potentially degrading the user experience. In addition, the lack of usage history analysis and effective feedback made system improvement difficult.

[0349] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0350] In this invention, the server includes means for collecting knowledge data from a database, means for feeding the collected knowledge data to a machine learning model and training it, means for receiving questions entered by the user through a terminal and sending them to the server, means for generating answers to the user's questions using the trained generative AI model, means for converting the generated answers into audio data using speech synthesis software, means for sending the generated audio data to the terminal and providing the audio data to the user, and means for saving the system usage history as a log and performing effectiveness analysis. This enables the user to efficiently and intuitively obtain useful information from the knowledge data.

[0351] A "database" is a computer system that systematically manages a collection of data and allows for efficient searching and updating.

[0352] "Knowledge data" refers to data that systematically organizes the knowledge and information held by a company or organization.

[0353] "Means of collection" refers to methods and technologies for gathering and acquiring specific information and storing it in a database or similar system.

[0354] A "machine learning model" is an algorithm or mathematical model that learns patterns from data and uses them to make predictions and classifications about unseen data.

[0355] "Training methods" refer to the methods and techniques used to execute the process by which a machine learning model learns from knowledge data.

[0356] A "generative AI model" is a trained model that uses artificial intelligence technology to generate appropriate outputs for specific inputs.

[0357] A "user question" is a question that a user enters through their device to find information or to solve a problem they want to resolve.

[0358] A "trained generative AI model" is an AI model that has been trained using knowledge data through a machine learning algorithm.

[0359] "Means for generating answers" refers to methods and technologies for generating appropriate information in response to a user's question and creating an answer.

[0360] "Speech synthesis software" is software that takes text data as input and generates speech that sounds like a human voice.

[0361] "Means of converting to audio data" refers to methods and technologies for converting generated text-formatted responses into audio data.

[0362] "Means of providing audio data to users" refers to methods and technologies for presenting and allowing users to listen to the generated audio data.

[0363] "Usage history" refers to a record of how the system was used, including logs of user questions and responses.

[0364] "Effectiveness analysis" is the process of analyzing system performance and user satisfaction based on collected usage history to identify areas for improvement.

[0365] The system of this invention mainly consists of three components: a server, a terminal, and a user. The roles of each component and specific operating procedures are described below.

[0366] First, the server collects knowledge data from a database. This collection is done via APIs from the company's internal knowledge database or external data sources. The acquired data is converted to JSON format and stored in a database such as MongoDB. The collected knowledge data is then trained using machine learning models (e.g., BERT or GPT-3). Python libraries such as TensorFlow and PyTorch are used for this training process. The trained generative AI model is then used to generate answers to user questions.

[0367] The user enters a question through the device's interface (e.g., a web browser). For example, they might enter a question like, "How do I create a new project?" into a text box. The device converts the user's question into JSON format and sends it to the server using an HTTP POST request. The server decodes the received question and inputs it into a generating AI model. The AI ​​model generates an appropriate answer based on a comprehensive knowledge database.

[0368] The generated responses are converted into audio data using text-to-speech software (e.g., Google Text-to-Speech API). The text-to-speech software takes text data as input and generates an audio file (e.g., MP3 format). The generated audio data is sent from the server to the terminal, which plays it using its built-in audio player to provide the user with the response.

[0369] Furthermore, the server stores system usage history as logs and performs effectiveness analysis. For example, log data is stored and analyzed using tools such as Elasticsearch. This allows for evaluation of system performance and user satisfaction, and identification of areas for improvement. Success stories are compiled into regular reports and used as proposals for implementation to other companies.

[0370] As a concrete example, let's consider a scenario where a project manager at a certain company asks about the procedure for launching a new project. In this case, the user enters the question "How do I create a new project?" into the terminal. The terminal sends this query to the server, which then sends it to a generative AI model. The model generates the appropriate procedure from its trained knowledge data, and the server sends that answer to speech synthesis software to generate audio data. The terminal plays the audio data received from the server, providing the project manager with the appropriate creation procedure in audio format.

[0371] Other examples of prompt statements include:

[0372] "Please tell me about best practices for project management."

[0373] "What are the key points when drafting contracts with clients?"

[0374] In this way, the system of the present invention realizes efficient and intuitive knowledge management and information provision to users.

[0375] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0376] Step 1: Collecting Knowledge Data

[0377] The server collects knowledge data from databases and external data sources. For example, the server uses HTTP requests to retrieve knowledge data within the company. The retrieved data is converted to JSON format and stored in a database such as MongoDB. The input is raw data retrieved from APIs and data sources, and the output is knowledge data in JSON format.

[0378] Step 2: Training the machine learning model

[0379] The server trains a generative AI model using the collected knowledge data. Specifically, it trains the model using Python's TensorFlow or PyTorch libraries. Data preprocessing (cleaning, normalization, etc.) is also performed during training. The input is knowledge data in JSON format and an initial machine learning model, and the output is a trained generative AI model.

[0380] Step 3: Enter user questions

[0381] The user enters the question through the terminal's interface. For example, they might type "How do I create a new project?" into a text box in a web browser. The input is the text-based question entered by the user, and the output is the question data in JSON format generated by the terminal.

[0382] Step 4: Sending the question to the server

[0383] The terminal converts the user's input into JSON format and sends it to the server using an HTTP POST request. The input is the JSON formatted question data generated by the terminal, and the output is the question request sent to the server.

[0384] Step 5: Generating answers using a generative AI model

[0385] The server decodes the received question and inputs it into a generative AI model. The AI ​​model generates an appropriate answer to the question based on the trained knowledge data. The input consists of the question data sent to the server and the trained generative AI model, and the output is the generated answer in text format.

[0386] Step 6: Speech synthesis processing

[0387] The server sends the generated text-based response to speech synthesis software, which converts it into audio data. Specifically, it calls the Google Text-to-Speech API to generate an audio file (e.g., in MP3 format). The input is the generated text-based response, and the output is the generated audio data.

[0388] Step 7: Providing the response to the user

[0389] The server sends the generated audio data to the terminal as an HTTP response. The terminal receives this and plays it using its built-in audio player. The input is the audio data sent from the server, and the output is the audio response presented to the user.

[0390] Step 8: Log usage history and analyze its effectiveness.

[0391] The server stores system usage history as logs and performs periodic effectiveness analyses. For example, it uses tools such as Elasticsearch to analyze log data and evaluate system performance and user satisfaction. The input is the usage history logs, and the output is analysis results and improvement suggestions.

[0392] (Application Example 1)

[0393] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0394] Traditional knowledge management systems typically rely on text-based interfaces for users to obtain answers to their questions, making intuitive and rapid information provision difficult. Furthermore, there is a lack of tools that allow store staff to instantly access knowledge data and provide accurate answers during customer service in physical stores. Therefore, improving the efficiency of customer service in physical stores and enhancing customer satisfaction are key challenges.

[0395] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0396] In this invention, the server includes means for collecting knowledge data from a database, means for feeding the collected knowledge data to a machine learning model and training it, means for generating answers to user questions using the trained model, means for converting the generated answers into voice data using speech synthesis software, means for providing the voice data while the user is wearing smart glasses, and means for running the speech synthesis software on the smart glasses when providing the voice data. This enables store clerks in physical stores to quickly and accurately provide answers to customer questions through smart glasses, thereby improving the efficiency of customer service and increasing customer satisfaction.

[0397] A "database" is an information system used to centrally store and manage knowledge data collected from both inside and outside a company.

[0398] "Knowledge data" refers to knowledge and information collected from data sources both inside and outside a company, and is used in question-answering systems.

[0399] A "machine learning model" is an algorithm that is trained on collected knowledge data to generate appropriate answers to user questions.

[0400] "Training" is the process of inputting collected knowledge data into a machine learning model and allowing that model to learn.

[0401] "Generated answers" refer to the information that a machine learning model outputs in response to a user's question.

[0402] "Speech synthesis software" is a program that converts text data into speech data.

[0403] "Smart glasses" are wearable devices that users wear and that provide information via displays and audio.

[0404] A "user" is a person who uses this system to ask questions and receive answers.

[0405] A "server" is a central computer that collects, processes, and stores data, and is responsible for providing various services.

[0406] As an embodiment of this invention, a system is constructed using the following configuration and procedure. The system mainly consists of a server, a terminal (smart glasses), a knowledge database, a machine learning model, and speech synthesis software.

[0407] First, the server periodically collects knowledge data from data sources both inside and outside the company. During this collection process, knowledge data is retrieved via APIs, converted into a common format (e.g., JSON), and stored in a knowledge database. This ensures data consistency and availability.

[0408] Next, the server uses the collected knowledge data to train a machine learning model. Through this training process, the machine learning model learns patterns and rules in the knowledge data and gains the ability to generate appropriate answers to user questions.

[0409] When a user wears smart glasses and enters a question, the device sends the question to a server. The server inputs the question into a generative AI model and generates an appropriate answer. The generated answer is sent to speech synthesis software, where the text data is converted into speech data. Specifically, the Google Text-to-Speech API can be used.

[0410] The generated audio data is sent from the server to the smart glasses, which then play it back to provide the user with an answer. This entire process allows the user to quickly and intuitively obtain useful information from the knowledge data.

[0411] Furthermore, the system saves usage history as logs and performs effectiveness analysis. Success stories are selected and documented as specific examples. These success stories are packaged as proposal materials for implementation in other companies. This packaged material is used as a guideline when other companies implement similar systems.

[0412] As a concrete example, consider a scenario where a store clerk in a physical store is wearing smart glasses while performing their duties. If a customer asks, "How is the sizing of these shoes?", an application built into the clerk's smart glasses recognizes the question and transmits it to a server. A generative AI model generates a response such as, "These shoes run a little smaller than usual, so we recommend going up one size," and provides this as voice data to the clerk's smart glasses. The clerk can then provide the customer with an appropriate answer through this voice data.

[0413] Examples of prompt statements include:

[0414] "Knowledge Data API Endpoint: https: / / example.com / api / knowledge"

[0415] "AI Model API Endpoint: https: / / example.com / api / ai_model"

[0416] Question: How is the sizing of these shoes?

[0417] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0418] Step 1:

[0419] The server periodically collects knowledge data. During this process, it retrieves necessary information from internal and external data sources via APIs, converts this data into a common format (such as JSON), and stores it in the knowledge database. The input is raw data obtained from data sources, and the output is knowledge data converted into the common format.

[0420] Step 2:

[0421] The server feeds the collected knowledge data into a machine learning model for training. Here, the input is the knowledge data stored in the knowledge database, and the output is the trained machine learning model. Each data point is input into the model, and an iterative training process is performed to teach the model patterns and rules in the knowledge data.

[0422] Step 3:

[0423] When a user wears smart glasses and inputs a question from a customer, the terminal (smart glasses) recognizes the question and sends it to the server. Here, the input is the question in the form of voice or text, and the output is the transfer of the question data to the server. This includes the specific operation of using the smart glasses' microphone and voice recognition function to convert the question into text.

[0424] Step 4:

[0425] The server inputs the received question into a generative AI model and generates an appropriate answer. Here, the input is the question text sent by the user, and the output is the generated answer text. The generative AI model executes an algorithm that generates the best possible answer to the question based on its pre-trained knowledge.

[0426] Step 5:

[0427] The server sends the generated response text to text-to-speech software, which converts it into audio data. Here, the input is the generated response text, and the output is the audio data. Specifically, the text is converted into an audio file using APIs such as the Google Text-to-Speech API.

[0428] Step 6:

[0429] The server sends the generated audio data to the smart glasses, and the device plays it back and provides it to the user. Here, the input is the audio data, and the output is the audio played back through the smart glasses. This includes the specific operation of playing the audio using the smart glasses' speaker.

[0430] Step 7:

[0431] The system saves the entire usage history as a log and performs effectiveness analysis. Here, the input is the usage history data, and the output is the analysis results. Through log saving and analysis, success stories are selected and implementation proposal materials are created for other companies.

[0432] This series of processing steps enables store staff in physical stores to provide quick and accurate answers to customer questions through smart glasses, thereby improving the efficiency of customer service and increasing customer satisfaction.

[0433] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0434] The system of the present invention includes means for collecting knowledge data from a database, feeding it to a machine learning model for training, and generating answers to user questions. It also includes means for converting the generated answers into audio data using speech synthesis software and providing the audio data to the user. By combining this system with an emotion engine, it is possible to recognize the user's emotions and provide optimal answers based on them.

[0435] First, let me explain how this system works. The server periodically collects knowledge data from the company's internal knowledge database and external data sources. For example, it retrieves data via an API, converts it to a common format (e.g., JSON), and stores it in the database.

[0436] Next, a process is executed to generate answers to user questions using a trained generative AI model. The server trains the machine learning model using pre-collected knowledge data. Through this training process, the model learns patterns and rules in the knowledge data and gains the ability to generate appropriate answers to user questions.

[0437] When a user enters a question, the device sends it to the server. The server inputs the question into a generative AI model and generates the optimal answer. At this time, the emotion engine analyzes the user's input and recognizes their emotions. The recognized emotion information is fed back to the generative AI model, which then generates the optimal answer.

[0438] The generated responses are sent to text-to-speech software, where the text data is converted into speech data. This is done, for example, using the Google Text-to-Speech API. The text-to-speech software adjusts the speech tone based on the emotion information obtained from the emotion engine and generates speech corresponding to each emotion.

[0439] The generated audio data is sent from the server to the terminal, which then plays it back to provide the user with an answer. This entire process allows the user to quickly and intuitively obtain useful information from knowledge data, while also receiving emotionally sensitive responses.

[0440] As a concrete example, consider a scenario where a customer support representative at a certain company asks, "How should I handle this complaint?" The user enters this question into a terminal and clicks the send button. The terminal sends this query to a server, which then sends it to a generative AI model. The model generates an appropriate response method from its trained knowledge data, and the server sends this response to speech synthesis software to generate audio data. In this process, an emotion engine recognizes the user's emotions, such as urgency and dissatisfaction, and generates a response in a corresponding voice tone. The terminal plays the audio data received from the server, providing the customer support representative with an appropriate response method in audio.

[0441] In this way, the system of the present invention realizes efficient and intuitive knowledge management and information provision to users, as well as enabling responses that take into account the user's feelings.

[0442] The following describes the processing flow.

[0443] Step 1:

[0444] Server: Accesses internal knowledge databases and external data sources, and collects knowledge data using APIs. For example, it retrieves data using HTTP requests and converts the data into JSON format.

[0445] Step 2:

[0446] Server: Stores the collected knowledge data in a database. For example, it uses a NoSQL database such as MongoDB to store data in JSON format.

[0447] Step 3:

[0448] Server: Initializes generative AI models for machine learning. Sets up the training environment and retrieves knowledge data from the database in batches.

[0449] Step 4:

[0450] Server: Trains a generative AI model using the acquired knowledge data. This process uses pairs of questions and their corresponding answers.

[0451] Step 5:

[0452] User: Enter your question from your device and click the submit button. For example, enter "How do I create a new project?"

[0453] Step 6:

[0454] Terminal: Sends the entered question to the server. Specifically, it passes the question data to the server using an HTTP POST request.

[0455] Step 7:

[0456] Server: Inputs questions received from users into a generative AI model to generate the optimal answer. During this process, the emotion engine analyzes the user's input and recognizes their emotions. For example, natural language processing techniques are used to extract the user's emotions from the text.

[0457] Step 8:

[0458] Server: Based on recognized emotion information, a generative AI model generates responses with a tone and content appropriate to that emotion. For example, if the user is expressing dissatisfaction, it will generate a more helpful and empathetic response.

[0459] Step 9:

[0460] Server: Retrieves the generated response and sends it to the text-to-speech software. For example, it sends a request to convert text into speech data using the Google Text-to-Speech API.

[0461] Step 10:

[0462] Server: Receives the voice data returned from the speech synthesis software and caches or temporarily stores it as needed. At this time, the speech synthesis software adjusts the voice tone based on emotional information.

[0463] Step 11:

[0464] Server: Sends audio data to the terminal. Specifically, it returns the audio data to the terminal as an HTTP response.

[0465] Step 12:

[0466] Terminal: Plays back received audio data and provides the user with an audio response. The user can listen to the audio and obtain the answer to their question. For example, the voice tone may be friendly and empathetic.

[0467] Step 13:

[0468] Server: Use history is saved as logs, and effectiveness analysis is performed. Specifically, log data stored in the database is analyzed, and successful cases are identified.

[0469] Step 14:

[0470] Server: Document and package success stories to create proposal materials for implementation at other companies. Specifically, compile them into presentation materials and digital booklets.

[0471] This series of processing steps enables the system of the present invention not only to achieve efficient and intuitive knowledge management and information provision to users, but also to respond in a way that takes users' emotions into consideration.

[0472] (Example 2)

[0473] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0474] Conventional knowledge management systems have struggled to provide quick and appropriate answers to user questions. Furthermore, they have been unable to provide answers that take user emotions into consideration, making it difficult to improve user satisfaction. This invention aims to solve these problems and provide a superior user experience.

[0475] In Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for collecting knowledge data from a database, means for feeding the collected knowledge data to a machine learning model and training it, means for generating answers to user questions using the trained model, means for converting the generated answers into voice data using speech synthesis software, means for analyzing the user's emotions, means for generating the optimal answer based on the analyzed emotional information, and means for providing the voice data to the user. This enables quick and appropriate answers to user questions, and further improves user satisfaction by providing answers that take the user's emotions into consideration.

[0476] A "database" is a software system for systematically organizing, storing, searching, updating, and managing information.

[0477] "Knowledge data" refers to data that companies and organizations have accumulated and stored in digital format, encompassing experience, knowledge, and information.

[0478] A "machine learning model" is a collection of algorithms that learn patterns and rules based on data and use them to make predictions and classifications on new data.

[0479] "Training" is the process of using knowledge data to teach a machine learning model patterns and rules.

[0480] A "generative AI model" is a model trained using machine learning that has the ability to generate answers to user questions.

[0481] "Speech synthesis software" is software that converts text data into speech data.

[0482] An "emotion engine" is a technology that analyzes user emotions from their input and actions and generates information based on those emotions.

[0483] An "API" is an interface for exchanging functions and data between different software programs.

[0484] A "data source" refers to an external system or service that provides data.

[0485] The "JSON format" is a text-based data format for representing data in a lightweight and highly readable format.

[0486] The system of the present invention includes means for collecting knowledge data from a database, feeding it to a machine learning model for training, and generating answers to user questions. It also includes means for converting the generated answers into audio data using speech synthesis software and providing the audio data to the user. Furthermore, this system incorporates an emotion engine that can recognize the user's emotions and provide the most appropriate answers based on those emotions.

[0487] First, the server collects knowledge data from the company's internal knowledge database and external data sources. Specifically, it retrieves data from external data sources via APIs, converts it to JSON format, and stores it in the database. Next, the server uses the collected knowledge data to train a generative AI model. This training process uses machine learning frameworks such as TensorFlow.

[0488] When a user enters a question, the device sends it to the server. The server inputs the received question into a generative AI model to generate the best possible answer. In this process, an emotion engine analyzes the user's input and recognizes their emotions. The recognized emotion information is fed back to the generative AI model, which then generates the best possible answer.

[0489] The generated response is sent to text-to-speech software (e.g., Google Text-to-Speech API), where the text data is converted into audio data. During this speech synthesis process, emotional information obtained from an emotion engine is used to adjust the voice tone. The generated audio data is then sent from the server to the device, which plays it back to provide the response to the user.

[0490] As a concrete example, consider a scenario where a customer support representative at a certain company asks, "How do I handle complaints?" The user enters this question into a device and clicks the send button. The device sends the query to a server, which passes it to a generative AI model. The generative AI model generates an appropriate answer from its trained knowledge data, and the server converts that answer into audio data using the Google Text-to-Speech API. In this process, an emotion engine recognizes the user's emotions, such as urgency and frustration, and generates an answer with an appropriate tone of voice. The device plays the audio data received from the server and provides the answer to the customer support representative.

[0491] Example of a prompt:

[0492] "Create a customer support system that generates responses to user inquiries about how to handle complaints. Please explain the process for outputting responses in an appropriate tone of voice, taking into account the user's emotional state and potential anxiety."

[0493] In this way, the system of the present invention realizes efficient and intuitive knowledge management and information provision to users, and further enables responses that take into account the user's feelings.

[0494] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0495] Step 1: Data Collection

[0496] The server collects knowledge data from the company's internal knowledge database and external data sources. Specifically, the server uses APIs to retrieve data from external data sources. This data is converted to JSON format and stored in the database. The input is raw data retrieved from the API, and the output is knowledge data converted to JSON format.

[0497] Step 2: Model Training

[0498] The server uses the collected knowledge data to train a generative AI model. Specifically, the server loads knowledge data from a database and performs data preprocessing. For training, it uses a machine learning framework such as TensorFlow, and optimizes the model by repeating multiple epochs. The input is knowledge data in JSON format, and the output is a trained generative AI model.

[0499] Step 3: Receiving questions from users

[0500] The user enters a question via the device. For example, they might enter "How should I handle complaints?" through a web browser's input form. When the user presses the submit button, the device sends the question to the server as an HTTP POST request. The input is the user's question text, and the output is the question data in HTTP request format.

[0501] Step 4: Analyzing the Question

[0502] The server inputs the received question into a generative AI model and generates an answer. Specifically, the server first parses the question text and inputs it into the generative AI model. During this process, the emotion engine analyzes the user's input and recognizes their emotions. The input is question data in HTTP request format, and the output is answer text with added emotion information.

[0503] Step 5: Generating the answer

[0504] The generative AI model generates the optimal answer based on the input question and sentiment information. The generated answer is temporarily stored in the server's memory. The input is question data including sentiment information, and the output is the generated answer text.

[0505] Step 6: Generate audio data

[0506] The server sends the generated response to speech synthesis software, which converts it into speech data. Specifically, the server uses the Google Text-to-Speech API to convert the text data into speech data. In this process, emotional information obtained from the emotion engine is reflected in the tone of voice. The input is the response text, and the output is emotion-adjusted speech data.

[0507] Step 7: Providing audio data

[0508] The generated audio data is sent from the server to the terminal, which then plays it back to provide a response to the user. Specifically, the terminal plays the audio data received as an HTTP response using the Audio API of the web browser. The input is the audio data sent from the server, and the output is the played audio.

[0509] (Application Example 2)

[0510] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0511] Traditional knowledge management systems often suffer from usability and poor information comprehension because they provide information without considering the user's emotions. Furthermore, systems for factory workers, in particular, require immediate responses in emergencies and high-stress situations, necessitating emotionally sensitive support. Therefore, there is a growing need for systems that recognize users' emotions in real time and provide optimal responses accordingly.

[0512] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0513] In this invention, the server includes means for collecting knowledge data from a database, means for feeding the collected knowledge data to a machine learning model and training it, means for generating answers to user questions using the trained model, means for converting the generated answers into voice data using speech synthesis software, means for providing the voice data to the user, means including an emotion engine that recognizes the user's emotions, and means for adjusting the voice tone based on emotion information from the emotion engine. This enables the rapid and appropriate provision of information that takes the user's emotions into consideration.

[0514] A "database" is an information storage system for systematically storing and managing knowledge data.

[0515] "Knowledge data" refers to a collection of information and knowledge used in a knowledge management system.

[0516] A "machine learning model" is an algorithm that learns from large amounts of data and uses that learning to make predictions and classifications for new inputs.

[0517] "Training methods" refer to the process of providing data to a machine learning model to allow it to learn and improve its performance.

[0518] "Generated answers" refer to information that a machine learning model generates based on the user's questions.

[0519] "Speech synthesis software" is a program that converts text data into speech data.

[0520] An "emotion engine" is a technology that recognizes a user's emotions in real time.

[0521] "Voice tone" refers to the characteristics of sound in audio data that express emotions and intentions.

[0522] A "server" is a computer system that collects and processes knowledge data and provides services to users.

[0523] A "user" is an individual or organization that uses a knowledge management system to retrieve information.

[0524] The system for carrying out this invention consists of multiple components. The main components include a database, a machine learning model, an emotion engine, speech synthesis software, and a server. The functions and roles of these components are described below.

[0525] The server periodically collects knowledge data from internal and external knowledge databases. The collected data is converted into a common format (e.g., JSON) and stored in the database. This knowledge data stored in the database serves as the basis for generating answers to user questions.

[0526] The collected knowledge data is used to train a machine learning model. The training process extracts patterns and rules from the large amount of knowledge data, enabling the model to generate appropriate answers to user questions. Because this training process requires significant computing resources, cloud computing environments are often used.

[0527] When a user enters a question, the device sends it to the server. The server inputs the question into a generative AI model to generate the best answer. At this time, the emotion engine analyzes the user's input and recognizes their emotions. The emotion engine uses a transformer model (e.g., DistilBERT) to determine the user's emotions. The recognized emotion information is fed back to the generative AI model, which generates an answer that takes the user's emotions into consideration.

[0528] The generated responses are sent to text-to-speech software, where the text data is converted into audio data. At this stage, the voice tone is adjusted based on emotional information obtained from the emotion engine, generating audio corresponding to each emotion. Google Cloud Text-to-Speech API is often used for this text-to-speech process. The generated audio data is sent from the server to the device, which then plays it back to provide the user with the response.

[0529] As a concrete example, consider a scenario where a factory worker asks, "What should I do if a machine stops?" The user inputs this question via a smartphone or smart glasses and sends it. The device sends this query to a server, which then sends it to a generative AI model. The model generates an appropriate response from its trained knowledge data, and the server sends this response to speech synthesis software to generate audio data. In this process, an emotion engine recognizes emotions such as stress and urgency from the user's input, and generates a response in a corresponding voice tone. The device plays the audio data received from the server, providing the factory worker with an appropriate response via voice.

[0530] The system of this invention realizes efficient and intuitive knowledge management and information delivery that takes user emotions into consideration. Users can receive timely and appropriate information and take appropriate action, especially in highly urgent situations.

[0531] Example of a prompt:

[0532] User question: "What should I do if the machine stops working?"

[0533] Generated response: "Try restarting the machine, and if it still doesn't work, contact maintenance."

[0534] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0535] Step 1:

[0536] The server collects knowledge data from internal and external knowledge databases. Specifically, it retrieves data via APIs, converts it to JSON format, and stores it in the database.

[0537] Input: Knowledge data

[0538] Data processing: Retrieve data using the API and convert it to JSON format.

[0539] Output: Knowledge data stored in the database

[0540] Step 2:

[0541] The server feeds the collected knowledge data into a machine learning model for training. During the training process, the model learns patterns and rules in the knowledge data, improving its accuracy.

[0542] Input: Knowledge data stored in the database

[0543] Data processing: Learning using machine learning algorithms

[0544] Output: Trained machine learning model

[0545] Step 3:

[0546] The user enters the question. Specifically, they enter the question in text format via a smartphone or smart glasses, and the device sends it to the server.

[0547] Input: User's question (text format)

[0548] Data processing: None

[0549] Output: User questions sent to the server

[0550] Step 4:

[0551] The server inputs the questions received from the user into a generating AI model to produce the optimal answer. In this process, the emotion engine analyzes the user's input and recognizes their emotions.

[0552] Input: User's question

[0553] Data processing: Emotion analysis using an emotion engine, response generation using a generative AI model.

[0554] Output: Generated responses and user sentiment information

[0555] Step 5:

[0556] Based on the emotional information obtained from the emotion engine, the server uses speech synthesis software to convert text data into speech data in order to adjust the voice tone of the generated response.

[0557] Input: Generated response, user sentiment information

[0558] Data processing: Speech synthesis (conversion from text to speech)

[0559] Output: Audio data of the adjusted voice tone

[0560] Step 6:

[0561] The server sends the generated audio data to the terminal. The terminal plays the received audio data and provides the user with a response.

[0562] Input: Audio data

[0563] Data processing: None

[0564] Output: Audio data provided to the user

[0565] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0566] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0567] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0568] [Third Embodiment]

[0569] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0570] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0571] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0572] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0573] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0574] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0575] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0576] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0577] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0578] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0579] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0580] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0581] The system of the present invention is implemented according to the following procedure.

[0582] First, knowledge data is collected. During this collection process, the server periodically retrieves necessary information from the company's internal knowledge database and external data sources. For example, knowledge data can be retrieved via an API, converted into a common format (e.g., JSON), and stored in the database. This ensures data consistency and availability.

[0583] Next, the process of generating answers to user questions is executed using a trained generative AI model. The server trains the machine learning model using knowledge data collected in advance as training data. Through this training process, the model learns patterns and rules in the knowledge data and gains the ability to generate appropriate answers to user questions.

[0584] When a user enters a question, the device sends it to a server. The server inputs the question into a generative AI and generates an appropriate answer. The generated answer is sent to speech synthesis software, where the text data is converted into speech data. This is done, for example, using the Google Text-to-Speech API.

[0585] The generated audio data is sent from the server to the terminal, which then plays it back to provide the user with an answer. This entire process allows the user to quickly and intuitively obtain useful information from the knowledge data.

[0586] Furthermore, the system saves usage history as logs and performs effectiveness analysis. Success stories are selected and documented as specific examples. These success stories are packaged as proposal materials for implementation to other companies. This packaged material is used as a guideline when other companies implement similar systems.

[0587] As a concrete example, consider a scenario where a project manager at a certain company asks about the procedure for launching a new project. The user enters the question, "How do I create a new project?" into the terminal. The terminal sends this query to the server, which then sends it to a generative AI model. The model generates the appropriate procedure from its trained knowledge data, and the server sends this answer to speech synthesis software, which generates audio data. The terminal plays the audio data received from the server, providing the project manager with the appropriate creation procedure in audio format.

[0588] In this way, the system of the present invention realizes efficient and intuitive knowledge management and information provision to users.

[0589] The following describes the processing flow.

[0590] Step 1:

[0591] Server: Accesses internal knowledge databases and external data sources, and collects knowledge data using APIs. Specifically, it retrieves data using HTTP requests and converts the data into JSON format.

[0592] Step 2:

[0593] Server: Stores the collected knowledge data in a database. For example, it uses a NoSQL database such as MongoDB to store data in JSON format.

[0594] Step 3:

[0595] Server: Initializes generative AI models for machine learning. Sets up the training environment and retrieves knowledge data from the database in batches.

[0596] Step 4:

[0597] Server: Trains a generative AI model using the acquired knowledge data. This process uses pairs of questions and their corresponding answers.

[0598] Step 5:

[0599] User: Enter your question from your device and click the submit button. For example, enter "How do I create a new project?"

[0600] Step 6:

[0601] Terminal: Sends the entered question to the server. Specifically, it passes the question data to the server using an HTTP POST request.

[0602] Step 7:

[0603] Server: Inputs questions received from users into a generative AI model and generates the optimal answer.

[0604] Step 8:

[0605] Server: Retrieves the generated response and sends it to the text-to-speech software. For example, it sends a request to convert text into speech data using the Google Text-to-Speech API.

[0606] Step 9:

[0607] Server: Receives the speech data returned from the speech synthesis software and caches or temporarily stores it as needed.

[0608] Step 10:

[0609] Server: Sends audio data to the terminal. Specifically, it returns the audio data to the terminal as an HTTP response.

[0610] Step 11:

[0611] Terminal: Plays back the received audio data and provides the user with an audio response. The user can listen to the audio and obtain the answer to the question.

[0612] Step 12:

[0613] Server: Use history is saved as logs, and effectiveness analysis is performed. Specifically, log data stored in the database is analyzed, and successful cases are identified.

[0614] Step 13:

[0615] Server: Document and package success stories to create proposal materials for implementation at other companies. Specifically, compile them into presentation materials and digital booklets.

[0616] Through this series of processing steps, the system of the present invention achieves efficient and intuitive knowledge management and information provision to users.

[0617] (Example 1)

[0618] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0619] Traditional knowledge database systems had problems with quickly obtaining the information users needed. Furthermore, text-based responses alone were difficult to understand intuitively, potentially degrading the user experience. In addition, the lack of usage history analysis and effective feedback made system improvement difficult.

[0620] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0621] In this invention, the server includes means for collecting knowledge data from a database, means for feeding the collected knowledge data to a machine learning model and training it, means for receiving questions entered by the user through a terminal and sending them to the server, means for generating answers to the user's questions using the trained generative AI model, means for converting the generated answers into audio data using speech synthesis software, means for sending the generated audio data to the terminal and providing the audio data to the user, and means for saving the system usage history as a log and performing effectiveness analysis. This enables the user to efficiently and intuitively obtain useful information from the knowledge data.

[0622] A "database" is a computer system that systematically manages a collection of data and allows for efficient searching and updating.

[0623] "Knowledge data" refers to data that systematically organizes the knowledge and information held by a company or organization.

[0624] "Means of collection" refers to methods and technologies for gathering and acquiring specific information and storing it in a database or similar system.

[0625] A "machine learning model" is an algorithm or mathematical model that learns patterns from data and uses them to make predictions and classifications about unseen data.

[0626] "Training methods" refer to the methods and techniques used to execute the process by which a machine learning model learns from knowledge data.

[0627] A "generative AI model" is a trained model that uses artificial intelligence technology to generate appropriate outputs for specific inputs.

[0628] A "user question" is a question that a user enters through their device to find information or to solve a problem they want to resolve.

[0629] A "trained generative AI model" is an AI model that has been trained using knowledge data through a machine learning algorithm.

[0630] "Means for generating answers" refers to methods and technologies for generating appropriate information in response to a user's question and creating an answer.

[0631] "Speech synthesis software" is software that takes text data as input and generates speech that sounds like a human voice.

[0632] "Means of converting to audio data" refers to methods and technologies for converting generated text-formatted responses into audio data.

[0633] "Means of providing audio data to users" refers to methods and technologies for presenting and allowing users to listen to the generated audio data.

[0634] "Usage history" refers to a record of how the system was used, including logs of user questions and responses.

[0635] "Effectiveness analysis" is the process of analyzing system performance and user satisfaction based on collected usage history to identify areas for improvement.

[0636] The system of this invention mainly consists of three components: a server, a terminal, and a user. The roles of each component and specific operating procedures are described below.

[0637] First, the server collects knowledge data from a database. This collection is done via APIs from the company's internal knowledge database or external data sources. The acquired data is converted to JSON format and stored in a database such as MongoDB. The collected knowledge data is then trained using machine learning models (e.g., BERT or GPT-3). Python libraries such as TensorFlow and PyTorch are used for this training process. The trained generative AI model is then used to generate answers to user questions.

[0638] The user enters a question through the device's interface (e.g., a web browser). For example, they might enter a question like, "How do I create a new project?" into a text box. The device converts the user's question into JSON format and sends it to the server using an HTTP POST request. The server decodes the received question and inputs it into a generating AI model. The AI ​​model generates an appropriate answer based on a comprehensive knowledge database.

[0639] The generated responses are converted into audio data using text-to-speech software (e.g., Google Text-to-Speech API). The text-to-speech software takes text data as input and generates an audio file (e.g., MP3 format). The generated audio data is sent from the server to the terminal, which plays it using its built-in audio player to provide the user with the response.

[0640] Furthermore, the server stores system usage history as logs and performs effectiveness analysis. For example, log data is stored and analyzed using tools such as Elasticsearch. This allows for evaluation of system performance and user satisfaction, and identification of areas for improvement. Success stories are compiled into regular reports and used as proposals for implementation to other companies.

[0641] As a concrete example, let's consider a scenario where a project manager at a certain company asks about the procedure for launching a new project. In this case, the user enters the question "How do I create a new project?" into the terminal. The terminal sends this query to the server, which then sends it to a generative AI model. The model generates the appropriate procedure from its trained knowledge data, and the server sends that answer to speech synthesis software to generate audio data. The terminal plays the audio data received from the server, providing the project manager with the appropriate creation procedure in audio format.

[0642] Other examples of prompt statements include:

[0643] "Please tell me about best practices for project management."

[0644] "What are the key points when drafting contracts with clients?"

[0645] In this way, the system of the present invention realizes efficient and intuitive knowledge management and information provision to users.

[0646] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0647] Step 1: Collecting Knowledge Data

[0648] The server collects knowledge data from databases and external data sources. For example, the server uses HTTP requests to retrieve knowledge data within the company. The retrieved data is converted to JSON format and stored in a database such as MongoDB. The input is raw data retrieved from APIs and data sources, and the output is knowledge data in JSON format.

[0649] Step 2: Training the machine learning model

[0650] The server trains a generative AI model using the collected knowledge data. Specifically, it trains the model using Python's TensorFlow or PyTorch libraries. Data preprocessing (cleaning, normalization, etc.) is also performed during training. The input is knowledge data in JSON format and an initial machine learning model, and the output is a trained generative AI model.

[0651] Step 3: Enter user questions

[0652] The user enters the question through the terminal's interface. For example, they might type "How do I create a new project?" into a text box in a web browser. The input is the text-based question entered by the user, and the output is the question data in JSON format generated by the terminal.

[0653] Step 4: Sending the question to the server

[0654] The terminal converts the user's input into JSON format and sends it to the server using an HTTP POST request. The input is the JSON formatted question data generated by the terminal, and the output is the question request sent to the server.

[0655] Step 5: Generating answers using a generative AI model

[0656] The server decodes the received question and inputs it into a generative AI model. The AI ​​model generates an appropriate answer to the question based on the trained knowledge data. The input consists of the question data sent to the server and the trained generative AI model, and the output is the generated answer in text format.

[0657] Step 6: Speech synthesis processing

[0658] The server sends the generated text-based response to speech synthesis software, which converts it into audio data. Specifically, it calls the Google Text-to-Speech API to generate an audio file (e.g., in MP3 format). The input is the generated text-based response, and the output is the generated audio data.

[0659] Step 7: Providing the response to the user

[0660] The server sends the generated audio data to the terminal as an HTTP response. The terminal receives this and plays it using its built-in audio player. The input is the audio data sent from the server, and the output is the audio response presented to the user.

[0661] Step 8: Log usage history and analyze its effectiveness.

[0662] The server stores system usage history as logs and performs periodic effectiveness analyses. For example, it uses tools such as Elasticsearch to analyze log data and evaluate system performance and user satisfaction. The input is the usage history logs, and the output is analysis results and improvement suggestions.

[0663] (Application Example 1)

[0664] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0665] Traditional knowledge management systems typically rely on text-based interfaces for users to obtain answers to their questions, making intuitive and rapid information provision difficult. Furthermore, there is a lack of tools that allow store staff to instantly access knowledge data and provide accurate answers during customer service in physical stores. Therefore, improving the efficiency of customer service in physical stores and enhancing customer satisfaction are key challenges.

[0666] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0667] In this invention, the server includes means for collecting knowledge data from a database, means for feeding the collected knowledge data to a machine learning model and training it, means for generating answers to user questions using the trained model, means for converting the generated answers into voice data using speech synthesis software, means for providing the voice data while the user is wearing smart glasses, and means for running the speech synthesis software on the smart glasses when providing the voice data. This enables store clerks in physical stores to quickly and accurately provide answers to customer questions through smart glasses, thereby improving the efficiency of customer service and increasing customer satisfaction.

[0668] A "database" is an information system used to centrally store and manage knowledge data collected from both inside and outside a company.

[0669] "Knowledge data" refers to knowledge and information collected from data sources both inside and outside a company, and is used in question-answering systems.

[0670] A "machine learning model" is an algorithm that is trained on collected knowledge data to generate appropriate answers to user questions.

[0671] "Training" is the process of inputting collected knowledge data into a machine learning model and allowing that model to learn.

[0672] "Generated answers" refer to the information that a machine learning model outputs in response to a user's question.

[0673] "Speech synthesis software" is a program that converts text data into speech data.

[0674] "Smart glasses" are wearable devices that users wear and that provide information via displays and audio.

[0675] A "user" is a person who uses this system to ask questions and receive answers.

[0676] A "server" is a central computer that collects, processes, and stores data, and is responsible for providing various services.

[0677] As an embodiment of this invention, a system is constructed using the following configuration and procedure. The system mainly consists of a server, a terminal (smart glasses), a knowledge database, a machine learning model, and speech synthesis software.

[0678] First, the server periodically collects knowledge data from data sources both inside and outside the company. During this collection process, knowledge data is retrieved via APIs, converted into a common format (e.g., JSON), and stored in a knowledge database. This ensures data consistency and availability.

[0679] Next, the server uses the collected knowledge data to train a machine learning model. Through this training process, the machine learning model learns patterns and rules in the knowledge data and gains the ability to generate appropriate answers to user questions.

[0680] When a user wears smart glasses and enters a question, the device sends the question to a server. The server inputs the question into a generative AI model and generates an appropriate answer. The generated answer is sent to speech synthesis software, where the text data is converted into speech data. Specifically, the Google Text-to-Speech API can be used.

[0681] The generated audio data is sent from the server to the smart glasses, which then play it back to provide the user with an answer. This entire process allows the user to quickly and intuitively obtain useful information from the knowledge data.

[0682] Furthermore, the system saves usage history as logs and performs effectiveness analysis. Success stories are selected and documented as specific examples. These success stories are packaged as proposal materials for implementation in other companies. This packaged material is used as a guideline when other companies implement similar systems.

[0683] As a concrete example, consider a scenario where a store clerk in a physical store is wearing smart glasses while performing their duties. If a customer asks, "How is the sizing of these shoes?", an application built into the clerk's smart glasses recognizes the question and transmits it to a server. A generative AI model generates a response such as, "These shoes run a little smaller than usual, so we recommend going up one size," and provides this as voice data to the clerk's smart glasses. The clerk can then provide the customer with an appropriate answer through this voice data.

[0684] Examples of prompt statements include:

[0685] "Knowledge Data API Endpoint: https: / / example.com / api / knowledge"

[0686] "AI Model API Endpoint: https: / / example.com / api / ai_model"

[0687] Question: How is the sizing of these shoes?

[0688] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0689] Step 1:

[0690] The server periodically collects knowledge data. During this process, it retrieves necessary information from internal and external data sources via APIs, converts this data into a common format (such as JSON), and stores it in the knowledge database. The input is raw data obtained from data sources, and the output is knowledge data converted into the common format.

[0691] Step 2:

[0692] The server feeds the collected knowledge data into a machine learning model for training. Here, the input is the knowledge data stored in the knowledge database, and the output is the trained machine learning model. Each data point is input into the model, and an iterative training process is performed to teach the model patterns and rules in the knowledge data.

[0693] Step 3:

[0694] When a user wears smart glasses and inputs a question from a customer, the terminal (smart glasses) recognizes the question and sends it to the server. Here, the input is the question in the form of voice or text, and the output is the transfer of the question data to the server. This includes the specific operation of using the smart glasses' microphone and voice recognition function to convert the question into text.

[0695] Step 4:

[0696] The server inputs the received question into a generative AI model and generates an appropriate answer. Here, the input is the question text sent by the user, and the output is the generated answer text. The generative AI model executes an algorithm that generates the best possible answer to the question based on its pre-trained knowledge.

[0697] Step 5:

[0698] The server sends the generated response text to text-to-speech software, which converts it into audio data. Here, the input is the generated response text, and the output is the audio data. Specifically, the text is converted into an audio file using APIs such as the Google Text-to-Speech API.

[0699] Step 6:

[0700] The server sends the generated audio data to the smart glasses, and the device plays it back and provides it to the user. Here, the input is the audio data, and the output is the audio played back through the smart glasses. This includes the specific operation of playing the audio using the smart glasses' speaker.

[0701] Step 7:

[0702] The system saves the entire usage history as a log and performs effectiveness analysis. Here, the input is the usage history data, and the output is the analysis results. Through log saving and analysis, success stories are selected and implementation proposal materials are created for other companies.

[0703] This series of processing steps enables store staff in physical stores to provide quick and accurate answers to customer questions through smart glasses, thereby improving the efficiency of customer service and increasing customer satisfaction.

[0704] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0705] The system of the present invention includes means for collecting knowledge data from a database, feeding it to a machine learning model for training, and generating answers to user questions. It also includes means for converting the generated answers into audio data using speech synthesis software and providing the audio data to the user. By combining this system with an emotion engine, it is possible to recognize the user's emotions and provide optimal answers based on them.

[0706] First, let me explain how this system works. The server periodically collects knowledge data from the company's internal knowledge database and external data sources. For example, it retrieves data via an API, converts it to a common format (e.g., JSON), and stores it in the database.

[0707] Next, a process is executed to generate answers to user questions using a trained generative AI model. The server trains the machine learning model using pre-collected knowledge data. Through this training process, the model learns patterns and rules in the knowledge data and gains the ability to generate appropriate answers to user questions.

[0708] When a user enters a question, the device sends it to the server. The server inputs the question into a generative AI model and generates the optimal answer. At this time, the emotion engine analyzes the user's input and recognizes their emotions. The recognized emotion information is fed back to the generative AI model, which then generates the optimal answer.

[0709] The generated responses are sent to text-to-speech software, where the text data is converted into speech data. This is done, for example, using the Google Text-to-Speech API. The text-to-speech software adjusts the speech tone based on the emotion information obtained from the emotion engine and generates speech corresponding to each emotion.

[0710] The generated audio data is sent from the server to the terminal, which then plays it back to provide the user with an answer. This entire process allows the user to quickly and intuitively obtain useful information from knowledge data, while also receiving emotionally sensitive responses.

[0711] As a concrete example, consider a scenario where a customer support representative at a certain company asks, "How should I handle this complaint?" The user enters this question into a terminal and clicks the send button. The terminal sends this query to a server, which then sends it to a generative AI model. The model generates an appropriate response method from its trained knowledge data, and the server sends this response to speech synthesis software to generate audio data. In this process, an emotion engine recognizes the user's emotions, such as urgency and dissatisfaction, and generates a response in a corresponding voice tone. The terminal plays the audio data received from the server, providing the customer support representative with an appropriate response method in audio.

[0712] In this way, the system of the present invention realizes efficient and intuitive knowledge management and information provision to users, as well as enabling responses that take into account the user's feelings.

[0713] The following describes the processing flow.

[0714] Step 1:

[0715] Server: Accesses internal knowledge databases and external data sources, and collects knowledge data using APIs. For example, it retrieves data using HTTP requests and converts the data into JSON format.

[0716] Step 2:

[0717] Server: Stores the collected knowledge data in a database. For example, it uses a NoSQL database such as MongoDB to store data in JSON format.

[0718] Step 3:

[0719] Server: Initializes generative AI models for machine learning. Sets up the training environment and retrieves knowledge data from the database in batches.

[0720] Step 4:

[0721] Server: Trains a generative AI model using the acquired knowledge data. This process uses pairs of questions and their corresponding answers.

[0722] Step 5:

[0723] User: Enter your question from your device and click the submit button. For example, enter "How do I create a new project?"

[0724] Step 6:

[0725] Terminal: Sends the entered question to the server. Specifically, it passes the question data to the server using an HTTP POST request.

[0726] Step 7:

[0727] Server: Inputs questions received from users into a generative AI model to generate the optimal answer. During this process, the emotion engine analyzes the user's input and recognizes their emotions. For example, natural language processing techniques are used to extract the user's emotions from the text.

[0728] Step 8:

[0729] Server: Based on recognized emotion information, a generative AI model generates responses with a tone and content appropriate to that emotion. For example, if the user is expressing dissatisfaction, it will generate a more helpful and empathetic response.

[0730] Step 9:

[0731] Server: Retrieves the generated response and sends it to the text-to-speech software. For example, it sends a request to convert text into speech data using the Google Text-to-Speech API.

[0732] Step 10:

[0733] Server: Receives the voice data returned from the speech synthesis software and caches or temporarily stores it as needed. At this time, the speech synthesis software adjusts the voice tone based on emotional information.

[0734] Step 11:

[0735] Server: Sends audio data to the terminal. Specifically, it returns the audio data to the terminal as an HTTP response.

[0736] Step 12:

[0737] Terminal: Plays back received audio data and provides the user with an audio response. The user can listen to the audio and obtain the answer to their question. For example, the voice tone may be friendly and empathetic.

[0738] Step 13:

[0739] Server: Use history is saved as logs, and effectiveness analysis is performed. Specifically, log data stored in the database is analyzed, and successful cases are identified.

[0740] Step 14:

[0741] Server: Document and package success stories to create proposal materials for implementation at other companies. Specifically, compile them into presentation materials and digital booklets.

[0742] This series of processing steps enables the system of the present invention not only to achieve efficient and intuitive knowledge management and information provision to users, but also to respond in a way that takes users' emotions into consideration.

[0743] (Example 2)

[0744] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0745] Conventional knowledge management systems have struggled to provide quick and appropriate answers to user questions. Furthermore, they have been unable to provide answers that take user emotions into consideration, making it difficult to improve user satisfaction. This invention aims to solve these problems and provide a superior user experience.

[0746] In Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for collecting knowledge data from a database, means for feeding the collected knowledge data to a machine learning model and training it, means for generating answers to user questions using the trained model, means for converting the generated answers into voice data using speech synthesis software, means for analyzing the user's emotions, means for generating the optimal answer based on the analyzed emotional information, and means for providing the voice data to the user. This enables quick and appropriate answers to user questions, and further improves user satisfaction by providing answers that take the user's emotions into consideration.

[0747] A "database" is a software system for systematically organizing, storing, searching, updating, and managing information.

[0748] "Knowledge data" refers to data that companies and organizations have accumulated and stored in digital format, encompassing experience, knowledge, and information.

[0749] A "machine learning model" is a collection of algorithms that learn patterns and rules based on data and use them to make predictions and classifications on new data.

[0750] "Training" is the process of using knowledge data to teach a machine learning model patterns and rules.

[0751] A "generative AI model" is a model trained using machine learning that has the ability to generate answers to user questions.

[0752] "Speech synthesis software" is software that converts text data into speech data.

[0753] An "emotion engine" is a technology that analyzes user emotions from their input and actions and generates information based on those emotions.

[0754] An "API" is an interface for exchanging functions and data between different software programs.

[0755] A "data source" refers to an external system or service that provides data.

[0756] The "JSON format" is a text-based data format for representing data in a lightweight and highly readable format.

[0757] The system of the present invention includes means for collecting knowledge data from a database, feeding it to a machine learning model for training, and generating answers to user questions. It also includes means for converting the generated answers into audio data using speech synthesis software and providing the audio data to the user. Furthermore, this system incorporates an emotion engine that can recognize the user's emotions and provide the most appropriate answers based on those emotions.

[0758] First, the server collects knowledge data from the company's internal knowledge database and external data sources. Specifically, it retrieves data from external data sources via APIs, converts it to JSON format, and stores it in the database. Next, the server uses the collected knowledge data to train a generative AI model. This training process uses machine learning frameworks such as TensorFlow.

[0759] When a user enters a question, the device sends it to the server. The server inputs the received question into a generative AI model to generate the best possible answer. In this process, an emotion engine analyzes the user's input and recognizes their emotions. The recognized emotion information is fed back to the generative AI model, which then generates the best possible answer.

[0760] The generated response is sent to text-to-speech software (e.g., Google Text-to-Speech API), where the text data is converted into audio data. During this speech synthesis process, emotional information obtained from an emotion engine is used to adjust the voice tone. The generated audio data is then sent from the server to the device, which plays it back to provide the response to the user.

[0761] As a concrete example, consider a scenario where a customer support representative at a certain company asks, "How do I handle complaints?" The user enters this question into a device and clicks the send button. The device sends the query to a server, which passes it to a generative AI model. The generative AI model generates an appropriate answer from its trained knowledge data, and the server converts that answer into audio data using the Google Text-to-Speech API. In this process, an emotion engine recognizes the user's emotions, such as urgency and frustration, and generates an answer with an appropriate tone of voice. The device plays the audio data received from the server and provides the answer to the customer support representative.

[0762] Example of a prompt:

[0763] "Create a customer support system that generates responses to user inquiries about how to handle complaints. Please explain the process for outputting responses in an appropriate tone of voice, taking into account the user's emotional state and potential anxiety."

[0764] In this way, the system of the present invention realizes efficient and intuitive knowledge management and information provision to users, and further enables responses that take into account the user's feelings.

[0765] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0766] Step 1: Data Collection

[0767] The server collects knowledge data from the company's internal knowledge database and external data sources. Specifically, the server uses APIs to retrieve data from external data sources. This data is converted to JSON format and stored in the database. The input is raw data retrieved from the API, and the output is knowledge data converted to JSON format.

[0768] Step 2: Model Training

[0769] The server uses the collected knowledge data to train a generative AI model. Specifically, the server loads knowledge data from a database and performs data preprocessing. For training, it uses a machine learning framework such as TensorFlow, and optimizes the model by repeating multiple epochs. The input is knowledge data in JSON format, and the output is a trained generative AI model.

[0770] Step 3: Receiving questions from users

[0771] The user enters a question via the device. For example, they might enter "How should I handle complaints?" through a web browser's input form. When the user presses the submit button, the device sends the question to the server as an HTTP POST request. The input is the user's question text, and the output is the question data in HTTP request format.

[0772] Step 4: Analyzing the Question

[0773] The server inputs the received question into a generative AI model and generates an answer. Specifically, the server first parses the question text and inputs it into the generative AI model. During this process, the emotion engine analyzes the user's input and recognizes their emotions. The input is question data in HTTP request format, and the output is answer text with added emotion information.

[0774] Step 5: Generating the answer

[0775] The generative AI model generates the optimal answer based on the input question and sentiment information. The generated answer is temporarily stored in the server's memory. The input is question data including sentiment information, and the output is the generated answer text.

[0776] Step 6: Generate audio data

[0777] The server sends the generated response to speech synthesis software, which converts it into speech data. Specifically, the server uses the Google Text-to-Speech API to convert the text data into speech data. In this process, emotional information obtained from the emotion engine is reflected in the tone of voice. The input is the response text, and the output is emotion-adjusted speech data.

[0778] Step 7: Providing audio data

[0779] The generated audio data is sent from the server to the terminal, which then plays it back to provide a response to the user. Specifically, the terminal plays the audio data received as an HTTP response using the Audio API of the web browser. The input is the audio data sent from the server, and the output is the played audio.

[0780] (Application Example 2)

[0781] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0782] Traditional knowledge management systems often suffer from usability and poor information comprehension because they provide information without considering the user's emotions. Furthermore, systems for factory workers, in particular, require immediate responses in emergencies and high-stress situations, necessitating emotionally sensitive support. Therefore, there is a growing need for systems that recognize users' emotions in real time and provide optimal responses accordingly.

[0783] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0784] In this invention, the server includes means for collecting knowledge data from a database, means for feeding the collected knowledge data to a machine learning model and training it, means for generating answers to user questions using the trained model, means for converting the generated answers into voice data using speech synthesis software, means for providing the voice data to the user, means including an emotion engine that recognizes the user's emotions, and means for adjusting the voice tone based on emotion information from the emotion engine. This enables the rapid and appropriate provision of information that takes the user's emotions into consideration.

[0785] A "database" is an information storage system for systematically storing and managing knowledge data.

[0786] "Knowledge data" refers to a collection of information and knowledge used in a knowledge management system.

[0787] A "machine learning model" is an algorithm that learns from large amounts of data and uses that learning to make predictions and classifications for new inputs.

[0788] "Training methods" refer to the process of providing data to a machine learning model to allow it to learn and improve its performance.

[0789] "Generated answers" refer to information that a machine learning model generates based on the user's questions.

[0790] "Speech synthesis software" is a program that converts text data into speech data.

[0791] An "emotion engine" is a technology that recognizes a user's emotions in real time.

[0792] "Voice tone" refers to the characteristics of sound in audio data that express emotions and intentions.

[0793] A "server" is a computer system that collects and processes knowledge data and provides services to users.

[0794] A "user" is an individual or organization that uses a knowledge management system to retrieve information.

[0795] The system for carrying out this invention consists of multiple components. The main components include a database, a machine learning model, an emotion engine, speech synthesis software, and a server. The functions and roles of these components are described below.

[0796] The server periodically collects knowledge data from internal and external knowledge databases. The collected data is converted into a common format (e.g., JSON) and stored in the database. This knowledge data stored in the database serves as the basis for generating answers to user questions.

[0797] The collected knowledge data is used to train a machine learning model. The training process extracts patterns and rules from the large amount of knowledge data, enabling the model to generate appropriate answers to user questions. Because this training process requires significant computing resources, cloud computing environments are often used.

[0798] When a user enters a question, the device sends it to the server. The server inputs the question into a generative AI model to generate the best answer. At this time, the emotion engine analyzes the user's input and recognizes their emotions. The emotion engine uses a transformer model (e.g., DistilBERT) to determine the user's emotions. The recognized emotion information is fed back to the generative AI model, which generates an answer that takes the user's emotions into consideration.

[0799] The generated responses are sent to text-to-speech software, where the text data is converted into audio data. At this stage, the voice tone is adjusted based on emotional information obtained from the emotion engine, generating audio corresponding to each emotion. Google Cloud Text-to-Speech API is often used for this text-to-speech process. The generated audio data is sent from the server to the device, which then plays it back to provide the user with the response.

[0800] As a concrete example, consider a scenario where a factory worker asks, "What should I do if a machine stops?" The user inputs this question via a smartphone or smart glasses and sends it. The device sends this query to a server, which then sends it to a generative AI model. The model generates an appropriate response from its trained knowledge data, and the server sends this response to speech synthesis software to generate audio data. In this process, an emotion engine recognizes emotions such as stress and urgency from the user's input, and generates a response in a corresponding voice tone. The device plays the audio data received from the server, providing the factory worker with an appropriate response via voice.

[0801] The system of this invention realizes efficient and intuitive knowledge management and information delivery that takes user emotions into consideration. Users can receive timely and appropriate information and take appropriate action, especially in highly urgent situations.

[0802] Example of a prompt:

[0803] User question: "What should I do if the machine stops working?"

[0804] Generated response: "Try restarting the machine, and if it still doesn't work, contact maintenance."

[0805] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0806] Step 1:

[0807] The server collects knowledge data from internal and external knowledge databases. Specifically, it retrieves data via APIs, converts it to JSON format, and stores it in the database.

[0808] Input: Knowledge data

[0809] Data processing: Retrieve data using the API and convert it to JSON format.

[0810] Output: Knowledge data stored in the database

[0811] Step 2:

[0812] The server feeds the collected knowledge data into a machine learning model for training. During the training process, the model learns patterns and rules in the knowledge data, improving its accuracy.

[0813] Input: Knowledge data stored in the database

[0814] Data processing: Learning using machine learning algorithms

[0815] Output: Trained machine learning model

[0816] Step 3:

[0817] The user enters the question. Specifically, they enter the question in text format via a smartphone or smart glasses, and the device sends it to the server.

[0818] Input: User's question (text format)

[0819] Data processing: None

[0820] Output: User questions sent to the server

[0821] Step 4:

[0822] The server inputs the questions received from the user into a generating AI model to produce the optimal answer. In this process, the emotion engine analyzes the user's input and recognizes their emotions.

[0823] Input: User's question

[0824] Data processing: Emotion analysis using an emotion engine, response generation using a generative AI model.

[0825] Output: Generated responses and user sentiment information

[0826] Step 5:

[0827] Based on the emotional information obtained from the emotion engine, the server uses speech synthesis software to convert text data into speech data in order to adjust the voice tone of the generated response.

[0828] Input: Generated response, user sentiment information

[0829] Data processing: Speech synthesis (conversion from text to speech)

[0830] Output: Audio data of the adjusted voice tone

[0831] Step 6:

[0832] The server sends the generated audio data to the terminal. The terminal plays the received audio data and provides the user with a response.

[0833] Input: Audio data

[0834] Data processing: None

[0835] Output: Audio data provided to the user

[0836] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0837] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0838] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0839] [Fourth Embodiment]

[0840] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0841] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0842] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0843] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0844] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0845] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0846] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0847] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0848] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0849] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0850] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0851] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0852] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0853] The system of the present invention is implemented according to the following procedure.

[0854] First, knowledge data is collected. During this collection process, the server periodically retrieves necessary information from the company's internal knowledge database and external data sources. For example, knowledge data can be retrieved via an API, converted into a common format (e.g., JSON), and stored in the database. This ensures data consistency and availability.

[0855] Next, the process of generating answers to user questions is executed using a trained generative AI model. The server trains the machine learning model using knowledge data collected in advance as training data. Through this training process, the model learns patterns and rules in the knowledge data and gains the ability to generate appropriate answers to user questions.

[0856] When a user enters a question, the device sends it to a server. The server inputs the question into a generative AI and generates an appropriate answer. The generated answer is sent to speech synthesis software, where the text data is converted into speech data. This is done, for example, using the Google Text-to-Speech API.

[0857] The generated audio data is sent from the server to the terminal, which then plays it back to provide the user with an answer. This entire process allows the user to quickly and intuitively obtain useful information from the knowledge data.

[0858] Furthermore, the system saves usage history as logs and performs effectiveness analysis. Success stories are selected and documented as specific examples. These success stories are packaged as proposal materials for implementation to other companies. This packaged material is used as a guideline when other companies implement similar systems.

[0859] As a concrete example, consider a scenario where a project manager at a certain company asks about the procedure for launching a new project. The user enters the question, "How do I create a new project?" into the terminal. The terminal sends this query to the server, which then sends it to a generative AI model. The model generates the appropriate procedure from its trained knowledge data, and the server sends this answer to speech synthesis software, which generates audio data. The terminal plays the audio data received from the server, providing the project manager with the appropriate creation procedure in audio format.

[0860] In this way, the system of the present invention realizes efficient and intuitive knowledge management and information provision to users.

[0861] The following describes the processing flow.

[0862] Step 1:

[0863] Server: Accesses internal knowledge databases and external data sources, and collects knowledge data using APIs. Specifically, it retrieves data using HTTP requests and converts the data into JSON format.

[0864] Step 2:

[0865] Server: Stores the collected knowledge data in a database. For example, it uses a NoSQL database such as MongoDB to store data in JSON format.

[0866] Step 3:

[0867] Server: Initializes generative AI models for machine learning. Sets up the training environment and retrieves knowledge data from the database in batches.

[0868] Step 4:

[0869] Server: Trains a generative AI model using the acquired knowledge data. This process uses pairs of questions and their corresponding answers.

[0870] Step 5:

[0871] User: Enter your question from your device and click the submit button. For example, enter "How do I create a new project?"

[0872] Step 6:

[0873] Terminal: Sends the entered question to the server. Specifically, it passes the question data to the server using an HTTP POST request.

[0874] Step 7:

[0875] Server: Inputs questions received from users into a generative AI model and generates the optimal answer.

[0876] Step 8:

[0877] Server: Retrieves the generated response and sends it to the text-to-speech software. For example, it sends a request to convert text into speech data using the Google Text-to-Speech API.

[0878] Step 9:

[0879] Server: Receives the speech data returned from the speech synthesis software and caches or temporarily stores it as needed.

[0880] Step 10:

[0881] Server: Sends audio data to the terminal. Specifically, it returns the audio data to the terminal as an HTTP response.

[0882] Step 11:

[0883] Terminal: Plays back the received audio data and provides the user with an audio response. The user can listen to the audio and obtain the answer to the question.

[0884] Step 12:

[0885] Server: Use history is saved as logs, and effectiveness analysis is performed. Specifically, log data stored in the database is analyzed, and successful cases are identified.

[0886] Step 13:

[0887] Server: Document and package success stories to create proposal materials for implementation at other companies. Specifically, compile them into presentation materials and digital booklets.

[0888] Through this series of processing steps, the system of the present invention achieves efficient and intuitive knowledge management and information provision to users.

[0889] (Example 1)

[0890] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0891] Traditional knowledge database systems had problems with quickly obtaining the information users needed. Furthermore, text-based responses alone were difficult to understand intuitively, potentially degrading the user experience. In addition, the lack of usage history analysis and effective feedback made system improvement difficult.

[0892] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0893] In this invention, the server includes means for collecting knowledge data from a database, means for feeding the collected knowledge data to a machine learning model and training it, means for receiving questions entered by the user through a terminal and sending them to the server, means for generating answers to the user's questions using the trained generative AI model, means for converting the generated answers into audio data using speech synthesis software, means for sending the generated audio data to the terminal and providing the audio data to the user, and means for saving the system usage history as a log and performing effectiveness analysis. This enables the user to efficiently and intuitively obtain useful information from the knowledge data.

[0894] A "database" is a computer system that systematically manages a collection of data and allows for efficient searching and updating.

[0895] "Knowledge data" refers to data that systematically organizes the knowledge and information held by a company or organization.

[0896] "Means of collection" refers to methods and technologies for gathering and acquiring specific information and storing it in a database or similar system.

[0897] A "machine learning model" is an algorithm or mathematical model that learns patterns from data and uses them to make predictions and classifications about unseen data.

[0898] "Training methods" refer to the methods and techniques used to execute the process by which a machine learning model learns from knowledge data.

[0899] A "generative AI model" is a trained model that uses artificial intelligence technology to generate appropriate outputs for specific inputs.

[0900] A "user question" is a question that a user enters through their device to find information or to solve a problem they want to resolve.

[0901] A "trained generative AI model" is an AI model that has been trained using knowledge data through a machine learning algorithm.

[0902] "Means for generating answers" refers to methods and technologies for generating appropriate information in response to a user's question and creating an answer.

[0903] "Speech synthesis software" is software that takes text data as input and generates speech that sounds like a human voice.

[0904] "Means of converting to audio data" refers to methods and technologies for converting generated text-formatted responses into audio data.

[0905] "Means of providing audio data to users" refers to methods and technologies for presenting and allowing users to listen to the generated audio data.

[0906] "Usage history" refers to a record of how the system was used, including logs of user questions and responses.

[0907] "Effectiveness analysis" is the process of analyzing system performance and user satisfaction based on collected usage history to identify areas for improvement.

[0908] The system of this invention mainly consists of three components: a server, a terminal, and a user. The roles of each component and specific operating procedures are described below.

[0909] First, the server collects knowledge data from a database. This collection is done via APIs from the company's internal knowledge database or external data sources. The acquired data is converted to JSON format and stored in a database such as MongoDB. The collected knowledge data is then trained using machine learning models (e.g., BERT or GPT-3). Python libraries such as TensorFlow and PyTorch are used for this training process. The trained generative AI model is then used to generate answers to user questions.

[0910] The user enters a question through the device's interface (e.g., a web browser). For example, they might enter a question like, "How do I create a new project?" into a text box. The device converts the user's question into JSON format and sends it to the server using an HTTP POST request. The server decodes the received question and inputs it into a generating AI model. The AI ​​model generates an appropriate answer based on a comprehensive knowledge database.

[0911] The generated responses are converted into audio data using text-to-speech software (e.g., Google Text-to-Speech API). The text-to-speech software takes text data as input and generates an audio file (e.g., MP3 format). The generated audio data is sent from the server to the terminal, which plays it using its built-in audio player to provide the user with the response.

[0912] Furthermore, the server stores system usage history as logs and performs effectiveness analysis. For example, log data is stored and analyzed using tools such as Elasticsearch. This allows for evaluation of system performance and user satisfaction, and identification of areas for improvement. Success stories are compiled into regular reports and used as proposals for implementation to other companies.

[0913] As a concrete example, let's consider a scenario where a project manager at a certain company asks about the procedure for launching a new project. In this case, the user enters the question "How do I create a new project?" into the terminal. The terminal sends this query to the server, which then sends it to a generative AI model. The model generates the appropriate procedure from its trained knowledge data, and the server sends that answer to speech synthesis software to generate audio data. The terminal plays the audio data received from the server, providing the project manager with the appropriate creation procedure in audio format.

[0914] Other examples of prompt statements include:

[0915] "Please tell me about best practices for project management."

[0916] "What are the key points when drafting contracts with clients?"

[0917] In this way, the system of the present invention realizes efficient and intuitive knowledge management and information provision to users.

[0918] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0919] Step 1: Collecting Knowledge Data

[0920] The server collects knowledge data from databases and external data sources. For example, the server uses HTTP requests to retrieve knowledge data within the company. The retrieved data is converted to JSON format and stored in a database such as MongoDB. The input is raw data retrieved from APIs and data sources, and the output is knowledge data in JSON format.

[0921] Step 2: Training the machine learning model

[0922] The server trains a generative AI model using the collected knowledge data. Specifically, it trains the model using Python's TensorFlow or PyTorch libraries. Data preprocessing (cleaning, normalization, etc.) is also performed during training. The input is knowledge data in JSON format and an initial machine learning model, and the output is a trained generative AI model.

[0923] Step 3: Enter user questions

[0924] The user enters the question through the terminal's interface. For example, they might type "How do I create a new project?" into a text box in a web browser. The input is the text-based question entered by the user, and the output is the question data in JSON format generated by the terminal.

[0925] Step 4: Sending the question to the server

[0926] The terminal converts the user's input into JSON format and sends it to the server using an HTTP POST request. The input is the JSON formatted question data generated by the terminal, and the output is the question request sent to the server.

[0927] Step 5: Generating answers using a generative AI model

[0928] The server decodes the received question and inputs it into a generative AI model. The AI ​​model generates an appropriate answer to the question based on the trained knowledge data. The input consists of the question data sent to the server and the trained generative AI model, and the output is the generated answer in text format.

[0929] Step 6: Speech synthesis processing

[0930] The server sends the generated text-based response to speech synthesis software, which converts it into audio data. Specifically, it calls the Google Text-to-Speech API to generate an audio file (e.g., in MP3 format). The input is the generated text-based response, and the output is the generated audio data.

[0931] Step 7: Providing the response to the user

[0932] The server sends the generated audio data to the terminal as an HTTP response. The terminal receives this and plays it using its built-in audio player. The input is the audio data sent from the server, and the output is the audio response presented to the user.

[0933] Step 8: Log usage history and analyze its effectiveness.

[0934] The server stores system usage history as logs and performs periodic effectiveness analyses. For example, it uses tools such as Elasticsearch to analyze log data and evaluate system performance and user satisfaction. The input is the usage history logs, and the output is analysis results and improvement suggestions.

[0935] (Application Example 1)

[0936] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0937] Traditional knowledge management systems typically rely on text-based interfaces for users to obtain answers to their questions, making intuitive and rapid information provision difficult. Furthermore, there is a lack of tools that allow store staff to instantly access knowledge data and provide accurate answers during customer service in physical stores. Therefore, improving the efficiency of customer service in physical stores and enhancing customer satisfaction are key challenges.

[0938] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0939] In this invention, the server includes means for collecting knowledge data from a database, means for feeding the collected knowledge data to a machine learning model and training it, means for generating answers to user questions using the trained model, means for converting the generated answers into voice data using speech synthesis software, means for providing the voice data while the user is wearing smart glasses, and means for running the speech synthesis software on the smart glasses when providing the voice data. This enables store clerks in physical stores to quickly and accurately provide answers to customer questions through smart glasses, thereby improving the efficiency of customer service and increasing customer satisfaction.

[0940] A "database" is an information system used to centrally store and manage knowledge data collected from both inside and outside a company.

[0941] "Knowledge data" refers to knowledge and information collected from data sources both inside and outside a company, and is used in question-answering systems.

[0942] A "machine learning model" is an algorithm that is trained on collected knowledge data to generate appropriate answers to user questions.

[0943] "Training" is the process of inputting collected knowledge data into a machine learning model and allowing that model to learn.

[0944] "Generated answers" refer to the information that a machine learning model outputs in response to a user's question.

[0945] "Speech synthesis software" is a program that converts text data into speech data.

[0946] "Smart glasses" are wearable devices that users wear and that provide information via displays and audio.

[0947] A "user" is a person who uses this system to ask questions and receive answers.

[0948] A "server" is a central computer that collects, processes, and stores data, and is responsible for providing various services.

[0949] As an embodiment of this invention, a system is constructed using the following configuration and procedure. The system mainly consists of a server, a terminal (smart glasses), a knowledge database, a machine learning model, and speech synthesis software.

[0950] First, the server periodically collects knowledge data from data sources both inside and outside the company. During this collection process, knowledge data is retrieved via APIs, converted into a common format (e.g., JSON), and stored in a knowledge database. This ensures data consistency and availability.

[0951] Next, the server uses the collected knowledge data to train a machine learning model. Through this training process, the machine learning model learns patterns and rules in the knowledge data and gains the ability to generate appropriate answers to user questions.

[0952] When a user wears smart glasses and enters a question, the device sends the question to a server. The server inputs the question into a generative AI model and generates an appropriate answer. The generated answer is sent to speech synthesis software, where the text data is converted into speech data. Specifically, the Google Text-to-Speech API can be used.

[0953] The generated audio data is sent from the server to the smart glasses, which then play it back to provide the user with an answer. This entire process allows the user to quickly and intuitively obtain useful information from the knowledge data.

[0954] Furthermore, the system saves usage history as logs and performs effectiveness analysis. Success stories are selected and documented as specific examples. These success stories are packaged as proposal materials for implementation in other companies. This packaged material is used as a guideline when other companies implement similar systems.

[0955] As a concrete example, consider a scenario where a store clerk in a physical store is wearing smart glasses while performing their duties. If a customer asks, "How is the sizing of these shoes?", an application built into the clerk's smart glasses recognizes the question and transmits it to a server. A generative AI model generates a response such as, "These shoes run a little smaller than usual, so we recommend going up one size," and provides this as voice data to the clerk's smart glasses. The clerk can then provide the customer with an appropriate answer through this voice data.

[0956] Examples of prompt statements include:

[0957] "Knowledge Data API Endpoint: https: / / example.com / api / knowledge"

[0958] "AI Model API Endpoint: https: / / example.com / api / ai_model"

[0959] Question: How is the sizing of these shoes?

[0960] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0961] Step 1:

[0962] The server periodically collects knowledge data. During this process, it retrieves necessary information from internal and external data sources via APIs, converts this data into a common format (such as JSON), and stores it in the knowledge database. The input is raw data obtained from data sources, and the output is knowledge data converted into the common format.

[0963] Step 2:

[0964] The server feeds the collected knowledge data into a machine learning model for training. Here, the input is the knowledge data stored in the knowledge database, and the output is the trained machine learning model. Each data point is input into the model, and an iterative training process is performed to teach the model patterns and rules in the knowledge data.

[0965] Step 3:

[0966] When a user wears smart glasses and inputs a question from a customer, the terminal (smart glasses) recognizes the question and sends it to the server. Here, the input is the question in the form of voice or text, and the output is the transfer of the question data to the server. This includes the specific operation of using the smart glasses' microphone and voice recognition function to convert the question into text.

[0967] Step 4:

[0968] The server inputs the received question into a generative AI model and generates an appropriate answer. Here, the input is the question text sent by the user, and the output is the generated answer text. The generative AI model executes an algorithm that generates the best possible answer to the question based on its pre-trained knowledge.

[0969] Step 5:

[0970] The server sends the generated response text to text-to-speech software, which converts it into audio data. Here, the input is the generated response text, and the output is the audio data. Specifically, the text is converted into an audio file using APIs such as the Google Text-to-Speech API.

[0971] Step 6:

[0972] The server sends the generated audio data to the smart glasses, and the device plays it back and provides it to the user. Here, the input is the audio data, and the output is the audio played back through the smart glasses. This includes the specific operation of playing the audio using the smart glasses' speaker.

[0973] Step 7:

[0974] The system saves the entire usage history as a log and performs effectiveness analysis. Here, the input is the usage history data, and the output is the analysis results. Through log saving and analysis, success stories are selected and implementation proposal materials are created for other companies.

[0975] This series of processing steps enables store staff in physical stores to provide quick and accurate answers to customer questions through smart glasses, thereby improving the efficiency of customer service and increasing customer satisfaction.

[0976] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0977] The system of the present invention includes means for collecting knowledge data from a database, feeding it to a machine learning model for training, and generating answers to user questions. It also includes means for converting the generated answers into audio data using speech synthesis software and providing the audio data to the user. By combining this system with an emotion engine, it is possible to recognize the user's emotions and provide optimal answers based on them.

[0978] First, let me explain how this system works. The server periodically collects knowledge data from the company's internal knowledge database and external data sources. For example, it retrieves data via an API, converts it to a common format (e.g., JSON), and stores it in the database.

[0979] Next, a process is executed to generate answers to user questions using a trained generative AI model. The server trains the machine learning model using pre-collected knowledge data. Through this training process, the model learns patterns and rules in the knowledge data and gains the ability to generate appropriate answers to user questions.

[0980] When a user enters a question, the device sends it to the server. The server inputs the question into a generative AI model and generates the optimal answer. At this time, the emotion engine analyzes the user's input and recognizes their emotions. The recognized emotion information is fed back to the generative AI model, which then generates the optimal answer.

[0981] The generated responses are sent to text-to-speech software, where the text data is converted into speech data. This is done, for example, using the Google Text-to-Speech API. The text-to-speech software adjusts the speech tone based on the emotion information obtained from the emotion engine and generates speech corresponding to each emotion.

[0982] The generated audio data is sent from the server to the terminal, which then plays it back to provide the user with an answer. This entire process allows the user to quickly and intuitively obtain useful information from knowledge data, while also receiving emotionally sensitive responses.

[0983] As a concrete example, consider a scenario where a customer support representative at a certain company asks, "How should I handle this complaint?" The user enters this question into a terminal and clicks the send button. The terminal sends this query to a server, which then sends it to a generative AI model. The model generates an appropriate response method from its trained knowledge data, and the server sends this response to speech synthesis software to generate audio data. In this process, an emotion engine recognizes the user's emotions, such as urgency and dissatisfaction, and generates a response in a corresponding voice tone. The terminal plays the audio data received from the server, providing the customer support representative with an appropriate response method in audio.

[0984] In this way, the system of the present invention realizes efficient and intuitive knowledge management and information provision to users, as well as enabling responses that take into account the user's feelings.

[0985] The following describes the processing flow.

[0986] Step 1:

[0987] Server: Accesses internal knowledge databases and external data sources, and collects knowledge data using APIs. For example, it retrieves data using HTTP requests and converts the data into JSON format.

[0988] Step 2:

[0989] Server: Stores the collected knowledge data in a database. For example, it uses a NoSQL database such as MongoDB to store data in JSON format.

[0990] Step 3:

[0991] Server: Initializes generative AI models for machine learning. Sets up the training environment and retrieves knowledge data from the database in batches.

[0992] Step 4:

[0993] Server: Trains a generative AI model using the acquired knowledge data. This process uses pairs of questions and their corresponding answers.

[0994] Step 5:

[0995] User: Enter your question from your device and click the submit button. For example, enter "How do I create a new project?"

[0996] Step 6:

[0997] Terminal: Sends the entered question to the server. Specifically, it passes the question data to the server using an HTTP POST request.

[0998] Step 7:

[0999] Server: Inputs questions received from users into a generative AI model to generate the optimal answer. During this process, the emotion engine analyzes the user's input and recognizes their emotions. For example, natural language processing techniques are used to extract the user's emotions from the text.

[1000] Step 8:

[1001] Server: Based on recognized emotion information, a generative AI model generates responses with a tone and content appropriate to that emotion. For example, if the user is expressing dissatisfaction, it will generate a more helpful and empathetic response.

[1002] Step 9:

[1003] Server: Retrieves the generated response and sends it to the text-to-speech software. For example, it sends a request to convert text into speech data using the Google Text-to-Speech API.

[1004] Step 10:

[1005] Server: Receives the voice data returned from the speech synthesis software and caches or temporarily stores it as needed. At this time, the speech synthesis software adjusts the voice tone based on emotional information.

[1006] Step 11:

[1007] Server: Sends audio data to the terminal. Specifically, it returns the audio data to the terminal as an HTTP response.

[1008] Step 12:

[1009] Terminal: Plays back received audio data and provides the user with an audio response. The user can listen to the audio and obtain the answer to their question. For example, the voice tone may be friendly and empathetic.

[1010] Step 13:

[1011] Server: Use history is saved as logs, and effectiveness analysis is performed. Specifically, log data stored in the database is analyzed, and successful cases are identified.

[1012] Step 14:

[1013] Server: Document and package success stories to create proposal materials for implementation at other companies. Specifically, compile them into presentation materials and digital booklets.

[1014] This series of processing steps enables the system of the present invention not only to achieve efficient and intuitive knowledge management and information provision to users, but also to respond in a way that takes users' emotions into consideration.

[1015] (Example 2)

[1016] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1017] Conventional knowledge management systems have struggled to provide quick and appropriate answers to user questions. Furthermore, they have been unable to provide answers that take user emotions into consideration, making it difficult to improve user satisfaction. This invention aims to solve these problems and provide a superior user experience.

[1018] In Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for collecting knowledge data from a database, means for feeding the collected knowledge data to a machine learning model and training it, means for generating answers to user questions using the trained model, means for converting the generated answers into voice data using speech synthesis software, means for analyzing the user's emotions, means for generating the optimal answer based on the analyzed emotional information, and means for providing the voice data to the user. This enables quick and appropriate answers to user questions, and further improves user satisfaction by providing answers that take the user's emotions into consideration.

[1019] A "database" is a software system for systematically organizing, storing, searching, updating, and managing information.

[1020] "Knowledge data" refers to data that companies and organizations have accumulated and stored in digital format, encompassing experience, knowledge, and information.

[1021] A "machine learning model" is a collection of algorithms that learn patterns and rules based on data and use them to make predictions and classifications on new data.

[1022] "Training" is the process of using knowledge data to teach a machine learning model patterns and rules.

[1023] A "generative AI model" is a model trained using machine learning that has the ability to generate answers to user questions.

[1024] "Speech synthesis software" is software that converts text data into speech data.

[1025] An "emotion engine" is a technology that analyzes user emotions from their input and actions and generates information based on those emotions.

[1026] An "API" is an interface for exchanging functions and data between different software programs.

[1027] A "data source" refers to an external system or service that provides data.

[1028] The "JSON format" is a text-based data format for representing data in a lightweight and highly readable format.

[1029] The system of the present invention includes means for collecting knowledge data from a database, feeding it to a machine learning model for training, and generating answers to user questions. It also includes means for converting the generated answers into audio data using speech synthesis software and providing the audio data to the user. Furthermore, this system incorporates an emotion engine that can recognize the user's emotions and provide the most appropriate answers based on those emotions.

[1030] First, the server collects knowledge data from the company's internal knowledge database and external data sources. Specifically, it retrieves data from external data sources via APIs, converts it to JSON format, and stores it in the database. Next, the server uses the collected knowledge data to train a generative AI model. This training process uses machine learning frameworks such as TensorFlow.

[1031] When a user enters a question, the device sends it to the server. The server inputs the received question into a generative AI model to generate the best possible answer. In this process, an emotion engine analyzes the user's input and recognizes their emotions. The recognized emotion information is fed back to the generative AI model, which then generates the best possible answer.

[1032] The generated response is sent to text-to-speech software (e.g., Google Text-to-Speech API), where the text data is converted into audio data. During this speech synthesis process, emotional information obtained from an emotion engine is used to adjust the voice tone. The generated audio data is then sent from the server to the device, which plays it back to provide the response to the user.

[1033] As a concrete example, consider a scenario where a customer support representative at a certain company asks, "How do I handle complaints?" The user enters this question into a device and clicks the send button. The device sends the query to a server, which passes it to a generative AI model. The generative AI model generates an appropriate answer from its trained knowledge data, and the server converts that answer into audio data using the Google Text-to-Speech API. In this process, an emotion engine recognizes the user's emotions, such as urgency and frustration, and generates an answer with an appropriate tone of voice. The device plays the audio data received from the server and provides the answer to the customer support representative.

[1034] Example of a prompt:

[1035] "Create a customer support system that generates responses to user inquiries about how to handle complaints. Please explain the process for outputting responses in an appropriate tone of voice, taking into account the user's emotional state and potential anxiety."

[1036] In this way, the system of the present invention realizes efficient and intuitive knowledge management and information provision to users, and further enables responses that take into account the user's feelings.

[1037] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1038] Step 1: Data Collection

[1039] The server collects knowledge data from the company's internal knowledge database and external data sources. Specifically, the server uses APIs to retrieve data from external data sources. This data is converted to JSON format and stored in the database. The input is raw data retrieved from the API, and the output is knowledge data converted to JSON format.

[1040] Step 2: Model Training

[1041] The server uses the collected knowledge data to train a generative AI model. Specifically, the server loads knowledge data from a database and performs data preprocessing. For training, it uses a machine learning framework such as TensorFlow, and optimizes the model by repeating multiple epochs. The input is knowledge data in JSON format, and the output is a trained generative AI model.

[1042] Step 3: Receiving questions from users

[1043] The user enters a question via the device. For example, they might enter "How should I handle complaints?" through a web browser's input form. When the user presses the submit button, the device sends the question to the server as an HTTP POST request. The input is the user's question text, and the output is the question data in HTTP request format.

[1044] Step 4: Analyzing the Question

[1045] The server inputs the received question into a generative AI model and generates an answer. Specifically, the server first parses the question text and inputs it into the generative AI model. During this process, the emotion engine analyzes the user's input and recognizes their emotions. The input is question data in HTTP request format, and the output is answer text with added emotion information.

[1046] Step 5: Generating the answer

[1047] The generative AI model generates the optimal answer based on the input question and sentiment information. The generated answer is temporarily stored in the server's memory. The input is question data including sentiment information, and the output is the generated answer text.

[1048] Step 6: Generate audio data

[1049] The server sends the generated response to speech synthesis software, which converts it into speech data. Specifically, the server uses the Google Text-to-Speech API to convert the text data into speech data. In this process, emotional information obtained from the emotion engine is reflected in the tone of voice. The input is the response text, and the output is emotion-adjusted speech data.

[1050] Step 7: Providing audio data

[1051] The generated audio data is sent from the server to the terminal, which then plays it back to provide a response to the user. Specifically, the terminal plays the audio data received as an HTTP response using the Audio API of the web browser. The input is the audio data sent from the server, and the output is the played audio.

[1052] (Application Example 2)

[1053] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1054] Traditional knowledge management systems often suffer from usability and poor information comprehension because they provide information without considering the user's emotions. Furthermore, systems for factory workers, in particular, require immediate responses in emergencies and high-stress situations, necessitating emotionally sensitive support. Therefore, there is a growing need for systems that recognize users' emotions in real time and provide optimal responses accordingly.

[1055] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1056] In this invention, the server includes means for collecting knowledge data from a database, means for feeding the collected knowledge data to a machine learning model and training it, means for generating answers to user questions using the trained model, means for converting the generated answers into voice data using speech synthesis software, means for providing the voice data to the user, means including an emotion engine that recognizes the user's emotions, and means for adjusting the voice tone based on emotion information from the emotion engine. This enables the rapid and appropriate provision of information that takes the user's emotions into consideration.

[1057] A "database" is an information storage system for systematically storing and managing knowledge data.

[1058] "Knowledge data" refers to a collection of information and knowledge used in a knowledge management system.

[1059] A "machine learning model" is an algorithm that learns from large amounts of data and uses that learning to make predictions and classifications for new inputs.

[1060] "Training methods" refer to the process of providing data to a machine learning model to allow it to learn and improve its performance.

[1061] "Generated answers" refer to information that a machine learning model generates based on the user's questions.

[1062] "Speech synthesis software" is a program that converts text data into speech data.

[1063] An "emotion engine" is a technology that recognizes a user's emotions in real time.

[1064] "Voice tone" refers to the characteristics of sound in audio data that express emotions and intentions.

[1065] A "server" is a computer system that collects and processes knowledge data and provides services to users.

[1066] A "user" is an individual or organization that uses a knowledge management system to retrieve information.

[1067] The system for carrying out this invention consists of multiple components. The main components include a database, a machine learning model, an emotion engine, speech synthesis software, and a server. The functions and roles of these components are described below.

[1068] The server periodically collects knowledge data from internal and external knowledge databases. The collected data is converted into a common format (e.g., JSON) and stored in the database. This knowledge data stored in the database serves as the basis for generating answers to user questions.

[1069] The collected knowledge data is used to train a machine learning model. The training process extracts patterns and rules from the large amount of knowledge data, enabling the model to generate appropriate answers to user questions. Because this training process requires significant computing resources, cloud computing environments are often used.

[1070] When a user enters a question, the device sends it to the server. The server inputs the question into a generative AI model to generate the best answer. At this time, the emotion engine analyzes the user's input and recognizes their emotions. The emotion engine uses a transformer model (e.g., DistilBERT) to determine the user's emotions. The recognized emotion information is fed back to the generative AI model, which generates an answer that takes the user's emotions into consideration.

[1071] The generated responses are sent to text-to-speech software, where the text data is converted into audio data. At this stage, the voice tone is adjusted based on emotional information obtained from the emotion engine, generating audio corresponding to each emotion. Google Cloud Text-to-Speech API is often used for this text-to-speech process. The generated audio data is sent from the server to the device, which then plays it back to provide the user with the response.

[1072] As a concrete example, consider a scenario where a factory worker asks, "What should I do if a machine stops?" The user inputs this question via a smartphone or smart glasses and sends it. The device sends this query to a server, which then sends it to a generative AI model. The model generates an appropriate response from its trained knowledge data, and the server sends this response to speech synthesis software to generate audio data. In this process, an emotion engine recognizes emotions such as stress and urgency from the user's input, and generates a response in a corresponding voice tone. The device plays the audio data received from the server, providing the factory worker with an appropriate response via voice.

[1073] The system of this invention realizes efficient and intuitive knowledge management and information delivery that takes user emotions into consideration. Users can receive timely and appropriate information and take appropriate action, especially in highly urgent situations.

[1074] Example of a prompt:

[1075] User question: "What should I do if the machine stops working?"

[1076] Generated response: "Try restarting the machine, and if it still doesn't work, contact maintenance."

[1077] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1078] Step 1:

[1079] The server collects knowledge data from internal and external knowledge databases. Specifically, it retrieves data via APIs, converts it to JSON format, and stores it in the database.

[1080] Input: Knowledge data

[1081] Data processing: Retrieve data using the API and convert it to JSON format.

[1082] Output: Knowledge data stored in the database

[1083] Step 2:

[1084] The server feeds the collected knowledge data into a machine learning model for training. During the training process, the model learns patterns and rules in the knowledge data, improving its accuracy.

[1085] Input: Knowledge data stored in the database

[1086] Data processing: Learning using machine learning algorithms

[1087] Output: Trained machine learning model

[1088] Step 3:

[1089] The user enters the question. Specifically, they enter the question in text format via a smartphone or smart glasses, and the device sends it to the server.

[1090] Input: User's question (text format)

[1091] Data processing: None

[1092] Output: User questions sent to the server

[1093] Step 4:

[1094] The server inputs the questions received from the user into a generating AI model to produce the optimal answer. In this process, the emotion engine analyzes the user's input and recognizes their emotions.

[1095] Input: User's question

[1096] Data processing: Emotion analysis using an emotion engine, response generation using a generative AI model.

[1097] Output: Generated responses and user sentiment information

[1098] Step 5:

[1099] Based on the emotional information obtained from the emotion engine, the server uses speech synthesis software to convert text data into speech data in order to adjust the voice tone of the generated response.

[1100] Input: Generated response, user sentiment information

[1101] Data processing: Speech synthesis (conversion from text to speech)

[1102] Output: Audio data of the adjusted voice tone

[1103] Step 6:

[1104] The server sends the generated audio data to the terminal. The terminal plays the received audio data and provides the user with a response.

[1105] Input: Audio data

[1106] Data processing: None

[1107] Output: Audio data provided to the user

[1108] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1109] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1110] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1111] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1112] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1113] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1114] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1115] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1116] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1117] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1118] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1119] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1120] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1121] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1122] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1123] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1124] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1125] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1126] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1127] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1128] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1129] The following is further disclosed regarding the embodiments described above.

[1130] (Claim 1)

[1131] A means of collecting knowledge data from a database,

[1132] A method for feeding collected knowledge data into a machine learning model and performing training,

[1133] A means for generating answers to user questions using a trained model,

[1134] A means of converting the generated response into speech data using speech synthesis software,

[1135] Means of providing audio data to users,

[1136] A system that includes this.

[1137] (Claim 2)

[1138] The system according to claim 1, which collects knowledge data through acquisition from APIs and data sources.

[1139] (Claim 3)

[1140] The system according to claim 1, which manages knowledge data stored in a database in JSON format.

[1141] "Example 1"

[1142] (Claim 1)

[1143] A means of collecting knowledge data from a database,

[1144] A method for feeding collected knowledge data into a machine learning model and performing training,

[1145] A means of receiving questions entered by the user via a terminal and sending them to a server,

[1146] A means for generating answers to user questions using a trained generative AI model,

[1147] A means of converting the generated response into speech data using speech synthesis software,

[1148] A means of transmitting the generated audio data to a terminal and providing the audio data to the user,

[1149] A means of saving the system usage history as a log and performing effectiveness analysis,

[1150] A system that includes this.

[1151] (Claim 2)

[1152] The system according to claim 1, which collects knowledge data through acquisition from APIs and data sources.

[1153] (Claim 3)

[1154] The system according to claim 1, which manages knowledge data stored in a database in a common format.

[1155] "Application Example 1"

[1156] (Claim 1)

[1157] A means of collecting knowledge data from a database,

[1158] A method for feeding collected knowledge data into a machine learning model and performing training,

[1159] A means for generating answers to user questions using a trained model,

[1160] A means of converting the generated response into speech data using speech synthesis software,

[1161] A means of providing audio data while the user is wearing smart glasses,

[1162] A means of running speech synthesis software on smart glasses when providing audio data,

[1163] A system that includes this.

[1164] (Claim 2)

[1165] The system according to claim 1, which collects knowledge data through acquisition from APIs and data sources.

[1166] (Claim 3)

[1167] The system according to claim 1, which manages knowledge data stored in a database in JSON format.

[1168] "Example 2 of combining an emotion engine"

[1169] (Claim 1)

[1170] A means of collecting knowledge data from a database,

[1171] A method for feeding collected knowledge data into a machine learning model and performing training,

[1172] A means for generating answers to user questions using a trained model,

[1173] A means of converting the generated response into speech data using speech synthesis software,

[1174] A means of analyzing user emotions,

[1175] A means of generating the optimal response based on analyzed emotional information,

[1176] Means of providing audio data to users,

[1177] A system that includes this.

[1178] (Claim 2)

[1179] The system according to claim 1, which collects knowledge data through acquisition from APIs and data sources.

[1180] (Claim 3)

[1181] The system according to claim 1, which manages knowledge data stored in a database in JSON format.

[1182] "Application example 2 when combining with an emotional engine"

[1183] (Claim 1)

[1184] A means of collecting knowledge data from a database,

[1185] A method for feeding collected knowledge data into a machine learning model and performing training,

[1186] A means for generating answers to user questions using a trained model,

[1187] A means of converting the generated response into speech data using speech synthesis software,

[1188] Means of providing audio data to users,

[1189] A means including an emotion engine that recognizes the user's emotions,

[1190] A means of adjusting voice tone based on emotional information from an emotion engine,

[1191] A system that includes this.

[1192] (Claim 2)

[1193] The system according to claim 1, which collects knowledge data through acquisition from APIs and data sources.

[1194] (Claim 3)

[1195] The system according to claim 1, which manages knowledge data stored in a database in JSON format. [Explanation of symbols]

[1196] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of collecting knowledge data from a database, A method for feeding collected knowledge data into a machine learning model and performing training, A means for generating answers to user questions using a trained model, A means of converting the generated response into speech data using speech synthesis software, Means of providing audio data to users, A system that includes this.

2. The system according to claim 1, which collects knowledge data through acquisition from APIs and data sources.

3. The system according to claim 1, which manages knowledge data stored in a database in JSON format.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A