system

The agent system for human resource management addresses the inefficiencies in existing systems by using language analysis and database search to provide immediate and accurate information, improving operational efficiency and reducing HR workload.

JP2026101370APending Publication Date: 2026-06-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-10
Publication Date
2026-06-22

AI Technical Summary

Technical Problem

Existing human resource management systems within enterprises face challenges in efficiently and accurately providing information to employees, leading to reduced work efficiency and increased workload due to complex information systems and the risk of overlooked information.

Method used

An agent system for human resource management that utilizes language analysis to understand user inquiries, searches internal databases for relevant information, and generates natural language responses, reducing the burden on HR departments and improving information accessibility.

Benefits of technology

Enables employees to quickly and accurately obtain necessary information, thereby enhancing operational efficiency and reducing the workload on HR departments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026101370000001_ABST
    Figure 2026101370000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of receiving user instructions as voice using natural language, A speech recognition means for converting the aforementioned instructions from speech to text, A language analysis means for analyzing the transcribed instructions and understanding their intent, Based on the understood intent, a means of using a knowledge base to search for information, A means for generating a response in natural language based on the aforementioned searched information, Means for providing the generated response in audio and visual formats, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including: receiving a user utterance; adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot; encoding the prompt; and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the personnel management within an enterprise, there is a problem that it is not easy to comprehensively search for a plurality of regulations and procedures and quickly and accurately access necessary information. For this reason, it takes time for employees to search for necessary information, which not only reduces work efficiency but also increases the workload of the HR department. In particular, in the complex information system within the company, there is also a high risk that some information will be overlooked.

Means for Solving the Problems

[0005] This invention provides an agent system for human resource management, comprising language analysis means that receive natural language inquiries from users, analyze them, and understand their intent. Furthermore, it includes generation means that search for relevant information from the company's internal database based on the analyzed intent, and generate an appropriate natural language response for the user based on the retrieved information. This generated response is transmitted to the user's terminal, enabling immediate information provision to the user. As a result, employees can quickly and accurately obtain the necessary information, and the workload of the HR department can be reduced.

[0006] "Human resource management" refers to activities that maximize the use of available resources within an organization and optimize their work, skills, and career development.

[0007] An "agent system" is software that automatically performs specific tasks based on user instructions and processes information efficiently.

[0008] "Users" refer to individuals or organizations that operate the system to search for information or perform tasks.

[0009] "Natural language" refers to linguistic forms based on the words and grammar that humans use on a daily basis.

[0010] "Means of receiving inquiries" refers to interfaces, devices, or programs for receiving questions or requests for information.

[0011] "Language analysis means" refers to a process or device that uses natural language processing technology to decipher the meaning and intent of a text.

[0012] An "internal database" is a data storage system used to structure, store, and manage information within a company or organization.

[0013] "Means of searching for information" refer to the technologies and methods used to find specific data or information from a database.

[0014] The "generation means" is a process or technology for creating new information or content based on underlying data and conditions.

[0015] The "terminal" is a computer or device for a user to access and operate the system.

Brief Explanation of Drawings

[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Mode for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0020] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] This invention is an agent system for supporting human resource management, which understands natural language, acquires, processes, and provides information to users, thereby assisting employees and HR departments within a company. Specific embodiments of this system are described below.

[0038] This system consists of servers, terminals, and users. This section details how each element functions.

[0039] The server is the core of this system, responsible for processing queries, acquiring data, and transmitting generated information. The server first receives queries sent from terminals. Next, it uses language analysis tools to process the query content using natural language processing and analyze the user's intent. External generative AI technologies can be incorporated into this analysis process, enabling highly accurate language understanding.

[0040] Furthermore, the server searches the internal database based on the analyzed intent and retrieves relevant information. This allows for the rapid collection of information necessary for the user. The retrieved information is organized into a user-friendly format and constructed as a response using a generation mechanism. This generated response is then sent to the terminal.

[0041] The terminal functions as an interface with the user. It transmits natural language queries entered by the user to the server, receives the response from the server, and displays it on the user interface. This allows the user to operate the system easily and intuitively and access the necessary information.

[0042] Users can use the system to ask questions related to company regulations and procedures and receive answers quickly. For example, if a user asks, "How do I apply for paid leave?", the user sends this question to the server via their terminal. The server then analyzes and retrieves information from the relevant company database, generates an answer, and displays it on the user's terminal.

[0043] This series of processes allows users to obtain accurate information quickly, thereby improving operational efficiency. In addition, it reduces the burden on the HR department of individually handling employee inquiries. In this way, the present invention provides innovative support for information management within a company.

[0044] The following describes the processing flow.

[0045] Step 1:

[0046] The user enters their inquiry in natural language through their device. For example, they might send a message such as, "How do I apply for leave?"

[0047] Step 2:

[0048] The terminal receives the inquiry and formats it as an API request to send to the server. This format includes the user's inquiry content and identification information.

[0049] Step 3:

[0050] The server receives an API request from the terminal. To analyze the received inquiry, it executes a language analysis method using a generative AI model.

[0051] Step 4:

[0052] The server performs language analysis to extract the intent of the query. In this process, the AI ​​understands the context and determines what the user is specifically asking for.

[0053] Step 5:

[0054] Based on the analyzed intent, the server searches the internal database for relevant data. For example, it queries the database to retrieve information related to leave requests.

[0055] Step 6:

[0056] The server is constructed using a generation method that organizes the acquired data into a user-friendly format and generates responses in natural language.

[0057] Step 7:

[0058] The server sends the generated response to the device. This response is returned to the user's device as an API response.

[0059] Step 8:

[0060] The terminal displays the response received from the server. The user interface visually presents the answer to the inquiry.

[0061] Step 9:

[0062] The user reviews the responses displayed on their device and uses them to determine tasks and actions. For example, they might follow the instructions for submitting a leave request.

[0063] (Example 1)

[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0065] Within a company, responding quickly and accurately to employee inquiries is crucial for improving operational efficiency and optimizing information management. However, traditional human resource management systems struggle to accurately grasp employee intentions and provide relevant information appropriately, resulting in a significant amount of time and effort being required to respond to inquiries.

[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0067] In this invention, the server includes a device for receiving inquiries from users, a language analysis device for analyzing the inquiries and understanding their intent, and a device for retrieving relevant information from an information storage device based on the understood intent. This enables the rapid and accurate acquisition and provision of information in response to user inquiries.

[0068] A "device for receiving inquiries from users" is a device that transmits inquiries made by users in natural language to the system.

[0069] A "language analysis device" is a device that analyzes received inquiries and understands the user's intent from their content.

[0070] An "information storage device" is a device that stores relevant information within a company and is used to hold data that is searchable.

[0071] A "generation device" is a device that generates natural language responses to be provided to users based on analyzed information.

[0072] A "display device" is a device used to visually present the generated response to the user.

[0073] A "terminal device" is a device used by users to input inquiries and receive and display responses from the server.

[0074] This agent system consists of servers, terminals, and users, and aims to improve information management and operational efficiency within a company by responding to user inquiries through natural language.

[0075] The server functions as the central hub of this system. The server receives inquiries sent from user terminals and processes their content using a language analysis device. An external generative AI model is used for this process. This model analyzes user inquiries and enables a high degree of understanding of intent. Based on the understood intent, the server retrieves relevant information from the information storage device, and the response is constructed by a generator. During this process, the information is formatted to be easily understood by the user.

[0076] The terminal provides an interface for users to input questions in natural language and receive and display answers from the server. This allows users to intuitively use the system and quickly obtain information.

[0077] Users can ask questions about company regulations and procedures through the system. For example, if a user asks, "How do I apply for paid leave?", they type the question on their terminal. The server analyzes the intent, retrieves relevant information, and generates a specific answer such as, "Download the application form from the company's portal site, get your supervisor's approval, and submit it to the Human Resources Department," which is then displayed on the terminal.

[0078] An example of a prompt message might be, "Please provide details about the company's remote work guidelines." This allows users to efficiently obtain information and streamline their daily work.

[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0080] Step 1:

[0081] The user uses a device to input a request in natural language. For example, they might input something like, "Please tell me how to apply for paid leave." This input generates digital data that conveys the user's intent to the device.

[0082] Step 2:

[0083] The terminal sends the user's inputted question to the server. Here, the question, formalized as digital data, is transferred to the server. In this process, the terminal appropriately packets the data and transmits it according to the communication protocol.

[0084] Step 3:

[0085] The server receives data sent from the terminal and analyzes the query using a language analysis device. Specifically, it uses a generative AI model to scrutinize the input data and understand the user's intent with high accuracy. Through this analysis, the server obtains specific instructions on what to search for.

[0086] Step 4:

[0087] The server searches the information storage device based on the analyzed intent. It finds data that matches the intended information and extracts the relevant data. At this time, it uses database queries to quickly retrieve the information.

[0088] Step 5:

[0089] The server uses a generator to construct responses based on the acquired data. The information is converted into natural language and organized into a format that is easy for the user to understand. For example, specific guidance such as "Download the application form from the company's portal site, get your supervisor's approval, and submit it to the Human Resources Department" is generated.

[0090] Step 6:

[0091] The server sends the generated response to the terminal. The terminal receives this information and prepares to display it to the user. The data is transmitted again in digital format, and the terminal displays the received data correctly.

[0092] Step 7:

[0093] The device displays the answer to the user. The user can check the answer on the device screen and obtain the necessary information. This allows the user to receive an answer to their inquiry quickly.

[0094] (Application Example 1)

[0095] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0096] Existing human resource management systems have the drawback of being difficult for general users to use because they require specialized operation and search skills to efficiently obtain the information users need. Furthermore, the limited means of providing information visually and audibly results in insufficient speed and convenience of information retrieval. To address these issues, there is a need for a system that provides a simple and intuitive interface, such as that of a household robot, allowing users to instantly obtain the information they need.

[0097] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0098] In this invention, the server includes means for receiving instructions from the user as voice in natural language, voice recognition means for converting the instructions from voice to text, and language analysis means for analyzing the transcribed instructions and understanding the user's intent. As a result, users can request information by voice without requiring any special skills, and the information is provided via voice and display, enabling quick and convenient information retrieval.

[0099] "Natural language" refers to the language that humans use on a daily basis, and the technology that enables dialogue by converting it into a form that computers can understand.

[0100] "Voice recognition means" refers to a technological device that analyzes voice input from a user and converts it into text.

[0101] "Linguistic analysis methods" refer to technologies and processes for analyzing textual instructions and understanding their intent.

[0102] A "knowledge base" is a database that stores specific information and serves as a foundation for searching and retrieving information.

[0103] A "generative artificial intelligence model" is an artificial intelligence technology that automatically creates text, images, and other elements according to a generation task.

[0104] "Means of providing information visually" refers to technological devices that visually display obtained information to the user through a screen or other means.

[0105] To implement this invention, a system utilizing a home voice-enabled robot is required. The system consists of a microphone with voice recognition capabilities, a display and speaker, and a backend server. The server resides in the cloud and utilizes Google® Cloud Speech-to-Text API, and OpenAI® GPT models for language analysis. An internal company database is used as the knowledge base.

[0106] The server uses speech recognition to convert user instructions from speech to text. A language analysis tool utilizes a GPT model to understand the intent behind the transcribed instructions. Based on the intent obtained through this analysis, relevant information is retrieved from a knowledge base. The retrieved information is provided to the user visually and aurally. Visually, the information is displayed on a screen; aurally, it is provided via audio through speakers.

[0107] For example, a user might say to the robot, "Please check the date of the next meeting." The robot transcribes the voice into text and sends it to a server. The server performs language analysis, retrieves the date of the next meeting from its knowledge base, and returns that information to the robot. The robot displays the retrieved date on its screen and informs the user verbally, "The next meeting is next Wednesday."

[0108] An example of a prompt message would be, "When is the next meeting scheduled?" Upon receiving this instruction, the AI ​​model on the server would analyze the request and provide relevant information to the user.

[0109] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0110] Step 1:

[0111] The user gives instructions to the robot by voice. The robot receives the user's voice input through its built-in microphone. This voice is then passed on to the next step as is.

[0112] Step 2:

[0113] The device uses the Google Cloud Speech-to-Text API to convert the user's speech into text data. This process analyzes the speech input and generates a corresponding text string, making it possible to send the user's intent as text to the server.

[0114] Step 3:

[0115] The server passes the received text data to a language analysis system. This system uses OpenAI's generative AI model to analyze the text content and process the data to understand the user's intent. As a result, specific information requests become clear.

[0116] Step 4:

[0117] The server searches a knowledge base to retrieve information based on the analyzed intent. It uses SQL queries to extract relevant data from the internal database. This data includes specific answers to the user's requests.

[0118] Step 5:

[0119] The server generates a response in natural language based on the acquired information. It utilizes language generation technology to format the information in a way that is easy for the user to understand. Next, the formatted response data is sent to the terminal.

[0120] Step 6:

[0121] The terminal provides the user with text responses received from the server, both visually and audibly. The response content is displayed on the screen, and the information is transmitted audibly via the speaker using speech synthesis technology. This allows the user to easily obtain the requested information.

[0122] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0123] This invention aims to improve the user experience by integrating an emotion engine for emotion recognition into an agent system for human resource management, thereby personalizing responses to employee inquiries. This system consists of three main components: a server, a terminal, and a user.

[0124] The server, as the central hub of the system, handles all data processing and analysis. The server processes natural language queries received from terminals. These queries are first passed to the emotion engine, which analyzes the user's emotions. The emotion engine, for example, identifies the user's emotional state (e.g., frustration, joy, excitement) based on keywords and expressions in the user's text.

[0125] Based on these analysis results, the server gains a more accurate understanding of the user's intent through language analysis tools. After understanding the intent, the server searches the internal database to gather necessary information. Then, reflecting the results of the emotion engine, the generation tools construct a response with appropriate tone and content. For example, for a dissatisfied user, a more polite and empathetic response will be generated.

[0126] The generated natural language response is sent from the server to the terminal. The terminal displays the received response to the user, allowing the user to verify it through an appropriate interface.

[0127] The terminal functions as an interface connecting the user and the server. It not only sends user-inputted information to the server but also displays received responses appropriately. Furthermore, it can change the display style, message colors, fonts, and other elements on the terminal based on the results of the emotion engine.

[0128] Users make inquiries using natural language and obtain various internal company information through the responses. By benefiting from emotion recognition capabilities, users can receive responses that are tailored to their feelings. For example, if a user makes an emotionally charged inquiry such as, "Why is this procedure so complicated?", the system can recognize the user's frustration and provide a response that includes ways to simplify the procedure and support information.

[0129] This system provides employees with a means to obtain necessary information without stress, improving the quality of communication in talent management. Furthermore, it can contribute to the efficiency of HR operations.

[0130] The following describes the processing flow.

[0131] Step 1:

[0132] Users make inquiries using natural language through their devices. For example, they might enter a specific question such as, "What is the progress of this project?"

[0133] Step 2:

[0134] The device formats the user's inquiry received by the device and sends it to the server as an API request.

[0135] Step 3:

[0136] The server receives the inquiry from the terminal and passes it to the emotion engine. The emotion engine analyzes the text and identifies the user's emotional state (e.g., excitement, doubt, anxiety).

[0137] Step 4:

[0138] The server uses language analysis tools to analyze the user's inquiry and understand their precise intent. This process includes using AI models to deepen the understanding of the context.

[0139] Step 5:

[0140] Based on the server's understanding of the intent, it searches the internal database to retrieve relevant information. For example, it might retrieve the latest report data regarding the progress of a project.

[0141] Step 6:

[0142] The server takes the emotion recognition results into account and generates a response with a tone and content appropriate to the user using a generation method. If the emotion is determined to be anxiety, words of encouragement may be included.

[0143] Step 7:

[0144] The server sends the generated response to the terminal. This response incorporates language and information that takes the user's emotions into consideration.

[0145] Step 8:

[0146] The device receives the submitted response and displays it on the user interface. The display style (e.g., color and font) may also be adjusted depending on the emotional state.

[0147] Step 9:

[0148] The user reviews the answers displayed on their device. Based on the information provided, the user decides on their next course of action and plans their next steps in the process.

[0149] (Example 2)

[0150] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0151] Existing human resource management systems struggle to respond to user inquiries in a way that reflects individual emotions and intentions, limiting their ability to improve the user experience. There is a need for a system that can respond appropriately while considering user emotions.

[0152] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0153] In this invention, the server includes means for receiving natural language queries from users, language analysis means for analyzing the queries and understanding their intent and emotions, and means for retrieving relevant information from an information storage area based on the understood intent and emotions. This enables personalized responses that correspond to the user's emotional state.

[0154] "User" refers to a person who makes natural language queries to the system.

[0155] "Natural language" refers to the language that humans use on a daily basis, and is the text or sound that is processed by computer programs.

[0156] An "inquiry" refers to a request expressed in natural language that a user sends to a system to obtain necessary information or a response.

[0157] "Language analysis means" refers to a system that processes received natural language queries and analyzes the user's intentions and emotions.

[0158] "Intention" refers to the purpose and evaluation of what the user is seeking from the system through natural language.

[0159] "Emotions" refer to the emotional states included in user inquiries, such as joy or dissatisfaction, which are the subjects of analysis.

[0160] "Information storage area" refers to databases and data storage used to store various types of information necessary to answer inquiries.

[0161] "Generation means" refers to a mechanism that generates natural language to respond to users based on analysis results.

[0162] "Device" refers to a physical or virtual device used by a user to interface with a system, such as a smartphone or computer.

[0163] This invention aims to improve the user experience in a human resource management agent system by personalizing responses to user inquiries. The system consists of three main components: a server, a terminal, and a user.

[0164] The server functions as the central hub of this system, deeply analyzing the content of incoming inquiries. Specifically, the server analyzes the user's natural language inquiries received from terminals through an emotion engine. The emotion engine uses natural language processing techniques to identify the user's emotional state based on keywords and expressions contained in the analyzed text. Through this analysis, the server gains a more accurate understanding of the user's inquiry intent. After understanding the intent, the server searches for the necessary information from its information storage area and constructs a response in an appropriate tone using a generative AI model. For example, if a user asks a question indicating confusion, the server will empathize with that emotion and provide a polite explanation.

[0165] The terminal acts as an interface connecting the user and the server. The terminal is responsible for sending information entered by the user to the server and receiving and displaying responses sent from the server. During display, the style and format are adjusted to be visually appealing, taking into account the analysis results of the emotion engine.

[0166] Users can use their devices to make inquiries in natural language and obtain the necessary information through the responses. This system allows users to receive responses that are tailored to their emotional state, resulting in a more satisfying experience.

[0167] For example, if a user asks, "What is the progress of this project?", the server will use an emotion engine to analyze the user's anticipated concerns and anxieties, and then form a detailed response based on the ongoing progress information. An example of a prompt to the generative AI model might be, "Please create a detailed project progress report for an anxious user."

[0168] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0169] Step 1:

[0170] The user inputs a natural language query via a terminal. This input is in text format, and the terminal receives a specific question such as, "What is the progress of this project?" The terminal processes the input text data into a digital format and sends it to the server.

[0171] Step 2:

[0172] The server passes the inquiry received from the terminal to the emotion engine. The input is the user's natural language text, and the emotion engine processes this text to analyze the emotional state (e.g., anxiety or curiosity). As a result of the analysis, emotional state data is generated. This data is used as output in subsequent processes.

[0173] Step 3:

[0174] The server uses emotional state data obtained from the emotion engine to understand the user's intent through linguistic analysis. The input data consists of the emotional state and the original query text. Analysis extracts the user's intent (e.g., "I want to get project progress information"). This yields intent data as output for the next process.

[0175] Step 4:

[0176] The server searches the information storage area based on intent data and retrieves relevant information. The input is the user's intent data, and a query is created to search for the corresponding database entry. The output includes detailed information related to project progress.

[0177] Step 5:

[0178] The server uses the acquired information and emotional state data to generate a response using a generative AI model. The input includes relevant information and emotional state. The generative AI model processes this information and outputs a response in natural language that takes emotions into consideration. The output response is structured in a way that will be well-received by the user.

[0179] Step 6:

[0180] The server sends the generated response to the terminal. The terminal visually displays the received response to the user. The input data is the generated response, and the terminal formats this response to be displayed in a readable format. For example, it might adjust the font size and color and display it to the user as "The project is progressing as planned. Please rest assured."

[0181] (Application Example 2)

[0182] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0183] In autonomous vehicles, there is a need to provide not only information in response to passengers' questions and requests, but also optimal responses that are sensitive to passengers' emotions. However, conventional automated response systems have difficulty understanding passengers' emotional states, resulting in decreased satisfaction both financially and emotionally. Therefore, the challenge is to provide a system that can alleviate the anxiety and stress experienced by passengers in autonomous vehicles and improve the in-vehicle experience.

[0184] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0185] In this invention, the server includes means for analyzing natural language using an external emotion recognition structure to identify the user's emotional state, means for retrieving relevant information and suggestions from a database based on the identified intentions and emotional state, and means for generating a natural language response optimized according to the user's situation based on the retrieved information and suggestions. This enables personalized responses that respond to the emotions of passengers in an autonomous vehicle.

[0186] "Users" refers to passengers who use this system to obtain services and information within an autonomous vehicle.

[0187] "Emotional state" refers to information indicating the user's emotional situation, obtained as a result of the emotion recognition structure analyzing the user's natural language.

[0188] "External emotion recognition structure" refers to analytical techniques that use external artificial intelligence or algorithms to identify a user's emotions through natural language processing.

[0189] A "natural language response" is a sentence-form response that is generated based on the analyzed intent and emotions of the user, and has a consistent context.

[0190] A "database" is an information management system that stores and makes searchable past inquiry history, proposal information, and other related data.

[0191] An "autonomous vehicle" refers to a vehicle that can drive automatically and is operated by making its own decisions using artificial intelligence technology.

[0192] The system that realizes this invention consists of three main elements: a server, a terminal, and a user. The server receives the user's natural language query as its primary input and analyzes the text information contained in that query. For the analysis, it uses sentiment analysis tools such as the Google Cloud Natural Language API to identify the user's emotional state. Based on this sentiment analysis, the server uses generative AI models such as the OpenAI GPT API to generate a personalized natural language response that matches the user's emotions.

[0193] The response is then sent to the user's device. The device implements a platform-appropriate UI / UX design to display the received response in an easily understandable visual format. For example, it provides visual feedback on the screen using Android® or iOS. In addition, the font, coloring, and text styling are adjusted according to the situation.

[0194] Ultimately, users can enjoy a reassuring experience in autonomous vehicles through the information and suggestions provided by this system. For example, if a user asks, "How far is it to the next service area?", and the server determines through sentiment analysis that the user is in a hurry, it will provide further information, including alternatives, in addition to the estimated arrival time.

[0195] Examples of prompt sentences include: "The passenger is indicating that they need to act quickly. Provide details related to location and distance, consider how you can reassure them, and generate an appropriate response."

[0196] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0197] Step 1:

[0198] The user enters a query in natural language into a terminal inside the vehicle. This input is converted into a format that can be sent from the terminal to the server. The terminal's input interface is designed to allow users to intuitively input text information.

[0199] Step 2:

[0200] The server uses the Google Cloud Natural Language API to process incoming queries and identify the user's emotional state. By analyzing the input text, it generates emotional information (e.g., feeling hurried, feeling anxious) as output based on keywords and context. This analysis result is stored for the next step.

[0201] Step 3:

[0202] Based on the obtained emotional state, the server uses the OpenAI GPT API to generate natural language responses appropriate to that emotional state. The model is executed using the emotional analysis results and prompt text as input, outputting response text in a style that meets the passenger's needs. This process generates emotionally sensitive tone and content.

[0203] Step 4:

[0204] The generated response message is sent from the server to the terminal. The terminal visually displays the response message on its user interface. This display uses fonts, colors, and layouts that are visually pleasing to the user.

[0205] Step 5:

[0206] The user reviews the displayed response and selects the next action as needed. Furthermore, if the provided information leads to further questions or requests, they can obtain new information by making another inquiry.

[0207] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0208] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0209] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0210] [Second Embodiment]

[0211] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0212] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0213] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0214] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0215] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0216] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0217] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0218] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0219] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0220] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0221] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0222] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0223] This invention is an agent system for supporting human resource management, which understands natural language, acquires, processes, and provides information to users, thereby assisting employees and HR departments within a company. Specific embodiments of this system are described below.

[0224] This system consists of servers, terminals, and users. This section details how each element functions.

[0225] The server is the core of this system, responsible for processing queries, acquiring data, and transmitting generated information. The server first receives queries sent from terminals. Next, it uses language analysis tools to process the query content using natural language processing and analyze the user's intent. External generative AI technologies can be incorporated into this analysis process, enabling highly accurate language understanding.

[0226] Furthermore, the server searches the internal database based on the analyzed intent and retrieves relevant information. This allows for the rapid collection of information necessary for the user. The retrieved information is organized into a user-friendly format and constructed as a response using a generation mechanism. This generated response is then sent to the terminal.

[0227] The terminal functions as an interface with the user. It transmits natural language queries entered by the user to the server, receives the response from the server, and displays it on the user interface. This allows the user to operate the system easily and intuitively and access the necessary information.

[0228] Users can use the system to ask questions related to company regulations and procedures and receive answers quickly. For example, if a user asks, "How do I apply for paid leave?", the user sends this question to the server via their terminal. The server then analyzes and retrieves information from the relevant company database, generates an answer, and displays it on the user's terminal.

[0229] This series of processes allows users to obtain accurate information quickly, thereby improving operational efficiency. In addition, it reduces the burden on the HR department of individually handling employee inquiries. In this way, the present invention provides innovative support for information management within a company.

[0230] The following describes the processing flow.

[0231] Step 1:

[0232] The user enters their inquiry in natural language through their device. For example, they might send a message such as, "How do I apply for leave?"

[0233] Step 2:

[0234] The terminal receives the inquiry and formats it as an API request to send to the server. This format includes the user's inquiry content and identification information.

[0235] Step 3:

[0236] The server receives an API request from the terminal. To analyze the received inquiry, it executes a language analysis method using a generative AI model.

[0237] Step 4:

[0238] The server performs language analysis to extract the intent of the query. In this process, the AI ​​understands the context and determines what the user is specifically asking for.

[0239] Step 5:

[0240] Based on the analyzed intent, the server searches the internal database for relevant data. For example, it queries the database to retrieve information related to leave requests.

[0241] Step 6:

[0242] The server is constructed using a generation method that organizes the acquired data into a user-friendly format and generates responses in natural language.

[0243] Step 7:

[0244] The server sends the generated response to the device. This response is returned to the user's device as an API response.

[0245] Step 8:

[0246] The terminal displays the response received from the server. The user interface visually presents the answer to the inquiry.

[0247] Step 9:

[0248] The user reviews the responses displayed on their device and uses them to determine tasks and actions. For example, they might follow the instructions for submitting a leave request.

[0249] (Example 1)

[0250] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0251] Within a company, responding quickly and accurately to employee inquiries is crucial for improving operational efficiency and optimizing information management. However, traditional human resource management systems struggle to accurately grasp employee intentions and provide relevant information appropriately, resulting in a significant amount of time and effort being required to respond to inquiries.

[0252] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0253] In this invention, the server includes a device for receiving inquiries from users, a language analysis device for analyzing the inquiries and understanding their intent, and a device for retrieving relevant information from an information storage device based on the understood intent. This enables the rapid and accurate acquisition and provision of information in response to user inquiries.

[0254] A "device for receiving inquiries from users" is a device that transmits inquiries made by users in natural language to the system.

[0255] A "language analysis device" is a device that analyzes received inquiries and understands the user's intent from their content.

[0256] An "information storage device" is a device that stores relevant information within a company and is used to hold data that is searchable.

[0257] A "generation device" is a device that generates natural language responses to be provided to users based on analyzed information.

[0258] A "display device" is a device used to visually present the generated response to the user.

[0259] A "terminal device" is a device used by users to input inquiries and receive and display responses from the server.

[0260] This agent system consists of servers, terminals, and users, and aims to improve information management and operational efficiency within a company by responding to user inquiries through natural language.

[0261] The server functions as the central hub of this system. The server receives inquiries sent from user terminals and processes their content using a language analysis device. An external generative AI model is used for this process. This model analyzes user inquiries and enables a high degree of understanding of intent. Based on the understood intent, the server retrieves relevant information from the information storage device, and the response is constructed by a generator. During this process, the information is formatted to be easily understood by the user.

[0262] The terminal provides an interface for users to input questions in natural language and receive and display answers from the server. This allows users to intuitively use the system and quickly obtain information.

[0263] Users can ask questions about company regulations and procedures through the system. For example, if a user asks, "How do I apply for paid leave?", they type the question on their terminal. The server analyzes the intent, retrieves relevant information, and generates a specific answer such as, "Download the application form from the company's portal site, get your supervisor's approval, and submit it to the Human Resources Department," which is then displayed on the terminal.

[0264] An example of a prompt message might be, "Please provide details about the company's remote work guidelines." This allows users to efficiently obtain information and streamline their daily work.

[0265] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0266] Step 1:

[0267] The user uses a device to input a request in natural language. For example, they might input something like, "Please tell me how to apply for paid leave." This input generates digital data that conveys the user's intent to the device.

[0268] Step 2:

[0269] The terminal sends the user's inputted question to the server. Here, the question, formalized as digital data, is transferred to the server. In this process, the terminal appropriately packets the data and transmits it according to the communication protocol.

[0270] Step 3:

[0271] The server receives data sent from the terminal and analyzes the query using a language analysis device. Specifically, it uses a generative AI model to scrutinize the input data and understand the user's intent with high accuracy. Through this analysis, the server obtains specific instructions on what to search for.

[0272] Step 4:

[0273] The server searches the information storage device based on the analyzed intent. It finds data that matches the intended information and extracts the relevant data. At this time, it uses database queries to quickly retrieve the information.

[0274] Step 5:

[0275] The server uses a generator to construct responses based on the acquired data. The information is converted into natural language and organized into a format that is easy for the user to understand. For example, specific guidance such as "Download the application form from the company's portal site, get your supervisor's approval, and submit it to the Human Resources Department" is generated.

[0276] Step 6:

[0277] The server sends the generated response to the terminal. The terminal receives this information and prepares to display it to the user. The data is transmitted again in digital format, and the terminal displays the received data correctly.

[0278] Step 7:

[0279] The device displays the answer to the user. The user can check the answer on the device screen and obtain the necessary information. This allows the user to receive an answer to their inquiry quickly.

[0280] (Application Example 1)

[0281] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0282] In existing human resource management systems, in order to efficiently obtain the information required by users, specialized operations and search skills are necessary, so there is a problem that it is difficult for general users to handle. In addition, since the means of providing information visually and auditorily are limited, the speed and convenience of obtaining information are insufficient. In order to solve this problem, there is a need to provide a system that provides a simple and intuitive interface for operations such as household robots, enabling users to immediately obtain the necessary information.

[0283] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0284] In this invention, the server includes means for receiving an instruction from a user as voice in natural language, voice recognition means for converting the instruction from voice to text, and language analysis means for analyzing the texturized instruction and grasping the intention. As a result, users do not need special skills, can request information by voice, and can obtain information quickly and conveniently by having information provided by voice and display.

[0285] "Natural language" is the language that humans use daily, and it is a technology that enables dialogue by converting it into a form that a computer can understand.

[0286] "Voice recognition means" is a technical device that analyzes voice input from a user and converts it into text.

[0287] "Language analysis means" is a technology or process for analyzing a texturized instruction and grasping its intention.

[0288] "Knowledge base" is a database that stores specific information and is the basis for use in search and information acquisition.

[0289] "Generative artificial intelligence model" is an artificial intelligence technology for automatically creating text, images, etc. according to a generation task.

[0290] "Means of providing information visually" refers to technological devices that visually display obtained information to the user through a screen or other means.

[0291] To implement this invention, a system utilizing a home voice-enabled robot is required. The system consists of a microphone with voice recognition capabilities, a display and speaker, and a backend server. The server resides in the cloud and utilizes the Google Cloud Speech-to-Text API, and OpenAI's GPT model for language analysis. An internal company database is used as the knowledge base.

[0292] The server uses speech recognition to convert user instructions from speech to text. A language analysis tool utilizes a GPT model to understand the intent behind the transcribed instructions. Based on the intent obtained through this analysis, relevant information is retrieved from a knowledge base. The retrieved information is provided to the user visually and aurally. Visually, the information is displayed on a screen; aurally, it is provided via audio through speakers.

[0293] For example, a user might say to the robot, "Please check the date of the next meeting." The robot transcribes the voice into text and sends it to a server. The server performs language analysis, retrieves the date of the next meeting from its knowledge base, and returns that information to the robot. The robot displays the retrieved date on its screen and informs the user verbally, "The next meeting is next Wednesday."

[0294] An example of a prompt message would be, "When is the next meeting scheduled?" Upon receiving this instruction, the AI ​​model on the server would analyze the request and provide relevant information to the user.

[0295] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0296] Step 1:

[0297] The user gives instructions to the robot by voice. The robot receives the user's voice input through its built-in microphone. This voice is then passed on to the next step as is.

[0298] Step 2:

[0299] The device uses the Google Cloud Speech-to-Text API to convert the user's speech into text data. This process analyzes the speech input and generates a corresponding text string, making it possible to send the user's intent as text to the server.

[0300] Step 3:

[0301] The server passes the received text data to a language analysis system. This system uses OpenAI's generative AI model to analyze the text content and process the data to understand the user's intent. As a result, specific information requests become clear.

[0302] Step 4:

[0303] The server searches a knowledge base to retrieve information based on the analyzed intent. It uses SQL queries to extract relevant data from the internal database. This data includes specific answers to the user's requests.

[0304] Step 5:

[0305] The server generates a response in natural language based on the acquired information. It utilizes language generation technology to format the information in a way that is easy for the user to understand. Next, the formatted response data is sent to the terminal.

[0306] Step 6:

[0307] The terminal provides the text answer received from the server to the user visually and auditorily. The content of the answer is displayed on the display, and information is transmitted in voice through the speaker by voice synthesis technology. Thus, the user can easily obtain the requested information.

[0308] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion recognition model 59 and perform specific processing using the user's emotion.

[0309] An object of the present invention is to personalize the response to an employee's inquiry and improve the user experience by incorporating an emotion engine for emotion recognition into an agent system for human resource management. This system is composed of three main components: a server, a terminal, and a user.

[0310] The server performs all data processing and analysis as the center of the system. The server processes the user's inquiry in natural language received from the terminal. The inquiry is first passed to the emotion engine, where the user's emotion is analyzed. The emotion engine identifies the user's emotional state (e.g., frustration, joy, excitement, etc.) based on keywords and expressions contained in the user's text, for example.

[0311] Based on this analysis result, the server more accurately understands the user's intention through the language analysis means. After the intention is grasped, the server searches the in-house database and collects the necessary information. Then, reflecting the results of the emotion engine, the generation means constructs a response in an appropriate tone and content. For example, for a user with dissatisfaction, a response using more polite and empathetic expressions is generated.

[0312] The generated answer in natural language is transmitted from the server to the terminal. The terminal displays the received response to the user so that the user can confirm it through an appropriate interface.

[0313] The terminal functions as an interface connecting the user and the server. It not only sends user-inputted information to the server but also displays received responses appropriately. Furthermore, it can change the display style, message colors, fonts, and other elements on the terminal based on the results of the emotion engine.

[0314] Users make inquiries using natural language and obtain various internal company information through the responses. By benefiting from emotion recognition capabilities, users can receive responses that are tailored to their feelings. For example, if a user makes an emotionally charged inquiry such as, "Why is this procedure so complicated?", the system can recognize the user's frustration and provide a response that includes ways to simplify the procedure and support information.

[0315] This system provides employees with a means to obtain necessary information without stress, improving the quality of communication in talent management. Furthermore, it can contribute to the efficiency of HR operations.

[0316] The following describes the processing flow.

[0317] Step 1:

[0318] Users make inquiries using natural language through their devices. For example, they might enter a specific question such as, "What is the progress of this project?"

[0319] Step 2:

[0320] The device formats the user's inquiry received by the device and sends it to the server as an API request.

[0321] Step 3:

[0322] The server receives the inquiry from the terminal and passes it to the emotion engine. The emotion engine analyzes the text and identifies the user's emotional state (e.g., excitement, doubt, anxiety).

[0323] Step 4:

[0324] The server uses language analysis tools to analyze the user's inquiry and understand their precise intent. This process includes using AI models to deepen the understanding of the context.

[0325] Step 5:

[0326] Based on the server's understanding of the intent, it searches the internal database to retrieve relevant information. For example, it might retrieve the latest report data regarding the progress of a project.

[0327] Step 6:

[0328] The server takes the emotion recognition results into account and generates a response with a tone and content appropriate to the user using a generation method. If the emotion is determined to be anxiety, words of encouragement may be included.

[0329] Step 7:

[0330] The server sends the generated response to the terminal. This response incorporates language and information that takes the user's emotions into consideration.

[0331] Step 8:

[0332] The device receives the submitted response and displays it on the user interface. The display style (e.g., color and font) may also be adjusted depending on the emotional state.

[0333] Step 9:

[0334] The user reviews the answers displayed on their device. Based on the information provided, the user decides on their next course of action and plans their next steps in the process.

[0335] (Example 2)

[0336] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0337] Existing human resource management systems struggle to respond to user inquiries in a way that reflects individual emotions and intentions, limiting their ability to improve the user experience. There is a need for a system that can respond appropriately while considering user emotions.

[0338] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0339] In this invention, the server includes means for receiving natural language queries from users, language analysis means for analyzing the queries and understanding their intent and emotions, and means for retrieving relevant information from an information storage area based on the understood intent and emotions. This enables personalized responses that correspond to the user's emotional state.

[0340] "User" refers to a person who makes natural language queries to the system.

[0341] "Natural language" refers to the language that humans use on a daily basis, and is the text or sound that is processed by computer programs.

[0342] An "inquiry" refers to a request expressed in natural language that a user sends to a system to obtain necessary information or a response.

[0343] "Language analysis means" refers to a system that processes received natural language queries and analyzes the user's intentions and emotions.

[0344] "Intention" refers to the purpose and evaluation of what the user is seeking from the system through natural language.

[0345] "Emotions" refer to the emotional states included in user inquiries, such as joy or dissatisfaction, which are the subjects of analysis.

[0346] "Information storage area" refers to databases and data storage used to store various types of information necessary to answer inquiries.

[0347] "Generation means" refers to a mechanism that generates natural language to respond to users based on analysis results.

[0348] "Device" refers to a physical or virtual device used by a user to interface with a system, such as a smartphone or computer.

[0349] This invention aims to improve the user experience in a human resource management agent system by personalizing responses to user inquiries. The system consists of three main components: a server, a terminal, and a user.

[0350] The server functions as the central hub of this system, deeply analyzing the content of incoming inquiries. Specifically, the server analyzes the user's natural language inquiries received from terminals through an emotion engine. The emotion engine uses natural language processing techniques to identify the user's emotional state based on keywords and expressions contained in the analyzed text. Through this analysis, the server gains a more accurate understanding of the user's inquiry intent. After understanding the intent, the server searches for the necessary information from its information storage area and constructs a response in an appropriate tone using a generative AI model. For example, if a user asks a question indicating confusion, the server will empathize with that emotion and provide a polite explanation.

[0351] The terminal acts as an interface connecting the user and the server. The terminal is responsible for sending information entered by the user to the server and receiving and displaying responses sent from the server. During display, the style and format are adjusted to be visually appealing, taking into account the analysis results of the emotion engine.

[0352] Users can use their devices to make inquiries in natural language and obtain the necessary information through the responses. This system allows users to receive responses that are tailored to their emotional state, resulting in a more satisfying experience.

[0353] For example, if a user asks, "What is the progress of this project?", the server will use an emotion engine to analyze the user's anticipated concerns and anxieties, and then form a detailed response based on the ongoing progress information. An example of a prompt to the generative AI model might be, "Please create a detailed project progress report for an anxious user."

[0354] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0355] Step 1:

[0356] The user inputs a natural language query via a terminal. This input is in text format, and the terminal receives a specific question such as, "What is the progress of this project?" The terminal processes the input text data into a digital format and sends it to the server.

[0357] Step 2:

[0358] The server passes the inquiry received from the terminal to the emotion engine. The input is the user's natural language text, and the emotion engine processes this text to analyze the emotional state (e.g., anxiety or curiosity). As a result of the analysis, emotional state data is generated. This data is used as output in subsequent processes.

[0359] Step 3:

[0360] The server uses emotional state data obtained from the emotion engine to understand the user's intent through linguistic analysis. The input data consists of the emotional state and the original query text. Analysis extracts the user's intent (e.g., "I want to get project progress information"). This yields intent data as output for the next process.

[0361] Step 4:

[0362] The server searches the information storage area based on intent data and retrieves relevant information. The input is the user's intent data, and a query is created to search for the corresponding database entry. The output includes detailed information related to project progress.

[0363] Step 5:

[0364] The server uses the acquired information and emotional state data to generate a response using a generative AI model. The input includes relevant information and emotional state. The generative AI model processes this information and outputs a response in natural language that takes emotions into consideration. The output response is structured in a way that will be well-received by the user.

[0365] Step 6:

[0366] The server sends the generated response to the terminal. The terminal visually displays the received response to the user. The input data is the generated response, and the terminal formats this response to be displayed in a readable format. For example, it might adjust the font size and color and display it to the user as "The project is progressing as planned. Please rest assured."

[0367] (Application Example 2)

[0368] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0369] In autonomous vehicles, there is a need to provide not only information in response to passengers' questions and requests, but also optimal responses that are sensitive to passengers' emotions. However, conventional automated response systems have difficulty understanding passengers' emotional states, resulting in decreased satisfaction both financially and emotionally. Therefore, the challenge is to provide a system that can alleviate the anxiety and stress experienced by passengers in autonomous vehicles and improve the in-vehicle experience.

[0370] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0371] In this invention, the server includes means for analyzing natural language using an external emotion recognition structure to identify the user's emotional state, means for retrieving relevant information and suggestions from a database based on the identified intentions and emotional state, and means for generating a natural language response optimized according to the user's situation based on the retrieved information and suggestions. This enables personalized responses that respond to the emotions of passengers in an autonomous vehicle.

[0372] "Users" refers to passengers who use this system to obtain services and information within an autonomous vehicle.

[0373] "Emotional state" refers to information indicating the user's emotional situation, obtained as a result of the emotion recognition structure analyzing the user's natural language.

[0374] "External emotion recognition structure" refers to analytical techniques that use external artificial intelligence or algorithms to identify a user's emotions through natural language processing.

[0375] A "natural language response" is a sentence-form response that is generated based on the analyzed intent and emotions of the user, and has a consistent context.

[0376] A "database" is an information management system that stores and makes searchable past inquiry history, proposal information, and other related data.

[0377] An "autonomous vehicle" refers to a vehicle that can drive automatically and is operated by making its own decisions using artificial intelligence technology.

[0378] The system that realizes this invention consists of three main elements: a server, a terminal, and a user. The server receives the user's natural language query as its primary input and analyzes the text information contained in that query. For the analysis, it uses sentiment analysis tools such as the Google Cloud Natural Language API to identify the user's emotional state. Based on this sentiment analysis, the server uses generative AI models such as the OpenAI GPT API to generate a personalized natural language response that matches the user's emotions.

[0379] The response is then sent to the user's device. The device implements platform-specific UI / UX design to display the received response visually and clearly. For example, using Android or iOS, visual feedback is provided on the screen. Furthermore, fonts, colors, and text styling are adjusted as needed.

[0380] Ultimately, users can enjoy a reassuring experience in autonomous vehicles through the information and suggestions provided by this system. For example, if a user asks, "How far is it to the next service area?", and the server determines through sentiment analysis that the user is in a hurry, it will provide further information, including alternatives, in addition to the estimated arrival time.

[0381] Examples of prompt sentences include: "The passenger is indicating that they need to act quickly. Provide details related to location and distance, consider how you can reassure them, and generate an appropriate response."

[0382] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0383] Step 1:

[0384] The user enters a query in natural language into a terminal inside the vehicle. This input is converted into a format that can be sent from the terminal to the server. The terminal's input interface is designed to allow users to intuitively input text information.

[0385] Step 2:

[0386] The server uses the Google Cloud Natural Language API to process incoming queries and identify the user's emotional state. By analyzing the input text, it generates emotional information (e.g., feeling hurried, feeling anxious) as output based on keywords and context. This analysis result is stored for the next step.

[0387] Step 3:

[0388] Based on the obtained emotional state, the server uses the OpenAI GPT API to generate natural language responses appropriate to that emotional state. The model is executed using the emotional analysis results and prompt text as input, outputting response text in a style that meets the passenger's needs. This process generates emotionally sensitive tone and content.

[0389] Step 4:

[0390] The generated response message is sent from the server to the terminal. The terminal visually displays the response message on its user interface. This display uses fonts, colors, and layouts that are visually pleasing to the user.

[0391] Step 5:

[0392] The user reviews the displayed response and selects the next action as needed. Furthermore, if the provided information leads to further questions or requests, they can obtain new information by making another inquiry.

[0393] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0394] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0395] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0396] [Third Embodiment]

[0397] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0398] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0399] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0400] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0401] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0402] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0403] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0404] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0405] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0406] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0407] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0408] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0409] This invention is an agent system for supporting human resource management, which understands natural language, acquires, processes, and provides information to users, thereby assisting employees and HR departments within a company. Specific embodiments of this system are described below.

[0410] This system consists of servers, terminals, and users. This section details how each element functions.

[0411] The server is the core of this system, responsible for processing queries, acquiring data, and transmitting generated information. The server first receives queries sent from terminals. Next, it uses language analysis tools to process the query content using natural language processing and analyze the user's intent. External generative AI technologies can be incorporated into this analysis process, enabling highly accurate language understanding.

[0412] Furthermore, the server searches the internal database based on the analyzed intent and retrieves relevant information. This allows for the rapid collection of information necessary for the user. The retrieved information is organized into a user-friendly format and constructed as a response using a generation mechanism. This generated response is then sent to the terminal.

[0413] The terminal functions as an interface with the user. It transmits natural language queries entered by the user to the server, receives the response from the server, and displays it on the user interface. This allows the user to operate the system easily and intuitively and access the necessary information.

[0414] Users can use the system to ask questions related to company regulations and procedures and receive answers quickly. For example, if a user asks, "How do I apply for paid leave?", the user sends this question to the server via their terminal. The server then analyzes and retrieves information from the relevant company database, generates an answer, and displays it on the user's terminal.

[0415] This series of processes allows users to obtain accurate information quickly, thereby improving operational efficiency. In addition, it reduces the burden on the HR department of individually handling employee inquiries. In this way, the present invention provides innovative support for information management within a company.

[0416] The following describes the processing flow.

[0417] Step 1:

[0418] The user enters their inquiry in natural language through their device. For example, they might send a message such as, "How do I apply for leave?"

[0419] Step 2:

[0420] The terminal receives the inquiry and formats it as an API request to send to the server. This format includes the user's inquiry content and identification information.

[0421] Step 3:

[0422] The server receives an API request from the terminal. To analyze the received inquiry, it executes a language analysis method using a generative AI model.

[0423] Step 4:

[0424] The server performs language analysis to extract the intent of the query. In this process, the AI ​​understands the context and determines what the user is specifically asking for.

[0425] Step 5:

[0426] Based on the analyzed intent, the server searches the internal database for relevant data. For example, it queries the database to retrieve information related to leave requests.

[0427] Step 6:

[0428] The server is constructed using a generation method that organizes the acquired data into a user-friendly format and generates responses in natural language.

[0429] Step 7:

[0430] The server sends the generated response to the device. This response is returned to the user's device as an API response.

[0431] Step 8:

[0432] The terminal displays the response received from the server. The user interface visually presents the answer to the inquiry.

[0433] Step 9:

[0434] The user reviews the responses displayed on their device and uses them to determine tasks and actions. For example, they might follow the instructions for submitting a leave request.

[0435] (Example 1)

[0436] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0437] Within a company, responding quickly and accurately to employee inquiries is crucial for improving operational efficiency and optimizing information management. However, traditional human resource management systems struggle to accurately grasp employee intentions and provide relevant information appropriately, resulting in a significant amount of time and effort being required to respond to inquiries.

[0438] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0439] In this invention, the server includes a device for receiving inquiries from users, a language analysis device for analyzing the inquiries and understanding their intent, and a device for retrieving relevant information from an information storage device based on the understood intent. This enables the rapid and accurate acquisition and provision of information in response to user inquiries.

[0440] A "device for receiving inquiries from users" is a device that transmits inquiries made by users in natural language to the system.

[0441] A "language analysis device" is a device that analyzes received inquiries and understands the user's intent from their content.

[0442] An "information storage device" is a device that stores relevant information within a company and is used to hold data that is searchable.

[0443] A "generation device" is a device that generates natural language responses to be provided to users based on analyzed information.

[0444] A "display device" is a device used to visually present the generated response to the user.

[0445] A "terminal device" is a device used by users to input inquiries and receive and display responses from the server.

[0446] This agent system consists of servers, terminals, and users, and aims to improve information management and operational efficiency within a company by responding to user inquiries through natural language.

[0447] The server functions as the central hub of this system. The server receives inquiries sent from user terminals and processes their content using a language analysis device. An external generative AI model is used for this process. This model analyzes user inquiries and enables a high degree of understanding of intent. Based on the understood intent, the server retrieves relevant information from the information storage device, and the response is constructed by a generator. During this process, the information is formatted to be easily understood by the user.

[0448] The terminal provides an interface for users to input questions in natural language and receive and display answers from the server. This allows users to intuitively use the system and quickly obtain information.

[0449] Users can ask questions about company regulations and procedures through the system. For example, if a user asks, "How do I apply for paid leave?", they type the question on their terminal. The server analyzes the intent, retrieves relevant information, and generates a specific answer such as, "Download the application form from the company's portal site, get your supervisor's approval, and submit it to the Human Resources Department," which is then displayed on the terminal.

[0450] An example of a prompt message might be, "Please provide details about the company's remote work guidelines." This allows users to efficiently obtain information and streamline their daily work.

[0451] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0452] Step 1:

[0453] The user uses a device to input a request in natural language. For example, they might input something like, "Please tell me how to apply for paid leave." This input generates digital data that conveys the user's intent to the device.

[0454] Step 2:

[0455] The terminal sends the user's inputted question to the server. Here, the question, formalized as digital data, is transferred to the server. In this process, the terminal appropriately packets the data and transmits it according to the communication protocol.

[0456] Step 3:

[0457] The server receives data sent from the terminal and analyzes the query using a language analysis device. Specifically, it uses a generative AI model to scrutinize the input data and understand the user's intent with high accuracy. Through this analysis, the server obtains specific instructions on what to search for.

[0458] Step 4:

[0459] The server searches the information storage device based on the analyzed intent. It finds data that matches the intended information and extracts the relevant data. At this time, it uses database queries to quickly retrieve the information.

[0460] Step 5:

[0461] The server uses a generator to construct responses based on the acquired data. The information is converted into natural language and organized into a format that is easy for the user to understand. For example, specific guidance such as "Download the application form from the company's portal site, get your supervisor's approval, and submit it to the Human Resources Department" is generated.

[0462] Step 6:

[0463] The server sends the generated response to the terminal. The terminal receives this information and prepares to display it to the user. The data is transmitted again in digital format, and the terminal displays the received data correctly.

[0464] Step 7:

[0465] The device displays the answer to the user. The user can check the answer on the device screen and obtain the necessary information. This allows the user to receive an answer to their inquiry quickly.

[0466] (Application Example 1)

[0467] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0468] Existing human resource management systems have the drawback of being difficult for general users to use because they require specialized operation and search skills to efficiently obtain the information users need. Furthermore, the limited means of providing information visually and audibly results in insufficient speed and convenience of information retrieval. To address these issues, there is a need for a system that provides a simple and intuitive interface, such as that of a household robot, allowing users to instantly obtain the information they need.

[0469] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0470] In this invention, the server includes means for receiving instructions from the user as voice in natural language, voice recognition means for converting the instructions from voice to text, and language analysis means for analyzing the transcribed instructions and understanding the user's intent. As a result, users can request information by voice without requiring any special skills, and the information is provided via voice and display, enabling quick and convenient information retrieval.

[0471] "Natural language" refers to the language that humans use on a daily basis, and the technology that enables dialogue by converting it into a form that computers can understand.

[0472] "Voice recognition means" refers to a technological device that analyzes voice input from a user and converts it into text.

[0473] "Linguistic analysis methods" refer to technologies and processes for analyzing textual instructions and understanding their intent.

[0474] A "knowledge base" is a database that stores specific information and serves as a foundation for searching and retrieving information.

[0475] A "generative artificial intelligence model" is an artificial intelligence technology that automatically creates text, images, and other elements according to a generation task.

[0476] "Means of providing information visually" refers to technological devices that visually display obtained information to the user through a screen or other means.

[0477] To implement this invention, a system utilizing a home voice-enabled robot is required. The system consists of a microphone with voice recognition capabilities, a display and speaker, and a backend server. The server resides in the cloud and utilizes the Google Cloud Speech-to-Text API, and OpenAI's GPT model for language analysis. An internal company database is used as the knowledge base.

[0478] The server uses speech recognition to convert user instructions from speech to text. A language analysis tool utilizes a GPT model to understand the intent behind the transcribed instructions. Based on the intent obtained through this analysis, relevant information is retrieved from a knowledge base. The retrieved information is provided to the user visually and aurally. Visually, the information is displayed on a screen; aurally, it is provided via audio through speakers.

[0479] For example, a user might say to the robot, "Please check the date of the next meeting." The robot transcribes the voice into text and sends it to a server. The server performs language analysis, retrieves the date of the next meeting from its knowledge base, and returns that information to the robot. The robot displays the retrieved date on its screen and informs the user verbally, "The next meeting is next Wednesday."

[0480] An example of a prompt message would be, "When is the next meeting scheduled?" Upon receiving this instruction, the AI ​​model on the server would analyze the request and provide relevant information to the user.

[0481] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0482] Step 1:

[0483] The user gives instructions to the robot by voice. The robot receives the user's voice input through its built-in microphone. This voice is then passed on to the next step as is.

[0484] Step 2:

[0485] The device uses the Google Cloud Speech-to-Text API to convert the user's speech into text data. This process analyzes the speech input and generates a corresponding text string, making it possible to send the user's intent as text to the server.

[0486] Step 3:

[0487] The server passes the received text data to a language analysis system. This system uses OpenAI's generative AI model to analyze the text content and process the data to understand the user's intent. As a result, specific information requests become clear.

[0488] Step 4:

[0489] The server searches a knowledge base to retrieve information based on the analyzed intent. It uses SQL queries to extract relevant data from the internal database. This data includes specific answers to the user's requests.

[0490] Step 5:

[0491] The server generates a response in natural language based on the acquired information. It utilizes language generation technology to format the information in a way that is easy for the user to understand. Next, the formatted response data is sent to the terminal.

[0492] Step 6:

[0493] The terminal provides the user with text responses received from the server, both visually and audibly. The response content is displayed on the screen, and the information is transmitted audibly via the speaker using speech synthesis technology. This allows the user to easily obtain the requested information.

[0494] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0495] This invention aims to improve the user experience by integrating an emotion engine for emotion recognition into an agent system for human resource management, thereby personalizing responses to employee inquiries. This system consists of three main components: a server, a terminal, and a user.

[0496] The server, as the central hub of the system, handles all data processing and analysis. The server processes natural language queries received from terminals. These queries are first passed to the emotion engine, which analyzes the user's emotions. The emotion engine, for example, identifies the user's emotional state (e.g., frustration, joy, excitement) based on keywords and expressions in the user's text.

[0497] Based on these analysis results, the server gains a more accurate understanding of the user's intent through language analysis tools. After understanding the intent, the server searches the internal database to gather necessary information. Then, reflecting the results of the emotion engine, the generation tools construct a response with appropriate tone and content. For example, for a dissatisfied user, a more polite and empathetic response will be generated.

[0498] The generated natural language response is sent from the server to the terminal. The terminal displays the received response to the user, allowing the user to verify it through an appropriate interface.

[0499] The terminal functions as an interface connecting the user and the server. It not only sends user-inputted information to the server but also displays received responses appropriately. Furthermore, it can change the display style, message colors, fonts, and other elements on the terminal based on the results of the emotion engine.

[0500] Users make inquiries using natural language and obtain various internal company information through the responses. By benefiting from emotion recognition capabilities, users can receive responses that are tailored to their feelings. For example, if a user makes an emotionally charged inquiry such as, "Why is this procedure so complicated?", the system can recognize the user's frustration and provide a response that includes ways to simplify the procedure and support information.

[0501] This system provides employees with a means to obtain necessary information without stress, improving the quality of communication in talent management. Furthermore, it can contribute to the efficiency of HR operations.

[0502] The following describes the processing flow.

[0503] Step 1:

[0504] Users make inquiries using natural language through their devices. For example, they might enter a specific question such as, "What is the progress of this project?"

[0505] Step 2:

[0506] The device formats the user's inquiry received by the device and sends it to the server as an API request.

[0507] Step 3:

[0508] The server receives the inquiry from the terminal and passes it to the emotion engine. The emotion engine analyzes the text and identifies the user's emotional state (e.g., excitement, doubt, anxiety).

[0509] Step 4:

[0510] The server uses language analysis tools to analyze the user's inquiry and understand their precise intent. This process includes using AI models to deepen the understanding of the context.

[0511] Step 5:

[0512] Based on the server's understanding of the intent, it searches the internal database to retrieve relevant information. For example, it might retrieve the latest report data regarding the progress of a project.

[0513] Step 6:

[0514] The server takes the emotion recognition results into account and generates a response with a tone and content appropriate to the user using a generation method. If the emotion is determined to be anxiety, words of encouragement may be included.

[0515] Step 7:

[0516] The server sends the generated response to the terminal. This response incorporates language and information that takes the user's emotions into consideration.

[0517] Step 8:

[0518] The device receives the submitted response and displays it on the user interface. The display style (e.g., color and font) may also be adjusted depending on the emotional state.

[0519] Step 9:

[0520] The user reviews the answers displayed on their device. Based on the information provided, the user decides on their next course of action and plans their next steps in the process.

[0521] (Example 2)

[0522] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0523] Existing human resource management systems struggle to respond to user inquiries in a way that reflects individual emotions and intentions, limiting their ability to improve the user experience. There is a need for a system that can respond appropriately while considering user emotions.

[0524] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0525] In this invention, the server includes means for receiving natural language queries from users, language analysis means for analyzing the queries and understanding their intent and emotions, and means for retrieving relevant information from an information storage area based on the understood intent and emotions. This enables personalized responses that correspond to the user's emotional state.

[0526] "User" refers to a person who makes natural language queries to the system.

[0527] "Natural language" refers to the language that humans use on a daily basis, and is the text or sound that is processed by computer programs.

[0528] An "inquiry" refers to a request expressed in natural language that a user sends to a system to obtain necessary information or a response.

[0529] "Language analysis means" refers to a system that processes received natural language queries and analyzes the user's intentions and emotions.

[0530] "Intention" refers to the purpose and evaluation of what the user is seeking from the system through natural language.

[0531] "Emotions" refer to the emotional states included in user inquiries, such as joy or dissatisfaction, which are the subjects of analysis.

[0532] "Information storage area" refers to databases and data storage used to store various types of information necessary to answer inquiries.

[0533] "Generation means" refers to a mechanism that generates natural language to respond to users based on analysis results.

[0534] "Device" refers to a physical or virtual device used by a user to interface with a system, such as a smartphone or computer.

[0535] This invention aims to improve the user experience in a human resource management agent system by personalizing responses to user inquiries. The system consists of three main components: a server, a terminal, and a user.

[0536] The server functions as the central hub of this system, deeply analyzing the content of incoming inquiries. Specifically, the server analyzes the user's natural language inquiries received from terminals through an emotion engine. The emotion engine uses natural language processing techniques to identify the user's emotional state based on keywords and expressions contained in the analyzed text. Through this analysis, the server gains a more accurate understanding of the user's inquiry intent. After understanding the intent, the server searches for the necessary information from its information storage area and constructs a response in an appropriate tone using a generative AI model. For example, if a user asks a question indicating confusion, the server will empathize with that emotion and provide a polite explanation.

[0537] The terminal acts as an interface connecting the user and the server. The terminal is responsible for sending information entered by the user to the server and receiving and displaying responses sent from the server. During display, the style and format are adjusted to be visually appealing, taking into account the analysis results of the emotion engine.

[0538] Users can use their devices to make inquiries in natural language and obtain the necessary information through the responses. This system allows users to receive responses that are tailored to their emotional state, resulting in a more satisfying experience.

[0539] For example, if a user asks, "What is the progress of this project?", the server will use an emotion engine to analyze the user's anticipated concerns and anxieties, and then form a detailed response based on the ongoing progress information. An example of a prompt to the generative AI model might be, "Please create a detailed project progress report for an anxious user."

[0540] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0541] Step 1:

[0542] The user inputs a natural language query via a terminal. This input is in text format, and the terminal receives a specific question such as, "What is the progress of this project?" The terminal processes the input text data into a digital format and sends it to the server.

[0543] Step 2:

[0544] The server passes the inquiry received from the terminal to the emotion engine. The input is the user's natural language text, and the emotion engine processes this text to analyze the emotional state (e.g., anxiety or curiosity). As a result of the analysis, emotional state data is generated. This data is used as output in subsequent processes.

[0545] Step 3:

[0546] The server uses emotional state data obtained from the emotion engine to understand the user's intent through linguistic analysis. The input data consists of the emotional state and the original query text. Analysis extracts the user's intent (e.g., "I want to get project progress information"). This yields intent data as output for the next process.

[0547] Step 4:

[0548] The server searches the information storage area based on intent data and retrieves relevant information. The input is the user's intent data, and a query is created to search for the corresponding database entry. The output includes detailed information related to project progress.

[0549] Step 5:

[0550] The server uses the acquired information and emotional state data to generate a response using a generative AI model. The input includes relevant information and emotional state. The generative AI model processes this information and outputs a response in natural language that takes emotions into consideration. The output response is structured in a way that will be well-received by the user.

[0551] Step 6:

[0552] The server sends the generated response to the terminal. The terminal visually displays the received response to the user. The input data is the generated response, and the terminal formats this response to be displayed in a readable format. For example, it might adjust the font size and color and display it to the user as "The project is progressing as planned. Please rest assured."

[0553] (Application Example 2)

[0554] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0555] In autonomous vehicles, there is a need to provide not only information in response to passengers' questions and requests, but also optimal responses that are sensitive to passengers' emotions. However, conventional automated response systems have difficulty understanding passengers' emotional states, resulting in decreased satisfaction both financially and emotionally. Therefore, the challenge is to provide a system that can alleviate the anxiety and stress experienced by passengers in autonomous vehicles and improve the in-vehicle experience.

[0556] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0557] In this invention, the server includes means for analyzing natural language using an external emotion recognition structure to identify the user's emotional state, means for retrieving relevant information and suggestions from a database based on the identified intentions and emotional state, and means for generating a natural language response optimized according to the user's situation based on the retrieved information and suggestions. This enables personalized responses that respond to the emotions of passengers in an autonomous vehicle.

[0558] "Users" refers to passengers who use this system to obtain services and information within an autonomous vehicle.

[0559] "Emotional state" refers to information indicating the user's emotional situation, obtained as a result of the emotion recognition structure analyzing the user's natural language.

[0560] "External emotion recognition structure" refers to analytical techniques that use external artificial intelligence or algorithms to identify a user's emotions through natural language processing.

[0561] A "natural language response" is a sentence-form response that is generated based on the analyzed intent and emotions of the user, and has a consistent context.

[0562] A "database" is an information management system that stores and makes searchable past inquiry history, proposal information, and other related data.

[0563] An "autonomous vehicle" refers to a vehicle that can drive automatically and is operated by making its own decisions using artificial intelligence technology.

[0564] The system that realizes this invention consists of three main elements: a server, a terminal, and a user. The server receives the user's natural language query as its primary input and analyzes the text information contained in that query. For the analysis, it uses sentiment analysis tools such as the Google Cloud Natural Language API to identify the user's emotional state. Based on this sentiment analysis, the server uses generative AI models such as the OpenAI GPT API to generate a personalized natural language response that matches the user's emotions.

[0565] The response is then sent to the user's device. The device implements platform-specific UI / UX design to display the received response visually and clearly. For example, using Android or iOS, visual feedback is provided on the screen. Furthermore, fonts, colors, and text styling are adjusted as needed.

[0566] Ultimately, users can enjoy a reassuring experience in autonomous vehicles through the information and suggestions provided by this system. For example, if a user asks, "How far is it to the next service area?", and the server determines through sentiment analysis that the user is in a hurry, it will provide further information, including alternatives, in addition to the estimated arrival time.

[0567] Examples of prompt sentences include: "The passenger is indicating that they need to act quickly. Provide details related to location and distance, consider how you can reassure them, and generate an appropriate response."

[0568] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0569] Step 1:

[0570] The user enters a query in natural language into a terminal inside the vehicle. This input is converted into a format that can be sent from the terminal to the server. The terminal's input interface is designed to allow users to intuitively input text information.

[0571] Step 2:

[0572] The server uses the Google Cloud Natural Language API to process incoming queries and identify the user's emotional state. By analyzing the input text, it generates emotional information (e.g., feeling hurried, feeling anxious) as output based on keywords and context. This analysis result is stored for the next step.

[0573] Step 3:

[0574] Based on the obtained emotional state, the server uses the OpenAI GPT API to generate natural language responses appropriate to that emotional state. The model is executed using the emotional analysis results and prompt text as input, outputting response text in a style that meets the passenger's needs. This process generates emotionally sensitive tone and content.

[0575] Step 4:

[0576] The generated response message is sent from the server to the terminal. The terminal visually displays the response message on its user interface. This display uses fonts, colors, and layouts that are visually pleasing to the user.

[0577] Step 5:

[0578] The user reviews the displayed response and selects the next action as needed. Furthermore, if the provided information leads to further questions or requests, they can obtain new information by making another inquiry.

[0579] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0580] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0581] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0582] [Fourth Embodiment]

[0583] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0584] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0585] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0586] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0587] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0588] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0589] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0590] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0591] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0592] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0593] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0594] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0595] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0596] This invention is an agent system for supporting human resource management, which understands natural language, acquires, processes, and provides information to users, thereby assisting employees and HR departments within a company. Specific embodiments of this system are described below.

[0597] This system consists of servers, terminals, and users. This section details how each element functions.

[0598] The server is the core of this system, responsible for processing queries, acquiring data, and transmitting generated information. The server first receives queries sent from terminals. Next, it uses language analysis tools to process the query content using natural language processing and analyze the user's intent. External generative AI technologies can be incorporated into this analysis process, enabling highly accurate language understanding.

[0599] Furthermore, the server searches the internal database based on the analyzed intent and retrieves relevant information. This allows for the rapid collection of information necessary for the user. The retrieved information is organized into a user-friendly format and constructed as a response using a generation mechanism. This generated response is then sent to the terminal.

[0600] The terminal functions as an interface with the user. It transmits natural language queries entered by the user to the server, receives the response from the server, and displays it on the user interface. This allows the user to operate the system easily and intuitively and access the necessary information.

[0601] Users can use the system to ask questions related to company regulations and procedures and receive answers quickly. For example, if a user asks, "How do I apply for paid leave?", the user sends this question to the server via their terminal. The server then analyzes and retrieves information from the relevant company database, generates an answer, and displays it on the user's terminal.

[0602] This series of processes allows users to obtain accurate information quickly, thereby improving operational efficiency. In addition, it reduces the burden on the HR department of individually handling employee inquiries. In this way, the present invention provides innovative support for information management within a company.

[0603] The following describes the processing flow.

[0604] Step 1:

[0605] The user enters their inquiry in natural language through their device. For example, they might send a message such as, "How do I apply for leave?"

[0606] Step 2:

[0607] The terminal receives the inquiry and formats it as an API request to send to the server. This format includes the user's inquiry content and identification information.

[0608] Step 3:

[0609] The server receives an API request from the terminal. To analyze the received inquiry, it executes a language analysis method using a generative AI model.

[0610] Step 4:

[0611] The server performs language analysis to extract the intent of the query. In this process, the AI ​​understands the context and determines what the user is specifically asking for.

[0612] Step 5:

[0613] Based on the analyzed intent, the server searches the internal database for relevant data. For example, it queries the database to retrieve information related to leave requests.

[0614] Step 6:

[0615] The server is constructed using a generation method that organizes the acquired data into a user-friendly format and generates responses in natural language.

[0616] Step 7:

[0617] The server sends the generated response to the device. This response is returned to the user's device as an API response.

[0618] Step 8:

[0619] The terminal displays the response received from the server. The user interface visually presents the answer to the inquiry.

[0620] Step 9:

[0621] The user reviews the responses displayed on their device and uses them to determine tasks and actions. For example, they might follow the instructions for submitting a leave request.

[0622] (Example 1)

[0623] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0624] Within a company, responding quickly and accurately to employee inquiries is crucial for improving operational efficiency and optimizing information management. However, traditional human resource management systems struggle to accurately grasp employee intentions and provide relevant information appropriately, resulting in a significant amount of time and effort being required to respond to inquiries.

[0625] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0626] In this invention, the server includes a device for receiving inquiries from users, a language analysis device for analyzing the inquiries and understanding their intent, and a device for retrieving relevant information from an information storage device based on the understood intent. This enables the rapid and accurate acquisition and provision of information in response to user inquiries.

[0627] A "device for receiving inquiries from users" is a device that transmits inquiries made by users in natural language to the system.

[0628] A "language analysis device" is a device that analyzes received inquiries and understands the user's intent from their content.

[0629] An "information storage device" is a device that stores relevant information within a company and is used to hold data that is searchable.

[0630] A "generation device" is a device that generates natural language responses to be provided to users based on analyzed information.

[0631] A "display device" is a device used to visually present the generated response to the user.

[0632] A "terminal device" is a device used by users to input inquiries and receive and display responses from the server.

[0633] This agent system consists of servers, terminals, and users, and aims to improve information management and operational efficiency within a company by responding to user inquiries through natural language.

[0634] The server functions as the central hub of this system. The server receives inquiries sent from user terminals and processes their content using a language analysis device. An external generative AI model is used for this process. This model analyzes user inquiries and enables a high degree of understanding of intent. Based on the understood intent, the server retrieves relevant information from the information storage device, and the response is constructed by a generator. During this process, the information is formatted to be easily understood by the user.

[0635] The terminal provides an interface for users to input questions in natural language and receive and display answers from the server. This allows users to intuitively use the system and quickly obtain information.

[0636] Users can ask questions about company regulations and procedures through the system. For example, if a user asks, "How do I apply for paid leave?", they type the question on their terminal. The server analyzes the intent, retrieves relevant information, and generates a specific answer such as, "Download the application form from the company's portal site, get your supervisor's approval, and submit it to the Human Resources Department," which is then displayed on the terminal.

[0637] An example of a prompt message might be, "Please provide details about the company's remote work guidelines." This allows users to efficiently obtain information and streamline their daily work.

[0638] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0639] Step 1:

[0640] The user uses a device to input a request in natural language. For example, they might input something like, "Please tell me how to apply for paid leave." This input generates digital data that conveys the user's intent to the device.

[0641] Step 2:

[0642] The terminal sends the user's inputted question to the server. Here, the question, formalized as digital data, is transferred to the server. In this process, the terminal appropriately packets the data and transmits it according to the communication protocol.

[0643] Step 3:

[0644] The server receives data sent from the terminal and analyzes the query using a language analysis device. Specifically, it uses a generative AI model to scrutinize the input data and understand the user's intent with high accuracy. Through this analysis, the server obtains specific instructions on what to search for.

[0645] Step 4:

[0646] The server searches the information storage device based on the analyzed intent. It finds data that matches the intended information and extracts the relevant data. At this time, it uses database queries to quickly retrieve the information.

[0647] Step 5:

[0648] The server uses a generator to construct responses based on the acquired data. The information is converted into natural language and organized into a format that is easy for the user to understand. For example, specific guidance such as "Download the application form from the company's portal site, get your supervisor's approval, and submit it to the Human Resources Department" is generated.

[0649] Step 6:

[0650] The server sends the generated response to the terminal. The terminal receives this information and prepares to display it to the user. The data is transmitted again in digital format, and the terminal displays the received data correctly.

[0651] Step 7:

[0652] The device displays the answer to the user. The user can check the answer on the device screen and obtain the necessary information. This allows the user to receive an answer to their inquiry quickly.

[0653] (Application Example 1)

[0654] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0655] Existing human resource management systems have the drawback of being difficult for general users to use because they require specialized operation and search skills to efficiently obtain the information users need. Furthermore, the limited means of providing information visually and audibly results in insufficient speed and convenience of information retrieval. To address these issues, there is a need for a system that provides a simple and intuitive interface, such as that of a household robot, allowing users to instantly obtain the information they need.

[0656] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0657] In this invention, the server includes means for receiving instructions from the user as voice in natural language, voice recognition means for converting the instructions from voice to text, and language analysis means for analyzing the transcribed instructions and understanding the user's intent. As a result, users can request information by voice without requiring any special skills, and the information is provided via voice and display, enabling quick and convenient information retrieval.

[0658] "Natural language" refers to the language that humans use on a daily basis, and the technology that enables dialogue by converting it into a form that computers can understand.

[0659] "Voice recognition means" refers to a technological device that analyzes voice input from a user and converts it into text.

[0660] "Linguistic analysis methods" refer to technologies and processes for analyzing textual instructions and understanding their intent.

[0661] A "knowledge base" is a database that stores specific information and serves as a foundation for searching and retrieving information.

[0662] A "generative artificial intelligence model" is an artificial intelligence technology that automatically creates text, images, and other elements according to a generation task.

[0663] "Means of providing information visually" refers to technological devices that visually display obtained information to the user through a screen or other means.

[0664] To implement this invention, a system utilizing a home voice-enabled robot is required. The system consists of a microphone with voice recognition capabilities, a display and speaker, and a backend server. The server resides in the cloud and utilizes the Google Cloud Speech-to-Text API, and OpenAI's GPT model for language analysis. An internal company database is used as the knowledge base.

[0665] The server uses speech recognition to convert user instructions from speech to text. A language analysis tool utilizes a GPT model to understand the intent behind the transcribed instructions. Based on the intent obtained through this analysis, relevant information is retrieved from a knowledge base. The retrieved information is provided to the user visually and aurally. Visually, the information is displayed on a screen; aurally, it is provided via audio through speakers.

[0666] For example, a user might say to the robot, "Please check the date of the next meeting." The robot transcribes the voice into text and sends it to a server. The server performs language analysis, retrieves the date of the next meeting from its knowledge base, and returns that information to the robot. The robot displays the retrieved date on its screen and informs the user verbally, "The next meeting is next Wednesday."

[0667] An example of a prompt message would be, "When is the next meeting scheduled?" Upon receiving this instruction, the AI ​​model on the server would analyze the request and provide relevant information to the user.

[0668] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0669] Step 1:

[0670] The user gives instructions to the robot by voice. The robot receives the user's voice input through its built-in microphone. This voice is then passed on to the next step as is.

[0671] Step 2:

[0672] The device uses the Google Cloud Speech-to-Text API to convert the user's speech into text data. This process analyzes the speech input and generates a corresponding text string, making it possible to send the user's intent as text to the server.

[0673] Step 3:

[0674] The server passes the received text data to a language analysis system. This system uses OpenAI's generative AI model to analyze the text content and process the data to understand the user's intent. As a result, specific information requests become clear.

[0675] Step 4:

[0676] The server searches a knowledge base to retrieve information based on the analyzed intent. It uses SQL queries to extract relevant data from the internal database. This data includes specific answers to the user's requests.

[0677] Step 5:

[0678] The server generates a response in natural language based on the acquired information. It utilizes language generation technology to format the information in a way that is easy for the user to understand. Next, the formatted response data is sent to the terminal.

[0679] Step 6:

[0680] The terminal provides the user with text responses received from the server, both visually and audibly. The response content is displayed on the screen, and the information is transmitted audibly via the speaker using speech synthesis technology. This allows the user to easily obtain the requested information.

[0681] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0682] This invention aims to improve the user experience by integrating an emotion engine for emotion recognition into an agent system for human resource management, thereby personalizing responses to employee inquiries. This system consists of three main components: a server, a terminal, and a user.

[0683] The server, as the central hub of the system, handles all data processing and analysis. The server processes natural language queries received from terminals. These queries are first passed to the emotion engine, which analyzes the user's emotions. The emotion engine, for example, identifies the user's emotional state (e.g., frustration, joy, excitement) based on keywords and expressions in the user's text.

[0684] Based on these analysis results, the server gains a more accurate understanding of the user's intent through language analysis tools. After understanding the intent, the server searches the internal database to gather necessary information. Then, reflecting the results of the emotion engine, the generation tools construct a response with appropriate tone and content. For example, for a dissatisfied user, a more polite and empathetic response will be generated.

[0685] The generated natural language response is sent from the server to the terminal. The terminal displays the received response to the user, allowing the user to verify it through an appropriate interface.

[0686] The terminal functions as an interface connecting the user and the server. It not only sends user-inputted information to the server but also displays received responses appropriately. Furthermore, it can change the display style, message colors, fonts, and other elements on the terminal based on the results of the emotion engine.

[0687] Users make inquiries using natural language and obtain various internal company information through the responses. By benefiting from emotion recognition capabilities, users can receive responses that are tailored to their feelings. For example, if a user makes an emotionally charged inquiry such as, "Why is this procedure so complicated?", the system can recognize the user's frustration and provide a response that includes ways to simplify the procedure and support information.

[0688] This system provides employees with a means to obtain necessary information without stress, improving the quality of communication in talent management. Furthermore, it can contribute to the efficiency of HR operations.

[0689] The following describes the processing flow.

[0690] Step 1:

[0691] Users make inquiries using natural language through their devices. For example, they might enter a specific question such as, "What is the progress of this project?"

[0692] Step 2:

[0693] The device formats the user's inquiry received by the device and sends it to the server as an API request.

[0694] Step 3:

[0695] The server receives the inquiry from the terminal and passes it to the emotion engine. The emotion engine analyzes the text and identifies the user's emotional state (e.g., excitement, doubt, anxiety).

[0696] Step 4:

[0697] The server uses language analysis tools to analyze the user's inquiry and understand their precise intent. This process includes using AI models to deepen the understanding of the context.

[0698] Step 5:

[0699] Based on the server's understanding of the intent, it searches the internal database to retrieve relevant information. For example, it might retrieve the latest report data regarding the progress of a project.

[0700] Step 6:

[0701] The server takes the emotion recognition results into account and generates a response with a tone and content appropriate to the user using a generation method. If the emotion is determined to be anxiety, words of encouragement may be included.

[0702] Step 7:

[0703] The server sends the generated response to the terminal. This response incorporates language and information that takes the user's emotions into consideration.

[0704] Step 8:

[0705] The device receives the submitted response and displays it on the user interface. The display style (e.g., color and font) may also be adjusted depending on the emotional state.

[0706] Step 9:

[0707] The user reviews the answers displayed on their device. Based on the information provided, the user decides on their next course of action and plans their next steps in the process.

[0708] (Example 2)

[0709] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0710] Existing human resource management systems struggle to respond to user inquiries in a way that reflects individual emotions and intentions, limiting their ability to improve the user experience. There is a need for a system that can respond appropriately while considering user emotions.

[0711] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0712] In this invention, the server includes means for receiving natural language queries from users, language analysis means for analyzing the queries and understanding their intent and emotions, and means for retrieving relevant information from an information storage area based on the understood intent and emotions. This enables personalized responses that correspond to the user's emotional state.

[0713] "User" refers to a person who makes natural language queries to the system.

[0714] "Natural language" refers to the language that humans use on a daily basis, and is the text or sound that is processed by computer programs.

[0715] An "inquiry" refers to a request expressed in natural language that a user sends to a system to obtain necessary information or a response.

[0716] "Language analysis means" refers to a system that processes received natural language queries and analyzes the user's intentions and emotions.

[0717] "Intention" refers to the purpose and evaluation of what the user is seeking from the system through natural language.

[0718] "Emotions" refer to the emotional states included in user inquiries, such as joy or dissatisfaction, which are the subjects of analysis.

[0719] "Information storage area" refers to databases and data storage used to store various types of information necessary to answer inquiries.

[0720] "Generation means" refers to a mechanism that generates natural language to respond to users based on analysis results.

[0721] "Device" refers to a physical or virtual device used by a user to interface with a system, such as a smartphone or computer.

[0722] This invention aims to improve the user experience in a human resource management agent system by personalizing responses to user inquiries. The system consists of three main components: a server, a terminal, and a user.

[0723] The server functions as the central hub of this system, deeply analyzing the content of incoming inquiries. Specifically, the server analyzes the user's natural language inquiries received from terminals through an emotion engine. The emotion engine uses natural language processing techniques to identify the user's emotional state based on keywords and expressions contained in the analyzed text. Through this analysis, the server gains a more accurate understanding of the user's inquiry intent. After understanding the intent, the server searches for the necessary information from its information storage area and constructs a response in an appropriate tone using a generative AI model. For example, if a user asks a question indicating confusion, the server will empathize with that emotion and provide a polite explanation.

[0724] The terminal acts as an interface connecting the user and the server. The terminal is responsible for sending information entered by the user to the server and receiving and displaying responses sent from the server. During display, the style and format are adjusted to be visually appealing, taking into account the analysis results of the emotion engine.

[0725] Users can use their devices to make inquiries in natural language and obtain the necessary information through the responses. This system allows users to receive responses that are tailored to their emotional state, resulting in a more satisfying experience.

[0726] For example, if a user asks, "What is the progress of this project?", the server will use an emotion engine to analyze the user's anticipated concerns and anxieties, and then form a detailed response based on the ongoing progress information. An example of a prompt to the generative AI model might be, "Please create a detailed project progress report for an anxious user."

[0727] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0728] Step 1:

[0729] The user inputs a natural language query via a terminal. This input is in text format, and the terminal receives a specific question such as, "What is the progress of this project?" The terminal processes the input text data into a digital format and sends it to the server.

[0730] Step 2:

[0731] The server passes the inquiry received from the terminal to the emotion engine. The input is the user's natural language text, and the emotion engine processes this text to analyze the emotional state (e.g., anxiety or curiosity). As a result of the analysis, emotional state data is generated. This data is used as output in subsequent processes.

[0732] Step 3:

[0733] The server uses emotional state data obtained from the emotion engine to understand the user's intent through linguistic analysis. The input data consists of the emotional state and the original query text. Analysis extracts the user's intent (e.g., "I want to get project progress information"). This yields intent data as output for the next process.

[0734] Step 4:

[0735] The server searches the information storage area based on intent data and retrieves relevant information. The input is the user's intent data, and a query is created to search for the corresponding database entry. The output includes detailed information related to project progress.

[0736] Step 5:

[0737] The server uses the acquired information and emotional state data to generate a response using a generative AI model. The input includes relevant information and emotional state. The generative AI model processes this information and outputs a response in natural language that takes emotions into consideration. The output response is structured in a way that will be well-received by the user.

[0738] Step 6:

[0739] The server sends the generated response to the terminal. The terminal visually displays the received response to the user. The input data is the generated response, and the terminal formats this response to be displayed in a readable format. For example, it might adjust the font size and color and display it to the user as "The project is progressing as planned. Please rest assured."

[0740] (Application Example 2)

[0741] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0742] In autonomous vehicles, there is a need to provide not only information in response to passengers' questions and requests, but also optimal responses that are sensitive to passengers' emotions. However, conventional automated response systems have difficulty understanding passengers' emotional states, resulting in decreased satisfaction both financially and emotionally. Therefore, the challenge is to provide a system that can alleviate the anxiety and stress experienced by passengers in autonomous vehicles and improve the in-vehicle experience.

[0743] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0744] In this invention, the server includes means for analyzing natural language using an external emotion recognition structure to identify the user's emotional state, means for retrieving relevant information and suggestions from a database based on the identified intentions and emotional state, and means for generating a natural language response optimized according to the user's situation based on the retrieved information and suggestions. This enables personalized responses that respond to the emotions of passengers in an autonomous vehicle.

[0745] "Users" refers to passengers who use this system to obtain services and information within an autonomous vehicle.

[0746] "Emotional state" refers to information indicating the user's emotional situation, obtained as a result of the emotion recognition structure analyzing the user's natural language.

[0747] "External emotion recognition structure" refers to analytical techniques that use external artificial intelligence or algorithms to identify a user's emotions through natural language processing.

[0748] A "natural language response" is a sentence-form response that is generated based on the analyzed intent and emotions of the user, and has a consistent context.

[0749] A "database" is an information management system that stores and makes searchable past inquiry history, proposal information, and other related data.

[0750] An "autonomous vehicle" refers to a vehicle that can drive automatically and is operated by making its own decisions using artificial intelligence technology.

[0751] The system that realizes this invention consists of three main elements: a server, a terminal, and a user. The server receives the user's natural language query as its primary input and analyzes the text information contained in that query. For the analysis, it uses sentiment analysis tools such as the Google Cloud Natural Language API to identify the user's emotional state. Based on this sentiment analysis, the server uses generative AI models such as the OpenAI GPT API to generate a personalized natural language response that matches the user's emotions.

[0752] The response is then sent to the user's device. The device implements platform-specific UI / UX design to display the received response visually and clearly. For example, using Android or iOS, visual feedback is provided on the screen. Furthermore, fonts, colors, and text styling are adjusted as needed.

[0753] Ultimately, users can enjoy a reassuring experience in autonomous vehicles through the information and suggestions provided by this system. For example, if a user asks, "How far is it to the next service area?", and the server determines through sentiment analysis that the user is in a hurry, it will provide further information, including alternatives, in addition to the estimated arrival time.

[0754] Examples of prompt sentences include: "The passenger is indicating that they need to act quickly. Provide details related to location and distance, consider how you can reassure them, and generate an appropriate response."

[0755] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0756] Step 1:

[0757] The user enters a query in natural language into a terminal inside the vehicle. This input is converted into a format that can be sent from the terminal to the server. The terminal's input interface is designed to allow users to intuitively input text information.

[0758] Step 2:

[0759] The server uses the Google Cloud Natural Language API to process incoming queries and identify the user's emotional state. By analyzing the input text, it generates emotional information (e.g., feeling hurried, feeling anxious) as output based on keywords and context. This analysis result is stored for the next step.

[0760] Step 3:

[0761] Based on the obtained emotional state, the server uses the OpenAI GPT API to generate natural language responses appropriate to that emotional state. The model is executed using the emotional analysis results and prompt text as input, outputting response text in a style that meets the passenger's needs. This process generates emotionally sensitive tone and content.

[0762] Step 4:

[0763] The generated response message is sent from the server to the terminal. The terminal visually displays the response message on its user interface. This display uses fonts, colors, and layouts that are visually pleasing to the user.

[0764] Step 5:

[0765] The user reviews the displayed response and selects the next action as needed. Furthermore, if the provided information leads to further questions or requests, they can obtain new information by making another inquiry.

[0766] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0767] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0768] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0769] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0770] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0771] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0772] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0773] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0774] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0775] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0776] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0777] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0778] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0779] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0780] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0781] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0782] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0783] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0784] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0785] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0786] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0787] The following is further disclosed regarding the embodiments described above.

[0788] (Claim 1)

[0789] In an agent system for managing human resources,

[0790] A means of receiving natural language inquiries from users,

[0791] A language analysis means for analyzing the aforementioned inquiry and understanding its intent,

[0792] Based on the understood intent, a means of searching for relevant information from the company's internal database,

[0793] A generation method for generating natural language responses to the user based on searched information,

[0794] A means for sending the generated response to the user's terminal,

[0795] A system that includes this.

[0796] (Claim 2)

[0797] The system according to claim 1, wherein the language analysis means analyzes the query using an external generative artificial intelligence model.

[0798] (Claim 3)

[0799] The system according to claim 1, wherein the retrieved information is identified based on the user's unique information.

[0800] "Example 1"

[0801] (Claim 1)

[0802] A device that receives inquiries from users,

[0803] A language analysis device that analyzes the aforementioned inquiry and understands its intent,

[0804] A device that retrieves relevant information from an information storage device based on the understood intent,

[0805] A generator that generates natural language responses for the user based on the searched information,

[0806] A device that transmits the generated response to the user's display device,

[0807] A terminal device that receives and displays user inquiries and generated responses,

[0808] A system that includes this.

[0809] (Claim 2)

[0810] The system according to claim 1, wherein the language analysis device analyzes the query using an external generative artificial intelligence model.

[0811] (Claim 3)

[0812] The system according to claim 1, wherein the retrieved information is identified based on user identification information.

[0813] "Application Example 1"

[0814] (Claim 1)

[0815] A means of receiving user instructions as voice using natural language,

[0816] A speech recognition means for converting the aforementioned instructions from speech to text,

[0817] A language analysis means for analyzing the transcribed instructions and understanding their intent,

[0818] Based on the understood intent, a means of using a knowledge base to search for information,

[0819] A means for generating a response in natural language based on the aforementioned searched information,

[0820] Means for providing the generated response in audio and visual formats,

[0821] A system that includes this.

[0822] (Claim 2)

[0823] The system according to claim 1, wherein the language analysis means analyzes the instructions using an external generative artificial intelligence model.

[0824] (Claim 3)

[0825] The system according to claim 1, wherein the retrieved information is identified based on the user's attribute information.

[0826] "Example 2 of combining an emotion engine"

[0827] (Claim 1)

[0828] A means of receiving natural language inquiries from users,

[0829] A language analysis means for analyzing the aforementioned inquiry and understanding the intent and emotion,

[0830] A means for retrieving relevant information from an information storage area based on understood intentions and emotions,

[0831] A generation method that generates natural language responses based on the searched information and according to the user's emotional state,

[0832] A means for transmitting and displaying the generated response on the user's device,

[0833] A system that includes this.

[0834] (Claim 2)

[0835] The system according to claim 1, wherein the language analysis means and generation means utilize an external generative artificial intelligence model to analyze the query and generate a response.

[0836] (Claim 3)

[0837] The system according to claim 1, wherein the retrieved information is identified based on the user's specific attributes and reflects the results of sentiment analysis.

[0838] "Application example 2 when combining with an emotional engine"

[0839] (Claim 1)

[0840] A means of analyzing natural language using an external emotion recognition structure in order to identify the emotional state of the user,

[0841] A means of searching for relevant information and suggestions from a database based on the understood intentions and emotional state,

[0842] A means for generating natural language responses optimized according to the user's situation, based on the searched information and suggested content,

[0843] A means for transmitting the generated response to be displayed on the user's communication terminal,

[0844] A system that includes this.

[0845] (Claim 2)

[0846] The system according to claim 1, wherein the external emotion recognition structure is supported by a generative artificial intelligence module.

[0847] (Claim 3)

[0848] The system according to claim 1, wherein the generated response is specifically adjusted for the passenger support function of an automated vehicle. [Explanation of Symbols]

[0849] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of receiving user instructions as voice using natural language, A speech recognition means for converting the aforementioned instructions from speech to text, A language analysis means for analyzing the transcribed instructions and understanding their intent, Based on the understood intent, a means of using a knowledge base to search for information, A means for generating a response in natural language based on the aforementioned searched information, Means for providing the generated response in audio and visual formats, A system that includes this.

2. The system according to claim 1, wherein the language analysis means analyzes the instructions using an external generative artificial intelligence model.

3. The system according to claim 1, wherein the retrieved information is identified based on the user's attribute information.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A