system

The information processing system addresses the challenge of understanding diverse religious texts by using natural language processing to generate annotations, facilitating efficient and secure multilingual insights.

JP2026071018APending Publication Date: 2026-04-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing systems lack an efficient method to quickly process and provide relevant annotations across diverse religious texts, requiring manual effort and failing to facilitate consistent understanding across various religions and languages.

Method used

An information processing system that utilizes natural language processing to search for relevant passages and generate annotations from user keywords, providing insights from multiple perspectives.

Benefits of technology

Enables users to efficiently understand religious texts from multiple angles by automatically generating annotations, supporting multilingual and secure data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071018000001_ABST
    Figure 2026071018000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A terminal means for receiving keywords entered by the user, A server means for searching for related document sets based on the aforementioned keywords, A terminal means that displays the search results obtained by the server means and automatically generates related annotations, An information processing system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Religious texts are diverse, and it takes time and effort to manually obtain and understand their vast information. In particular, it is very difficult to find relevant passages among multiple documents based on a specific theme and compare and understand them with each other. Also, in this process, a consistent understanding across various religions and languages is required, but there is no existing efficient system that meets this need. Therefore, there is a need for a system that can quickly process this online and provide relevant annotations.

Means for Solving the Problems

[0005] To solve the above problems, the present invention provides an information processing system including a terminal and a server that receive keywords entered by a user and quickly search for related document sets based on those keywords. This system uses natural language processing technology to highly analyze the relationships within the document set and provides the user with relevant passages. In addition, the server automatically generates annotations and presents them to the user clearly as information including multiple perspectives. As a result, the user can deepen their understanding of the target topic from multiple angles and efficiently.

[0006] A "user" is an entity that uses an information processing system to search for specific information and retrieve the results.

[0007] A "keyword" is a string of characters that a user enters as the starting point for a search, and it is an element that serves as a criterion for identifying related document groups.

[0008] "Terminal means" refers to the functions of the device or software that a user uses to input keywords and receive search results.

[0009] A "server means" is a central processing unit that searches a set of documents based on received keywords, generates results, and transmits them to a terminal means.

[0010] A "document collection" is a data set that includes religious and other literature, and is a collection of information that is searched for.

[0011] "Natural language processing" is a technology that enables computers to understand and interpret language data, and it is a method used to improve the accuracy of searches.

[0012] "Automatic generation" refers to the process of creating annotations using a program without human intervention.

[0013] Annotations are explanations or descriptions added to search results that help users understand the information provided.

[0014] An "information processing system" is a general term for a group of devices and software that control digital data input / output and calculations. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention is implemented as an information processing system. The user connects to the system via a terminal and inputs keywords for a search. The terminal processes the input keywords and sends the data to the server. The server uses the received data to search for documents related to the keywords. The search process employed here utilizes natural language processing technology to efficiently extract relevant information from the document set.

[0037] The server returns the search results to the terminal as a set of data. This result includes passages related to the user's specified theme, with automatically generated annotations that supplement their meaning and background as needed. This annotation generation utilizes pre-set rules and a knowledge base, enabling interpretation from multiple perspectives.

[0038] The terminal receives data sent from the server and displays it to the user. The information is organized, highlighted, and visually supported to facilitate user comprehension. This allows the user to gain knowledge based on multiple religious and philosophical perspectives.

[0039] As a concrete example, if a user enters the keyword "compassion" into their device, the server receives it and searches its database for relevant documents. For instance, it might extract passages related to "compassion" from the scriptures and related literature of a particular religion. The server also generates related annotations explaining how the concept of compassion is understood in different religions. This information is clearly presented on the device, allowing the user to gain deeper insights.

[0040] This system is multilingual, providing value to users from diverse cultural backgrounds. Furthermore, the server ensures data integrity and privacy through secure connections, providing a safe environment for users to search for information. This makes it an indispensable tool for religious and philosophical education settings and individual researchers.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The user enters a specific keyword using the terminal. The terminal prepares to store this input internally and triggers its transmission to the server.

[0044] Step 2:

[0045] The terminal sends the keywords entered by the user to the server. At this time, the data is appropriately packaged according to the format and protocol of the transmitted data.

[0046] Step 3:

[0047] The server analyzes the received keywords and initiates a process to search for related documents. This uses natural language processing techniques to analyze the context of the keywords and perform the search.

[0048] Step 4:

[0049] The server selects passages extracted from relevant documents as search results. These results are then formatted to be easily understood by the user.

[0050] Step 5:

[0051] The server automatically generates annotations, adding background information and interpretations relevant to the search results. Using AI technology, it generates annotations that include multiple perspectives.

[0052] Step 6:

[0053] The server sends the formatted search results and generated annotations to the terminal. At this time, it verifies the integrity and security of the data and ensures that the transmission is completed successfully.

[0054] Step 7:

[0055] The terminal receives information from the server and displays it to the user. The information is easy to read and visually organized, allowing the user to immediately utilize the presented information.

[0056] (Example 1)

[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0058] Conventional information retrieval systems make it difficult for users to effectively search for specific information and obtain additional information that includes various perspectives. Therefore, there is a challenge in that it is difficult for users to acquire knowledge from diverse viewpoints.

[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0060] In this invention, the server includes information terminal means for a user to input information, computing device means for processing the information and searching for related data, and information terminal means for displaying the information searched by the computing device means and automatically generating related additional information. This makes it possible for the user to search for information from various perspectives and efficiently obtain related additional information.

[0061] An "information terminal means" is a device for users to input information, and typically provides an interface for sending data to a server.

[0062] A "computation device" is a device that processes received information and searches for and analyzes related data. This includes functions that efficiently handle information using language processing technology.

[0063] "Additional information" refers to information automatically generated in relation to search results, intended to deepen the user's understanding by providing multiple perspectives.

[0064] This invention aims to achieve efficient data processing and information provision in an information retrieval system. First, the user uses a terminal to input information and enters what they want to search for. For example, they can enter the text "What is compassion?". The terminal receives this input and sends the data to the server in an appropriate format.

[0065] The server performs computational processing based on the received data. This processing includes searching for and analyzing relevant information using natural language processing techniques. Specifically, it uses algorithms such as TF-IDF and BERT to extract the most relevant information from the database for the input keywords. It can also utilize generative AI models to create additional information and annotations related to the search results.

[0066] The generated information is sent to the device and displayed to the user. The device organizes this information and highlights particularly important parts. Furthermore, to provide information from multiple perspectives, it is possible to include different interpretations of additional information. This allows users to gain deeper insights, not just acquire data.

[0067] As a concrete example, a prompt message for a user to input the keyword "compassion" might be: "Use natural language processing to search for how the concept of compassion is interpreted in different religions, extract relevant documents, and generate explanations." In this way, users can easily acquire a wide range of knowledge and different perspectives.

[0068] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0069] Step 1:

[0070] The user uses a terminal to input information. The user inputs keywords in text format, such as words like "compassion." The input data is formatted by the terminal and prepared for transmission to the server. This format often uses JSON.

[0071] Step 2:

[0072] The terminal sends keyword data entered by the user to the server. This communication is conducted via HTTP requests, ensuring a secure connection. The entered data arrives at the server, where analysis begins.

[0073] Step 3:

[0074] The server performs a document search within the database based on the received keywords. This search utilizes natural language processing techniques, employing algorithms such as TF-IDF and BERT to extract highly relevant information. The input is the search keywords, and the output is a list of related documents.

[0075] Step 4:

[0076] The server generates annotations related to the search results. A generative AI model is activated, referencing a knowledge base to interpret information from multiple perspectives and create annotations. The input is a list of search result documents, and the output is annotations and additional information for those documents.

[0077] Step 5:

[0078] The server sends the organized search results and generated annotations to the terminal. The data is again sent in JSON format or similar, and communication takes place via a secure protocol. This ensures data privacy and integrity.

[0079] Step 6:

[0080] The terminal receives information from the server and prepares it for display to the user. At this stage, the information is highlighted and visual effects are added to make it easier for the user to understand. The displayed content includes the searched information and its annotations, allowing the user to gain deeper knowledge based on it.

[0081] (Application Example 1)

[0082] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0083] In modern society, there is a need for information processing systems that can understand information from diverse perspectives and reference it quickly and accurately. However, existing systems lack sufficient means to visually present and aid in the understanding of the multifaceted information that users seek. Therefore, there is a need to provide a system that makes it easy for users to obtain visually emphasized information and deepen their understanding from multiple perspectives.

[0084] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0085] In this invention, the server includes a device means for receiving a concept input by a user, a computer means for searching for a set of related information based on the concept, and a device means for presenting the results searched by the computer means and automatically generating related supplementary information. This makes it possible to deepen understanding from multiple perspectives through visually highlighted information.

[0086] A "user" is an entity that inputs information and receives output from a system.

[0087] A "concept" is a collection of information that a user inputs for a system to interpret and process.

[0088] "Reception" refers to the process by which a device takes in user input into the system.

[0089] "Device means" refers to hardware or software components that receive user input and present system output to the user.

[0090] "Searching" is the process by which a computer finds relevant information from a database or knowledge base.

[0091] An "information set" is a collection of data and knowledge related to user input.

[0092] "Computing means" refers to a computing device or server system for searching and processing a set of information.

[0093] "Presentation" refers to the act of a device visually showing information or results to a user.

[0094] "Supplemental information" refers to explanations and annotations added to the information that has been searched, and is data that helps users to understand the information more deeply.

[0095] "Automatic generation" refers to the process by which a system creates additional information or annotations without human intervention.

[0096] "Visual display means" refers to display devices or interfaces used to visually present information to the user.

[0097] To realize this application, the system consists of multiple elements, including a device, a computer system, and a visual display device.

[0098] First, smart glasses are used as a device for the user to input concepts. This device receives concepts from the user using voice input technology. The received information is then transmitted to a server via the internet.

[0099] Next, the server functions as a computing system and employs machine learning techniques. Specifically, it uses the Google® Cloud Natural Language API to quickly search for information based on user-inputted concepts. The server retrieves relevant data from databases and automatically generates supplementary information from Wikipedia and other knowledge bases as needed.

[0100] The server then sends the search results back to the smart glasses. The smart glasses act as a visual display device, providing the user with visually highlighted information. This allows the user to gain a deeper understanding from multiple perspectives.

[0101] As a concrete example, a user might want to learn about "non-violence" and use voice input into smart glasses. The server searches for relevant philosophical and religious documents and provides automatically generated annotations from various perspectives. Visual presentations allow the user to compare different viewpoints and deepen their understanding.

[0102] An example of a prompt sentence generated using an AI model is: "Summarize and present interpretations of nonviolence from the perspectives of Buddhism, Hinduism, and modern ethics."

[0103] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0104] Step 1:

[0105] Users provide voice input through smart glasses. Specifically, they specify keywords or concepts they are interested in by voice. The smart glasses use voice recognition technology to convert the input into text data and send that text data to a server via the internet. The input data is the voice keywords specified by the user, and the output is the keywords in text format sent to the server.

[0106] Step 2:

[0107] The server uses machine learning techniques to search for relevant information based on the received text data. It uses the Google Cloud Natural Language API to explore document sets related to the input keywords. Data processing includes keyword-based information retrieval and extraction of important sentences. The input is keywords in text format, and the output is the searched set of relevant information.

[0108] Step 3:

[0109] The server automatically generates additional supplementary information based on the search results. It utilizes a knowledge base to perform data calculations that form annotations from multiple perspectives. Specifically, it collects data from Wikipedia and similar sources to generate supplementary information that highlights different interpretations. The input is the searched related information, and the output is the annotated information.

[0110] Step 4:

[0111] The server sends the generated annotated information back to the smart glasses. The user can view the received information in real time through the visual highlighting function. Specifically, the smart glasses display relevant information on the HUD in the user's field of view, visualizing highlighted text and graphics. The input is the annotated information from the server, and the output is the information displayed to the user in an easy-to-read format.

[0112] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0113] This invention provides information tailored to the user's emotional state by combining an emotion engine with an information processing system. The user inputs keywords using a terminal, and this information is transmitted to a server via the terminal. The server searches for related documents based on the aforementioned keywords and automatically extracts relevant passages using natural language processing technology.

[0114] Here, an emotion engine is added, introducing an additional process to analyze the user's emotional state. Based on the user's input patterns and operation history, the emotion engine infers whether the user is experiencing a particular emotion. Based on this inference, the server adjusts the content and presentation method of search results and automatically generated annotations.

[0115] For example, if a user searches for "comfort," the emotion engine will determine that the user is feeling down. The server will then prioritize extracting passages with encouraging and comforting themes, and further highlight their content. The emotion engine supports customization to more effectively deliver information that matches the user's emotions.

[0116] This system supports multiple languages, offering global applicability. In terms of security, communication between the server and the terminal is encrypted, ensuring user privacy while providing accurate and personalized information.

[0117] With the system configuration described above, the present invention not only provides educational value to religious educational institutions and research institutions, but also enables the provision of information that takes into account the deep learning and individual spiritual needs of individual believers and learners.

[0118] The following describes the processing flow.

[0119] Step 1:

[0120] The user enters a specific keyword through the terminal. The terminal temporarily records the input and then prepares to transfer it to the server.

[0121] Step 2:

[0122] The terminal sends the recorded keywords to the server and packages them according to a protocol that ensures the integrity and security of the communication.

[0123] Step 3:

[0124] The server analyzes the received keywords and uses them as matching criteria for searching the document set. Natural language processing techniques are applied to efficiently extract relevant passages from the document set.

[0125] Step 4:

[0126] The emotion engine analyzes user input and past operation history to estimate the emotions the user may be experiencing. It then determines an information provision policy that takes this emotional state into account.

[0127] Step 5:

[0128] The server adjusts the retrieved search results and annotations based on the analysis results from the sentiment engine. Specifically, it performs processes such as highlighting passages that correspond to the user's emotions.

[0129] Step 6:

[0130] The server sends search results and adjusted annotations to the terminal. This process verifies data integrity while enabling immediate display to the user.

[0131] Step 7:

[0132] The device accurately receives information from the server and presents the data in a visually organized format that is easy for the user to understand. This allows the user to obtain information that corresponds to their individual emotional state.

[0133] (Example 2)

[0134] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0135] In modern information processing systems, providing information tailored to the user's emotional state is challenging. Conventional information retrieval systems often provide uniform information without considering the user's psychological state, resulting in a failure to meet the user's essential needs. This invention aims to solve this problem by developing a system that understands the user's emotions and provides information accordingly.

[0136] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0137] In this invention, the server includes means for receiving concepts input by the user, a computer for searching for relevant information resources based on the concepts, and an analysis device for inferring the user's emotional state. This allows for the adjustment of information provided according to the user's emotional state, enabling personalized information delivery.

[0138] A "device" is a machine or a similar system designed to perform a specific function.

[0139] A "concept" is a word, phrase, or collection of words or phrases used to communicate a particular thought or piece of information to others.

[0140] "Information resources" refer to data and knowledge that exist in various forms and may contain the information that users are looking for.

[0141] A "computer" is an electronic device or system used to process data and perform calculations.

[0142] "Notes" are additional explanations or explanatory texts accompanying information or data, provided to aid understanding.

[0143] An "analytical device" is a system that analyzes data and information and uses the results to draw specific conclusions or make inferences.

[0144] "Emotional state" refers to an internal state that indicates a user's psychological and emotional tendencies and mood.

[0145] This invention is an information processing system that provides information tailored to the user's emotional state. The specific implementation method is described below.

[0146] The user first inputs a specific concept using a terminal. This concept is in text format and includes keywords related to what the user wants to research or the information they need. Upon receiving the input, the terminal sends the information to the server using an encrypted communication protocol (e.g., SSL / TLS) to maintain security.

[0147] The server uses a computer to find relevant data from its internally held information resources based on the received concepts. The software technology used here employs models with natural language processing capabilities (e.g., generative AI models such as BERT and GPT). Leveraging these models, it has the ability to automatically extract relevant information.

[0148] Furthermore, the server uses an analysis device called an emotion engine to analyze the user's emotional state. This involves a process of evaluating the user's past activity history and current input data to identify specific emotions. For example, if the concept "seeking encouragement" is input, it is determined that the user's input indicates a depressed state.

[0149] Based on the results of this sentiment analysis, the server adjusts the priority and content of the information presented to the user. This adjustment includes processing the selected information to optimize it for the user's state.

[0150] For example, if a user enters a specific concept seeking "comfort," the server's emotion engine infers that the user is feeling down. The server then uses natural language processing technology to prioritize extracting passages and articles that provide a sense of reassurance and improve the user's mood, and sends that content to the device.

[0151] Through this process, users can receive information optimized to align with their own emotional state. An example of a prompt using a generative AI model is: "Based on the concept entered by user A, please generate optimized information while considering their emotional state."

[0152] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0153] Step 1:

[0154] The user uses the terminal to input concepts related to what they want to research. The input concepts are displayed on the terminal screen and simultaneously processed as digital data. The input format is regular text, which is then prepared to be sent to the server as the initial data.

[0155] Step 2:

[0156] The terminal transmits the input concept to the server over the network. During this process, the data is encrypted using encryption protocols such as SSL / TLS to ensure security during transmission. The output of this process is encrypted concept data.

[0157] Step 3:

[0158] The server decrypts the received encrypted data and extracts the original concept. Based on this concept, the server launches a search engine to find related data from information resources. Here, the input is the decrypted concept, and the output is the identifier of the related information resource.

[0159] Step 4:

[0160] The server uses a generative AI model equipped with natural language processing technology to automatically extract appropriate passages from relevant information resources. The input is the identifier of the searched information resource, and the output is an optimized list of passages to be presented to the user. In this process, the model performs text summarization and keyword analysis.

[0161] Step 5:

[0162] An emotion engine embedded in the server analyzes input concepts and past interactions to infer the user's emotional state. The input consists of the user's operation history and input concepts, and the output is an inferred emotional state of the user. This process involves algorithmic data analysis.

[0163] Step 6:

[0164] The server adjusts the priority of the optimized passage list based on the sentiment engine's predictions. The input is the sentiment state and the passage list, and the output is the re-prioritized passage list. Specific operations include sorting and filtering the list.

[0165] Step 7:

[0166] The server sends the adjusted information to the terminal, which then presents this information to the user through the user interface. The input here is the adjusted passage list, and the output is the information displayed on the user's screen. Specifically, the font and color on the UI may change according to the user's emotions.

[0167] (Application Example 2)

[0168] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0169] Improving the customer experience is a crucial challenge in modern brick-and-mortar stores. Traditional customer service methods have made it difficult for staff to accurately understand customer emotions and respond accordingly. In particular, providing nuanced responses that respond to customer emotions quickly is challenging, which can lead to decreased customer satisfaction. To solve this problem, there is a need for a system that provides appropriate responses based on customer emotions in real time.

[0170] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0171] In this invention, the server includes an analysis device for inferring the user's emotional state, a control device for adjusting the content of search results and automatically generated annotations based on the analysis device's inference, and a device for providing appropriate responses to staff in real time. This enables quick and appropriate customer service responses that correspond to the customer's emotional state.

[0172] A "user" is a person who uses this system to input and retrieve information.

[0173] "Keywords" are words or phrases that users use to search for information.

[0174] A "device" is a collection of hardware or software designed to perform a specific function.

[0175] A "data set" is a collection of multiple data points that contain specific information.

[0176] A "processing device" is a device that analyzes input data and generates output according to its intended purpose.

[0177] A "display device" is a device used to visually present information to a user.

[0178] An "analytical device" is a device used to examine data and find meaning or patterns in it.

[0179] A "control device" is a device that manages the entire system and directs the operation of each component.

[0180] "Staff" refers to those who interact with customers through this system.

[0181] A "response" is a system's reaction to input from a user or customer.

[0182] The system that realizes this application example is based on devices used by users in physical stores. When a user wears smart glasses or a head-mounted display and interacts with a customer, these devices capture the customer's facial expressions and speech in real time. A server receives this data and analyzes it using emotion recognition AI and natural language processing (NLP) technology.

[0183] Specifically, the system uses the Google Cloud Vision API to analyze customer facial expressions and processes customer speech using OpenAI's GPT generative AI model. This allows the server to infer the customer's emotional state and generate an appropriate response based on that. This response is then displayed on a device worn by the staff to facilitate smooth customer service.

[0184] As a concrete example of its use, if a customer expresses dissatisfaction with a product, the system detects that emotion and suggests a response in real time to the staff, such as "apologize immediately and ask how they can help." This enables a quick and appropriate response to the customer, thereby improving customer satisfaction.

[0185] The following are specific examples of prompt statements to be input to a generative AI model:

[0186] "The customer appears tired. Please think of and offer words that will help them feel at ease and relax."

[0187] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0188] Step 1:

[0189] The user captures the customer's facial expressions and voice during interactions with the customer through smart glasses or a head-mounted display. This data is transmitted in real time from the user's device to a server. The input is the customer's facial expression and voice data, and the output is the data transmission to the server.

[0190] Step 2:

[0191] The server analyzes the received facial expression data using the Google Cloud Vision API. The data processing performed here involves converting the facial expression data into structured data and adding emotion tags (e.g., anger, happiness, dissatisfaction). The input is facial expression data, and the output is data with emotion tags.

[0192] Step 3:

[0193] The server converts received audio data into text and performs natural language processing on the text using OpenAI's generative AI model. Based on specific keywords and context, it analyzes the customer's intent and emotions. The input is text converted from audio data, and the output is an inference of emotional state based on text analysis.

[0194] Step 4:

[0195] The server integrates the results from steps 2 and 3 to determine the customer's overall emotional state. Based on this determination, it generates an appropriate response for the user. The input is emotion-tagged data and text analysis results, and the output is the text as a proposed response.

[0196] Step 5:

[0197] The generated response is transmitted to the user's smart glasses or head-mounted display for visual presentation. Based on this, the user can respond to the customer quickly and appropriately. The input is the text of the response, and the output is the presentation of information to the user.

[0198] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0199] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0200] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0201] [Second Embodiment]

[0202] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0203] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0204] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0205] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0206] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0207] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0208] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0209] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0210] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0211] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0212] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0213] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0214] This invention is implemented as an information processing system. The user connects to the system via a terminal and inputs keywords for a search. The terminal processes the input keywords and sends the data to the server. The server uses the received data to search for documents related to the keywords. The search process employed here utilizes natural language processing technology to efficiently extract relevant information from the document set.

[0215] The server returns the search results to the terminal as a set of data. This result includes passages related to the user's specified theme, with automatically generated annotations that supplement their meaning and background as needed. This annotation generation utilizes pre-set rules and a knowledge base, enabling interpretation from multiple perspectives.

[0216] The terminal receives data sent from the server and displays it to the user. The information is organized, highlighted, and visually supported to facilitate user comprehension. This allows the user to gain knowledge based on multiple religious and philosophical perspectives.

[0217] As a concrete example, if a user enters the keyword "compassion" into their device, the server receives it and searches its database for relevant documents. For instance, it might extract passages related to "compassion" from the scriptures and related literature of a particular religion. The server also generates related annotations explaining how the concept of compassion is understood in different religions. This information is clearly presented on the device, allowing the user to gain deeper insights.

[0218] This system is multilingual, providing value to users from diverse cultural backgrounds. Furthermore, the server ensures data integrity and privacy through secure connections, providing a safe environment for users to search for information. This makes it an indispensable tool for religious and philosophical education settings and individual researchers.

[0219] The following describes the processing flow.

[0220] Step 1:

[0221] The user enters a specific keyword using the terminal. The terminal prepares to store this input internally and triggers its transmission to the server.

[0222] Step 2:

[0223] The terminal sends the keywords entered by the user to the server. At this time, the data is appropriately packaged according to the format and protocol of the transmitted data.

[0224] Step 3:

[0225] The server analyzes the received keywords and initiates a process to search for related documents. This uses natural language processing techniques to analyze the context of the keywords and perform the search.

[0226] Step 4:

[0227] The server selects passages extracted from relevant documents as search results. These results are then formatted to be easily understood by the user.

[0228] Step 5:

[0229] The server automatically generates annotations, adding background information and interpretations relevant to the search results. Using AI technology, it generates annotations that include multiple perspectives.

[0230] Step 6:

[0231] The server sends the formatted search results and generated annotations to the terminal. At this time, it verifies the integrity and security of the data and ensures that the transmission is completed successfully.

[0232] Step 7:

[0233] The terminal receives information from the server and displays it to the user. The information is easy to read and visually organized, allowing the user to immediately utilize the presented information.

[0234] (Example 1)

[0235] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0236] Conventional information retrieval systems make it difficult for users to effectively search for specific information and obtain additional information that includes various perspectives. Therefore, there is a challenge in that it is difficult for users to acquire knowledge from diverse viewpoints.

[0237] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0238] In this invention, the server includes information terminal means for a user to input information, computing device means for processing the information and searching for related data, and information terminal means for displaying the information searched by the computing device means and automatically generating related additional information. This makes it possible for the user to search for information from various perspectives and efficiently obtain related additional information.

[0239] An "information terminal means" is a device for users to input information, and typically provides an interface for sending data to a server.

[0240] A "computation device" is a device that processes received information and searches for and analyzes related data. This includes functions that efficiently handle information using language processing technology.

[0241] "Additional information" refers to information automatically generated in relation to search results, intended to deepen the user's understanding by providing multiple perspectives.

[0242] This invention aims to achieve efficient data processing and information provision in an information retrieval system. First, the user uses a terminal to input information and enters what they want to search for. For example, they can enter the text "What is compassion?". The terminal receives this input and sends the data to the server in an appropriate format.

[0243] The server performs computational processing based on the received data. This processing includes searching for and analyzing relevant information using natural language processing techniques. Specifically, it uses algorithms such as TF-IDF and BERT to extract the most relevant information from the database for the input keywords. It can also utilize generative AI models to create additional information and annotations related to the search results.

[0244] The generated information is sent to the device and displayed to the user. The device organizes this information and highlights particularly important parts. Furthermore, to provide information from multiple perspectives, it is possible to include different interpretations of additional information. This allows users to gain deeper insights, not just acquire data.

[0245] As a concrete example, a prompt message for a user to input the keyword "compassion" might be: "Use natural language processing to search for how the concept of compassion is interpreted in different religions, extract relevant documents, and generate explanations." In this way, users can easily acquire a wide range of knowledge and different perspectives.

[0246] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0247] Step 1:

[0248] The user uses a terminal to input information. The user inputs keywords in text format, such as words like "compassion." The input data is formatted by the terminal and prepared for transmission to the server. This format often uses JSON.

[0249] Step 2:

[0250] The terminal sends keyword data entered by the user to the server. This communication is conducted via HTTP requests, ensuring a secure connection. The entered data arrives at the server, where analysis begins.

[0251] Step 3:

[0252] The server performs a document search within the database based on the received keywords. This search utilizes natural language processing techniques, employing algorithms such as TF-IDF and BERT to extract highly relevant information. The input is the search keywords, and the output is a list of related documents.

[0253] Step 4:

[0254] The server generates annotations related to the search results. A generative AI model is activated, referencing a knowledge base to interpret information from multiple perspectives and create annotations. The input is a list of search result documents, and the output is annotations and additional information for those documents.

[0255] Step 5:

[0256] The server sends the organized search results and generated annotations to the terminal. The data is again sent in JSON format or similar, and communication takes place via a secure protocol. This ensures data privacy and integrity.

[0257] Step 6:

[0258] The terminal receives information from the server and prepares it for display to the user. At this stage, the information is highlighted and visual effects are added to make it easier for the user to understand. The displayed content includes the searched information and its annotations, allowing the user to gain deeper knowledge based on it.

[0259] (Application Example 1)

[0260] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0261] In modern society, there is a need for information processing systems that can understand information from diverse perspectives and reference it quickly and accurately. However, existing systems lack sufficient means to visually present and aid in the understanding of the multifaceted information that users seek. Therefore, there is a need to provide a system that makes it easy for users to obtain visually emphasized information and deepen their understanding from multiple perspectives.

[0262] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0263] In this invention, the server includes a device means for receiving a concept input by a user, a computer means for searching for a set of related information based on the concept, and a device means for presenting the results searched by the computer means and automatically generating related supplementary information. This makes it possible to deepen understanding from multiple perspectives through visually highlighted information.

[0264] A "user" is an entity that inputs information and receives output from a system.

[0265] A "concept" is a collection of information that a user inputs for a system to interpret and process.

[0266] "Reception" refers to the process by which a device takes in user input into the system.

[0267] "Device means" refers to hardware or software components that receive user input and present system output to the user.

[0268] "Searching" is the process by which a computer finds relevant information from a database or knowledge base.

[0269] An "information set" is a collection of data and knowledge related to user input.

[0270] "Computing means" refers to a computing device or server system for searching and processing a set of information.

[0271] "Presentation" refers to the act of a device visually showing information or results to a user.

[0272] "Supplemental information" refers to explanations and annotations added to the information that has been searched, and is data that helps users to understand the information more deeply.

[0273] "Automatic generation" refers to the process by which a system creates additional information or annotations without human intervention.

[0274] "Visual display means" refers to display devices or interfaces used to visually present information to the user.

[0275] To realize this application, the system consists of multiple elements, including a device, a computer system, and a visual display device.

[0276] First, smart glasses are used as a device for the user to input concepts. This device receives concepts from the user using voice input technology. The received information is then transmitted to a server via the internet.

[0277] Next, the server functions as a computing system and employs machine learning techniques. Specifically, it uses the Google Cloud Natural Language API to quickly search for information based on user-inputted concepts. The server retrieves relevant data from databases and automatically generates supplementary information from Wikipedia and other knowledge bases as needed.

[0278] After that, the server returns the search results to the smart glasses. The smart glasses serve as a visual display device and provide the user with visually highlighted information. This enables the user to gain a deep understanding from a multi-faceted perspective.

[0279] As a specific example, if the user wants to know about "non-violence" and makes a voice input to the smart glasses. The server searches for relevant philosophical and religious documents and provides annotations automatically generated from each perspective. Through visual presentation, the user can compare different perspectives and deepen their knowledge.

[0280] An example of a prompt sentence using a generative AI model is "Summarize and present the interpretation of non-violence from the perspectives of Buddhism, Hinduism, and modern ethics."

[0281] The flow of specific processing in Application Example 1 will be described using FIG. 12.

[0282] Step 1:

[0283] The user makes a voice input through the smart glasses. Specifically, the user verbally specifies keywords or concepts of interest. The smart glasses use voice recognition technology to convert the input into text data and transmit that text data to the server via the Internet. The input data is the voice keyword specified by the user, and the output is the keyword in text form transmitted to the server.

[0284] Step 2:

[0285] Based on the received text data, the server utilizes machine learning technology to search for relevant information. Using the Google Cloud Natural Language API, it searches for a group of documents related to the input keyword. As data processing, information retrieval based on the keyword and extraction of important sentences are performed. The input is the keyword in text form, and the output is the group of relevant information retrieved.

[0286] Step 3:

[0287] Based on the search results, the server automatically generates additional supplementary information. Utilizing the knowledge base, it performs data operations to form annotations from multiple perspectives. Specifically, it collects data from Wikipedia and similar documents and generates supplementary information that clarifies different interpretation viewpoints. The input is the group of retrieved related information, and the output is the group of information with added annotations.

[0288] Step 4:

[0289] The server transmits the generated group of information with annotations back to the smart glasses. The user can view the received information in real time through the visual highlighting function. As a specific operation, the smart glasses display the related information on the HUD in the field of vision and visualize the emphasized text and graphics. The input is the group of information with annotations from the server, and the output is the information presented in an easy-to-view manner to the user.

[0290] Furthermore, an emotion engine for estimating the user's emotions may be combined. That is, the specific processing unit 290 may estimate the user's emotions using the emotion identification model 59 and perform specific processing using the user's emotions.

[0291] The present invention provides information according to the user's emotional state by combining an emotion engine with an information processing system. The user inputs keywords using a terminal, and that information is transmitted by the terminal to the server. The server searches for related document groups based on the aforementioned keywords and automatically extracts related passages using natural language processing technology.

[0292] Here, since an emotion engine is added, an additional process for analyzing the user's emotional state intervenes. From the user's input patterns, operation history, etc., the emotion engine infers whether the user has a specific emotion. Based on this inference result, the server adjusts the search results, the content and presentation method of the automatically generated annotations.

[0293] For example, if a user searches for "comfort," the emotion engine will determine that the user is feeling down. The server will then prioritize extracting passages with encouraging and comforting themes, and further highlight their content. The emotion engine supports customization to more effectively deliver information that matches the user's emotions.

[0294] This system supports multiple languages, offering global applicability. In terms of security, communication between the server and the terminal is encrypted, ensuring user privacy while providing accurate and personalized information.

[0295] With the system configuration described above, the present invention not only provides educational value to religious educational institutions and research institutions, but also enables the provision of information that takes into account the deep learning and individual spiritual needs of individual believers and learners.

[0296] The following describes the processing flow.

[0297] Step 1:

[0298] The user enters a specific keyword through the terminal. The terminal temporarily records the input and then prepares to transfer it to the server.

[0299] Step 2:

[0300] The terminal sends the recorded keywords to the server and packages them according to a protocol that ensures the integrity and security of the communication.

[0301] Step 3:

[0302] The server analyzes the received keywords and uses them as matching criteria for searching the document set. Natural language processing techniques are applied to efficiently extract relevant passages from the document set.

[0303] Step 4:

[0304] The emotion engine analyzes the user's input and past operation history, and estimates the emotions that the user may have. An information provision policy considering this emotional state is determined.

[0305] Step 5:

[0306] Based on the analysis results by the emotion engine, the server adjusts the acquired search results and annotations. Specifically, processing such as emphasizing messages corresponding to the user's emotions is performed.

[0307] Step 6:

[0308] The server sends the search results and adjusted annotations to the terminal. In this process, while confirming the integrity of the data, immediate display to the user is enabled.

[0309] Step 7:

[0310] The terminal accurately receives the information from the server and presents visually organized data so that the user can easily understand it. As a result, the user can obtain information corresponding to the individual emotional state.

[0311] (Example 2)

[0312] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0313] In modern information processing systems, it is difficult to realize information provision suitable for the user's emotional state. In ordinary information search systems, the psychological state faced by the user is often not considered, and only uniform information is provided. As a result, the essential needs of the user may not be met. This invention attempts to solve this problem by developing a system that understands the user's emotions and provides information accordingly.

[0314] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0315] In this invention, the server includes means for receiving concepts input by the user, a computer for searching for relevant information resources based on the concepts, and an analysis device for inferring the user's emotional state. This allows for the adjustment of information provided according to the user's emotional state, enabling personalized information delivery.

[0316] A "device" is a machine or a similar system designed to perform a specific function.

[0317] A "concept" is a word, phrase, or collection of words or phrases used to communicate a particular thought or piece of information to others.

[0318] "Information resources" refer to data and knowledge that exist in various forms and may contain the information that users are looking for.

[0319] A "computer" is an electronic device or system used to process data and perform calculations.

[0320] "Notes" are additional explanations or explanatory texts accompanying information or data, provided to aid understanding.

[0321] An "analytical device" is a system that analyzes data and information and uses the results to draw specific conclusions or make inferences.

[0322] "Emotional state" refers to an internal state that indicates a user's psychological and emotional tendencies and mood.

[0323] This invention is an information processing system that provides information tailored to the user's emotional state. The specific implementation method is described below.

[0324] The user first inputs a specific concept using a terminal. This concept is in text format and includes keywords related to what the user wants to research or the information they need. Upon receiving the input, the terminal sends the information to the server using an encrypted communication protocol (e.g., SSL / TLS) to maintain security.

[0325] The server uses a computer to find relevant data from its internally held information resources based on the received concepts. The software technology used here employs models with natural language processing capabilities (e.g., generative AI models such as BERT and GPT). Leveraging these models, it has the ability to automatically extract relevant information.

[0326] Furthermore, the server uses an analysis device called an emotion engine to analyze the user's emotional state. This involves a process of evaluating the user's past activity history and current input data to identify specific emotions. For example, if the concept "seeking encouragement" is input, it is determined that the user's input indicates a depressed state.

[0327] Based on the results of this sentiment analysis, the server adjusts the priority and content of the information presented to the user. This adjustment includes processing the selected information to optimize it for the user's state.

[0328] For example, if a user enters a specific concept seeking "comfort," the server's emotion engine infers that the user is feeling down. The server then uses natural language processing technology to prioritize extracting passages and articles that provide a sense of reassurance and improve the user's mood, and sends that content to the device.

[0329] Through this process, users can receive information optimized to align with their own emotional state. An example of a prompt using a generative AI model is: "Based on the concept entered by user A, please generate optimized information while considering their emotional state."

[0330] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0331] Step 1:

[0332] The user uses the terminal to input concepts related to what they want to research. The input concepts are displayed on the terminal screen and simultaneously processed as digital data. The input format is regular text, which is then prepared to be sent to the server as the initial data.

[0333] Step 2:

[0334] The terminal transmits the input concept to the server over the network. During this process, the data is encrypted using encryption protocols such as SSL / TLS to ensure security during transmission. The output of this process is encrypted concept data.

[0335] Step 3:

[0336] The server decrypts the received encrypted data and extracts the original concept. Based on this concept, the server launches a search engine to find related data from information resources. Here, the input is the decrypted concept, and the output is the identifier of the related information resource.

[0337] Step 4:

[0338] The server uses a generative AI model equipped with natural language processing technology to automatically extract appropriate passages from relevant information resources. The input is the identifier of the searched information resource, and the output is an optimized list of passages to be presented to the user. In this process, the model performs text summarization and keyword analysis.

[0339] Step 5:

[0340] An emotion engine embedded in the server analyzes input concepts and past interactions to infer the user's emotional state. The input consists of the user's operation history and input concepts, and the output is an inferred emotional state of the user. This process involves algorithmic data analysis.

[0341] Step 6:

[0342] The server adjusts the priority of the optimized passage list based on the sentiment engine's predictions. The input is the sentiment state and the passage list, and the output is the re-prioritized passage list. Specific operations include sorting and filtering the list.

[0343] Step 7:

[0344] The server sends the adjusted information to the terminal, which then presents this information to the user through the user interface. The input here is the adjusted passage list, and the output is the information displayed on the user's screen. Specifically, the font and color on the UI may change according to the user's emotions.

[0345] (Application Example 2)

[0346] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0347] Improving the customer experience is a crucial challenge in modern brick-and-mortar stores. Traditional customer service methods have made it difficult for staff to accurately understand customer emotions and respond accordingly. In particular, providing nuanced responses that respond to customer emotions quickly is challenging, which can lead to decreased customer satisfaction. To solve this problem, there is a need for a system that provides appropriate responses based on customer emotions in real time.

[0348] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0349] In this invention, the server includes an analysis device for inferring the user's emotional state, a control device for adjusting the content of search results and automatically generated annotations based on the analysis device's inference, and a device for providing appropriate responses to staff in real time. This enables quick and appropriate customer service responses that correspond to the customer's emotional state.

[0350] A "user" is a person who uses this system to input and retrieve information.

[0351] "Keywords" are words or phrases that users use to search for information.

[0352] A "device" is a collection of hardware or software designed to perform a specific function.

[0353] A "data set" is a collection of multiple data points that contain specific information.

[0354] A "processing device" is a device that analyzes input data and generates output according to its intended purpose.

[0355] A "display device" is a device used to visually present information to a user.

[0356] An "analytical device" is a device used to examine data and find meaning or patterns in it.

[0357] A "control device" is a device that manages the entire system and directs the operation of each component.

[0358] "Staff" refers to those who interact with customers through this system.

[0359] A "response" is a system's reaction to input from a user or customer.

[0360] The system that realizes this application example is based on devices used by users in physical stores. When a user wears smart glasses or a head-mounted display and interacts with a customer, these devices capture the customer's facial expressions and speech in real time. A server receives this data and analyzes it using emotion recognition AI and natural language processing (NLP) technology.

[0361] Specifically, the system uses the Google Cloud Vision API to analyze customer facial expressions and processes customer speech using OpenAI's generative AI model, GPT. This allows the server to infer the customer's emotional state and generate an appropriate response based on that. This response is then displayed on a device worn by the staff to facilitate smooth customer service.

[0362] As a concrete example of its use, if a customer expresses dissatisfaction with a product, the system detects that emotion and suggests a response in real time to the staff, such as "apologize immediately and ask how they can help." This enables a quick and appropriate response to the customer, thereby improving customer satisfaction.

[0363] The following are specific examples of prompt statements to be input to a generative AI model:

[0364] "The customer appears tired. Please think of and offer words that will help them feel at ease and relax."

[0365] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0366] Step 1:

[0367] The user captures the customer's facial expressions and voice during interactions with the customer through smart glasses or a head-mounted display. This data is transmitted in real time from the user's device to a server. The input is the customer's facial expression and voice data, and the output is the data transmission to the server.

[0368] Step 2:

[0369] The server analyzes the received facial expression data using the Google Cloud Vision API. The data processing performed here involves converting the facial expression data into structured data and adding emotion tags (e.g., anger, happiness, dissatisfaction). The input is facial expression data, and the output is data with emotion tags.

[0370] Step 3:

[0371] The server converts received audio data into text and performs natural language processing on the text using OpenAI's generative AI model. Based on specific keywords and context, it analyzes the customer's intent and emotions. The input is text converted from audio data, and the output is an inference of emotional state based on text analysis.

[0372] Step 4:

[0373] The server integrates the results from steps 2 and 3 to determine the customer's overall emotional state. Based on this determination, it generates an appropriate response for the user. The input is emotion-tagged data and text analysis results, and the output is the text as a proposed response.

[0374] Step 5:

[0375] The generated response is transmitted to the user's smart glasses or head-mounted display for visual presentation. Based on this, the user can respond to the customer quickly and appropriately. The input is the text of the response, and the output is the presentation of information to the user.

[0376] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0377] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0378] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0379] [Third Embodiment]

[0380] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0381] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0382] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0383] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0384] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0385] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0386] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0387] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0388] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0389] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0390] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0391] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0392] This invention is implemented as an information processing system. The user connects to the system via a terminal and inputs keywords for a search. The terminal processes the input keywords and sends the data to the server. The server uses the received data to search for documents related to the keywords. The search process employed here utilizes natural language processing technology to efficiently extract relevant information from the document set.

[0393] The server returns the search results to the terminal as a set of data. This result includes passages related to the user's specified theme, with automatically generated annotations that supplement their meaning and background as needed. This annotation generation utilizes pre-set rules and a knowledge base, enabling interpretation from multiple perspectives.

[0394] The terminal receives data sent from the server and displays it to the user. The information is organized, highlighted, and visually supported to facilitate user comprehension. This allows the user to gain knowledge based on multiple religious and philosophical perspectives.

[0395] As a concrete example, if a user enters the keyword "compassion" into their device, the server receives it and searches its database for relevant documents. For instance, it might extract passages related to "compassion" from the scriptures and related literature of a particular religion. The server also generates related annotations explaining how the concept of compassion is understood in different religions. This information is clearly presented on the device, allowing the user to gain deeper insights.

[0396] This system is multilingual, providing value to users from diverse cultural backgrounds. Furthermore, the server ensures data integrity and privacy through secure connections, providing a safe environment for users to search for information. This makes it an indispensable tool for religious and philosophical education settings and individual researchers.

[0397] The following describes the processing flow.

[0398] Step 1:

[0399] The user enters a specific keyword using the terminal. The terminal prepares to store this input internally and triggers its transmission to the server.

[0400] Step 2:

[0401] The terminal sends the keywords entered by the user to the server. At this time, the data is appropriately packaged according to the format and protocol of the transmitted data.

[0402] Step 3:

[0403] The server analyzes the received keywords and initiates a process to search for related documents. This uses natural language processing techniques to analyze the context of the keywords and perform the search.

[0404] Step 4:

[0405] The server selects passages extracted from relevant documents as search results. These results are then formatted to be easily understood by the user.

[0406] Step 5:

[0407] The server automatically generates annotations, adding background information and interpretations relevant to the search results. Using AI technology, it generates annotations that include multiple perspectives.

[0408] Step 6:

[0409] The server sends the formatted search results and generated annotations to the terminal. At this time, it verifies the integrity and security of the data and ensures that the transmission is completed successfully.

[0410] Step 7:

[0411] The terminal receives information from the server and displays it to the user. The information is easy to read and visually organized, allowing the user to immediately utilize the presented information.

[0412] (Example 1)

[0413] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0414] Conventional information retrieval systems make it difficult for users to effectively search for specific information and obtain additional information that includes various perspectives. Therefore, there is a challenge in that it is difficult for users to acquire knowledge from diverse viewpoints.

[0415] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0416] In this invention, the server includes information terminal means for a user to input information, computing device means for processing the information and searching for related data, and information terminal means for displaying the information searched by the computing device means and automatically generating related additional information. This makes it possible for the user to search for information from various perspectives and efficiently obtain related additional information.

[0417] An "information terminal means" is a device for users to input information, and typically provides an interface for sending data to a server.

[0418] A "computation device" is a device that processes received information and searches for and analyzes related data. This includes functions that efficiently handle information using language processing technology.

[0419] "Additional information" refers to information automatically generated in relation to search results, intended to deepen the user's understanding by providing multiple perspectives.

[0420] This invention aims to achieve efficient data processing and information provision in an information retrieval system. First, the user uses a terminal to input information and enters what they want to search for. For example, they can enter the text "What is compassion?". The terminal receives this input and sends the data to the server in an appropriate format.

[0421] The server performs computational processing based on the received data. This processing includes searching for and analyzing relevant information using natural language processing techniques. Specifically, it uses algorithms such as TF-IDF and BERT to extract the most relevant information from the database for the input keywords. It can also utilize generative AI models to create additional information and annotations related to the search results.

[0422] The generated information is sent to the device and displayed to the user. The device organizes this information and highlights particularly important parts. Furthermore, to provide information from multiple perspectives, it is possible to include different interpretations of additional information. This allows users to gain deeper insights, not just acquire data.

[0423] As a concrete example, a prompt message for a user to input the keyword "compassion" might be: "Use natural language processing to search for how the concept of compassion is interpreted in different religions, extract relevant documents, and generate explanations." In this way, users can easily acquire a wide range of knowledge and different perspectives.

[0424] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0425] Step 1:

[0426] The user uses a terminal to input information. The user inputs keywords in text format, such as words like "compassion." The input data is formatted by the terminal and prepared for transmission to the server. This format often uses JSON.

[0427] Step 2:

[0428] The terminal sends keyword data entered by the user to the server. This communication is conducted via HTTP requests, ensuring a secure connection. The entered data arrives at the server, where analysis begins.

[0429] Step 3:

[0430] The server performs a document search within the database based on the received keywords. This search utilizes natural language processing techniques, employing algorithms such as TF-IDF and BERT to extract highly relevant information. The input is the search keywords, and the output is a list of related documents.

[0431] Step 4:

[0432] The server generates annotations related to the search results. A generative AI model is activated, referencing a knowledge base to interpret information from multiple perspectives and create annotations. The input is a list of search result documents, and the output is annotations and additional information for those documents.

[0433] Step 5:

[0434] The server sends the organized search results and generated annotations to the terminal. The data is again sent in JSON format or similar, and communication takes place via a secure protocol. This ensures data privacy and integrity.

[0435] Step 6:

[0436] The terminal receives information from the server and prepares it for display to the user. At this stage, the information is highlighted and visual effects are added to make it easier for the user to understand. The displayed content includes the searched information and its annotations, allowing the user to gain deeper knowledge based on it.

[0437] (Application Example 1)

[0438] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0439] In modern society, there is a need for information processing systems that can understand information from diverse perspectives and reference it quickly and accurately. However, existing systems lack sufficient means to visually present and aid in the understanding of the multifaceted information that users seek. Therefore, there is a need to provide a system that makes it easy for users to obtain visually emphasized information and deepen their understanding from multiple perspectives.

[0440] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0441] In this invention, the server includes a device means for receiving a concept input by a user, a computer means for searching for a set of related information based on the concept, and a device means for presenting the results searched by the computer means and automatically generating related supplementary information. This makes it possible to deepen understanding from multiple perspectives through visually highlighted information.

[0442] A "user" is an entity that inputs information and receives output from a system.

[0443] A "concept" is a collection of information that a user inputs for a system to interpret and process.

[0444] "Reception" refers to the process by which a device takes in user input into the system.

[0445] "Device means" refers to hardware or software components that receive user input and present system output to the user.

[0446] "Searching" is the process by which a computer finds relevant information from a database or knowledge base.

[0447] An "information set" is a collection of data and knowledge related to user input.

[0448] "Computing means" refers to a computing device or server system for searching and processing a set of information.

[0449] "Presentation" refers to the act of a device visually showing information or results to a user.

[0450] "Supplemental information" refers to explanations and annotations added to the information that has been searched, and is data that helps users to understand the information more deeply.

[0451] "Automatic generation" refers to the process by which a system creates additional information or annotations without human intervention.

[0452] "Visual display means" refers to display devices or interfaces used to visually present information to the user.

[0453] To realize this application, the system consists of multiple elements, including a device, a computer system, and a visual display device.

[0454] First, smart glasses are used as a device for the user to input concepts. This device receives concepts from the user using voice input technology. The received information is then transmitted to a server via the internet.

[0455] Next, the server functions as a computing system and employs machine learning techniques. Specifically, it uses the Google Cloud Natural Language API to quickly search for information based on user-inputted concepts. The server retrieves relevant data from databases and automatically generates supplementary information from Wikipedia and other knowledge bases as needed.

[0456] The server then sends the search results back to the smart glasses. The smart glasses act as a visual display device, providing the user with visually highlighted information. This allows the user to gain a deeper understanding from multiple perspectives.

[0457] As a concrete example, a user might want to learn about "non-violence" and use voice input into smart glasses. The server searches for relevant philosophical and religious documents and provides automatically generated annotations from various perspectives. Visual presentations allow the user to compare different viewpoints and deepen their understanding.

[0458] An example of a prompt sentence generated using an AI model is: "Summarize and present interpretations of nonviolence from the perspectives of Buddhism, Hinduism, and modern ethics."

[0459] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0460] Step 1:

[0461] Users provide voice input through smart glasses. Specifically, they specify keywords or concepts they are interested in by voice. The smart glasses use voice recognition technology to convert the input into text data and send that text data to a server via the internet. The input data is the voice keywords specified by the user, and the output is the keywords in text format sent to the server.

[0462] Step 2:

[0463] The server uses machine learning techniques to search for relevant information based on the received text data. It uses the Google Cloud Natural Language API to explore document sets related to the input keywords. Data processing includes keyword-based information retrieval and extraction of important sentences. The input is keywords in text format, and the output is the searched set of relevant information.

[0464] Step 3:

[0465] The server automatically generates additional supplementary information based on the search results. It utilizes a knowledge base to perform data calculations that form annotations from multiple perspectives. Specifically, it collects data from Wikipedia and similar sources to generate supplementary information that highlights different interpretations. The input is the searched related information, and the output is the annotated information.

[0466] Step 4:

[0467] The server sends the generated annotated information back to the smart glasses. The user can view the received information in real time through the visual highlighting function. Specifically, the smart glasses display relevant information on the HUD in the user's field of view, visualizing highlighted text and graphics. The input is the annotated information from the server, and the output is the information displayed to the user in an easy-to-read format.

[0468] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0469] This invention provides information tailored to the user's emotional state by combining an emotion engine with an information processing system. The user inputs keywords using a terminal, and this information is transmitted to a server via the terminal. The server searches for related documents based on the aforementioned keywords and automatically extracts relevant passages using natural language processing technology.

[0470] Here, an emotion engine is added, introducing an additional process to analyze the user's emotional state. Based on the user's input patterns and operation history, the emotion engine infers whether the user is experiencing a particular emotion. Based on this inference, the server adjusts the content and presentation method of search results and automatically generated annotations.

[0471] For example, if a user searches for "comfort," the emotion engine will determine that the user is feeling down. The server will then prioritize extracting passages with encouraging and comforting themes, and further highlight their content. The emotion engine supports customization to more effectively deliver information that matches the user's emotions.

[0472] This system supports multiple languages, offering global applicability. In terms of security, communication between the server and the terminal is encrypted, ensuring user privacy while providing accurate and personalized information.

[0473] With the system configuration described above, the present invention not only provides educational value to religious educational institutions and research institutions, but also enables the provision of information that takes into account the deep learning and individual spiritual needs of individual believers and learners.

[0474] The following describes the processing flow.

[0475] Step 1:

[0476] The user enters a specific keyword through the terminal. The terminal temporarily records the input and then prepares to transfer it to the server.

[0477] Step 2:

[0478] The terminal sends the recorded keywords to the server and packages them according to a protocol that ensures the integrity and security of the communication.

[0479] Step 3:

[0480] The server analyzes the received keywords and uses them as matching criteria for searching the document set. Natural language processing techniques are applied to efficiently extract relevant passages from the document set.

[0481] Step 4:

[0482] The emotion engine analyzes user input and past operation history to estimate the emotions the user may be experiencing. It then determines an information provision policy that takes this emotional state into account.

[0483] Step 5:

[0484] The server adjusts the retrieved search results and annotations based on the analysis results from the sentiment engine. Specifically, it performs processes such as highlighting passages that correspond to the user's emotions.

[0485] Step 6:

[0486] The server sends search results and adjusted annotations to the terminal. This process verifies data integrity while enabling immediate display to the user.

[0487] Step 7:

[0488] The device accurately receives information from the server and presents the data in a visually organized format that is easy for the user to understand. This allows the user to obtain information that corresponds to their individual emotional state.

[0489] (Example 2)

[0490] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0491] In modern information processing systems, providing information tailored to the user's emotional state is challenging. Conventional information retrieval systems often provide uniform information without considering the user's psychological state, resulting in a failure to meet the user's essential needs. This invention aims to solve this problem by developing a system that understands the user's emotions and provides information accordingly.

[0492] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0493] In this invention, the server includes means for receiving concepts input by the user, a computer for searching for relevant information resources based on the concepts, and an analysis device for inferring the user's emotional state. This allows for the adjustment of information provided according to the user's emotional state, enabling personalized information delivery.

[0494] A "device" is a machine or a similar system designed to perform a specific function.

[0495] A "concept" is a word, phrase, or collection of words or phrases used to communicate a particular thought or piece of information to others.

[0496] "Information resources" refer to data and knowledge that exist in various forms and may contain the information that users are looking for.

[0497] A "computer" is an electronic device or system used to process data and perform calculations.

[0498] "Notes" are additional explanations or explanatory texts accompanying information or data, provided to aid understanding.

[0499] An "analytical device" is a system that analyzes data and information and uses the results to draw specific conclusions or make inferences.

[0500] "Emotional state" refers to an internal state that indicates a user's psychological and emotional tendencies and mood.

[0501] This invention is an information processing system that provides information tailored to the user's emotional state. The specific implementation method is described below.

[0502] The user first inputs a specific concept using a terminal. This concept is in text format and includes keywords related to what the user wants to research or the information they need. Upon receiving the input, the terminal sends the information to the server using an encrypted communication protocol (e.g., SSL / TLS) to maintain security.

[0503] The server uses a computer to find relevant data from its internally held information resources based on the received concepts. The software technology used here employs models with natural language processing capabilities (e.g., generative AI models such as BERT and GPT). Leveraging these models, it has the ability to automatically extract relevant information.

[0504] Furthermore, the server uses an analysis device called an emotion engine to analyze the user's emotional state. This involves a process of evaluating the user's past activity history and current input data to identify specific emotions. For example, if the concept "seeking encouragement" is input, it is determined that the user's input indicates a depressed state.

[0505] Based on the results of this sentiment analysis, the server adjusts the priority and content of the information presented to the user. This adjustment includes processing the selected information to optimize it for the user's state.

[0506] For example, if a user enters a specific concept seeking "comfort," the server's emotion engine infers that the user is feeling down. The server then uses natural language processing technology to prioritize extracting passages and articles that provide a sense of reassurance and improve the user's mood, and sends that content to the device.

[0507] Through this process, users can receive information optimized to align with their own emotional state. An example of a prompt using a generative AI model is: "Based on the concept entered by user A, please generate optimized information while considering their emotional state."

[0508] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0509] Step 1:

[0510] The user uses the terminal to input concepts related to what they want to research. The input concepts are displayed on the terminal screen and simultaneously processed as digital data. The input format is regular text, which is then prepared to be sent to the server as the initial data.

[0511] Step 2:

[0512] The terminal transmits the input concept to the server over the network. During this process, the data is encrypted using encryption protocols such as SSL / TLS to ensure security during transmission. The output of this process is encrypted concept data.

[0513] Step 3:

[0514] The server decrypts the received encrypted data and extracts the original concept. Based on this concept, the server launches a search engine to find related data from information resources. Here, the input is the decrypted concept, and the output is the identifier of the related information resource.

[0515] Step 4:

[0516] The server uses a generative AI model equipped with natural language processing technology to automatically extract appropriate passages from relevant information resources. The input is the identifier of the searched information resource, and the output is an optimized list of passages to be presented to the user. In this process, the model performs text summarization and keyword analysis.

[0517] Step 5:

[0518] An emotion engine embedded in the server analyzes input concepts and past interactions to infer the user's emotional state. The input consists of the user's operation history and input concepts, and the output is an inferred emotional state of the user. This process involves algorithmic data analysis.

[0519] Step 6:

[0520] The server adjusts the priority of the optimized passage list based on the sentiment engine's predictions. The input is the sentiment state and the passage list, and the output is the re-prioritized passage list. Specific operations include sorting and filtering the list.

[0521] Step 7:

[0522] The server sends the adjusted information to the terminal, which then presents this information to the user through the user interface. The input here is the adjusted passage list, and the output is the information displayed on the user's screen. Specifically, the font and color on the UI may change according to the user's emotions.

[0523] (Application Example 2)

[0524] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0525] Improving the customer experience is a crucial challenge in modern brick-and-mortar stores. Traditional customer service methods have made it difficult for staff to accurately understand customer emotions and respond accordingly. In particular, providing nuanced responses that respond to customer emotions quickly is challenging, which can lead to decreased customer satisfaction. To solve this problem, there is a need for a system that provides appropriate responses based on customer emotions in real time.

[0526] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0527] In this invention, the server includes an analysis device for inferring the user's emotional state, a control device for adjusting the content of search results and automatically generated annotations based on the analysis device's inference, and a device for providing appropriate responses to staff in real time. This enables quick and appropriate customer service responses that correspond to the customer's emotional state.

[0528] A "user" is a person who uses this system to input and retrieve information.

[0529] "Keywords" are words or phrases that users use to search for information.

[0530] A "device" is a collection of hardware or software designed to perform a specific function.

[0531] A "data set" is a collection of multiple data points that contain specific information.

[0532] A "processing device" is a device that analyzes input data and generates output according to its intended purpose.

[0533] A "display device" is a device used to visually present information to a user.

[0534] An "analytical device" is a device used to examine data and find meaning or patterns in it.

[0535] A "control device" is a device that manages the entire system and directs the operation of each component.

[0536] "Staff" refers to those who interact with customers through this system.

[0537] A "response" is a system's reaction to input from a user or customer.

[0538] The system that realizes this application example is based on devices used by users in physical stores. When a user wears smart glasses or a head-mounted display and interacts with a customer, these devices capture the customer's facial expressions and speech in real time. A server receives this data and analyzes it using emotion recognition AI and natural language processing (NLP) technology.

[0539] Specifically, the system uses the Google Cloud Vision API to analyze customer facial expressions and processes customer speech using OpenAI's generative AI model, GPT. This allows the server to infer the customer's emotional state and generate an appropriate response based on that. This response is then displayed on a device worn by the staff to facilitate smooth customer service.

[0540] As a concrete example of its use, if a customer expresses dissatisfaction with a product, the system detects that emotion and suggests a response in real time to the staff, such as "apologize immediately and ask how they can help." This enables a quick and appropriate response to the customer, thereby improving customer satisfaction.

[0541] The following are specific examples of prompt statements to be input to a generative AI model:

[0542] "The customer appears tired. Please think of and offer words that will help them feel at ease and relax."

[0543] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0544] Step 1:

[0545] The user captures the customer's facial expressions and voice during interactions with the customer through smart glasses or a head-mounted display. This data is transmitted in real time from the user's device to a server. The input is the customer's facial expression and voice data, and the output is the data transmission to the server.

[0546] Step 2:

[0547] The server analyzes the received facial expression data using the Google Cloud Vision API. The data processing performed here involves converting the facial expression data into structured data and adding emotion tags (e.g., anger, happiness, dissatisfaction). The input is facial expression data, and the output is data with emotion tags.

[0548] Step 3:

[0549] The server converts received audio data into text and performs natural language processing on the text using OpenAI's generative AI model. Based on specific keywords and context, it analyzes the customer's intent and emotions. The input is text converted from audio data, and the output is an inference of emotional state based on text analysis.

[0550] Step 4:

[0551] The server integrates the results from steps 2 and 3 to determine the customer's overall emotional state. Based on this determination, it generates an appropriate response for the user. The input is emotion-tagged data and text analysis results, and the output is the text as a proposed response.

[0552] Step 5:

[0553] The generated response is transmitted to the user's smart glasses or head-mounted display for visual presentation. Based on this, the user can respond to the customer quickly and appropriately. The input is the text of the response, and the output is the presentation of information to the user.

[0554] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0555] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0556] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0557] [Fourth Embodiment]

[0558] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0559] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0560] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0561] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0562] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0563] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0564] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0565] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0566] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0567] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0568] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0569] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0570] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0571] This invention is implemented as an information processing system. The user connects to the system via a terminal and inputs keywords for a search. The terminal processes the input keywords and sends the data to the server. The server uses the received data to search for documents related to the keywords. The search process employed here utilizes natural language processing technology to efficiently extract relevant information from the document set.

[0572] The server returns the search results to the terminal as a set of data. This result includes passages related to the user's specified theme, with automatically generated annotations that supplement their meaning and background as needed. This annotation generation utilizes pre-set rules and a knowledge base, enabling interpretation from multiple perspectives.

[0573] The terminal receives data sent from the server and displays it to the user. The information is organized, highlighted, and visually supported to facilitate user comprehension. This allows the user to gain knowledge based on multiple religious and philosophical perspectives.

[0574] As a concrete example, if a user enters the keyword "compassion" into their device, the server receives it and searches its database for relevant documents. For instance, it might extract passages related to "compassion" from the scriptures and related literature of a particular religion. The server also generates related annotations explaining how the concept of compassion is understood in different religions. This information is clearly presented on the device, allowing the user to gain deeper insights.

[0575] This system is multilingual, providing value to users from diverse cultural backgrounds. Furthermore, the server ensures data integrity and privacy through secure connections, providing a safe environment for users to search for information. This makes it an indispensable tool for religious and philosophical education settings and individual researchers.

[0576] The following describes the processing flow.

[0577] Step 1:

[0578] The user enters a specific keyword using the terminal. The terminal prepares to store this input internally and triggers its transmission to the server.

[0579] Step 2:

[0580] The terminal sends the keywords entered by the user to the server. At this time, the data is appropriately packaged according to the format and protocol of the transmitted data.

[0581] Step 3:

[0582] The server analyzes the received keywords and initiates a process to search for related documents. This uses natural language processing techniques to analyze the context of the keywords and perform the search.

[0583] Step 4:

[0584] The server selects passages extracted from relevant documents as search results. These results are then formatted to be easily understood by the user.

[0585] Step 5:

[0586] The server automatically generates annotations, adding background information and interpretations relevant to the search results. Using AI technology, it generates annotations that include multiple perspectives.

[0587] Step 6:

[0588] The server sends the formatted search results and generated annotations to the terminal. At this time, it verifies the integrity and security of the data and ensures that the transmission is completed successfully.

[0589] Step 7:

[0590] The terminal receives information from the server and displays it to the user. The information is easy to read and visually organized, allowing the user to immediately utilize the presented information.

[0591] (Example 1)

[0592] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0593] Conventional information retrieval systems make it difficult for users to effectively search for specific information and obtain additional information that includes various perspectives. Therefore, there is a challenge in that it is difficult for users to acquire knowledge from diverse viewpoints.

[0594] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0595] In this invention, the server includes information terminal means for a user to input information, computing device means for processing the information and searching for related data, and information terminal means for displaying the information searched by the computing device means and automatically generating related additional information. This makes it possible for the user to search for information from various perspectives and efficiently obtain related additional information.

[0596] An "information terminal means" is a device for users to input information, and typically provides an interface for sending data to a server.

[0597] A "computation device" is a device that processes received information and searches for and analyzes related data. This includes functions that efficiently handle information using language processing technology.

[0598] "Additional information" refers to information automatically generated in relation to search results, intended to deepen the user's understanding by providing multiple perspectives.

[0599] This invention aims to achieve efficient data processing and information provision in an information retrieval system. First, the user uses a terminal to input information and enters what they want to search for. For example, they can enter the text "What is compassion?". The terminal receives this input and sends the data to the server in an appropriate format.

[0600] The server performs computational processing based on the received data. This processing includes searching for and analyzing relevant information using natural language processing techniques. Specifically, it uses algorithms such as TF-IDF and BERT to extract the most relevant information from the database for the input keywords. It can also utilize generative AI models to create additional information and annotations related to the search results.

[0601] The generated information is sent to the device and displayed to the user. The device organizes this information and highlights particularly important parts. Furthermore, to provide information from multiple perspectives, it is possible to include different interpretations of additional information. This allows users to gain deeper insights, not just acquire data.

[0602] As a concrete example, a prompt message for a user to input the keyword "compassion" might be: "Use natural language processing to search for how the concept of compassion is interpreted in different religions, extract relevant documents, and generate explanations." In this way, users can easily acquire a wide range of knowledge and different perspectives.

[0603] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0604] Step 1:

[0605] The user uses a terminal to input information. The user inputs keywords in text format, such as words like "compassion." The input data is formatted by the terminal and prepared for transmission to the server. This format often uses JSON.

[0606] Step 2:

[0607] The terminal sends keyword data entered by the user to the server. This communication is conducted via HTTP requests, ensuring a secure connection. The entered data arrives at the server, where analysis begins.

[0608] Step 3:

[0609] The server performs a document search within the database based on the received keywords. This search utilizes natural language processing techniques, employing algorithms such as TF-IDF and BERT to extract highly relevant information. The input is the search keywords, and the output is a list of related documents.

[0610] Step 4:

[0611] The server generates annotations related to the search results. A generative AI model is activated, referencing a knowledge base to interpret information from multiple perspectives and create annotations. The input is a list of search result documents, and the output is annotations and additional information for those documents.

[0612] Step 5:

[0613] The server sends the organized search results and generated annotations to the terminal. The data is again sent in JSON format or similar, and communication takes place via a secure protocol. This ensures data privacy and integrity.

[0614] Step 6:

[0615] The terminal receives information from the server and prepares it for display to the user. At this stage, the information is highlighted and visual effects are added to make it easier for the user to understand. The displayed content includes the searched information and its annotations, allowing the user to gain deeper knowledge based on it.

[0616] (Application Example 1)

[0617] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0618] In modern society, there is a need for information processing systems that can understand information from diverse perspectives and reference it quickly and accurately. However, existing systems lack sufficient means to visually present and aid in the understanding of the multifaceted information that users seek. Therefore, there is a need to provide a system that makes it easy for users to obtain visually emphasized information and deepen their understanding from multiple perspectives.

[0619] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0620] In this invention, the server includes a device means for receiving a concept input by a user, a computer means for searching for a set of related information based on the concept, and a device means for presenting the results searched by the computer means and automatically generating related supplementary information. This makes it possible to deepen understanding from multiple perspectives through visually highlighted information.

[0621] A "user" is an entity that inputs information and receives output from a system.

[0622] A "concept" is a collection of information that a user inputs for a system to interpret and process.

[0623] "Reception" refers to the process by which a device takes in user input into the system.

[0624] "Device means" refers to hardware or software components that receive user input and present system output to the user.

[0625] "Searching" is the process by which a computer finds relevant information from a database or knowledge base.

[0626] An "information set" is a collection of data and knowledge related to user input.

[0627] "Computing means" refers to a computing device or server system for searching and processing a set of information.

[0628] "Presentation" refers to the act of a device visually showing information or results to a user.

[0629] "Supplemental information" refers to explanations and annotations added to the information that has been searched, and is data that helps users to understand the information more deeply.

[0630] "Automatic generation" refers to the process by which a system creates additional information or annotations without human intervention.

[0631] "Visual display means" refers to display devices or interfaces used to visually present information to the user.

[0632] To realize this application, the system consists of multiple elements, including a device, a computer system, and a visual display device.

[0633] First, smart glasses are used as a device for the user to input concepts. This device receives concepts from the user using voice input technology. The received information is then transmitted to a server via the internet.

[0634] Next, the server functions as a computing system and employs machine learning techniques. Specifically, it uses the Google Cloud Natural Language API to quickly search for information based on user-inputted concepts. The server retrieves relevant data from databases and automatically generates supplementary information from Wikipedia and other knowledge bases as needed.

[0635] The server then sends the search results back to the smart glasses. The smart glasses act as a visual display device, providing the user with visually highlighted information. This allows the user to gain a deeper understanding from multiple perspectives.

[0636] As a concrete example, a user might want to learn about "non-violence" and use voice input into smart glasses. The server searches for relevant philosophical and religious documents and provides automatically generated annotations from various perspectives. Visual presentations allow the user to compare different viewpoints and deepen their understanding.

[0637] An example of a prompt sentence generated using an AI model is: "Summarize and present interpretations of nonviolence from the perspectives of Buddhism, Hinduism, and modern ethics."

[0638] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0639] Step 1:

[0640] Users provide voice input through smart glasses. Specifically, they specify keywords or concepts they are interested in by voice. The smart glasses use voice recognition technology to convert the input into text data and send that text data to a server via the internet. The input data is the voice keywords specified by the user, and the output is the keywords in text format sent to the server.

[0641] Step 2:

[0642] The server uses machine learning techniques to search for relevant information based on the received text data. It uses the Google Cloud Natural Language API to explore document sets related to the input keywords. Data processing includes keyword-based information retrieval and extraction of important sentences. The input is keywords in text format, and the output is the searched set of relevant information.

[0643] Step 3:

[0644] The server automatically generates additional supplementary information based on the search results. It utilizes a knowledge base to perform data calculations that form annotations from multiple perspectives. Specifically, it collects data from Wikipedia and similar sources to generate supplementary information that highlights different interpretations. The input is the searched related information, and the output is the annotated information.

[0645] Step 4:

[0646] The server sends the generated annotated information back to the smart glasses. The user can view the received information in real time through the visual highlighting function. Specifically, the smart glasses display relevant information on the HUD in the user's field of view, visualizing highlighted text and graphics. The input is the annotated information from the server, and the output is the information displayed to the user in an easy-to-read format.

[0647] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0648] This invention provides information tailored to the user's emotional state by combining an emotion engine with an information processing system. The user inputs keywords using a terminal, and this information is transmitted to a server via the terminal. The server searches for related documents based on the aforementioned keywords and automatically extracts relevant passages using natural language processing technology.

[0649] Here, an emotion engine is added, introducing an additional process to analyze the user's emotional state. Based on the user's input patterns and operation history, the emotion engine infers whether the user is experiencing a particular emotion. Based on this inference, the server adjusts the content and presentation method of search results and automatically generated annotations.

[0650] For example, if a user searches for "comfort," the emotion engine will determine that the user is feeling down. The server will then prioritize extracting passages with encouraging and comforting themes, and further highlight their content. The emotion engine supports customization to more effectively deliver information that matches the user's emotions.

[0651] This system supports multiple languages, offering global applicability. In terms of security, communication between the server and the terminal is encrypted, ensuring user privacy while providing accurate and personalized information.

[0652] With the system configuration described above, the present invention not only provides educational value to religious educational institutions and research institutions, but also enables the provision of information that takes into account the deep learning and individual spiritual needs of individual believers and learners.

[0653] The following describes the processing flow.

[0654] Step 1:

[0655] The user enters a specific keyword through the terminal. The terminal temporarily records the input and then prepares to transfer it to the server.

[0656] Step 2:

[0657] The terminal sends the recorded keywords to the server and packages them according to a protocol that ensures the integrity and security of the communication.

[0658] Step 3:

[0659] The server analyzes the received keywords and uses them as matching criteria for searching the document set. Natural language processing techniques are applied to efficiently extract relevant passages from the document set.

[0660] Step 4:

[0661] The emotion engine analyzes user input and past operation history to estimate the emotions the user may be experiencing. It then determines an information provision policy that takes this emotional state into account.

[0662] Step 5:

[0663] The server adjusts the retrieved search results and annotations based on the analysis results from the sentiment engine. Specifically, it performs processes such as highlighting passages that correspond to the user's emotions.

[0664] Step 6:

[0665] The server sends search results and adjusted annotations to the terminal. This process verifies data integrity while enabling immediate display to the user.

[0666] Step 7:

[0667] The device accurately receives information from the server and presents the data in a visually organized format that is easy for the user to understand. This allows the user to obtain information that corresponds to their individual emotional state.

[0668] (Example 2)

[0669] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0670] In modern information processing systems, providing information tailored to the user's emotional state is challenging. Conventional information retrieval systems often provide uniform information without considering the user's psychological state, resulting in a failure to meet the user's essential needs. This invention aims to solve this problem by developing a system that understands the user's emotions and provides information accordingly.

[0671] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0672] In this invention, the server includes means for receiving concepts input by the user, a computer for searching for relevant information resources based on the concepts, and an analysis device for inferring the user's emotional state. This allows for the adjustment of information provided according to the user's emotional state, enabling personalized information delivery.

[0673] A "device" is a machine or a similar system designed to perform a specific function.

[0674] A "concept" is a word, phrase, or collection of words or phrases used to communicate a particular thought or piece of information to others.

[0675] "Information resources" refer to data and knowledge that exist in various forms and may contain the information that users are looking for.

[0676] A "computer" is an electronic device or system used to process data and perform calculations.

[0677] "Notes" are additional explanations or explanatory texts accompanying information or data, provided to aid understanding.

[0678] An "analytical device" is a system that analyzes data and information and uses the results to draw specific conclusions or make inferences.

[0679] "Emotional state" refers to an internal state that indicates a user's psychological and emotional tendencies and mood.

[0680] This invention is an information processing system that provides information tailored to the user's emotional state. The specific implementation method is described below.

[0681] The user first inputs a specific concept using a terminal. This concept is in text format and includes keywords related to what the user wants to research or the information they need. Upon receiving the input, the terminal sends the information to the server using an encrypted communication protocol (e.g., SSL / TLS) to maintain security.

[0682] The server uses a computer to find relevant data from its internally held information resources based on the received concepts. The software technology used here employs models with natural language processing capabilities (e.g., generative AI models such as BERT and GPT). Leveraging these models, it has the ability to automatically extract relevant information.

[0683] Furthermore, the server uses an analysis device called an emotion engine to analyze the user's emotional state. This involves a process of evaluating the user's past activity history and current input data to identify specific emotions. For example, if the concept "seeking encouragement" is input, it is determined that the user's input indicates a depressed state.

[0684] Based on the results of this sentiment analysis, the server adjusts the priority and content of the information presented to the user. This adjustment includes processing the selected information to optimize it for the user's state.

[0685] For example, if a user enters a specific concept seeking "comfort," the server's emotion engine infers that the user is feeling down. The server then uses natural language processing technology to prioritize extracting passages and articles that provide a sense of reassurance and improve the user's mood, and sends that content to the device.

[0686] Through this process, users can receive information optimized to align with their own emotional state. An example of a prompt using a generative AI model is: "Based on the concept entered by user A, please generate optimized information while considering their emotional state."

[0687] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0688] Step 1:

[0689] The user uses the terminal to input concepts related to what they want to research. The input concepts are displayed on the terminal screen and simultaneously processed as digital data. The input format is regular text, which is then prepared to be sent to the server as the initial data.

[0690] Step 2:

[0691] The terminal transmits the input concept to the server over the network. During this process, the data is encrypted using encryption protocols such as SSL / TLS to ensure security during transmission. The output of this process is encrypted concept data.

[0692] Step 3:

[0693] The server decrypts the received encrypted data and extracts the original concept. Based on this concept, the server launches a search engine to find related data from information resources. Here, the input is the decrypted concept, and the output is the identifier of the related information resource.

[0694] Step 4:

[0695] The server uses a generative AI model equipped with natural language processing technology to automatically extract appropriate passages from relevant information resources. The input is the identifier of the searched information resource, and the output is an optimized list of passages to be presented to the user. In this process, the model performs text summarization and keyword analysis.

[0696] Step 5:

[0697] An emotion engine embedded in the server analyzes input concepts and past interactions to infer the user's emotional state. The input consists of the user's operation history and input concepts, and the output is an inferred emotional state of the user. This process involves algorithmic data analysis.

[0698] Step 6:

[0699] The server adjusts the priority of the optimized passage list based on the sentiment engine's predictions. The input is the sentiment state and the passage list, and the output is the re-prioritized passage list. Specific operations include sorting and filtering the list.

[0700] Step 7:

[0701] The server sends the adjusted information to the terminal, which then presents this information to the user through the user interface. The input here is the adjusted passage list, and the output is the information displayed on the user's screen. Specifically, the font and color on the UI may change according to the user's emotions.

[0702] (Application Example 2)

[0703] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0704] Improving the customer experience is a crucial challenge in modern brick-and-mortar stores. Traditional customer service methods have made it difficult for staff to accurately understand customer emotions and respond accordingly. In particular, providing nuanced responses that respond to customer emotions quickly is challenging, which can lead to decreased customer satisfaction. To solve this problem, there is a need for a system that provides appropriate responses based on customer emotions in real time.

[0705] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0706] In this invention, the server includes an analysis device for inferring the user's emotional state, a control device for adjusting the content of search results and automatically generated annotations based on the analysis device's inference, and a device for providing appropriate responses to staff in real time. This enables quick and appropriate customer service responses that correspond to the customer's emotional state.

[0707] A "user" is a person who uses this system to input and retrieve information.

[0708] "Keywords" are words or phrases that users use to search for information.

[0709] A "device" is a collection of hardware or software designed to perform a specific function.

[0710] A "data set" is a collection of multiple data points that contain specific information.

[0711] A "processing device" is a device that analyzes input data and generates output according to its intended purpose.

[0712] A "display device" is a device used to visually present information to a user.

[0713] An "analytical device" is a device used to examine data and find meaning or patterns in it.

[0714] A "control device" is a device that manages the entire system and directs the operation of each component.

[0715] "Staff" refers to those who interact with customers through this system.

[0716] A "response" is a system's reaction to input from a user or customer.

[0717] The system that realizes this application example is based on devices used by users in physical stores. When a user wears smart glasses or a head-mounted display and interacts with a customer, these devices capture the customer's facial expressions and speech in real time. A server receives this data and analyzes it using emotion recognition AI and natural language processing (NLP) technology.

[0718] Specifically, the system uses the Google Cloud Vision API to analyze customer facial expressions and processes customer speech using OpenAI's generative AI model, GPT. This allows the server to infer the customer's emotional state and generate an appropriate response based on that. This response is then displayed on a device worn by the staff to facilitate smooth customer service.

[0719] As a concrete example of its use, if a customer expresses dissatisfaction with a product, the system detects that emotion and suggests a response in real time to the staff, such as "apologize immediately and ask how they can help." This enables a quick and appropriate response to the customer, thereby improving customer satisfaction.

[0720] The following are specific examples of prompt statements to be input to a generative AI model:

[0721] "The customer appears tired. Please think of and offer words that will help them feel at ease and relax."

[0722] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0723] Step 1:

[0724] The user captures the customer's facial expressions and voice during interactions with the customer through smart glasses or a head-mounted display. This data is transmitted in real time from the user's device to a server. The input is the customer's facial expression and voice data, and the output is the data transmission to the server.

[0725] Step 2:

[0726] The server analyzes the received facial expression data using the Google Cloud Vision API. The data processing performed here involves converting the facial expression data into structured data and adding emotion tags (e.g., anger, happiness, dissatisfaction). The input is facial expression data, and the output is data with emotion tags.

[0727] Step 3:

[0728] The server converts received audio data into text and performs natural language processing on the text using OpenAI's generative AI model. Based on specific keywords and context, it analyzes the customer's intent and emotions. The input is text converted from audio data, and the output is an inference of emotional state based on text analysis.

[0729] Step 4:

[0730] The server integrates the results from steps 2 and 3 to determine the customer's overall emotional state. Based on this determination, it generates an appropriate response for the user. The input is emotion-tagged data and text analysis results, and the output is the text as a proposed response.

[0731] Step 5:

[0732] The generated response is transmitted to the user's smart glasses or head-mounted display for visual presentation. Based on this, the user can respond to the customer quickly and appropriately. The input is the text of the response, and the output is the presentation of information to the user.

[0733] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0734] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0735] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0736] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0737] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0738] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0739] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0740] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0741] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0742] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0743] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0744] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0745] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0746] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0747] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0748] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0749] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0750] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0751] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0752] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0753] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0754] The following is further disclosed regarding the embodiments described above.

[0755] (Claim 1)

[0756] A terminal means for receiving keywords entered by the user,

[0757] A server means for searching for related document sets based on the aforementioned keywords,

[0758] A terminal means that displays the search results obtained by the server means and automatically generates related annotations,

[0759] An information processing system that includes this.

[0760] (Claim 2)

[0761] The information processing system according to claim 1, wherein the server means searches the document group using natural language processing technology.

[0762] (Claim 3)

[0763] The information processing system according to claim 1, wherein the annotation is generated to include multiple viewpoints.

[0764] "Example 1"

[0765] (Claim 1)

[0766] An information terminal means for the user to input information,

[0767] A computing device means for processing the aforementioned information and searching for related data,

[0768] Information terminal means that displays the information retrieved by the aforementioned computing device means and automatically generates related additional information,

[0769] A system that includes this.

[0770] (Claim 2)

[0771] The system according to claim 1, wherein the computing device means retrieves the data using language processing technology.

[0772] (Claim 3)

[0773] The system according to claim 1, wherein the additional information is generated to include multiple perspectives.

[0774] "Application Example 1"

[0775] (Claim 1)

[0776] A device means for receiving a concept input by a user,

[0777] A computing means for searching for related information sets based on the above concept,

[0778] A device means that presents the results searched by the aforementioned computer means and automatically generates related supplementary information,

[0779] A visual display means for visually highlighting the results to the user,

[0780] A system that includes this.

[0781] (Claim 2)

[0782] The system according to claim 1, wherein the computing means searches the information set using machine learning techniques.

[0783] (Claim 3)

[0784] The system according to claim 1, wherein the supplementary information is generated to include multiple perspectives.

[0785] "Example 2 of combining an emotion engine"

[0786] (Claim 1)

[0787] A device that receives concepts input by the user,

[0788] A computer for searching for relevant information resources based on the above concept,

[0789] A device that displays the search results obtained by the aforementioned computer and automatically generates related notes,

[0790] An analytical device for inferring the emotional state of a user,

[0791] A device that adjusts the information presented based on the prediction results of the aforementioned analytical device,

[0792] A system that includes this.

[0793] (Claim 2)

[0794] The system according to claim 1, wherein the computer searches for the information resources using natural language processing technology and further generates a prompt sentence using a generative AI model.

[0795] (Claim 3)

[0796] The system according to claim 1, wherein the aforementioned note is customized to match the user's emotional state.

[0797] "Application example 2 when combining with an emotional engine"

[0798] (Claim 1)

[0799] A device that receives keywords entered by the user,

[0800] A processing device for searching for a related data set based on the aforementioned keywords,

[0801] A display device that displays the results searched by the aforementioned processing device and automatically generates related annotations,

[0802] An analytical device for inferring the emotional state of a user,

[0803] A control device that adjusts the content of search results and automatically generated annotations based on the predictions of the aforementioned analysis device,

[0804] A device that provides staff with appropriate responses in real time,

[0805] A system that includes this.

[0806] (Claim 2)

[0807] The system according to claim 1, wherein the processing device searches the data set using natural language processing technology.

[0808] (Claim 3)

[0809] The system according to claim 1, wherein the annotations are generated to include multiple perspectives and are customized based on emotional states. [Explanation of Symbols]

[0810] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A terminal means for receiving keywords entered by the user, A server means for searching for related document sets based on the aforementioned keywords, A terminal means that displays the search results obtained by the server means and automatically generates related annotations, An information processing system that includes this.

2. The information processing system according to claim 1, wherein the server means searches the document group using natural language processing technology.

3. The information processing system according to claim 1, wherein the annotation is generated to include multiple viewpoints.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A