system
The system addresses the lack of personalized and emotional response in existing systems by using natural language processing and emotion recognition to provide tailored religious support, enhancing user interaction and accuracy through feedback integration.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-16
- Publication Date
- 2026-04-28
AI Technical Summary
Existing systems fail to provide quick, accurate, and personalized responses to religious doubts and personal troubles, lacking resources for diverse religious backgrounds and failing to consider user emotions and feedback for improved accuracy.
A system that acquires user input as text or voice, analyzes it using natural language processing based on religious texts, generates appropriate responses, and presents them visually or audibly, while incorporating feedback for improvement, and includes emotion recognition to tailor responses to individual emotional states.
Enables quick, accurate, and personalized responses to religious questions and personal concerns, improving user experience through emotional state analysis and feedback integration.
Smart Images

Figure 2026071016000001_ABST
Abstract
Description
Technical Field
[0001] The technology of this disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, there has been a problem of insufficient resources that can quickly and accurately respond to religious doubts and personal troubles. Face-to-face consultations with experts have time and geographical constraints, and it is difficult for many people to obtain the information and spiritual support they need. In addition, there is a problem that counselors with knowledge of appropriately understanding and applying a large number of religious doctrines are limited. Therefore, there is a need to develop a system that allows users with diverse religious backgrounds to easily and efficiently obtain the necessary information.
Means for Solving the Problems
[0005] This invention provides means for acquiring user input and transmitting it as text data, as well as means for analyzing user input using natural language processing technology based on multiple religious texts and generating an appropriate response. Furthermore, by providing means for visually or audibly presenting this generated response to the user, it is possible to respond quickly and accurately to a wide range of religious questions and personal concerns. In addition, by providing means for converting voice input into text data and means for collecting user feedback to improve the system's response generation means, a more user-friendly and reliable platform is provided.
[0006] A "user" refers to an individual who uses the system to seek advice on religious questions or personal problems.
[0007] "Input" refers to information about questions or concerns that the user provides to the system in text or voice.
[0008] "Text data" refers to character information that has been converted into a digital format from user input.
[0009] "Means of transmission" refers to the function for transferring input data from the terminal to the server.
[0010] "Religious texts" refer to books and documents that contain doctrines and teachings related to various religions.
[0011] "Natural language processing technology" refers to the technology used to understand and analyze human language using computers.
[0012] "Analysis" refers to the entire process of understanding the meaning and intent of input data using natural language processing techniques.
[0013] "Response" refers to information from the system, including explanations and advice in response to user questions or concerns.
[0014] "The means for generating" refers to the function for creating an appropriate response based on the analyzed user input.
[0015] "The means for presenting" refers to the function for providing the generated response to the user in a visible or audible form.
[0016] "Voice input" refers to the form in which the user conveys questions or concerns to the system by voice.
[0017] "The means for converting" refers to the technology for converting voice data into text data.
[0018] "Feedback" refers to the evaluation or opinion given by the user regarding the system's response.
[0019] "The means for utilization in improvement" refers to the function for improving the system's performance and user experience based on the collected feedback.
Brief Description of the Drawings
[0020] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8]It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Embodiments for Carrying Out the Invention
[0021] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0022] First, the terms used in the following description will be explained.
[0023] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0024] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0025] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0026] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0028] [First Embodiment]
[0029] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0030] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0031] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0032] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0033] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0035] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0036] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0037] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0038] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0039] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0040] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0041] This invention provides a system that offers a platform for users to consult about religious questions or personal concerns. This system operates online, allowing users to input questions via text or voice, and the system then provides responses based on religious teachings. The embodiments of this system are described in detail below.
[0042] First, the user accesses the platform using a device. The device receives user input and sends the consultation details to the server in text or voice format. If voice input is received, the device converts the voice into text data and sends it to the server.
[0043] The server is equipped with a natural language processing engine for analyzing incoming text data. Using this engine, the server analyzes the input text and determines its intent and sentiment. Based on this analysis, the server consults a database of multiple religious texts and extracts the most relevant information. Based on this information, the server generates the most appropriate response for the user.
[0044] The generated response is sent from the server to the terminal. The terminal then presents the response to the user. This presentation can be displayed visually as text, or audibly using speech synthesis technology.
[0045] As a concrete example, consider a scenario where a user enters a question such as, "How can one maintain mental stability in modern times?" The server analyzes this question, and, if necessary, references relevant religious teachings, such as, "There are ways to maintain mental balance through daily gratitude and meditation," to generate a response. This response is then provided to the user, allowing them to ask further questions or provide feedback.
[0046] In addition, the system includes a function to collect user feedback, which the server analyzes and uses to improve the accuracy of future responses. Through this process, it becomes possible to provide users with more personalized, accurate, and useful information.
[0047] Therefore, this system has the ability to provide appropriate support to users regarding their religious questions and personal concerns, regardless of time or place.
[0048] The following describes the processing flow.
[0049] Step 1:
[0050] The user accesses the platform using their device and creates or logs in an account. The device retrieves the entered authentication information and sends it to the server.
[0051] Step 2:
[0052] The server compares the received authentication information with the database to determine whether authentication was successful. If successful, it generates an authentication token and sends it back to the terminal.
[0053] Step 3:
[0054] The user enters their religious questions or personal concerns into the device as text or voice. The device receives the input and converts it into text data if it is voice.
[0055] Step 4:
[0056] The terminal sends text data to the server. The server inputs the received consultation content into a natural language processing engine and analyzes the intent and emotions.
[0057] Step 5:
[0058] Based on the analysis results, the server consults a religious text database and extracts relevant information. The server then generates a response based on that information.
[0059] Step 6:
[0060] The server sends the generated response to the terminal. The terminal displays the response on its screen and, if audio output is selected, presents it audibly using speech synthesis technology.
[0061] Step 7:
[0062] Users can enter feedback or additional questions regarding the response they receive. The terminal collects this information and sends it to the server.
[0063] Step 8:
[0064] The server stores the collected feedback in a database and uses it as data to improve the accuracy of response generation.
[0065] (Example 1)
[0066] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0067] The challenge lies in providing a platform where users can receive appropriate support for religious questions and personal concerns, regardless of time or location. Furthermore, existing systems have been criticized for lacking accuracy and personalization in their responses based on user input. Additionally, there is a need to accurately process user voice input while ensuring security, and to improve response accuracy through feedback.
[0068] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0069] In this invention, the server includes means for acquiring user input and transmitting it as information data, means for analyzing the user input using natural language processing technology based on multiple religious data and generating an appropriate response, means for presenting the generated response to the user visually or audibly, and means for acquiring user feedback and improving the accuracy of the response generation means. This makes it possible to provide the user with a more appropriate and personalized response. Furthermore, by accurately processing voice input while improving the security of communication, the user experience can be improved.
[0070] A "user" refers to someone who accesses the system to ask religious questions or seek advice on personal problems.
[0071] "Input" refers to information that a user provides to the system through their device, and includes information in text or audio format.
[0072] "Information data" refers to data that has been processed from user input and converted into a digital format that can be transmitted by the system.
[0073] "Religious data" refers to a collection of information, including religious teachings and texts, that the system references when responding to user inquiries.
[0074] "Natural language processing technology" refers to algorithms and methods that systems use to interpret user input and analyze its meaning and intent.
[0075] "Response" refers to the reply or information that the system generates based on user input.
[0076] "Presenting visually or audibly" refers to a method of presentation where the response is displayed to the user as text on the screen or heard as audio.
[0077] "Feedback" refers to information that users convey to the system regarding their evaluations and opinions on the responses provided.
[0078] "Communication technology" refers to the technologies and protocols used to securely transmit user input data from a terminal to a server.
[0079] This invention is a system that uses information technology to provide appropriate responses to users' religious questions and personal concerns. The detailed configuration for implementing this system is described below.
[0080] The user first accesses the online platform using a terminal. This terminal is responsible for receiving text or voice input from the user. If voice input is provided, the terminal uses speech recognition software (e.g., a speech recognition API) to convert the speech into text. The converted text data is then sent to the server as information data. Secure communication technology (e.g., HTTPS protocol) is used for this transmission.
[0081] Next, the server analyzes the received data using natural language processing techniques (e.g., machine learning models). The server analyzes the grammatical and semantic structure of the text entered by the user to infer the user's intent and emotions. During this analysis, it understands the user's questions and searches for relevant "religious data." This data consists of doctrines and texts stored in a database.
[0082] The server then runs a natural language generation model (e.g., a generative AI model) to generate a response. This model receives the analyzed input and generates an appropriate and specific response to the user's question. The server is designed to take user feedback into consideration during this process to improve the accuracy of the response.
[0083] The generated response is sent from the server to the terminal, which can then present it to the user either visually as text or audibly using speech synthesis technology (e.g., a speech conversion API). This method of presenting the response is designed to enhance system flexibility and ensure easy user comprehension.
[0084] As a concrete example, consider a scenario where a user inputs the question, "How can one maintain mental stability in modern times?" The server analyzes this question and generates an appropriate response based on religious teachings that suggest daily gratitude and meditation are helpful in maintaining mental balance. This response has the potential to offer the user further reflection and new perspectives.
[0085] Another example of a prompt for a generative AI model is: "The user has asked about how to maintain mental balance. Refer to relevant religious teachings and generate an appropriate response."
[0086] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0087] Step 1:
[0088] Users use their devices to access online platforms and input their religious questions or personal concerns in text or voice format. The input data is received by the device and processed according to its format. In the case of voice input, the device uses speech recognition to convert the voice into text data and prepares it as information data.
[0089] Step 2:
[0090] The device transmits text data converted by speech recognition or text data entered by the user to the server using secure communication technology. Privacy is maintained through data encryption during transmission. The output is the information data sent to the server.
[0091] Step 3:
[0092] The server analyzes the received data using natural language processing techniques. The input is user text data; by analyzing this data, the server understands its grammatical and semantic structure and performs data processing to infer the user's intentions and emotions. The output is the analysis result.
[0093] Step 4:
[0094] The server searches for relevant information from religious databases based on the analysis results. This search process applies an efficient search algorithm to extract the most relevant data. The input is the analysis results, and the output is the relevant information.
[0095] Step 5:
[0096] The server generates an appropriate response using a natural language generation model based on relevant information. In this process, a generative AI model operates to create appropriate and specific answers to the user's questions. The input is relevant information, and the output is the generated response.
[0097] Step 6:
[0098] The server sends a generated response to the terminal. The terminal either displays this response as text on the screen or conveys it to the user audibly using speech synthesis. This process allows the user to receive the response visually or audibly. The input is the generated response, and the output is the response displayed or presented audibly to the user.
[0099] Step 7:
[0100] The user sends feedback on the response provided through the terminal. This feedback is collected as data on the terminal and sent back to the server. This feedback is then analyzed by the server to help improve the accuracy of the response. The input is the user's feedback, and the output is the data used by the server for accuracy improvement.
[0101] (Application Example 1)
[0102] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0103] In modern society, users face a great deal of stress and anxiety in their daily lives. In such circumstances, there is a need for a system that can understand the user's mental state in real time and provide appropriate guidance. However, current technology lacks the means to accurately analyze a user's emotional state and provide religious or philosophical guidance tailored to their individual circumstances.
[0104] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0105] In this invention, the server includes means for acquiring user input and transmitting it as text data; means for analyzing the user input using natural language processing technology based on multiple information resources and generating an appropriate response; means for presenting the generated response to the user visually or audibly; means for measuring the user's biometric information and analyzing their emotional state; and means for generating and presenting appropriate guidance based on the emotional state. This enables real-time support for the user's mental health and the provision of personalized information.
[0106] "User input" refers to information provided by the user in voice or text format.
[0107] "Text data" refers to strings of information expressed in a format that can be processed by a computer.
[0108] An "information resource" is a collection of knowledge and data, including information based on specific teachings or instruction.
[0109] "Natural language processing technology" is a technology that enables computers to understand and analyze human language.
[0110] An "appropriate response" is information, including solutions and guidance, that is generated based on user input.
[0111] "Presenting visually or audibly" refers to a method of conveying information through the user's sight or hearing.
[0112] "Biometric information" refers to data that indicates the user's physical condition, including pulse rate and facial expressions.
[0113] "Analyzing emotional state" is the process of estimating a user's mental state and emotions from their biometric information.
[0114] "Generating and presenting guidance" means creating and providing advice and information tailored to the user's situation.
[0115] This invention provides a system that allows users to consult about religious questions or personal concerns using a terminal. The terminal receives input from the user as voice or text data and sends it to a server. In the case of voice input, the terminal first converts the voice into text data. Software used includes Google® Cloud Speech-to-Text API, among others.
[0116] The server analyzes the received text data using natural language processing techniques to determine the intent and sentiment of the input text. This analysis utilizes open-source natural language processing libraries such as NLTK and Transformers. After analysis, the server consults multiple information resources to extract the information most relevant to the user's input. Based on this information, a generative AI model generates appropriate guidance.
[0117] Next, the server sends the generated instructions to the terminal, which then presents them to the user visually or audibly. Presentation methods include text display on a screen and auditory presentation using a speech synthesis system. Amazon Polly, for example, can be used for speech synthesis.
[0118] Furthermore, this system has the function of measuring the user's biometric information and analyzing their emotional state. Sensors installed in the terminal detect pulse rate and facial expressions, and a server processes this data to determine the user's emotional state. Based on the determination, appropriate guidance is generated.
[0119] For example, if a user asks, "How can I maintain peace of mind?", the system will generate and provide guidance such as, "There are ways to maintain mental balance through meditation and daily gratitude."
[0120] Example prompt for the AI model: "What kind of guidance would be appropriate if the user is feeling anxious?"
[0121] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0122] Step 1:
[0123] The device accepts user input. The user enters religious questions or concerns in voice or text format, and the device receives this input data. In the case of voice input, the device uses the Google Cloud Speech-to-Text API to convert the voice into text data. The input is then sent to the server as text data.
[0124] Step 2:
[0125] The server analyzes the received text data. Using the received text data as input, the server performs natural language processing using NLTK and Transformers to determine the user's intent and emotions. As a result of the analysis, tags representing the user's question intent and the sentiment analysis results are output.
[0126] Step 3:
[0127] The server references information resources and extracts appropriate information. Based on the analysis results, the server searches multiple information resources (including religious teachings and philosophical guidance) and selects the most relevant information. In this step, a search is performed using the prompt "What guidance is appropriate when a user is feeling anxious?". As a result, relevant guidance information is output.
[0128] Step 4:
[0129] The server generates instruction using a generated AI model. Using the extracted information as input, the generated AI model generates instruction appropriate for the user. Specific instruction content is provided as output.
[0130] Step 5:
[0131] The server sends the generated instruction to the terminal. Using the generated instruction as input, the server sends data to the terminal. The instruction data arrives at the terminal as output.
[0132] Step 6:
[0133] The device presents instructions to the user. The device presents the received instruction data to the user audibly using text display on the screen or speech synthesis. Amazon Polly can be used for this presentation.
[0134] Step 7:
[0135] The device measures the user's biometric information and analyzes their emotional state. The device's sensors capture the user's pulse and facial expressions, and transmit this data to a server. The server analyzes this biometric information to determine the user's emotional state. The output is an evaluation of the emotional state.
[0136] Step 8:
[0137] The server generates further guidance based on the emotional state and sends it to the device. Based on the evaluation of the emotional state, the server generates additional guidance as needed. Based on this, new guidance data is provided to the device.
[0138] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0139] This invention provides a system that takes into account the user's emotional state when addressing their religious questions or personal concerns. This system operates on an online platform and enables more personalized information delivery through user interaction.
[0140] When a user accesses the platform using their device, they first undergo login authentication. Users can then input religious questions or personal concerns via text and voice. In the case of voice input, the device converts the voice into text data and sends it to the server.
[0141] The server analyzes each input using a natural language processing engine to recognize the user's intent and emotional state. This analysis utilizes an emotion engine, which extracts emotional information from the input text to understand the user's feelings. For example, if words or phrases indicating stress or anxiety are included, the server recognizes them and adjusts the response accordingly.
[0142] Based on the analysis results, the server consults multiple religious text databases and selects the information most appropriate for the user. The tone and content of the response are flexibly adjusted according to the user's emotional state. For example, if the server determines the user is depressed, comforting words and encouraging messages will be selected.
[0143] The generated response is sent from the server to the terminal and presented to the user. Presentation methods include not only text display but also auditory feedback using speech synthesis technology. The user can provide feedback on the response, which the terminal reports to the server.
[0144] The emotion engine uses machine learning based on this feedback data to improve the accuracy of emotion recognition. For example, when a user inputs "I feel lonely," the emotion engine accurately recognizes the emotion of "loneliness," and as a result, the server generates a response such as "You are not alone; many people support you." In this way, the present invention provides a new form of religious support by realizing information provision that is attentive to the user's emotions.
[0145] The following describes the processing flow.
[0146] Step 1:
[0147] The user accesses the platform via their device, enters their account information, and logs in. The device retrieves this authentication information and sends it to the server.
[0148] Step 2:
[0149] The server compares the received authentication information with the authentication server or database to determine whether the login was successful. If authentication is successful, it starts a session and prepares it for use by the user.
[0150] Step 3:
[0151] Users send their religious questions or personal concerns to the device using either text or voice input. If voice input is selected, the device converts the voice into text data.
[0152] Step 4:
[0153] The terminal sends the converted text data to the server. The server receives this data, uses natural language processing (NLP) to analyze the input content, and extracts the intent.
[0154] Step 5:
[0155] The server uses an emotion engine to recognize emotional states from text data. Here, emotions are classified as positive, negative, or neutral to understand the user's feelings.
[0156] Step 6:
[0157] Based on the analysis results, the server accesses multiple religious text databases and extracts relevant teachings and information. Simultaneously, it adjusts the tone and context of its responses based on the perceived emotional state.
[0158] Step 7:
[0159] The server packages the generated response in text and audio formats and sends it to the terminal. This allows the user to provide feedback in their preferred format.
[0160] Step 8:
[0161] The terminal presents the received response to the user. If it is in text format, it is displayed on the screen; if it is in audio format, it provides auditory feedback using speech synthesis technology.
[0162] Step 9:
[0163] Users can input their thoughts and feedback on the response. The device sends this feedback to the server, which uses the data to train the emotion engine and further improve its accuracy.
[0164] (Example 2)
[0165] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0166] In society, it is not uncommon for individuals to have religious questions or personal concerns, but there are limited systems that adequately address these. In particular, when emotional considerations are necessary, simple information provision is insufficient. Responses provided without considering emotional states fail to adequately provide the reassurance and support that users seek. Therefore, there is a need for systems that incorporate emotion recognition technology to provide more personalized responses.
[0167] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0168] In this invention, the server includes means for acquiring user input and transmitting it as information, means for analyzing the user input using natural language processing technology based on the document and generating an appropriate response, means for displaying or audibly presenting the generated response to the user, means for analyzing the user's emotional state and adjusting the response content based on the results, and means for automatically generating a response using a generative AI model. This makes it possible to appropriately provide responses that take the user's emotions into consideration and to provide religious support more effectively.
[0169] A "user" is an individual who uses the system to resolve religious questions or personal problems.
[0170] "Input" refers to information that a user transmits to a system using a device, and can be in text or audio format.
[0171] "Information" refers to data obtained by converting user input into a digital format, which is then analyzed and processed within the system.
[0172] A "document" is a collection of text data containing religious or related content used to generate a response.
[0173] "Natural language processing technology" refers to computer program techniques used to analyze user input and understand its meaning and intent.
[0174] An "appropriate response" refers to information or messages that reflect the user's emotional state and intentions based on their input.
[0175] A "generative AI model" is a type of program that uses machine learning techniques to autonomously generate text and voice responses.
[0176] "Emotional state" refers to information that indicates the psychological or emotional condition of a user.
[0177] "Presenting by display or sound" refers to a method of providing information to a user visually or audibly.
[0178] This invention is a system that addresses users' religious questions and personal concerns and provides responses. Users access an online platform using a terminal and input their concerns or questions. Input can be in text or voice format; in the case of voice input, the terminal uses speech recognition technology to convert it into text. For this speech recognition, speech recognition software is used as a standard service.
[0179] The server receives input data sent from the terminal and performs analysis using natural language processing technology. During the analysis process, a generative AI model is used to recognize the user's intent and emotional state. This model may utilize a natural language model such as OpenAI®. To understand the emotional state, an emotion analysis algorithm is employed to evaluate the user's emotions. This makes it possible to dynamically adjust the response content according to the user's psychological state.
[0180] The server references multiple religious text databases based on the analysis results and selects the information most relevant to the user. The selected information is then used to generate a response tailored to the user's emotional state and intentions. This response is automatically generated by a generative AI model, enabling a more human-like interaction.
[0181] The generated response is sent from the server to the terminal and presented to the user. In addition to text display, voice feedback is also provided using speech synthesis technology. A speech synthesis engine is used for speech synthesis, allowing the user to have a more interactive experience.
[0182] Users provide feedback on the responses they receive via their devices, and the server uses this feedback data to perform machine learning to improve the accuracy of sentiment analysis. This improves the overall response accuracy of the system, enabling the provision of more accurate information.
[0183] For example, if a user inputs "I feel lonely," the system recognizes the emotional state and generates a response such as "You are not alone; many people support you." Furthermore, by presenting the AI model with prompts such as "How should I respond if the user inputs 'I feel anxious'?", it is possible to train the model to provide appropriate responses.
[0184] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0185] Step 1:
[0186] The user accesses the online platform using a terminal and logs in. The user enters their ID and password, goes through an authentication process, and is granted access to the platform. In this step, the input is the user's authentication information, and the output is a message indicating authentication success or failure. Specifically, the authentication server verifies the user information.
[0187] Step 2:
[0188] The user inputs a religious question or concern. The input method is either text or voice; in the case of voice, the device uses speech recognition software to convert the voice into text data. In this step, the input is the user's question or concern, either voice or text, and the output is the converted text data. Specifically, the speech recognition engine converts the voice to text in real time.
[0189] Step 3:
[0190] The terminal sends the converted text data to the server. The server passes the received data to a natural language processing engine, which analyzes the user's intent and emotional state. This analysis uses a generative AI model and an emotion recognition algorithm. The input is text data, and the output is data indicating the user's intent and emotional state. Specifically, the natural language processing engine analyzes the text and categorizes its intent.
[0191] Step 4:
[0192] The server selects appropriate information from multiple religious text databases based on the analysis results. A generative AI model is used to generate responses tailored to the user's emotional state. The input is the analyzed intent and emotional state, and the output is a customized response text. Specifically, database queries are executed within the server to retrieve the relevant text.
[0193] Step 5:
[0194] The server sends the generated response to the terminal. The terminal presents it to the user using text display or speech synthesis technology. For speech presentation, a speech synthesis engine is used. The input is the generated response text, and the output is visual or auditory feedback to the user. Specifically, the terminal converts the text into speech so that the user can hear it.
[0195] Step 6:
[0196] The user provides feedback on the presented response. The device sends this feedback back to the server. The server uses the feedback data to perform machine learning to improve the accuracy of the sentiment recognition engine. The input is the user's feedback, and the output is the improved sentiment analysis model. Specifically, the server analyzes the feedback data and updates the algorithm of the generative AI model.
[0197] (Application Example 2)
[0198] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0199] Existing information delivery systems often struggle to adequately consider users' personal needs and emotional states. As a result, they can only provide consistent information in response to user questions, failing to offer appropriate responses tailored to individual emotions and interests, thus limiting the user experience. Especially in online business transactions and consultations, there is a need for systems that understand user emotions and provide personalized suggestions.
[0200] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0201] In this invention, the server includes means for acquiring user input and transmitting it as information data; means for analyzing the user input using natural language processing technology based on multiple religious texts and generating an appropriate response; means for presenting the generated response to the user visually or audibly; means for analyzing the user's emotional state and adjusting the response content accordingly; and means for suggesting information on products and services of interest to the user along with the appropriate response. This makes it possible to provide highly personalized responses and suggestions that are tailored to the user's individual emotions and interests.
[0202] "Means for acquiring user input and transmitting it as information data" refers to a function for collecting information provided by the user and transmitting it as digital data to a processing device or server.
[0203] "A means of analyzing user input using natural language processing technology based on multiple religious texts and generating appropriate responses" refers to a function that uses a wide range of religious literature as a database, interprets input information using a computer, and creates information accordingly.
[0204] "Means for presenting the generated response to the user visually or audibly" refers to functions for providing the created response to the user through sight or hearing, and this includes display devices and audio output devices.
[0205] "Means for analyzing the user's emotional state and adjusting the response content based on that" refers to a function that determines the user's psychological state and optimizes the content and expression of the response accordingly.
[0206] "A means of suggesting information about products and services that users are interested in, along with appropriate responses," refers to a function that selects and recommends products and services that meet the user's needs and interests.
[0207] "Means of converting user voice input into information data" refers to technology that transforms a user's words from speech into digital representations.
[0208] "Means for collecting user feedback and using it to improve response generation and suggestion provision means" refers to technologies for gathering user reactions and opinions and using them to improve the performance and functionality of the system.
[0209] In implementing this invention, smart glasses or head-mounted displays are used as user interface devices. These devices are capable of presenting information in accordance with the user's visual and auditory perception. Specific hardware includes smart glasses devices, voice input devices, and wireless communication devices for data transmission.
[0210] First, the user provides input through an interface device. This is provided as voice input and is captured by the device's built-in microphone. The captured voice data is then converted into text data by the device's speech recognition module. In this process, speech recognition software such as Google Cloud Speech-to-Text is used as an example.
[0211] Next, the generated text data is sent to the server. Here, the server analyzes the input data using a natural language processing engine. This analysis includes, for example, language analysis techniques using spaCy. Furthermore, a custom model utilizing TENSORFLOW® is used for sentiment analysis to determine the user's emotional state.
[0212] Based on the analysis results, the server generates an appropriate response. The response is tailored based on multiple religious texts and personalized information. The generated response's tone and content are modified according to the user's emotional state. Specific products or services may also be suggested based on the user's interests.
[0213] Ultimately, this response is presented to the user visually or audibly. Presentation methods include text display on a screen or voice output using speech synthesis software (e.g., Amazon Polly).
[0214] As a concrete example of its use, if a user is in a situation where they are "looking for a vegan recipe book and feeling a little tired," this system will take the user's emotional state into consideration and provide optimal book information while suggesting products that have a relaxing effect.
[0215] An example of a prompt message would be, "I'm looking for a vegan recipe book and I'm a little tired. What books would you recommend?"
[0216] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0217] Step 1:
[0218] The user provides voice input via an interface device.
[0219] The voice provided by the user is captured by the device's built-in microphone. This input data is in audio format and is processed as input to the speech recognition module.
[0220] Step 2:
[0221] The device converts the user's voice into text data.
[0222] Speech recognition software is used to analyze audio data, and as a result, text data is generated. This conversion process uses lexical analysis and acoustic models to accurately transcribe the user's speech into text.
[0223] Step 3:
[0224] The terminal sends the converted text data to the server.
[0225] The generated text data is transferred to the server using wireless communication technology. This data contains information including the user's intent and the content of their question.
[0226] Step 4:
[0227] The server analyzes the text data using a natural language processing engine.
[0228] The server uses natural language processing techniques to analyze the language structure of the received text data and interpret the user's intent. Semantic and sentiment analysis are performed to recognize the user's psychological state and requests.
[0229] Step 5:
[0230] The server generates a response based on the user's emotional state.
[0231] Based on the analyzed data, the server generates an appropriate response. This response utilizes an emotion engine to adjust its tone and content, and may include relevant products or services as needed.
[0232] Step 6:
[0233] The server sends the generated response to the terminal.
[0234] The generated response is then transmitted back to the terminal via wireless communication. This output data contains information presented to the user.
[0235] Step 7:
[0236] The device presents a response to the user, either visually or audibly.
[0237] The device presents the received data to the user using text display on the screen or speech synthesis functionality. This allows the user to acquire information through sight or hearing.
[0238] Step 8:
[0239] Users provide feedback on the response.
[0240] Users provide feedback by entering their opinions and impressions about the information presented. This information is used to continuously improve the system.
[0241] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0242] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0243] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0244] [Second Embodiment]
[0245] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0246] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0247] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0248] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0249] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0250] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0251] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0252] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0253] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0254] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0255] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0256] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0257] This invention provides a system that offers a platform for users to consult about religious questions or personal concerns. This system operates online, allowing users to input questions via text or voice, and the system then provides responses based on religious teachings. The embodiments of this system are described in detail below.
[0258] First, the user accesses the platform using a device. The device receives user input and sends the consultation details to the server in text or voice format. If voice input is received, the device converts the voice into text data and sends it to the server.
[0259] The server is equipped with a natural language processing engine for analyzing incoming text data. Using this engine, the server analyzes the input text and determines its intent and sentiment. Based on this analysis, the server consults a database of multiple religious texts and extracts the most relevant information. Based on this information, the server generates the most appropriate response for the user.
[0260] The generated response is sent from the server to the terminal. The terminal then presents the response to the user. This presentation can be displayed visually as text, or audibly using speech synthesis technology.
[0261] As a concrete example, consider a scenario where a user enters a question such as, "How can one maintain mental stability in modern times?" The server analyzes this question, and, if necessary, references relevant religious teachings, such as, "There are ways to maintain mental balance through daily gratitude and meditation," to generate a response. This response is then provided to the user, allowing them to ask further questions or provide feedback.
[0262] In addition, the system includes a function to collect user feedback, which the server analyzes and uses to improve the accuracy of future responses. Through this process, it becomes possible to provide users with more personalized, accurate, and useful information.
[0263] Therefore, this system has the ability to provide appropriate support to users regarding their religious questions and personal concerns, regardless of time or place.
[0264] The following describes the processing flow.
[0265] Step 1:
[0266] The user accesses the platform using their device and creates or logs in an account. The device retrieves the entered authentication information and sends it to the server.
[0267] Step 2:
[0268] The server compares the received authentication information with the database to determine whether authentication was successful. If successful, it generates an authentication token and sends it back to the terminal.
[0269] Step 3:
[0270] The user enters their religious questions or personal concerns into the device as text or voice. The device receives the input and converts it into text data if it is voice.
[0271] Step 4:
[0272] The terminal sends text data to the server. The server inputs the received consultation content into a natural language processing engine and analyzes the intent and emotions.
[0273] Step 5:
[0274] Based on the analysis results, the server consults a religious text database and extracts relevant information. The server then generates a response based on that information.
[0275] Step 6:
[0276] The server sends the generated response to the terminal. The terminal displays the response on its screen and, if audio output is selected, presents it audibly using speech synthesis technology.
[0277] Step 7:
[0278] Users can enter feedback or additional questions regarding the response they receive. The terminal collects this information and sends it to the server.
[0279] Step 8:
[0280] The server stores the collected feedback in a database and uses it as data to improve the accuracy of response generation.
[0281] (Example 1)
[0282] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".
[0283] It is an issue to provide a platform that enables users to receive appropriate support regardless of time or place for religious questions or personal troubles. Also, it has been pointed out that the accuracy and personalization of responses based on user input in previous systems are insufficient. Furthermore, while ensuring security, it is necessary to accurately process the user's voice input and improve the response accuracy using feedback.
[0284] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0285] In this invention, the server includes means for acquiring the user's input and transmitting it as information data, means for analyzing the user's input using natural language processing technology based on a plurality of religious data and generating an appropriate response, means for visually or auditorily presenting the generated response to the user, and means for acquiring the user's feedback and improving the accuracy of the response generation means. As a result, it becomes possible to provide a more appropriate and personalized response to the user. Also, by accurately processing the voice input while improving the security of communication, the user experience can be improved.
[0286] The "user" refers to a person who accesses the system and consults about religious questions or personal troubles.
[0287] The "input" is information provided by the user to the system through the terminal and includes text or voice forms.
[0288] "Information data" refers to data that has been processed from user input and converted into a digital format that can be transmitted by the system.
[0289] "Religious data" refers to a collection of information, including religious teachings and texts, that the system references when responding to user inquiries.
[0290] "Natural language processing technology" refers to algorithms and methods that systems use to interpret user input and analyze its meaning and intent.
[0291] "Response" refers to the reply or information that the system generates based on user input.
[0292] "Presenting visually or audibly" refers to a method of presentation where the response is displayed to the user as text on the screen or heard as audio.
[0293] "Feedback" refers to information that users convey to the system regarding their evaluations and opinions on the responses provided.
[0294] "Communication technology" refers to the technologies and protocols used to securely transmit user input data from a terminal to a server.
[0295] This invention is a system that uses information technology to provide appropriate responses to users' religious questions and personal concerns. The detailed configuration for implementing this system is described below.
[0296] The user first accesses the online platform using a terminal. This terminal is responsible for receiving text or voice input from the user. If voice input is provided, the terminal uses speech recognition software (e.g., a speech recognition API) to convert the speech into text. The converted text data is then sent to the server as information data. Secure communication technology (e.g., HTTPS protocol) is used for this transmission.
[0297] Next, the server analyzes the received data using natural language processing techniques (e.g., machine learning models). The server analyzes the grammatical and semantic structure of the text entered by the user to infer the user's intent and emotions. During this analysis, it understands the user's questions and searches for relevant "religious data." This data consists of doctrines and texts stored in a database.
[0298] The server then runs a natural language generation model (e.g., a generative AI model) to generate a response. This model receives the analyzed input and generates an appropriate and specific response to the user's question. The server is designed to take user feedback into consideration during this process to improve the accuracy of the response.
[0299] The generated response is sent from the server to the terminal, which can then present it to the user either visually as text or audibly using speech synthesis technology (e.g., a speech conversion API). This method of presenting the response is designed to enhance system flexibility and ensure easy user comprehension.
[0300] As a concrete example, consider a scenario where a user inputs the question, "How can one maintain mental stability in modern times?" The server analyzes this question and generates an appropriate response based on religious teachings that suggest daily gratitude and meditation are helpful in maintaining mental balance. This response has the potential to offer the user further reflection and new perspectives.
[0301] Another example of a prompt for a generative AI model is: "The user has asked about how to maintain mental balance. Refer to relevant religious teachings and generate an appropriate response."
[0302] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0303] Step 1:
[0304] The user uses the terminal to access the online platform and inputs their religious questions or personal troubles in text or voice form. The input data is received by the terminal and processed according to its form. In the case of voice input, the terminal uses the voice recognition function to convert the voice into text data and prepares it as information data.
[0305] Step 2:
[0306] The terminal sends the text data converted by voice recognition or the text data input by the user to the server using a secure communication technology. During this transmission, the data is encrypted to maintain privacy. The output is the information data sent to the server.
[0307] Step 3:
[0308] The server analyzes the received information data using natural language processing technology. The input is the user's text data, and by analyzing it, data processing is performed to understand the grammatical and semantic structure of the data and infer the user's intentions and emotions. The output is the analysis result.
[0309] Step 4:
[0310] The server searches for relevant information from the religious database based on the analyzed result. In this search process, an efficient search algorithm is applied to extract the most relevant data. The input is the analysis result, and the output is the relevant information.
[0311] Step 5:
[0312] The server generates an appropriate response using the natural language generation model based on the relevant information. In this process, the generation AI model operates to create an appropriate and specific answer to the user's question. The input is the relevant information, and the output is the generated response.
[0313] Step 6:
[0314] The server sends a generated response to the terminal. The terminal either displays this response as text on the screen or conveys it to the user audibly using speech synthesis. This process allows the user to receive the response visually or audibly. The input is the generated response, and the output is the response displayed or presented audibly to the user.
[0315] Step 7:
[0316] The user sends feedback on the response provided through the terminal. This feedback is collected as data on the terminal and sent back to the server. This feedback is then analyzed by the server to help improve the accuracy of the response. The input is the user's feedback, and the output is the data used by the server for accuracy improvement.
[0317] (Application Example 1)
[0318] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0319] In modern society, users face a great deal of stress and anxiety in their daily lives. In such circumstances, there is a need for a system that can understand the user's mental state in real time and provide appropriate guidance. However, current technology lacks the means to accurately analyze a user's emotional state and provide religious or philosophical guidance tailored to their individual circumstances.
[0320] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0321] In this invention, the server includes means for acquiring user input and transmitting it as text data; means for analyzing the user input using natural language processing technology based on multiple information resources and generating an appropriate response; means for presenting the generated response to the user visually or audibly; means for measuring the user's biometric information and analyzing their emotional state; and means for generating and presenting appropriate guidance based on the emotional state. This enables real-time support for the user's mental health and the provision of personalized information.
[0322] "User input" refers to information provided by the user in voice or text format.
[0323] "Text data" refers to strings of information expressed in a format that can be processed by a computer.
[0324] An "information resource" is a collection of knowledge and data, including information based on specific teachings or instruction.
[0325] "Natural language processing technology" is a technology that enables computers to understand and analyze human language.
[0326] An "appropriate response" is information, including solutions and guidance, that is generated based on user input.
[0327] "Presenting visually or audibly" refers to a method of conveying information through the user's sight or hearing.
[0328] "Biometric information" refers to data that indicates the user's physical condition, including pulse rate and facial expressions.
[0329] "Analyzing emotional state" is the process of estimating a user's mental state and emotions from their biometric information.
[0330] "Generating and presenting guidance" means creating and providing advice and information tailored to the user's situation.
[0331] This invention provides a system that allows users to consult about religious questions or personal concerns using a terminal. The terminal receives input from the user as voice or text data and sends it to a server. In the case of voice input, the terminal first converts the voice into text data. Software used includes the Google Cloud Speech-to-Text API.
[0332] The server analyzes the received text data using natural language processing techniques to determine the intent and sentiment of the input text. This analysis utilizes open-source natural language processing libraries such as NLTK and Transformers. After analysis, the server consults multiple information resources to extract the information most relevant to the user's input. Based on this information, a generative AI model generates appropriate guidance.
[0333] Next, the server sends the generated instructions to the terminal, which then presents them to the user visually or audibly. Presentation methods include text display on a screen and auditory presentation using a speech synthesis system. Amazon Polly, for example, can be used for speech synthesis.
[0334] Furthermore, this system has the function of measuring the user's biometric information and analyzing their emotional state. Sensors installed in the terminal detect pulse rate and facial expressions, and a server processes this data to determine the user's emotional state. Based on the determination, appropriate guidance is generated.
[0335] For example, if a user asks, "How can I maintain peace of mind?", the system will generate and provide guidance such as, "There are ways to maintain mental balance through meditation and daily gratitude."
[0336] Example prompt for the AI model: "What kind of guidance would be appropriate if the user is feeling anxious?"
[0337] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0338] Step 1:
[0339] The device accepts user input. The user enters religious questions or concerns in voice or text format, and the device receives this input data. In the case of voice input, the device uses the Google Cloud Speech-to-Text API to convert the voice into text data. The input is then sent to the server as text data.
[0340] Step 2:
[0341] The server analyzes the received text data. Using the received text data as input, the server performs natural language processing using NLTK and Transformers to determine the user's intent and emotions. As a result of the analysis, tags representing the user's question intent and the sentiment analysis results are output.
[0342] Step 3:
[0343] The server references information resources and extracts appropriate information. Based on the analysis results, the server searches multiple information resources (including religious teachings and philosophical guidance) and selects the most relevant information. In this step, a search is performed using the prompt "What guidance is appropriate when a user is feeling anxious?". As a result, relevant guidance information is output.
[0344] Step 4:
[0345] The server generates instruction using a generated AI model. Using the extracted information as input, the generated AI model generates instruction appropriate for the user. Specific instruction content is provided as output.
[0346] Step 5:
[0347] The server sends the generated instruction to the terminal. Using the generated instruction as input, the server sends data to the terminal. The instruction data arrives at the terminal as output.
[0348] Step 6:
[0349] The device presents instructions to the user. The device presents the received instruction data to the user audibly using text display on the screen or speech synthesis. Amazon Polly can be used for this presentation.
[0350] Step 7:
[0351] The device measures the user's biometric information and analyzes their emotional state. The device's sensors capture the user's pulse and facial expressions, and transmit this data to a server. The server analyzes this biometric information to determine the user's emotional state. The output is an evaluation of the emotional state.
[0352] Step 8:
[0353] The server generates further guidance based on the emotional state and sends it to the device. Based on the evaluation of the emotional state, the server generates additional guidance as needed. Based on this, new guidance data is provided to the device.
[0354] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0355] This invention provides a system that takes into account the user's emotional state when addressing their religious questions or personal concerns. This system operates on an online platform and enables more personalized information delivery through user interaction.
[0356] When a user accesses the platform using their device, they first undergo login authentication. Users can then input religious questions or personal concerns via text and voice. In the case of voice input, the device converts the voice into text data and sends it to the server.
[0357] The server analyzes each input using a natural language processing engine to recognize the user's intent and emotional state. This analysis utilizes an emotion engine, which extracts emotional information from the input text to understand the user's feelings. For example, if words or phrases indicating stress or anxiety are included, the server recognizes them and adjusts the response accordingly.
[0358] Based on the analysis results, the server consults multiple religious text databases and selects the information most appropriate for the user. The tone and content of the response are flexibly adjusted according to the user's emotional state. For example, if the server determines the user is depressed, comforting words and encouraging messages will be selected.
[0359] The generated response is sent from the server to the terminal and presented to the user. Presentation methods include not only text display but also auditory feedback using speech synthesis technology. The user can provide feedback on the response, which the terminal reports to the server.
[0360] The emotion engine uses machine learning based on this feedback data to improve the accuracy of emotion recognition. For example, when a user inputs "I feel lonely," the emotion engine accurately recognizes the emotion of "loneliness," and as a result, the server generates a response such as "You are not alone; many people support you." In this way, the present invention provides a new form of religious support by realizing information provision that is attentive to the user's emotions.
[0361] The following describes the processing flow.
[0362] Step 1:
[0363] The user accesses the platform via their device, enters their account information, and logs in. The device retrieves this authentication information and sends it to the server.
[0364] Step 2:
[0365] The server compares the received authentication information with the authentication server or database to determine whether the login was successful. If authentication is successful, it starts a session and prepares it for use by the user.
[0366] Step 3:
[0367] Users send their religious questions or personal concerns to the device using either text or voice input. If voice input is selected, the device converts the voice into text data.
[0368] Step 4:
[0369] The terminal sends the converted text data to the server. The server receives this data, uses natural language processing (NLP) to analyze the input content, and extracts the intent.
[0370] Step 5:
[0371] The server uses an emotion engine to recognize emotional states from text data. Here, emotions are classified as positive, negative, or neutral to understand the user's feelings.
[0372] Step 6:
[0373] Based on the analysis results, the server accesses multiple religious text databases and extracts relevant teachings and information. Simultaneously, it adjusts the tone and context of its responses based on the perceived emotional state.
[0374] Step 7:
[0375] The server packages the generated response in text and audio formats and sends it to the terminal. This allows the user to provide feedback in their preferred format.
[0376] Step 8:
[0377] The terminal presents the received response to the user. If it is in text format, it is displayed on the screen; if it is in audio format, it provides auditory feedback using speech synthesis technology.
[0378] Step 9:
[0379] Users can input their thoughts and feedback on the response. The device sends this feedback to the server, which uses the data to train the emotion engine and further improve its accuracy.
[0380] (Example 2)
[0381] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0382] In society, it is not uncommon for individuals to have religious questions or personal concerns, but there are limited systems that adequately address these. In particular, when emotional considerations are necessary, simple information provision is insufficient. Responses provided without considering emotional states fail to adequately provide the reassurance and support that users seek. Therefore, there is a need for systems that incorporate emotion recognition technology to provide more personalized responses.
[0383] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0384] In this invention, the server includes means for acquiring user input and transmitting it as information, means for analyzing the user input using natural language processing technology based on the document and generating an appropriate response, means for displaying or audibly presenting the generated response to the user, means for analyzing the user's emotional state and adjusting the response content based on the results, and means for automatically generating a response using a generative AI model. This makes it possible to appropriately provide responses that take the user's emotions into consideration and to provide religious support more effectively.
[0385] A "user" is an individual who uses the system to resolve religious questions or personal problems.
[0386] "Input" refers to information that a user transmits to a system using a device, and can be in text or audio format.
[0387] "Information" refers to data obtained by converting user input into a digital format, which is then analyzed and processed within the system.
[0388] A "document" is a collection of text data containing religious or related content used to generate a response.
[0389] "Natural language processing technology" refers to computer program techniques used to analyze user input and understand its meaning and intent.
[0390] An "appropriate response" refers to information or messages that reflect the user's emotional state and intentions based on their input.
[0391] A "generative AI model" is a type of program that uses machine learning techniques to autonomously generate text and voice responses.
[0392] "Emotional state" refers to information that indicates the psychological or emotional condition of a user.
[0393] "Presenting by display or sound" refers to a method of providing information to a user visually or audibly.
[0394] This invention is a system that addresses users' religious questions and personal concerns and provides responses. Users access an online platform using a terminal and input their concerns or questions. Input can be in text or voice format; in the case of voice input, the terminal uses speech recognition technology to convert it into text. For this speech recognition, speech recognition software is used as a standard service.
[0395] The server receives input data sent from the terminal and performs analysis using natural language processing technology. During the analysis process, a generative AI model is used to recognize the user's intent and emotional state. This model may utilize a natural language model such as OpenAI. To understand the emotional state, an emotion analysis algorithm is employed to evaluate the user's emotions. This allows for dynamic adjustment of responses according to the user's psychological state.
[0396] The server references multiple religious text databases based on the analysis results and selects the information most relevant to the user. The selected information is then used to generate a response tailored to the user's emotional state and intentions. This response is automatically generated by a generative AI model, enabling a more human-like interaction.
[0397] The generated response is sent from the server to the terminal and presented to the user. In addition to text display, voice feedback is also provided using speech synthesis technology. A speech synthesis engine is used for speech synthesis, allowing the user to have a more interactive experience.
[0398] Users provide feedback on the responses they receive via their devices, and the server uses this feedback data to perform machine learning to improve the accuracy of sentiment analysis. This improves the overall response accuracy of the system, enabling the provision of more accurate information.
[0399] For example, if a user inputs "I feel lonely," the system recognizes the emotional state and generates a response such as "You are not alone; many people support you." Furthermore, by presenting the AI model with prompts such as "How should I respond if the user inputs 'I feel anxious'?", it is possible to train the model to provide appropriate responses.
[0400] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0401] Step 1:
[0402] The user accesses the online platform using a terminal and logs in. The user enters their ID and password, goes through an authentication process, and is granted access to the platform. In this step, the input is the user's authentication information, and the output is a message indicating authentication success or failure. Specifically, the authentication server verifies the user information.
[0403] Step 2:
[0404] The user inputs a religious question or concern. The input method is either text or voice; in the case of voice, the device uses speech recognition software to convert the voice into text data. In this step, the input is the user's question or concern, either voice or text, and the output is the converted text data. Specifically, the speech recognition engine converts the voice to text in real time.
[0405] Step 3:
[0406] The terminal sends the converted text data to the server. The server passes the received data to a natural language processing engine, which analyzes the user's intent and emotional state. This analysis uses a generative AI model and an emotion recognition algorithm. The input is text data, and the output is data indicating the user's intent and emotional state. Specifically, the natural language processing engine analyzes the text and categorizes its intent.
[0407] Step 4:
[0408] The server selects appropriate information from multiple religious text databases based on the analysis results. A generative AI model is used to generate responses tailored to the user's emotional state. The input is the analyzed intent and emotional state, and the output is a customized response text. Specifically, database queries are executed within the server to retrieve the relevant text.
[0409] Step 5:
[0410] The server sends the generated response to the terminal. The terminal presents it to the user using text display or speech synthesis technology. For speech presentation, a speech synthesis engine is used. The input is the generated response text, and the output is visual or auditory feedback to the user. Specifically, the terminal converts the text into speech so that the user can hear it.
[0411] Step 6:
[0412] The user provides feedback on the presented response. The device sends this feedback back to the server. The server uses the feedback data to perform machine learning to improve the accuracy of the sentiment recognition engine. The input is the user's feedback, and the output is the improved sentiment analysis model. Specifically, the server analyzes the feedback data and updates the algorithm of the generative AI model.
[0413] (Application Example 2)
[0414] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0415] Existing information delivery systems often struggle to adequately consider users' personal needs and emotional states. As a result, they can only provide consistent information in response to user questions, failing to offer appropriate responses tailored to individual emotions and interests, thus limiting the user experience. Especially in online business transactions and consultations, there is a need for systems that understand user emotions and provide personalized suggestions.
[0416] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0417] In this invention, the server includes means for acquiring user input and transmitting it as information data; means for analyzing the user input using natural language processing technology based on multiple religious texts and generating an appropriate response; means for presenting the generated response to the user visually or audibly; means for analyzing the user's emotional state and adjusting the response content accordingly; and means for suggesting information on products and services of interest to the user along with the appropriate response. This makes it possible to provide highly personalized responses and suggestions that are tailored to the user's individual emotions and interests.
[0418] "Means for acquiring user input and transmitting it as information data" refers to a function for collecting information provided by the user and transmitting it as digital data to a processing device or server.
[0419] "A means of analyzing user input using natural language processing technology based on multiple religious texts and generating appropriate responses" refers to a function that uses a wide range of religious literature as a database, interprets input information using a computer, and creates information accordingly.
[0420] "Means for presenting the generated response to the user visually or audibly" refers to functions for providing the created response to the user through sight or hearing, and this includes display devices and audio output devices.
[0421] "Means for analyzing the user's emotional state and adjusting the response content based on that" refers to a function that determines the user's psychological state and optimizes the content and expression of the response accordingly.
[0422] "A means of suggesting information about products and services that users are interested in, along with appropriate responses," refers to a function that selects and recommends products and services that meet the user's needs and interests.
[0423] "Means of converting user voice input into information data" refers to technology that transforms a user's words from speech into digital representations.
[0424] "Means for collecting user feedback and using it to improve response generation and suggestion provision means" refers to technologies for gathering user reactions and opinions and using them to improve the performance and functionality of the system.
[0425] In implementing this invention, smart glasses or head-mounted displays are used as user interface devices. These devices are capable of presenting information in accordance with the user's visual and auditory perception. Specific hardware includes smart glasses devices, voice input devices, and wireless communication devices for data transmission.
[0426] First, the user provides input through an interface device. This is provided as voice input and is captured by the device's built-in microphone. The captured voice data is then converted into text data by the device's speech recognition module. In this process, speech recognition software such as Google Cloud Speech-to-Text is used as an example.
[0427] Next, the generated text data is sent to the server. Here, the server analyzes the input data using a natural language processing engine. This analysis includes, for example, language analysis techniques using spaCy. Furthermore, a custom model using TensorFlow is employed for sentiment analysis to determine the user's emotional state.
[0428] Based on the analysis results, the server generates an appropriate response. The response is tailored based on multiple religious texts and personalized information. The generated response's tone and content are modified according to the user's emotional state. Specific products or services may also be suggested based on the user's interests.
[0429] Ultimately, this response is presented to the user visually or audibly. Presentation methods include text display on a screen or voice output using speech synthesis software (e.g., Amazon Polly).
[0430] As a concrete example of its use, if a user is in a situation where they are "looking for a vegan recipe book and feeling a little tired," this system will take the user's emotional state into consideration and provide optimal book information while suggesting products that have a relaxing effect.
[0431] An example of a prompt message would be, "I'm looking for a vegan recipe book and I'm a little tired. What books would you recommend?"
[0432] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0433] Step 1:
[0434] The user provides voice input via an interface device.
[0435] The voice provided by the user is captured by the device's built-in microphone. This input data is in audio format and is processed as input to the speech recognition module.
[0436] Step 2:
[0437] The device converts the user's voice into text data.
[0438] Speech recognition software is used to analyze audio data, and as a result, text data is generated. This conversion process uses lexical analysis and acoustic models to accurately transcribe the user's speech into text.
[0439] Step 3:
[0440] The terminal sends the converted text data to the server.
[0441] The generated text data is transferred to the server using wireless communication technology. This data contains information including the user's intent and the content of their question.
[0442] Step 4:
[0443] The server analyzes the text data using a natural language processing engine.
[0444] The server uses natural language processing techniques to analyze the language structure of the received text data and interpret the user's intent. Semantic and sentiment analysis are performed to recognize the user's psychological state and requests.
[0445] Step 5:
[0446] The server generates a response based on the user's emotional state.
[0447] Based on the analyzed data, the server generates an appropriate response. This response utilizes an emotion engine to adjust its tone and content, and may include relevant products or services as needed.
[0448] Step 6:
[0449] The server sends the generated response to the terminal.
[0450] The generated response is then transmitted back to the terminal via wireless communication. This output data contains information presented to the user.
[0451] Step 7:
[0452] The device presents a response to the user, either visually or audibly.
[0453] The device presents the received data to the user using text display on the screen or speech synthesis functionality. This allows the user to acquire information through sight or hearing.
[0454] Step 8:
[0455] Users provide feedback on the response.
[0456] Users provide feedback by entering their opinions and impressions about the information presented. This information is used to continuously improve the system.
[0457] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0458] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0459] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0460] [Third Embodiment]
[0461] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0462] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0463] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0464] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0465] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0466] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0467] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0468] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0469] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0470] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0471] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0472] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0473] This invention provides a system that offers a platform for users to consult about religious questions or personal concerns. This system operates online, allowing users to input questions via text or voice, and the system then provides responses based on religious teachings. The embodiments of this system are described in detail below.
[0474] First, the user accesses the platform using a device. The device receives user input and sends the consultation details to the server in text or voice format. If voice input is received, the device converts the voice into text data and sends it to the server.
[0475] The server is equipped with a natural language processing engine for analyzing incoming text data. Using this engine, the server analyzes the input text and determines its intent and sentiment. Based on this analysis, the server consults a database of multiple religious texts and extracts the most relevant information. Based on this information, the server generates the most appropriate response for the user.
[0476] The generated response is sent from the server to the terminal. The terminal then presents the response to the user. This presentation can be displayed visually as text, or audibly using speech synthesis technology.
[0477] As a concrete example, consider a scenario where a user enters a question such as, "How can one maintain mental stability in modern times?" The server analyzes this question, and, if necessary, references relevant religious teachings, such as, "There are ways to maintain mental balance through daily gratitude and meditation," to generate a response. This response is then provided to the user, allowing them to ask further questions or provide feedback.
[0478] In addition, the system includes a function to collect user feedback, which the server analyzes and uses to improve the accuracy of future responses. Through this process, it becomes possible to provide users with more personalized, accurate, and useful information.
[0479] Therefore, this system has the ability to provide appropriate support to users regarding their religious questions and personal concerns, regardless of time or place.
[0480] The following describes the processing flow.
[0481] Step 1:
[0482] The user accesses the platform using their device and creates or logs in an account. The device retrieves the entered authentication information and sends it to the server.
[0483] Step 2:
[0484] The server compares the received authentication information with the database to determine whether authentication was successful. If successful, it generates an authentication token and sends it back to the terminal.
[0485] Step 3:
[0486] The user enters their religious questions or personal concerns into the device as text or voice. The device receives the input and converts it into text data if it is voice.
[0487] Step 4:
[0488] The terminal sends text data to the server. The server inputs the received consultation content into a natural language processing engine and analyzes the intent and emotions.
[0489] Step 5:
[0490] Based on the analysis results, the server consults a religious text database and extracts relevant information. The server then generates a response based on that information.
[0491] Step 6:
[0492] The server sends the generated response to the terminal. The terminal displays the response on its screen and, if audio output is selected, presents it audibly using speech synthesis technology.
[0493] Step 7:
[0494] Users can enter feedback or additional questions regarding the response they receive. The terminal collects this information and sends it to the server.
[0495] Step 8:
[0496] The server stores the collected feedback in a database and uses it as data to improve the accuracy of response generation.
[0497] (Example 1)
[0498] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0499] The challenge lies in providing a platform where users can receive appropriate support for religious questions and personal concerns, regardless of time or location. Furthermore, existing systems have been criticized for lacking accuracy and personalization in their responses based on user input. Additionally, there is a need to accurately process user voice input while ensuring security, and to improve response accuracy through feedback.
[0500] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0501] In this invention, the server includes means for acquiring user input and transmitting it as information data, means for analyzing the user input using natural language processing technology based on multiple religious data and generating an appropriate response, means for presenting the generated response to the user visually or audibly, and means for acquiring user feedback and improving the accuracy of the response generation means. This makes it possible to provide the user with a more appropriate and personalized response. Furthermore, by accurately processing voice input while improving the security of communication, the user experience can be improved.
[0502] A "user" refers to someone who accesses the system to ask religious questions or seek advice on personal problems.
[0503] "Input" refers to information that a user provides to the system through their device, and includes information in text or audio format.
[0504] "Information data" refers to data that has been processed from user input and converted into a digital format that can be transmitted by the system.
[0505] "Religious data" refers to a collection of information, including religious teachings and texts, that the system references when responding to user inquiries.
[0506] "Natural language processing technology" refers to algorithms and methods that systems use to interpret user input and analyze its meaning and intent.
[0507] "Response" refers to the reply or information that the system generates based on user input.
[0508] "Presenting visually or audibly" refers to a method of presentation where the response is displayed to the user as text on the screen or heard as audio.
[0509] "Feedback" refers to information that users convey to the system regarding their evaluations and opinions on the responses provided.
[0510] "Communication technology" refers to the technologies and protocols used to securely transmit user input data from a terminal to a server.
[0511] This invention is a system that uses information technology to provide appropriate responses to users' religious questions and personal concerns. The detailed configuration for implementing this system is described below.
[0512] The user first accesses the online platform using a terminal. This terminal is responsible for receiving text or voice input from the user. If voice input is provided, the terminal uses speech recognition software (e.g., a speech recognition API) to convert the speech into text. The converted text data is then sent to the server as information data. Secure communication technology (e.g., HTTPS protocol) is used for this transmission.
[0513] Next, the server analyzes the received data using natural language processing techniques (e.g., machine learning models). The server analyzes the grammatical and semantic structure of the text entered by the user to infer the user's intent and emotions. During this analysis, it understands the user's questions and searches for relevant "religious data." This data consists of doctrines and texts stored in a database.
[0514] The server then runs a natural language generation model (e.g., a generative AI model) to generate a response. This model receives the analyzed input and generates an appropriate and specific response to the user's question. The server is designed to take user feedback into consideration during this process to improve the accuracy of the response.
[0515] The generated response is sent from the server to the terminal, which can then present it to the user either visually as text or audibly using speech synthesis technology (e.g., a speech conversion API). This method of presenting the response is designed to enhance system flexibility and ensure easy user comprehension.
[0516] As a concrete example, consider a scenario where a user inputs the question, "How can one maintain mental stability in modern times?" The server analyzes this question and generates an appropriate response based on religious teachings that suggest daily gratitude and meditation are helpful in maintaining mental balance. This response has the potential to offer the user further reflection and new perspectives.
[0517] Another example of a prompt for a generative AI model is: "The user has asked about how to maintain mental balance. Refer to relevant religious teachings and generate an appropriate response."
[0518] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0519] Step 1:
[0520] Users use their devices to access online platforms and input their religious questions or personal concerns in text or voice format. The input data is received by the device and processed according to its format. In the case of voice input, the device uses speech recognition to convert the voice into text data and prepares it as information data.
[0521] Step 2:
[0522] The device transmits text data converted by speech recognition or text data entered by the user to the server using secure communication technology. Privacy is maintained through data encryption during transmission. The output is the information data sent to the server.
[0523] Step 3:
[0524] The server analyzes the received data using natural language processing techniques. The input is user text data; by analyzing this data, the server understands its grammatical and semantic structure and performs data processing to infer the user's intentions and emotions. The output is the analysis result.
[0525] Step 4:
[0526] The server searches for relevant information from religious databases based on the analysis results. This search process applies an efficient search algorithm to extract the most relevant data. The input is the analysis results, and the output is the relevant information.
[0527] Step 5:
[0528] The server generates an appropriate response using a natural language generation model based on relevant information. In this process, a generative AI model operates to create appropriate and specific answers to the user's questions. The input is relevant information, and the output is the generated response.
[0529] Step 6:
[0530] The server sends a generated response to the terminal. The terminal either displays this response as text on the screen or conveys it to the user audibly using speech synthesis. This process allows the user to receive the response visually or audibly. The input is the generated response, and the output is the response displayed or presented audibly to the user.
[0531] Step 7:
[0532] The user sends feedback on the response provided through the terminal. This feedback is collected as data on the terminal and sent back to the server. This feedback is then analyzed by the server to help improve the accuracy of the response. The input is the user's feedback, and the output is the data used by the server for accuracy improvement.
[0533] (Application Example 1)
[0534] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0535] In modern society, users face a great deal of stress and anxiety in their daily lives. In such circumstances, there is a need for a system that can understand the user's mental state in real time and provide appropriate guidance. However, current technology lacks the means to accurately analyze a user's emotional state and provide religious or philosophical guidance tailored to their individual circumstances.
[0536] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0537] In this invention, the server includes means for acquiring user input and transmitting it as text data; means for analyzing the user input using natural language processing technology based on multiple information resources and generating an appropriate response; means for presenting the generated response to the user visually or audibly; means for measuring the user's biometric information and analyzing their emotional state; and means for generating and presenting appropriate guidance based on the emotional state. This enables real-time support for the user's mental health and the provision of personalized information.
[0538] "User input" refers to information provided by the user in voice or text format.
[0539] "Text data" refers to strings of information expressed in a format that can be processed by a computer.
[0540] An "information resource" is a collection of knowledge and data, including information based on specific teachings or instruction.
[0541] "Natural language processing technology" is a technology that enables computers to understand and analyze human language.
[0542] An "appropriate response" is information, including solutions and guidance, that is generated based on user input.
[0543] "Presenting visually or audibly" refers to a method of conveying information through the user's sight or hearing.
[0544] "Biometric information" refers to data that indicates the user's physical condition, including pulse rate and facial expressions.
[0545] "Analyzing emotional state" is the process of estimating a user's mental state and emotions from their biometric information.
[0546] "Generating and presenting guidance" means creating and providing advice and information tailored to the user's situation.
[0547] This invention provides a system that allows users to consult about religious questions or personal concerns using a terminal. The terminal receives input from the user as voice or text data and sends it to a server. In the case of voice input, the terminal first converts the voice into text data. Software used includes the Google Cloud Speech-to-Text API.
[0548] The server analyzes the received text data using natural language processing techniques to determine the intent and sentiment of the input text. This analysis utilizes open-source natural language processing libraries such as NLTK and Transformers. After analysis, the server consults multiple information resources to extract the information most relevant to the user's input. Based on this information, a generative AI model generates appropriate guidance.
[0549] Next, the server sends the generated instructions to the terminal, which then presents them to the user visually or audibly. Presentation methods include text display on a screen and auditory presentation using a speech synthesis system. Amazon Polly, for example, can be used for speech synthesis.
[0550] Furthermore, this system has the function of measuring the user's biometric information and analyzing their emotional state. Sensors installed in the terminal detect pulse rate and facial expressions, and a server processes this data to determine the user's emotional state. Based on the determination, appropriate guidance is generated.
[0551] For example, if a user asks, "How can I maintain peace of mind?", the system will generate and provide guidance such as, "There are ways to maintain mental balance through meditation and daily gratitude."
[0552] Example prompt for the AI model: "What kind of guidance would be appropriate if the user is feeling anxious?"
[0553] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0554] Step 1:
[0555] The device accepts user input. The user enters religious questions or concerns in voice or text format, and the device receives this input data. In the case of voice input, the device uses the Google Cloud Speech-to-Text API to convert the voice into text data. The input is then sent to the server as text data.
[0556] Step 2:
[0557] The server analyzes the received text data. Using the received text data as input, the server performs natural language processing using NLTK and Transformers to determine the user's intent and emotions. As a result of the analysis, tags representing the user's question intent and the sentiment analysis results are output.
[0558] Step 3:
[0559] The server references information resources and extracts appropriate information. Based on the analysis results, the server searches multiple information resources (including religious teachings and philosophical guidance) and selects the most relevant information. In this step, a search is performed using the prompt "What guidance is appropriate when a user is feeling anxious?". As a result, relevant guidance information is output.
[0560] Step 4:
[0561] The server generates instruction using a generated AI model. Using the extracted information as input, the generated AI model generates instruction appropriate for the user. Specific instruction content is provided as output.
[0562] Step 5:
[0563] The server sends the generated instruction to the terminal. Using the generated instruction as input, the server sends data to the terminal. The instruction data arrives at the terminal as output.
[0564] Step 6:
[0565] The device presents instructions to the user. The device presents the received instruction data to the user audibly using text display on the screen or speech synthesis. Amazon Polly can be used for this presentation.
[0566] Step 7:
[0567] The device measures the user's biometric information and analyzes their emotional state. The device's sensors capture the user's pulse and facial expressions, and transmit this data to a server. The server analyzes this biometric information to determine the user's emotional state. The output is an evaluation of the emotional state.
[0568] Step 8:
[0569] The server generates further guidance based on the emotional state and sends it to the device. Based on the evaluation of the emotional state, the server generates additional guidance as needed. Based on this, new guidance data is provided to the device.
[0570] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0571] This invention provides a system that takes into account the user's emotional state when addressing their religious questions or personal concerns. This system operates on an online platform and enables more personalized information delivery through user interaction.
[0572] When a user accesses the platform using their device, they first undergo login authentication. Users can then input religious questions or personal concerns via text and voice. In the case of voice input, the device converts the voice into text data and sends it to the server.
[0573] The server analyzes each input using a natural language processing engine to recognize the user's intent and emotional state. This analysis utilizes an emotion engine, which extracts emotional information from the input text to understand the user's feelings. For example, if words or phrases indicating stress or anxiety are included, the server recognizes them and adjusts the response accordingly.
[0574] Based on the analysis results, the server consults multiple religious text databases and selects the information most appropriate for the user. The tone and content of the response are flexibly adjusted according to the user's emotional state. For example, if the server determines the user is depressed, comforting words and encouraging messages will be selected.
[0575] The generated response is sent from the server to the terminal and presented to the user. Presentation methods include not only text display but also auditory feedback using speech synthesis technology. The user can provide feedback on the response, which the terminal reports to the server.
[0576] The emotion engine uses machine learning based on this feedback data to improve the accuracy of emotion recognition. For example, when a user inputs "I feel lonely," the emotion engine accurately recognizes the emotion of "loneliness," and as a result, the server generates a response such as "You are not alone; many people support you." In this way, the present invention provides a new form of religious support by realizing information provision that is attentive to the user's emotions.
[0577] The following describes the processing flow.
[0578] Step 1:
[0579] The user accesses the platform via their device, enters their account information, and logs in. The device retrieves this authentication information and sends it to the server.
[0580] Step 2:
[0581] The server compares the received authentication information with the authentication server or database to determine whether the login was successful. If authentication is successful, it starts a session and prepares it for use by the user.
[0582] Step 3:
[0583] Users send their religious questions or personal concerns to the device using either text or voice input. If voice input is selected, the device converts the voice into text data.
[0584] Step 4:
[0585] The terminal sends the converted text data to the server. The server receives this data, uses natural language processing (NLP) to analyze the input content, and extracts the intent.
[0586] Step 5:
[0587] The server uses an emotion engine to recognize emotional states from text data. Here, emotions are classified as positive, negative, or neutral to understand the user's feelings.
[0588] Step 6:
[0589] Based on the analysis results, the server accesses multiple religious text databases and extracts relevant teachings and information. Simultaneously, it adjusts the tone and context of its responses based on the perceived emotional state.
[0590] Step 7:
[0591] The server packages the generated response in text and audio formats and sends it to the terminal. This allows the user to provide feedback in their preferred format.
[0592] Step 8:
[0593] The terminal presents the received response to the user. If it is in text format, it is displayed on the screen; if it is in audio format, it provides auditory feedback using speech synthesis technology.
[0594] Step 9:
[0595] Users can input their thoughts and feedback on the response. The device sends this feedback to the server, which uses the data to train the emotion engine and further improve its accuracy.
[0596] (Example 2)
[0597] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0598] In society, it is not uncommon for individuals to have religious questions or personal concerns, but there are limited systems that adequately address these. In particular, when emotional considerations are necessary, simple information provision is insufficient. Responses provided without considering emotional states fail to adequately provide the reassurance and support that users seek. Therefore, there is a need for systems that incorporate emotion recognition technology to provide more personalized responses.
[0599] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0600] In this invention, the server includes means for acquiring user input and transmitting it as information, means for analyzing the user input using natural language processing technology based on the document and generating an appropriate response, means for displaying or audibly presenting the generated response to the user, means for analyzing the user's emotional state and adjusting the response content based on the results, and means for automatically generating a response using a generative AI model. This makes it possible to appropriately provide responses that take the user's emotions into consideration and to provide religious support more effectively.
[0601] A "user" is an individual who uses the system to resolve religious questions or personal problems.
[0602] "Input" refers to information that a user transmits to a system using a device, and can be in text or audio format.
[0603] "Information" refers to data obtained by converting user input into a digital format, which is then analyzed and processed within the system.
[0604] A "document" is a collection of text data containing religious or related content used to generate a response.
[0605] "Natural language processing technology" refers to computer program techniques used to analyze user input and understand its meaning and intent.
[0606] An "appropriate response" refers to information or messages that reflect the user's emotional state and intentions based on their input.
[0607] A "generative AI model" is a type of program that uses machine learning techniques to autonomously generate text and voice responses.
[0608] "Emotional state" refers to information that indicates the psychological or emotional condition of a user.
[0609] "Presenting by display or sound" refers to a method of providing information to a user visually or audibly.
[0610] This invention is a system that addresses users' religious questions and personal concerns and provides responses. Users access an online platform using a terminal and input their concerns or questions. Input can be in text or voice format; in the case of voice input, the terminal uses speech recognition technology to convert it into text. For this speech recognition, speech recognition software is used as a standard service.
[0611] The server receives input data sent from the terminal and performs analysis using natural language processing technology. During the analysis process, a generative AI model is used to recognize the user's intent and emotional state. This model may utilize a natural language model such as OpenAI. To understand the emotional state, an emotion analysis algorithm is employed to evaluate the user's emotions. This allows for dynamic adjustment of responses according to the user's psychological state.
[0612] The server references multiple religious text databases based on the analysis results and selects the information most relevant to the user. The selected information is then used to generate a response tailored to the user's emotional state and intentions. This response is automatically generated by a generative AI model, enabling a more human-like interaction.
[0613] The generated response is sent from the server to the terminal and presented to the user. In addition to text display, voice feedback is also provided using speech synthesis technology. A speech synthesis engine is used for speech synthesis, allowing the user to have a more interactive experience.
[0614] Users provide feedback on the responses they receive via their devices, and the server uses this feedback data to perform machine learning to improve the accuracy of sentiment analysis. This improves the overall response accuracy of the system, enabling the provision of more accurate information.
[0615] For example, if a user inputs "I feel lonely," the system recognizes the emotional state and generates a response such as "You are not alone; many people support you." Furthermore, by presenting the AI model with prompts such as "How should I respond if the user inputs 'I feel anxious'?", it is possible to train the model to provide appropriate responses.
[0616] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0617] Step 1:
[0618] The user accesses the online platform using a terminal and logs in. The user enters their ID and password, goes through an authentication process, and is granted access to the platform. In this step, the input is the user's authentication information, and the output is a message indicating authentication success or failure. Specifically, the authentication server verifies the user information.
[0619] Step 2:
[0620] The user inputs a religious question or concern. The input method is either text or voice; in the case of voice, the device uses speech recognition software to convert the voice into text data. In this step, the input is the user's question or concern, either voice or text, and the output is the converted text data. Specifically, the speech recognition engine converts the voice to text in real time.
[0621] Step 3:
[0622] The terminal sends the converted text data to the server. The server passes the received data to a natural language processing engine, which analyzes the user's intent and emotional state. This analysis uses a generative AI model and an emotion recognition algorithm. The input is text data, and the output is data indicating the user's intent and emotional state. Specifically, the natural language processing engine analyzes the text and categorizes its intent.
[0623] Step 4:
[0624] The server selects appropriate information from multiple religious text databases based on the analysis results. A generative AI model is used to generate responses tailored to the user's emotional state. The input is the analyzed intent and emotional state, and the output is a customized response text. Specifically, database queries are executed within the server to retrieve the relevant text.
[0625] Step 5:
[0626] The server sends the generated response to the terminal. The terminal presents it to the user using text display or speech synthesis technology. For speech presentation, a speech synthesis engine is used. The input is the generated response text, and the output is visual or auditory feedback to the user. Specifically, the terminal converts the text into speech so that the user can hear it.
[0627] Step 6:
[0628] The user provides feedback on the presented response. The device sends this feedback back to the server. The server uses the feedback data to perform machine learning to improve the accuracy of the sentiment recognition engine. The input is the user's feedback, and the output is the improved sentiment analysis model. Specifically, the server analyzes the feedback data and updates the algorithm of the generative AI model.
[0629] (Application Example 2)
[0630] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0631] Existing information delivery systems often struggle to adequately consider users' personal needs and emotional states. As a result, they can only provide consistent information in response to user questions, failing to offer appropriate responses tailored to individual emotions and interests, thus limiting the user experience. Especially in online business transactions and consultations, there is a need for systems that understand user emotions and provide personalized suggestions.
[0632] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0633] In this invention, the server includes means for acquiring user input and transmitting it as information data; means for analyzing the user input using natural language processing technology based on multiple religious texts and generating an appropriate response; means for presenting the generated response to the user visually or audibly; means for analyzing the user's emotional state and adjusting the response content accordingly; and means for suggesting information on products and services of interest to the user along with the appropriate response. This makes it possible to provide highly personalized responses and suggestions that are tailored to the user's individual emotions and interests.
[0634] "Means for acquiring user input and transmitting it as information data" refers to a function for collecting information provided by the user and transmitting it as digital data to a processing device or server.
[0635] "A means of analyzing user input using natural language processing technology based on multiple religious texts and generating appropriate responses" refers to a function that uses a wide range of religious literature as a database, interprets input information using a computer, and creates information accordingly.
[0636] "Means for presenting the generated response to the user visually or audibly" refers to functions for providing the created response to the user through sight or hearing, and this includes display devices and audio output devices.
[0637] "Means for analyzing the user's emotional state and adjusting the response content based on that" refers to a function that determines the user's psychological state and optimizes the content and expression of the response accordingly.
[0638] "A means of suggesting information about products and services that users are interested in, along with appropriate responses," refers to a function that selects and recommends products and services that meet the user's needs and interests.
[0639] "Means of converting user voice input into information data" refers to technology that transforms a user's words from speech into digital representations.
[0640] "Means for collecting user feedback and using it to improve response generation and suggestion provision means" refers to technologies for gathering user reactions and opinions and using them to improve the performance and functionality of the system.
[0641] In implementing this invention, smart glasses or head-mounted displays are used as user interface devices. These devices are capable of presenting information in accordance with the user's visual and auditory perception. Specific hardware includes smart glasses devices, voice input devices, and wireless communication devices for data transmission.
[0642] First, the user provides input through an interface device. This is provided as voice input and is captured by the device's built-in microphone. The captured voice data is then converted into text data by the device's speech recognition module. In this process, speech recognition software such as Google Cloud Speech-to-Text is used as an example.
[0643] Next, the generated text data is sent to the server. Here, the server analyzes the input data using a natural language processing engine. This analysis includes, for example, language analysis techniques using spaCy. Furthermore, a custom model using TensorFlow is employed for sentiment analysis to determine the user's emotional state.
[0644] Based on the analysis results, the server generates an appropriate response. The response is tailored based on multiple religious texts and personalized information. The generated response's tone and content are modified according to the user's emotional state. Specific products or services may also be suggested based on the user's interests.
[0645] Ultimately, this response is presented to the user visually or audibly. Presentation methods include text display on a screen or voice output using speech synthesis software (e.g., Amazon Polly).
[0646] As a concrete example of its use, if a user is in a situation where they are "looking for a vegan recipe book and feeling a little tired," this system will take the user's emotional state into consideration and provide optimal book information while suggesting products that have a relaxing effect.
[0647] An example of a prompt message would be, "I'm looking for a vegan recipe book and I'm a little tired. What books would you recommend?"
[0648] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0649] Step 1:
[0650] The user provides voice input via an interface device.
[0651] The voice provided by the user is captured by the device's built-in microphone. This input data is in audio format and is processed as input to the speech recognition module.
[0652] Step 2:
[0653] The device converts the user's voice into text data.
[0654] Speech recognition software is used to analyze audio data, and as a result, text data is generated. This conversion process uses lexical analysis and acoustic models to accurately transcribe the user's speech into text.
[0655] Step 3:
[0656] The terminal sends the converted text data to the server.
[0657] The generated text data is transferred to the server using wireless communication technology. This data contains information including the user's intent and the content of their question.
[0658] Step 4:
[0659] The server analyzes the text data using a natural language processing engine.
[0660] The server uses natural language processing techniques to analyze the language structure of the received text data and interpret the user's intent. Semantic and sentiment analysis are performed to recognize the user's psychological state and requests.
[0661] Step 5:
[0662] The server generates a response based on the user's emotional state.
[0663] Based on the analyzed data, the server generates an appropriate response. This response utilizes an emotion engine to adjust its tone and content, and may include relevant products or services as needed.
[0664] Step 6:
[0665] The server sends the generated response to the terminal.
[0666] The generated response is then transmitted back to the terminal via wireless communication. This output data contains information presented to the user.
[0667] Step 7:
[0668] The device presents a response to the user, either visually or audibly.
[0669] The device presents the received data to the user using text display on the screen or speech synthesis functionality. This allows the user to acquire information through sight or hearing.
[0670] Step 8:
[0671] Users provide feedback on the response.
[0672] Users provide feedback by entering their opinions and impressions about the information presented. This information is used to continuously improve the system.
[0673] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0674] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0675] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0676] [Fourth Embodiment]
[0677] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0678] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0679] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0680] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0681] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0682] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0683] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0684] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0685] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0686] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0687] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0688] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0689] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0690] This invention provides a system that offers a platform for users to consult about religious questions or personal concerns. This system operates online, allowing users to input questions via text or voice, and the system then provides responses based on religious teachings. The embodiments of this system are described in detail below.
[0691] First, the user accesses the platform using a device. The device receives user input and sends the consultation details to the server in text or voice format. If voice input is received, the device converts the voice into text data and sends it to the server.
[0692] The server is equipped with a natural language processing engine for analyzing incoming text data. Using this engine, the server analyzes the input text and determines its intent and sentiment. Based on this analysis, the server consults a database of multiple religious texts and extracts the most relevant information. Based on this information, the server generates the most appropriate response for the user.
[0693] The generated response is sent from the server to the terminal. The terminal then presents the response to the user. This presentation can be displayed visually as text, or audibly using speech synthesis technology.
[0694] As a concrete example, consider a scenario where a user enters a question such as, "How can one maintain mental stability in modern times?" The server analyzes this question, and, if necessary, references relevant religious teachings, such as, "There are ways to maintain mental balance through daily gratitude and meditation," to generate a response. This response is then provided to the user, allowing them to ask further questions or provide feedback.
[0695] In addition, the system includes a function to collect user feedback, which the server analyzes and uses to improve the accuracy of future responses. Through this process, it becomes possible to provide users with more personalized, accurate, and useful information.
[0696] Therefore, this system has the ability to provide appropriate support to users regarding their religious questions and personal concerns, regardless of time or place.
[0697] The following describes the processing flow.
[0698] Step 1:
[0699] The user accesses the platform using their device and creates or logs in an account. The device retrieves the entered authentication information and sends it to the server.
[0700] Step 2:
[0701] The server compares the received authentication information with the database to determine whether authentication was successful. If successful, it generates an authentication token and sends it back to the terminal.
[0702] Step 3:
[0703] The user enters their religious questions or personal concerns into the device as text or voice. The device receives the input and converts it into text data if it is voice.
[0704] Step 4:
[0705] The terminal sends text data to the server. The server inputs the received consultation content into a natural language processing engine and analyzes the intent and emotions.
[0706] Step 5:
[0707] Based on the analysis results, the server consults a religious text database and extracts relevant information. The server then generates a response based on that information.
[0708] Step 6:
[0709] The server sends the generated response to the terminal. The terminal displays the response on its screen and, if audio output is selected, presents it audibly using speech synthesis technology.
[0710] Step 7:
[0711] Users can enter feedback or additional questions regarding the response they receive. The terminal collects this information and sends it to the server.
[0712] Step 8:
[0713] The server stores the collected feedback in a database and uses it as data to improve the accuracy of response generation.
[0714] (Example 1)
[0715] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0716] The challenge lies in providing a platform where users can receive appropriate support for religious questions and personal concerns, regardless of time or location. Furthermore, existing systems have been criticized for lacking accuracy and personalization in their responses based on user input. Additionally, there is a need to accurately process user voice input while ensuring security, and to improve response accuracy through feedback.
[0717] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0718] In this invention, the server includes means for acquiring user input and transmitting it as information data, means for analyzing the user input using natural language processing technology based on multiple religious data and generating an appropriate response, means for presenting the generated response to the user visually or audibly, and means for acquiring user feedback and improving the accuracy of the response generation means. This makes it possible to provide the user with a more appropriate and personalized response. Furthermore, by accurately processing voice input while improving the security of communication, the user experience can be improved.
[0719] A "user" refers to someone who accesses the system to ask religious questions or seek advice on personal problems.
[0720] "Input" refers to information that a user provides to the system through their device, and includes information in text or audio format.
[0721] "Information data" refers to data that has been processed from user input and converted into a digital format that can be transmitted by the system.
[0722] "Religious data" refers to a collection of information, including religious teachings and texts, that the system references when responding to user inquiries.
[0723] "Natural language processing technology" refers to algorithms and methods that systems use to interpret user input and analyze its meaning and intent.
[0724] "Response" refers to the reply or information that the system generates based on user input.
[0725] "Presenting visually or audibly" refers to a method of presentation where the response is displayed to the user as text on the screen or heard as audio.
[0726] "Feedback" refers to information that users convey to the system regarding their evaluations and opinions on the responses provided.
[0727] "Communication technology" refers to the technologies and protocols used to securely transmit user input data from a terminal to a server.
[0728] This invention is a system that uses information technology to provide appropriate responses to users' religious questions and personal concerns. The detailed configuration for implementing this system is described below.
[0729] The user first accesses the online platform using a terminal. This terminal is responsible for receiving text or voice input from the user. If voice input is provided, the terminal uses speech recognition software (e.g., a speech recognition API) to convert the speech into text. The converted text data is then sent to the server as information data. Secure communication technology (e.g., HTTPS protocol) is used for this transmission.
[0730] Next, the server analyzes the received data using natural language processing techniques (e.g., machine learning models). The server analyzes the grammatical and semantic structure of the text entered by the user to infer the user's intent and emotions. During this analysis, it understands the user's questions and searches for relevant "religious data." This data consists of doctrines and texts stored in a database.
[0731] The server then runs a natural language generation model (e.g., a generative AI model) to generate a response. This model receives the analyzed input and generates an appropriate and specific response to the user's question. The server is designed to take user feedback into consideration during this process to improve the accuracy of the response.
[0732] The generated response is sent from the server to the terminal, which can then present it to the user either visually as text or audibly using speech synthesis technology (e.g., a speech conversion API). This method of presenting the response is designed to enhance system flexibility and ensure easy user comprehension.
[0733] As a concrete example, consider a scenario where a user inputs the question, "How can one maintain mental stability in modern times?" The server analyzes this question and generates an appropriate response based on religious teachings that suggest daily gratitude and meditation are helpful in maintaining mental balance. This response has the potential to offer the user further reflection and new perspectives.
[0734] Another example of a prompt for a generative AI model is: "The user has asked about how to maintain mental balance. Refer to relevant religious teachings and generate an appropriate response."
[0735] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0736] Step 1:
[0737] Users use their devices to access online platforms and input their religious questions or personal concerns in text or voice format. The input data is received by the device and processed according to its format. In the case of voice input, the device uses speech recognition to convert the voice into text data and prepares it as information data.
[0738] Step 2:
[0739] The device transmits text data converted by speech recognition or text data entered by the user to the server using secure communication technology. Privacy is maintained through data encryption during transmission. The output is the information data sent to the server.
[0740] Step 3:
[0741] The server analyzes the received data using natural language processing techniques. The input is user text data; by analyzing this data, the server understands its grammatical and semantic structure and performs data processing to infer the user's intentions and emotions. The output is the analysis result.
[0742] Step 4:
[0743] The server searches for relevant information from religious databases based on the analysis results. This search process applies an efficient search algorithm to extract the most relevant data. The input is the analysis results, and the output is the relevant information.
[0744] Step 5:
[0745] The server generates an appropriate response using a natural language generation model based on relevant information. In this process, a generative AI model operates to create appropriate and specific answers to the user's questions. The input is relevant information, and the output is the generated response.
[0746] Step 6:
[0747] The server sends a generated response to the terminal. The terminal either displays this response as text on the screen or conveys it to the user audibly using speech synthesis. This process allows the user to receive the response visually or audibly. The input is the generated response, and the output is the response displayed or presented audibly to the user.
[0748] Step 7:
[0749] The user sends feedback on the response provided through the terminal. This feedback is collected as data on the terminal and sent back to the server. This feedback is then analyzed by the server to help improve the accuracy of the response. The input is the user's feedback, and the output is the data used by the server for accuracy improvement.
[0750] (Application Example 1)
[0751] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0752] In modern society, users face a great deal of stress and anxiety in their daily lives. In such circumstances, there is a need for a system that can understand the user's mental state in real time and provide appropriate guidance. However, current technology lacks the means to accurately analyze a user's emotional state and provide religious or philosophical guidance tailored to their individual circumstances.
[0753] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0754] In this invention, the server includes means for acquiring user input and transmitting it as text data; means for analyzing the user input using natural language processing technology based on multiple information resources and generating an appropriate response; means for presenting the generated response to the user visually or audibly; means for measuring the user's biometric information and analyzing their emotional state; and means for generating and presenting appropriate guidance based on the emotional state. This enables real-time support for the user's mental health and the provision of personalized information.
[0755] "User input" refers to information provided by the user in voice or text format.
[0756] "Text data" refers to strings of information expressed in a format that can be processed by a computer.
[0757] An "information resource" is a collection of knowledge and data, including information based on specific teachings or instruction.
[0758] "Natural language processing technology" is a technology that enables computers to understand and analyze human language.
[0759] An "appropriate response" is information, including solutions and guidance, that is generated based on user input.
[0760] "Presenting visually or audibly" refers to a method of conveying information through the user's sight or hearing.
[0761] "Biometric information" refers to data that indicates the user's physical condition, including pulse rate and facial expressions.
[0762] "Analyzing emotional state" is the process of estimating a user's mental state and emotions from their biometric information.
[0763] "Generating and presenting guidance" means creating and providing advice and information tailored to the user's situation.
[0764] This invention provides a system that allows users to consult about religious questions or personal concerns using a terminal. The terminal receives input from the user as voice or text data and sends it to a server. In the case of voice input, the terminal first converts the voice into text data. Software used includes the Google Cloud Speech-to-Text API.
[0765] The server analyzes the received text data using natural language processing techniques to determine the intent and sentiment of the input text. This analysis utilizes open-source natural language processing libraries such as NLTK and Transformers. After analysis, the server consults multiple information resources to extract the information most relevant to the user's input. Based on this information, a generative AI model generates appropriate guidance.
[0766] Next, the server sends the generated instructions to the terminal, which then presents them to the user visually or audibly. Presentation methods include text display on a screen and auditory presentation using a speech synthesis system. Amazon Polly, for example, can be used for speech synthesis.
[0767] Furthermore, this system has the function of measuring the user's biometric information and analyzing their emotional state. Sensors installed in the terminal detect pulse rate and facial expressions, and a server processes this data to determine the user's emotional state. Based on the determination, appropriate guidance is generated.
[0768] For example, if a user asks, "How can I maintain peace of mind?", the system will generate and provide guidance such as, "There are ways to maintain mental balance through meditation and daily gratitude."
[0769] Example prompt for the AI model: "What kind of guidance would be appropriate if the user is feeling anxious?"
[0770] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0771] Step 1:
[0772] The device accepts user input. The user enters religious questions or concerns in voice or text format, and the device receives this input data. In the case of voice input, the device uses the Google Cloud Speech-to-Text API to convert the voice into text data. The input is then sent to the server as text data.
[0773] Step 2:
[0774] The server analyzes the received text data. Using the received text data as input, the server performs natural language processing using NLTK and Transformers to determine the user's intent and emotions. As a result of the analysis, tags representing the user's question intent and the sentiment analysis results are output.
[0775] Step 3:
[0776] The server references information resources and extracts appropriate information. Based on the analysis results, the server searches multiple information resources (including religious teachings and philosophical guidance) and selects the most relevant information. In this step, a search is performed using the prompt "What guidance is appropriate when a user is feeling anxious?". As a result, relevant guidance information is output.
[0777] Step 4:
[0778] The server generates instruction using a generated AI model. Using the extracted information as input, the generated AI model generates instruction appropriate for the user. Specific instruction content is provided as output.
[0779] Step 5:
[0780] The server sends the generated instruction to the terminal. Using the generated instruction as input, the server sends data to the terminal. The instruction data arrives at the terminal as output.
[0781] Step 6:
[0782] The device presents instructions to the user. The device presents the received instruction data to the user audibly using text display on the screen or speech synthesis. Amazon Polly can be used for this presentation.
[0783] Step 7:
[0784] The device measures the user's biometric information and analyzes their emotional state. The device's sensors capture the user's pulse and facial expressions, and transmit this data to a server. The server analyzes this biometric information to determine the user's emotional state. The output is an evaluation of the emotional state.
[0785] Step 8:
[0786] The server generates further guidance based on the emotional state and sends it to the device. Based on the evaluation of the emotional state, the server generates additional guidance as needed. Based on this, new guidance data is provided to the device.
[0787] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0788] This invention provides a system that takes into account the user's emotional state when addressing their religious questions or personal concerns. This system operates on an online platform and enables more personalized information delivery through user interaction.
[0789] When a user accesses the platform using their device, they first undergo login authentication. Users can then input religious questions or personal concerns via text and voice. In the case of voice input, the device converts the voice into text data and sends it to the server.
[0790] The server analyzes each input using a natural language processing engine to recognize the user's intent and emotional state. This analysis utilizes an emotion engine, which extracts emotional information from the input text to understand the user's feelings. For example, if words or phrases indicating stress or anxiety are included, the server recognizes them and adjusts the response accordingly.
[0791] Based on the analysis results, the server consults multiple religious text databases and selects the information most appropriate for the user. The tone and content of the response are flexibly adjusted according to the user's emotional state. For example, if the server determines the user is depressed, comforting words and encouraging messages will be selected.
[0792] The generated response is sent from the server to the terminal and presented to the user. Presentation methods include not only text display but also auditory feedback using speech synthesis technology. The user can provide feedback on the response, which the terminal reports to the server.
[0793] The emotion engine uses machine learning based on this feedback data to improve the accuracy of emotion recognition. For example, when a user inputs "I feel lonely," the emotion engine accurately recognizes the emotion of "loneliness," and as a result, the server generates a response such as "You are not alone; many people support you." In this way, the present invention provides a new form of religious support by realizing information provision that is attentive to the user's emotions.
[0794] The following describes the processing flow.
[0795] Step 1:
[0796] The user accesses the platform via their device, enters their account information, and logs in. The device retrieves this authentication information and sends it to the server.
[0797] Step 2:
[0798] The server compares the received authentication information with the authentication server or database to determine whether the login was successful. If authentication is successful, it starts a session and prepares it for use by the user.
[0799] Step 3:
[0800] Users send their religious questions or personal concerns to the device using either text or voice input. If voice input is selected, the device converts the voice into text data.
[0801] Step 4:
[0802] The terminal sends the converted text data to the server. The server receives this data, uses natural language processing (NLP) to analyze the input content, and extracts the intent.
[0803] Step 5:
[0804] The server uses an emotion engine to recognize emotional states from text data. Here, emotions are classified as positive, negative, or neutral to understand the user's feelings.
[0805] Step 6:
[0806] Based on the analysis results, the server accesses multiple religious text databases and extracts relevant teachings and information. Simultaneously, it adjusts the tone and context of its responses based on the perceived emotional state.
[0807] Step 7:
[0808] The server packages the generated response in text and audio formats and sends it to the terminal. This allows the user to provide feedback in their preferred format.
[0809] Step 8:
[0810] The terminal presents the received response to the user. If it is in text format, it is displayed on the screen; if it is in audio format, it provides auditory feedback using speech synthesis technology.
[0811] Step 9:
[0812] Users can input their thoughts and feedback on the response. The device sends this feedback to the server, which uses the data to train the emotion engine and further improve its accuracy.
[0813] (Example 2)
[0814] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0815] In society, it is not uncommon for individuals to have religious questions or personal concerns, but there are limited systems that adequately address these. In particular, when emotional considerations are necessary, simple information provision is insufficient. Responses provided without considering emotional states fail to adequately provide the reassurance and support that users seek. Therefore, there is a need for systems that incorporate emotion recognition technology to provide more personalized responses.
[0816] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0817] In this invention, the server includes means for acquiring user input and transmitting it as information, means for analyzing the user input using natural language processing technology based on the document and generating an appropriate response, means for displaying or audibly presenting the generated response to the user, means for analyzing the user's emotional state and adjusting the response content based on the results, and means for automatically generating a response using a generative AI model. This makes it possible to appropriately provide responses that take the user's emotions into consideration and to provide religious support more effectively.
[0818] A "user" is an individual who uses the system to resolve religious questions or personal problems.
[0819] "Input" refers to information that a user transmits to a system using a device, and can be in text or audio format.
[0820] "Information" refers to data obtained by converting user input into a digital format, which is then analyzed and processed within the system.
[0821] A "document" is a collection of text data containing religious or related content used to generate a response.
[0822] "Natural language processing technology" refers to computer program techniques used to analyze user input and understand its meaning and intent.
[0823] An "appropriate response" refers to information or messages that reflect the user's emotional state and intentions based on their input.
[0824] A "generative AI model" is a type of program that uses machine learning techniques to autonomously generate text and voice responses.
[0825] "Emotional state" refers to information that indicates the psychological or emotional condition of a user.
[0826] "Presenting by display or sound" refers to a method of providing information to a user visually or audibly.
[0827] This invention is a system that addresses users' religious questions and personal concerns and provides responses. Users access an online platform using a terminal and input their concerns or questions. Input can be in text or voice format; in the case of voice input, the terminal uses speech recognition technology to convert it into text. For this speech recognition, speech recognition software is used as a standard service.
[0828] The server receives input data sent from the terminal and performs analysis using natural language processing technology. During the analysis process, a generative AI model is used to recognize the user's intent and emotional state. This model may utilize a natural language model such as OpenAI. To understand the emotional state, an emotion analysis algorithm is employed to evaluate the user's emotions. This allows for dynamic adjustment of responses according to the user's psychological state.
[0829] The server references multiple religious text databases based on the analysis results and selects the information most relevant to the user. The selected information is then used to generate a response tailored to the user's emotional state and intentions. This response is automatically generated by a generative AI model, enabling a more human-like interaction.
[0830] The generated response is sent from the server to the terminal and presented to the user. In addition to text display, voice feedback is also provided using speech synthesis technology. A speech synthesis engine is used for speech synthesis, allowing the user to have a more interactive experience.
[0831] Users provide feedback on the responses they receive via their devices, and the server uses this feedback data to perform machine learning to improve the accuracy of sentiment analysis. This improves the overall response accuracy of the system, enabling the provision of more accurate information.
[0832] For example, if a user inputs "I feel lonely," the system recognizes the emotional state and generates a response such as "You are not alone; many people support you." Furthermore, by presenting the AI model with prompts such as "How should I respond if the user inputs 'I feel anxious'?", it is possible to train the model to provide appropriate responses.
[0833] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0834] Step 1:
[0835] The user accesses the online platform using a terminal and logs in. The user enters their ID and password, goes through an authentication process, and is granted access to the platform. In this step, the input is the user's authentication information, and the output is a message indicating authentication success or failure. Specifically, the authentication server verifies the user information.
[0836] Step 2:
[0837] The user inputs a religious question or concern. The input method is either text or voice; in the case of voice, the device uses speech recognition software to convert the voice into text data. In this step, the input is the user's question or concern, either voice or text, and the output is the converted text data. Specifically, the speech recognition engine converts the voice to text in real time.
[0838] Step 3:
[0839] The terminal sends the converted text data to the server. The server passes the received data to a natural language processing engine, which analyzes the user's intent and emotional state. This analysis uses a generative AI model and an emotion recognition algorithm. The input is text data, and the output is data indicating the user's intent and emotional state. Specifically, the natural language processing engine analyzes the text and categorizes its intent.
[0840] Step 4:
[0841] The server selects appropriate information from multiple religious text databases based on the analysis results. A generative AI model is used to generate responses tailored to the user's emotional state. The input is the analyzed intent and emotional state, and the output is a customized response text. Specifically, database queries are executed within the server to retrieve the relevant text.
[0842] Step 5:
[0843] The server sends the generated response to the terminal. The terminal presents it to the user using text display or speech synthesis technology. For speech presentation, a speech synthesis engine is used. The input is the generated response text, and the output is visual or auditory feedback to the user. Specifically, the terminal converts the text into speech so that the user can hear it.
[0844] Step 6:
[0845] The user provides feedback on the presented response. The device sends this feedback back to the server. The server uses the feedback data to perform machine learning to improve the accuracy of the sentiment recognition engine. The input is the user's feedback, and the output is the improved sentiment analysis model. Specifically, the server analyzes the feedback data and updates the algorithm of the generative AI model.
[0846] (Application Example 2)
[0847] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0848] Existing information delivery systems often struggle to adequately consider users' personal needs and emotional states. As a result, they can only provide consistent information in response to user questions, failing to offer appropriate responses tailored to individual emotions and interests, thus limiting the user experience. Especially in online business transactions and consultations, there is a need for systems that understand user emotions and provide personalized suggestions.
[0849] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0850] In this invention, the server includes means for acquiring user input and transmitting it as information data; means for analyzing the user input using natural language processing technology based on multiple religious texts and generating an appropriate response; means for presenting the generated response to the user visually or audibly; means for analyzing the user's emotional state and adjusting the response content accordingly; and means for suggesting information on products and services of interest to the user along with the appropriate response. This makes it possible to provide highly personalized responses and suggestions that are tailored to the user's individual emotions and interests.
[0851] "Means for acquiring user input and transmitting it as information data" refers to a function for collecting information provided by the user and transmitting it as digital data to a processing device or server.
[0852] "A means of analyzing user input using natural language processing technology based on multiple religious texts and generating appropriate responses" refers to a function that uses a wide range of religious literature as a database, interprets input information using a computer, and creates information accordingly.
[0853] "Means for presenting the generated response to the user visually or audibly" refers to functions for providing the created response to the user through sight or hearing, and this includes display devices and audio output devices.
[0854] "Means for analyzing the user's emotional state and adjusting the response content based on that" refers to a function that determines the user's psychological state and optimizes the content and expression of the response accordingly.
[0855] "A means of suggesting information about products and services that users are interested in, along with appropriate responses," refers to a function that selects and recommends products and services that meet the user's needs and interests.
[0856] "Means of converting user voice input into information data" refers to technology that transforms a user's words from speech into digital representations.
[0857] "Means for collecting user feedback and using it to improve response generation and suggestion provision means" refers to technologies for gathering user reactions and opinions and using them to improve the performance and functionality of the system.
[0858] In implementing this invention, smart glasses or head-mounted displays are used as user interface devices. These devices are capable of presenting information in accordance with the user's visual and auditory perception. Specific hardware includes smart glasses devices, voice input devices, and wireless communication devices for data transmission.
[0859] First, the user provides input through an interface device. This is provided as voice input and is captured by the device's built-in microphone. The captured voice data is then converted into text data by the device's speech recognition module. In this process, speech recognition software such as Google Cloud Speech-to-Text is used as an example.
[0860] Next, the generated text data is sent to the server. Here, the server analyzes the input data using a natural language processing engine. This analysis includes, for example, language analysis techniques using spaCy. Furthermore, a custom model using TensorFlow is employed for sentiment analysis to determine the user's emotional state.
[0861] Based on the analysis results, the server generates an appropriate response. The response is tailored based on multiple religious texts and personalized information. The generated response's tone and content are modified according to the user's emotional state. Specific products or services may also be suggested based on the user's interests.
[0862] Ultimately, this response is presented to the user visually or audibly. Presentation methods include text display on a screen or voice output using speech synthesis software (e.g., Amazon Polly).
[0863] As a concrete example of its use, if a user is in a situation where they are "looking for a vegan recipe book and feeling a little tired," this system will take the user's emotional state into consideration and provide optimal book information while suggesting products that have a relaxing effect.
[0864] An example of a prompt message would be, "I'm looking for a vegan recipe book and I'm a little tired. What books would you recommend?"
[0865] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0866] Step 1:
[0867] The user provides voice input via an interface device.
[0868] The voice provided by the user is captured by the device's built-in microphone. This input data is in audio format and is processed as input to the speech recognition module.
[0869] Step 2:
[0870] The device converts the user's voice into text data.
[0871] Speech recognition software is used to analyze audio data, and as a result, text data is generated. This conversion process uses lexical analysis and acoustic models to accurately transcribe the user's speech into text.
[0872] Step 3:
[0873] The terminal sends the converted text data to the server.
[0874] The generated text data is transferred to the server using wireless communication technology. This data contains information including the user's intent and the content of their question.
[0875] Step 4:
[0876] The server analyzes the text data using a natural language processing engine.
[0877] The server uses natural language processing techniques to analyze the language structure of the received text data and interpret the user's intent. Semantic and sentiment analysis are performed to recognize the user's psychological state and requests.
[0878] Step 5:
[0879] The server generates a response based on the user's emotional state.
[0880] Based on the analyzed data, the server generates an appropriate response. This response utilizes an emotion engine to adjust its tone and content, and may include relevant products or services as needed.
[0881] Step 6:
[0882] The server sends the generated response to the terminal.
[0883] The generated response is then transmitted back to the terminal via wireless communication. This output data contains information presented to the user.
[0884] Step 7:
[0885] The device presents a response to the user, either visually or audibly.
[0886] The device presents the received data to the user using text display on the screen or speech synthesis functionality. This allows the user to acquire information through sight or hearing.
[0887] Step 8:
[0888] Users provide feedback on the response.
[0889] Users provide feedback by entering their opinions and impressions about the information presented. This information is used to continuously improve the system.
[0890] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0891] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0892] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0893] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0894] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0895] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0896] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0897] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0898] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0899] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0900] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0901] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0902] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0903] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0904] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0905] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0906] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0907] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0908] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0909] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0910] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0911] The following is further disclosed regarding the embodiments described above.
[0912] (Claim 1)
[0913] A means of obtaining user input and sending it as text data,
[0914] A means for analyzing user input using natural language processing techniques based on multiple religious texts and generating an appropriate response,
[0915] A means of presenting the generated response to the user visually or audibly,
[0916] A system that includes this.
[0917] (Claim 2)
[0918] The system according to claim 1, further comprising means for converting voice input from a user into text data.
[0919] (Claim 3)
[0920] The system according to claim 1, further comprising means for collecting user feedback and using it to improve the response generation means.
[0921] "Example 1"
[0922] (Claim 1)
[0923] A means of obtaining user input and transmitting it as information data,
[0924] A means for analyzing user input using natural language processing technology based on multiple religious data and generating an appropriate response,
[0925] Means for presenting the generated response to the user visually or audibly,
[0926] A means for obtaining user feedback and improving the accuracy of the response generation means,
[0927] A system that includes this.
[0928] (Claim 2)
[0929] The system according to claim 1, further comprising means for converting voice input from a user into information data.
[0930] (Claim 3)
[0931] The system according to claim 1, further comprising means for transmitting user input to a server using secure communication technology.
[0932] "Application Example 1"
[0933] (Claim 1)
[0934] A means of obtaining user input and sending it as text data,
[0935] A means for analyzing user input using natural language processing technology based on multiple information resources and generating an appropriate response,
[0936] A means of presenting the generated response to the user visually or audibly,
[0937] A means of measuring the user's biometric information and analyzing their emotional state,
[0938] A means of generating and presenting appropriate guidance based on emotional state,
[0939] A system that includes this.
[0940] (Claim 2)
[0941] The system according to claim 1, further comprising means for converting voice input from a user into text data.
[0942] (Claim 3)
[0943] The system according to claim 1, further comprising means for collecting user feedback and using it to improve the response generation means.
[0944] "Example 2 of combining an emotion engine"
[0945] (Claim 1)
[0946] A means of obtaining user input and transmitting it as information,
[0947] A means for analyzing user input using natural language processing technology based on a document and generating an appropriate response,
[0948] A means of presenting the generated response to the user, either by displaying it or by sound.
[0949] A means for analyzing the user's emotional state and adjusting the response based on the results,
[0950] A means of automatically generating a response using a generative AI model,
[0951] A system that includes this.
[0952] (Claim 2)
[0953] The system according to claim 1, further comprising means for converting voice input into information.
[0954] (Claim 3)
[0955] The system according to claim 1, further comprising means for collecting user feedback and using it to improve the response generation means.
[0956] "Application example 2 when combining with an emotional engine"
[0957] (Claim 1)
[0958] A means of obtaining user input and transmitting it as information data,
[0959] A means for analyzing user input using natural language processing techniques based on multiple religious texts and generating an appropriate response,
[0960] A means of presenting the generated response to the user visually or audibly,
[0961] A means for analyzing the user's emotional state and adjusting the response content based on that,
[0962] A means of suggesting information about products and services that users are interested in, along with appropriate responses,
[0963] An information processing system that includes this.
[0964] (Claim 2)
[0965] The information processing system according to claim 1, further comprising means for converting voice input from a user into information data.
[0966] (Claim 3)
[0967] The information processing system according to claim 1, further comprising means for collecting user feedback and using it to improve response generation means and suggestion provision means. [Explanation of symbols]
[0968] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of obtaining user input and sending it as text data, A means for analyzing user input using natural language processing techniques based on multiple religious texts and generating an appropriate response, A means of presenting the generated response to the user visually or audibly, A system that includes this.
2. The system according to claim 1, further comprising means for converting voice input from a user into text data.
3. The system according to claim 1, further comprising means for collecting user feedback and using it to improve the response generation means.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A