Computer and information processing method
The system addresses the inflexibility of conventional chatbots by using a large-scale language model and database servers to generate dynamic responses based on user actions, improving interaction flexibility and engagement.
Patent Information
- Application Number
- PCT/JP2025/013582
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-03
- Filing Date
- 2025-04-03
- Publication Date
- 2025-10-09
AI Technical Summary
Conventional chatbots lack flexibility in their responses, as they rely on pre-defined databases and cannot adapt to user interactions effectively.
A computer system that generates input sentences based on user actions, utilizing a large-scale language model and database servers to create dynamic responses, allowing for more flexible and personalized interactions.
Enables chatbots to provide more natural and adaptable responses by generating input sentences based on user actions, enhancing user engagement and interaction flexibility.
Smart Images

Figure JP2025013582_09102025_PF_FP_ABST
Abstract
Description
Computer and information processing method
[0001] The present disclosure relates to computers and information processing methods related to chatbots.
[0002] Chatbots are computer programs that can converse with humans in natural language and are used in a wide range of fields, including customer service, information search, education, and entertainment. In recent years, the use of artificial intelligence and other technologies has enabled chatbots to perform more complex tasks and provide more natural and fluent interactions with humans. Patent Document 1 discloses a method for acquiring suggested response items based on user input or events in an embedded application and providing corresponding commands. Specifically, Japanese Patent No. 6718028 (JP6718028B) discloses a method for displaying suggested response items between user devices and providing associated commands based on events occurring in chat or an embedded application. Furthermore, the suggested response items are displayed in a chat interface or an embedded interface, and specific commands can be executed according to a user's selection.
[0003] Conventional chatbots could only provide answers based on a database that stores combinations of user input sentences and answers, which meant that they had limited flexibility in their answers.
[0004] The present disclosure has been made in consideration of these points, and aims to provide a computer and an information processing method that can increase the flexibility of responses by creating input sentences based on information corresponding to the user's actions.
[0005] The computer of the present disclosure comprises a memory and a processor, and by executing a computer program stored in the memory, the processor receives action information regarding an action performed by a user on an application screen displayed on a user terminal, or on a chatbot screen that is a screen for operating a chatbot and is different from the application screen; generates an input sentence corresponding to the action based on the received action information; obtains information on a response to the generated input sentence; and transmits display information for displaying the obtained response on the chatbot screen to the user terminal.
[0006] In the computer of the present disclosure, when obtaining information about an answer to the generated input sentence, the generated input sentence may be sent to a large-scale language model, and the answer to the input sentence calculated by the large-scale language model may be received from the large-scale language model, thereby obtaining the information about the answer.
[0007] In addition, when an input indicating consent to the use of the chatbot is made on the user terminal, an action detection program that detects the action performed by the user on the user terminal is executed on the user terminal, and the action information regarding the action performed by the user on the application screen detected by the action detection program may be transmitted from the user terminal to the processor.
[0008] Further, the action is an action in which the user selects an image displayed on the application screen and drops it onto the chatbot screen, and when generating an input sentence corresponding to the action, if the dropping action is detected, a base answer to a pre-input sentence generated based on the dropped image and the action information is obtained, and the input sentence is generated based on the base answer and the action information.
[0009] Furthermore, when generating an input sentence corresponding to the action, if the dropping action is detected, the pre-input sentence generated based on the dropped image and the action information may be sent to a first database server, and the base answer to the pre-input sentence extracted by the first database server may be received from the first database server, thereby obtaining information on the base answer.
[0010] Furthermore, when sending display information for displaying the acquired answer on the chatbot screen to the user terminal, an instruction requesting approval of the acquired base answer may be sent to the user terminal, and when generating an input sentence according to the action, if information regarding approval of the base answer is received from the user terminal, the input sentence may be generated based on the approved base answer and the action information.
[0011] Furthermore, when generating an input sentence corresponding to the action, specification information of the user terminal may be further acquired, and the input sentence may be generated by referring to the specification information in addition to the base answer and the action information.
[0012] Furthermore, the action may be an action in which the user accesses a second database server that stores the information displayed on the application screen, and when generating an input sentence corresponding to the action, the action information regarding the action to be accessed and the information stored in the second database server to be accessed may be obtained based on the content entered by the user on the chatbot screen, and the input sentence may be generated based on the content entered by the user, the action information, and the information stored in the second database server to be accessed.
[0013] Furthermore, the type of the second database server to be accessed from among a plurality of types of the second database servers may be determined according to the type of the action.
[0014] In addition, the processor may acquire information about the user from the user terminal by executing a computer program stored in the memory, and when generating an input sentence corresponding to the action, may generate the input sentence corresponding to the action by referring to the acquired information about the user in addition to the received action information.
[0015] Furthermore, when obtaining information on an answer to the generated input statement, the generated input statement may be sent to a third database server in which data on the set input statement and the set answer are stored in association with each other, and the answer corresponding to the input statement extracted based on the data on the set input statement and the set answer stored in the third database server may be received from the third database server, thereby obtaining information on the answer.
[0016] In addition, the memory may store data on a set input sentence and a set answer in association with each other, and when obtaining information on an answer to the generated input sentence, the information on the answer corresponding to the input sentence may be obtained based on the data on the set input sentence and the set answer stored in the memory.
[0017] In addition, the processor may execute a computer program stored in the memory to send an instruction to an administrator terminal inquiring about the content of the input sentence corresponding to the content of the action in the received action information, and perform learning by associating the content of the input sentence received from the administrator terminal with the content of the action, and when generating an input sentence corresponding to the action, upon receiving the action information, may generate the input sentence corresponding to the action based on the content of the learning that has been performed.
[0018] Furthermore, when generating an input sentence corresponding to the action, a plurality of input sentences corresponding to the action may be generated, and the generated plurality of input sentences may be displayed on the user terminal, thereby making it possible to select one of the plurality of input sentences or to create a new input sentence on the user terminal.
[0019] The computer of the present disclosure comprises a memory and a processor, and by executing a computer program stored in the memory, the processor receives action information regarding an action performed by a user on a user terminal, generates an input sentence corresponding to the action based on the received action information, obtains information on a response to the generated input sentence, and transmits display information to the user terminal for displaying the obtained response on a chatbot screen.When an input indicating consent to use of the chatbot is made on the user terminal, an action detection program is executed on the user terminal, which detects actions other than input to the chatbot performed by the user on the user terminal, and the action information regarding the actions other than input to the chatbot performed by the user, detected by the action detection program, is transmitted from the user terminal to the processor.
[0020] The information processing method disclosed herein is an information processing method performed by a computer having a memory in which a computer program is stored and a processor that executes the computer program stored in the memory, and includes the steps of: receiving action information regarding an action performed by a user on an application screen displayed on a user terminal and a chatbot screen that is a screen for operating a chatbot and is different from the application screen; generating an input sentence corresponding to the action based on the received action information; obtaining information on a response to the generated input sentence; and transmitting display information to the user terminal for displaying the obtained response on the chatbot screen.
[0021] The information processing method disclosed herein is an information processing method performed by a computer having a memory in which a computer program is stored and a processor that executes the computer program stored in the memory, and includes the steps of: receiving action information regarding an action performed by a user on a user terminal; generating an input sentence corresponding to the action based on the received action information; obtaining information on an answer to the generated input sentence; transmitting display information to the user terminal for displaying the obtained answer on a chatbot screen; when an input indicating consent to use of the chatbot is made on the user terminal, an action detection program that detects actions other than input to the chatbot performed by the user on the user terminal is executed on the user terminal; and the action information regarding the actions other than input to the chatbot performed by the user detected by the action detection program is transmitted from the user terminal to the processor.
[0022] 1 is a block diagram showing a configuration of a system according to an embodiment of the present disclosure. FIG. 1 is a block diagram showing a configuration of a user terminal in the system shown in FIG. 1. FIG. 2 is a block diagram showing a configuration of a tag management server in the system shown in FIG. 1. FIG. 3 is a block diagram showing a configuration of a web server in the system shown in FIG. 1. FIG. 4 is a block diagram showing a configuration of a second database server in the system shown in FIG. 1. FIG. 5 is a block diagram showing a configuration of a chat server in the system shown in FIG. 1. FIG. 6 is a block diagram showing a configuration of a language model server in the system shown in FIG. 1. FIG. 7 is a block diagram showing a configuration of a first database server in the system shown in FIG. 1. FIG. 8 is a chart showing the flow of information between components when a user agrees to use a chatbot on a user terminal in the system shown in FIG. 1. FIG. 9 is a chart showing the flow of information between components when an action is detected on a user terminal in the system shown in FIG. 1. FIG. 10 is a flowchart showing the operation of a chat server in the system shown in FIG. 1. FIG. 11 is a diagram showing the contents of a screen displayed on a display unit of a user terminal in the system shown in FIG. 1. FIG. 12 is a diagram showing the contents of a screen displayed on a display unit of a user terminal in the system shown in FIG. 1.
[0023] Hereinafter, an embodiment of the present disclosure will be described with reference to the drawings. Figures 1 to 11 are diagrams showing a system 1 according to this embodiment and the components of this system 1. Figures 12 to 16 are diagrams showing the contents of a screen displayed on a display unit of a user terminal in the system shown in Figure 1.
[0024] [Overall Configuration of System 1] As shown in FIG. 1 , the system 1 of this embodiment includes a tag management server 20, a web server 30, a second database server 40, a chat server 50, a language model server 60 (large-scale language model), a first database server 70, an administrator terminal 80, and a third database server 90. The system 1 of this embodiment enables a user to use a chatbot via a browser on a user terminal 10, such as a personal computer, a PC tablet, or a smartphone owned by the user. The tag management server 20 is communicatively connected to the user terminal 10 via a communication network such as the Internet. The web server 30 is communicatively connected to each of the user terminal 10 and the second database server 40 via a communication network such as the Internet. The chat server 50 is communicatively connected to each of the user terminal 10, the language model server 60, the first database server 70, the administrator terminal 80, and the third database server 90 via a communication network such as the Internet. Each component of the system 1 and the user terminal 10 will be described in detail below.
[0025] [Configuration of User Terminal 10] The configuration of the user terminal 10 will be described using Figure 2. As described above, the user terminal 10 includes, but is not limited to, a personal computer, a PC tablet, a smartphone, etc. As shown in Figure 2, the user terminal 10 has a control unit 11 that functions as a processor, a display unit 13, an operation unit 14, an imaging unit 15, a microphone 16, a storage unit 18 that functions as a memory, and a communication unit 19. The control unit 11 is connected to each of the display unit 13, the operation unit 14, the imaging unit 15, the microphone 16, the storage unit 18, and the communication unit 19 via a bus 11a.
[0026] The control unit 11 is composed of a microcomputer including a CPU and semiconductor memory, and controls the operation of the user terminal 10 by executing a computer program (hereinafter simply referred to as a program) stored in the storage unit 18. The display unit 13 is composed of, for example, a liquid crystal display, and functions as a means for displaying various information. The display unit 13 displays various information in response to instructions from the control unit 11. The operation unit 14 functions as a means for inputting various instructions by the user. For example, a keyboard, a mouse, a touch panel, etc. are used as the operation unit 14. When the operation unit 14 is a touch panel, such a touch panel is superimposed on the display unit 13, and an operation signal is input to the control unit 11 when the user touches the touch panel.
[0027] The imaging unit 15 is, for example, a camera, and captures images or videos by capturing an object. The microphone 16 converts voices emitted by the user and sound waves in the air around the user terminal 10 into electrical signals. The storage unit 18 is composed of a hard disk drive (HDD), random access memory (RAM), read-only memory (ROM), solid state drive (SSD), or the like. The storage unit 18 stores programs executed by the control unit 11. The storage unit 18 also stores information input by the user via the operation unit 14, images and videos captured by the imaging unit 15, electrical signals converted from sound waves by the microphone 16, and the like. The communication unit 19 includes a communication interface that transmits and receives various data between the control unit 11 and external devices such as the tag management server 20, the web server 30, and the chat server 50 via a communication network.
[0028] [Configuration of Tag Management Server 20] The configuration of the tag management server 20 will be described with reference to FIG. 3. The tag management server 20 is a platform that can manage website tags by combining predetermined commands without directly writing code, and can cause the user terminal 10 to execute Javascript as an action detection program. Specifically, by entering code that accesses the tag management server 20 into a website displayed on the display unit 13 of the user terminal 10, action detection and chatbot execution via the tag management server 20 are possible. For example, by executing Javascript as an action detection program obtained from the tag management server 20 on the user terminal 10, a specific action on the user terminal 10 is detected, and information about the detected action is sent to the chat server 50. Note that the action detection program, such as Javascript, executed by the user terminal 10 detects actions other than input to a chatbot by the user on the user terminal 10. 3, the tag management server 20 has a control unit 21 that functions as a processor, a storage unit 28 that functions as a memory, and a communication unit 29. The control unit 21 is connected to each of the storage unit 28 and the communication unit 29 via a bus 21a.
[0029] The control unit 21 is configured by a computer including a CPU and a semiconductor memory, and controls the operation of the tag management server 20 by executing a program stored in the storage unit 28. Specifically, the control unit 21 functions as a reception unit 22 and a transmission unit 23 by executing the program stored in the storage unit 28. When the reception unit 22 receives access from the user terminal 10, the transmission unit 23 transmits tag information to the user terminal 10 that is the sender of the access information.
[0030] The storage unit 28 is configured with an HDD, RAM, ROM, SSD, etc. The storage unit 28 stores programs executed by the control unit 21. The storage unit 28 also stores tag information to be transmitted to the user terminal 10.
[0031] The communication unit 29 includes a communication interface that transmits and receives various data between the control unit 21 and the user terminal 10 via a communication network.
[0032] [Configuration of Web Server 30] The configuration of the web server 30 will be described with reference to FIG. 4. The web server 30 is managed by a business (client) that sells products or provides services to users. The web server 30 displays websites for selling products or providing services on a browser or the like displayed on the display unit 13 of the user terminal 10. Websites displayed on the display unit 13 of the user terminal 10 by the web server 30 include various sites, such as e-commerce sites and accommodation reservation sites. When a chatbot is running on a website, the display unit 13 of the user terminal 10 accessing the website includes an application screen and a chatbot screen, which is a screen for operating the chatbot and is different from the application screen. The application screen refers to a display area that displays content different from the chatbot on the website accessed by the user. The chatbot screen refers to a display area on the website where the chatbot is displayed. The chatbot screen may be displayed on the website or elsewhere. Specifically, the window displaying the website may have both an application display area and a chatbot display area, or two windows may be displayed on the display unit 13: one displaying the application screen and the other displaying the chatbot screen.
[0033] As shown in FIG. 4 , the web server 30 includes a control unit 31 functioning as a processor, a storage unit 38 functioning as a memory, and a communication unit 39. The control unit 31 is connected to the storage unit 38 and the communication unit 39 via a bus 31a. The control unit 31 is configured as a microcomputer including a CPU and semiconductor memory, and controls the operation of the web server 30 by executing a program stored in the storage unit 38. Specifically, the control unit 31 functions as a reception unit 32 and a transmission unit 33 by executing a program stored in the storage unit 38. When the reception unit 32 receives URL information from the user terminal 10, the transmission unit 33 transmits to the user terminal 10 a display instruction for a website corresponding to the URL information received by the reception unit 32 and stored in the storage unit 38. This causes the website to be displayed in a browser or the like displayed on the display unit 13 of the user terminal 10. Furthermore, when the user inputs consent to use of the chatbot on the user terminal 10, the reception unit 32 receives consent information from the user terminal 10. The consent information is information indicating that the user has consented. When the acceptance means 32 accepts the consent information, the transmission means 33 transmits to the user terminal 10 an instruction to display a chatbot screen and an instruction to access the tag management server 20. As a result, the chatbot screen is displayed on the website displayed on the display unit 13 of the user terminal 10. Furthermore, when the acceptance means 32 accepts login information from the user terminal 10, it transmits access information to the second database server 40.
[0034] The storage unit 38 is configured with a HDD, RAM, ROM, SSD, etc. The storage unit 38 stores programs executed by the control unit 31. The storage unit 38 also stores various website information (specifically, HTML information that constructs websites, etc.) in association with URL information.
[0035] The communication unit 39 includes a communication interface that transmits and receives various data between the control unit 31, the user terminal 10, and the second database server 40 via a communication network.
[0036] [Configuration of Second Database Server 40] The configuration of the second database server 40 will be described with reference to FIG. 5. The second database server 40 is managed by the client and stores a membership database, product information, service information, and the like related to the client's web server. The membership database, product information, service information, and the like stored in the second database server 40 are linked to each website. For example, if the website displayed on the display unit 13 of the user terminal 10 by the web server 30 is an e-commerce site, the membership database uses the name, address, telephone number, past product purchase history, and e-commerce site browsing history of the user of the e-commerce site. Furthermore, if the website displayed on the display unit 13 of the user terminal 10 by the web server 30 is an e-commerce site, the product information uses the name, specification information, sales price, past sales quantity, and the like of the product sold on the e-commerce site. Furthermore, if the website displayed on the display unit 13 of the user terminal 10 by the web server 30 is an accommodation reservation site, the member database includes the name, address, telephone number, past accommodation usage history, and browsing history of the accommodation reservation site user of the accommodation reservation site. Furthermore, if the website displayed on the display unit 13 of the user terminal 10 by the web server 30 is an accommodation reservation site, the service information includes the name, address, URL of the accommodation website, accommodation price, and other information of accommodations available for reservation on the reservation site. Although only one second database server 40 is illustrated in FIG. 1 , multiple types of second database servers 40 may be communicatively connected to the web server 30, and the type of second database server 40 to be accessed may be determined from the multiple types of second database servers 40 depending on the type of user action (described later). As shown in FIG. 5 , the second database server 40 includes a control unit 41 functioning as a processor, a storage unit 48 functioning as a memory, and a communication unit 49. The control unit 41 is connected to each of the storage unit 48 and the communication unit 49 via a bus 41a.
[0037] The control unit 41 is configured by a microcomputer including a CPU and semiconductor memory, and controls the operation of the second database server 40 by executing a program stored in the storage unit 48. Specifically, the control unit 41 functions as a reception means 42 and a transmission means 43 by executing a program stored in the storage unit 48. When the reception means 42 receives access information from the web server 30, the transmission means 43 transmits the user's personal information, product information, service information, etc. to the web server 30 based on the member database, product information, service information, etc. stored in the storage unit 48.
[0038] The storage unit 48 is composed of a HDD, RAM, ROM, SSD, etc. The storage unit 48 stores programs executed by the control unit 41. The storage unit 48 also stores a membership database, product information, service information, etc. linked to each website.
[0039] The communication unit 49 includes a communication interface that transmits and receives various data between the control unit 41 and the web server 30 via a communication network.
[0040] [Configuration of Chat Server 50] The configuration of the chat server 50 will be described using Fig. 6. The chat server 50 is managed by a business (administrator) that provides chatbot services, and manages chatbots on a website displayed on the display unit 13 of the user terminal 10. As shown in Fig. 6, the chat server 50 has a control unit 51 that functions as a processor, a storage unit 58 that functions as memory, and a communication unit 59. The control unit 51 is connected to each of the storage unit 58 and the communication unit 59 via a bus 51a.
[0041] The control unit 51 is composed of a microcomputer including a CPU and a semiconductor memory, and controls the operation of the chat server 50 by executing a program stored in the storage unit 58. Specifically, the control unit 51 functions as a reception unit 52, a user information acquisition unit 53, an input sentence generation unit 54, a response acquisition unit 55, and a transmission unit 56 by executing the program stored in the storage unit 58. The reception unit 52 receives action information and the like from the user terminal 10. The user information acquisition unit 53 acquires information about the user from the user terminal 10. The input sentence generation unit 54 generates an input sentence corresponding to the action based on the action information and the like received by the reception unit 52. The response acquisition unit 55 acquires information about a response to the input sentence generated by the input sentence generation unit 54. The transmission unit 56 transmits display information to the user terminal 10 for displaying the response acquired by the response acquisition unit 55 on the chatbot screen. Details of the functions of each of these units 52, 53, 54, 55, and 56 will be described later.
[0042] The storage unit 58 is configured with a HDD, RAM, ROM, SSD, etc. The storage unit 58 stores programs executed by the control unit 51.
[0043] The communication unit 59 includes a communication interface that transmits and receives various data between the control unit 51 and each of the user terminal 10, the language model server 60, the first database server 70, the administrator terminal 80, and the third database server 90 via a communication network.
[0044] [Configuration of Language Model Server 60] The configuration of the language model server 60 will be described with reference to FIG. 7 . The language model server 60 is a server designed to be able to generate replies quickly in response to requests from the chat server 50 using, for example, a large-scale language model (LLM). A large-scale language model is a machine learning model for natural language processing, and is a model that has the function of generating sentences based on input information. For example, the large-scale language model is a machine learning model based on a Transformer model with a self-attention mechanism. An example of the Transformer model is a GPT (Generative Pre-trained Transformer) model. A large-scale language model is trained using a large dataset, enabling general-purpose sentence generation tasks. For example, a model trained using a large dataset using the GPT-3 or GPT-4 algorithm developed by OpenAI (registered trademark) (for example, a GPT model generated by combining training using a large dataset and reinforcement learning using a reward prediction model) can be applied as the large-scale language model. The language model server 60 hosts a model, invokes the model in response to a request, and returns the results. The language model server 60 is also capable of reusing a model that has been trained once. As shown in FIG. 7 , the language model server 60 includes a control unit 61 that functions as a processor, a storage unit 68 that functions as a memory, and a communication unit 69. The control unit 61 is connected to the storage unit 68 and the communication unit 69 via a bus 61 a. The language model server 60 is also designed to vectorize input text using a vector conversion model and output vector values. The vector conversion model is a machine learning model that converts input text into coordinate information in multiple dimensions (e.g., 1,536 dimensions) based on parameters generated by training. For example, the vector conversion model is an embedding API model developed by OpenAI (registered trademark).
[0045] The control unit 61 is configured by a microcomputer including a CPU and a semiconductor memory, and controls the operation of the language model server 60 by executing a program stored in the storage unit 68. Specifically, the control unit 61 functions as a reception unit 62, a response generation unit 63, and a transmission unit 64 by executing the program stored in the storage unit 68. When the reception unit 62 receives an input sentence from the chat server 50, the response generation unit 63 generates a response using, for example, OpenAI (registered trademark), and the transmission unit 64 transmits the generated response to the chat server 50.
[0046] The storage unit 68 is configured with an HDD, RAM, ROM, SSD, etc. The storage unit 68 stores programs executed by the control unit 61. The storage unit 68 also stores a huge amount of learned data.
[0047] The communication unit 69 includes a communication interface that transmits and receives various data between the control unit 61 and the chat server 50 via a communication network.
[0048] [Configuration of First Database Server 70] The configuration of the first database server 70 will be described with reference to FIG. 8 . The first database server 70 is managed by a business operator (administrator) that provides chatbot services, and stores vector values vectorized by the language model server 60. Note that a vector value refers to coordinate information in multiple dimensions (e.g., 1,536 dimensions). For example, a vector value is a feature that indicates magnitude and direction in multiple dimensions. As shown in FIG. 8 , the first database server 70 includes a control unit 71 that functions as a processor, a storage unit 78 that functions as a memory, and a communication unit 79. The control unit 71 is connected to each of the storage unit 78 and the communication unit 79 via a bus 71 a. For example, the vector values stored in the first database server 70 are values obtained by converting prepared anticipated questions into vector values by the language model server 60. The first database server 70 also stores pairs of anticipated questions converted into vector values and base answers that serve as anticipated answers to the anticipated questions. The vector values stored in the first database server 70 may be vector values that have been vectorized by an information processing device other than the language model server 60 .
[0049] The control unit 71 is configured as a computer including a CPU and semiconductor memory, and controls the operation of the first database server 70 by executing a program stored in the storage unit 78. Specifically, the control unit 71 functions as a receiving unit 72 and a transmitting unit 73 by executing a program stored in the storage unit 78. When the receiving unit 72 receives a vector search request from the chat server 50, the transmitting unit 73 transmits to the chat server 50 a base answer corresponding to an approximate value of the vector value based on the vector value stored in the storage unit 78. Here, the vector search is a process of calculating the distance between pieces of information contained in the vectors to calculate the approximate value. For example, the vector search is a process of determining the cosine similarity between the vector value to be searched included in the vector search request and multiple stored vector values, and transmitting the base answer that is the pair of vector values with the highest cosine similarity as the search result to the chat server 50.
[0050] The storage unit 78 is configured with a HDD, RAM, ROM, SSD, etc. The storage unit 78 stores programs executed by the control unit 71. The storage unit 78 also stores base answers and vector values in an associated state.
[0051] The communication unit 79 includes a communication interface that transmits and receives various data between the control unit 71 and the chat server 50 via a communication network.
[0052] The configurations of the administrator terminal 80 and the third database server 90 will be described in detail later.
[0053] [Operation of System 1] Next, the operation of system 1 according to this embodiment will be described with reference to Figs. 9 to 16. Fig. 9 is a chart showing the flow of information between the components when a user agrees to use a chatbot on the user terminal 10 in system 1 shown in Fig. 1, and Fig. 10 is a chart showing the flow of information between the components when an action is detected on the user terminal 10 in system 1 shown in Fig. 1. Fig. 11 is a flowchart showing the operation of chat server 50 in system 1 shown in Fig. 1. Figs. 12 to 16 are diagrams showing the contents of screens displayed on display unit 13 of user terminal 10 in system 1 shown in Fig. 1, respectively.
[0054] First, the flow of information between the components of the system 1 shown in FIG. 1 when a user agrees to use a chatbot on the user terminal 10 will be described with reference to FIG. 9.
[0055] When a user accesses a predetermined website using a browser or the like displayed on the display unit 13 of the user terminal 10, the receiving means 32 of the web server 30 receives URL information of the predetermined website from the user terminal 10. The transmitting means 33 transmits to the user terminal 10 a display instruction for the website corresponding to the URL information received by the receiving means 32, which is stored in the memory unit 38. As a result, the predetermined website is displayed on the browser or the like displayed on the display unit 13 of the user terminal 10. Furthermore, when the user agrees to use the chatbot on this website, consent information is transmitted from the user terminal 10 to the web server 30. When the receiving means 32 of the web server 30 receives the consent information from the user terminal 10, the transmitting means 33 transmits to the user terminal 10 an instruction to display a chatbot screen and an instruction to access the tag management server 20. As a result, the chatbot screen is displayed on the website displayed on the display unit 13 of the user terminal 10. Note that if the user does not agree to use the chatbot on the website displayed on the display unit 13 of the user terminal 10, the chatbot screen will not be displayed on the website. Furthermore, an access command to the tag management server 20 is sent to the user terminal 10, thereby allowing the user terminal 10 to access the tag management server 20. When the receiving means 22 of the tag management server 20 receives access from the user terminal 10, the transmitting means 23 transmits tag information (e.g., information including an action detection program) to the user terminal 10 that is the sender of the access information. Then, when the tag information transmitted from the tag management server 20 to the user terminal 10 is read by the user terminal 10, communication between the user terminal 10 and the chat server 50 becomes possible. Specifically, when the user terminal 10 executes Javascript as the action detection program acquired from the tag management server 20, a specific action in the user terminal 10 is detected, and the detected action information is transmitted to the chat server 50. Furthermore, the chatbot displayed on the website is executed.
[0056] When a user logs in to their personal page on a specific website displayed on the display unit 13 of the user terminal 10 by, for example, entering a login ID and password, the login information is sent from the user terminal 10 to the web server 30, allowing the user to access their personal page. When the accepting means 32 of the web server 30 accepts the login information, the transmitting means 33 sends access information to the second database server 40, thereby enabling the user to access the member database, product information, service information, etc. linked to the website in the second database server 40. In this way, when the web server 30 obtains the user's personal information, product information, and service information from the second database server 40, the information is encrypted using a public key issued by the operator (administrator) providing the chatbot service. The user's personal information, product information, etc. encrypted by the web server 30 are sent to the user terminal 10. The operation of encrypting the user's personal information, product information, service information, etc. acquired from the second database server 40 by the web server 30 and transmitting the information to the user terminal 10 as described above may be performed when a user action, which will be described later, is detected, or may be performed in advance before the action is detected. Performing the operation in advance before the action is detected has the advantage that the series of processes can be performed quickly, but it also has the problem that the capacity of the storage unit 18 of the user terminal 10 must be increased because the encrypted user's personal information, product information, etc. must be stored in the storage unit 18.
[0057] Next, the flow of information between the components and the operation of the chat server 50 when an action is detected by the user terminal 10 in the system 1 shown in FIG. 1 will be described with reference to FIGS. 10 and 11. FIG.
[0058] When a user performs a specific action on a website displayed on the display unit 13 of the user terminal 10, the specific action on the user terminal 10 is detected by JavaScript, an action detection program provided by the tag management server 20 and executed on the user terminal 10, and information about the detected action is sent to the chat server 50. Here, specific actions include an action of selecting specific text, image, or video on the website displayed on the user terminal 10, an action of selecting a specific image on the website displayed on the user terminal 10 and dropping it onto the chatbot screen, an action of accessing a database that stores information displayed on the application screen, an action of scrolling a website on the browser, an action of transitioning from one website to another on the browser, etc. The action of selecting an image or video also includes the act of hovering the pointer over an image or video. Furthermore, user actions detected by Javascript include a first action that is detected even if the user does not input any questions or the like on the chatbot screen, and a second action that is detected when the user inputs a question or the like on the chatbot screen after the user has performed a predetermined action. Details of such first and second actions will be described later. Note that user actions detected by Javascript are not limited to those described above, but also include various actions other than those described above. When a user action is detected by Javascript, the action information and the user's personal information, product information, etc. encrypted by the web server 30 are sent to the chat server 50.
[0059] When the user terminal 10 and the chat server 50 are in a communicable state ("YES" in step S1 of FIG. 11), the accepting means 52 of the chat server 50 accepts action information, etc. ("YES" in step S2 of FIG. 11), generates a pre-input sentence based on the accepted action information, etc. and predetermined pre-input generation information, and the transmitting means 56 of the chat server 50 transmits an instruction to the language model server 60 to vectorize the pre-input sentence (step S3 of FIG. 11). When the accepting means 62 of the language model server 60 accepts the instruction, the answer generating unit 63 obtains vector values by vectorizing the pre-input sentence. The transmitting means 64 of the language model server 60 transmits the vector information (specifically, vector values) generated by the answer generating unit 63 to the chat server 50. As a result, when the receiving means 52 of the chat server 50 receives the vector information (step S4 in FIG. 11 ), the transmitting means 56 transmits to the first database server 70 a vector search request to search the first database server 70 for vector values that approximate the vector values of the pre-input sentence generated by the language model server 60 (step S5 in FIG. 11 ). When the receiving means 72 of the first database server 70 receives the vector search request from the chat server 50, it extracts vector values that approximate the vector values in the received vector search request from the storage unit 78, and the transmitting means 73 of the first database server 70 transmits a base answer corresponding to this approximate value to the chat server 50. As a result, the receiving means 52 of the chat server 50 receives from the first database server 70 a base answer that approximates the vector value transmitted from the language model server 60 (step S6 in FIG. 11 ). In another aspect, the control unit 51 of the chat server 50 may determine, based on predetermined pre-input generation information, information necessary for a base answer from the user's actions detected by Javascript and the content of the question entered by the user on the chatbot screen, and request and acquire the information necessary for the base answer (encrypted user personal information, product information, etc.) from the user terminal 10. The pre-input generation information is information including rules, etc. necessary for generating a pre-input sentence based on the accepted action.For example, the pre-input generation information is a rule that, when the received action is "an action in which a user selects an image displayed on an application screen and drops the image onto a chatbot screen," reads a predetermined sentence, such as "Please tell me information about {image}," from the content of the action, obtains metadata of the dropped image (e.g., part name A), and combines it with the sentence to generate a pre-input sentence, such as "Please tell me information about part name A." Note that the information may be a rule-based model or a machine learning model. For example, a rule that "the action of dropping an image determines that the user wants to know information about the image" may be stored in a large-scale language model in advance, and the rule that "the action content (the user performed the action of dropping an image), the metadata of the image dropped into the action content is 'part name A', generate a possible question for the user" may be input to the large-scale language model to generate a pre-input sentence.
[0060] Furthermore, in the chat server 50, the control unit 51 decrypts the encrypted information using an encryption key issued by the operator (administrator) providing the chatbot service (step S7 in FIG. 11 ). In this manner, the user information acquisition means 53 acquires the user's personal information. Note that information such as the specifications of the user terminal 10 itself may be acquired by Javascript executed on the user terminal 10, and the acquired information may be transmitted from the user terminal 10 to the chat server 50 by Javascript, thereby allowing the user information acquisition means 53 to acquire the information such as the specifications of the user terminal 10 itself. Then, the input sentence generation means 54 generates an input sentence to be sent to the language model server 60 based on the acquired base answer, action information, and decrypted information (step S8 in FIG. 11 ). The transmission means 56 transmits the input sentence generated by the input sentence generation means 54 to the language model server 60 (step S9 in FIG. 11 ). When the receiving means 62 of the language model server 60 receives an input sentence, the answer generating unit 63 generates an answer corresponding to the input sentence, and the transmitting means 64 transmits the answer generated by the answer generating unit 63 to the chat server 50. When the receiving means 52 receives an answer from the language model server 60 in this manner, the answer acquiring means 55 acquires the answer from the language model server 60 corresponding to the input sentence generated by the input sentence generating means 54 (step S10 in FIG. 11 ). Then, the transmitting means 56 of the chat server 50 transmits an instruction to the user terminal 10 to display the answer information generated by the language model server 60 (step S11 in FIG. 11 ). As a result, the answer generated by the language model server 60 is displayed on the chatbot screen of the website displayed on the display unit 13 of the user terminal 10.
[0061] A specific example of the operation of the system 1 will be described in more detail using the display screens on the display unit 13 of the user terminal 10 shown in FIGS.
[0062] First, we will explain an example of a first action detected by Javascript even if the user does not enter any questions or other information on the chatbot screen. Figure 12 is a diagram showing a display screen on the display unit 13 of the user terminal 10 when an e-commerce site for users to purchase electrical appliances and their parts is displayed as a website provided by the web server 30. The left area 13a of this display screen is an application screen that displays the product names, model numbers, images 13c, specifications, etc. of the electrical appliances and their parts available for purchase, and the right area 13b of the display screen shown in Figure 12 is a chatbot screen. As mentioned above, this chatbot screen is not displayed if the user does not agree to use the chatbot on the website displayed on the display unit 13 of the user terminal 10. If the user does not consent to the use of the chatbot, the chatbot will not detect the user's actions and will be activated to respond to the user's questions entered through the user terminal 10 based solely on information stored in the first database server 70 and the third database server 90 (described later) (i.e., without using information stored in the second database server 40). When a chatbot screen is displayed on a website, the user can input questions to the chatbot screen using the operation unit 14. For example, if the user enters a question on the chatbot screen using the operation unit 14 to inquire about detailed information about an electrical appliance or its components displayed on the application screen, the answer to the input question will be displayed on the chatbot screen. When a user logs in to their personal page on a website such as the one shown in FIG. 12, the web server 30 encrypts the user's personal information, product information, etc., and the encrypted information is transmitted from the web server 30 to the user terminal 10, where it is temporarily stored in the storage unit 18 of the user terminal 10.
[0063] Furthermore, when a user selects an image 13c of an electrical appliance or its component displayed on an application screen on the display unit 13 of the user terminal 10 with the cursor and drops it onto the chatbot screen as shown in FIG. 12 , this action is detected by Javascript as first action information. Then, Javascript determines information to send to the chat server 50 based on the first action information, and the action information based on the first action information, information about the image 13c, and encrypted personal information, product information, etc. are sent from the user terminal 10 to the chat server 50. The information about the image 13c includes metadata assigned to the image 13c (e.g., the title and description of the image 13c). The encrypted personal information includes the user's past purchase history of electrical appliances and components. The encrypted product information includes information such as the specifications and price of the product displayed on the application screen. In addition, information such as the specifications of the user terminal 10 itself is obtained by Javascript executed on the user terminal 10, and the obtained information is sent from the user terminal 10 to the chat server 50 by Javascript, whereby information such as the specifications of the user terminal 10 itself is obtained by the user information acquisition means 53.
[0064] When the accepting means 52 of the chat server 50 accepts action information and the like from the user terminal 10, the transmitting means 56 of the chat server 50 transmits to the language model server 60 a command to vectorize a pre-input sentence (e.g., a sentence such as "I want to know the specifications of image 13c") generated based on the action information and the like accepted by the accepting means 52 (specifically, the action of selecting with the cursor an image 13c of an electrical appliance or its component displayed on the application screen and dropping it on the chatbot screen, information on image 13c, and identification information of the product corresponding to image 13c) and the pre-input generation information. When the accepting means 62 of the language model server 60 accepts the command, the answer generating unit 63 obtains a vector value by vectorizing the pre-input sentence generated based on the action information and the like. Thereafter, when the accepting means 52 of the chat server 50 accepts the vector value as vector information, the transmitting means 56 transmits to the first database server 70 a vector search request to search the first database server 70 for a vector value approximate to the vector value generated by the language model server 60. When the receiving means 72 of the first database server 70 receives a vector search request from the chat server 50, it extracts vector values that are approximate to the vector value in the received vector search request from the vector values stored in the storage unit 78, and the transmitting means 73 of the first database server 70 transmits a base answer corresponding to this approximate value to the chat server 50. Here, for example, an answer such as "The specifications of the product corresponding to the image 13c dropped on the chatbot screen are A" is obtained as the base answer corresponding to the approximate value.
[0065] In the chat server 50, the control unit 51 decrypts encrypted information using an encryption key issued by the operator (administrator) providing the chatbot service. The input sentence generation means 54 generates an input sentence to the language model server 60 based on predetermined input generation information, based on the acquired base answer (specifically, the specifications of the product corresponding to the image 13c dropped on the chatbot screen), the information on the image 13c, the action information, and the decrypted information (specifically, the product information, etc.). For example, "The user wants to know the specifications of the image 13c," "The specifications of the product corresponding to the image 13c are A," and "The user has previously purchased product B." The input generation information includes rules and other information necessary for generating an input sentence based on the accepted action. For example, the input generation information is a rule that, when the received action is "an action in which a user selects an image displayed on the application screen and drops the image onto the chatbot screen," the input generation information reads a predetermined sentence, "Please tell me information about {image}," from the content of the action, obtains metadata (e.g., part name A) of the dropped image, and combines it with the sentence to generate an input sentence, "Please tell me information about part name A." When the sending means 56 sends the input sentence generated by the input sentence generation means 54 to the language model server 60, the answer generation unit 63 of the language model server 60 generates an answer corresponding to the input sentence, and the sending means 64 sends the answer generated by the answer generation unit 63 to the chat server 50. Then, the sending means 56 of the chat server 50 sends an instruction to the user terminal 10 to display the answer information generated by the language model server 60. As a result, the answer generated by the language model server 60 is displayed on the chatbot screen of the website displayed on the display unit 13 of the user terminal 10, as shown in FIG. 13 . Specifically, an answer obtained by comparing the specifications of the product corresponding to the image 13c dropped onto the chatbot screen with the specifications of the user terminal 10 itself is displayed on the chatbot screen.Furthermore, the language model server 60 creates a message on the chatbot screen that prompts the user to select whether or not to display a page introducing multiple parts that will work with the specifications of the user terminal 10, and the created message is displayed on the chatbot screen (see reference numeral 13d in FIG. 13). This allows the user to select whether or not to display a page on the chatbot screen that introduces multiple parts that will work with the specifications of the user terminal 10. Here, if the user selects to display a page that introduces multiple parts that will work with the specifications of the user terminal 10, a page 13e that introduces multiple such parts is displayed on the chatbot screen, as shown in FIG. 14.
[0066] Furthermore, when a user drags a specific character (e.g., the character "core block") displayed on an application screen on the screen of the display unit 13 of the user terminal 10 as shown in Figures 12 and 14, such an action is detected by Javascript as first action information. Then, the action information and encrypted personal information, product information, etc. are transmitted from the user terminal 10 to the chat server 50. When the receiving means 52 of the chat server 50 receives the action information, etc. from the user terminal 10, the transmitting means 56 of the chat server 50 transmits to the language model server 60 a command to vectorize the action information, etc. received by the receiving means 52 (specifically, information that the user has dragged a specific character displayed on the application screen and the content of the dragged character). When the receiving means 62 of the language model server 60 receives the command, the answer generating unit 63 obtains a vector value by vectorizing the action information, etc. Thereafter, when the accepting means 52 of the chat server 50 accepts the vector value as vector information, the transmitting means 56 transmits to the first database server 70 a vector search request for searching the first database server 70 for a vector value that is approximate to the vector value generated by the language model server 60. When the accepting means 72 of the first database server 70 accepts the vector search request from the chat server 50, the accepting means 72 extracts, from the vector values stored in the storage unit 78, a vector value that is approximate to the vector value in the accepted vector search request, and the transmitting means 73 of the first database server 70 transmits a base answer corresponding to this approximate value to the chat server 50. Here, as the base answer corresponding to the approximate value, for example, an answer may be obtained that provides an explanation of the dragged character and acquires specification information of the user terminal 10 itself related to the dragged character.
[0067] In the chat server 50, the control unit 51 decrypts the encrypted information using an encryption key issued by the operator (administrator) providing the chatbot service. The input sentence generation means 54 generates an input sentence for the language model server 60 based on the acquired base answer (specifically, an explanation of the dragged character and acquisition of specification information related to the dragged character of the user terminal 10 itself), action information, and decrypted information (specifically, product information, etc.). When the transmission means 56 transmits the input sentence generated by the input sentence generation means 54 to the language model server 60, the answer generation unit 63 of the language model server 60 generates an answer corresponding to the input sentence, and the transmission means 64 transmits the answer generated by the answer generation unit 63 to the chat server 50. The transmission means 56 of the chat server 50 then transmits an instruction to the user terminal 10 to display the answer information generated by the language model server 60. As a result, the answer generated by the language model server 60 is displayed on the chatbot screen of the website displayed on the display unit 13 of the user terminal 10, as shown in FIG. 15 . Specifically, an explanation of the dragged character is provided, and an answer comparing the specification information of the dragged character on the user terminal 10 itself with the specification information of the dragged character on the product displayed on the application screen is displayed on the chatbot screen.
[0068] Next, an example of a second action detected by Javascript when the user inputs a question or the like on the chatbot screen after performing a predetermined action on the chatbot screen is described. FIG. 16 illustrates a display screen when a reservation site for making reservations for accommodations is displayed on the display unit 13 of the user terminal 10 as a website provided by the web server 30. The left area 13a of the display screen is an application screen that displays the names of available accommodations, descriptions of the accommodations, accommodation prices, addresses, maps, etc., while the right area 13b of the display screen shown in FIG. 16 is a chatbot screen for operating the chatbot, which is different from the application screen. As described above, this chatbot screen is not displayed if the user does not agree to use the chatbot on the website displayed on the display unit 13 of the user terminal 10. When the chatbot screen is displayed on the website, the user can input a question into the chatbot screen using the operation unit 14. When a user logs in to their personal page on a website such as that shown in Figure 16, the user's personal information, accommodation service information, etc. are encrypted by the web server 30, and the encrypted information is sent from the web server 30 to the user terminal 10, where it is temporarily stored in the memory unit 18 of the user terminal 10.
[0069] When a user displays a page for a specific accommodation facility on the screen of the display unit 13 of the user terminal 10 as shown in FIG. 16 and inputs the sentence "I would like to make a reservation" on the chatbot screen, this action is detected by Javascript as second action information. The action information, encrypted personal information, service information, etc. are then transmitted from the user terminal 10 to the chat server 50. The encrypted personal information includes the user's past usage history of accommodation facilities. The encrypted service information includes information such as the content of the accommodation facility description displayed on the application screen, accommodation fees, and address. When the reception means 52 of the chat server 50 receives the action information, etc. from the user terminal 10, the transmission means 56 of the chat server 50 transmits to the language model server 60 a command to vectorize the action information, etc. received by the reception means 52 (the user's action of displaying a page for a specific accommodation facility and inputting the sentence "I would like to make a reservation" on the chatbot screen). When the receiving means 62 of the language model server 60 receives a command, the answer generating unit 63 obtains a vector value by vectorizing the action information, etc. Thereafter, when the receiving means 52 of the chat server 50 receives the vector value as vector information, the transmitting means 56 transmits to the first database server 70 a vector search request for searching the first database server 70 for a vector value that approximates the vector value generated by the language model server 60. When the receiving means 72 of the first database server 70 receives the vector search request from the chat server 50, the receiving means 72 extracts, from the vector values stored in the storage unit 78, a vector value that approximates the vector value in the received vector search request, and the transmitting means 73 of the first database server 70 transmits a base answer corresponding to this approximate value to the chat server 50. Here, the base answer corresponding to the approximate value may be, for example, an answer that prompts the user to select whether they would like to reserve an accommodation displayed on the application screen, whether they would like to reserve an accommodation suggested based on the user's past accommodation usage history, or neither.
[0070] In addition, in the chat server 50, the control unit 51 decrypts encrypted information using an encryption key issued by the operator (administrator) providing the chatbot service. The input sentence generation means 54 generates an input sentence to the language model server 60 based on the acquired base answer (an answer prompting the user to select whether they would like to reserve an accommodation displayed on the application screen, an accommodation suggested based on the user's past accommodation usage history, or neither), action information, and decrypted information (specifically, the user's past accommodation usage history, service information of the accommodation displayed on the application screen, etc.). When the transmission means 56 transmits the input sentence generated by the input sentence generation means 54 to the language model server 60, the response generation unit 63 of the language model server 60 generates a response corresponding to the input sentence, and the transmission means 64 transmits the response generated by the response generation unit 63 to the chat server 50. The transmission means 56 of the chat server 50 then transmits an instruction to the user terminal 10 to display the response information generated by the language model server 60. As a result, as shown in FIG. 16 , the answer generated by the language model server 60 is displayed on the chatbot screen of the website displayed on the display unit 13 of the user terminal 10. Specifically, an option 13f is displayed on the chatbot screen, prompting the user to select whether they want to reserve the accommodation displayed on the application screen, whether they want to reserve an accommodation suggested based on the user's past accommodation usage history, or neither. This allows the user to select on the chatbot screen whether they want to reserve the accommodation displayed on the application screen, whether they want to reserve an accommodation suggested based on the user's past accommodation usage history, or neither. Here, if the user selects to reserve an accommodation suggested based on the user's past accommodation usage history, a message for making a reservation for this accommodation (e.g., a message to confirm the reservation date) is displayed on the chatbot screen based on the decoded information, as shown in FIG. 16 .
[0071] 16, on the screen of the display unit 13 of the user terminal 10, the input sentence generation means 54 generates a plurality of input sentences corresponding to the action (for example, "I would like to reserve Kirana Garden," "I would like to reserve Bayside Garden," etc.), and displays the generated plurality of input sentences on the user terminal 10, thereby enabling the user to select one input sentence from the plurality of input sentences or to create a new input sentence on the user terminal 10. This allows the user to obtain a more appropriate answer from the chatbot by selecting one input sentence from the plurality of input sentences or creating a new input sentence.
[0072] [Summary of Configuration and Operation of the Present Embodiment] According to the program, computer (specifically, chat server 50), system 1 combining web server 30 and chat server 50, and information processing method of the present embodiment configured as described above, the program causes control unit 51 of chat server 50 to function as reception means 52, input statement generation means 54, response acquisition means 55, and transmission means 56. The reception means 52 receives action information related to an action performed by a user on an application screen displayed on user terminal 10, or a chatbot screen, which is a screen for operating the chatbot and different from the application screen. Based on the action information received by the reception means 52, input statement generation means 54 generates an input statement corresponding to the action, and response acquisition means 55 acquires information on a response to the input statement generated by input statement generation means 54. The transmission means 56 transmits display information for displaying the response acquired by response acquisition means 55 on the chatbot screen to the user terminal 10. According to such a program, chat server 50, system 1, and information processing method, input sentences can be created based on information corresponding to user actions, thereby increasing the flexibility of responses. More specifically, conventional chatbots can only provide responses based on a database that stores combinations of user input sentences and responses, which has the problem of low response flexibility. However, according to the program, chat server 50, system 1, and information processing method of the present embodiment, action information related to actions performed by the user on the application screen can be referenced when creating an input sentence, thereby increasing the flexibility of responses compared to when responses are based on a database that stores combinations of user input sentences and responses.
[0073] Furthermore, in the program, computer (specifically, chat server 50), system 1, and information processing method of the present embodiment, as described above, the answer obtaining means 55 transmits the input sentence generated by the input sentence generation means 54 to the language model server 60 and obtains answer information by receiving from the language model server 60 an answer to the input sentence calculated by the language model server 60. This makes it possible to improve the accuracy of answers to input sentences by using the language model server 60. Note that in another aspect of the present embodiment, the language model server 60 as a large-scale language model may be provided inside the chat server 50, rather than being provided separately from the chat server 50. Furthermore, the answer obtaining means 55 is not limited to transmitting the input sentence generated by the input sentence generation means 54 to the language model server 60 and receiving from the language model server 60 an answer to the input sentence calculated by the language model server 60 to obtain answer information. As long as an input sentence is created based on information corresponding to a user's action, an answer may be obtained from the input sentence by a method other than inputting the input sentence to the language model server 60.
[0074] Furthermore, in the program, computer (specifically, chat server 50), system 1, and information processing method of this embodiment, as described above, when consent to the use of a chatbot is input into the user terminal 10, an action detection program that detects actions taken by the user on the user terminal 10 is executed on the user terminal 10, and action information regarding actions taken by the user on the application screen, detected by the action detection program, is transmitted from the user terminal 10 to the chat server 50. This makes it possible for the action detection program to reliably detect actions taken by the user on the application screen. Note that in this embodiment, the action detection program is not limited to Javascript, and various other programs may be used.
[0075] Furthermore, in the program, computer (specifically, chat server 50), system 1, and information processing method of this embodiment, as described above, the action is an action (first action) in which a user selects an image displayed on an application screen and drops it onto the chatbot screen. When the input sentence generation means 54 detects the dropping action, it acquires a base answer for a pre-input sentence generated based on the dropped image and action information, and generates an input sentence based on the base answer and action information. This allows the input sentence generation means 54 to generate an input sentence with high accuracy. Note that, as described above, "based on the dropped image" means based on metadata (image title and description) assigned to the image. Furthermore, the input sentence is transmitted from the chat server 50 to the language model server 60, and the pre-input sentence is transmitted from the chat server 50 to the first database server 70 to create the input sentence.
[0076] Furthermore, at this time, when the input sentence generation means 54 detects a dropping action, it transmits a pre-input sentence generated based on the dropped image and action information to the first database server 70, and acquires information on the base answer by receiving from the first database server 70 a base answer to the pre-input sentence extracted by the first database server 70. In this case, by using the base answer stored in the first database server 70, the input sentence generation means 54 can generate an input sentence with even greater accuracy.
[0077] Furthermore, the sending means 56 may be configured to send to the user terminal 10 an instruction requesting approval of the base answer acquired by the input sentence generation means 54, and the input sentence generation means 54 may be configured to generate an input sentence based on the approved base answer and action information when receiving information regarding the approval of the base answer from the user terminal 10. In this case, by obtaining the user's approval of the base answer, the input sentence generation means 54 can generate an input sentence with even greater accuracy.
[0078] Furthermore, as described above, the input sentence generation means 54 may further acquire specification information of the user terminal 10 and generate an input sentence by referring to the specification information in addition to the base answer and action information. In this case, by referring to the specification information of the user terminal 10, the input sentence generation means 54 can generate an input sentence with even greater accuracy.
[0079] Furthermore, in the program, computer (specifically, chat server 50), system 1, and information processing method of this embodiment, as described above, the action is an action (second action) of the user accessing the second database server 40 that stores the information displayed on the application screen, and the input sentence generation means 54 acquires action information related to the action to be accessed based on the content entered by the user on the chatbot screen and information stored in the accessed second database server 40, and generates an input sentence based on the content entered by the user, the action information, and the information stored in the accessed second database server 40. In this case, by referring to the information stored in the second database server 40, the input sentence generation means 54 can generate an input sentence with greater accuracy.
[0080] In this case, the type of second database server 40 to be accessed may be determined from a plurality of types of second database servers 40 according to the type of action. In this case, by obtaining information from the second database server 40 corresponding to the user's action on the user terminal 10, the obtained information can be more appropriate.
[0081] Furthermore, in the program, computer (specifically, chat server 50), system 1, and information processing method of this embodiment, as described above, the program causes control unit 51 to further function as user information acquisition means 53, which acquires information about the user from user terminal 10, and input sentence generation means 54 may generate an input sentence according to the action by referring to the information about the user acquired by user information acquisition means 53 in addition to the action information accepted by acceptance means 52. In this case, by referring to the information about the user acquired by user information acquisition means 53, input sentence generation means 54 can generate an input sentence with higher accuracy.
[0082] Furthermore, in the program, computer (specifically, chat server 50), system 1, and information processing method of this embodiment, as described above, the input sentence generation means 54 generates a plurality of input sentences according to the action, and displays the generated plurality of input sentences on the user terminal 10, thereby enabling the user to select one input sentence from the plurality of input sentences or to create a new input sentence on the user terminal 10. This allows the user to obtain a more appropriate answer from the chatbot by selecting one input sentence from the plurality of input sentences or creating a new input sentence.
[0083] Furthermore, in the program, computer (specifically, chat server 50), system 1, and information processing method of this embodiment, when an input indicating consent to use of the chatbot is made on the user terminal 10, an action detection program (e.g., Javascript) that detects actions other than input to the chatbot made by the user on the user terminal 10 is executed on the user terminal 10, and action information regarding actions other than input to the chatbot made by the user, detected by the action detection program, is sent from the user terminal 10 to the chat server 50. With such a program, chat server 50, system 1, and information processing method, too, the flexibility of responses can be increased by creating input sentences based on information corresponding to the user's actions.
[0084] [Other Aspects of the Present Embodiment] The program, computer (specifically, chat server 50), system 1, and information processing method according to the present embodiment are not limited to the aspects described above, and various modifications can be made.
[0085] For example, in the above description, the answer obtaining means 55 transmits the input sentence generated by the input sentence generation means 54 to the language model server 60 and receives from the language model server 60 the answer to the input sentence calculated by the language model server 60, thereby obtaining answer information. However, this embodiment is not limited to this configuration. As another configuration, a third database server 90, in which data on the set input sentence and the set answer are stored in association with each other, may be connected to the chat server 50 so as to be able to communicate with the chat server 50. The answer obtaining means 55 may transmit the input sentence generated by the input sentence generation means 54 to the third database server 90, and obtain answer information by receiving from the third database server 90 the answer corresponding to the input sentence extracted based on the data on the set input sentence and the set answer stored in the third database server 90. In another aspect, instead of sending the answer sent from the third database server 90 to the chat server 50 as is to the user terminal 10 and displaying it on the chatbot, the answer sent from the third database server 90 to the chat server 50 and the question by the user sent from the user terminal 10 to the chat server 50 may be sent to the language model server 60. In this case, the language model server 60 generates an answer based on the question by the user and the answer obtained from the third database server 90, and the answer generated by the language model server 60 is sent from the chat server 50 to the user terminal 10 and displayed on the chatbot.
[0086] In yet another embodiment, data on the set input sentence and the set answer may be stored in association with each other in the memory unit 58 of the chat server 50, and the answer acquisition means 55 may acquire information on the answer corresponding to the input sentence based on the data on the set input sentence and the set answer stored in the memory unit 58.
[0087] Furthermore, the method for generating an input sentence by the input sentence generation means 54 is not limited to the method using the vector information stored in the storage unit 78 of the first database server 70 as described above. Various other methods can be used as the method for generating an input sentence by the input sentence generation means 54 as long as they are based on the action information accepted by the acceptance means 52.
[0088] Furthermore, in the program, computer (specifically, chat server 50), system 1, and information processing method of the present embodiment, the program may further cause control unit 51 to function as learning means 57. For example, when an administrator logs in to an administration page rather than a personal page using an administrator account on a website displayed on a display unit (not shown) of administrator terminal 80, such as a personal computer owned by the administrator, the learning means 57 can cause the chatbot to learn responses to input sentences. Specifically, the learning means 57 of chat server 50 transmits an instruction to administrator terminal 80 inquiring about the content of an input sentence corresponding to the content of an action in the action information received by receiving means 52, and performs learning by associating the content of the input sentence received from administrator terminal 80 with the content of the action. Then, when receiving means 52 receives action information, input sentence generation means 54 generates an input sentence corresponding to the action based on the content of the learning performed by learning means 57. Even in this case, by learning by associating the content of the input sentence received from the administrator terminal 80 with the content of the action, the input sentence generation means 54 will be able to generate highly accurate input sentences, and the answer acquisition means 55 will be able to obtain highly accurate answers.
[0089] In yet another embodiment, instead of or in addition to an action performed in a browser or the like displayed on the display unit 13 of the user terminal 10, an actual action (e.g., turning one's head, closing one's eyes, speaking, etc.) of the user operating the user terminal 10 may be detected, and information about the detected action may be transmitted from the user terminal 10 to the chat server 50, causing the input sentence generation means 54 to generate an input sentence corresponding to the action based on the action information, and the answer acquisition means 55 to acquire information about an answer to the input sentence generated by the input sentence generation means 54. Such an actual action of the user operating the user terminal 10 is detected by the imaging unit 15, microphone 16, etc. of the user terminal 10. In this case, the user can obtain an answer from the chatbot based on the actual action of the user operating the user terminal 10 in addition to an action performed in a browser or the like displayed on the display unit 13 of the user terminal 10, thereby further improving the convenience of the chatbot.
[0090] DESCRIPTION OF SYMBOLS 1 System 10 User terminal 11 Control unit 11a Bus 13 Display unit 14 Operation unit 15 Imaging unit 16 Microphone 18 Memory unit 19 Communication unit 20 Tag management server 21 Control unit 21a Bus 22 Receiving means 23 Transmission means 28 Memory unit 29 Communication unit 30 Web server 31 Control unit 31a Bus 32 Receiving means 33 Transmission means 38 Memory unit 39 Communication unit 40 Second database server 41 Control unit 41a Bus 42 Receiving means 43 Transmission means 48 Memory unit 49 Communication unit 50 Chat server 51 Control unit 51a Bus 52 Receiving means 53 User information acquisition means 54 Input sentence generation means 55 Answer acquisition means 56 Transmission means 57 Learning means 58 Memory unit 59 Communication unit 60 Language model server 61 Control unit 61a Bus 62 Receiving means 63 Answer generation unit 64 Transmission means 68 Storage unit 69 Communication unit 70 First database server 71 Control unit 71a Bus 72 Receiving means 73 Transmission means 78 Storage unit 79 Communication unit 80 Administrator terminal 90 Third database server
Claims
1. A computer comprising a memory and a processor, wherein the processor executes a computer program stored in the memory to: receive action information regarding an action performed by a user on an application screen displayed on a user terminal, or on a chatbot screen that is a screen for operating a chatbot and is different from the application screen; generate an input sentence corresponding to the action based on the received action information; obtain information on a response to the generated input sentence; and transmit display information for displaying the obtained response on the chatbot screen to the user terminal.
2. The computer of claim 1, wherein when obtaining information about an answer to the generated input sentence, the computer transmits the generated input sentence to a large-scale language model, and obtains the information about the answer by receiving from the large-scale language model the answer to the input sentence calculated by the large-scale language model.
3. A computer as described in claim 1, wherein an action detection program is executed on the user terminal when consent to the use of a chatbot is input into the user terminal, the action detection program detecting the action performed by the user on the user terminal, and the action information regarding the action performed by the user on the application screen detected by the action detection program is transmitted from the user terminal to the processor.
4. The computer of claim 1, wherein the action is an action in which the user selects an image displayed on the application screen and drops it onto the chatbot screen, and when generating an input sentence corresponding to the action, if the dropping action is detected, a base answer to a pre-input sentence generated based on the dropped image and the action information is obtained, and the input sentence is generated based on the base answer and the action information.
5. A computer as described in claim 4, wherein when generating an input sentence corresponding to the action, if the dropping action is detected, the pre-input sentence generated based on the dropped image and the action information is sent to a first database server, and information on the base answer is obtained by receiving from the first database server the base answer to the pre-input sentence extracted by the first database server.
6. A computer as described in claim 4, wherein when sending display information for displaying the obtained answer on the chatbot screen to the user terminal, an instruction requesting approval of the obtained base answer is sent to the user terminal, and when generating an input sentence corresponding to the action, if information regarding approval of the base answer is received from the user terminal, the input sentence is generated based on the approved base answer and the action information.
7. The computer of claim 4, wherein when generating an input sentence corresponding to the action, the computer further acquires specification information of the user terminal, and generates the input sentence by referring to the specification information in addition to the base answer and the action information.
8. The computer of claim 1, wherein the action is an action of the user accessing a second database server that stores information displayed on the application screen, and when generating an input sentence corresponding to the action, the computer obtains the action information regarding the action to be accessed and information stored in the second database server to be accessed based on the content entered by the user on the chatbot screen, and generates the input sentence based on the content entered by the user, the action information, and the information stored in the second database server to be accessed.
9. The computer according to claim 8, wherein the type of said second database server to be accessed from among a plurality of types of said second database servers is determined according to the type of said action.
10. The computer of claim 1, wherein the processor acquires information about the user from the user terminal by executing a computer program stored in the memory, and when generating an input sentence corresponding to the action, the processor references the acquired information about the user in addition to the received action information to generate the input sentence corresponding to the action.
11. A computer as described in claim 1, wherein when obtaining information on an answer to the generated input sentence, the generated input sentence is sent to a third database server in which data on the set input sentence and the set answer are stored in association with each other, and the answer corresponding to the input sentence extracted based on the data on the set input sentence and the set answer stored in the third database server is received from the third database server, thereby obtaining information on the answer.
12. A computer as described in claim 1, wherein the memory stores data on a set input sentence and a set answer in association with each other, and when obtaining information on an answer to the generated input sentence, the information on the answer corresponding to the input sentence is obtained based on the data on the set input sentence and the set answer stored in the memory.
13. The computer of claim 1, wherein the processor executes a computer program stored in the memory to send an instruction to an administrator terminal inquiring about the content of the input sentence in response to the content of the action in the received action information, and learns by associating the content of the input sentence received from the administrator terminal with the content of the action, and when generating an input sentence corresponding to the action, upon receiving the action information, generates the input sentence corresponding to the action based on the content of the learning that has been performed.
14. The computer of claim 1, wherein when generating an input sentence corresponding to the action, a plurality of input sentences corresponding to the action are generated, and the generated plurality of input sentences are displayed on the user terminal, thereby allowing the user to select one of the plurality of input sentences or to create a new input sentence on the user terminal.
15. A computer comprising a memory and a processor, wherein the processor executes a computer program stored in the memory to: receive action information regarding an action performed by a user on a user terminal; generate an input sentence corresponding to the action based on the received action information; obtain information on a response to the generated input sentence; send display information to the user terminal for displaying the obtained response on a chatbot screen; when an input indicating consent to use of a chatbot is made on the user terminal, an action detection program is executed on the user terminal to detect actions performed by the user on the user terminal other than input to the chatbot; and the action information regarding the actions performed by the user other than input to the chatbot, detected by the action detection program, is sent from the user terminal to the processor.
16. An information processing method carried out by a computer having a memory in which a computer program is stored and a processor that executes the computer program stored in the memory, the information processing method comprising: receiving action information regarding an action performed by a user on an application screen displayed on a user terminal, or a chatbot screen that is a screen for operating a chatbot and is different from the application screen; generating an input sentence corresponding to the action based on the received action information; obtaining information on a response to the generated input sentence; and transmitting display information for displaying the obtained response on the chatbot screen to the user terminal.
17. An information processing method performed by a computer having a memory in which a computer program is stored and a processor that executes the computer program stored in the memory, comprising the steps of: receiving action information regarding an action performed by a user on a user terminal; generating an input sentence corresponding to the action based on the received action information; obtaining information on a response to the generated input sentence; transmitting display information for displaying the obtained response on a chatbot screen to the user terminal; when an input indicating consent to use of a chatbot is made on the user terminal, an action detection program that detects actions other than input to the chatbot performed by the user on the user terminal is executed on the user terminal; and the action information regarding the actions other than input to the chatbot performed by the user that is detected by the action detection program is transmitted from the user terminal to the processor.
Citation Information
Patent Citations
Method for providing commodity information and electronic equipment
CN116977013A
PROGRAM, COMPUTER AND INFORMATION PROCESSING METHOD
JP7418766B1
Customer service assistance system and customer service assistance method
WO2019186678A1