Information processing device, terminal device, information processing system, and information processing method

The information processing system addresses the limitation of conventional systems by incorporating linguistic and image inputs to generate conversational responses and monitor task progress, enabling efficient task management through conversational interactions.

WO2026053788A1PCT designated stage Publication Date: 2026-03-12SHARP KK
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Conventional information processing systems are unable to respond to linguistic inputs from users, limiting their functionality to determining work processes and not enabling conversational interactions.

Method used

An information processing system that includes a language input unit, an image input unit, a response generation unit, and a response control unit, which analyze linguistic and image inputs to generate responses based on a task list, allowing for conversational interactions and task progress monitoring.

Benefits of technology

Enables conversational responses to verbal inputs, providing task progress monitoring and management beyond mere work process determination, enhancing user interaction and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025029817_12032026_PF_FP_ABST
    Figure JP2025029817_12032026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure enables correspondence not limited to determination of a work process, by performing conversational response to a language input from a user. An information processing system (100) comprises: a language input unit that receives a language input from a user; an image input unit that acquires a captured image of the user executing a task including a plurality of processes; and a response control unit that receives a response related to the progress of the task by the user, the response being generated on the basis of analysis results of the language input and the image, as well as a task list, and performs output control.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, terminal device, information processing system, and information processing method

[0001] The present disclosure relates to an information processing device, a terminal device, an information processing system, and an information processing method that respond to a user's language input.

[0002] In recent years, various technologies for managing user tasks have been proposed. For example, an information processing device described in Patent Literature 1 acquires an image group consisting of multiple images of a subject captured consecutively by an imaging unit, and detects multiple analysis points corresponding to multiple body parts of the subject from each image in the acquired image group to determine the task of the subject. The information processing device then includes a processing unit that determines which task the subject was performing out of multiple tasks included in a predetermined task process based on time-series changes in the image group of the detected analysis points.

[0003] Japanese Patent Application Publication No. 2024-031497

[0004] However, such conventional techniques cannot respond to linguistic input from the user because they detect a "specific analysis point" and determine the current work process.

[0005] An object of one aspect of the present disclosure is to enable a system to respond in a conversational manner to a linguistic input from a user, thereby enabling a response that is not limited to determining a work process.

[0006] In order to solve the above problem, an information processing device according to one aspect of the present disclosure includes a language input unit that accepts language input from a user, an image input unit that acquires an image of the user performing a task including multiple steps, and a response control unit that receives a response regarding the progress of the task by the user, which response is generated based on the results of analyzing at least one of the language input and the image and a task list indicating the multiple steps, and controls the output of the response.

[0007] In order to solve the above-mentioned problems, an information processing system according to one aspect of the present disclosure includes a language input unit that accepts language input from a user, an image input unit that acquires an image of the user performing a task including multiple steps, a response generation unit that generates a response regarding the progress of the task by the user based on the results of analyzing at least one of the language input and the image and a task list indicating the multiple steps, and a response control unit that receives the response generated by the response generation unit and controls the output of the response.

[0008] In order to solve the above-mentioned problem, an information processing method according to one aspect of the present disclosure is executed by a computer and includes the steps of accepting linguistic input from a user, acquiring an image of the user performing a task including multiple steps, receiving a response regarding the user's progress in the task, the response being generated based on the results of analyzing at least one of the linguistic input and the image and a task list indicating the multiple steps, and controlling the output of the response.

[0009] According to one aspect of the present disclosure, the system provides a conversational response to a verbal input from a user, thereby enabling responses that are not limited to determining the work process.

[0010] FIG. 1 is a block diagram illustrating a configuration of an information processing system according to a first embodiment. FIG. 2 is a flowchart illustrating an example of a monitoring process in the information processing system according to the first embodiment. FIG. 3 is a flowchart illustrating an example of a process for each predetermined interval in the information processing system according to the first embodiment. FIG. 4 is a flowchart illustrating an example of a process for a language input in the information processing system according to the first embodiment. FIG. 5 is a flowchart illustrating an example of a process for acquiring vital information in the information processing system. FIG. 6 is a block diagram illustrating a configuration of an information processing system according to a second embodiment. FIG. 7 is a perspective view illustrating a terminal device according to the second embodiment.

[0011] First Embodiment Hereinafter, one embodiment of the present disclosure will be described in detail.

[0012] FIG. 1 is a block diagram illustrating an example of the configuration of an information processing system 100 according to a first embodiment. The information processing system 100 is an AI (Artificial Intelligence) conversation system. As shown in FIG. 1 , the information processing system 100 includes a microphone 11, a camera 12, a fingerprint sensor 14 (sensor), an output device 15, and a storage unit 16. The information processing system 100 also includes a voice input unit 21 (language input unit), an image input unit 22, an authentication unit 25, a context analysis unit 26, an LLM (Large Language Model) determination unit 27, a simple response generation unit 28, a response control unit 29, a response output unit 30, a conversation history management unit 31, an experience information generation unit 32, an imaging control unit 33, and a task list management unit 34. The information processing system 100 also includes a first LLM 51 and a second LLM 52.

[0013] (Input / Output Devices) The microphone 11 is a voice input device that accepts voice input from the user to the information processing system 100. The camera 12 is an imaging device that captures images in accordance with user operations or predetermined setting conditions. There are no particular limitations on the type, placement, number, imaging range, and imaging conditions of the cameras 12. The fingerprint sensor 14 is a fingerprint detection device that detects the user's fingerprint.

[0014] The output device 15 is a device that transmits information to the user. Information is transmitted to the user in response to a voice input by the user. The output device 15 is, for example, an image display device and a speaker. The image display device is a device that displays responses to the user's conversation, displays images captured by the camera 12, and notifies the user by display. The speaker is an audio output device that outputs audio responses to the user's conversation and notifies the user by voice. The output device 15 may also be equipped with a vibration device for emphasizing notifications to the user, a transmission device for emergency contact, etc.

[0015] (Storage Unit, Audio Input Unit, and Image Input Unit) The storage unit 16 records information necessary for controlling the information processing system 100. The information processing system 100 may be communicably connected to an external storage device serving as the storage unit 16. That is, the storage unit 16 may be provided outside the information processing system 100.

[0016] The voice input unit 21 functions as a language input unit that accepts voice input (language input) from the user via the microphone 11. The voice input unit 21 converts the input voice into text data. Alternatively, the user may input language by text input. When inputting language by text input, a text input unit that accepts text as language input via a touch panel or a text input device such as a smartphone that is communicably connected to the context analysis unit 26 may be provided. In the following description, it is assumed that the user inputs language by voice.

[0017] The image input unit 22 acquires images from the camera 12. Unless otherwise specified in this disclosure, images include still images and video images. The images acquired by the image input unit 22 include images of a user performing a task including multiple steps. The image input unit 22 may convert the acquired images into embedded data (text data). The images may be converted into text data and incorporated into a conversation response request or a task list creation request sent to the first LLM 51 or the second LLM 52. The image input unit 22 may convert the images into text data indicating a natural language understandable by the first LLM 51 or the second LLM 52.

[0018] There are several known methods for converting image data into text data. In the field of AI, the following known methods can be appropriately selected or combined for use.

[0019] For example, the conversion method may be Base64 encoding. Base64 encoding is a method for converting binary data into ASCII text and is widely used when handling binary data such as image files in text format. Base64 encoding is often used when embedding images as data URIs in applications, and can also be used in the information processing system 100.

[0020] The conversion method may also be hexadecimal encoding. Hexadecimal encoding is a method of converting binary data into a hexadecimal string. Hexadecimal encoding is generally used more often when visualizing binary data for debugging or data analysis than for images, and is rarely used in the information processing system 100. However, if the image provided is a CG image or, in particular, a mechanical design image, it is easy to capture the characteristics of the image and may be used depending on the application.

[0021] Alternatively, the conversion method may be URL encoding. URL encoding is often used in applications such as embedding images directly into HTML, and is widely used because it is easy to redisplay the converted image. In the information processing system 100, it is preferable that the LLM be an easy-to-use format, so it is not particularly selected as a conversion method for the chat system. However, URL encoding may be used in applications where it is important to redisplay the input image on an image display device.

[0022] Although it is not a direct conversion method, there is also a method called JSON encoding. This method encodes binary data using Base64 and stores the result in JSON format. Currently available open LLMs and closed LLMs that support many image modes support this method, so it can be used effectively in the information processing system 100.

[0023] In any case, the conversion method can be selected based on the learning method of the LLM to be used, as long as it is in a format that can be interpreted by the LLM that ultimately generates the response. Various encoding methods are well known, and conversion can be performed at the input stage of various LLMs. Therefore, the method can be selected taking into account resources, response time, and other qualities. The important thing is that these conversion methods convert image data into text data, thereby providing a method for the context analysis unit 26 and storage unit 16 to handle image input in the same way as normal conversational input.

[0024] Furthermore, the image input unit 22 outputs the input image from the camera 12 to the context analysis unit 26. The timing of image acquisition and output by the image input unit 22 may be set at a predetermined time interval (interval) by, for example, an interval timer.

[0025] (Authentication Unit) The authentication unit 25 performs personal authentication of the user based on information from the fingerprint sensor 14. The authentication unit 25 functions as a sensor information input unit that acquires personal authentication information as sensor information from the fingerprint sensor 14. The authentication unit 25 compares the fingerprint of the user registered in advance with the fingerprint acquired from the fingerprint sensor 14, and determines whether the acquired fingerprint is that of the registered user. If the acquired fingerprint is that of the registered user, the authentication unit 25 verifies the user's personal ID or the organization to which the user belongs, and grants access rights to necessary information.

[0026] The authentication unit 25 outputs the determination result to the context analysis unit 26. The context analysis unit 26 may change the personal information used for context analysis depending on whether the user is a registered user or a guest user.

[0027] 1 illustrates a fingerprint sensor 14 as the personal authentication sensor, but other personal authentication sensors such as an iris authentication sensor, a voiceprint authentication sensor, or a vein sensor may also be used. Personal authentication may also be performed using an image of the user's face or voice. For example, the authentication unit 25 may perform personal authentication of the user through iris authentication or voiceprint authentication. Of course, the authentication unit 25 may also perform authentication using a combination of multiple sensors.

[0028] (Context Analysis Unit) The context analysis unit 26 performs a context analysis of a general conversation. At this time, the context analysis unit 26 may perform the context analysis using past conversation history and / or personal information of the user. The context analysis by the context analysis unit 26 may be, for example, an analysis using a small-scale language model that extracts keywords from the user's language input based on the past history and organizes correlations between the keywords. Alternatively, the context analysis may include a process of determining attributes of the language input by analyzing keywords extracted from the user's language input using information in a database.

[0029] The attributes of the language input are simple tags corresponding to the content of the language input, such as a task monitoring request, task classification, task result evaluation, question, greeting, impression, analysis request, request, or knowledge field, etc. The LLM determination unit 27 can select a response generation unit based on these tags.

[0030] Task classification is used to clarify restrictions on the purpose or conditions for accomplishing a task by categorizing and classifying the characteristics of the tasks that the information processing system 100 will handle. By referring to this classification, the LLM can supplement information such as the accuracy of each process, the time required to accomplish the task, whether or not it can be interrupted, the resources that can be allocated to the task, and the importance of completing the task (e.g., whether it is okay to give up depending on the situation), enabling more efficient and effective management.

[0031] There are various well-known management methods for classifying tasks, but since conversations are processed with the LLM, it is not necessary to use very strict definitions, and in one embodiment of the present invention, for example, the following classification may be used: Here, an explanation is given based on everyday life, but of course, the higher-level basic classifications and task constraints can be considered similarly for other purposes, such as business or ceremonies.

[0032] Tasks may be categorized as, for example, "basic tasks - daily tasks." Daily tasks may include, for example, preparing meals or shopping every day. Tasks may also be categorized as "basic tasks - regular tasks." Regular tasks may include, for example, organizing monthly household accounts, weekend cleaning, or organizing recorded programs. Tasks may also be categorized as "project tasks." Project tasks may include, for example, preparing food or venues for special parties, safety checks before going out, seasonal cleaning, and navigation for family trips. Tasks may also be categorized as "emergency tasks." Emergency tasks may include, for example, responding to water leaks, malfunctioning home appliances, and responding to injuries or accidents involving children. Tasks may also be categorized as "emergency responses." Emergency responses may include, for example, responding to unexpected visitors or unexpected family events. Tasks may also be categorized as "maintenance." Maintenance may include, for example, inspecting and servicing home appliances or vehicles.

[0033] By categorizing tasks and providing information in this way, the LLM can respond in line with the importance of the task and the user's feelings, such as enthusiasm or impatience. Of course, these do not need to be defined strictly, and the context analysis unit 26 can supplement information for responding in line with the user's intentions as tags, in conjunction with reference to the conversation history, etc.

[0034] The context analysis unit 26 functions as a task list creation instruction unit that outputs a task list creation request to the LLM determination unit 27 and instructs the language model to create a task list. In the present disclosure, a language model is a machine learning model such as a generative artificial intelligence (AI) that can analyze and generate language, and an example of such a language model is an LLM. In the information processing system 100, a first LLM 51 and a second LLM 52 are used as language models.

[0035] A task list is information indicating a task that includes multiple steps. A task list may also include information such as the task classification, task name, and a visual description of the task, in addition to the content of each step of the task. The visual description of the task may be, for example, an image of a cooking recipe or a work manual that was used to create the task list.

[0036] The context analysis unit 26 acquires task information relating to tasks performed by the user, and instructs the language model via the LLM determination unit 27 to create a task list based on the task information.

[0037] The context analysis unit 26 analyzes the language input and, when it determines that the user has requested task monitoring, generates a task list creation request. When the context analysis unit 26 outputs the generated task list creation request to the LLM determination unit 27, it may record information indicating that the task list has been created in the storage unit 16. Note that, as described below, the context analysis unit 26 may not generate a task list creation request when a suspended task is resumed or when the storage unit 16 records a similar task that has been performed in the past. In this case, the context analysis unit 26 may obtain and use a task list related to the suspended task or a previously performed task from the storage unit 16 via the conversation history management unit 31. When a task similar to the task for which the user has requested monitoring is recorded in the storage unit 16, the context analysis unit 26 may generate an improved task list creation request based on past evaluations or task correction history from past task executions.

[0038] When language input and / or image input are received after the task list is created, the context analysis unit 26 outputs an analysis request based on the task list, including the input contents, to the LLM determination unit 27. The analysis request may include, for example, basic instruction information for instructing a language model to perform image analysis of an image acquired by the context analysis unit 26. The basic instruction is a typical instruction that is frequently used when instructing a language model to perform image analysis of the image, and is an instruction that is predetermined at the system design stage. The basic instruction information is, for example, text information for instructing an analysis of which process in the task list the image corresponds to.

[0039] An example of the specific content of an analysis request including such basic instruction information is as follows: "Analyze the input image and determine which step in the task list the current situation corresponds to. If the current situation is appropriate, confirm with the user and explain the work content of the next step in the task list. If there is no step in the task list that corresponds to the current situation, estimate the current situation and confirm with the user. If there is language input from the user, compare it with the situation and, if it is normal, explain the work content of the next step. If it is not normal, appropriately modify the task list and confirm with the user." By including such instructions in the analysis request, the response generation unit selected by the LLM determination unit 27 can perform an analysis tailored to the purpose while reducing the number of questions to the user. Furthermore, as in the example of the analysis request described above, the context analysis unit 26 functions as a task list modification instruction unit that instructs the language model to modify the task list based on the progress of the tasks.

[0040] The context analysis unit 26 may store the content of such an analysis request itself, or may have the content stored in the task list management unit 34, which will be described later. From the viewpoint of centralized management of task lists, it is preferable that the context analysis unit 26 store the content of the analysis request in the task list management unit 34. In other words, the task list management unit 34 may function as a task list modification instruction unit.

[0041] The modification of the task list may include modification of each step according to the progress of the task, or modification of the prerequisites for the entire task. The prerequisites for the entire task are conditions that must be considered when performing the task, and may include conditions that define the goal of the task, such as the size of the deliverable of the task and the time required to complete it, conditions that define the attributes of the user who will perform the task (e.g., whether the user is a child or an adult, male or female), and conditions that define the environment for performing the task (e.g., time of day, location, available equipment, etc.).

[0042] As a specific prerequisite for the entire task, for example, if the task is cooking, the size of the product may be information indicating the number of servings, and the available equipment may be information on cooking utensils such as frying pans and knives. Also, it may be information on restrictions on ingredients or seasonings, which are imposed when children are present. In this way, it is preferable that the task list includes information on the prerequisites for the entire task.

[0043] The context analysis unit 26 or the task list management unit 34 serves as a task list correction instruction unit that, when information indicating a change in the prerequisites is included in the language input and / or image from the user, instructs the language model to correct the task list based on the information. More specifically, the task list correction instruction unit instructs the language model to correct the information on the prerequisites included in the task list based on the information indicating the change in the prerequisites.

[0044] Specifically, the task list correction instruction unit instructs the user to correct the task list in response to information such as a change in the number of people dining, information about dietary restrictions of members, or information about an alternative method when the intended cooking equipment cannot be found.

[0045] Furthermore, when the language input and / or image input reveal that the wrong seasoning or too much seasoning has been added, making it necessary to change the original task, the task list correction instruction unit may similarly instruct the language model to correct the premise information. Furthermore, if the correction is large-scale, the task list correction instruction unit may create a proposed correction to the task list and generate a request for confirmation from the user.

[0046] As an example of a large-scale correction, if the first LLM 51 determines that too much salt has been added, the task list correction instruction unit can create a task list correction proposal to bulk up the entire dish, add more vegetables, etc. Also, if the first LLM 51 determines that a user has mistakenly added too much vinegar while making boiled pork, the task list correction instruction unit may create a task list correction proposal to change the dish to a tomato-based stew.

[0047] The context analysis unit 26 may determine whether a historical task list has been created in the past for a task that is the same as or similar to the task classification obtained by the context analysis. The historical task list is a task list created in monitoring a past task that is different from the task currently being monitored or that is scheduled to be monitored.

[0048] If a historical task list exists, the context analysis unit 26 may include information about the historical task list in a task list creation request or modification request and output the request to the LLM determination unit 27 .

[0049] The context analysis unit 26 functions as a request generation unit that generates a conversation response request based on the user's language input and / or image input. The context analysis unit 26 outputs the generated conversation response request to the LLM determination unit 27.

[0050] In addition to the language input and the results of the context analysis, the conversation response request may include an acquired image, data in which the image is converted into embedded data, a conversation history, experience information, or a historical task list, or a combination of these. The context analysis unit 26 outputs the generated conversation response request to the LLM determination unit 27. That is, the context analysis unit 26 may analyze the user's request or task status based on the history information (conversation history, history task list, and experience information) and the current language input and / or image input, and transmit the analysis result to the LLM determination unit 27 as a conversation response request.

[0051] If the user's authentication information (personal information) is present, the context analysis unit 26 may perform context analysis using the user's own conversation history and, if necessary, group (family, business, etc.) conversation history from past history information. If the context does not depend on personal information, the previous history information can be referenced. The history information stores the conversation user ID or conversation group ID, the conversation text itself, and a history task list, and the context analysis unit 26 may perform context analysis using this information.

[0052] Furthermore, before generating a conversation response request, the context analysis unit 26 may determine whether to generate the request based on the user's authority. For example, if the question input by the user's voice is one for which the user does not have the authority to obtain an answer, the context analysis unit 26 may determine not to generate a conversation response request. Alternatively, the context analysis unit 26 may generate a request for generating a response such as "I cannot answer (with a reason depending on the situation)."

[0053] (LLM Determination Unit) The LLM determination unit 27 determines to which of the multiple response generation units the conversation response request should be sent based on the content of the conversation response request. The response generation units will be described later. The LLM determination unit 27 sends the conversation response request to at least one of the multiple response generation units. The LLM determination unit 27 may send the conversation response request to multiple response generation units. The context analysis unit 26 and the LLM determination unit 27 may be implemented as a single block.

[0054] (Response Generation Unit and Imaging Control Unit) The first LLM 51, the second LLM 52, and the simple response generation unit 28 function as response generation units that generate responses to a user's language input and / or image input. The LLM determination unit 27 determines to which of the first LLM 51, the second LLM 52, and the simple response generation unit 28 a conversational response request should be sent, based on the request content such as a conversational response request. Furthermore, the LLM determination unit 27 may transmit a conversational response request that includes confidential information that should not be sent to an external server to the simple response generation unit 28, without sending it to the external server, for example.

[0055] Furthermore, for example, if the first LLM 51 is located on an internal server, the LLM determination unit 27 may select the first LLM 51 as the destination of a conversation response request containing confidential information that should not be sent to an external server. Each LLM may be tagged with an attribute indicating its location (affiliated organization) and / or its handling authority for confidential information. The LLM determination unit 27 may refer to the attribute tag of each response generator in addition to the content of the response request to select a response generator suitable for generating a response.

[0056] In this disclosure, when there is no need to limit the type of response generator, the LLM and the simple response generator are collectively referred to simply as the response generator. The number of LLMs is not limited to two and may be three or more. The LLM determination unit 27 can select the optimal LLM by switching between multiple LLMs as needed.

[0057] The simple response generator 28 may be a response generator that generates simple responses and is provided in a terminal that includes the LLM determination unit 27. Meanwhile, the first LLM 51 and the second LLM 52 may be response generators that generate complex responses and are provided in a server or the like external to the terminal that includes the LLM determination unit 27. Furthermore, the first LLM 51 and the second LLM 52 may be LLMs that differ in the generation of responses to conversational response requests. Specifically, the first LLM 51 and the second LLM 52 may be language models that differ in the type and amount of trained data, the maximum input size that can be processed, the response time, and / or the response accuracy.

[0058] In the present disclosure, at least one of the LLMs functioning as a response generator (described as the first LLM 51 in this embodiment) is a language model corresponding to task monitoring. The attribute of the first LLM 51 is set to "task correspondence," and the LLM determination unit 27 selects the first LLM 51 in response to a request related to task monitoring.

[0059] The first LLM 51 analyzes the image captured by the camera 12 acquired by the image input unit 22. The first LLM 51 may perform object recognition in the image and cause the imaging control unit 33 to control the camera 12 to capture an image of the target. The objects that can be recognized are those that have been learned in advance to suit various use cases, but it is preferable that information about the corresponding objects is shared with the context analysis unit 26 and an appropriate search target is selected.

[0060] The first LLM 51 may detect an imaging target specified by the context analysis unit 26 in the image, generate information indicating the position of the imaging target in the image, and modify the task list.

[0061] Here, the relationship between the output of the first LLM 51 and the imaging control unit 33 will be explained in more detail. When the first LLM 51 initially creates a task list, the task list includes not only a description of each step but also imaging instructions for achieving desirable imaging in each step. The task list includes imaging instructions for each step, and the task list management unit 34 outputs imaging instructions for the current or next step to the imaging control unit 33 according to the current task status. This eliminates the need for the first LLM 51 to create desirable imaging instructions for each step every time, and the imaging control unit 33 does not need to wait for the generation of new imaging instructions each time. The imaging instructions generated by the first LLM 51 may be instructions that reflect the difference from the immediately preceding status of each step, or may be direct instructions based on a prior understanding of the work environment. Furthermore, if a situation arises in which the prerequisite information included in the task list needs to be changed due to linguistic input from the user or image input information from the camera 12, the first LLM 51 may regenerate imaging instructions in accordance with the situation and the input information and modify the task list based on the regenerated imaging instructions.

[0062] The imaging control unit 33 controls the camera 12 based on the positional relationship between the user and the camera 12, the current status in the task list, and the user's linguistic input. The current status in the task list includes not only the position of the current task but also the imaging instruction created by the response generation unit and recorded corresponding to the task position. The imaging control unit 33 converts the operation instruction for the camera 12 received from the context analysis unit 26 into an actual control command for the camera 12. For example, in response to an operation instruction for the camera 12 to move downward, the imaging control unit 33 performs control such as changing the vertical angle of the camera 12 by -20°.

[0063] The task of converting the user's intention for the camera 12, obtained from the context analysis of the linguistic input, into camera control parameters and commands is a relatively easy text conversion task that can be achieved by simple AI or rule-based processing. Therefore, this task may be performed by the simple response generator 28. In this case, the context analyzer 26 simply transmits the linguistic input to the LLM determiner 27 along with tags such as "Destination: simple response generator, Mission: camera control condition generation, Speech: No."

[0064] Before and / or after the imaging control unit 33 controls the camera 12, the first LLM 51 may analyze the image of the camera 12 after the control to detect the imaging target in the image, and output information indicating the position of the imaging target in the image to the imaging control unit 33 once or multiple times via an imaging instruction on a task list managed by the task list management unit 34.

[0065] The imaging control unit 33 may adjust the control of the camera 12 at least once based on the position or size of the imaging target detected from the image in the imaging range of the camera 12 .

[0066] For example, when receiving an instruction to "take a picture of the user's hands," the imaging control unit 33 may control the camera 12 taking into account the positional relationship between the user and the camera 12, and then control the camera 12 again so that the image of the user's hands detected from the image is in the center of the imaging range.

[0067] The control of the camera 12 by the imaging control unit 33 may be such that the imaging target can be appropriately imaged, and the imaging control unit 33 may adjust only the imaging target and the size of the image of the imaging target.

[0068] The imaging control unit 33 may have a simple object recognition function. In this case, the imaging control unit 33 can more appropriately control the camera 12 depending on the situation using an imaging object instruction and feedback from the camera 12 via the image input unit 22. Here, the imaging object instruction is an instruction specifying an object to be imaged, and may be an instruction from the task list management unit 34 based on a task list, or may be an instruction from the first LLM 51. Using the object recognition function, the imaging control unit 33 may change the imaging instruction generated by the first LLM 51 and recorded in the task list management unit 34 into an imaging object instruction and a size specification on the screen. Furthermore, the first LLM 51 may generate information on the imaging object instruction and the size specification on the screen as an imaging instruction.

[0069] The imaging control unit 33 may be configured to receive input from the camera 12 directly from the image input unit 22. That is, the imaging control unit 33 may be able to obtain feedback from the camera 12 via the image input unit 22. In this manner, when the imaging control unit 33 is configured to obtain an image directly from the image input unit 22, the context analysis unit 26 outputs an imaging instruction to the imaging control unit 33 by referencing the image. The imaging control unit 33 can control the camera 12 more appropriately depending on the situation by using recognition of the object to be imaged and feedback from the camera 12. Specifically, the imaging control unit 33 can automatically capture an image in which the target object is included at an appropriate size in the center by using object recognition and feedback from the camera 12.

[0070] Furthermore, by controlling the camera 12 through feedback from the camera 12, the response generation unit reduces the number of times it generates control information for the camera 12, allowing the first LLM 51 to allocate resources to task management and response generation for the conversation, thereby enabling the first LLM 51 to realize a higher quality conversation.

[0071] By using object recognition in this way, the first LLM 51 and the imaging control unit 33 can provide feedback on imaging conditions and imaging results, and by repeatedly controlling the camera 12, optimal imaging conditions can be achieved.

[0072] Furthermore, the imaging control unit 33 has a function of ignoring the object recognition function and following a direct user instruction such as "turn 30 degrees left." That is, the context analysis unit 26 may interpret the linguistic input and output an imaging instruction to the imaging control unit 33. This configuration enables the imaging process by the camera 12 to be performed at a higher speed.

[0073] As an example of capturing images in accordance with the progress of a task, if the user's linguistic input includes an instruction such as "I'm going to cook now, so please help me," the image capturing control unit 33 controls the camera 12 to capture images of the user and the user's surroundings. If the images acquired by the image input unit 22 include a page in a cooking magazine that contains a recipe, the context analysis unit 26 outputs an instruction to the image capturing control unit 33 to read the recipe from the cooking magazine. Specifically, the context analysis unit 26 outputs an image capturing instruction to the image capturing control unit 33 to enlarge the cooking magazine and capture the image so that the page is centered. Alternatively, the context analysis unit 26 may simply instruct the camera 12 to "read the recipe from the cooking magazine," and the image capturing control unit 33 may interpret the instruction and cause the camera 12 to capture the image.

[0074] In response to the instruction, the imaging control unit 33 outputs control information for the camera 12 to the camera 12 so that the camera 12 captures an image of the presented cooking magazine, taking into consideration the relative positions of the cooking magazine and the camera 12. The control information for the camera 12 may be, for example, information such as the attitude (imaging axis), angle of view, imaging frame, shutter speed, imaging resolution, and / or whether or not image stabilization is enabled.

[0075] Furthermore, if an electronic file containing task information, such as a work manual saved by the user, exists in the storage unit 16, the first LLM 51 may create a task list by referencing the electronic file. The format of the electronic file is not particularly limited, and may be, for example, an image format, a portable document format (PDF), or a text format. In this manner, the first LLM 51 preferably creates a task list from primary information containing task information, such as an input image or electronic file. The primary information containing task information may include information input by the user to the information processing system 100 as information indicating the content of the task, and information indicating a task identified based on the user's linguistic input (specific task information). The specific task information is information that specifically indicates the content of a commonly performed task designated by the user through linguistic input. For example, based on the linguistic input "I want to make curry," the first LLM 51 may retrieve a standard curry recipe from a database or the like as the primary information.

[0076] If the image acquired by the image input unit 22 does not contain content indicating task information, the context analysis unit 26 may generate a task list creation instruction without inputting an image and output it to the LLM determination unit 27. The first LLM 51 that receives the task list creation instruction may be unable to acquire an electronic file containing task information and may determine that an image containing task information needs to be input. In preparation for this case, the context analysis unit 26 may include, in advance, a request in the task list creation instruction, as necessary, for generating a response that prompts the user to show the cooking recipe to the camera 12.

[0077] A specific example of the process leading up to the creation of such a task list is shown below. Upon receiving a cooking supervision instruction, the context analysis unit 26 interprets the instruction and outputs a request to the LLM determination unit 27 stating, "I'd like you to assist me in cooking. This may be a task." The first LLM 51, upon receiving the request, generates a response, such as the following, inquiring about the details of the dish: "I understand. You're cooking, aren't you? What would you like to cook? If I know the name of the dish, I'll search for the recipe. If you have a recipe, please show it in front of the camera. I'll help you." In response to the response, the user may bring a cooking magazine in front of the camera, saying, for example, "This is it." The context analysis unit 26 outputs a request to the LLM determination unit 27 to create a cooking assistant task list and confirm the necessary conditions for the task. The context analysis unit 26 also outputs an imaging instruction to the imaging control unit 33 to enlarge the magazine and capture it so that the magazine is centered.

[0078] Once the camera 12 captures an image that allows the user to understand the recipe contents, the first LLM 51 creates a task list by determining the procedure and status of each step and creating image capture instructions for each step. Once the task list is created, the first LLM 51 begins task management using the task list and generates the following response: "Yes, I understand the recipe. Let's make it together. Are there any changes to your taste preferences or the number of people dining?" Thereafter, unless there is a particular problem, the image input unit 22 acquires images captured by the camera 12 at predetermined intervals. Furthermore, the context analysis unit 26 outputs instructions to the LLM determination unit 27 to analyze the input image as an automated task and to modify the task list as necessary. While executing the automated task, the first LLM 51 appropriately generates responses such as "We're going to do XX now. Please prepare YY and put it in ZZ." according to the progress of the task. The context analysis unit 26 also responds separately when it receives linguistic input from the user.

[0079] Conventional task management systems semi-automatically request task information from the user when managing tasks. On the other hand, the information processing system 100 according to the present disclosure allows the user to select a method for obtaining task information necessary for creating a task list from among many options other than requesting task information from the user so as to create a natural conversation.

[0080] The first LLM 51 creates a task list containing specific work processes based on the acquired task information, with content that is easy for the first LLM 51 to understand and that is also understandable to the user. The created task list has content that allows the first LLM 51 itself to interpret the content of each task included in the task information.

[0081] For example, if one item in a cooking recipe is "Add powdered sugar and lemon juice to a bowl," the first LLM 51 creates a task list including specific steps, such as "Add powdered sugar (dry white powder) to an empty bowl. Next, add lemon juice (liquid) to the bowl containing the dry white powder." In this way, the first LLM 51 functions as a task list creation unit.

[0082] A more detailed example of a task list including specific steps is as follows. If a cooking recipe includes the steps "1: Add powdered sugar to a bowl" and "2: Next, add lemon juice," even if the first LLM 51 stores the recipe as is and analyzes an image from camera 12 of the process of adding lemon juice, it would be impossible to determine the correct way to add the lemon juice from that single image. In other words, it would be difficult for the first LLM 51 to understand, from the original recipe, what step "next" in the recipe follows, or where the lemon juice should be added. The first LLM 51 supplements and converts the content of the recipe, along with a photo commentary in the recipe if necessary, into a clear, understandable expression for the first LLM 51, such as "2: Add lemon juice (liquid) to a bowl containing dry white flour." This task list allows the first LLM 51 to determine whether the process of adding lemon juice, "2: Next, add lemon juice," is being properly completed, even from a single image of the process of adding lemon juice.

[0083] By creating a task list that includes such specific process steps, the first LLM 51 makes it easy to identify the corresponding process step in the task list even from only a single image taken at a certain point during work.

[0084] After creating the task list, the first LLM 51 analyzes the images acquired by the image input unit 22 based on the progress of the tasks in accordance with the task list. The first LLM 51 may also analyze the user's linguistic input, or a combination of the linguistic input and the images, based on the progress of the tasks in accordance with the task list. Based on the analysis results, the first LLM 51 generates a conversational response (response) including instructions for the next step and an analysis of the current situation. In other words, the first LLM 51 generates a response regarding the user's progress of the tasks based on the results of analyzing at least one of the linguistic input and the images, and a task list showing multiple steps.

[0085] In this way, the first LLM 51 is a language model that can analyze images acquired by the image input unit 22 and interpret the analysis results in accordance with a task list. The first LLM 51 combines information interpreted from the image with the user's linguistic input, and can output responses such as instructions for tasks to be performed and an analysis of the current situation, as well as a modified task list as needed. In this way, the first LLM 51 also functions as a task list modification unit.

[0086] For example, if the user's verbal input includes a question about what to do when it is difficult to work according to the task list, the first LLM 51 may generate an alternative for the next step in the task list as a response.

[0087] Taking cooking as a specific example, for example, a user may find that an ingredient they had in mind is unavailable while cooking, or that they have mistakenly selected a different ingredient. In such a case, if the first LLM 51 receives a linguistic input from the user such as, "Actually, I don't have XX, what should I do?", it can suggest appropriate substitute ingredients, for example, as follows:

[0088] If you are making curry, you might suggest "beef → chicken or pork" and "potato → sweet potato or pumpkin"; if you are making nikujaga (meat and potato stew), you might suggest "pork → chicken or beef" and "potato → daikon radish or sweet potato"; if you are making okonomiyaki, you might suggest "cabbage → lettuce or Chinese cabbage" and "pork → chicken or shrimp." Or, for example, if you cannot find a whisk, the first LLM 51 might suggest using chopsticks or a fork instead.

[0089] Furthermore, if the user's linguistic input includes a question about how to interrupt a task, the first LLM 51 may determine whether the task can be interrupted in the current situation. If the first LLM 51 determines that the task can be interrupted in the current situation, it notifies the user of this fact and a method for interrupting the task via the response output unit 30. If the first LLM 51 determines that it is not desirable to interrupt the task in the current situation, it may notify the user of the steps up to the nearest task in the task list that can be interrupted via the response output unit 30.

[0090] When a user interrupts a task, the first LLM 51 may wait if it determines from the analysis results up to that point that the user will resume the task in a short time. Alternatively, if it determines that the user has stopped the task and will not resume it immediately, the first LLM 51 may terminate processing using the current task list. At this time, the first LLM 51 may record the current task list and progress status in the storage unit 16. The first LLM 51 can retrieve and reuse the task list recorded in the storage unit 16 when necessary, such as when it detects that the user has resumed the task.

[0091] The first LLM 51 may analyze the progress of the task based on the user's linguistic input and / or image input, and when it determines that all steps included in the task list have been completed, may generate a response asking the user for an evaluation of the work results. If subsequent linguistic input from the user includes information indicating the evaluation, the context analysis unit 26 may link the information indicating the evaluation with information on the corresponding task list and output the information indicating the evaluation to the conversation history management unit 31. In other words, the task list may include information on the user's evaluation of the results of executing the task.

[0092] Such a task list, including information on the evaluation of completed tasks, is recorded in the storage unit 16. The first LLM 51 can use the recorded task list as a historical task list in subsequent task monitoring processes. That is, when creating a task list, the first LLM 51 can create a task list that is optimized for the user and in line with the user's preferences by referencing the historical task list information obtained via the conversation history management unit 31.

[0093] (Response Control Unit and Response Output Unit) The response control unit 29 receives the response generated by the response generation unit and controls the output of the response. Specifically, when there are multiple responses, the response control unit 29 may control the output order of the responses. Furthermore, when multiple responses can be integrated or summarized, the response control unit 29 may integrate or summarize the multiple responses. Furthermore, the response control unit 29 may generate a response message by formatting the response generated by the response generation unit so that it is a natural conversation in response to the user's language input. Furthermore, the response control unit 29 may output the generated response message to the conversation history management unit 31. The response control unit 29 outputs the generated response message to the response output unit 30.

[0094] For example, when there is a response from the simple response generation unit 28 and a response from the LLM, the response control unit 29 may prioritize and output the response from the simple response generation unit 28 first. Also, for example, when the response from the LLM contains a tag that restricts speech, or a tag such as characters that are not suitable for pronunciation, vital information, or attributes, the response control unit 29 may extract and format the speechable portion and output it to the response output unit 30.

[0095] The response output unit 30 outputs the response message generated by the response control unit 29 via the output device 15. Furthermore, if the context analysis unit 26 does not generate a conversational response request, the response output unit 30 may output a response message indicating that there will be no response to the user's language input.

[0096] In the present disclosure, the response output unit 30 outputs the proposed next step work content based on the task list and the current analysis result as a response message. The conversation history management unit 31 records the response message in the storage unit 16. This allows the information processing system 100 to have a conversation with the user while reviewing the task progress to date.

[0097] (Conversation History Management Unit) The conversation history management unit 31 records and manages, as additional information, a conversation history including authentication information and a history task list, which is a task list created for each task that has been monitored to date, in addition to conversation text including language input and / or image input and responses to these inputs. This additional information is managed by appropriate tagging. For example, if any of the additional information has not changed from the previous time or if the context analysis unit 26 does not need any of the additional information, the conversation history management unit 31 does not need to record the additional information in the conversation history.

[0098] For example, for information that anyone can access, the conversation history management unit 31 may not need to record authentication information in the conversation history. In this way, the conversation history management unit 31 may select additional information to manage and record based on the results of the context analysis unit 26. By managing the additional information appropriately in this way, the conversation history management unit 31 can appropriately manage the additional information without wasting memory resources used for the conversation history.

[0099] The conversation history management unit 31 may manage the corrected task list, the language input and / or image input that triggered the correction, and the response thereto as a conversation history in association with each other.

[0100] The conversation history and the history task list are recorded in the storage unit 16. The conversation history management unit 31 can access the storage unit 16 as necessary and output the conversation history and the history task list to the context analysis unit 26.

[0101] Note that past conversation history data becomes large as the usage time and frequency of the information processing system 100 increases. Therefore, storing all conversation history data is undesirable from the perspective of increasing memory resources and data processing time. Of course, memory capacity and memory access speeds are still improving year by year, and it is highly likely that in the future it will be possible to record virtually all conversations, if not all. However, at the time a system is implemented, available resources are limited, and it is desirable to be able to efficiently manage conversation history using limited resources. Therefore, the conversation history management unit 31 may have a function to maintain the data stored in the memory unit 16 at an appropriate size. There are several methods for maintaining an appropriate data size for conversation history, and any of these methods can be applied to the information processing system 100.

[0102] The simplest method is to set a limit on the data size of the conversation history in advance, and delete the oldest data when the limit is exceeded. This method is easy to implement and is effective in reducing the size. However, this method has the problem that because old information is automatically deleted, it is difficult to maintain consistency in the conversation, especially with old information. Therefore, this method is suitable for applications where consistency in the conversation is sufficient over a relatively short period of time, such as one day's worth of data.

[0103] Another method is to set an upper limit on the data size of the conversation history and use a separate LLM (which can be the same model) to summarize it and reduce it to a predetermined amount of data. This method removes meaningless conversation history through summarization, increasing the ratio of valid conversations in the record, and therefore makes it possible to keep important conversations, even if they are old, for a relatively long time.

[0104] Furthermore, a suitable method for the information processing system 100 is to use Retrieval-Augmented Generation (RAG). RAG is an approach that combines information retrieval and generative models to generate more accurate answers to user questions. Below, we will briefly explain the basic operation of RAG, from creating embedded data to the search method.

[0105] 1. Creating Embedded Data: First, the conversation history is divided into segments of a predetermined size or a predetermined period, such as one day's worth of data, and each segment is summarized. The summarized data is then converted into vectors (embedded data) using a known embedding model, such as a pre-trained model like BERT, RoBERTa, or Sentence-BERT. An appropriate index may then be constructed. Furthermore, image data can be separated during the summaries and indexes, and stored, for example, on a cloud or on other inexpensive, high-capacity media where the image data can be referenced by the index. This improves memory utilization efficiency and enables longer-term storage. Furthermore, by including a brief description of the separated images in the summaries, older images can be recalled from the conversation as needed.

[0106] 2. Processing of User Questions: A user's question is converted into vector data using the same model as the history data. Using inner product calculations, for example, the top three most similar history data are extracted and sent to the context analysis unit 26 and the LLM, which is the response generation unit. Furthermore, if there is no data showing a similarity above a predetermined threshold, it can be prevented from being sent. This prevents the generation of an incorrect answer due to being influenced by less relevant information. Therefore, no matter how large the history data, the context analysis and response generation can utilize history data that is closely related to the question with a practical data size.

[0107] Furthermore, the conversation history management unit may delete the vectorized data itself in order to manage the size of the history data itself. As mentioned above, the deletion method may be to delete the oldest data first, or to appropriately re-summarize the data. Even in this case, important data can be retained for a much longer period of time than if non-vectorized information were deleted.

[0108] This method using RAG allows the contents of conversation history data to be retained relatively accurately and for a long period of time, making it particularly suitable for use in the information processing system 100. However, in order for RAG to function effectively, a certain limit on the number of histories and memory capacity is required. Therefore, a method using RAG or another method may be selected depending on the application and available resources. There are also several well-known methods for appropriately managing memory, and any of these may be used. Furthermore, multiple methods for managing history may be used in combination.

[0109] In either method, the conversation history is organized at an appropriate timing during breaks in the conversation. However, in order to maintain the consistency of the most recent conversation, it is preferable that the most recent conversation history, for example, a conversation history of about 10 turns, is neither compressed by summarization nor deleted.

[0110] (Task List Management Unit) The task list management unit 34 manages the progress information of the current tasks in the currently running task list in response to the response output from the language model. For example, the task list management may include a process of outputting instructions to modify, suspend, or resume the task list to the LLM determination unit 27. That is, the task list management unit 34 functions as a task list modification instruction unit, a task list suspension instruction unit, and a task list resume instruction unit.

[0111] Furthermore, in managing the task list, instructions specific to each step in the task list according to the progress of the task may be output to the LLM determination unit 27 or the imaging control unit 33 .

[0112] An example of an instruction specific to each step in the task list is an instruction to sequentially update information on the progress stage and execution status of the task in the task list. Specifically, the task list management unit 34 outputs an instruction to modify the task list to the LLM determination unit 27 so as to add or update information on the progress stage and execution status of the task to the task list according to the progress of the task. In other words, the task list may include information indicating the current progress stage and execution status of the task.

[0113] Another example of an instruction specific to each step in the task list is a control instruction for the camera 12 corresponding to each task in the task list. The task list management unit 34 can retain information about the control instructions for the camera 12 corresponding to each step in the task list. The control instructions for the camera 12 are control instructions for capturing an image of an object that can suitably analyze the user's work content in each step in the task list. The task list may include, in at least one stage of the task, information about an image capture instruction that instructs the object corresponding to the task. The task list management unit 34 can output a control instruction for the camera 12 to the image capture control unit 33 based on the image capture instruction information included in the task list. In this way, by including an image capture instruction in the task list and creating the image capture instruction when the task list is initially created, the first LLM 51 can efficiently support the task without having to generate image capture instructions sequentially as the task progresses. Furthermore, by being able to issue image capture instructions to the image capture control unit 33 through the task list, the first LLM 51 can generate an image capture instruction only when a change to the image capture instruction is necessary.

[0114] The image capture instruction information may be generated from task information indicating work steps such as a cooking recipe and included in the created task list when the task list is created by the first LLM 51. Furthermore, the task list management unit 34 may output a task list correction instruction to the LLM determination unit 27, based on historical information of past task monitoring including the historical task list, so as to add image capture instruction information appropriate for each step to the task list.

[0115] Examples of the objects to be imaged include the user's hands, the positions of cooking utensils, and the positions of cooking ingredients. Depending on the content of the process, the task list management unit 34 may output a control instruction for the camera 12 to the image capture control unit 33 so as to capture images of multiple objects to be imaged.

[0116] With this configuration, the task list management unit 34 can cause the camera 12 to capture an appropriate image without having to perform image analysis processing such as object recognition by the first LLM 51 or the like at each stage of the task.

[0117] Another example of an instruction specific to each step in the task list is an analysis request for the task list. The task list management unit 34 can generate and store optimized analysis requests useful for task list management. The analysis requests are analysis requests for each input based on the task list that the context analysis unit 26 generates when a linguistic input and / or an image input is received, and examples of these include the "analysis request including basic instructions" described above in the description of the context analysis unit 26. Analysis requests can sometimes be optimized depending on the content of each step in the task list. By storing such optimized analysis requests for each step, the task list management unit 34 can quickly output appropriate analysis requests to the LLM determination unit 27 while reducing the processing load of the analysis by the context analysis unit 26.

[0118] For example, in addition to the examples described above, the following requests can be exemplified as process-specific analysis requests optimized for each process. In the cooking process of boiling ingredients, it is sometimes important to confirm that the water is boiling before adding the ingredients. For example, when boiling pasta, always add salt to the boiling water before adding the pasta, so that the pasta does not stick and is deliciously cooked. Therefore, in the process of boiling pasta, a request to confirm that the water in the pot is boiling before adding the ingredients can be an example of the analysis request. The first LLM 51, which receives this analysis request, generates a response with the following content, for example, "Check the state of the water in the pot and let us know when it boils. When we give you the signal, please add salt and add the pasta."

[0119] Furthermore, for example, in meat dishes, it may be necessary to confirm that the meat is thoroughly cooked after the heating process. For example, in the heating process of chicken, it is important to cook the chicken until its color changes from pink to white and the center is completely white. Similarly, in the heating process of beef or pork, the meat can be safely eaten by thoroughly cooking it until its color changes from red to brown. Because insufficient cooking increases the risk of food poisoning, it is preferable for the information processing system 100 to make suggestions to reduce such risks in task management. Therefore, in the heating process of pork, a request to confirm that the pork has turned brown is an example of the analysis request. The first LLM 51 that receives the analysis request generates a response with the content, for example, "It is important to cook the pork thoroughly here. I'll check the color of the pork. Please wait a little longer."

[0120] (Experience Information Generation Unit) The experience information generation unit 32 generates experience information in which information indicating the user's situation when a task is executed is associated with a task list corresponding to the execution of the task. Examples of information indicating the user's situation include location information, an image, and time. The conversation history management unit 31 may record and manage the experience information in the storage unit 16 by including it in a history task list.

[0121] The experience information generation unit 32 may store the generated experience information directly in the storage unit 16, and the experience information generation unit 32 may manage the experience information. In other words, the experience information generation unit 32 may function as an experience information management unit.

[0122] The experience information is, for example, information indicating when (time), where (location information), what (image) the user looked at, and what task the user performed. One example of the experience information may include information identifying a subject in an image or a description of the image.

[0123] In the present disclosure, the experience information generating unit 32 is not an essential component.

[0124] (Example of Processing Flow) FIG. 2 is a flowchart showing an example of a monitoring process (information processing method) in the information processing system 100.

[0125] The voice input unit 21 receives a voice input from a user that includes a task monitoring request (S1). The context analysis unit 26 analyzes the context of the input language (S2). When the context analysis unit 26 detects that the input language includes a task monitoring request, it outputs a task list creation request to the LLM determination unit 27 (S3).

[0126] Based on the content of the task list creation request, the LLM determination unit 27 selects an appropriate LLM from the response generation unit (S4). In this embodiment, the LLM determination unit 27 selects the first LLM 51 as an LLM suitable for processing related to task monitoring. Based on the acquired task list creation request, the selected first LLM 51 creates a task list corresponding to the task requested to be monitored by the user (S5). At this time, the first LLM 51 may refer to task information included in a cooking recipe or the like in the acquired image, or may refer to an electronic file or the like recorded in the storage unit 16.

[0127] The image input unit 22 captures an image via the image capture control unit 33 at predetermined intervals or in response to language input from the user (S6). A specific description of the image capture process in S6 will be given later. The captured image is sent from the context analysis unit 26 to the first LLM 51 via the LLM determination unit 27.

[0128] The first LLM 51 analyzes the language input and / or image input acquired from the LLM determination unit 27 and interprets it in accordance with the task list (S7). Based on the interpretation result, the first LLM 51 determines whether the task list needs to be revised (S8). If it is determined that the task list needs to be revised (Yes in S8), the first LLM 51 revises the task list in accordance with the progress of the tasks, etc. (S9). If it is determined that the task list does not need to be revised (No in S8), the first LLM 51 does not revise the task list.

[0129] The first LLM 51 generates a response including a response to be taken by the user and an analysis of the current situation, etc., in accordance with the progress of the task interpreted from the linguistic input and / or image input, and outputs the response to the response control unit 29 (S10). Based on the output, the response control unit 29 outputs a response to the user via the output device 15.

[0130] The determination in S8 as to whether or not the task list needs to be modified may be made by the task list management unit 34 based on the output content from the first LLM 51 in S10.

[0131] Furthermore, if the first LLM 51 interprets the linguistic input and / or image input and determines that the task is progressing smoothly and that there is nothing to communicate to the user, it may not need to generate a response. This eliminates the need for the information processing system 100 to generate utterances at predetermined intervals, thereby avoiding unnecessary utterances and allowing the user to be appropriately notified of necessary information. Furthermore, if the first LLM 51 determines that there is nothing to communicate to the user, it may only generate simple utterances, such as "That's going well" or "That's good," indicating that the work is progressing smoothly.

[0132] The first LLM 51 determines whether the user has completed the task based on the language input and / or image input (S11). The first LLM 51 may determine whether the user has completed the task based on whether the last step or a specific step in the task list has been performed, or may determine whether all steps in the task list have been completed.

[0133] If it is determined that the user has not completed the task (No in S11), the information processing system 100 returns to S6 and continues to monitor the task. If it is determined that the user has completed the task (Yes in S11), the first LLM 51 congratulates the user by notifying them of the completion of the task, and generates a response asking for an evaluation of the task execution results, and outputs the response to the response control unit 29 (S12).

[0134] The information processing system 100 may wait until it receives feedback on the evaluation from the user. Furthermore, if the information processing system 100 determines that the evaluation result cannot be obtained immediately due to the time the user will be eating the dish, the information processing system 100 may temporarily terminate the processing. In this case, when the first LLM 51 later receives linguistic input related to the evaluation from the user, the first LLM 51 may identify the corresponding task list recorded in the storage unit 16 and add the evaluation information.

[0135] A specific example of the imaging process in S6 will now be described with reference to FIGS.

[0136] 3 is a flowchart showing an example of the imaging process at predetermined intervals in the process of S6 of the information processing system 100. When the task monitoring process is started by creating a task list, the image input unit 22 captures images at predetermined intervals via the imaging control unit 33 (S6).

[0137] First, the image input unit 22 determines whether a predetermined interval has elapsed since the previous language input or image input from the user (S21). If it is determined that the predetermined interval has not elapsed (No in S21), the image input unit 22 waits until the predetermined interval has elapsed. If it is determined that the predetermined interval has elapsed (Yes in S21), the image input unit 22 outputs an image capture instruction to the imaging control unit 33 (S22). The imaging control unit 33 controls the camera 12 to capture an image (S23).

[0138] Here, an example is shown in which the image input unit 22 determines whether a predetermined interval has elapsed. However, this determination may be made without using system resources, by having a set interval timer device (not shown) generate an interrupt and automatically generate an image capture instruction for the image capture control unit 33. This allows for more system resources when each interval elapses, allowing the information processing system 100 to more fluently and accurately engage in conversations unrelated to task monitoring.

[0139] The first LLM 51 analyzes the image input obtained by capturing an image and interprets it in accordance with the task list (S24). Note that the process of S24 is an example of the process of S7 described above.

[0140] The predetermined interval may be set by the first LLM 51 based on the content of the task, or may be set by the task list management unit 34. The first LLM 51 may also create a task list that includes information about the predetermined interval, and in this case, the image input unit 22 may acquire the information about the predetermined interval included in the task list and perform imaging.

[0141] 4 is a flowchart showing an example of an image capturing process in response to a language input from a user in the process of S6 of the information processing system 100. Here, an example of a processing flow when the user asks, "I don't have a whisk, what should I do?" while monitoring a cooking task will be described.

[0142] The voice input unit 21 accepts language input from the user (S31). The context analysis unit 26 analyzes the context of the input language (S32). The context analysis unit 26 detects through the analysis of the context that a question requiring imaging is included, and outputs an imaging instruction to the imaging control unit 33 (S33). The imaging control unit 33 controls the camera 12 to capture an image (S34).

[0143] The first LLM 51 analyzes the language input and image input received from the context analysis unit 26 via the LLM determination unit 27 and interprets them in accordance with the task list (S35). Here, the first LLM 51 generates an alternative task plan for when a whisk is not used in the current task. For example, if the first LLM 51 determines that the current task can be performed using a mixing tool other than a whisk, it generates a response to notify the user of this.

[0144] In addition, the first LLM 51 proposes an alternative plan and modifies the task list as needed to proceed with the task based on the alternative plan (S36). In the above example, the first LLM 51 may add to the task list a process for checking the finished product when stirring with a device other than a whisk, along with a modification indicating that the current process was performed with a device other than a whisk.

[0145] The first LLM 51 outputs the response including the generated corrected task list and the alternative to the response control unit 29 (S37). Note that the processing of S35 to S37 is an example of the processing of S7 to S10 described above.

[0146] (Modifications) Various modifications of the present embodiment are conceivable. As shown in FIG. 1 , an information processing system 100 a according to one modification of the present embodiment further includes a vital sensor 13 (sensor), a vital information input unit 23 (sensor input unit), and a vital information analysis unit 24.

[0147] The vital sensor 13 acquires biological information of the user and outputs a signal corresponding to the biological information. Examples of the biological information acquired by the vital sensor 13 include body temperature, pulse rate, blood pressure, sweating, and respiratory rate.

[0148] The vital information input unit 23 functions as a sensor information input unit that acquires the user's vital information as sensor information. The vital information input unit 23 organizes data such as pulse rate and blood pressure into a form that can be understood as biological data by classifying the output of the vital sensor 13 into specified items and numerical values. The information organized in this manner is referred to as first vital information. The first vital information may be text data.

[0149] For example, converting the output of an infrared sensor into a temperature value that humans can understand is called "calibration" or "data conversion." These processes are procedures that convert the sensor's output into an accurate physical quantity (in this case, temperature). This process involves correcting or adjusting the sensor's output voltage or digital signal to convert it into an actual temperature.

[0150] In the information processing system 100a, the data conversion part is included in the processing by the vital information input unit 23. On the other hand, the part corresponding to the calibration is performed at the stage of creating the information processing system 100a or at the stage of first starting up the information processing system 100a, and is stored as basic data for the vital information input unit 23. However, it goes without saying that the processing by the vital information input unit 23 of the information processing system 100a may include a calibration function.

[0151] Here, we will provide a basic explanation of everything from calibration to data conversion. Of course, specific operations may differ depending on the type of sensor. However, since the relationship between the sensor and physical quantity is well-known, it can be handled using well-known methods.

[0152] Accurately establishing the relationship between a sensor's output and the actual physical quantity (temperature) essentially involves the following processes: ・Measurement of a reference point: Measure the sensor's output in a known temperature environment to obtain a reference point. ・Determining the relationship: Determine the relationship between the measured output and the actual temperature (e.g., linear or non-linear). ・Tuning: Adjust the sensor's output to match the actual temperature. This may mean changing the sensor's hardware or software settings. ・Verification: Verify that the sensor's output is accurate after calibration. This may involve re-measuring in a different temperature environment.

[0153] Data conversion uses calibration data to convert the sensor's output (e.g., analog voltage or digital value) into a human-understandable format (e.g., a temperature number). In other words, data conversion is the mapping of the sensor's output to a meaningful physical quantity. This mapping involves calculations and corrections based on the sensor's characteristics.

[0154] For example, in temperature conversion using an infrared sensor, the infrared sensor measures the intensity of infrared radiation emitted from an object and outputs that data as a voltage or digital signal (vital sensor 13). Converting this output to temperature requires the following steps: Understand the sensor's characteristics: Check the sensor's datasheet to understand the relationship between output and temperature. For example, check how the output voltage corresponds to temperature over a specific range. Obtain reference data: Measure the sensor's output in a known temperature environment to obtain reference data. For example, record the output at a known temperature point, such as the freezing point or boiling point. Create a conversion formula: Create a conversion formula to calculate temperature from the sensor's output. This may involve linear interpolation, nonlinear interpolation, or fitting a calibration curve. Then, the sensor's output and conversion formula can be used to convert data such as "Temperature: 36.5°C."

[0155] The vital information analysis unit 24 analyzes the user's physical condition based on the first vital information. The vital information analysis unit 24 may also analyze the user's mental state based on the first vital information. The analysis results of the user's physical condition may be indicated by, for example, whether any abnormality, such as fever, arrhythmia, hyperventilation, hypothermia, extreme sweating, abnormally fast pulse (tachycardia), abnormally slow pulse (bradycardia), high blood pressure, or low blood pressure, is detected, or whether the user is in a normal state with no abnormality detected. In addition to obvious abnormalities, for example, for a cooking task, the information on the analysis results of the physical condition can be used to modify the cooking method, such as reducing salt intake, for a user with persistent high blood pressure. Alternatively, for an exercise task, the analysis results of the physical condition can be used to adjust the pace, etc.

[0156] Examples of mental states include excitement, anger, sadness, joy, pleasure, and suffering. Information indicating these analysis results is referred to as second vital information. The first vital information and the second vital information may also be collectively referred to as vital information. The vital information analysis unit 24 analyzes the first vital information and derives the second vital information using a pre-set database indicating the correspondence between the first vital information and the second vital information. The second vital information may be text data. The vital information analysis unit 24 sends the second vital information to the context analysis unit 26. The vital information analysis unit 24 may also send the first vital information to the context analysis unit 26.

[0157] The vital information analysis unit 24 may be a small-scale AI (CNN, RNN) trained for the purpose of deriving the second vital information, or may be a judgment circuit based on an appropriate database. When the vital information analysis unit 24 is the above-mentioned AI, it is considered that the first vital information (physiological indicators) and the second vital information (mental state) are closely related during learning, and a learning or judgment database can be created based on this information.

[0158] Here, the relationship between the first vital information and the second vital information will be explained using some general principles, but the second vital information may be created using a unique index.

[0159] For example, the relationship between body temperature and mental state has been shown to increase with increasing stress or anxiety. In other cases, it is known that body temperature tends to become unstable in people with depression. This relationship between body temperature and mental state has been well-known through numerous studies, and by applying appropriate weighting, it is possible to infer a person's mental state.

[0160] There is also a great deal of research and clinical data on the relationship between pulse (heart rate) and mental state, and it is known that stress, excitement, or anxiety can cause the pulse to speed up, while relaxation can slow the pulse.

[0161] It is also known that breathing rate increases in response to stress, anxiety, or fear, and becomes deeper and slower in a relaxed state.

[0162] By learning this information in advance and processing it statistically, mental state can be determined using simple AI or a database.

[0163] The information processing system 100a is a system for conducting conversations. The information processing system 100a may include a system (not shown) that enables the LLM to infer the user's mental state from the content of the conversation and feeds back a combination of the first and / or second vital information and the mental state determined by the conversation to the vital information analysis unit 24. By including a feedback system, the information processing system 100a can respond to situations specific to the user and can support the generation of conversations that are more tailored to the user.

[0164] The vital information analysis unit 24 may be capable of detecting vital information that is life-threatening to the user or vital information that has been preset as information that should trigger an alert. When the vital information analysis unit 24 detects the vital information, the vital information analysis unit 24 may make an emergency call to a medical institution without performing any other processing. Although not shown in the drawings, the emergency call may include contact via a public line, a loud voice or alarm sound notification requesting help from those in the vicinity, or a message function to a pre-designated party.

[0165] The vital information analysis unit 24 outputs second vital information, which is the analysis result of the vital information, to the context analysis unit 26. The context analysis unit 26 outputs the second vital information together with the language input and the image input to the LLM determination unit 27. The first LLM 51 may change the work content suggested in response or may suggest interrupting the task, depending on the second vital information and / or the first vital information of the user.

[0166] Furthermore, the context analysis unit 26 or the task list management unit 34 may serve as a task list modification instruction unit and instruct the first LLM 51 to modify the task list based on the vital information via the LLM determination unit 27. The first LLM 51 can modify the task list based on the input task list modification request or its own analysis results to include the user's second vital information and / or first vital information in the task list.

[0167] With this configuration, after the user has performed a task, a task list is obtained that includes the analysis results of vital information such as the user's physical and mental state, and information about the tasks that were actually performed. By recording this task list as a historical task list in the storage unit 16, the first LLM 51 can create a task list and suggest work contents that are appropriate for the user's current physical and mental state, taking into account the analysis results of the user's vital information.

[0168] As described above, the information based on the sensor information, such as the first vital sign information and the second vital sign information, may be text data, which allows the information processing system 100a according to the present embodiment to handle the information based on the sensor information together with linguistic input.

[0169] (Processing Flow According to Modification) FIG. 5 is a flowchart showing an example of processing for acquiring vital information in the information processing system 100a.

[0170] The vital information input unit 23 acquires first vital information as vital information of the user from the vital sensor 13 (S41). The vital information input unit 23 may acquire the vital information in real time. Alternatively, the vital information input unit 23 may acquire the vital information at the timing of image capture by the image input unit 22 at predetermined intervals, or at the timing of verbal input from the user.

[0171] The vital information analysis unit 24 analyzes the first vital information acquired from the vital information input unit 23 (S42). The vital information analysis unit 24 generates second vital information indicating the user's physical condition, mental state, etc. as a result of analyzing the first vital information.

[0172] The first LLM 51 determines whether the second vital information and / or the first vital information acquired from the vital information analysis unit 24 via the context analysis unit 26 indicates that the user's physical or mental state is normal (S43). If the second vital information indicates some abnormality (Yes in S43), the first LLM 51 modifies the task list so that the second vital information indicating the abnormality is included in the task list (S44).

[0173] The first LLM 51 may also present an alternative task to the user based on the details of the abnormality contained in the second vital sign information. For example, an alternative task may be proposed to change the recipe to make the food less seasoned when the second vital sign information indicates an abnormality in fever. In this case, the first LLM 51 may modify the details of the process for which the alternative task is presented in the task list.

[0174] For example, the first LLM 51 may suggest an alternative to reducing the salt content of a recipe when detecting an increase in blood pressure from the second vital information. Furthermore, for example, when detecting a deviation from normal mental state from the second vital information, the first LLM 51 may suggest an alternative to adding ingredients that are effective in calming the mood if the person is in an excited state (including excessive stress or anxiety). Examples of ingredients that are effective in calming the mood include adding nuts (almonds, cashews), green leafy vegetables (spinach, kale), beans (black beans, lentils), and herbs such as lavender or valerian root. Conversely, if the person is in a depressed state, the first LLM 51 may suggest an alternative to adding ingredients that are effective in lifting the mood. Examples of foods that are useful for elevating mood include turkey, chicken, eggs, dairy products (yogurt, cheese), bananas, salmon, tuna, egg yolks, fortified milk, folic acid (spinach, asparagus, beans), vitamin B12 (shellfish, chicken, dairy products), etc. Alternatively, the first LLM 51 may suggest alternatives that include adjusting the sweetness to calm the mood or increasing the spiciness (e.g., chili peppers) to elevate the mood.

[0175] If the second vital information is determined to indicate that the user's physical condition is normal (No in S43), the first LLM 51 can skip modifying the task list related to the second vital information in S44.

[0176] In addition, the first LLM 51 may modify the task list to include the second vital information at each point in time, regardless of the content of the second vital information, each time the second vital information is obtained or each time the task process progresses.

[0177] [Embodiment 2] Another embodiment of the present disclosure will be described below. For ease of explanation, the same reference numerals will be used to designate components having the same functions as those described in the above embodiment, and the description thereof will not be repeated.

[0178] 6 is a block diagram illustrating the configuration of an information processing system 200 according to embodiment 2. As shown in FIG. 6, the information processing system 200 includes an information processing device 210 and a server 220.

[0179] The information processing device 210 is a conversation terminal that includes elements other than the first LLM 51 and the second LLM 52 in the information processing system 100a. In particular, the information processing device 210 integrates the microphone 11, camera 12, vital sign sensor 13, and fingerprint sensor 14 into a single device. This allows accurate information about the user to be acquired with a high degree of certainty. Furthermore, integrating these input devices reduces the possibility of the user's personal information being leaked.

[0180] The server 220 includes the first LLM 51 and the second LLM 52 in the information processing system 100a. A publicly available server or a secure server managed by an individual or organization may be used as the server 220. In FIG. 6 , a single server 220 includes the first LLM 51 and the second LLM 52. However, in the information processing system 200, the server including the first LLM 51 and the server including the second LLM 52 may be separate servers.

[0181] If the server 220 manages multiple LLMs that exist inside or outside the server 220, the server 220 may have a function to call a publicly accessible server from among the LLMs it manages. For example, if the first LLM 51 is located inside the server 220 and the second LLM 52 is located on a different server, the server 220 may access the second LLM 52 via the different server.

[0182] The server 220 may have a function for switching the LLMs it manages. In this case, the server 220 may manage attribute tags of the LLMs it manages and share the attribute information with the LLM determination unit 27. This allows the LLM determination unit 27 to obtain information for selecting an appropriate LLM from currently available LLMs. The server 220 may have information on switchable LLMs and servers containing LLMs, or may be configured to obtain information from a separate database.

[0183] In the example shown in FIG. 6 , the voice input unit 21, image input unit 22, vital information input unit 23, vital information analysis unit 24, authentication unit 25, context analysis unit 26, LLM determination unit 27, simple response generation unit 28, response control unit 29, response output unit 30, conversation history management unit 31, experience information generation unit 32, imaging control unit 33, and task list management unit 34 are all present on the same information processing device 210. These units may be integrated into the information processing device 210 in a state where they are integrated on a chip. Alternatively, these units may be distributed and arranged on a secure server or the like. The information processing system 200 may be realized as a program executed by computers provided in the information processing device 210 and the server 220.

[0184] 7 is a perspective view illustrating the terminal device 230. The terminal device 230 is a portable electronic device equipped with the information processing device 210.

[0185] As shown in Fig. 7, the terminal device 230 includes a mounting unit 235. The mounting unit 235 is a member for mounting the terminal device 230 around the user's neck. The mounting unit 235 has an open-ring shape that can be hooked around the user's neck. With this configuration, the user's movements are less likely to be hindered even when the terminal device 230 is mounted.

[0186] The attachment unit 235 has a microphone 11. Specifically, the microphone 11 is disposed at one end of the attachment unit 235, which has an open ring shape. This position is near the user's mouth when the user wears the terminal device 230 around their neck. Therefore, the user can easily input voice via the microphone 11 while wearing the terminal device 230 around their neck.

[0187] The mounting unit 235 has a camera 12. Specifically, the camera 12 is arranged near the end of the open-ring-shaped mounting unit 235 so as to face outward. This allows the orientation of the camera 12 to roughly match the orientation of the user's face when the user is wearing the terminal device 230 around their neck. This makes it possible for the camera 12 to capture what the user sees.

[0188] The attachment unit 235 has a vital sensor 13. Specifically, the vital sensor 13 is disposed inside a portion of the attachment unit 235 that is open-ring shaped, opposite to the open portion. As a result, when the user wears the terminal device 230 around their neck, the vital sensor 13 comes into contact with the user's neck. Because the vital sensor 13 is integrated with the attachment unit 235, vital information of the user wearing the terminal device 230 can be reliably acquired, and vital information can be acquired whenever necessary.

[0189] The attachment part 235 has a fingerprint sensor 14. Specifically, the fingerprint sensor 14 is disposed at the end of the open-ring-shaped attachment part 235 opposite to the end where the microphone 11 is disposed. This position allows the user to easily touch the fingerprint sensor 14 with their finger when the user is wearing the terminal device 230 around their neck. Therefore, the user can easily perform personal authentication.

[0190] The wearing unit 235 has an output device 15. Specifically, speakers serving as the output device 15 are arranged on each of the left and right sides of the wearing unit 235, which has an open ring shape, when the open part is in the front. These positions are near the user's ears when the user wears the terminal device 230 around their neck. The user can easily hear the output from the output device 15 when wearing the terminal device 230 around their neck.

[0191] The terminal device 230 may allow the user to specify which first vital information should be acquired or which first vital information should not be acquired in the vital information input unit 23. For example, this may be the case when the user intends to monitor only their own blood pressure. The terminal device 230 may allow the user to specify which first vital information should be acquired or which first vital information should not be acquired by voice input. Alternatively, the terminal device 230 may allow the user to specify which first vital information should be acquired or which first vital information should not be acquired via another terminal such as a smartphone.

[0192] The terminal device 230 may also be configured to be able to turn the function of the vital sensor 13 on and off. For example, if the vital sensor 13 has functions for measuring body temperature, pulse rate, blood pressure, sweat rate, and respiratory rate, the user may be able to select not to measure one or more of these. In this case, the attachment unit 235 may be provided with an interface, such as a button, for turning the function of the vital sensor 13 on and off. Furthermore, the user's instructions regarding the operation of the terminal device 230 may be input to the terminal device 230 via another terminal device, such as a smartphone, that transmits the user's instructions to the terminal device 230. That is, the terminal device 230 may itself be provided with an interface that allows the user to configure settings related to the acquisition of sensor information, or the interface may be provided by using a device capable of communicating with the terminal device 230.

[0193] [Example of implementation by software] The functions of an information processing device (hereinafter referred to as the "device") can be realized by an information processing program for causing a computer to function as the device, and a program for causing a computer to function as each control block of the device (in particular, the voice input unit 21, the image input unit 22, the vital information input unit 23, the vital information analysis unit 24, the authentication unit 25, the context analysis unit 26, the LLM determination unit 27, the simple response generation unit 28, the response control unit 29, the response output unit 30, the conversation history management unit 31, the experience information generation unit 32, the imaging control unit 33, and the task list management unit 34).

[0194] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., a memory) as hardware for executing the program. The functions described in each of the above embodiments are realized by executing the program using the control device and storage device.

[0195] The program may be non-transitory and may be recorded on one or more computer-readable recording media. The recording media may or may not be included in the device. In the latter case, the program may be supplied to the device via any wired or wireless transmission medium.

[0196] In addition, some or all of the functions of each of the control blocks can be realized by logic circuits. For example, integrated circuits in which logic circuits that function as each of the control blocks are formed are also included in the scope of the present disclosure. In addition, the functions of each of the control blocks can also be realized by, for example, a quantum computer.

[0197] Furthermore, each process described in each of the above embodiments may be executed by AI (Artificial Intelligence). In this case, the AI ​​may run on the control device or on another device (for example, an edge computer or a cloud server).

[0198] The present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present disclosure. Furthermore, new technical features can be formed by combining the technical means disclosed in each embodiment.

[0199] DESCRIPTION OF SYMBOLS 11 Microphone 12 Camera 13 Vital sensor (sensor) 15 Output device 21 Voice input unit (language input unit) 22 Image input unit 23 Vital information input unit 24 Vital information analysis unit 25 Authentication unit 26 Context analysis unit (task list creation instruction unit, task list correction instruction unit) 27 LLM determination unit 28 Simple response generation unit (response generation unit) 29 Response control unit 30 Response output unit 31 Conversation history management unit 32 Experience information generation unit (experience information management unit) 33 Imaging control unit (image analysis unit) 34 Task list management unit (task list correction instruction unit, task list interruption instruction unit, task list restart instruction unit) 51 First LLM (response generation unit, language model) 52 Second LLM (response generation unit, language model) 100, 100a, 200 Information processing system 210 Information processing device 230 Terminal device

Claims

1. An information processing device comprising: a language input unit that accepts language input from a user; an image input unit that acquires images of the user performing a task including multiple steps; and a response control unit that receives a response regarding the progress of the task by the user, generated based on the results of analyzing at least one of the language input and the images and a task list indicating the multiple steps, and controls the output of the response.

2. The information processing device according to claim 1, further comprising a task list creation instruction unit that acquires task information relating to tasks to be performed by the user and instructs a language model to create the task list based on the task information.

3. The information processing device according to claim 2, further comprising a task list correction instruction unit that instructs the language model to correct the task list based on the progress of the tasks.

4. The information processing device according to claim 2 or 3, wherein the task list includes information on basic instructions for instructing the language model to perform image analysis of the image.

5. An information processing device according to any one of claims 1 to 4, wherein the task list includes imaging instruction information that indicates an imaging target corresponding to the task in at least one stage of the task.

6. The information processing device according to any one of claims 1 to 5, wherein the task list includes information indicating the current progress stage and execution status of the tasks.

7. An information processing device according to claim 3, wherein the task list correction instruction unit instructs the language model to correct the task list based on information indicating a change in a prerequisite condition for the task, the information being included in the language input and / or the image.

8. The information processing device according to claim 7, wherein the task list correction instruction unit instructs the language model to correct the information on the prerequisites included in the task list based on information indicating a change in the prerequisites.

9. The information processing device according to any one of claims 1 to 8, wherein the task list includes information on an evaluation from the user regarding the results of executing the tasks.

10. The information processing device according to any one of claims 2 to 4, 7 and 8, wherein the language model creates the task list from primary information including the task information.

11. A terminal device comprising the information processing device according to claim 1.

12. The terminal device according to claim 11, comprising: a microphone for capturing the user's voice; a camera for capturing the image; and an output device for outputting the response generated by the response control unit.

13. The terminal device according to claim 11 or 12, comprising a sensor that acquires sensor information including vital information.

14. The terminal device according to claim 13, wherein the information processing device comprises a task list correction instruction unit that instructs a language model to correct the task list based on the vital information.

15. An information processing program for causing a computer to function as the information processing device according to claim 1, the information processing program causing a computer to function as the language input unit, the image input unit, and the response control unit.

16. An information processing system comprising: a language input unit that accepts language input from a user; an image input unit that acquires images of the user performing a task including multiple steps; a response generation unit that generates a response regarding the progress of the task by the user based on the results of analyzing at least one of the language input and the images and a task list indicating the multiple steps; and a response control unit that receives the response generated by the response generation unit and controls the output of the response.

17. An information processing method executed by a computer, comprising: a step of accepting linguistic input from a user; a step of acquiring an image of the user performing a task including a plurality of steps; and a step of receiving a response regarding the progress of the task by the user, the response being generated based on the results of analyzing at least one of the linguistic input and the image and a task list indicating the plurality of steps, and controlling the output of the response.

Citation Information

Patent Citations

  • Content provision support method and server device

    JP2016209030A

  • Information processing device, information processing method, program and distribution system

    WO2024161686A1