Information processing device, terminal device, information processing program, information processing system, and information processing method
The information processing apparatus with a moving camera and imaging control system addresses the limitation of conventional imaging by autonomously capturing subjects relevant to user conversations, improving the relevance and responsiveness of captured images.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2026-03-19
AI Technical Summary
Conventional imaging apparatuses cannot autonomously image a subject corresponding to a conversation, limiting the ability to capture images relevant to user interactions.
An information processing apparatus communicably connected to a camera that moves with the user, incorporating a language input unit and an imaging control unit to control the camera based on positional relationship and user input, enabling autonomous subject imaging.
Enables the autonomous capture of images relevant to user interactions, enhancing the responsiveness and relevance of captured subjects in conversation systems.
Smart Images

Figure JP2025029813_19032026_PF_FP_ABST
Abstract
Description
Information Processing Apparatus, Terminal Apparatus, Information Processing Program, Information Processing System, and Information Processing Method
[0001] The present disclosure relates to an information processing apparatus that responds to a user's language input, and a terminal apparatus including the information processing apparatus, etc.
[0002] Patent Document 1 discloses an imaging apparatus realized as a digital still camera or a digital video camera. The imaging apparatus is configured to evaluate the expression of a person's face in a captured image, automatically release the shutter according to the degree of evaluation, and record still image data. As the evaluation of the expression, it is disclosed that an index called a smile score is used to evaluate the degree to which the expression is a smile or not.
[0003] Japanese Patent Application Laid-Open No. 2008-042319
[0004] In a conversation system that converses with a user by language input, a more natural conversation can be conducted by incorporating image data capturing a subject based on the user's interest into the conversation. Although a conventional imaging apparatus can automatically release the shutter, the user selects the subject. Therefore, the conventional imaging apparatus could not autonomously image a subject corresponding to the conversation.
[0005] One aspect of the present disclosure aims to realize an information processing apparatus capable of autonomously imaging a subject corresponding to a conversation.
[0006] In order to solve the above problems, an information processing apparatus according to one aspect of the present disclosure is communicably connected to a camera that moves as the user moves, and includes a language input unit that receives a language input from the user, and an imaging control unit that controls the camera based on the positional relationship between the user and the camera and the language input.
[0007] An information processing system according to one aspect of the present disclosure is an information processing system including a camera that moves in conjunction with the movement of a user and an information processing device that is communicably connected to the camera, wherein the information processing device comprises a language input unit that receives language input from the user and an imaging control unit that controls the camera based on the positional relationship between the user and the camera and the language input.
[0008] An information processing method according to one aspect of the present disclosure is an information processing method used in an information processing device that is communicably connected to a camera that moves in conjunction with the movement of a user, and includes the steps of receiving language input from the user and controlling the camera based on the positional relationship between the user and the camera and the language input.
[0009] According to one aspect of this disclosure, it becomes possible to autonomously image a subject in response to a conversation.
[0010] This is a block diagram illustrating the configuration of the information processing system according to Embodiment 1. This is a flowchart showing an example of the processing of the information processing system according to Embodiment 1. This is a flowchart showing another example of the processing of the information processing system according to Embodiment 1. This is a flowchart showing another example of the processing of the information processing system according to Embodiment 1. This is a block diagram illustrating the configuration of the information processing system according to Embodiment 2. This is a perspective view illustrating a terminal device according to Embodiment 2.
[0011] [Embodiment 1] Hereinafter, an embodiment of the present disclosure will be described in detail with reference to the drawings. In the drawings, the same or substantially the same components will be denoted by the same reference numerals and will not be repeated in the description.
[0012] The terms "First," "Second," etc., used in this disclosure are used to distinguish one component from another, and are not intended to limit the number, order, or priority of such components. For example, the presence of "First Element" and "Second Element" does not mean that only two elements, "First Element" and "Second Element," are adopted, nor does it mean that "First Element" must precede "Second Element."
[0013] Figure 1 is a block diagram illustrating the configuration of an information processing system 100 according to Embodiment 1. The information processing system 100 is an AI (Artificial Intelligence) conversation system.
[0014] As shown in Figure 1, the information processing system 100 includes a microphone 11, a camera 12, a fingerprint sensor 14, an output device 15, and a storage unit 16. The information processing system 100 also includes a voice input unit 21 (language input unit), an image input unit 22, an imaging control unit 23, an image analysis unit 24, an authentication unit 25, a context analysis unit 26, an LLM (Large Language Model) determination unit 27, a simple response generation unit 28, a response control unit 29, a response output unit 30 (output unit), a conversation history management unit 31, and an experience information generation unit 32. Furthermore, the information processing system 100 includes a first LLM 51 and a second LLM 52. The information processing system 100 may also include a position sensor (not shown) to acquire experience information, which will be described later.
[0015] Microphone 11 is a voice input device that receives voice input from a user to the information processing system 100.
[0016] Camera 12 has an image sensor such as a CCD (charge-coupled device) or CMOS (complementary metal oxide semiconductor). Camera 12 is an imaging device that captures images desired by the user by changing imaging conditions (angle of view, frame rate, etc.) mainly in accordance with commands from the imaging control unit 23. There are no particular limitations on the type or number of cameras 12.
[0017] In this embodiment, the camera 12 is described as a device that moves with the user. Such a camera 12 may be attached to the user's body, or it may be a camera carried by the user (for example, a camera hanging around the user's neck). In addition to being hung around the user's neck, the camera 12 may also be fixed to a bag, suitcase, or the like that the user carries.
[0018] Furthermore, camera 12 may be installed on a mobile mechanism capable of autonomously tracking a moving user. Examples of such mobile mechanisms include drones and walking robots. Alternatively, camera 12 may be installed on a mobile mechanism operated primarily by the user. Examples of such mobile mechanisms include vehicles, motorcycles, and bicycles. In any of these configurations, camera 12 will move in accordance with the user's movement.
[0019] The fingerprint sensor 14 is a fingerprint detection device that detects the user's fingerprint.
[0020] The output device 15 is a device that transmits information to the user. Information is transmitted to the user as a response to voice input from the user. The output device 15 is, for example, an image display device and a speaker. The image display device is a device for displaying responses to the user's conversation, displaying images captured by the camera 12, and providing notifications to the user via display. The speaker is an audio output device for outputting voice responses to the user's conversation and providing notifications to the user via voice. The output device 15 may also be equipped with a vibration device to emphasize notifications to the user, or a transmitter for emergency contact, etc.
[0021] The storage unit 16 stores information necessary for controlling the information processing system 100. The information processing system 100 may be connected to an external storage device acting as the storage unit 16 in a communicative manner. In other words, the storage unit 16 may be located outside the information processing system 100.
[0022] The voice input unit 21 (language input unit) functions as a language input unit that receives voice input (language input) from the user via the microphone 11. The voice input unit 21 converts the input voice into text data. Note that input to the information processing system 100 is not limited to voice input from the user. For example, the user may input language to the information processing system 100 by text input. When language input is performed by text input, a text input unit can be provided that receives the text as language input via a touch panel or a text input device such as a smartphone that is connected to the context analysis unit 26 in a communicative manner. In the following description, it will be assumed that the user performs language input by voice.
[0023] The image input unit 22 acquires an image from the camera 12 and converts it into embedded data (text data). This conversion is performed to make the image into a data format that the first LLM 51 or the second LLM 52 can understand. The image captured by the camera 12 may be a still image or a video. That is, the output from the camera 12 to the image input unit 22 may be a still image or a video.
[0024] Several conversion methods are known for handling image data as text data. In the field of AI, the following known methods can be appropriately selected or combined for optimal use.
[0025] For example, the conversion method may be Base64 encoding. Base64 encoding is a method for converting binary data into ASCII text and is widely used when handling binary data such as image files in text format. Base64 encoding is often used when embedding images as data URIs in applications and can also be used in the information processing system 100.
[0026] Furthermore, the conversion method may be hexadecimal encoding. Hexadecimal encoding is a method of converting binary data into a hexadecimal string. Hexadecimal encoding is generally used to visualize binary data for debugging or data analysis rather than images, and is rarely used in the information processing system 100. However, if the provided image is a computer graphics image or, in particular, a mechanical design drawing, it may be used to capture the characteristics of the image, depending on the application.
[0027] Furthermore, the conversion method may be URL encoding. URL encoding is widely used because it is often used for purposes such as directly embedding images in HTML, and it is easy to display the converted image again. In the information processing system 100, it is preferable that LLM be in an easily usable format, so it is not particularly selected as a conversion method for the conversation system. However, in applications where it is important to display the input image again on an image display device, URL encoding may be adopted.
[0028] Furthermore, although not a direct conversion method, there is also the JSON encoding method. This method involves Base64 encoding binary data and storing the result in JSON format. Many currently published open LLMs and closed LLMs that support various image modes support this method, and therefore can be suitably used in the information processing system 100.
[0029] In any case, the conversion method only needs to be in a format that the LLM that ultimately generates the response can interpret, and it can be selected in accordance with the learning method of the LLM used. Various encoding methods are publicly known, and conversion can also be performed at the input stage of various LLMs, so they should be selected considering the resources and quality such as response time. The important thing is that by converting image data into text data using these conversion methods, the context analysis unit 26 and the storage unit 16 can handle image input in the same way as normal conversation input.
[0030] The authentication unit 25 performs personal authentication of the user based on information from the fingerprint sensor 14. The authentication unit 25 functions as a sensor information input unit that acquires personal authentication information as sensor information from the fingerprint sensor 14. The authentication unit 25 compares the fingerprint of a user that has been registered in advance with the fingerprint acquired from the fingerprint sensor 14 and determines whether the acquired fingerprint is the fingerprint of a registered user. If the authentication unit 25 determines that the acquired fingerprint is the fingerprint of a registered user, it verifies the user's personal ID or the organization to which they belong, or grants access rights to necessary information.
[0031] The authentication unit 25 outputs the determination result to the context analysis unit 26. The context analysis unit 26 may change the personal information used for context analysis depending on whether the user is a registered user or a guest user.
[0032] Figure 1 shows a fingerprint sensor 14 as the personal authentication sensor, but other personal authentication sensors such as iris recognition sensors, voiceprint recognition sensors, and vein sensors may also be used. Furthermore, user authentication may be performed using an image of the user's face or their voice. For example, the authentication unit 25 may perform user authentication using iris recognition or voiceprint recognition. Of course, authentication may also be performed using multiple sensors.
[0033] (Context Analysis Unit 26) In this embodiment, the context analysis unit 26 has two main functions. The first function (hereinafter referred to as the first function) is to analyze the context as an AI conversation system, to determine the user's intent from past conversation history and the content of voice input, to generate a conversation response request along with necessary information (images, sensor information, etc.) based on the determination result, and to output it to the LLM determination unit 27. The second function (hereinafter referred to as the second function) is to analyze the user's intent and context, extract imaging conditions for the image to be input, and output the extracted imaging conditions to the imaging control unit 23. The "image to be input" here can be rephrased as "the image desired by the user." The second function also includes generating instructions for controlling the camera 12 based on the imaging conditions for the image to be input, and outputting these instructions to the imaging control unit 23. The details of the first and second functions will be explained below, in the order of the first function followed by the second function.
[0034] (First function of context analysis unit 26) As its first function, the context analysis unit 26 performs contextual analysis of a general conversation. At this time, the context analysis unit 26 may perform contextual analysis using past conversation history and / or the user's personal information. The context analysis unit 26 outputs a conversation response request that reflects the contextual analysis to the LLM determination unit 27.
[0035] The contextual analysis performed by the contextual analysis unit 26 may be, for example, an analysis using a small-scale language model that extracts keywords from the user's language input based on past conversation history and organizes the correlations between keywords. Such contextual analysis may include a process to determine the attributes of the language input by analyzing the keywords extracted from the user's language input using information from a database.
[0036] Language input attributes are simple tags corresponding to the content of the language input, such as questions, greetings, comments, requests for analysis, demands, or knowledge areas. The LLM determination unit 27 can select a response generation unit based on these tags.
[0037] If user authentication information (personal information) exists, the context analysis unit 26 may perform context analysis by using the user's own conversations and, if necessary, the conversation history of groups (family, business groups, etc.) from past conversation history. If the context does not involve personal information, the previous conversation history can be referenced. The conversation history contains the conversation user ID or conversation group ID and the conversation text itself, and the context analysis unit 26 may use this information to perform context analysis.
[0038] Furthermore, the context analysis unit 26 may determine whether or not to generate the aforementioned conversation response request based on the user's permissions before generating the request. For example, if the user's voice input question is one that the user does not have the authority to answer, the context analysis unit 26 may determine not to generate a conversation response request. Alternatively, the context analysis unit 26 may generate a request to produce a response such as "I cannot answer (with a reason depending on the situation)."
[0039] Thus, the context analysis unit 26 functions as a request generation unit that generates a conversation response request based on the user's language input. The conversation response request may include, in addition to the user's language input and the results of context analysis, image data converted into embedded data, past conversation history or experience information, or a combination thereof. The context analysis unit 26 outputs the generated conversation response request to the LLM determination unit 27. That is, the context analysis unit 26 may analyze the user's request based on the conversation history (conversation content and experience information), the current conversation input, and the input image, and send the analysis result to the LLM determination unit as a conversation response request.
[0040] The LLM determination unit 27 determines, based on the content of the conversation response request, which of the multiple response generation units to send the conversation response request to. The response generation units will be described later. The LLM determination unit 27 sends the conversation response request to at least one of the multiple response generation units. The LLM determination unit 27 may send the conversation response request to multiple response generation units. The context analysis unit 26 and the LLM determination unit 27 may be implemented in a single block.
[0041] The first LLM 51, the second LLM 52, and the simple response generation unit 28 function as response generation units that generate responses to the user's language input. The LLM determination unit 27 determines, based on the content of the conversational response request, which of the first LLM 51, the second LLM 52, and the simple response generation unit 28 to send the conversational response request to. In this disclosure, when it is not necessary to limit the type of response generation unit, the LLMs and simple response generation units are collectively referred to simply as response generation units. The number of LLMs is not limited to two, but may be three or more. Based on the determination of the LLM determination unit 27, multiple LLMs can be switched as needed to select the optimal LLM.
[0042] The simple response generation unit 28 may be a response generation unit that generates simple responses and is provided in a terminal equipped with an LLM determination unit 27. On the other hand, the first LLM 51 and the second LLM 52 may be response generation units that generate complex responses and are provided in a server or the like outside the terminal equipped with the LLM determination unit 27. Furthermore, the first LLM 51 and the second LLM 52 may be LLMs that differ in their ability to generate responses to conversational response requests. Specifically, the first LLM 51 and the second LLM 52 may be language models that differ in the type and amount of learned data, the maximum input size that can be processed, the response time, and / or the accuracy of the response.
[0043] The LLM determination unit 27 selects a response generation unit suitable for generating a response to a conversation response request according to the content of the conversation response request and sends the conversation response request. In addition, the LLM determination unit 27 may, for example, send a conversation response request containing confidential information that should not be sent to an external server to the simple response generation unit 28 instead of sending it to the external server.
[0044] Further, for example, when the first LLM 51 is in an internal server, the LLM determination unit 27 may select the first LLM 51 as the transmission destination of a conversation response request including confidential information that should not be sent to an external server. Each LLM may be tagged with attributes such as its location (affiliated organization) and / or handling authority for confidential information. The LLM determination unit 27 may refer to the attribute tags of each response generation unit in accordance with the content of the response request and select a response generation unit suitable for generating a response.
[0045] The response control unit 29 receives the response generated in response to the conversation response request and performs output control of the response. Specifically, when there are a plurality of responses, the response control unit 29 may control their output order. Also, when a plurality of responses can be integrated or summarized, the response control unit 29 may integrate or summarize the plurality of responses. Further, the response control unit 29 may generate a response message in which the response generated by the response generation unit is formatted to form a natural conversation with the user's voice input. Also, the response control unit 29 may output the generated response message to the conversation history management unit 31. The response control unit 29 outputs the generated response message to the response output unit 30.
[0046] For example, when there is a response from the simple response generation unit 28 and a response from the LLM, the response control unit 29 may prioritize and output the response from the simple response generation unit 28 first. Also, for example, when there is a tag restricting speech in the response from the LLM, or when tags such as characters not suitable for pronunciation, vital information, or attributes are included, the response control unit 29 may extract and format the speakable part and output it to the response output unit 30.
[0047] The response output unit 30 outputs the response message generated by the response control unit 29 via the output device 15. Also, when the context analysis unit 26 does not generate a conversation response request, the response output unit 30 may output a message indicating that it will not respond to the user's voice input.
[0048] (Second Function of Context Analysis Unit 26) Next, the second function of the context analysis unit 26 will be described. As the second function, the context analysis unit 26 analyzes the user's intention and context based on the user's language input, and extracts imaging conditions for the image to be input. The imaging conditions extracted in this way are converted into commands that can be understood by the imaging control unit 23. Hereinafter, the command may be referred to as an "operation command". The command is output together with the input image to the imaging control unit 23. If the imaging control unit 23 can directly acquire an image from the image input unit 22, the object output from the context analysis unit 26 to the imaging control unit 23 may be only the command. Although the conversion into a command is described in Japanese for convenience of explanation, appropriate coding such as a programming language may be executed. For example, when the user's intention is "Watch me while I cut vegetables from now on", the command may be "Turn the camera 12 downward so that my hands are in the frame".
[0049] Such a command "Turn the camera 12 downward" is an example of generating a command that the imaging control unit 23 can directly convert into a camera control command. For the purpose of explanation, this type of command is referred to as a "direct command".
[0050] Also, in another example, as will be described later, for the same user intention, a command specifying an imaging object such as "Center the vegetables in the frame" may be generated and output to the imaging control unit 23 equipped with the image analysis unit 24. For the purpose of explanation, this type of command is referred to as an "object command".
[0051] The configuration of the imaging control unit 23 and the context of the user language input may determine whether to generate direct commands or object commands. In the information processing system 100 equipped with the image analysis unit 24, object commands often yield more appropriate results. Furthermore, if the output device 15 allows the user to directly check the acquired image, direct commands may be generated by repeatedly using phrases such as "down further," "a little more to the right," or "zoom in a bit." Also, even after the camera 12 has been operated by object commands, fine adjustments may be made using direct commands. In other words, the context analysis unit 26 automatically determines the type of command, but if the language input from the user leans towards more direct commands, the possibility of generating more direct commands through context analysis increases.
[0052] The imaging control unit 23 converts the commands received from the context analysis unit 26 into actual camera control commands. "Converting the commands received from the context analysis unit 26 into actual camera control commands" means, as a concrete example, if a command is generated directly, changing the vertical angle to -20° in response to the command "downwards". Examples of commands output from the imaging control unit 23 to the camera 12 include commands to adjust the field of view (wide-angle, telephoto, etc.), commands to control the camera's orientation (pitch angle, yaw angle, roll angle, etc.), commands to change the frame rate (e.g., 1 fps to 120 fps), commands to control the shutter speed, commands to switch the image stabilization mode ON / OFF, and commands to change the resolution. Depending on the functional limitations of the camera 12, the camera 12 may accept only some of these commands, or it may accept commands other than those listed above. Furthermore, the imaging control unit 23 may be configured to acquire images from the image input unit 22 in real time. In other words, the imaging control unit 23 may be configured to receive only commands from the context analysis unit 26 and images from the image input unit 22. Furthermore, if the imaging control unit 23 uses actual camera images instead of text-based image data, the imaging control unit 23 may be configured to directly acquire images from the camera 12 in real time. The camera 12 changes the imaging conditions and takes images according to the commands from the imaging control unit 23.
[0053] The imaging control unit 23 controls the camera 12 based on the positional relationship between the user and the camera 12, and the user's language input. For example, as described above, if the user's language input was intended to mean "I'm going to cut vegetables now, so please watch me," the context analysis unit 26 inputs a command to the imaging control unit 23 to "point the camera 12 downwards so that the user's hands are visible." In response to this instruction, the imaging control unit 23 takes into account the positional relationship between the user and the camera 12 and outputs control information to the camera 12 to adjust the camera 12's orientation, field of view, imaging frame, shutter speed, imaging resolution, and / or whether or not to use image stabilization, so that the camera 12 captures the user's hands.
[0054] The image analysis unit 24 analyzes the image from the camera 12 acquired by the image input unit 22, in accordance with the response (object command from the context analysis unit 26) corresponding to the user's language input. In particular, the image analysis unit 24 performs object recognition in the image. Here, an object refers to the subject of imaging. For example, as mentioned above, if the user's intention is something like "I'm going to cut vegetables now, so please watch," the objects would be specified as "user's hand" (or "user's left hand" if the user is right-handed), "vegetables," "radish (vegetable)," "knife," "cutting board," etc. The types of objects that can be specified depend on the types of objects that the image analysis unit 24 supports, but this list of supported objects is shared with the context analysis unit 26, and it is preferable to select objects that have a high probability of being recognized. For object recognition, a high-performance MM-LLM (MultiModal Large Language Model) may be used, or, considering cost and judgment time, a small-scale AI such as a CNN or RNN specialized for object recognition may be used. For example, using a high-performance MM-LLM might enable sophisticated recognition, such as "vegetables being cut by the user," but such recognition often takes time, and this delay can be a problem when the purpose is to monitor potentially dangerous tasks. On the other hand, small-scale AIs such as CNNs and RNNs specialized in object recognition require specifying specific and limited objects, such as "kitchen knife," but they can perform recognition at high speed. The objects that can be recognized are those that have been pre-trained to suit various use cases, but as mentioned above, it is preferable that the corresponding object list is shared with the context analysis unit 26 and that appropriate recognition targets are selected.
[0055] When using a small-scale AI for object recognition, object recognition can be easily performed using only a portion of the resources of the edge device. In this embodiment, it is not necessary to recognize objects with high accuracy; for example, it is sufficient to detect the presence of the target object (e.g., a human face, a human hand, a red board (sign), a dog, a cat, a car, a sign, etc.). The image analysis unit 24 outputs the area occupied by the object as rectangular coordinates on the screen, which allows for the determination of the direction in which the orientation of the camera 12 should be corrected, the size of the object within the imaging range, and so on. In other words, by using an object-recognition-capable AI, it is possible to determine whether the target object exists within the imaging range, and if so, in which direction from the center of the camera 12 it is located, and the relative size of the object to the imaging range. In this embodiment, the object recognition AI is described as having a fixed range of objects that it can recognize. This is because an AI capable of handling many general-purpose objects would be large in size and processing time, which is undesirable for operation on edge devices. The information processing system 100 may also have a function to download and update AI capable of recognizing necessary objects from the cloud. Furthermore, the information processing system 100 may include an interface for explicitly specifying the object to be recognized and updating the AI.
[0056] In this way, the image analysis unit 24 detects the object to be captured in the image that corresponds to the user's language input. The image analysis unit 24 outputs information to the imaging control unit 23 indicating the presence of the object to be captured in the image, the position of the object to be captured, and the size of the object to be captured. Using this information, the imaging control unit 23 can directly convert object commands into commands.
[0057] Furthermore, if the image analysis unit 24 does not find an object to be imaged within the imaging range, it may generate a message such as "The target ○○ (e.g., vegetables) was not found within the imaging range" and output it to the response output unit 30.
[0058] Although the output path is not shown in the diagram, the image analysis unit 24 may also output to the context analysis unit 26 that there is no object to be imaged. In this case, the context analysis unit 26 may, upon receiving this result, specify a new object, or generate a conversational response request to ask the user to generate a response such as, "The target shooting range could not be determined. Could you please specify what you would like to monitor in more detail?" or "Adjusting the camera position. Please place the vegetables on the cutting board."
[0059] The imaging control unit 23 controls the field of view, orientation, etc. of the camera 12 to capture an image that matches the user's intention, based on information indicating the presence, position, and size of the object to be imaged detected by the image analysis unit 24, and commands received from the context analysis unit 26. Specifically, the imaging control unit 23 converts object commands into direct commands based on the information from the image analysis unit 24. Furthermore, the direct commands are converted into more specific operation commands in accordance with the control system of the camera 12 itself and output to the camera 12.
[0060] Furthermore, the imaging control unit 23 may be equipped with a function that allows it to ignore the object recognition function and follow direct user instructions, such as "turn 30 degrees to the left."
[0061] Furthermore, the imaging control unit 23 may adjust the camera 12's field of view, orientation, etc., at least once based on the position or size of the object to be captured within the camera 12's imaging range. For example, if the imaging control unit 23 receives a command to "point the camera 12 downwards so that the user's hands are in the frame," it may control the camera 12 considering the positional relationship between the user and the camera 12, and then control the camera 12 again so that the user's hands, detected by the image analysis unit 24, are in the center of the imaging range.
[0062] The camera 12 is controlled by the imaging control unit 23 only needs to be able to properly image the target object, and the imaging control unit 23 may only adjust the size of the target object and the image of the target object. By utilizing object recognition in this way, feedback between imaging conditions and imaging results becomes possible, and by repeating this control, it becomes possible to accurately capture images that match the user's intentions.
[0063] Before and / or after the imaging control unit 23 controls the camera 12, the image analysis unit 24 may analyze the image captured by the camera 12 after the control to detect the object to be captured from the image, and may output information indicating the presence of the object to be captured in the image, the position of the object to be captured, and the size of the object to the imaging control unit 23 once or more times.
[0064] The image analysis unit 24 may acquire images directly from the image input unit 22. Alternatively, the image analysis unit 24 may be provided as one of the functions of the imaging control unit 23.
[0065] The image analysis unit 24 may analyze the image acquired from the camera 12 based on the response corresponding to the user's language input and inform the user of the analysis result via the response output unit 30. Furthermore, if the image analysis unit 24 finds that any of the multiple still images (or videos) acquired continuously (in other words, over time) from the camera 12 contain an image (or frame in the case of a video) that satisfies the conditions indicated by the language input, it may inform the user of this fact via the response output unit 30. For example, the image analysis unit 24 may output a message via the response output unit 30 such as, "Your left hand (the specified object) is currently in front of the camera. Is this camera position acceptable?" Needless to say, the specified object in this embodiment is an object determined by the context analysis unit 26, and not one specified by the user in language input. Therefore, a camera angle different from the user's intention may be set. Consequently, the user can provide language input for correction, such as, "Hey, that's wrong. Show the radish," or "Having my left hand in the center isn't good. Show a little more to the right." The context analysis unit 26 can modify direct commands to the imaging control unit 23 in response to new language input.
[0066] Furthermore, when a user asks a question such as "What is visible in front?" during language input, the context analysis unit 26 determines that this is a question for image analysis, not an imaging instruction. Therefore, the context analysis unit 26 creates a conversational response request such as "Please briefly describe what is in the image (+ embedded data of the input image)," which is then determined by the LLM capable of image analysis, and a description of what is currently being captured is output from the output device 15. This allows the user to check the current imaging status of the camera 12 at any time.
[0067] Furthermore, if the image analysis unit 24 utilizes a high-performance MM-LLM, the image analysis unit 24 can be used to address the above-mentioned questions. However, the image analysis function becomes more advanced and general-purpose, making it difficult to execute on edge devices. These choices can be adjusted during system design according to available resources and objectives, and may be adjusted as appropriate. However, as already mentioned, from the standpoint of cost and response time, it is preferable to use a small-scale, limited object recognition AI for the AI used in the image analysis unit 24.
[0068] In this embodiment, the context analysis unit 26 has been described as sending a command to the imaging control unit 23, but the system is not limited to this. For example, the context analysis unit 26 may send user input to the LLM determination unit 27 along with tags such as "Destination: Simple response generation unit 28, Mission: Generate conditions for camera control, Speech: Do not speak". In this case, the simple response generation unit 28 will send a command to the imaging control unit 23. Converting the user's intentions regarding the camera 12 obtained from the context into camera-controllable parameters and commands is a relatively simple text conversion task and can be realized with a small-scale language model or rule-based processing, so the command to the imaging control unit 23 does not necessarily have to be performed by the context analysis unit 26.
[0069] The commands to be sent to the imaging control unit 23 may be generated by the LLM, for example, the first LLM 51. In this case, the first LLM 51 is assumed to have capabilities suitable for camera imaging, such as "image analysis and camera control," and to be tagged with attributes.
[0070] When the context analysis unit 26 determines from the context that new adjustments to the camera 12 are needed in response to user input, it sends a conversational response request consisting of tags such as "Destination: LLM, Mission: Generate conditions for camera control, Speech: None" + "Watch me as I cut vegetables" + user input. The LLM determination unit 27 uses tags such as camera control to select and send the first LLM 51 as the appropriate LLM. In response to the conversational response request, the first LLM 51 can infer the optimal object selection, display position, size, etc. from the performance information and input of the image analysis unit 24 and generate an object command.
[0071] Thus, the response generation unit may generate the command to the camera 12. In this case, although it differs from Figure 1, the generated command may be sent to the imaging control unit 23 via the response control unit 29.
[0072] The conversation history management unit 31 may record and manage authentication information and input image information as additional information in the conversation history, in addition to the conversation text which includes language input by the user and the response to that language input. This additional information may be managed by appropriate tagging. For example, if there is no change in any of the additional information from the previous time, or if the context analysis unit 26 does not need any of the additional information, the conversation history management unit 31 does not need to record that additional information in the conversation history. As another example, for information that is accessible to anyone, the conversation history management unit 31 does not need to record authentication information in the conversation history. As yet another example, when a normal conversation is taking place, the conversation history management unit 31 does not need to record input image information in the conversation history. As yet another example, when a normal conversation is taking place but the input image information changes suddenly, the conversation history management unit 31 may add and record the input image information in the conversation history. In this way, the conversation history management unit 31 may select the input image information, etc. to be managed and recorded based on the results of the context analysis unit 26. By appropriately managing additional information in this way, the conversation history management unit 31 can manage the conversation history appropriately without wasting memory resources used for it.
[0073] The conversation history is recorded in the storage unit 16. The conversation history management unit 31 can access the storage unit 16 as needed and pass the conversation history to the context analysis unit 26. Furthermore, the conversation history management unit 31 may include a description of the image generated by the response generation unit in the conversation history.
[0074] Furthermore, the conversation history management unit 31 may modify the database based on the conversation history. In other words, the conversation history management unit 31 may function as a database modification unit that modifies the database based on the user's conversation history.
[0075] Furthermore, past conversation history becomes a large amount of data as the usage time and frequency of the information processing system 100 increase. For this reason, saving all conversation history is undesirable from the standpoint of increasing memory resources and data processing time. Of course, memory capacity and memory access speed are still improving year by year, and it is highly likely that in the future it will be possible to record virtually all conversations, if not all, that are actually used. However, at the time the system is implemented, the available resources are finite, and it is desirable to be able to efficiently manage the conversation history with finite resources. Therefore, the conversation history management unit 31 may have a function to maintain the data stored in the storage unit 16 at an appropriate size. There are several methods for maintaining an appropriate data size for the conversation history, and any of these methods can be applied to the information processing system 100.
[0076] The simplest method is to set a predetermined limit on the data size of the conversation history and delete older data when the limit is exceeded. This method is easy to implement and reliably reduces the size. However, this method has the problem of not being able to maintain consistency in conversations, especially with older information, because it automatically deletes old information. Therefore, this method is suitable for applications where it is sufficient to maintain consistency in conversations over a relatively short period, such as one day's worth of data.
[0077] Another method involves setting a limit on the data size of the conversation history and using a separate LLM (which can be the same model) from the one used for conversations to summarize it, thereby reducing the data size to a predetermined amount. This method increases the proportion of useful conversations in the record by removing meaningless conversation history through summarization, allowing important conversations to be retained for a relatively long period, even if they are old.
[0078] Furthermore, a suitable method for the information processing system 100 is to utilize RAG (Retrieval-Augmented Generation). RAG is an approach that combines information retrieval and generation models to generate more accurate answers to user questions. The following briefly explains the basic operation of RAG, from the creation of embedded data to the search method.
[0079] 1. Creating Embedded Data First, the conversation history is divided into data for a predetermined upper size or a predetermined period, for example, one day's worth of data, and each is summarized. The summarized individual data is converted into vectors (embedded data) using a known embedding model, such as a pre-trained model like BERT, RoBERTa, or Sentence-BERT. At this time, an appropriate index may be constructed. In addition, image data can be separated during the summarization and indexing process and stored, for example, on the cloud where the image data can be referenced by the index, or on another inexpensive, high-capacity medium. This improves memory utilization efficiency and allows for longer-term storage. Furthermore, by including a brief description of the separated images in the summary, older images can be recalled from the conversation as needed.
[0080] 2. Processing User Questions: User questions are converted into vector data using the same model as the historical data. Through methods such as dot product calculations, the top three most similar historical data points are extracted and sent to the context analysis unit 26 and the response generation unit (LLM). Furthermore, if no data showing a similarity above a predetermined level is found, it can be omitted. This prevents the generation of incorrect answers due to being influenced by irrelevant information. This ensures that even with large historical data sets, context analysis and response generation can utilize historical data that is practical in terms of data size and relevant to the question.
[0081] Furthermore, the conversation history management unit 31 may delete the vectorized data itself in order to manage the size of the history data. As already mentioned, the deletion method may be to delete the oldest data first, or to appropriately resummarize the data. Even in this case, important data can be retained for a much longer period than when this is done for non-vectorized information.
[0082] This method of using RAG makes it possible to retain the contents of conversation history data relatively accurately and over a long period of time, and is particularly suitable for use in the information processing system 100. However, in order for RAG to function effectively, a certain limit on the number of history entries and memory capacity is required. Therefore, depending on the application and available resources, one should choose to use RAG or another method. In addition, there are several known methods for properly managing memory, and these may be used. Furthermore, multiple methods for managing history may be used in combination.
[0083] In either method, the conversation history will be organized at appropriate times, taking advantage of breaks in the conversation. However, in order to maintain the consistency of the most recent conversation, it is preferable that the most recent conversation history, for example, the last 10 turns, is neither compressed through summarization nor deleted.
[0084] The experience information generation unit 32 generates experience information that associates the user's vital information at a certain point in time with information indicating the user's condition at the time the vital information was acquired. Examples of information indicating the user's condition include location information, images, and time. The conversation history management unit 31 records and manages the experience information in the storage unit 16, including it in the conversation history. That is, the conversation history management unit 31 may function as an experience information management unit that manages vital information as a result of determining the user's mental state, together with information indicating the user's condition, as experience information. The conversation history management unit 31 may also manage the experience information as part of the conversation history. Such vital information may be acquired by, for example, a vital sensor (not shown). Examples of vital information include body temperature, pulse rate, blood pressure, sweating, and respiratory rate.
[0085] Alternatively, the experience information generation unit 32 may directly store the generated experience information in the storage unit 16 and manage the experience information itself. In other words, the experience information generation unit 32 may function as an experience information management unit.
[0086] Experiential information refers to information that indicates, for example, when (time), where (location information), what the user saw (image), and what mental state they were in. Experiential information may also include information identifying the subject in an image or a description of the image.
[0087] The experience information generation unit 32 is not an essential component and may be omitted from the information processing system 100.
[0088] (Example of processing flow) Figure 2 is a flowchart showing an example of processing (information processing method) in the information processing system 100.
[0089] In the information processing system 100, the voice input unit 21 receives voice input from the user (S11). When the voice input unit 21 receives voice input, the context analysis unit 26 acquires the conversation history with the user (S12) and uses the conversation history to analyze the context of the input voice (S13). Based on the context analysis results, the context analysis unit 26 generates a conversation response request (S14). The function of the context analysis unit 26 described here corresponds to the first function described above.
[0090] Based on the content of the conversation response request, the LLM determination unit 27 selects a response generation unit and sends the conversation response request (S15). The response generation unit selected by the LLM determination unit 27 generates a response to the conversation response request (S16). Based on the content of the response generated by the response generation unit, the response control unit 29 generates a response message to output to the user (S17). The conversation history management unit 31 updates the user's conversation history with the response message generated by the response control unit 29 (S18). The response output unit 30 outputs the response message to the user via the output device 15 (S19). Note that the conversation history update (S18) may occur after the output of the response message (S19).
[0091] Furthermore, the context analysis unit 26 may omit the processing in step S12 (i.e., without acquiring the conversation history) and analyze the context of the input speech without using the conversation history in the processing of step S13. For example, when the information processing system 100 is first used, there may be no conversation history. Moreover, there may be cases where the user explicitly states that they are about to start a new conversation, such as saying "By the way, changing the subject," "Let's talk about something new," or "Putting aside what we've talked about so far," and there is no need to refer to past information. Similarly, in the processes shown in the flowcharts of Figures 3 and 4 described later, the processing in steps S22 and S32 may be omitted (without acquiring the conversation history), and the context of the input speech may be analyzed without using the conversation history in the processing of steps S23 and S33.
[0092] (Other examples of processing flow) Next, with reference to the flowcharts in Figures 3 and 4, another example of processing in the information processing system 100 will be shown. The difference between the processing shown in Figures 3 and 4 and the processing shown in Figure 2 lies in the function of the context analysis unit 26. In the processing shown in Figures 3 and 4, the second function of the context analysis unit 26 will be explained.
[0093] In step S21, the voice input unit 21 receives voice input from the user. Processing proceeds to step S22, where the context analysis unit 26 acquires the conversation history with the user. Processing proceeds to step S23, where the context analysis unit 26 uses the conversation history with the user to analyze the context of the input voice. Here, let's assume that as a result of processing in step S23, the user's intention is analyzed to be something like, "I'm going to cut vegetables now, so please watch me." In other words, the context analysis unit 26 determines that it is a mission that requires the control of the camera 12 to track a specific object, and decides that it is necessary to generate an operation command for the camera 12. Based on the content of the mission and the corresponding object list of the image analysis unit 24, an object command is generated, and the target object is set to, for example, "the user's hand."
[0094] In this case, in step S24, the context analysis unit 26 extracts imaging conditions for the camera 12 based on the user's language input and generates operation commands for the camera 12 based on the extracted imaging conditions. The process proceeds to S25, where the imaging control unit 23 controls the camera 12 based on the operation commands generated in step S24. Specifically, the imaging control unit 23 takes into account the positional relationship between the user and the camera 12 and outputs control information to the camera 12 to adjust the camera 12's orientation, field of view, imaging frame, shutter speed, imaging resolution, and / or whether or not to use image stabilization, so that the camera 12 images the user's hands.
[0095] The process proceeds to step S26, where the image analysis unit 24 acquires the image captured in step S25 and detects the "user's hand," which is the object to be captured, from the image. The image analysis unit 24 may output a camera control command to the image capture control unit 23 so that the "user's hand" in the image is positioned in the center of the imaging range (screen). Alternatively, the image analysis unit 24 may output a camera control command to the image capture control unit 23 so that the "user's hand" occupies about half of the imaging range.
[0096] The process proceeds to step S27, where the imaging control unit 23 controls the camera 12 based on the camera control command from the image analysis unit 24, either so that the "user's hand" is positioned in the center of the imaging range, or so that the "user's hand" occupies about half of the imaging range. Either one of these controls may be performed, or both may be performed. Furthermore, the processes in steps S26 to S27 may be performed multiple times as needed.
[0097] The image analysis unit 24 may detect basic abnormalities while monitoring. If an abnormality is detected (YES in S28), the image analysis unit 24 notifies the user of a warning message via the response output unit 30 (S29). An example of a "basic abnormality" is when the direction of the camera 12 suddenly changes or when an obstacle is placed. When such a basic abnormality is detected, the image analysis unit 24 may notify the user via the response output unit 30 that the target object can no longer be tracked within the imaging range of the camera 12. The notification in this case may contain a message such as, "The camera's state has changed, and it is no longer possible to photograph the target object. Please reset the shooting conditions." The system may be configured to periodically perform the processes in steps S26 to S27 in order to respond to such basic abnormalities.
[0098] The context analysis unit 26 modifies and stores a basic command to the response generation unit so that it can warn of danger without relying on the user's voice input when the "user monitoring" mission is issued. The modified basic command is something like, "Analyze the given image, determine if a person is in danger or in a dangerous situation, and issue a warning if there is danger. If there is no danger, respond with null." The basic command is modified and stored when a monitoring mission is in progress.
[0099] Furthermore, the context analysis unit 26 may acquire images from the image input unit 22 every 10 seconds, for example, using an interval timer. The timing of image acquisition should be determined considering the content of the mission and the overall processing capacity of the system. For example, in a mission such as "Tell me if you see a beautiful flower while walking," images may be acquired every minute. For example, in a mission such as "Tell me if it's dangerous while cutting vegetables," images may be acquired every second. A request for image acquisition using an interval timer is generated even if there is no voice language input from the user. Although an interval timer was used here, any sensor system that can control image input at the necessary timing and notify the context analysis unit 26 can be similarly used. The LLM determination unit 27 calls the first LLM 51 in response to this request. The first LLM 51 is an LLM with attributes such as "image analysis, high-speed response, and simple response tendency."
[0100] The first LLM51 generates responses such as "Your finger position is dangerous," "You're holding the knife incorrectly," or "You've cut your finger. Please follow the instructions for stopping the bleeding. We will contact you again in case of emergency." If no dangerous situation is detected, it does not need to respond. In other words, if the system does not point out a problem, the user can continue to concentrate on their work without being distracted by voice prompts.
[0101] Figure 3 illustrates the case where the user's intention is something like, "I'm going to cut vegetables now, so please watch me." Figure 4 illustrates the case where the user's intention is something like, "Please let me know when the display on that signage changes." The processes in steps S31 to S32 shown in Figure 4 are the same as the processes in steps S21 to S22 shown in Figure 3, so we will omit the explanation and start the explanation from step S33 in Figure 4.
[0102] As a result of the processing in step S33, it is determined that the user's intention is something like, "Please let me know when the display on that signage changes."
[0103] In step S34, the context analysis unit 26 extracts imaging conditions for the camera 12 based on the user's language input and generates operation commands for the camera 12 based on the extracted imaging conditions. The process proceeds to step S35, where the imaging control unit 23 controls the camera 12 based on the operation commands generated in step S34. Specifically, the imaging control unit 23 takes into account the positional relationship between the user and the camera 12 and outputs control information to the camera 12 to adjust the camera 12's orientation, field of view, imaging frame, shutter speed, imaging resolution, and / or whether or not to use image stabilization, so that the camera 12 can image the signage.
[0104] The process proceeds to step S36, where the image analysis unit 24 acquires the image captured in step S35 and detects the "signage" that is the target of the image capture from the image. The image analysis unit 24 may output a camera control command to the image capture control unit 23 so that the "signage" shown in the image is located in the center of the image capture range (screen). Alternatively, the image analysis unit 24 may output a camera control command to the image capture control unit 23 so that the "signage" occupies most of the image capture range, for example, about 80%.
[0105] The process proceeds to step S37, where the imaging control unit 23 controls the camera 12 based on the camera control command from the image analysis unit 24, either so that the "signage" is positioned in the center of the imaging range, or so that the "signage" occupies about half of the imaging range. Either one of these controls may be performed, or both may be performed. Furthermore, the processes in steps S36 to S37 may be performed multiple times as needed.
[0106] In the process shown in Figure 4, similar to the process in Figure 3, the context analysis unit 26 may acquire the signage image from the image input unit 22 every 5 seconds, for example, using an interval timer. Let's assume that the first LLM 51 is called, similar to the process in Figure 3. When the first LLM 51 detects a change in the signage display (YES in S38), it generates a response such as "The signage display has changed" (S39). If no change in the signage display is detected, it does not need to respond.
[0107] It is also possible that the user's intention is "to be notified when the signage (traffic light) in front changes from red to blue." A typical object recognition AI can simultaneously check both "red signage" and "blue signage," and can detect when the red signage disappears from the imaging range and the blue signage appears. In other words, such detection can be performed by the image analysis unit 24 equipped with a typical object recognition AI. This detection of state change may be used as the processing in step S38. An example of an output message in this case would be "The signage color has turned blue" (S39).
[0108] As explained in Figures 3 and 4, instructions for centering the target object in the camera 12 include camera control such as (1) detecting the target object, (2) controlling the camera's orientation so that the target object is centered, and (3) adjusting the target object to an appropriate size according to the purpose. However, this is not the only option. For example, it may be sufficient to simply specify the target object to be detected and its size, or the operation may be specified in detail.
[0109] (Effects) As described above, according to one embodiment of the present disclosure, the following effects can be obtained.
[0110] The information processing device in this embodiment is communicatively connected to a camera 12 that moves in conjunction with the user's movement. The information processing device includes a language input unit that receives language input from the user, and an imaging control unit 23 that controls the camera 12 based on the positional relationship between the user and the camera 12, and the language input. The "language input unit" referred to here corresponds to the voice input unit 21 described above.
[0111] The functions described in the above-mentioned conventional technology presuppose that the subject is already within the shooting range, such as capturing the subject through the viewfinder at the initial stage of imaging, and do not capture the subject based on vague instructions from the user. In contrast, with the above configuration, the user's intent and context are analyzed based on the user's language input, and the camera 12 is controlled based on the analysis results. This makes it possible to capture an image that matches the user's intent, that is, the object desired by the user, even if the user cannot confirm the camera 12's imaging screen with a viewfinder or monitor. In other words, with the above configuration, the information processing device in this embodiment can autonomously capture a subject in response to a conversation with the user.
[0112] The information processing device may also include a context analysis unit 26 that extracts imaging conditions for the camera 12 to image the target based on language input and generates operation commands for the camera 12 based on the imaging conditions. The imaging control unit 23 may also control the camera 12 based on the positional relationship between the user and the camera 12, and the operation commands generated by the context analysis unit 26.
[0113] If the user's intention is something like, "I'm going to cut vegetables now, so please watch me," the context analysis unit 26 determines that the object to be imaged is the "user's hand" and extracts the imaging conditions necessary to image the "user's hand." Based on the extracted imaging conditions, the context analysis unit 26 outputs an operation command to the imaging control unit 23, such as "point the camera 12 downwards so that the hand is visible." The imaging control unit 23 converts the operation command received from the context analysis unit 26 into an actual camera control command. For example, in response to the operation command "downwards," the imaging control unit 23 changes the vertical angle to -20°. This makes it possible to image the "user's hand" as desired by the user.
[0114] The information processing device may also further include an image analysis unit 24 that analyzes images acquired from the camera 12 based on language input. The imaging control unit 23 may control the camera 12 based on the positional relationship between the user and the camera 12 and language input, as well as the analysis results from the image analysis unit 24.
[0115] When the image analysis unit 24 detects the "user's hand," which is the object to be captured, from the image captured by the camera 12, it also detects the position information of the "user's hand" within the imaging range (screen). The image analysis unit 24 may output a camera control command to the imaging control unit 23 so that the "user's hand" in the image is located in the center of the imaging range. Alternatively, the image analysis unit 24 may output a camera control command to the imaging control unit 23 so that the "user's hand" occupies about half of the imaging range. By controlling the field of view, orientation, etc., of the camera 12, the imaging control unit 23 can capture the "user's hand" in the center of the imaging range and / or occupy about half of the imaging range. With such imaging, it becomes possible to appropriately capture the object desired by the user (in this case, the "user's hand").
[0116] Furthermore, the image analysis unit 24 may notify the user via the output unit if the image acquired from the camera 12 satisfies the conditions indicated by the language input. The "output unit" here corresponds to the response output unit 30 described above. The "conditions indicated by the language input" here means the user's intention. Also, "satisfying the conditions indicated by the language input" means, for example, if the user's intention was something like "I'm going to cut vegetables now, so please watch," then the image is of the "user's hand."
[0117] With the above configuration, the user can confirm that the desired object (in this case, "the user's hand") is being captured in the image.
[0118] Furthermore, the image analysis unit 24 may, in response to user instructions, inform the user via the output unit of information indicating the target being captured by the camera 12 at the time the instruction was given.
[0119] As previously explained, the system analyzes the user's intent and context to capture the object the user desires. However, if the user's intent is misinterpreted, there is a risk of capturing an object other than the one the user desires. Therefore, by informing the user of information indicating the target object being captured at the time of the instruction (for example, a voice input such as "check camera"), the user can confirm whether or not the desired object is being captured. The method of notification is not particularly limited and may be by voice or on the screen, but since it is assumed that the user may not be looking at the display screen, notification by voice is preferable.
[0120] Furthermore, the context analysis unit 26 may generate a response request (conversation response request) based on the user's language input. The information processing device may also include a response control unit 29 that receives a response generated in response to a response request generated by the context analysis unit 26 and controls the output of the response.
[0121] According to the above configuration, responses are made that take into account the images the user is interested in, making it possible to provide useful information or draw attention to the user through conversation.
[0122] Furthermore, if the language input from the user satisfies predetermined conditions, the context analysis unit 26 may modify the response request (conversation response request) regardless of whether or not there is subsequent language input from the user. An example of "when the language input from the user satisfies predetermined conditions" is when the user intends to monitor a potentially dangerous task. An example of a "potentially dangerous task," in this embodiment, is cutting vegetables with a knife. As described above, when the "user monitoring" mission is issued, the context analysis unit 26 may modify the basic command to the response generation unit so that it can warn of danger without relying on the user's voice input. The response control unit 29 may then receive the response generated in accordance with the modified response request and control the output of the response. An example of such a response output, as described above, would be something like, "Your finger is in a dangerous position," "You're holding the knife incorrectly," or "You've cut your finger. Please follow the instructions for stopping the bleeding. We will also make an emergency call."
[0123] With the above configuration, during monitoring missions, the system determines whether or not a danger has occurred regardless of whether or not the user has provided language input. Therefore, even if the user is focused on their work and forgets instructions, the system can alert the user if a danger has occurred. On the other hand, if no danger has occurred, no special response is given, allowing the user to continue working with concentration.
[0124] [Embodiment 2] Embodiment 2 of the present disclosure will be described below. For the sake of convenience of explanation, components having the same function as those described in Embodiment 1 above will be denoted by the same reference numerals, and their descriptions will not be repeated.
[0125] Figure 5 is a block diagram illustrating the configuration of the information processing system 200 according to Embodiment 2. As shown in Figure 5, the information processing system 200 includes an information processing device 210 and a server 220.
[0126] The information processing device 210 is a conversational terminal that includes elements other than the first LLM 51 and second LLM 52 in the information processing system 100. In particular, in the information processing device 210, the microphone 11, camera 12, and fingerprint sensor 14 are integrated into a single device. This allows for the acquisition of accurate information about the user with high precision. Furthermore, the integration of these input devices reduces the possibility of leakage of the user's personal information.
[0127] Server 220 includes the first LLM 51 and the second LLM 52 in the information processing system 100. Server 220 may be a publicly accessible server or a secure server managed by an individual or organization. In Figure 5, a single server 220 includes the first LLM 51 and the second LLM 52. However, in the information processing system 200, the server containing the first LLM 51 and the server containing the second LLM 52 may be separate entities.
[0128] If server 220 manages multiple LLMs located inside or outside of server 220, server 220 may have a function to call publicly accessible servers from among the LLMs under its management. For example, if the first LLM 51 is inside server 220 and the second LLM 52 is on another server, server 220 may use the second LLM 52 via the other server.
[0129] Server 220 may have a function to switch the LLMs it manages. In this case, server 220 may manage attribute tags of the LLMs it manages and share this attribute information with the LLM determination unit 27. This allows the LLM determination unit 27 to obtain information to select the appropriate LLM from the currently available LLMs. Server 220 may be configured to own or obtain from another database information about switchable LLMs and servers containing LLMs.
[0130] In the example shown in Figure 5, the voice input unit 21, image input unit 22, image analysis unit 24, imaging control unit 23, authentication unit 25, context analysis unit 26, LLM determination unit 27, simple response generation unit 28, response control unit 29, response output unit 30, conversation history management unit 31, and experience information generation unit 32 are all located on the same information processing device 210. These units may be integrated on a chip and combined into a single information processing device 210. Alternatively, these units may be distributed and located on a secure server or the like. The information processing system 200 may be implemented as a program executed by computers provided on the information processing device 210 and the server 220.
[0131] Figure 6 is a perspective view illustrating a terminal device 230. The terminal device 230 is a portable electronic device equipped with an information processing device 210.
[0132] As shown in Figure 6, the terminal device 230 includes a mounting portion 235. The mounting portion 235 is a component for attaching the terminal device 230 to the user's neck. The mounting portion 235 has an open ring shape that can be hooked onto the user's neck. With this configuration, the user's movements are less likely to be hindered even when the terminal device 230 is attached.
[0133] The attachment portion 235 has a microphone 11. Specifically, the microphone 11 is located at one end of the attachment portion 235, which has an open ring shape. This position is near the user's mouth when the user is wearing the terminal device 230 around their neck. Therefore, the user can easily input voice via the microphone 11 while wearing the terminal device 230 around their neck.
[0134] The attachment portion 235 has a camera 12. Specifically, the camera 12 is positioned near the end of the attachment portion 235, which has an open ring shape, so as to face outwards. As a result, the orientation of the camera 12 substantially coincides with the orientation of the user's face when the user is wearing the terminal device 230 around their neck. Therefore, it is possible to capture images of what the user sees with the camera 12. Such a camera 12 will move as the user moves.
[0135] The attachment portion 235 has a fingerprint sensor 14. Specifically, the fingerprint sensor 14 is located at the end of the attachment portion 235, which has an open ring shape, opposite to the end where the microphone 11 is located. This position makes it easy for the user to touch the fingerprint sensor 14 with their finger when the terminal device 230 is attached to their neck. Therefore, the user can easily perform personal authentication.
[0136] The attachment portion 235 has an output device 15. Specifically, speakers, which serve as output devices 15, are positioned on both the left and right sides of the attachment portion 235, which has an open ring shape, when the open ring portion is facing forward. This position is near the user's ears when the user is wearing the terminal device 230 around their neck. The user can easily hear the output from the output devices 15 while wearing the terminal device 230 around their neck.
[0137] Thus, the terminal device 230, which is equipped with an information processing device 210 integrally formed with the camera 12 and worn around the user's neck, can continue to image the object desired by the user even if the user moves. As a result, the user can obtain an image of the desired object with just one instruction, even if they move to a different location after the instruction.
[0138] In addition, user instructions regarding the operation of the terminal device 230 may be input to the terminal device 230 via another terminal device, such as a smartphone.
[0139] Furthermore, the terminal device 230 may be provided with a mechanism that allows the camera 12 to be activated, paused, etc., by tap input.
[0140] [Other Embodiments] The lens portion of the camera 12 may be equipped with an LED laser (Light Emitting Diode). The LED laser is installed or controlled to emit laser light in the imaging direction of the camera 12. As a result, a laser marker is projected onto the imaging center that the camera 12 is currently imaging, so that the user can know whether or not the camera 12 is imaging an object of their choice.
[0141] In the above-described embodiment, a method was explained in which the user's intent and context are analyzed based on the user's language input, and an image of the object desired by the user is captured. However, such a camera control function based on contextual analysis may be configured to be ON / OFF. When the camera control function is set to OFF, the imaging control unit 23 may control the camera 12 to capture an image from a fixed viewpoint, regardless of the user's language input. This is useful, for example, in scenes where it is desired to collect a video log.
[0142] In the terminal device 230 described above, the camera 12 and the information processing device 210 were described as being integrally formed, but the device is not limited to this, and the camera 12 and the information processing device 210 do not have to be integrally formed. For example, as described above, if the camera 12 is installed on a drone, it is sufficient that the image input unit 22, the imaging control unit 23, etc., and such an external camera are connected by a predetermined interface. The interface is not particularly limited, but wireless communication such as Wi-Fi® or Bluetooth® may be used.
[0143] Furthermore, users may indirectly specify the object they desire as the subject of the photograph. For example, a user can indirectly specify the object they desire as the subject by giving a voice command such as, "Take a photograph of what is shown in the area I am pointing to."
[0144] [Example of implementation by software] The functions of an information processing device (hereinafter referred to as "device") can be realized by an information processing program that causes a computer to function as the device, and which causes a computer to function as each control block of the device (particularly the voice input unit 21, image input unit 22, imaging control unit 23, image analysis unit 24, authentication unit 25, context analysis unit 26, LLM determination unit 27, simple response generation unit 28, response control unit 29, response output unit 30, conversation history management unit 31, and experience information generation unit 32).
[0145] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., memory) as hardware for executing the program. By executing the program using this control device and storage device, the functions described in each of the embodiments are realized.
[0146] The above program may be recorded on one or more computer-readable recording media, not temporary ones. These recording media may or may not be provided by the above device. In the latter case, the program may be supplied to the above device via any wired or wireless transmission medium.
[0147] Furthermore, some or all of the functions of each of the above control blocks can also be implemented by logic circuits. For example, an integrated circuit in which logic circuits functioning as each of the above control blocks are formed is also included in the scope of this disclosure. In addition, it is also possible to implement the functions of each of the above control blocks by, for example, a quantum computer.
[0148] Furthermore, each process described in the above embodiments may be performed by AI (Artificial Intelligence). In this case, the AI may operate on the control device described above, or it may operate on another device (for example, an edge computer or a cloud server).
[0149] This disclosure is not limited to the embodiments described above, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of this disclosure. Furthermore, new technical features can be formed by combining the technical means disclosed in each embodiment.
[0150] 14 Fingerprint sensor 21 Voice input unit (language input unit) 22 Image input unit 23 Image capture control unit 24 Image analysis unit 25 Authentication unit 26 Context analysis unit 27 LLM determination unit 28 Simple response generation unit 29 Response control unit 30 Response output unit (output unit) 31 Conversation history management unit 32 Experience information generation unit 51 First LLM 52 Second LLM 100, 200 Information processing system 210 Information processing device 230 Terminal device
Claims
1. An information processing device comprising: a language input unit that is communicatively connected to a camera that moves in conjunction with the user's movement and receives language input from the user; and an imaging control unit that controls the camera based on the positional relationship between the user and the camera and the language input.
2. The information processing apparatus according to claim 1, further comprising a context analysis unit that extracts imaging conditions for the camera to image an object based on the language input and generates an operation command for the camera based on the imaging conditions, wherein the imaging control unit controls the camera based on the positional relationship between the user and the camera and the operation command.
3. The information processing apparatus according to claim 1 or 2, further comprising an image analysis unit that analyzes an image acquired from the camera based on the language input, wherein the imaging control unit controls the camera based on the analysis results by the image analysis unit in addition to the positional relationship and the language input.
4. The information processing apparatus according to any one of claims 1 to 3, further comprising an image analysis unit that analyzes an image acquired from the camera based on the language input, wherein the image analysis unit notifies the user via an output unit when the image acquired from the camera satisfies the conditions indicated by the language input.
5. The information processing apparatus according to claim 3 or 4, wherein the image analysis unit, in response to the user's instructions, notifies the user via an output unit of information indicating the object being captured by the camera at the time the instructions were given.
6. The information processing apparatus according to claim 2, wherein the context analysis unit generates a response request based on the language input, receives a response generated in accordance with the response request, and controls the output of the response.
7. The information processing apparatus according to claim 6, wherein, if the language input satisfies predetermined conditions, the context analysis unit modifies the response request regardless of whether there is any subsequent language input from the user, and the response control unit receives the response generated in accordance with the modified response request and controls the output of the response.
8. A terminal device comprising the information processing device described in any one of claims 1 to 7.
9. The terminal device according to claim 8, comprising the camera.
10. An information processing program for causing a computer to function as an information processing device according to claim 1, wherein the information processing program causes the computer to function as the language input unit and the imaging control unit.
11. An information processing system including a camera that moves in conjunction with the movement of a user, and an information processing device that is communicatively connected to the camera, wherein the information processing device comprises a language input unit that receives language input from the user, and an imaging control unit that controls the camera based on the positional relationship between the user and the camera and the language input.
12. An information processing method used in an information processing device that is communicatively connected to a camera that moves in conjunction with the movement of a user, the method comprising: receiving language input from the user; and controlling the camera based on the positional relationship between the user and the camera, and the language input.
Citation Information
Patent Citations
Information processing device, information processing method, and program
JP2017156511A
Imaging apparatus, control method, and information processing program
JP2018085579A
User support system
JP2019101766A