Information processing device, information processing method, and program

The information processing device improves image task inference accuracy in large-scale language models by utilizing chronologically continuous and scheduled task information, addressing the challenge of low inference accuracy due to lacking features in image data content.

JP7780058B1Active Publication Date: 2025-12-03KDDI CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025157355
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-12-03
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

The accuracy of inference regarding image data in large-scale language models is affected by the image content, particularly for data lacking necessary features, leading to low inference accuracy.

Method used

An information processing device that includes a storage unit for imaging task information, a data acquisition unit, and an inference result acquisition unit, which utilizes chronologically continuous and scheduled task information to improve inference accuracy by inputting prompts to a large-scale language model, incorporating imaging task information and scheduled task details.

Benefits of technology

Enhances the accuracy of image task inference by leveraging prior information on the chronological order and scheduled tasks, even for image data with low initial accuracy, by using a large-scale language model effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007780058000001_ABST
    Figure 0007780058000001_ABST
Patent Text Reader

Abstract

Improve the accuracy of inference about image data in large-scale language models. [Solution] A storage unit (10) stores imaging task information that associates an image identifier for identifying image data capturing an image of a person performing one of a plurality of predetermined tasks, information indicating the task captured in the image data, and the image capture time of the image data. A data acquisition unit (120) acquires a target data group including a plurality of image data, each associated with an image capture time, each image data including image data capturing an image of a person performing one of the plurality of tasks. An inference result acquisition unit (121) acquires an inference result including a result of inputting into the large-scale language model a prompt including imaging task information for causing a large-scale language model to infer which of the plurality of tasks is captured in each image data included in the target data group, and the target data group.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and a program, and in particular to an image recognition technology. [Background technology]

[0002] In recent years, generative AI (Artificial Intelligence), particularly large language models (LLMs), have rapidly developed and are being used in a variety of applications. For example, Patent Document 1 discloses a technology for identifying the location where image data was captured by utilizing a large language model. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 7633733 Summary of the Invention [Problem to be solved by the invention]

[0004] Even if the performance of large-scale language models improves, the accuracy of inference regarding image data is affected by the image content of the input image data, etc. For example, the accuracy of inference may be low for image data that lacks the features necessary for inference of large-scale language models.

[0005] The present invention has been made in view of these points, and aims to provide a technique for improving the accuracy of inference regarding image data in a large-scale language model. [Means for solving the problem]

[0006] A first aspect of the present invention is an information processing device including: a storage unit that stores imaging task information that associates an image identifier for identifying image data capturing an image of a person performing one of a plurality of predetermined tasks, information indicating the task captured in the image data, and an imaging time of the image data; a data acquisition unit that acquires a target data group including a plurality of image data each associated with an imaging time, each image data including image data capturing an image of a person performing one of the plurality of tasks; and an inference result acquisition unit that acquires an inference result including a result of inputting the target data group into the large-scale language model and performing inference, the prompt including the imaging task information.

[0007] The data acquisition unit may acquire a series of image data that are chronologically continuous as the image data included in the target data group, and the inference result acquisition unit may input the prompt including the imaging task information for one or more pieces of image data that are chronologically continuous with the series of image data into the large-scale language model.

[0008] The storage unit may store the imaging task information by referring to the inference result acquired by the inference result acquisition unit.

[0009] The inference result acquisition unit may input the prompt, which contains aggregated information about the imaging task information associated with imaging times that are a predetermined period after the imaging time of each of the image data included in the target data group, into the large-scale language model.

[0010] The inference result acquisition unit may input, to the large-scale language model, scheduled task information in which time information is associated with a task that is scheduled to be performed at the time indicated by the time information.

[0011] When performing inference on the image data in which a task having a similar task among the plurality of tasks is captured, the inference result acquisition unit may input the prompt including an instruction to use the scheduled task information for inference to the large-scale language model.

[0012] When the scheduled task information contains a task whose implementation period and number of times are predetermined and the inference result acquisition unit performs inference on the image data associated with an imaging time within the implementation period, the inference result acquisition unit may input the prompt to the large-scale language model, the prompt including an instruction to use the imaging task information and the scheduled task information associated with an imaging time within the implementation period for inference.

[0013] The storage section may include a degree of certainty of the association between the image data and the information indicating the task in the imaging task information.

[0014] The data acquisition unit may acquire, as the image data included in the target data group, image data corresponding to one or a plurality of pieces of imaging task information that are chronologically continuous and have a confidence level lower than a predetermined threshold, and the inference result acquisition unit may input, to the large-scale language model, the prompt including the imaging task information for the image data before and after the image data.

[0015] The data acquisition unit may acquire image data corresponding to one or multiple pieces of chronologically consecutive imaging task information for which information indicating the task is unknown as the image data contained in the target data group, and the inference result acquisition unit may input the prompt including the imaging task information before and after the imaging task information into the large-scale language model.

[0016] When the inference result acquisition unit acquires the inference result for the target data group including the image data corresponding to the imaging task information, the memory unit may change the imaging task information by referring to the inference result, and the data acquisition unit may acquire, as the image data included in the target data group, the image data corresponding to one or multiple pieces of imaging task information in chronological order in which the information indicating the task has been changed, and the image data before and after the image data.

[0017] A second aspect of the present invention is an information processing method, which includes the steps of: reading, from a storage unit, imaging task information associating an image identifier for identifying image data capturing an image of a person performing one of a plurality of predetermined tasks, information indicating the task captured in the image data, and an imaging time of the image data; acquiring a target data group including a plurality of image data, each associated with an imaging time, each image data showing an image of a person performing one of the plurality of tasks; and acquiring an inference result including a result of inputting, into the large-scale language model, a prompt including the imaging task information, the target data group.

[0018] A third aspect of the present invention is a program that causes a computer to perform the following functions: read out from a storage unit an image identifier for identifying image data capturing an image of a person performing one of a plurality of predetermined tasks, imaging task information associating information indicating the task captured in the image data with the image capture time of the image data, acquire a target data group including a plurality of image data each associated with an image capture time, each image data showing an image of a person performing one of the plurality of tasks, and acquire an inference result including a result of inputting the target data group into the large-scale language model and performing inference.

[0019] In order to provide this program or to update a part of the program, a computer-readable recording medium on which this program is recorded may be provided, or this program may be transmitted over a communication line.

[0020] Any combination of the above components, and any transformation of the present invention into a method, device, system, computer program, data structure, recording medium, etc., are also valid aspects of the present invention. [Effects of the Invention]

[0021] According to the present invention, it is possible to provide a technique for improving the accuracy of inference regarding image data in a large-scale language model. [Brief explanation of the drawings]

[0022] [Figure 1] FIG. 2 is a schematic diagram illustrating an overview of a process executed by an information processing device according to an embodiment. [Figure 2] FIG. 1 is a diagram schematically illustrating a functional configuration of an information processing device according to an embodiment. [Figure 3] 3 is a diagram schematically showing a data structure of imaging task information stored in a storage unit; FIG. [Figure 4] FIG. 2 is a diagram schematically illustrating a data structure of scheduled task information to be input to a large-scale language model according to an embodiment. [Figure 5] 10 is a flowchart illustrating a flow of information processing executed by an information processing device according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0023] <Outline of the embodiment> For example, an information processing device according to an embodiment analyzes a series of image data captured at intervals of several seconds of the store and back area of ​​a convenience store (hereinafter referred to as "convenience store"), and infers whether each image data captures the performance of one of a plurality of predetermined tasks, such as customer service, cleaning, or ordering. Specifically, the information processing device according to an embodiment acquires, as a target data group, image data to be processed from the series of image data captured at intervals of several seconds of the store and back area of ​​the convenience store. The information processing device uses a large-scale language model to infer which of the plurality of tasks is captured in each of the image data included in the target data group.

[0024] In this specification, inferring which of multiple predetermined tasks is captured in image data is referred to as “image task inference.” Note that a large-scale language model that uses image data as input and performs inference on image data is sometimes called a multimodal large language model (MLLM) or a vision language model (VLM), but in this specification it will be simply referred to as a large-scale language model.

[0025] Fig. 1 is a schematic diagram for explaining an overview of processing executed by an information processing device 1 according to an embodiment. Fig. 1 shows an example in which image data to be processed by the information processing device 1 is a series of image data in chronological order starting with image data I1, and image data I121 shows, for example, an image of a convenience store clerk picking up a product in front of a display shelf.

[0026] Hereinafter, with reference to FIG. 1, an outline of the processes executed by the information processing device 1 according to the embodiment will be explained in the order of (1) to (3), and these numbers correspond to (1) to (3) in FIG.

[0027] (1) The information processing device 1 reads out a series of imaging task information from imaging task information D1 to imaging task information D120 and transcribes it into the prompt P. Here, the imaging task information is information about image data for which image task inference has already been performed. Specifically, the imaging task information is information in which, for example, an image identifier, a task name, and an imaging time are associated. In the example shown in FIG. 1, the imaging task information is a series of information in chronological order from imaging task information D1 to imaging task information D120. The imaging task information D1 has an image identifier IID001 and a task name of "front display" (a task of moving a product at the back of a display shelf to the front). The imaging task information D120 has an image identifier IID120 and a task name of "cleaning."

[0028] (2) The information processing device 1 inputs the target data group T and the prompt P into the large-scale language model M. The target data group T includes image data acquired by the information processing device 1 as a target for image task inference. In the example shown in FIG. 1, the target data group T includes five pieces of image data with image identifiers IID121 to IID125. The prompt P is a prompt for causing the large-scale language model M to perform image task inference for each piece of image data in the target data group T, and includes imaging task information as described above.

[0029] (3) The information processing device 1 acquires an inference result R, which is the result of image task inference output by the large-scale language model M. For example, if the inference result R is that the task captured in image data I121 with an image identifier IID121 is cleaning, the inference result R is expressed in a format such as "IID121, cleaning, (capture time of image data I121)".

[0030] In this way, the information processing device 1 according to the embodiment performs image task inference for each image data included in the target data group T. As described above, in the example shown in FIG. 1, image data I121 captures an image of a convenience store clerk picking up a product in front of a display shelf. In this case, it may not be possible to determine whether the task being performed is cleaning the display shelf or stocking the product from the image data I121 alone. If image data close in time to image data I121 captures an image of the same convenience store clerk holding cleaning tools and cleaning the product being picked up in image data I121, there is a high probability that the task being performed in image data I121 is cleaning.

[0031] Even if image data capturing an image of a product being cleaned is not included in the target data group T, the information processing device 1 according to the embodiment can infer that the task being performed for the image data I121 is cleaning because the prompt P includes the imaging task information D120 immediately before the image data I121 and the task name of the imaging task information D120 is cleaning. In this way, the information processing device 1 according to the embodiment can improve the accuracy of inference for each of a series of image data in chronological order by utilizing the a priori information that the image data to be inferred is in chronological order.

[0032] <Functional configuration of information processing device 1 according to the embodiment> FIG. 2 is a diagram schematically illustrating the functional configuration of an information processing device 1 according to an embodiment. The information processing device 1 includes a storage unit 10, a communication unit 11, and a control unit 12. In FIG. 2, arrows indicate main data flows, and there may be data flows not shown in FIG. 2. In FIG. 2, each functional block indicates a configuration in functional units, rather than a configuration in hardware (device) units. Therefore, the functional blocks shown in FIG. 2 may be implemented in a single device, or may be implemented separately in multiple devices. Data may be exchanged between functional blocks via any means, such as a data bus, a network, or a portable storage medium.

[0033] The storage unit 10 is a large-capacity storage device such as a ROM (Read Only Memory) that stores the BIOS (Basic Input Output System) of the computer that realizes the information processing device 1, a RAM (Random Access Memory) that serves as the working area of ​​the information processing device 1, an HDD (Hard Disk Drive) or an SSD (Solid State Drive) that stores various information such as an OS (Operating System), application programs, and imaging task information that is referenced when the application programs are executed.

[0034] The communication unit 11 is a communication interface for the information processing device 1 to communicate with external devices, and is realized by a known communication module such as a LAN (Local Area Network) module or a Wi-Fi (registered trademark) module. Hereinafter, in this specification, description of the communication unit 11 may be omitted on the assumption that communication between the information processing device 1 and external devices is via the communication unit 11.

[0035] The control unit 12 is a processor such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), or NPU (Neural network Processing Unit) of the information processing device 1, and functions as a data acquisition unit 120 and an inference result acquisition unit 121 by executing a program stored in the memory unit 10.

[0036] 2 shows an example in which the information processing device 1 is configured as a single device. However, the information processing device 1 may be realized by a plurality of processors, memories, and other computing resources, such as a cloud computing system. In this case, each unit constituting the control unit 12 is realized by at least one of a plurality of different processors executing a program.

[0037] The storage unit 10 stores imaging task information that associates an image identifier for identifying image data that captures the state of performing one of a plurality of predetermined tasks, information indicating the task captured in the image data, and the time the image data was captured.

[0038] Here, we will explain "multiple predefined tasks" (hereinafter referred to as "task definitions") and "image data." First, in the above example, an example of operational tasks at a convenience store, which is a retail business, was given as the subject of imaging of "image data." However, this is merely an example, and various business operations in various industries, such as manufacturing and logistics, may be the subject of imaging. Imaging means include, but are not limited to, surveillance cameras and smart glasses worn by the task performer. Imaging may be performed periodically at predetermined time intervals, such as every few seconds or minutes, or may be performed as an event trigger based on the detection of motion, sound, etc.

[0039] Next, a "task definition" defines a set of tasks that may be captured in image data to be processed by the information processing device 1. Any of the tasks included in the task definition is associated with the image data as a result of image task inference by the information processing device 1. The task definition is information that associates, for example, a task identifier for identifying the task with a task name for each task included in the task set. A task description that explains an overview of the task may also be associated with the task definition. For example, the task description in the previous section could be "a task to move a product from the back of a display shelf to the front."

[0040] The industry, business field, and granularity of the task definition may be freely set depending on the subject of image data capture and the intended use of the information processing device 1. For example, in the convenience store example mentioned above, the task name used in the task definition may be a generic task name common to the retail industry in general, or a task name specific to a particular convenience store chain. Furthermore, the granularity of the task definition may be relatively more abstract, such as interpersonal tasks or non-interpersonal tasks, or relatively more specific, such as cleaning display shelves or cleaning aisles. Specific methods for using task definitions by the information processing device 1 will be described later.

[0041] 3 is a diagram schematically illustrating the data structure of the imaging task information stored in the storage unit 10. In the example illustrated in FIG. 3, the imaging task information is information that associates an "image identifier" for identifying image data, a "task identifier" and a "task name" that indicate the task captured in the image data, the "imaging time" of the image data, and the "certainty" of the image task inference for the image data. As described above, the imaging task information is information about image data for which image task inference has already been performed. Here, the performed image task inference may be an inference performed by the information processing device 1, an inference performed by an information processing device different from the information processing device 1, or an inference performed by a human.

[0042] In the example shown in Figure 3, the image data with image identifier IID001 was captured at "2025 / 7 / 24 13:17:00." This image data contains an image of a task with task identifier task_003 and task name "Maechen." The seven images with image identifiers IID001 to IID007 were each captured at 10-second intervals. The "certainty" and the case where the task name is "unknown," such as image identifier "IID004," will be discussed later.

[0043] The task description described above may be associated as information indicating the task. Furthermore, the inference basis for image task inference regarding the image data may also be associated as information indicating the task. For example, if the task name is cleaning, the inference basis may be "because it was observed that the person was cleaning the display shelf where the sweets were placed with cleaning tools." The information indicating the task may be information included in the task definition, or may be information newly generated by the information processing device 1 or the large-scale language model M based on the task definition and the image data to be processed.

[0044] The data acquisition unit 120 acquires a target data group T including a plurality of image data each associated with an image capture time. Here, each image data included in the target data group T captures an image of the performance of one of the tasks in the task definition. The image data included in the target data group T is basically image data for which image task inference has not yet been performed, but as will be described later, in some cases it may be image data for which image task inference has already been performed.

[0045] The data acquisition unit 120 may limit the target data set T to be acquired, taking into consideration the memory capacity and processing performance of the information processing device 1, an upper limit on the number of input tokens of the large-scale language model M, and the like. For example, image data may be acquired every predetermined time unit, such as one minute or five minutes, and an upper limit may be set to a maximum of 30 pieces of image data to be acquired. This reduces the processing load on the information processing device 1 and the large-scale language model M, enabling stable data processing.

[0046] The inference result acquisition unit 121 reads out the imaging task information stored in the storage unit 10. The inference result acquisition unit 121 inputs a prompt P and the target data group T to the large-scale language model M for performing image task inference for each piece of image data included in the target data group T. The prompt P includes the imaging task information.

[0047] Here, the large-scale language model M receives text data and image data as input and analyzes and infers the content of the data. The large-scale language model M includes, but is not limited to, existing models that are publicly available, models that are fine-tuned from existing models, and models that are independently constructed by users of the information processing device 1.

[0048] The prompt P to be input to the large-scale language model M is generated, for example, by the inference result acquisition unit 121. Specifically, for example, a template of the prompt P is stored in the storage unit 10, and the inference result acquisition unit 121 reads out the template of the prompt P and imaging task information from the storage unit 10 and embeds the imaging task information in the template of the prompt P to generate the prompt P.

[0049] The prompt template includes instructions to the large-scale language model M, such as, "Please infer which task in the task definition is captured in each piece of image data included in the target data group T. The image data to be inferred is a series of image data in chronological order. When making the inference, please use the capture time of the image data, information about the image data in the target data group T, and the following imaging task information. Please output the inference result R in the format of the imaging task information." The content of the instructions can be adjusted depending on the large-scale language model used, and for example, depending on the large-scale language model, a relatively simple instruction such as "Please infer which task is captured in each piece of image data included in the target data group T" may be used.

[0050] As the task definition is described in the example of prompt P above, task definition information is required for image task inference. The task definition information may be input to the large-scale language model M together with the prompt P and the target data set T each time.

[0051] Furthermore, for example, using a known technique such as prompt caching, the task definition may be automatically supplemented using cached context information or the like after the initial image task inference of the large-scale language model M. Also, a large-scale language model M may be constructed that is trained to input a prompt P and a target data group T and output an inference result R. This allows the information processing device 1 to eliminate the need for the user to input a task definition each time.

[0052] The task definition may be applied to multiple industries or multiple business fields, rather than being limited to one. For example, if a conglomerate, a business structure in which multiple companies from different industries form a single corporate group, uses the information processing device 1, task definitions that can be used for all industries or business fields operated by the conglomerate may be created. In the convenience store example mentioned above, the task definition may also include tasks for various stores, including those from different convenience store chains, as well as tasks from other industries, such as transportation and pharmacies. In addition to traditional retail tasks, convenience stores already perform tasks such as assisting with prescription acceptance and handling mail, and the number of tasks they perform may expand in the future. Some convenience stores also offer relatively unique services.

[0053] This makes it possible to create a more general task definition that can handle a wider range of convenience store tasks that may arise now and in the future. In particular, by using data related to various convenience stores and operations in a variety of industries in training the large-scale language model M, the information processing device 1 is expected to obtain relatively more accurate inference results R from the large-scale language model M even for new tasks that may appear in the future that are not included in the task definition.

[0054] The inference result acquisition unit 121 acquires an inference result R including the result of inference by the large-scale language model M. The inference result R includes, for example, information equivalent to the imaging task information. Specifically, for each image data included in the target data group T, the inference result R is information in which an "image identifier," a "task identifier" and a "task name" that are information indicating the task, and the "imaging time" of the data are associated with each other. As with the imaging task information, the information indicating the task may be associated with a task description and an inference basis.

[0055] As described above, by utilizing the prior information that the image data to be inferred is in chronological order, the information processing device 1 can improve the inference accuracy of image task inference for each of a series of image data in chronological order. Specifically, even for image data for which the accuracy of image task inference is low when used alone, the information processing device 1 can improve the inference accuracy of image task inference by having the large-scale language model M use image data and imaging task information captured at close times. In particular, even if the amount of image data that can be included in the target data group T at one time is limited due to an upper limit on the number of input tokens of the large-scale language model M, the information processing device 1 can improve the inference accuracy of image task inference by having the large-scale language model M use imaging task information that is relatively small in size.

[0056] More specifically, for example, for image data with low inference accuracy on its own, if it is known from image data captured at nearby times and imaging task information that the task performed in the five minutes before and after the image data capture time is cleaning, then the probability that the task captured in the image data is also cleaning increases. Note that causes of low inference accuracy on its own are not limited, but include, for example, a lack of features necessary for inference or noise in the image.

[0057] Note that the data acquisition unit 120 may acquire the target data group T, and the inference result acquisition unit 121 may acquire the inference result R of image task inference for the image data to be inferred, in a state in which the inference result acquisition unit 121 has not yet read out the imaging task information stored in the storage unit 10. For example, the information processing device 1 can be used even before image task inference is performed on the image data to be inferred and there is no imaging task information stored in the storage unit 10. This allows the information processing device 1 to use image data for which image task inference has not yet been performed, thereby expanding the range of use.

[0058] It is preferable that the data acquisition unit 120 acquires a series of image data that are chronologically continuous as the image data included in the target data group T. It is also preferable that the inference result acquisition unit 121 inputs to the large-scale language model M a prompt P that includes imaging task information for one or more image data that are chronologically continuous with the series of image data. This increases the probability that image data and imaging task information that can improve the inference accuracy of image task inference for the image data to be inferred will be included, and therefore the information processing device 1 is expected to further improve the inference accuracy of image task inference using the large-scale language model M. It is to be noted that the imaging task information may be chronologically future or past with respect to the image data included in the target data group T, or may include both past and future.

[0059] Furthermore, the storage unit 10 may process the inference result R acquired by the inference result acquisition unit 121 into the format of imaging task information and store it as new imaging task information. This allows the information processing device 1 to improve the inference accuracy of image task inference by having the large-scale language model M use the inference result R of the latest image task inference in the next and subsequent inferences.

[0060] Furthermore, when the imaging time associated with the imaging task information is a predetermined period of time after the imaging time of each piece of image data included in the target data group T, the inference result acquisition unit 121 may input a prompt P including an aggregated version of the imaging task information into the large-scale language model M. For example, when the imaging task information before aggregation is "IID300, 2025 / 7 / 24 13:25:00, cleaning" and the task name of the imaging task information for the next 25 minutes is "cleaning," the inference result acquisition unit 121 aggregates the information as, for example, "IID300 to IID450, 2025 / 7 / 24 13:25:00 to 13:50:00, cleaning."

[0061] Here, the predetermined period is a period that the inference result acquisition unit 121 refers to in order to determine whether or not to perform aggregation. This period may be determined experimentally in consideration of the size of the imaging task information, the upper limit of the number of input tokens of the large-scale language model M, optimization of inference accuracy, optimization of response time, etc., and may be, for example, five minutes. This allows the information processing device 1 to realize organization of imaging task information taking into account time series while complying with the input restrictions of the large-scale language model M, thereby further improving the inference accuracy of image task inference.

[0062] Furthermore, the inference result acquisition unit 121 may input scheduled task information, which associates time information with a task scheduled to be performed at that time, together with the prompt P and the target data group T to the large-scale language model M. This allows the large-scale language model M to use the scheduled task information, for example, when there is a task scheduled to be performed before or after the image capture time of the image data to be inferred, thereby enabling the information processing device 1 to further improve the inference accuracy of image task inference. The scheduled task information may simply be input to the large-scale language model M, or the prompt P may include an instruction such as "Please also use the scheduled task information when performing inference." Whether to include an instruction regarding the scheduled task information and what kind of instruction to include may be determined through experiments depending on the large-scale language model M used by the information processing device 1.

[0063] FIG. 4 is a diagram schematically illustrating the data structure of scheduled task information input to the large-scale language model M according to an embodiment. In the example shown in FIG. 4, the scheduled task information is information in which a "schedule identifier" for identifying each scheduled task information, a "schedule start time" and a "schedule end time" representing time information, a "task identifier" and a "task name" representing the task to be performed, an "implementation period" and a "number of times" are associated with each other. The "implementation period" and the "number of times" will be described later. The scheduled task information may be a schedule corresponding to a specific day, week, month, year, etc., or may be a periodic or conditional schedule based on a day of the week, a day off, a public holiday, etc.

[0064] In the example shown in FIG. 4, the scheduled task information with the schedule identifier schd_001 has a scheduled start time of "6:00:00" and a scheduled end time of "6:30:00." A task called "displaying the morning paper" with the task identifier task_005 is associated with this scheduled task information as a task scheduled to be performed. As shown by the scheduled task information with the schedule identifier schd_002 having a scheduled start time of "8:00:00," the scheduled task information does not have to be consecutive in chronological order. Furthermore, as shown by the scheduled task information shown in FIG. 4, the scheduled task information with the schedule identifier schd_002 and the scheduled task information with the schedule identifier schd_003 are scheduled to be performed between 8:30:00 and 9:00:00, so multiple tasks scheduled to be performed in the same time period may exist.

[0065] Furthermore, when performing inference on image data in which a task whose task definition has a similar task is captured, the inference result acquisition unit 121 may input a prompt P including an instruction to use scheduled task information for inference to the large-scale language model M. This clarifies the conditions under which the large-scale language model M uses scheduled task information for inference, and enables the information processing device 1 to improve the inference accuracy in image task inference and the processing efficiency related to inference.

[0066] Examples of similar tasks include the task of displaying morning papers and the task of displaying evening papers shown in FIG. 4. The task of displaying morning papers is similar to the task of displaying evening papers, as it can be difficult to determine whether the morning papers or evening papers are being displayed based on image data alone, and vice versa. In the example shown in FIG. 4, the task of displaying morning papers has a schedule identifier of schd_001, a scheduled start time of "6:00:00," and a scheduled end time of "6:30:00." For example, if image data captured at 6:15:00 captures a scene of displaying morning papers, performing image task inference on the image data may not be able to determine whether the image data represents a morning paper display or an evening paper display. In such cases, the large-scale language model M can use the scheduled task information shown in FIG. 4 for inference, thereby increasing the probability of inferring that the image data captures a scene of displaying morning papers.

[0067] In this case, the prompt P may include an instruction such as, for example, "When performing inference on image data in which a task having a similar task in its task definition is captured, please use the scheduled task information for inference." The content of the instruction may be adjusted depending on the large-scale language model used by the information processing device 1. For example, depending on the large-scale language model, a relatively simple instruction such as, "When similar tasks exist, please also use the scheduled task information." Alternatively, depending on the large-scale language model, a relatively detailed instruction such as, "When the similarity between task vectors beta obtained by converting each task included in the task definition into a vector is calculated, if inference is performed on image data in which a task having a similarity equal to or greater than a predetermined threshold is captured, please use the scheduled task information for inference." is preferable.

[0068] Furthermore, when the scheduled task information includes a task whose implementation period and number of times are predetermined and inference is to be performed on image data captured within the implementation period, the inference result acquisition unit 121 may input a prompt P including an instruction to use, for inference, the scheduled task information and imaging task information associated with the imaging time within the implementation period, to the large-scale language model M. This clarifies the conditions under which the large-scale language model M uses the scheduled task information for inference, and enables the information processing device 1 to improve the inference accuracy in image task inference and improve the processing efficiency related to inference.

[0069] Here, an example of a task for which the implementation period and number of times are predetermined in the scheduled task information is the task of displaying the morning paper shown in Fig. 4. As described with reference to Fig. 4, the scheduled task information is information that associates not only a "scheduled identifier," "scheduled start time," "scheduled end time," "task identifier," and "task name," but also an "implementation period" and "number of times."

[0070] Here, the "execution period" is the period of time determined as the time zone during which the task implementer will implement the task. In contrast, the "planned period" is the period between the "planned start time" and the "planned end time," and is the period during which the implementer plans to actually implement the task. Generally, the implementer determines the planned period as all or part of the implementation period and plans the implementation of the task. The "number of times" is the number of times the task implementer should implement the task during the implementation period.

[0071] In the example shown in FIG. 4, the task "displaying morning newspapers" has a scheduled identifier of schd_001, an implementation period of "3:00:00-9:00:00," and a count of "1." For example, if image data captured at 7:00:00 shows a display of weekly magazines, even if image task inference is performed on the image data, it may not be possible to determine whether the display is a morning newspaper or a weekly magazine. In such a case, if the task name of the imaging task information associated with the imaging times from 6:00:00 to 6:15:00 is "displaying morning newspapers," the information processing device 1 can increase the likelihood that the large-scale language model M will infer that the task captured in the image data is displaying weekly magazines by causing the large-scale language model M to use the imaging task information and the scheduled task information shown in FIG. 4.

[0072] In this case, the prompt P may include an instruction such as, for example, "If there is a task whose implementation period and number of times are defined in the scheduled task information, and if inference is to be made on image data captured within the implementation period, please use the scheduled task information and the imaging task information associated with the imaging time within the implementation period for inference." The content of the instruction may be adjusted depending on the large-scale language model to be used, and, for example, depending on the large-scale language model, a relatively simple instruction such as, "If there is scheduled task information whose implementation period and number of times are defined, please use the scheduled task information and the imaging task information for the period when inferring image data during the period."

[0073] The storage unit 10 may also include in the imaging task information a degree of certainty of the association between image data and information indicating a task. The degree of certainty is information output by the entity performing the image task inference as a degree of confidence the entity performing the image task inference has in the accuracy of the result. As described with reference to FIG. 3 , the imaging task information is information that associates not only an “image identifier,” a “task identifier,” a “task name,” and “imaging time,” but also a “certainty.” In the example shown in FIG. 3 , the image data with the image identifier IID001 has the task name “Pre-Changing” and a certainty of “0.9,” so the certainty that the task captured in the image data is Pre-Changing is 0.9. The certainty of the imaging task information with the image identifier IID003 is set to a value of “0.6.” In the example shown in FIG. 3 , the certainty is expressed as a real value ranging from 0 to 1, with a minimum value of 0 and a maximum value of 1, indicating a higher degree of certainty in the accuracy of the result.

[0074] This allows the large-scale language model M to utilize information on the degree of certainty, and the information processing device 1 is expected to prevent a decrease in the inference accuracy of image task inference. Without information on the degree of certainty of imaging task information, the large-scale language model M will use information with low and high certainty equally, which could result in a decrease in inference accuracy. Therefore, by allowing the large-scale language model M to utilize information on the degree of certainty, it is expected that the degree of use of each piece of imaging task information can be adjusted by weighting each piece of imaging task information according to the degree of certainty, etc. Therefore, the information processing device 1 is expected to prevent a decrease in the inference accuracy of image task inference.

[0075] At this time, the inference result acquisition unit 121 may input a prompt P including an instruction to adjust the degree of use of the imaging task information in image task inference according to the degree of certainty of the imaging task information to the large-scale language model M. This allows the information processing device 1 to explicitly instruct the large-scale language model M to adjust the degree of use of the imaging task information according to the degree of certainty, and more reliably prevent a decrease in the inference accuracy of the image task inference.

[0076] The above mainly describes the case where the image data included in the target data group T is image data for which image task inference has not been performed. Next, we will explain the case where the image data included in the target data group T is image data for which image task inference has already been performed.

[0077] The data acquiring unit 120 may acquire image data corresponding to imaging task information whose certainty is lower than a predetermined threshold as image data included in the target data group T. At this time, if the certainty of each imaging task information that is chronologically consecutive to the imaging task information is also lower than the predetermined threshold, the data acquiring unit 120 collectively acquires image data corresponding to a series of imaging task information whose certainty is lower than the predetermined threshold.

[0078] Here, the predetermined threshold value for the confidence level is a threshold value used by the data acquisition unit 120 to determine whether or not to perform image task inference again on the image data. The threshold value may be determined by experiment, taking into consideration the tendency of the large-scale language model M to assign confidence levels, the purpose of use of the information processing device 1, the desired time to complete the entire process, etc., and may be, for example, "0.7."

[0079] Then, the inference result acquisition unit 121 preferably inputs a prompt P including imaging task information for image data before and after the image data acquired by the data acquisition unit 120 to the large-scale language model M. Considering the acquisition method of the image data, the certainty of the imaging task information for image data before and after the image data included in the prompt P is equal to or greater than a predetermined threshold.

[0080] As a result, for image data corresponding to one piece of imaging task information or pieces of imaging task information that are consecutive in time series and have a certainty level lower than a predetermined threshold, the information processing device 1 can improve the inference accuracy of image task inference by having the large-scale language model M use imaging task information that is consecutive in time series and has a certainty level higher than the threshold to perform image task inference again.

[0081] For example, suppose the data acquisition unit 120 acquires a target data group T sequentially from past image data in chronological order and causes the large-scale language model M to perform image task inference. For each image data item included in a certain target data group T, the large-scale language model M infers that it is cleaning with a confidence level of 0.3, even though it is unclear whether the image data includes a cleaning task or a previous task. If the result of image task inference by the large-scale language model M for each image data item included in the next target data group T is previous with a confidence level of 0.9, then the image data inferred as cleaning with a confidence level of 0.3 is likely to include the previous task. Therefore, the information processing device 1 can increase the likelihood that the large-scale language model M will infer that the image data includes a previous state by causing the large-scale language model M to perform image task inference again using the above method.

[0082] Furthermore, the data acquiring unit 120 may acquire image data corresponding to imaging task information for which information indicating a task is unknown as image data included in the target data group T. At this time, if information indicating a task of each piece of imaging task information that is chronologically consecutive to the imaging task information is unknown, image data corresponding to a series of imaging task information for which information indicating a task is unknown is acquired collectively.

[0083] Here, imaging task information in which the information indicating the task is unknown refers to, for example, imaging task information in which the large-scale language model M cannot infer that the task captured in the image data is any task in the task definition, and therefore the information indicating the task is associated with "unknown." Alternatively, imaging task information in which the information indicating the task is unknown refers to, for example, imaging task information in which the task definition includes a task "unknown," and the large-scale language model M cannot infer that the task captured in the image data is any task other than "unknown," and therefore the information indicating the task is associated with "unknown." In the example shown in FIG. 3, the image data with the image identifier IID004 has the task name "unknown." Note that the certainty of the image data with the image identifier IID004 is "-," indicating that it is not set, but the certainty may be set to "0."

[0084] Then, the inference result acquisition unit 121 preferably inputs a prompt P including imaging task information before and after the imaging task information corresponding to the image data acquired by the data acquisition unit 120 to the large-scale language model M. In consideration of the method for acquiring the image data, the imaging task information before and after the imaging task information included in the prompt P is a task different from the task for which the information indicating the task is unknown.

[0085] As a result, for image data corresponding to one piece of imaging task information or pieces of imaging task information that are consecutive in time series and for which the information indicating the task is unknown, the information processing device 1 can improve the inference accuracy of image task inference by having the large-scale language model M use imaging task information that is a different task that comes before or after in time series and for which the information indicating the task is unknown, and perform image task inference again.

[0086] For example, suppose the data acquisition unit 120 acquires a target data group T sequentially from past image data in chronological order and performs image task inference using the large-scale language model M. For each piece of image data included in a certain target data group T, the large-scale language model M is unclear as to whether a cleaning task or a previous task is captured, and ultimately infers that the task name is unknown. If the result of image task inference for each piece of image data included in the next target data group T is previous, there is a high probability that the image data inferred as unknown contains an image of the previous task. Therefore, the information processing device 1 can increase the probability that the large-scale language model M will infer that a previous state has been captured by having the large-scale language model M perform image task inference again using the above method.

[0087] Furthermore, when the inference result acquisition unit 121 acquires an inference result R for a target data group T including image data corresponding to the imaging task information, the storage unit 10 may change the imaging task information by referring to the inference result R. This case can be rephrased as a case where the inference result R acquired by the inference result acquisition unit 121 includes a result of image task inference for image data corresponding to the imaging task information already stored in the storage unit 10. Furthermore, specifically, the change of the imaging task information by the storage unit 10 means, for example, the storage unit 10 refers to information about the image data in the inference result R, compares it with the imaging task information stored therein, and, if there is any information that differs as a result of the comparison, changes the information that differs based on the information in the inference result R. Note that, in the example shown in FIG. 3, the information that may differ is the "task identifier" and "task name," which are information indicating the task, and the "certainty level."

[0088] The data acquisition unit 120 preferably acquires image data corresponding to the imaging task information whose task-indicating information has been changed, and image data before and after the image data, as image data included in the target data group T. At this time, if the task-indicating information of each imaging task information chronologically consecutive to the imaging task information whose task-indicating information has been changed has also been changed, the data acquisition unit 120 acquires image data corresponding to the series of imaging task information whose task-indicating information has been changed, and image data before and after the image data. Considering the method of acquiring the image data, the task-indicating information of the oldest image data and the futureest image data in chronological order among the image data acquired by the data acquisition unit 120 has not been changed.

[0089] As a result, the information processing device 1 can improve the inference accuracy of image task inference by having the large-scale language model M use image data corresponding to one piece of imaging task information or chronologically consecutive pieces of imaging task information for which the information indicating the task has been changed, and having the large-scale language model M perform image task inference again for image data corresponding to imaging task information that comes before or after in the chronological order and for which the information indicating the task has not yet been changed, thereby attempting to extend the information indicating the changed task to the image data before and after it.

[0090] <Processing flow of information processing method executed by information processing device 1> 5 is a flowchart for explaining the flow of information processing executed by the information processing device 1 according to the embodiment. The processing in this flowchart starts, for example, when the information processing device 1 is started.

[0091] The inference result acquisition unit 121 reads out imaging task information from the storage unit 10 (S1). The data acquisition unit 120 acquires a target data group T including image data to be inferred (S2).

[0092] The inference result acquisition unit 121 inputs a prompt P and a target data group T to the large-scale language model M (S3). The prompt P is a prompt for causing the large-scale language model M to perform image task inference for each image data of the target data group T, and also includes imaging task information. The prompt P is generated by, for example, the inference result acquisition unit 121.

[0093] The inference result acquisition unit 121 acquires the inference result R including the result of the image task inference output by the large-scale language model M (S4). When the inference result acquisition unit 121 acquires the inference result R, the processing in this flowchart ends.

[0094] <Advantages of the information processing device 1 according to the embodiment> As described above, according to the information processing device 1 according to the embodiment, the accuracy of inference regarding image data in the large-scale language model M can be improved.

[0095] Furthermore, this invention will make it possible to contribute to Goal 9 of the United Nations' Sustainable Development Goals (SDGs), which is "Build resilient infrastructure, promote inclusive and sustainable industrialization, and promote innovation and resilience."

[0096] The present invention has been described above using embodiments, but the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. For example, all or part of the device can be configured by functionally or physically distributing or integrating any unit. Furthermore, new embodiments resulting from any combination of multiple embodiments are also included in the embodiments of the present invention. The effects of the new embodiments resulting from the combination also have the effects of the original embodiments. [Explanation of symbols]

[0097] 1. Information processing device 10...Storage section 11. Communications Department 12 Control section 120 Data acquisition section 121...Inference result acquisition unit M···Large-scale language model

Claims

1. a storage unit that stores imaging task information that associates an image identifier for identifying image data that captures a state in which one of a plurality of predetermined tasks is being performed, information indicating the task captured in the image data, and an imaging time of the image data; a data acquisition unit that acquires a target data group including a plurality of image data pieces each associated with an image capture time, each image data piece depicting the performance of one of the plurality of tasks; an inference result acquisition unit that acquires an inference result including a result of inference by inputting the target data group into the large-scale language model a prompt for inferring which of the plurality of tasks is imaged in each of the image data included in the target data group, the prompt including the imaging task information; and the target data group; An information processing device comprising:

2. the data acquisition unit acquires a series of image data that are continuous in time series as the image data included in the target data group; the inference result acquisition unit inputs the prompt, which includes the imaging task information for one or more pieces of image data that are time-series consecutive to the series of image data, into the large-scale language model; The information processing device according to claim 1 .

3. the storage unit stores the imaging task information by referring to the inference result acquired by the inference result acquisition unit. The information processing device according to claim 1 .

4. the inference result acquisition unit inputs the prompt, which includes aggregated information, into the large-scale language model for the imaging task information associated with an imaging time that is a predetermined period after the imaging time of each of the image data included in the target data group; The information processing device according to claim 1 .

5. the inference result acquisition unit inputs scheduled task information, which associates time information with a task scheduled to be performed at the time indicated by the time information, into the large-scale language model; The information processing device according to claim 1 .

6. When performing inference on the image data in which a task having a similar task among the plurality of tasks is captured, the inference result acquisition unit inputs the prompt including an instruction to use the scheduled task information for inference to the large-scale language model. The information processing device according to claim 5 .

7. When there is a task whose implementation period and number of times are predetermined in the scheduled task information and inference is to be performed on the image data associated with an imaging time within the implementation period, the inference result acquisition unit inputs the prompt including an instruction to use the imaging task information associated with an imaging time within the implementation period and the scheduled task information for inference to the large-scale language model. The information processing device according to claim 5 .

8. the storage unit includes a degree of certainty of an association between the image data and the information indicating the task in the imaging task information; The information processing device according to claim 1 .

9. the data acquisition unit acquires, as the image data included in the target data group, image data corresponding to one piece of imaging task information or a plurality of pieces of imaging task information that are continuous in time series and have a certainty level lower than a predetermined threshold; the inference result acquisition unit inputs the prompt including the imaging task information for the image data before and after the image data into the large-scale language model; The information processing device according to claim 8 .

10. the data acquisition unit acquires, as the image data included in the target data group, the image data corresponding to one piece of imaging task information or a plurality of pieces of imaging task information that are continuous in time series, for which information indicating the task is unknown; the inference result acquisition unit inputs the prompt including the imaging task information before and after the imaging task information into the large-scale language model; The information processing device according to claim 8 .

11. When the inference result acquisition unit acquires the inference result for the target data group including the image data corresponding to the imaging task information, the storage unit changes the imaging task information by referring to the inference result; the data acquisition unit acquires, as the image data included in the target data group, the image data corresponding to one piece of imaging task information or a plurality of pieces of imaging task information that are chronologically continuous, in which the information indicating the task has been changed, and the image data before and after the image data.

11. The information processing device according to claim 9 or 10.

12. The processor: reading from a storage unit imaging task information that associates an image identifier for identifying image data capturing an image of a state in which one of a plurality of predetermined tasks is being performed, information indicating the task captured in the image data, and an imaging time of the image data; acquiring a target data group including a plurality of image data pieces each associated with an image capture time, each image data piece depicting an image of the person performing one of the plurality of tasks; a step of acquiring an inference result including a result of inputting a prompt including the imaging task information and the target data group into the large-scale language model to cause the large-scale language model to infer which of the plurality of tasks is imaged in each of the image data included in the target data group; An information processing method that performs the above.

13. On the computer, a function of reading from a storage unit an image identifier for identifying image data in which a state of performing one of a plurality of predetermined tasks is captured, information indicating the task captured in the image data, and imaging task information in which the imaging time of the image data is associated with each other; a function of acquiring a target data group including a plurality of image data pieces each associated with an image capture time, each image data piece capturing an image of a state in which one of the plurality of tasks is being performed; a function of acquiring an inference result including a result of inputting the target data group into the large-scale language model a prompt for inferring which of the plurality of tasks is imaged in each of the image data included in the target data group, the prompt including the image capture task information; and the target data group; A program to make this happen.

Citation Information

Patent Citations

  • Artificially intelligent assistant for work protocols

    JP2025121863A

  • Location identification system and program

    JP7633733B1