Information processing device, information processing method, and program
The information processing device improves large-scale language model inference accuracy by identifying high and low confidence data groups, prompting re-inference, and updating task classification information, addressing low accuracy issues in image data inference.
Patent Information
- Application Number
- JP2025157356
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-09-22
AI Technical Summary
The accuracy of inference in large-scale language models is affected by the image content of the input image data, particularly for images lacking necessary features, leading to low inference accuracy.
An information processing device that identifies high and low confidence data groups in a chronological series of image data, acquires a target data group including a re-inference data group with low confidence, prompts the large-scale language model for re-inference, and modifies task classification information based on the results to improve accuracy.
Enhances the accuracy of image data inference in large-scale language models by utilizing chronological order and prior information, especially for images with initially low confidence levels.
Smart Images

Figure 0007766222000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, an information processing method, and a program, and in particular to an image recognition technology. [Background technology]
[0002] In recent years, generative AI (Artificial Intelligence), particularly large language models (LLMs), have rapidly developed and are being used in a variety of applications. For example, Patent Document 1 discloses a technology for identifying the location where image data was captured by utilizing a large language model. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 7633733 Summary of the Invention [Problem to be solved by the invention]
[0004] Even if the performance of large-scale language models improves, the accuracy of inference regarding image data is affected by the image content of the input image data, etc. For example, the accuracy of inference may be low for image data that lacks the features necessary for inference of large-scale language models.
[0005] The present invention has been made in view of these points, and aims to provide a technique for improving the accuracy of inference regarding image data in a large-scale language model. [Means for solving the problem]
[0006] A first aspect of the present invention is an information processing device. The information processing device includes a storage unit that stores task classification information in which a plurality of image data in which a state of performing a task is captured is associated with the capture time of each image data, the task name of the task inferred from each image data by a large-scale language model, and the confidence of the inference; an identification unit that references the task classification information and identifies (1) a high confidence data group, which is one or more image data whose confidence is equal to or greater than a predetermined threshold, whose task name is the same, and which are chronologically consecutive, and (2) a low confidence data group, which is other image data and is one or more image data that are chronologically consecutive; and a device that identifies the high confidence data group and the high confidence data group. The system includes a data acquisition unit that acquires a target data group including a re-inference data group, which is the low-confidence data group that is chronologically continuous with a data group; a prompt that causes the large-scale language model to infer whether the task name of the task captured in each of the image data that constitutes the re-inference data group is the task name associated with the image data that constitutes the high-confidence data group included in the target data group or is different from the task name; an inference control unit that inputs the re-inference data group into the large-scale language model and acquires the inference results; and a modification unit that changes the task classification information in the memory unit by referring to the results acquired by the inference control unit.
[0007] The identification unit may further identify the low-certainty data group as (1) a low-certainty data group consisting of one or more image data whose certainty is less than a predetermined threshold, whose task name is the same, and which are chronologically consecutive, and (2) an unclassified data group consisting of one or more image data whose task name is inferred to be unknown and which are chronologically consecutive.
[0008] The data acquisition unit may acquire the target data group including the high confidence data group with the high confidence data group having the high confidence level and the re-inference data group when the high confidence data groups chronologically precede and follow the re-inference data group and the high confidence data groups have different confidence levels.
[0009] When the inference control unit obtains a negative result that the task name is different, the data acquisition unit does not need to acquire the target data group that includes both the high-confidence data group that includes the image data used to input the negative result and the re-inference data group that includes the image data, and may repeat the identification by the identification unit, acquisition by the data acquisition unit, acquisition by the inference control unit, and modification by the modification unit until there are no more target data groups for the data acquisition unit to acquire.
[0010] The identification unit may further identify a task name data group, which is one or more pieces of image data having the same task name and that are chronologically consecutive, the data acquisition unit may further acquire a renamed data group including the task name data group, and the inference control unit may further acquire a prompt for causing the large-scale language model to infer a task name that embodies the task name associated with the task name data group that constitutes the renamed data group, and the result of inference obtained by inputting the renamed data group into the large-scale language model.
[0011] The identification unit may further identify a task name data group, which is one or more image data having the same task name and which are chronologically consecutive, and the data acquisition unit may further acquire a renamed data group including multiple chronologically consecutive task name data groups, and the inference control unit may further acquire a prompt to cause the large-scale language model to infer a common task name for changing each task name associated with the task name data groups that make up the renamed data group, and the result of inference by inputting the renamed data group into the large-scale language model.
[0012] The inference control unit may input a list of task names that lists examples of the task names to be output to the large-scale language model.
[0013] The inference control unit may input the prompt to the large-scale language model, the prompt including an instruction to infer the common task name for changing each task name when the tasks associated with the task name data group that constitutes the rename data group are similar to each other.
[0014] The inference control unit may input the prompt, which includes an instruction to keep the number of types of task names included in the task classification information within a predetermined range, to the large-scale language model.
[0015] The data acquisition unit may not need to acquire the renamed data group input into the large-scale language model again as the renamed data group, and may repeat the identification by the identification unit, acquisition by the data acquisition unit, acquisition by the inference control unit, and modification by the modification unit until there are no more renamed data groups to acquire by the data acquisition unit.
[0016] A second aspect of the present invention is an information processing method, which includes the steps of: reading from a storage unit task classification information in which a plurality of image data in which a state of performing a task is captured, the image capture time of each image data, the task name of the task inferred from each image data by a large-scale language model, and the degree of certainty of the inference; and referring to the task classification information, identifying (1) a high-certainty data group which is one or more image data whose degree of certainty is equal to or greater than a predetermined threshold, whose task name is the same, and which are chronologically consecutive; and (2) a low-certainty data group which is other image data and is one or more image data which are chronologically consecutive; and The method includes the steps of: acquiring a target data group including a data group and a re-inference data group, which is the low-certainty data group that is continuous in time series with the high-certainty data group; prompting the large-scale language model to infer whether the task name of the task captured in each of the image data that constitutes the re-inference data group is the task name associated with the image data that constitutes the high-certainty data group included in the target data group or different from the task name; inputting the re-inference data group into the large-scale language model to acquire the inference results; and changing the task classification information in the memory unit by referring to the acquired results.
[0017] A third aspect of the present invention is a program that is programmed on a computer to: read from a storage unit task classification information that associates a plurality of image data in which a state of performing a task is captured, the image capture time of each image data, the task name of the task inferred from each image data by a large-scale language model, and the confidence of the inference; and, by referencing the task classification information, identify (1) a high confidence data group that is one or more image data whose confidence is equal to or greater than a predetermined threshold, whose task name is the same, and that are chronologically consecutive; and (2) a low confidence data group that is other image data and is one or more image data that are chronologically consecutive; and The system realizes the following functions: a function to acquire a target data group including a data group and a re-inference data group which is the low-certainty data group whose time series is continuous with the high-certainty data group; a prompt to cause the large-scale language model to infer whether the task name of the task captured in each of the image data constituting the re-inference data group is the task name associated with the image data constituting the high-certainty data group contained in the target data group or different from the task name; a function to input the re-inference data group into the large-scale language model and acquire the inference results; and a function to change the task classification information in the memory unit by referring to the acquired results.
[0018] In order to provide this program or to update a part of the program, a computer-readable recording medium on which this program is recorded may be provided, or this program may be transmitted over a communication line.
[0019] Any combination of the above components, and any transformation of the present invention into a method, device, system, computer program, data structure, recording medium, etc., are also valid aspects of the present invention. [Effects of the Invention]
[0020] According to the present invention, it is possible to provide a technique for improving the accuracy of inference regarding image data in a large-scale language model. [Brief explanation of the drawings]
[0021] [Figure 1] FIG. 2 is a schematic diagram illustrating an overview of a process executed by an information processing device according to an embodiment. [Figure 2] FIG. 1 is a diagram schematically illustrating a functional configuration of an information processing device according to an embodiment. [Figure 3] FIG. 2 is a diagram schematically illustrating a data structure of task classification information stored in a storage unit. [Figure 4] FIG. 2 is a diagram schematically illustrating a data structure of task classification information stored in a storage unit. [Figure 5] 10 is a schematic diagram for explaining a target data group that is not acquired by a data acquisition unit. FIG. [Figure 6] FIG. 2 is a diagram schematically illustrating a data structure of task classification information stored in a storage unit. [Figure 7] 10 is a flowchart illustrating a flow of information processing executed by an information processing device according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0022] <Outline of the embodiment> The information processing device according to the embodiment analyzes a series of image data captured at intervals of a few seconds showing tasks related to convenience store operations, such as customer service, cleaning, and ordering, being performed in the store and back area of a convenience store (hereinafter referred to as "convenience store"), and infers the tasks captured in each image data.
[0023] In this specification, inferring a task captured in image data is referred to as “image task inference.” Note that a large-scale language model that uses image data as input and performs inference on image data is sometimes called a multimodal large language model (MLLM) or a vision language model (VLM), but in this specification it will be simply referred to as a large-scale language model.
[0024] The image data to be processed by the information processing device according to the embodiment is a series of image data in a time series. The information processing device stores task classification information that associates each image data, the capture time of each image data, a task name output as a result of image task inference performed by a large-scale language model for each image data, and a confidence level of the image task inference by the large-scale language model. Here, the confidence level is information output by the large-scale language model that performed the image task inference as a degree of confidence that the result is correct.
[0025] Fig. 1 is a schematic diagram for explaining an overview of processing executed by an information processing device 1 according to an embodiment. In the example shown in Fig. 1, the image data to be processed includes, for example, image data having a task name of "front display" (a task of moving a product from the back of a display shelf to the front) in which a convenience store clerk is captured picking up a product in front of the display shelf, and image data having a task name of "cleaning" in which a convenience store clerk is captured cleaning a product with cleaning tools.
[0026] Hereinafter, with reference to FIG. 1, an outline of the processes executed by the information processing device 1 according to the embodiment will be explained in the order of (1) to (6), and these numbers correspond to (1) to (6) in FIG.
[0027] (1) The information processing device 1 refers to the task classification information and identifies a high-certainty data group H, which is a data group with a relatively high degree of certainty, and a low-certainty data group L, which is a data group with a relatively low degree of certainty. Each piece of image data constituting the high-certainty data group H is one or more pieces of image data that are consecutive in time series, and the task names of the pieces of image data are the same. Each piece of image data constituting the low-certainty data group L is one or more pieces of image data that are consecutive in time series, and the task names of the pieces of image data are not necessarily the same.
[0028] In the example shown in Figure 1, the low-confidence data group L is made up of two chronologically consecutive image data pieces, and the task names of the two image data pieces are the same as above. In addition, the high-confidence data group H is chronologically consecutive to the low-confidence data group L. The high-confidence data group H is made up of four chronologically consecutive image data pieces, and the task names of the four image data pieces are the same as above.
[0029] (2) The information processing device 1 acquires, as the target data group T, one high-confidence data group H and one low-confidence data group L that is chronologically continuous with the high-confidence data group H. The one low-confidence data group L is a re-inference data group R that is the target of re-image task inference using the large-scale language model M. In the example shown in FIG. 1 , the information processing device 1 acquires, as the target data group T, the high-confidence data group H and the re-inference data group R, which is the low-confidence data group L.
[0030] (3) The information processing device 1 reads information from the target data group T and task classification information related to the image data included in the target data group T, and generates a prompt P using the information. The prompt P is a prompt for causing the large-scale language model M to perform image task inference for each piece of image data that constitutes the re-inference data group R. Figure 1 shows an example in which the prompt P is text data with the content "Is the task being performed for each image in the re-inference data group the same as that in the high confidence data group, 'cleaning'?"
[0031] (4) The information processing device 1 inputs the prompt P and the re-inference data group R, which is the low confidence data group L, into the large-scale language model M.
[0032] (5) The information processing device 1 acquires an inference result E, which is a result of image task inference output by the large-scale language model M. The inference result E has a format equivalent to, for example, task classification information.
[0033] (6) The information processing device 1 compares the inference result E with the task classification information related to the image data constituting the re-inference data group R, and if there is any information that differs as a result of the comparison, changes the information according to the information of the inference result E.
[0034] In this way, the information processing device 1 according to the embodiment performs image task inference for each piece of image data constituting the re-inference data group R. As described above, in the example shown in FIG. 1 , the image data constituting the low-certainty data group L captures an image of a convenience store clerk picking up a product in front of a display shelf. In this case, it may be difficult to determine from the image data alone that the task being performed is a pre-invention task, and as a result, the confidence level of the task determined to be a pre-invention task may be low. For example, consider a case where the task name of the high-certainty data group H, which is chronologically continuous with the low-certainty data group L, is cleaning, and an object resembling a cleaning tool drawer is captured in the image data constituting the low-certainty data group L. In this case, the act of picking up a product captured in each piece of image data constituting the low-certainty data group L is one step in the cleaning task, and there is a high probability that the task captured in the image data is cleaning.
[0035] In such a case, the information processing device 1 according to the embodiment has a high probability of inferring that cleaning is the task captured in each piece of image data constituting the low certainty data group L. In this way, the information processing device 1 according to the embodiment can improve the accuracy of inference for each piece of a series of image data in chronological order by utilizing the a priori information that the image data to be inferred is in chronological order.
[0036] <Functional configuration of information processing device 1 according to the embodiment> FIG. 2 is a diagram schematically illustrating the functional configuration of an information processing device 1 according to an embodiment. The information processing device 1 includes a storage unit 10, a communication unit 11, and a control unit 12. In FIG. 2, arrows indicate main data flows, and there may be data flows not shown in FIG. 2. In FIG. 2, each functional block indicates a configuration in functional units, rather than a configuration in hardware (device) units. Therefore, the functional blocks shown in FIG. 2 may be implemented in a single device, or may be implemented separately in multiple devices. Data may be exchanged between functional blocks via any means, such as a data bus, a network, or a portable storage medium.
[0037] The storage unit 10 is a large-capacity storage device such as a ROM (Read Only Memory) that stores the BIOS (Basic Input Output System) of the computer that realizes the information processing device 1, a RAM (Random Access Memory) that serves as the working area of the information processing device 1, an HDD (Hard Disk Drive) or an SSD (Solid State Drive) that stores the OS (Operating System), application programs, and various information referenced when the application programs are executed.
[0038] The communication unit 11 is a communication interface for the information processing device 1 to communicate with external devices, and is realized by a known communication module such as a LAN (Local Area Network) module or a Wi-Fi (registered trademark) module. Hereinafter, in this specification, description of the communication unit 11 may be omitted on the assumption that communication between the information processing device 1 and external devices is via the communication unit 11.
[0039] The control unit 12 is a processor such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), or NPU (Neural network Processing Unit) of the information processing device 1, and functions as an identification unit 120, a data acquisition unit 121, an inference control unit 122, and a change unit 123 by executing programs stored in the memory unit 10.
[0040] 2 shows an example in which the information processing device 1 is configured as a single device. However, the information processing device 1 may be realized by a plurality of processors, memories, and other computing resources, such as a cloud computing system. In this case, each unit constituting the control unit 12 is realized by at least one of a plurality of different processors executing a program.
[0041] The memory unit 10 stores task classification information that associates multiple image data showing the state of performing a task with the time when each image data was captured, and the task name and inference confidence level output as a result of image task inference performed by the large-scale language model M from each image data.
[0042] Here, image data will be described. First, in the above example, an example of an operational task at a convenience store, which is a retail business, was shown as the image capturing target of the image data. However, this is merely an example, and various business operations in various industries, such as manufacturing and logistics, may be captured as the image capturing target. Examples of image capturing means include, but are not limited to, a surveillance camera and smart glasses worn by the task performer. The image capturing may be performed periodically at predetermined time intervals, such as every few seconds or minutes, or may be performed as an event trigger based on the detection of motion, sound, or the like.
[0043] Here, the large-scale language model M receives text data and image data as input and analyzes and infers the content of the data. The large-scale language model M includes, but is not limited to, existing models that are publicly available, models that are fine-tuned from existing models, and models that are independently constructed by users of the information processing device 1.
[0044] 3 is a diagram schematically illustrating the data structure of task classification information stored in the storage unit 10. In the example illustrated in FIG. 3, the task classification information is information that associates an "image identifier" for identifying image data, "image data," the "capture time" of the image data, a "task name" output as a result of image task inference performed by the large-scale language model M on the image data, the "certainty" of the image task inference by the large-scale language model M, a "certainty data group identifier," and a "certainty data group type." The "certainty data group identifier" and "certainty data group type," as well as a case where the task name is "unknown," such as for the image identifier "IID005," will be described later.
[0045] In the example shown in Figure 3, image data "20250724_131700.jpg" with image identifier IID001 was captured at "2025 / 7 / 24 13:17:00." The task name of the task captured in the image data is "Previous Chang," and the confidence level is "0.9." The eight images with image identifiers IID001 to IID008 were each captured at 10-second intervals. In the example shown in Figure 3, the confidence level is expressed as a real value ranging from 0 as the minimum value to 1 as the maximum value, with the larger the value, the higher the likelihood of the result being correct.
[0046] The task classification information may further be associated with a task description that provides an overview of the task name. For example, the task description for the previous task may be "a task to move products from the back of the display shelf to the front." The task classification information may also be associated with the inference basis for the image task inference regarding the image data. For example, if the task name is cleaning, the inference basis may be "because it was observed that the person was cleaning the display shelf where the sweets were placed with cleaning tools."
[0047] Furthermore, the task classification information may further be associated with an image caption describing the content of the image data, or may be associated with an image caption instead of the image data. For example, an image caption for image data constituting the low-confidence data group L illustrated in FIG. 1 may be, "This image shows a male uniformed employee standing in front of a display shelf of snacks in a convenience store, gently lifting a package of snacks he has taken out in his right hand to inspect them. Products have been temporarily removed from some shelves, creating empty display space. Lighting, a surveillance camera, and another display shelf can be seen in the background." The image caption may be generated using a known technique, such as image captioning using a large-scale language model.
[0048] The identification unit 120 refers to the task classification information and identifies (1) a high-certainty data group H, which is one or more image data whose certainty is equal to or greater than a predetermined threshold, has the same task name, and is chronologically consecutive, and (2) a low-certainty data group L, which is other image data and is one or more image data that is chronologically consecutive.
[0049] Here, the predetermined threshold value for the confidence level is a threshold value used by the identification unit 120 to determine whether or not the image data constitutes the high confidence level data group H. Furthermore, since the low confidence level data group L is subject to image task inference again using the large-scale language model M, the predetermined threshold value for the confidence level can also be said to be a threshold value used by the identification unit 120 to determine whether or not image task inference should be performed again on the image data. The threshold value may be determined by experiment taking into consideration the tendency of the large-scale language model M to assign confidence levels, the purpose of use of the information processing device 1, the desired time to complete the entire process, etc., and is, for example, "0.7."
[0050] The results identified by the identification unit 120 are stored in association with task classification information. The results identified by the identification unit 120 may be stored in the data structure of the task classification information, or may be stored in a data structure different from the data structure of the task classification information. FIG. 3 shows an example in which the results identified by the identification unit 120 are stored in the data structure of the task classification information. Specifically, as described with reference to FIG. 3, the task classification information is information in which, in addition to an “image identifier,” “image data,” “capture time,” “task name,” and “certainty,” a “certainty data group identifier” and a “certainty data group type” are associated. The “certainty data group identifier” is identification information for identifying each of the high-certainty data group H and the low-certainty data group L identified by the identification unit 120. The “certainty data group type” indicates the type of the data group (hereinafter referred to as “certainty data group”) identified using the certainty and the task name, and is set to, for example, either a “high-certainty data group” or a “low-certainty data group.”
[0051] FIG. 3 is an example in which the predetermined threshold for certainty is "0.7." Image data with an image identifier IID001 (hereinafter, image data with an image identifier X will be referred to as "image data of X." Note that X is any of the image identifiers listed in FIGS. 3 to 6) has a task name of "Pre-Cleaning" and a certainty of "0.9." Image data with IID002 has a task name of "Cleaning" and a certainty of "0.9." Because the task names of the image data with IID001 and the image data with IID002 are different, only the image data with IID001 constitutes a high certainty data group H with a certainty data group identifier of DS001 and a certainty data group type of "high certainty data group."
[0052] Furthermore, the task name of the image data IID002 and the image data IID003 are the same, "cleaning," and both have a certainty of "0.9." The task name of the image data IID004 is "cleaning," the same as the image data ID003, but the certainty is "0.2," which is smaller than the predetermined threshold. Therefore, the image data IID002 and the image data IID003 constitute a high certainty data group H, whose certainty data group identifier is DS002 and whose certainty data group type is "high certainty data group."
[0053] 3, the image data IID004, the image data IID005, and the image data IID006 have different task names, but their respective certainties are lower than a predetermined threshold. Furthermore, the certainty of the image data IID003 and the image data IID007 is equal to or higher than a predetermined threshold. Therefore, the image data IID004, the image data IID005, and the image data IID006 constitute a low certainty data group L whose certainty data group identifier is DS003 and whose certainty data group type is "low certainty data group."
[0054] The data acquisition unit 121 acquires a target data group T including a high confidence data group H and a re-inference data group R which is a low confidence data group L that is continuous with the high confidence data group H in time series.
[0055] In the example shown in Figure 3, the target data group T may be acquired by using the certainty data group with the certainty data group identifier DS002 (hereinafter, the certainty data group with the certainty data group identifier X will be referred to as the "certainty data group of X", where X is any of the certainty data group identifiers shown in Figures 3 to 5) as the high certainty data group H, and the certainty data group of DS003 as the re-inference data group R, which is the low certainty data group L. Alternatively, the certainty data group of DS004 may be acquired as the high certainty data group H instead of the certainty data group of DS002.
[0056] The inference control unit 122 reads out the task classification information stored in the memory unit 10. The inference control unit 122 inputs to the large-scale language model M a prompt P and the re-inference data group R to cause the large-scale language model M to infer whether the task name of the task captured in each image data constituting the re-inference data group R is the task name associated with the image data constituting the high-confidence data group H included in the target data group T or is different from the task name.
[0057] The prompt P to be input to the large-scale language model M is generated, for example, by the inference control unit 122. Specifically, for example, a template of the prompt P is stored in the storage unit 10, and the inference control unit 122 reads the template of the prompt P and task classification information from the storage unit 10 and embeds the task classification information in the template of the prompt P to generate the prompt P.
[0058] The inference control unit 122 may cause the large-scale language model M to infer whether the task name of the task captured in the image data for each image data constituting the re-inference data group R is the task name associated with the image data constituting the high confidence data group H or is different from the task name. In this case, a result is output for each image data constituting the re-inference data group R that the task name is the task name or is different from the task name.
[0059] Furthermore, the inference control unit 122 may cause the large-scale language model M to infer whether the task name of the task captured in the image data for all of the image data constituting the re-inference data group R is the same as or different from the task name associated with the image data constituting the high-confidence data group H. In this case, a result indicating that the task name is the same or different from the task name is output for all of the image data constituting the re-inference data group R. When performing image task inference for all of the image data, the inference control unit 122 may, for example, output a result indicating that the task name is the same when it is inferred that the task name is the same for all of the image data. In this case, for example, the inference control unit 122 may output a result indicating that the task name is the same when it is inferred that a certain number or more of the image data among the entire image data is the same as the task name. Here, the certain number may specifically be, for example, one or a majority of the image data that are the subject of image task inference.
[0060] The inference control unit 122 may further input, to the large-scale language model M, any one or more of the above-mentioned task description, inference basis, and image caption for the image data constituting the high-confidence data group H included in the target data group T. Of the task description, inference basis, and image caption, the information input to the large-scale language model M is preferably associated with task classification information. This provides more basis for the large-scale language model M to infer whether the image data constituting the re-inference data group R and the image data constituting the high-confidence data group H perform the same task or different tasks, and the information processing device 1 can improve the inference accuracy of image task inference.
[0061] The inference control unit 122 may further input image data constituting the high confidence data group H included in the target data group T to the large-scale language model M. This provides even more basis for the large-scale language model M to infer whether the image data constituting the re-inference data group R and the image data constituting the high confidence data group H perform the same task or different tasks, and the information processing device 1 can improve the inference accuracy of image task inference.
[0062] The inference control unit 122 may input image captions for the image data constituting the re-inference data group R to the large-scale language model M, instead of the re-inference data group R. This makes it possible to expect improvements in the efficiency and speed of input / output, analysis, inference, and other processes in the information processing device 1 and the large-scale language model M, since the data size of the image captions, which are text data, is smaller than the data size of the image data.
[0063] The inference control unit 122 acquires the inference result E, which is the result of inference by the large-scale language model M. The inference result E is, for example, information in which an "image identifier," a "task name," and a "certainty" of the image task inference in this process are associated with each image data included in the re-inference data group R. The inference result E may further be associated with a task description and an inference basis.
[0064] The change unit 123 changes the task classification information in the storage unit 10 by referring to the result acquired by the inference control unit 122. The change unit 123 compares the inference result E with the task classification information related to the image data constituting the re-inference data group R, and if there is different information as a result of the comparison, changes the information based on the information in the inference result E.
[0065] The change unit 123 may update the "certainty data group identifier" and "certainty data group type" related to the image data constituting the re-inference data group R. This allows the storage unit 10 to keep the "certainty data group identifier" and "certainty data group type" up to date and maintain consistency within the task classification information. The change unit 123 may not need to update the "certainty data group identifier" and "certainty data group type" related to the image data constituting the re-inference data group R. This allows the change unit 123 to reduce the amount of processing performed. The change unit 123 may also initialize the "certainty data group identifier" and "certainty data group type" related to the image data constituting the re-inference data group R. This allows the information processing device 1 to determine whether or not the image data is undergoing re-image task inference by referring to the "certainty data group identifier" and "certainty data group type" of the image data.
[0066] When the modification unit 123 updates the "certainty data group type" for image data constituting the re-inference data group R, it may set either a "high certainty data group" or a "low certainty data group" based on the aforementioned predetermined threshold for certainty and the certainty included in the inference result E for the image data. When the modification unit 123 updates the "certainty data group identifier" for image data constituting the re-inference data group R, for example, if the image data can be included in an existing certainty data group, it may set the "certainty data group identifier" for the existing certainty data group. In this case, if the image data cannot be included in the existing certainty data group, the modification unit 123 may set a new "certainty data group identifier."
[0067] 3, updating of the confidence data group identifier will be described using an example in which a large-scale language model M performs image task inference, with the confidence data group of DS002 as the high confidence data group H and the confidence data group of DS003 as the re-inference data group R. Assume that the result of the image task inference is that the task name of the image data of IID005 is "cleaning" and the confidence is "0.9", and the task name and confidence of the image data of IID004 and IID006 are the same as before the image task inference in this process.
[0068] In this case, the image data of IID005 is a high-certainty data group H, and although it has the same task name as the high-certainty data group H of DS002, they are not chronologically continuous, so the image data of IID005 cannot be included in the high-certainty data group H of DS002. Therefore, the image data of IID005 cannot be included in any other existing certainty data groups, so a new certainty data group identifier, such as "DS100," is set. Similarly, the image data of IID004 and the image data of IID006 are not chronologically continuous, so a new "certainty data group identifier" is set for either image data. For example, the certainty data group identifier of the image data of IID004 is the existing "DS003," and a new certainty data group identifier, such as "DS101," is set for the image data of IID006.
[0069] As described above, by utilizing the prior information that the image data to be inferred is in chronological order, the information processing device 1 can improve the inference accuracy of image task inference for each of a series of image data in chronological order. Specifically, even for image data with a low degree of certainty in the initial image task inference, the information processing device 1 can improve the inference accuracy of image task inference by having the large-scale language model M infer whether the task name of the task captured in the image data with a low degree of certainty in the initial image task inference is the task name of image data whose image capture times are consecutive in chronological order and whose degree of certainty in the image task inference is relatively high.
[0070] In particular, the information processing device 1 is expected to further improve the inference accuracy of image task inference when the amount of image data that can be subjected to image task inference at one time is limited due to, for example, an upper limit on the number of input tokens of the large-scale language model M. In this case, for example, suppose there are 30 pieces of image data that capture images of someone performing a cleaning task at 10-second intervals, and image task inference is performed on the 29 pieces of image data that are chronologically leading at once, and image task inference is performed separately on the remaining piece of image data.
[0071] In this case, the confidence level of the image task inference performed on one of the 30 pieces of image data may be low, and the output task name may be incorrect. Reasons for low inference accuracy include, but are not limited to, a lack of features necessary for inference or noise in the image. For the image data with low confidence level, if the task classification information of the 29 pieces of image data preceding it in time series indicates that cleaning was the task performed for more than four minutes before the image data was captured, and the confidence level of this inference is high, then a second image task inference will likely be able to infer that the task captured in the image data is also cleaning.
[0072] The identification unit 120 may further identify the low-certainty data group L as (1) a low-certainty data group L that is one or more image data whose certainty is less than a predetermined threshold, has the same task name, and is chronologically consecutive, and (2) an unclassified data group that is one or more image data whose task name is inferred to be unknown and is chronologically consecutive.
[0073] Here, image data for which the task name is inferred to be unknown refers to image data for which, for example, when image task inference is performed on the image data using the large-scale language model M, it is not possible to infer what task is captured in the image data, and so the result "unknown" is output as the task name. In the example shown in FIG. 3, the image data with IID005 has the task name "unknown." Note that the certainty factor of the image data with IID005 is "0," but "-" may be set to indicate that the certainty factor is not set.
[0074] In this specification, further identification by the identification unit 120 based on the three types of data, i.e., the high-certainty data group H, the low-certainty data group L, and the unclassified data group, will be referred to as "identification by three types." Also, hereinafter in this specification, identification by the identification unit 120 based on the two types of data, i.e., the high-certainty data group H and the low-certainty data group L, will be referred to as "identification by two types." Note that FIG. 3 shows an example in which the identification results by two types are stored and included in the data structure of the task classification information. Identification by two types has been described above with reference to FIG. 3, etc. Also, in this specification, when there is no need to distinguish between the low-certainty data group L identified by identification by two types and the low-certainty data group L identified by identification by three types, they will simply be referred to as the low-certainty data group L.
[0075] FIG. 4 is a diagram schematically illustrating the data structure of task classification information stored in the storage unit 10. FIG. 4 illustrates an example in which the identification results based on three types are stored in the same data structure of task classification information as FIG. 3. Specifically, the task classification information is information that associates an "image identifier," "image data," "capture time," "task name," and "certainty," as well as a "certainty data group identifier" and a "certainty data group type." The "certainty data group identifier" is identification information for identifying each of the certainty data groups identified by the identification unit 120. As described above, the certainty data group is a data group identified using the certainty and task name. The "certainty data group type" indicates the type of the certainty data group, and is set to, for example, one of a "high certainty data group," a "low certainty data group," and an "unclassified data group."
[0076] Identification by three types will be described with reference to FIG. 4. FIG. 4 shows an example in which the predetermined threshold for certainty is "0.7," as in FIG. 3, and the identification of the high certainty data group H is the same as in FIG. 3. The certainty of the image data of IID004 is less than 0.7. Furthermore, there is no image data that is chronologically consecutive with the image data of IID004, has a certainty of less than 0.7, and has the same task name. Therefore, a certainty data group with a certainty data group identifier of DS003 and a certainty data group type of "low certainty data group" is composed only of the image data of IID004. Similarly, a certainty data group with a certainty data group identifier of DS005 and a certainty data group type of "low certainty data group" is composed only of the image data of IID006. The task name of the image data of IID005 is unknown, and there is no chronologically consecutive image data with an unknown task name. Therefore, a certainty data group with a certainty data group identifier of DS004 and a certainty data group type of "unclassified data group" is composed of only the image data of IID005.
[0077] The results of identifying the same task classification information using two types are shown in Figure 3. In the example shown in Figure 3, three pieces of image data with image identifiers IID004 to IID006 constitute a certainty data group with a certainty data group identifier of DS003 and a certainty data group type of "low certainty data group."
[0078] As a result, the image data constituting the low-confidence data group L, which is the re-inference data group R, are limited to image data all having the same task name or all having an unknown task name. The further chronologically the image data constituting the low-confidence data group L, which is the re-inference data group R, is from the high-confidence data group H included in the target data group T, the lower the likelihood that it has the task name of the image data constituting the high-confidence data group H. Furthermore, among the image data constituting the re-inference data group R, image data with a task name different from the task name of image data chronologically adjacent to the high-confidence data group H included in the target data group T has a relatively high likelihood that the task has been switched from that adjacent image data. Therefore, image data with a task name different from the task name of the adjacent image data has a relatively low likelihood of having the task name of the image data constituting the high-confidence data group H.
[0079] Therefore, the information processing device 1 can prevent image data with a lower probability of being the task name of image data constituting the high-confidence data group H included in the target data group T from being input into the large-scale language model M. As a result, the information processing device 1 can improve the efficiency and speed of processing such as input / output, analysis, and inference in the information processing device 1 and the large-scale language model M.
[0080] The data acquisition unit 121 may acquire a target data group T including a high-certainty data group H with a high degree of certainty and the re-inference data group R when the high-certainty data groups H are both chronologically preceding and following the re-inference data group R and the certainty of each high-certainty data group H is different.
[0081] As a result, the higher the confidence level, the higher the probability that it is associated with the correct task name, so the information processing device 1 can input a high confidence data group H that is associated with a more correct task name into the large-scale language model M, thereby improving the inference accuracy of image task inference.
[0082] When the high confidence data groups H are both located chronologically before and after the re-inference data group R and the confidence levels of the respective high confidence data groups H are the same, the data acquisition unit 121 may acquire a target data group T including any one of the high confidence data groups H and the re-inference data group R based on a predetermined rule. The predetermined rule may be, for example, acquiring a target data group T including a chronologically earlier high confidence data group H and the re-inference data group R.
[0083] When the inference control unit 122 obtains a negative result that the task name is different, the data acquisition unit 121 does not need to acquire the target data group T that includes both the high-confidence data group H that includes the image data used to input the negative result and the re-inference data group R that includes the image data.
[0084] FIG. 5 is a schematic diagram illustrating a target data group T that is not acquired by the data acquisition unit 121. FIG. 5(a) schematically illustrates the data structure of task classification information before the information processing device 1 causes the large-scale language model M to perform image task inference again. FIG. 5(b) schematically illustrates the data structure of the inference result E of the image task inference that the information processing device 1 causes the large-scale language model M to perform. FIG. 5(c) schematically illustrates the data structure of task classification information in which all items are up to date after the information processing device 1 causes the large-scale language model M to perform the image task inference. Below, with reference to FIG. 5, a case in which the data acquisition unit 121 does not acquire the target data group T will be described. Note that FIG. 5 illustrates an example in which the predetermined threshold for the confidence level is "0.7".
[0085] In the example shown in FIG. 5(a), let us assume that the confidence data group of DS001 is the high confidence data group H, and the confidence data group of DS002 is the re-inference data group R, which is the low confidence data group L, and that the large-scale language model M performs image task inference again. In this example, the task name of the image data of IID001 that constitutes the confidence data group of DS001, which is the high confidence data group H, is "Previous". Also in this example, the task names of the image data of IID002 and IID003 that constitute the confidence data group of DS002 are both "Cleaning".
[0086] Figure 5(b) shows an example of the inference result E of the image task inference. In the example shown in Figure 5(b), the inference result E shows that the task name of the image data of IID002 is "Pre-Chem" and the confidence level is "0.8", and the task name of the image data of IID003 is "Cleaning" and the confidence level is "0.2". In this example, the inference result E for the image data of IID003 is a negative result in that the task name and confidence level are the same as those before the image task inference was performed, and are different from the task name "Pre-Chem" in the high confidence level data group H.
[0087] Figure 5(c) shows a schematic diagram of the data structure of the task classification information after the task inference has been performed. In the example shown in Figure 5(c), the certainty data group of DS001, whose certainty data group type is "high certainty data group," is composed of image data IID001 and image data IID002, and the task name of the image data IID002 is "pre-entry" and the certainty is "0.8." In this example, the certainty data group of DS002, whose certainty data group type is "low certainty data group," is composed only of image data IID003, and the task name of the image data IID003 is "cleaning" and the certainty is "0.2."
[0088] In the case shown in Fig. 5, the image data used to input a negative result for the image data of IID003 are three pieces of image data with image identifiers IID001 to IID003. Also, in the example shown in Fig. 5(c), both the confidence data group of DS001 and the confidence data group of DS002 include image data of any of the three pieces of image data with image identifiers IID001 to IID003. Therefore, in the example shown in Fig. 5(c), the data acquisition unit 121 does not acquire a target data group T that includes both the confidence data group of DS001 as the high confidence data group H and the confidence data group of DS002 as the re-inference data group R.
[0089] 5(c), the high-certainty data group H, which is the confidence data group of DS003, does not include the three image data with image identifiers IID001 to IID003, which are the image data used to input the negative result. Therefore, in this case as well, the data acquisition unit 121 acquires, as the high-certainty data group H, a target data group T that includes both the confidence data group of DS003 and the confidence data group of DS002 as the re-inference data group R.
[0090] Identification by the identification unit 120, acquisition by the data acquisition unit 121, acquisition by the inference control unit 122, and modification by the modification unit 123 are repeated until there is no target data group T to be acquired by the data acquisition unit 121.
[0091] As a result, the information processing device 1 repeats a specific operation until there is no more image data to be subjected to further image task inference. This allows the information processing device 1 to improve processing efficiency and speed, for example, by reducing overhead related to the start and end of processing.
[0092] The information processing device 1 may set an upper limit to the number of repetitions. The upper limit may be determined by experiment, taking into consideration the quantity of image data associated with the task classification information in the storage unit 10, the desired time to complete the entire process, and the like, and may be, for example, "100." This allows the information processing device 1 to prevent a situation in which image data that is the target of re-image task inference continues to exist, making it impossible to complete the process.
[0093] The above mainly describes a case where the information processing device 1 improves the inference accuracy of image task inference for each of a series of image data in a time series by utilizing a priori information that the image data to be inferred is in a time series. Next, a case where the task name output as a result of image task inference is changed depending on the purpose of use of the information processing device 1 will be described. For example, in the convenience store example mentioned above, the granularity of the task name can be relatively more abstract, such as interpersonal work or non-interpersonal work, or relatively more specific, such as cleaning display shelves or cleaning aisles. First, a description will be given of changing the task name output as a result of image task inference to a more specific task name.
[0094] The identifying unit 120 further identifies a task name data group, which is one or more pieces of image data that have the same task name and are chronologically consecutive.
[0095] The task name data group identified by the identification unit 120 is stored in association with task classification information. The task name data group identified by the identification unit 120 may be stored as being included in the data structure of the task classification information, or may be stored as a data structure different from the data structure of the task classification information.
[0096] Fig. 6 is a diagram schematically illustrating the data structure of task classification information stored in the storage unit 10. Fig. 6 shows an example in which a task name data group is stored as part of the data structure of the task classification information. In the example shown in Fig. 6, the task classification information is information in which a "task name data group identifier," which is identification information for identifying the task name data group further identified by the identification unit 120, is associated with the "image identifier," "image data," "capture time," "task name," and "certainty factor" included in the task classification information described with reference to Fig. 3.
[0097] In the example shown in Figure 6, the task name data group with the task name data group identifier NDS001 is made up of three image data with image identifiers IID001 to IID003. These three image data have the same task name, "Preparing cleaning tools and tidying up," and are chronologically consecutive. The image data IID004, which is chronologically consecutive to the image data IID003, has the task name "Cleaning execution," and therefore constitutes a different task name data group from the task name data group with the task name data group identifier NDS001.
[0098] The data acquisition unit 121 further acquires a rename data group including a task name data group.
[0099] The inference control unit 122 inputs the renamed data group and a prompt P to cause the large-scale language model M to infer a task name that embodies the task name associated with the task name data group that constitutes the renamed data group to the large-scale language model M. The inference control unit 122 acquires an inference result E that is the result of the inference by the large-scale language model M.
[0100] Here, the prompt P to be input to the large-scale language model M is generated, for example, by the inference control unit 122. Specifically, for example, a template of the prompt P is stored in the storage unit 10, and the inference control unit 122 reads the template of the prompt P and task classification information from the storage unit 10. Then, for example, the inference control unit 122 generates the prompt P by embedding the task classification information in the template of the prompt P. Note that in addition to the prompt P and the renamed data group, the inference control unit 122 is required to input, to the large-scale language model M, the task name associated with the task name data group that constitutes the renamed data group. Therefore, for example, the inference control unit 122 may include, in the prompt P, the task name associated with the task name data group that constitutes the renamed data group.
[0101] In the example shown in FIG. 6, for example, the data acquisition unit 121 acquires a task name data group having a task name data group identifier of NDS002 as a renaming data group (hereinafter, a task name data group having a task name data group identifier of X will be referred to as the "task name data group of X." Note that X is one of the task name data group identifiers shown in FIG. 6). In this case, the prompt P specifically has content such as, for example, "Please infer a task name that embodies the task name "Cleaning" of the task captured in the input image data." In this case, the inference control unit 122 acquires an inference result E that includes, for example, the task name "Cleaning the display shelves."
[0102] As a result, in cases where the task name output by the large-scale language model M is at an abstract level relative to the intended use of the information processing device 1, the information processing device 1 can have the large-scale language model M infer a more specific task name, thereby improving the usefulness of the task classification information that reflects the processing results of the information processing device 1.
[0103] The above has described changing the task name output as a result of image task inference to a more specific task name, among other cases where the task name output as a result of image task inference is changed depending on the purpose of use of the information processing device 1. Next, we will explain changing each task name of multiple class name data groups to one common task name.
[0104] The data acquisition unit 121 may further acquire a rename data group including a plurality of task name data groups that are chronologically continuous.
[0105] The data acquisition unit 121 may limit the renamed data set to be acquired, taking into consideration the memory capacity and processing performance of the information processing device 1, an upper limit on the number of input tokens of the large-scale language model M, etc. For example, the data acquisition unit 121 may set an upper limit on the number of image data items to be acquired, such as a maximum of 30. This reduces the processing load on the information processing device 1 and the large-scale language model M, enabling stable data processing.
[0106] The data acquisition unit 121 may also limit the target data set T to be acquired depending on the purpose of use of the information processing device 1 and the characteristics of the business, etc., that are the subject of image data capture. For example, suppose the purpose of use of the information processing device 1 is to divide the content of the business to be captured into relatively short tasks that can be completed within 10 minutes and visualize the progress status and progress rate of each task. In this case, the data acquisition unit 121 may acquire the renamed data set such that, for example, the interval between the capture time of the oldest image data and the capture time of the newest image data among the image data included in the renamed data set does not exceed 10 minutes. This allows the information processing device 1 to improve the usefulness of the task classification information that reflects the processing results of the information processing device 1.
[0107] The inference control unit 122 inputs the renamed data group and a prompt P to the large-scale language model M to cause the large-scale language model M to infer a common task name for changing each task name associated with the task name data group that constitutes the renamed data group. Then, the inference control unit 122 obtains an inference result E that is the result of inference by the large-scale language model M. Note that in addition to the prompt P and the renamed data group, the inference control unit 122 must also input each task name associated with the task name data group that constitutes the renamed data group to the large-scale language model M. Therefore, the inference control unit 122 may, for example, include each task name associated with the task name data group that constitutes the renamed data group in the prompt P.
[0108] 6, for example, the data acquisition unit 121 acquires the task name data group NDS001, the task name data group NDS002, and the task name data group NDS003 as the renaming data group. In this case, the prompt P specifically has content such as, for example, "Please infer the task name common to the tasks captured in the input image data." In this case, the inference control unit 122 acquires the inference result E including, for example, the task name "cleaning."
[0109] As a result, in cases where the task names output by the large-scale language model M are divided into tasks that can be completed in a shorter time than the intended use of the information processing device 1, the information processing device 1 can have the large-scale language model M infer task names that are more suited to the intended use, thereby improving the usefulness of the task classification information that reflects the processing results of the information processing device 1.
[0110] The inference control unit 122 may input a task name list that lists examples of task names to be output into the large-scale language model M. Even if the business that is the subject of the image data capture is the same, the task names included in the business may differ if the granularity of the inferred task or the division unit based on the estimated completion time, as described above, is different. Furthermore, even for the same task, for example, in the convenience store example described above, the task names may differ depending on whether the task name is a general-purpose task name common to the retail industry as a whole or a task name specific to a particular convenience store chain. Therefore, the inference control unit 122 may input a task name list that lists task names that indicate the granularity, versatility, etc. of task names that match the purpose of use of the information processing device 1 into the large-scale language model M.
[0111] This allows the information processing device 1 to have the large-scale language model M infer task names that are more suited to the intended use, thereby improving the usefulness of task classification information that reflects the processing results of the information processing device 1.
[0112] The inference control unit 122 may input a prompt P to the large-scale language model M, which includes an instruction to infer a common task name for changing each task name, when the tasks associated with the task name data group that constitutes the rename data group are similar to each other.
[0113] 6, for example, it is assumed that the data acquisition unit 121 acquires the task name data group of NDS001, the task name data group of NDS002, and the task name data group of NDS003 as the renaming data group. In this case, the prompt P specifically has content such as, for example, "If the task "Prepare cleaning tools and tidy up" and the task "Perform cleaning" are similar, please infer the task name that is common to the tasks captured in the input image data."
[0114] This allows the information processing device 1 to cause the large-scale language model M to infer task names that are more suited to the intended use of the information processing device 1, thereby improving the usefulness of task classification information that reflects the processing results of the information processing device 1. Specifically, when the large-scale language model M infers a common task name for tasks that it determines are dissimilar, the common task name may have a more abstract granularity than the intended use of the information processing device 1. For example, more specifically, it is assumed that the large-scale language model M determines that the task of cleaning and the task of customer service are dissimilar. In this case, the large-scale language model M may infer a relatively more abstract granularity task name, such as "convenience store operation," as the common task name for the tasks. The information processing device 1 can prevent the inference of such a relatively more abstract granularity common task name when such a relatively more abstract granularity task name is not suited to the intended use.
[0115] The inference control unit 122 may input to the large-scale language model M a prompt P including an instruction to make the number of types of task names included in the task classification information fall within a predetermined range.
[0116] Here, the predetermined range is a range for the large-scale language model M to determine whether or not the large-scale language model M infers a common task name. The predetermined range may be determined through experiments to obtain a granularity of task names that matches the purpose of use, taking into consideration the purpose of use of the information processing device 1 and the characteristics of the business that is the subject of image data capture, and may be, for example, "a range from 100 to 120." This allows the information processing device 1 to cause the large-scale language model M to infer task names that are more suited to the purpose of use of the information processing device 1, thereby improving the usefulness of task classification information that reflects the processing results of the information processing device 1.
[0117] The inference control unit 122 may input a prompt P including information on all task names included in the task classification information to the large-scale language model M. In this way, the information processing device 1 can make the large-scale language model M use information on all task names included in the task classification information, thereby more reliably keeping the number of types of task names included in the task classification information within a predetermined range, and improving the usefulness of the task classification information in which the processing results of the information processing device 1 are reflected.
[0118] The data acquisition unit 121 does not need to acquire the renamed data group input to the large-scale language model M again as a renamed data group.
[0119] 6, for example, suppose the data acquisition unit 121 acquires the task name data group of NDS001 and the task name data group of NDS002 as the renamed data groups, but does not infer a common task name because the tasks in each task name data group are not similar to each other. In this case, the data acquisition unit 121 does not subsequently acquire the task name data group of NDS001 or the task name data group of NDS002 as the renamed data groups. Note that because the combinations of task name data groups constituting the renamed data groups are different, even in this case, the data acquisition unit 121 acquires the task name data group of NDS001, the task name data group of NDS002, and the task name data group of NDS003 as the renamed data groups.
[0120] The identification by the identification unit 120, the acquisition by the data acquisition unit 121, the acquisition by the inference control unit 122, and the change by the change unit 123 may be repeated until there are no renamed data groups to be acquired by the data acquisition unit 121.
[0121] As a result, the information processing device 1 repeats a specific operation until there is no more image data for which a common name is to be inferred, thereby enabling the information processing device 1 to improve processing efficiency and speed, for example, by reducing overhead related to the start and end of processing.
[0122] The information processing device 1 may set an upper limit to the number of repetitions. The upper limit may be determined experimentally taking into consideration the quantity of image data associated with the task classification information in the storage unit 10, the desired time to complete the entire process, and the like, and may be, for example, "100." This allows the information processing device 1 to prevent a situation in which the renamed data group that is the subject of inference continues to exist, making it impossible to complete the process.
[0123] <Processing flow of information processing method executed by information processing device 1> 7 is a flowchart for explaining the flow of information processing executed by the information processing device 1 according to the embodiment. The processing in this flowchart starts, for example, when the information processing device 1 is started.
[0124] The inference control unit 122 reads out task classification information from the storage unit 10 (S1). The identification unit 120 refers to the task classification information stored in the storage unit 10 and identifies two or three types of confidence data groups by identifying by two types or by identifying by three types (S2). The data acquisition unit 121 acquires a target data group T including a re-inference data group R composed of image data that is the target of re-image task inference (S3).
[0125] The inference control unit 122 inputs the prompt P and the re-inference data group R to the large-scale language model M (S4). The prompt P is a prompt for causing the large-scale language model M to perform image task inference for each piece of image data that constitutes the re-inference data group R. The prompt P is generated by, for example, the inference control unit 122.
[0126] The inference control unit 122 acquires the inference result E including the result of the image task inference output by the large-scale language model M (S5). The change unit 123 refers to the inference result E and changes the task classification information stored in the storage unit 10 (S6). When the change unit 123 changes the task classification information, the processing in this flowchart ends.
[0127] <Advantages of the information processing device 1 according to the embodiment> As described above, according to the information processing device 1 according to the embodiment, the accuracy of inference regarding image data in the large-scale language model M can be improved.
[0128] Furthermore, this invention will make it possible to contribute to Goal 9 of the United Nations' Sustainable Development Goals (SDGs), which is "Build resilient infrastructure, promote inclusive and sustainable industrialization, and promote innovation and resilience."
[0129] The present invention has been described above using embodiments, but the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. For example, all or part of the device can be configured by functionally or physically distributing or integrating any unit. Furthermore, new embodiments resulting from any combination of multiple embodiments are also included in the embodiments of the present invention. The effects of the new embodiments resulting from the combination also have the effects of the original embodiments. [Explanation of symbols]
[0130] 1. Information processing device 10...Storage section 11. Communications Department 12 Control section 120...Specific part 121 Data Acquisition Unit 122 Inference control unit 123...Changes M···Large-scale language model
Claims
1. a storage unit that stores task classification information that associates a plurality of image data in which a state of performing a task is captured, the image capture time of each image data, the task name of the task inferred from each image data by a large-scale language model, and the degree of certainty of the inference; an identification unit that refers to the task classification information and identifies (1) a high-certainty data group that is one or more of the image data whose certainty is equal to or greater than a predetermined threshold, whose task name is the same, and which are consecutive in time series, and (2) a low-certainty data group that is other image data and is one or more of the image data which are consecutive in time series; a data acquisition unit that acquires a target data group including the high confidence data group and a re-inference data group that is the low confidence data group that is time-series continuous with the high confidence data group; a prompt for causing the large-scale language model to infer whether the task name of the task captured in each of the image data constituting the re-inference data group is the task name associated with the image data constituting the high confidence data group included in the target data group or different from the task name; and an inference control unit that inputs the re-inference data group into the large-scale language model and acquires the inference result. a change unit that changes the task classification information in the storage unit by referring to the result obtained by the inference control unit; An information processing device comprising:
2. The identification unit further identifies, as the low-certainty data group, (1) a low-certainty data group in which the certainty is less than a predetermined threshold, the task name is the same, and the image data is one or more pieces of image data that are chronologically consecutive, and (2) an unclassified data group in which the task name is inferred to be unknown and the image data is one or more pieces of image data that are chronologically consecutive. The information processing device according to claim 1 .
3. the data acquisition unit acquires the target data group including the high confidence data group with the high confidence data group having the high confidence data group with the high confidence data group having the high confidence data group and the re-inference data group when the high confidence data group is both chronologically preceding and following the re-inference data group and the high confidence data group has a different confidence level; 3. The information processing device according to claim 1.
4. When the inference control unit obtains a negative result that the task name is different from the task name, the data acquisition unit does not acquire the target data group that includes both the high confidence data group that includes the image data used to input the negative result and the re-inference data group that includes the image data, The identification by the identification unit, the acquisition by the data acquisition unit, the acquisition by the inference control unit, and the change by the change unit are repeated until there are no more target data groups to be acquired by the data acquisition unit.
3. The information processing device according to claim 1.
5. the identifying unit further identifies a task name data group, which is one or more pieces of image data having the same task name and which are chronologically consecutive; the data acquisition unit further acquires a renaming data group including the task name data group; the inference control unit further acquires a prompt for causing the large-scale language model to infer a task name that embodies the task name associated with the task name data group constituting the renamed data group, and a result of inference performed by inputting the renamed data group into the large-scale language model; 3. The information processing device according to claim 1.
6. the identifying unit further identifies a task name data group, which is one or more pieces of image data having the same task name and which are chronologically consecutive; the data acquisition unit further acquires a renaming data group including a plurality of the task name data groups that are chronologically continuous; the inference control unit further acquires a prompt for causing the large-scale language model to infer a common task name for changing each task name associated with the task name data group constituting the renamed data group, and a result of inference performed by inputting the renamed data group into the large-scale language model; The information processing device according to claim 1 .
7. the inference control unit inputs a list of task names, which lists examples of the output task names, into the large-scale language model; The information processing device according to claim 6 .
8. the inference control unit inputs the prompt to the large-scale language model, the prompt including an instruction to infer the common task name for changing the names of the tasks associated with the task name data group constituting the rename data group when the tasks are similar to each other; The information processing device according to claim 6 .
9. the inference control unit inputs the prompt, which includes an instruction to make the number of types of task names included in the task classification information fall within a predetermined range, to the large-scale language model; The information processing device according to claim 6 .
10. the data acquisition unit does not acquire the renamed data group input to the large-scale language model again as the renamed data group, repeating the identification by the identification unit, the acquisition by the data acquisition unit, the acquisition by the inference control unit, and the change by the change unit until there are no more renamed data groups to be acquired by the data acquisition unit; 10. The information processing device according to claim 8.
11. The processor: reading from a storage unit task classification information that associates a plurality of image data in which a state of performing a task is captured, the image capture time of each image data, the task name of the task inferred from each image data by a large-scale language model, and the degree of certainty of the inference; A step of referring to the task classification information to identify (1) a high-certainty data group, which is one or more pieces of image data whose certainty is equal to or greater than a predetermined threshold, whose task name is the same, and which are consecutive in time series, and (2) a low-certainty data group, which is other image data and is one or more pieces of image data that are consecutive in time series; acquiring a target data group including the high confidence data group and a re-inference data group, which is the low confidence data group that is time-series continuous with the high confidence data group; a step of inputting the re-inference data group into the large-scale language model and acquiring an inference result, and a prompt for inferring whether the task name of the task captured in each of the image data constituting the re-inference data group is the task name associated with the image data constituting the high confidence data group included in the target data group or different from the task name; and changing the task classification information in the storage unit by referring to the acquired result; An information processing method that performs the above.
12. On the computer, a function of reading from a storage unit task classification information that associates multiple image data in which a task is being performed, the image capture time of each image data, the task name of the task inferred from each image data by a large-scale language model, and the degree of certainty of the inference; A function of referring to the task classification information to identify (1) a high-certainty data group, which is one or more pieces of image data whose certainty is equal to or greater than a predetermined threshold, whose task name is the same, and which are consecutive in time series, and (2) a low-certainty data group, which is other image data and is one or more pieces of image data that are consecutive in time series; a function of acquiring a target data group including the high confidence data group and a re-inference data group, which is the low confidence data group whose time series is continuous with the high confidence data group; a prompt for causing the large-scale language model to infer whether the task name of the task captured in each of the image data constituting the re-inference data group is the task name associated with the image data constituting the high confidence data group included in the target data group or different from the task name; and a function for inputting the re-inference data group into the large-scale language model and acquiring the inference result; a function of changing the task classification information in the storage unit by referring to the acquired result; A program to make this happen.
Citation Information
Patent Citations
Artificially intelligent assistant for work protocols
JP2025121863A
Location identification system and program
JP7633733B1