Identification device, identification program, and identification method
The use of large-scale language models for zero-shot classification across different modalities addresses the cost and time issues of conventional methods, enhancing identification accuracy by utilizing common sense and multiple data types.
Patent Information
- Application Number
- JP2024105636
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2026-01-16
Smart Images

Figure 2026006561000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an identification device, an identification program, and an identification method. [Background technology]
[0002] In recent years, technological development of large-scale models such as large-scale language models (LLM) and large-scale visual language models (VLM) has progressed. For example, Non-Patent Document 1 proposes preparing 43,741 combinations of tactile data, image data, and annotations, and using the prepared combinations to fine-tune the large-scale visual language model LLaMA2 using LoRA. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Letian Fu, et al. “A Touch, Vision, and Language Dataset for Multimodal Alignment”, [online], [Retrieved June 17, 2020], Internet<URL:https: / / tactile-vlm.github.io> [Non-patent document 2] Sai Shashank Kalakonda, et al. “Action-GPT: Leveraging Large-scale Language Models for Improved and Generalized Action Generation”, [online], [Retrieved June 17, 2024], Internet<URL:https: / / arxiv.org / abs / 2211.15603> [Non-patent document 3] OpenAI, “GPT-4 Technical Report”, [online], [Retrieved June 17, 2024], Internet<URL:https: / / arxiv.org / abs / 2303.08774> [Non-patent document 4] Zhengyuan Yang, et al. “The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)”, [online], [Retrieved June 17, 2024], Internet<URL:https: / / arxiv.org / abs / 2309.17421> [Non-patent document 5] Swarup Ranjan Behera, et al. “AQUALLM: Audio Question Answering Data Generation Using Large Language Models”, [online], [Retrieved June 17, 2024], Internet<URL:https: / / arxiv.org / abs / 2312.17343v1> [Non-patent document 6] Aaron van den Oord, et al. “Neural Discrete Representation Learning”, [online], [Retrieved June 17, 2024], Internet<URL:https: / / arxiv.org / abs / 1711.00937> Summary of the Invention [Problem to be solved by the invention]
[0004] The present inventors have found that the above-mentioned conventional methods have the following problems. Conventional methods build large-scale models corresponding to modalities for which no large-scale models have been obtained (tactile data in the method of Non-Patent Document 1). Fine-tuning a large-scale model for each new modality requires time and effort, such as collecting training samples, which increases costs. In other words, with conventional methods, building a classification system using a large-scale model for any modality can be expensive.
[0005] In one aspect, the present invention has been made in view of the above circumstances, and an object of the present invention is to provide a technique for reducing the cost required to build an identification system in any modality. [Means for solving the problem]
[0006] In order to solve the above-mentioned problems, the present invention employs the following configuration. The above configurations can be combined as appropriate.
[0007] According to one aspect of the present invention, an identification device includes a control unit configured to: obtain a text representation of a first sample of a first data type, the first sample being generated by observing an object with a first sensor, according to a category to which the first sample belongs; provide a prompt to a large-scale language model, the prompt including a command statement for instructing the large-scale language model to identify a feature in the object and the text representation of the first sample; and obtain an answer indicating a result of identifying the feature from the large-scale language model; and output information about the obtained result.
[0008] Some studies, such as Non-Patent Document 2, have reported that large-scale language models can acquire common sense. This configuration utilizes this property to achieve zero-shot classification. That is, a text representation corresponding to the category of a first sample is provided to the large-scale language model. The large-scale language model can derive a feature classification result from the category of the first sample in reference to common sense. Therefore, this configuration makes it possible to achieve zero-shot classification of target features without fine-tuning the large-scale language model. This reduces the cost of building a classification system for any modality.
[0009] In the identification device according to the above aspect, the prompt may further include a list of labels that are candidates for the identification result of the feature, and identifying the feature may be configured by selecting a label corresponding to the feature from the list. With this configuration, when an identification range is given, information indicating the identification range can be provided to the large-scale language model as a list, thereby narrowing the range of the identification result. This can be expected to improve identification accuracy compared to when the identification range is unlimited.
[0010] In the identification device according to the above aspect, the control unit may further be configured to acquire second samples of a second data type different from the first data type, the second samples being generated by observing an object with a second sensor. The large-scale language model may be configured to accept input of the second data type samples. The prompt may further include the second samples.
[0011] Large-scale language models that can accept input of data other than text, such as large-scale visual language models, already exist. This configuration can reduce the cost of building a classification system when using such large-scale language models that can accept input of data in modalities other than text. Furthermore, the use of such large-scale language models can be expected to improve classification accuracy.
[0012] In the classification device according to the above aspect, the second data type may be image data. With this configuration, it is possible to reduce the cost required to build a classification system in a situation where a large-scale visual language model is used.
[0013] In the identification device according to the above aspect, the first data type may be tactile data, and the second data type may be image data. With this configuration, it is possible to reduce the cost of building an identification system when identifying features using tactile data and image data.
[0014] In the identification device according to the above aspect, the first data type may be temperature data, and the second data type may be image data. With this configuration, when identifying features using temperature data and image data, it is possible to reduce the cost required to build an identification system. It is possible.
[0015] In the identification device according to the above aspect, the first data type may be sound data, and the second data type may be image data. With this configuration, when identifying features using sound data and image data, it is possible to reduce the cost required to build an identification system.
[0016] In the identification device according to the above aspect, the first data type may be point cloud data, and the second data type may be image data. With this configuration, when identifying features using point cloud data and image data, it is possible to reduce the cost required to build an identification system.
[0017] In the identification device according to the above aspect, the first data type may be weight data, and the second data type may be image data. With this configuration, when identifying features using weight data and image data, it is possible to reduce the cost required to build an identification system.
[0018] In the identification device according to the above aspect, the first data type may be odor data, and the second data type may be image data. With this configuration, when identifying features using odor data and image data, it is possible to reduce the cost required to build an identification system.
[0019] In the identification device according to the above aspect, the first data type may be inertial characteristic data, and the second data type may be image data. With this configuration, when identifying features using the inertial characteristic data and the image data, it is possible to reduce the cost required to build an identification system.
[0020] In the identification device according to the above aspect, the command statement may include a first subsentence instructing to derive a first interim result identifying the feature from the text representation of the first sample, a second subsentence instructing to derive a second interim result identifying the feature from the second sample, and a third subsentence instructing to derive a result identifying the feature based on the first interim result and the second interim result. This configuration makes it possible to prevent the modality of either the first data type or the second data type from being ignored during the identification process. As a result, improved identification accuracy can be expected.
[0021] In the identification device according to the above aspect, the control unit may further be configured to acquire the first sample and determine a category to which the acquired first sample belongs. Determining the category to which the first sample belongs may include converting the first sample into a feature value using a trained encoder generated by machine learning, and comparing the feature value obtained from the first sample with a reference value obtained by converting each training sample used in the machine learning into the feature using the trained encoder to extract one or more categories to which the first sample may belong from among multiple categories assigned to the training samples. The text representation of the first sample may include text representations of the one or more extracted categories. This configuration makes it possible to accurately identify the category to which the first sample belongs. This allows an accurate text representation to be provided to a large-scale language model, thereby ensuring identification accuracy.
[0022] The present invention is not limited to the above-described identification device (information processing device). As another embodiment of the identification device according to each of the above aspects, one aspect of the present invention is an identification device including all or part of each of the above configurations. The information may be an information processing method (identification method) that realizes the above, a program, or a storage medium that stores such a program and is readable by a machine such as a computer. A storage medium that is readable by a machine such as a computer may be a non-transitory medium that stores information such as a program by electrical, magnetic, optical, mechanical, or chemical action. Non-transitory storage media may include storage media (CDs, DVDs, semiconductor memories, etc.), auxiliary storage devices of computers, external storage devices connected to computers, etc.
[0023] For example, an identification program according to an aspect of the present invention may be a program for causing a computer to execute an identification method, which may include the steps of: acquiring a text representation of a first sample of a first data type, the first sample being generated by observing an object with a first sensor, according to a category to which the first sample belongs; providing a prompt to a large-scale language model, the prompt including a command statement for instructing the large-scale language model to identify a feature in the object and the text representation of the first sample, thereby acquiring an answer indicating a result of identifying the feature from the large-scale language model; and outputting information about the acquired result.
[0024] In the identification program according to the above aspect, the prompt may further include a list of labels that are candidates for identifying the feature, and identifying the feature may be performed by selecting a label corresponding to the feature from the list.
[0025] In the identification program according to the above aspect, the identification method may further include acquiring second samples of a second data type different from the first data type, the second samples being generated by observing an object with a second sensor. The large-scale language model may be configured to accept input of the second data type samples. The prompt may further include the second samples.
[0026] In the identification program according to the above aspect, the instruction statement may include a first sub-sentence instructing to derive a first interim result identifying the feature from the text representation of the first sample, a second sub-sentence instructing to derive a second interim result identifying the feature from the second sample, and a third sub-sentence instructing to derive a result identifying the feature based on the first interim result and the second interim result.
[0027] Also, for example, an identification method according to an aspect of the present invention may be executed by a computer, and may include the steps of: acquiring a text representation of a first sample of a first data type, the first sample being generated by observing an object with a first sensor, according to a category to which the first sample belongs; providing a prompt to a large-scale language model, the prompt including a command statement for instructing the large-scale language model to identify a feature in the object and the text representation of the first sample, thereby acquiring an answer indicating a result of identifying the feature from the large-scale language model; and outputting information about the acquired result.
[0028] In the identification method according to the above aspect, the prompt may further include a list of labels that are candidates for identifying the feature, and identifying the feature may be performed by selecting a label that corresponds to the feature from the list.
[0029] The identification method according to the above aspect may further include acquiring second samples of a second data type different from the first data type, the second samples being generated by observing an object with a second sensor. The large-scale language model may be configured to accept input of the samples of the second data type. The prompt may further include the second samples.
[0030] In the identification method according to the above aspect, the command sentence may include a first sub-sentence instructing to derive a first interim result identifying the feature from the text representation of the first sample, a second sub-sentence instructing to derive a second interim result identifying the feature from the second sample, and a third sub-sentence instructing to derive a result identifying the feature based on the first interim result and the second interim result. [Effects of the Invention]
[0031] According to the present invention, it is possible to provide a technique for reducing the cost required to build an identification system in any modality. [Brief explanation of the drawings]
[0032] [Figure 1] FIG. 1 shows a schematic diagram of an example of a situation in which the present invention is applied. [Figure 2] FIG. 2 shows a schematic example of a prompt. [Figure 3] FIG. 3 schematically shows an example of a method for classifying the category of the first sample. [Figure 4] FIG. 4 shows a schematic diagram of an example of a large-scale language model. [Figure 5] FIG. 5 shows a schematic example of a command statement. [Figure 6] FIG. 6 shows a schematic diagram of a first example to which the present invention is applied. [Figure 7] FIG. 7 is a diagram showing an example of a second case to which the present invention is applied. [Figure 8] FIG. 8 shows a schematic diagram of a third example to which the present invention is applied. [Figure 9] FIG. 9 is a schematic diagram showing an example of a fourth case to which the present invention is applied. [Figure 10] FIG. 10 is a schematic diagram showing an example of a fifth case to which the present invention is applied. [Figure 11] FIG. 11 is a diagram showing an example of a sixth case to which the present invention is applied. [Figure 12]FIG. 12 is a diagram showing an example of a seventh case to which the present invention is applied. [Figure 13] FIG. 13 is a diagram illustrating an example of a hardware configuration of the identification device. [Figure 14] FIG. 14 is a diagram illustrating an example of the software configuration of the identification device. [Figure 15] FIG. 15 is a flowchart illustrating an example of a processing procedure of the identification device. [Figure 16] FIG. 16 shows the configuration of an identification system in the embodiment. [Figure 17] FIG. 17 shows the prompts in the example. [Figure 18] FIG. 18 shows the prompt in the reference example. [Figure 19] FIG. 19 shows the items used to collect the training dataset used in the experiments. [Figure 20] FIG. 20 shows the food (real food) and food replicas used to collect the first dataset used in the experiment. [Figure 21] FIG. 21 shows the food (real objects) used to collect the second dataset used in the experiment. [Figure 22] FIG. 22 shows the classification accuracy (results) of the first data set according to the reference example. [Figure 23] FIG. 23 shows the classification accuracy (results) of the first data set according to the embodiment. [Figure 24] FIG. 24 shows the classification accuracy (results) of the second data set according to the reference example. [Figure 25] FIG. 25 shows the classification accuracy (results) of the second data set according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0033] An embodiment according to one aspect of the present invention will be described below with reference to the drawings. However, the embodiment described below is merely an example of the present invention in all respects. Various improvements or modifications may be made without departing from the scope of the present invention. In implementing the present invention, a specific configuration according to the embodiment may be appropriately adopted. Note that, although data appearing in this embodiment is described in natural language, more specifically, it is specified using pseudo-language, commands, parameters, machine language, etc. that can be recognized by a computer.
[0034] §1 Application Examples FIG. 1 is a schematic diagram showing an example of a situation in which the present invention is applied. are one or more computers configured to perform a classification task using a large-scale language model 6. In this embodiment, a first sample 35 of a first data type 30 is generated by observing an object with a first sensor S1. The classification device 1 obtains a text representation 55 of the first sample 35 according to a category to which the first sample 35 belongs. The classification device 1 provides the large-scale language model 6 with a prompt 60 including a command statement 50 for instructing the large-scale language model 6 to identify features in the object and the text representation 55 of the first sample 35, thereby obtaining an answer 65 indicating the result of identifying the features from the large-scale language model 6. The classification device 1 outputs information related to the obtained result.
[0035] The large-scale language model 6 can acquire common sense during the process of machine learning of a huge amount of data. In this embodiment, the first sample 35 is converted into a text representation 55 for input to the large-scale language model 6. The text representation 55 corresponding to the category of the first sample 35 is provided to the large-scale language model 6 as a prompt 60. When generating an answer 65 corresponding to the command sentence 50, the large-scale language model 6 can derive a feature identification result from the category of the first sample 35 using the acquired common sense. Therefore, according to this embodiment, by converting the first sample 35 of the first data type 30 into the text representation 55, zero-shot identification of features in a target can be achieved without fine-tuning the large-scale language model 6. This reduces the cost of building an identification system for any modality.
[0036] [Identification task] The identification task may include any task that classifies the features of an object. The object may be selected as appropriate depending on the embodiment, for example, from among objects, living things (people, etc.), etc. The object may be an actual entity or a virtual entity. The features may include, for example, attributes, states, etc. Attributes are static features such as type. States are dynamically changeable features such as the presence or absence of a malfunction, the state of a chemical reaction, hardness, temperature, and posture. Identifying may include regressing the degree of match (likelihood, etc.) to the features of the object to be identified.
[0037] [prompt] The prompt 60 is an input to the large-scale language model 6. The prompt 60 may be input directly to the large-scale language model 6, or may be pre-processed before being input to the large-scale language model 6. In this embodiment, the prompt 60 is configured to include the command sentence 50 and the text representation 55 of the first sample 35.
[0038] The command statement 50 is composed of a text expression that instructs the large-scale language model 6 to identify features in an object. The content and expression of the command statement 50 are not particularly limited as long as they are configured to indicate an identification command, and may be determined appropriately depending on the embodiment. In one example, the command statement 50 may be composed of text of a command statement instructing the identification of features contained in a sample.
[0039] The text expression 55 is configured to indicate the category to which the first sample 35 belongs. The content of the text expression 55 is not particularly limited as long as it corresponds to the result of categorization of the first sample 35, and may be determined depending on the embodiment. In one example, the text expression 55 may be configured with text indicating that the text expression 55 resembles a type that fits the category. Note that the number of categories extracted as the category to which the first sample 35 belongs may be one or more. In the latter case, in categorization of the first sample 35, multiple categories to which the first sample 35 is likely to belong may be extracted.
[0040] The categories may be defined as appropriate. The categories may be set manually or automatically using a method such as clustering. For example, the categories may be set for the first sample 35 The category may be defined according to features (such as object names) that appear in samples (such as training samples) used to generate a computational model (such as an encoder, described below) used for category classification. The features used to set the category and the features to be identified in the identification task may be unrelated to each other, or may be at least partially related to each other. In one example, the category set for the first sample 35 may be related to answer candidates for the identification task (candidates for feature identification results). For example, if the identification task is to identify the type of item, the category for the first sample 35 may be set according to the type of item. However, each value of the category set for the first sample 35 does not necessarily overlap with the answer candidates for the identification task. For example, if the identification task is to identify the type of food, the category set for the first sample 35 may or may not at least partially include the type of food to be identified. Furthermore, the category set for the first sample 35 may include types of items other than food. It is assumed that the large-scale language model 6 is more likely to apply common sense to commonly used terms than to terms used within a limited range. Therefore, when the text expression 55 is composed of text (or its tokens) including the name of a category, it is desirable that the name (value) of the category be defined in general terms, including academic terms. For example, the general terms may be terms registered in any dictionary. This is expected to improve the accuracy of identification.
[0041] In this embodiment, the configuration of the prompt 60 may be determined appropriately depending on the embodiment, as long as it includes the command statement 50 and the text representation 55 of the first sample 35. The prompt 60 may be composed of only the command statement 50 and the first sample 35, or may further include any information (data) other than the command statement 50 and the text representation 55 of the first sample 35.
[0042] FIG. 2 schematically illustrates an example of a prompt 60 according to this embodiment. As illustrated in FIG. 2, the prompt 60 may further include, in addition to the command statement 50 and the text expression 55, a list 59 of candidate labels for feature identification results. The list 59 is an example of arbitrary information. In this case, identifying a feature may be performed by selecting a label from the list 59 that corresponds to a feature included in the target. The command statement 50 may be configured to instruct the large-scale language model 6 to at least one of select a corresponding label from the list 59 and regress the degree of correspondence to each label.
[0043] According to one example of this embodiment, when an identification range is given, the range of the identification result can be narrowed by providing information indicating the identification range as a list 59 to the large-scale language model 6. This can be expected to improve the accuracy of identification compared to when the identification range is unlimited.
[0044] (Data format) In a typical example, the prompt 60 (command sentence 50, text expression 55, list 59, etc.) may be composed of text data. Text data may be associated with each category that is a candidate for the first sample 35 to belong to, and the text expression 55 may be composed of text data corresponding to the classification result of the category to which the first sample 35 belongs. The text data associated with each category may be text including the name of the category (e.g., text indicating similarity to a type that fits the category). However, as long as the data format of the prompt 60 can be input to the large-scale language model 6, the data format of the prompt 60 is not limited to this example and may be selected appropriately depending on the embodiment. In another example, at least a portion of the prompt 60 may be composed of one or more tokens corresponding to the text data.
[0045] For example, when the text expression 55 is composed of tokens, as in the above example, Each category to which the sample 35 can belong may have associated text data. The text representation 55 may be obtained from the text data according to the classification result of the category, and then converted into tokens by a tokenizer. In a simple example, each category The text data associated with the category may be text including the name of each category, and the text representation 55 may be obtained by converting text including a description of the category to which the first sample 35 belongs into tokens.
[0046] Furthermore, for example, by associating the result of category classification with a value of a token, the token (text representation 55) may be obtained directly according to the result of category classification of the first sample 35 without using text data. A token value may be assigned in advance to each category that is a candidate to which the first sample 35 belongs. The token value of each category may be set appropriately. In a simple example, the token value of each category may be obtained by converting text (such as a name) containing a description of each category into tokens using a tokenizer.
[0047] When the number of categories extracted as the category to which the first sample 35 belongs is one, the tokens obtained corresponding to the extracted category may be input to the large-scale language model 6 as the text representation 55. When the number of categories extracted as the category to which the first sample 35 belongs is multiple, statistics of the tokens obtained corresponding to each of the multiple extracted categories may be calculated, and the calculated statistics may be input to the large-scale language model 6 as the text representation 55. The statistics may be arbitrarily selected from, for example, the mean, the median, etc. In a simple example, the statistics may be formed by a simple average of the values of the tokens in each extracted category. In another example, the values of the tokens in each category may be weighted based on the degree (likelihood, etc.) that the first sample 35 belongs to each category. The statistics may also be formed by a weighted average of the values of the tokens in each extracted category.
[0048] Note that the conversion into tokens is not limited to the method using a tokenizer, and may be performed by other methods, such as using a text-token correspondence table. The token values may be optimized for each category for the identification task using a method such as machine learning. The same applies to the command statement 50 and the list 59, which may be composed of the above tokens.
[0049] [Category Classification] The method for classifying the category of the first sample 35 is not particularly limited and may be selected appropriately depending on the embodiment. The method for classifying the category may employ any method, such as a method using machine learning, an analytical method, or other methods using comparison operations. The machine learning may include, for example, supervised learning, unsupervised learning, etc. The unsupervised learning may include, for example, learning of an autoencoder, clustering, etc. The autoencoder may include, for example, a modified autoencoder such as a VAE (Variational Auto-Encoder) or a VQ-VAE (Vector Quantized-Variational Auto-Encoder). The clustering may be performed, for example, by using a spectral The analytical methods may include, for example, principal component analysis, non-negative matrix factorization (NMF), independent component analysis (ICA), etc. The classification of categories may involve the use of a method based on each method. Any computational model may be used. Depending on the method adopted for categorization, the categories may be referred to by other terms such as classes, clusters, etc.
[0050] In one example of this embodiment, the classification device 1 acquires a first sample 35 and determines a category to which the acquired first sample 35 belongs. Determining the category to which the first sample 35 belongs involves converting the first sample 35 into feature values using a trained encoder generated by machine learning, and comparing the feature values obtained from the first sample 35 with reference values obtained by converting each training sample used in machine learning into a feature using the trained encoder, thereby determining the categories assigned to each training sample. The method may include extracting one or more categories from the list of categories to which the first sample 35 may belong. The text representation 55 of the first sample 35 may include a text representation of the one or more extracted categories. The trained encoder is an example of a computational model used for category classification.
[0051] FIG. 3 schematically illustrates an example of a method for classifying a category to which a first sample 35 belongs using the trained encoder. In the example of FIG. 3, an autoencoder is assumed to be used as a framework for generating a trained encoder (encoder EN). The autoencoder is composed of an encoder EN and a decoder DE. The encoder EN is configured to convert input samples into features. The decoder DE is configured to reconstruct the input samples from the features obtained by the encoder EN (i.e., generate reconstructed samples). Note that the configuration of the autoencoder is not limited to this example and may be changed as appropriate depending on the type of autoencoder used.
[0052] The encoder EN and the decoder DE may each be configured with a machine learning model. The machine learning model has one or more computational parameters that can be adjusted by machine learning. The one or more computational parameters are used for the desired inference computation. The machine learning model may be configured with, for example, a neural network, a support vector machine, a regression model, or other functional formula (computational model). When at least one of the encoder EN and the decoder DE includes a neural network, the structure of the neural network is not particularly limited and may be determined appropriately depending on the embodiment. The structure of the neural network may be specified, for example, by the number of layers from the input layer to the output layer, the type of each layer, the number of nodes (neurons) included in each layer, the connection relationships between the nodes in each layer, etc. The neural network may include any mechanism such as a recurrent structure, a self-attention mechanism, or an autoregressive model. The neural network may include any layer such as a fully connected layer, a convolutional layer, a pooling layer, a deconvolutional layer, an unpooling layer, a normalization layer, a dropout layer, or a long short-term memory (LSTM). The neural network may include any type of model, such as a diffusion model, a transformer model, or a generative model. The weights of the connections between the nodes included in the neural network and the thresholds of the nodes are examples of computational parameters. The machine learning method may be appropriately selected depending on the embodiment of the machine learning model to be adopted (e.g., backpropagation).
[0053] In the example of FIG. 3, machine learning (unsupervised learning) of the autoencoder is performed in the training (learning) stage. Machine learning involves adjusting (optimizing) the values of calculation parameters using training samples. Machine learning of the autoencoder may be performed as appropriate depending on the embodiment. As an example of machine learning processing, when the encoder EN and the decoder DE are configured as neural networks, training samples TR may be provided to the encoder EN, and forward calculation processing of the encoder EN may be performed. The training samples TR, like the first samples 35, are samples of the first data type 30, and are generated by observing an object using the first sensor S1. The number of training samples TR used for training may be determined as appropriate depending on the embodiment. Each training sample TR may be collected as appropriate. As a result of the calculation processing of the encoder EN, a feature F0 may be obtained from the encoder EN (i.e., the training sample TR is converted into a feature F0). The obtained feature F0 may be provided to the decoder DE, and forward calculation processing of the decoder DE may be performed. As a result of the calculation processing of the decoder DE, a reconstructed sample RC may be obtained from the decoder DE. The loss may be calculated by calculating the reconstruction error (difference) between the reconstructed sample RC and the corresponding training sample TR. The reconstruction error is one example of the loss. The loss in the autoencoder may further include errors other than the reconstruction error, such as a regularization term, a vector quantization error, and a commitment error. Then, the values of the operation parameters of the encoder EN and the decoder DE may be adjusted (optimized) so that the calculated loss is small. In the example of FIG. 3, the gradient of the reconstruction error is calculated. Then, the calculated gradients may be backpropagated to calculate errors in the values of the calculation parameters of the decoder DE and the encoder EN. The values of the calculation parameters of the decoder DE and the encoder EN may be updated based on the calculated errors. The degree to which the values of the calculation parameters are updated may be adjusted by a learning rate. A series of processes from providing a training sample TR to adjusting the values of the calculation parameters (calculation parameter adjustment process) may be repeatedly executed until a predetermined condition is met, such as the loss becoming less than a threshold or a predetermined number of repetitions. As a result of repeatedly executing this adjustment process, a trained encoder EN and a trained decoder DE can be generated.
[0054] The training samples TR may be prepared for each category. When preparing the training samples TR, multiple categories may be defined as appropriate. In one example, each category may be defined according to a feature (such as a type of object) commonly appearing in the training samples TR to which it belongs. Each category may be defined manually or automatically using a method such as clustering. Each training sample TR is assigned to at least one of the multiple categories.
[0055] After the above machine learning is completed, each training sample TR may be provided to a trained encoder EN, and the trained encoder EN may execute a calculation process to convert each training sample TR used in the machine learning into a feature F0. This allows a reference value R0 that can be used for category classification in the inference stage to be obtained. That is, the value of the feature F0 calculated from each training sample TR using the trained encoder EN may be collected as a reference value R0 to be used for category classification. Multiple reference values R0 may be collected for each category, and each category (reference value R0) may be associated with a text expression such as the name of the category. This allows a database that can be used for conversion between categories (reference values R0) and text to be constructed.
[0056] Note that at least a part of the pre-processing, including the machine learning process for generating the trained encoder EN and the process for collecting the reference value R0 for each category, may be executed on the classification device 1 or may be executed on an external computer other than the classification device 1. When the machine learning process is executed on an external computer, in one example, the classification device 1 may acquire the trained encoder EN directly or indirectly from the external computer at any timing and by any method. Acquiring it indirectly means acquiring it via a storage medium, another computer, or the like. In another example, the trained encoder EN may be pre-installed in the classification device 1. Similarly, when the reference value R0 for each category is generated by an external computer, in one example, the classification device 1 may acquire the reference value R0 for each category directly or indirectly from the external computer. In another example, the reference value R0 for each category may be pre-installed in the classification device 1.
[0057] On the other hand, in the example of FIG. 3 , in the inference stage, the classification device 1 may acquire a first sample 35. The method for acquiring the first sample 35 may be selected appropriately depending on the embodiment. In one example, the classification device 1 may acquire the first sample 35 directly or indirectly from the first sensor S1. The first sample 35 may be acquired in response to an operator's operation or automatically. The classification device 1 may convert the acquired first sample 35 into a feature (value F1) using a trained encoder EN. The classification device 1 may compare the feature value F1 obtained from the first sample 35 with a reference value R0 obtained from each training sample TR. Depending on the result of this comparison, the classification device 1 may extract one or more categories to which the first sample 35 may belong from among multiple categories assigned to each training sample TR. In one example, the classification device 1 may search for a neighborhood of the value F1 in a database (latent space) of reference values R0. In this neighborhood search, the classification device 1 may extract k categories (k is 1 or more) by picking up categories to which the reference values R0 belong, starting from the reference value R0 closest to the value F1. The classification device 1 may acquire the extracted k categories as candidates for the category to which the first sample 35 belongs. In this way, the classification device 1 may extract multiple From these categories, one or more categories can be extracted as candidates to which the first sample 35 belongs. That is, the identification device 1 can obtain a determination (classification) result of the category to which the acquired first sample 35 belongs. The identification device 1 may construct a text representation 55 of the first sample 35 using text representations of the one or more extracted categories. According to this example of the present embodiment, it is possible to properly identify the category to which the first sample 35 belongs. This makes it possible to provide the large-scale language model 6 with a proper text representation 55, and as a result, it is expected that the identification accuracy can be ensured.
[0058] Note that the text expression may be a fixed value or a variable value among reference values R0 belonging to the same category. In the latter case, the text expression may be associated with each reference value R0. Associating the text expression with a category may include associating the text expression with the reference value R0.
[0059] The trained encoder EN may also be generated by methods other than the machine learning of the autoencoder. In another example, the trained encoder EN may be generated by supervised learning. In yet another example, the trained encoder EN may be generated by an analytical method. For example, the trained encoder EN may be composed of principal component vectors obtained by principal component analysis of multiple training samples TR.
[0060] Furthermore, the classification device 1 may classify the category to which the first sample 35 belongs by directly comparing the first sample 35 and the training sample TR, rather than comparing features in a latent space. That is, the classification device 1 may compare each sample (35, TR) without compressing them into features. The classification device 1 can extract one or more categories to which the first sample 35 belongs using a method similar to the feature-based method described above, except that the objects to be compared are replaced by each sample (35, TR) instead of each value (F1, R0). Note that when this method is adopted, each training sample TR may be read as a reference sample, etc.
[0061] Furthermore, the method for classifying the category of the first sample 35 is not limited to the method based on the feature or sample comparison. In another example, a trained classifier configured to derive a category classification result from a sample may be generated. The trained classifier may be generated by machine learning, such as unsupervised learning (e.g., clustering) or supervised learning. The trained classifier is an example of a computational model used for category classification. The identification device 1 may provide the acquired first sample 35 to the trained classifier and execute computational processing of the trained classifier to obtain a category classification result for the first sample 35 from the trained classifier. The identification device 1 may generate a text representation 55 according to the acquired classification result.
[0062] Regardless of which of the above methods is adopted as a method for classifying the category of the first sample 35, at least a part of the pre-processing such as machine learning may be executed on the identification device 1 or may be executed on an external computer other than the identification device 1. When data used for classification (reference samples, trained classifiers, etc.) is generated by an external computer, the identification device 1 may obtain the data directly or indirectly from the external computer.
[0063] In one example of the above embodiment, the identification device 1 executes a series of processes from acquiring the first sample 35 to generating the text representation 55 in accordance with the classification result of the first sample 35. However, the entity that executes at least part of this series of processes does not have to be the identification device 1, and may be an external computer other than the identification device 1. In another example, the external computer may execute the series of processes from acquiring the first sample 35 to generating the text representation 55. In this case, the identification device 1 may directly or indirectly acquire the text representation 55 corresponding to the first sample 35 from the external computer.
[0064] [Large-scale language model] As long as an answer can be generated from text, the configuration of the large-scale language model 6 is not particularly limited and may be determined appropriately depending on the embodiment. A known model proposed in Non-Patent Document 3 or the like may be adopted as the large-scale language model 6. Furthermore, the large-scale language model 6 may be configured to further accept input of data other than text in addition to text, such as a large-scale visual language model (Non-Patent Document 4, etc.) or an Audio Question Answering Model (Non-Patent Document 5, etc.).
[0065] FIG. 4 schematically illustrates an example of a large-scale language model 6 according to this embodiment. As illustrated in the example of FIG. 4, the large-scale language model 6 may be configured to accept input of a sample of a second data type 40 different from the first data type 30. In response to this, the identification device 1 may further acquire a second sample 45 of the second data type 40. The second sample 45 may be generated by observing an object using a second sensor S2. A method for acquiring the second sample 45 may be selected appropriately depending on the embodiment. In one example, the identification device 1 may acquire the second sample 45 directly or indirectly from the second sensor S2. The prompt 60 may further include the second sample 45 in addition to the command sentence 50 and the text expression 55. The second sample 45 may be included in the prompt 60 as is, or may be included in the prompt 60 after applying any preprocessing. According to this example of the present embodiment, when a large-scale language model capable of accepting data input of a modality other than text (the second data type 40) is used as the large-scale language model 6, the cost of building an identification system can be reduced. Furthermore, by using such a large-scale language model, it is possible to expect improvement in recognition accuracy. Note that, in one example of this embodiment, the form shown in FIG. 2 may also be adopted. That is, the prompt 60 may be configured to include a command sentence 50, a text representation 55 of the first sample 35, a second sample 45, and a list 59.
[0066] (Data type / sensor) The first data type 30 and the second data type 40 may be appropriately selected from data types other than text. When adopting the form of FIG. 4 in which the large-scale language model 6 is configured to be able to accept samples of the second data type 40, it is desirable that the first data type 30 be selected from a minor data type for which large-scale language models that can accept samples (first samples 35) are not generally available. In particular, it is desirable that the first data type 30 be selected from a data type for which large-scale language models that can accept samples are not provided as commercial services. A minor data type may be a data type for which fine-tuning a large-scale language model is costly due to factors such as the lack of a large amount of data or the lack of standardized sensor standards.
[0067] As an example, the first data type 30 may be sensing data such as tactile data, temperature data, point cloud data, weight data, smell data, inertial property data, electromagnetic data, etc. The first sensor S1 may be configured with a tactile sensor, a temperature sensor (such as a thermometer), a point cloud sensor, a load sensor (such as a weigh scale or load meter), a smell sensor, an inertial property measuring instrument, an electromagnetic sensor (such as an ammeter, voltmeter, or magnetometer), etc. The point cloud sensor may include, for example, a LiDAR (light detection and ranging), an MMS (Mobile Mapping System), a depth sensor, an ultrasonic sensor, an infrared sensor, a radar, etc.
[0068] On the other hand, the second data type 40 is preferably selected from major data types for which large-scale language models expanded to be able to accept samples (second samples 45) are widespread due to factors such as unified sensor standards and the existence of a large amount of data. In particular, the second data type 40 is preferably selected from data types for which large-scale language models expanded to be able to accept samples are provided as commercial services. For example, the second data type 40 may be at least one of image data D1 and sound data D2. S2 may be configured with at least one of an image sensor (such as an RGB camera) and a microphone.
[0069] In one example, by adopting a large-scale visual language model as the large-scale language model 6, the second data type 40 may be image data D1. The image data D1 may be composed of still images or moving images. According to one example of the present embodiment, it is possible to reduce the cost of building a classification system when using a large-scale visual language model. Note that the large-scale visual language model may include a visual question answering model, an open vocabulary object detection model, an open vocabulary object segmentation model, etc.
[0070] In another example, the second data type 40 may be audio data D2 by employing an Audio Question Answering Model as the large-scale language model 6. According to this example of the present embodiment, it is possible to reduce the cost required to build a classification system when using an Audio Question Answering Model.
[0071] However, the relationship between the first data type 30 and the second data type 40 need not be limited to this example. Not only in cases where the configuration of FIG. 4 is not adopted, but also in cases where the configuration of FIG. 4 is adopted, the first data type 30 may be selected from major data types for which large-scale language models capable of accepting samples are widely available. For example, the first data type 30 may be at least one of image data and sound data. Accordingly, the first sensor S1 may be composed of at least one of an image sensor and a microphone.
[0072] In one example, in response to the selection of image data as the second data type 40, a data type other than image data, such as audio data, whether minor or major, may be selected as the first data type 30. This allows the large-scale language model 6, which can accept text and samples of the second data type 40 (second samples 45), to reflect samples of the first data type 30 (first samples 35), which are not originally acceptable, in the classification task. As a result, improvement in classification accuracy can be expected.
[0073] Note that any data type other than those described above may be adopted for each data type (30, 40). Each sensor (S1, S2) is any machine configured to observe an object and generate data indicating the observation results. Each sensor (S1, S2) may be configured by a computer. In one example, a sample generated by any calculation process of the computer may be treated as at least one of the first sample 35 and the second sample 45. Furthermore, each sensor (S1, S2) may be at least one of a real sensor and a virtual sensor. Observing may include simulating.
[0074] Furthermore, the number of first data types 30 does not have to be limited to one and may be two or more. The number of first samples 35 for each data type does not have to be one and may be two or more. When multiple first data types 30 are set, the number of first samples 35 applied to a single classification task may be the same among the first data types 30, or may at least partially differ among the first data types 30. A text representation 55 may be obtained for each first sample 35. The text representations 55 of each first sample 35 may be integrated within a prompt 60. As long as the large-scale language model 6 can accept it, the number of second data types 40 does not have to be limited to one and may be two or more. The number of second samples 45 for each data type does not have to be one and may be two or more.
[0075] Also, the first data type 30 may be configured by a combination of a plurality of data types (modalities). The large-scale language model 6 may be configured to accept input of each sample and calculate features from each input sample. In this way, multiple data types may be regarded as one first data type. If the large-scale language model 6 is capable of accepting input of a video including audio, the second data type 40 may also be configured as a combination of multiple data types. For example, if the large-scale language model 6 is configured to be capable of accepting input of a video including audio, the second data type 40 may be configured as a combination of audio data and image data.
[0076] Furthermore, in a typical example, data obtained from the same type of sensor may be considered to be the same data type. However, the definition of the data type is not limited to this example. In another example, even if the sensor is the same type, data acquired under different conditions, such as different sensing methods, may be considered to be different data types. For example, when tactile data is measured using a tactile sensor provided in a gripper, a data type may be defined for each gripping method of the gripper (i.e., tactile data acquired using different gripping methods may be considered to be different data types). Accordingly, the first samples 35 of the multiple first data types 30 that are different from each other may be acquired from the same sensor. Furthermore, in a typical example, the first sensor S1 and the second sensor S2 may be different from each other. In another example, depending on the definition of the first data type 30 and the second data type 40, the first sensor S1 and the second sensor S2 may at least partially overlap. For example, when operating to acquire data under different conditions as different data types, the first sensor S1 and the second sensor S2 may be the same sensor.
[0077] The category of the first sample 35 may be defined to have a correlation with the second data type 40, or may be defined without a correlation with the second data type 40. In one example, the category of the first sample 35 is preferably defined so as to maximize the reduction in uncertainty of the correct answer (candidate answer) of the classification task for the sample of the second data type 40 when the category of the first sample 35 is given. For example, assuming that the category of the first sample 35 is "X1," the sample of the second data type 40 is "X2," and the correct answer of the classification task is "Y," the difference "H(Y;X2)-H(Y;X1,X2)" between the relative entropy of the correct answer of the classification task for the sample of the second data type 40 and the relative entropy of the correct answer of the classification task for the category of the first sample 35 and the sample of the second data type 40 can be used as an index of the reduction in uncertainty. The relative entropy may be expressed in other ways, such as information gain or Kullback-Leibler distance. Samples of the first data type 30 and the second data type 40 may be collected, and the category of the first sample 35 may be defined such that "H(Y;X2)-H(Y;X1,X2)" exceeds a threshold value for the collected samples. The threshold value may be set as appropriate.
[0078] (command text) If the large-scale language model 6 is configured to further receive input of samples of the second data type 40, the command statement 50 may be configured to instruct identification of features in the object from the text representation 55 of the first sample 35 and the second sample 45. If configured in this way, the content of the command statement 50 may be determined appropriately depending on the embodiment.
[0079] Fig. 5 schematically shows an example of a command statement 50 according to this embodiment. In the example of Fig. 5, a list of object types is provided as a list 59, and the command statement 50 and the list 59 are assumed to be composed of text data. As shown in the example of Fig. 5, the command statement 50 may include a presupposition subsentence 500, a first subsentence 501, a second subsentence 502, and a third subsentence 503.
[0080] The preamble 500 is configured to indicate the content of the identification task to be performed. The preamble 500 may be omitted. The first subsentence 501 may be configured to instruct deriving a first interim result that identifies features from the text representation 55 of the first sample 35. In one example, the first subsentence 501 indicates that the first interim result is to be derived from the text representation 55 of the first sample 35. The second sub-sentence 502 may further include an instruction to ignore the second sample 45 in the process of deriving the second interim result ("Ignore the 2nd sample at this stage." in FIG. 5 ). The second sub-sentence 502 may be configured to instruct deriving a second interim result that identifies features from the second sample 45. In one example, the second sub-sentence 502 may further include an instruction to ignore the text representation 55 of the first sample 35 in the process of deriving the second interim result from the second sample 45 ("Ignore the 1st sample at this stage." in FIG. 5 ). The third sub-sentence 503 may be configured to instruct deriving a result (final result) that identifies features based on the first interim result and the second interim result.
[0081] According to one example of the present embodiment, the command statement 50 includes a first partial sentence 501, a second partial sentence 502, and a third partial sentence 503, thereby making it possible to prevent the modality of either the first data type 30 or the second data type 40 from being ignored during the classification process. Furthermore, the first partial sentence 501 further includes an instruction to ignore the second sample 45, and the second partial sentence 502 further includes an instruction to ignore the text representation 55 of the first sample 35, thereby making it possible to reliably reflect the modality of each of the first data type 30 and the second data type 40 in the classification process. As a result, an improvement in classification accuracy can be expected.
[0082] The content of the command statement 50 is not limited to the example in FIG. 5 and may be modified as appropriate depending on the embodiment. For example, the command statement 50 shown in FIG. 5 does not include any sub-statements other than the sub-statements 500-503. However, the configuration of the command statement 50 is not limited to this example and may include further sub-statements other than the sub-statements 500-503. In another example, the command statement 50 may further include a sub-statement that specifies the output format of the answer 65.
[0083] 5, the command statement 50 includes, from top to bottom, a presupposed partial statement 500, a first partial statement 501, a second partial statement 502, and a third partial statement 503. However, the order in which the partial statements 500 to 503 are written is not limited to this example and may be changed as appropriate depending on the embodiment. In another example, the second partial statement 502 may be placed before the first partial statement 501.
[0084] Also, in the first sub-sentence 501, the instruction to ignore the second sample 45 may be omitted. In the second sub-sentence 502, the instruction to ignore the text representation 55 of the first sample 35 may be omitted.
[0085] Furthermore, the content of the descriptions in each of the partial sentences 500-503 need not be limited to the example in FIG. 5, and may be modified as appropriate depending on the embodiment. For example, in FIG. 5, the first partial sentence 501 and the second partial sentence 502 instruct to evaluate the degree of belonging to each category using a score of 0-10. The score assignment result is the provisional result of the classification. However, the score range and the format of the provisional result need not be limited to this example. The score range may be set arbitrarily. The first partial sentence 501 and the second partial sentence 502 may be configured to instruct the derivation of the provisional result in a format other than a score. The list 59 may be omitted, and the description of the classification task in the command statement 50 may be modified accordingly.
[0086] (Deployment location) The large-scale language model 6 may be deployed in any location. In one example, the large-scale language model 6 may be located in the identification device 1. In this case, the identification device 1 provides a prompt 60 to the large-scale language model 6 and executes calculation processing of the large-scale language model 6, thereby obtaining an answer 65 indicating the result of identifying the feature from the large-scale language model 6. In another example, the large-scale language model 6 may be deployed in an external computer other than the identification device 1. In this case, providing the prompt 60 to the large-scale language model 6 is configured by providing the external computer with a request to execute calculation processing of the large-scale language model 6 together with the prompt 60. In response to the request received from the identification device 1, the external computer may provide a prompt 60 to the large-scale language model 6 and execute calculation processing of the large-scale language model 6. As a result, the external computer may generate an answer 65 indicating the result of identifying the feature. The identification device 1 may obtain the generated answer 65 directly or indirectly from the external computer.
[0087] [Specific example] This embodiment is applicable to various situations in which a classification task is performed. Furthermore, the example of this embodiment shown in FIG. 4 is applicable to various situations in which a classification task is solved using two or more types of data (first data type 30, second data type 40). Each data type (30, 40) may be selected appropriately depending on the application situation, classification task, etc. The example of this embodiment shown in FIG. 4 may be applied to at least one of the following first, second, third, fourth, fifth, sixth, and seventh cases. Specific application situations of the embodiment shown in FIG. 4 will be exemplified below for each case.
[0088] (1) First Case 6 is a diagram illustrating an example of a first case to which this embodiment is applied. The first case is an example of a situation in which this embodiment is applied to identify features of an object T1 from tactile data C1 and image data D1.
[0089] As shown in FIG. 6, the first data type 30 may be tactile data C1. A sample (first sample 35) of the tactile data C1 may be generated by a tactile sensor S11. The tactile sensor S11 is an example of the first sensor S1. The second data type 40 may be image data D1. A sample (second sample 45) of the image data D1 may be generated by an image sensor S21. The image sensor S21 is an example of the second sensor S2. The tactile sensor S11 and the image sensor S21 may be appropriately disposed in a location where the target T1 can be observed.
[0090] The identification device 1 may construct a prompt 60 with the command sentence 50, a text representation 55 of the first sample 35 of the tactile data C1, and the second sample 45 of the image data D1. In the first case, the identification device 1 may construct the prompt 60 to further include a list 59. The identification device 1 may provide the obtained prompt 60 to the large-scale language model 6 to obtain an answer 65 from the large-scale language model 6 indicating the results of identifying features in the target T1.
[0091] The tactile and image-based identification task may be performed for any purpose. The target T1 may be determined appropriately depending on the embodiment. In one example, the target T1 may be a workpiece of the robot device R1, and identifying the characteristics may include identifying the type of the target T1. The tactile sensor S11 and the image sensor S21 may be appropriately positioned so as to be able to observe the workpiece (target T1) of the robot device R1. For example, the tactile sensor S11 may be attached to a robot hand such as a gripper, and the image sensor S21 may be positioned either inside or outside the robot device R1 so that the working range of the robot device R1 is within the imaging range.
[0092] The result of identifying the type of the object T1 may be used to control the robot device R1. For example, consider a scenario in which the robot device R1 is caused to perform a task of gripping the object T1 with a gripper and transporting the object T1 to a destination. In this scenario, at the initial stage of gripping the object T1 with the gripper, a first sample 35 of tactile data C1 may be obtained from the tactile sensor S11 arranged on the gripper. At this initial stage, the gripper may be controlled to grip the object T1 with a force that is too weak to lift the object T1 for transporting it, but is light enough to obtain tactile data C1 that can be used for the identification task. A second sample 45 of the image data D1 may be acquired at any timing. The identification device 1 may generate a prompt 60 from the obtained first sample 35 and second sample 45, and provide the generated prompt 60 to the large-scale language model 6 to obtain a result of identifying the type of the object T1. The identification device 1 identifies the type of the object T1. Depending on the result, the identification device 1 may determine the gripping force to be used by the gripper when lifting the object T1, and may issue a command to the robot device R1 to grip the object T1 with the determined force and transport the object T1. Alternatively, the identification device 1 may cause the robot device R1 to determine the force to be used when lifting the object T1 by providing the robot device R1 with the identification result of the type of the object T1.
[0093] The type of the robot device R1 is not particularly limited and may be appropriately selected depending on the embodiment. In one example, the robot device R1 may be, for example, an industrial robot used in a production line, an autonomous robot configured to operate autonomously, or a mobile robot configured to move. The industrial robot may be, for example, a vertical articulated robot, a horizontal articulated robot (SCARA robot), a parallel link robot, or an orthogonal robot. The autonomous robot may be, for example, a humanoid robot, a guide robot, an agricultural robot, a nursing robot, a security robot, a transport robot (including a food distribution robot), or the like. The mobile robot may be, for example, a cleaning robot, the above-mentioned autonomous robot (including a mobile robot) configured to move, a vehicle configured to be autonomously driven, an air vehicle capable of autonomous flight (such as a drone), or a ship capable of autonomous navigation (such as a ship or submarine). The robot device R1 may be operated manually.
[0094] According to the first example, it is possible to reduce the cost of building a classification system when identifying features using tactile data C1 and image data D1. Furthermore, when using a large-scale visual language model as the large-scale language model 6, it is possible to expect an improvement in classification accuracy by incorporating tactile data C1 (text representation 55) in the classification task in addition to image data D1.
[0095] (2) Second Case 7 is a schematic diagram illustrating a second example of the application of this embodiment. The second example is an example of the application of this embodiment to identify the characteristics of an object T2 from temperature data C2 and image data D1.
[0096] As shown in FIG. 7, the first data type 30 may be temperature data C2. A sample (first sample 35) of the temperature data C2 may be generated by a temperature sensor S12. The temperature sensor S12 is an example of the first sensor S1. The temperature sensor S12 may be either a contact type or a non-contact type. As in the first example, the second data type 40 may be image data D1, and a sample (second sample 45) of the image data D1 may be generated by an image sensor S21. The temperature sensor S12 and the image sensor S21 may be appropriately positioned in a location where the object T2 can be observed.
[0097] The identification device 1 may construct a prompt 60 including a command sentence 50, a text representation 55 of the first sample 35 of the temperature data C2, and a second sample 45 of the image data D1. In the second case, the identification device 1 may also construct the prompt 60 to further include a list 59. The identification device 1 may provide the obtained prompt 60 to the large-scale language model 6 to obtain an answer 65 from the large-scale language model 6 indicating the results of identifying features in the target T2.
[0098] The temperature and image identification task may be performed for any purpose. The target T2 may be determined appropriately depending on the embodiment. In one example, the target T2 may be a target of a chemical experiment (e.g., a cell culture solution, etc.). If the target T2 is a cell culture solution, the identification task may include identifying a range of solutions suitable for cell culture from the temperature and image. The identification result may be fed back to the cell culture.
[0099] In another example, the target T2 may be a cooking target. The identification task may include identifying the state of the cooking process, such as whether the target T2 is in a suitable state for cooking, from the temperature and the image. When cooking is performed by a robotic device, the identification result may be fed back to the cooking operation (for example, whether to proceed to the next step may be determined depending on the identification result).
[0100] In another example, the target T2 may be a metal (such as an alloy) to be processed. The identification task may include identifying the state of the target T2 during processing from the temperature and the image. If the metal processing is performed by a robotic device, the processing operation of the robotic device may be determined according to the identification result.
[0101] In another example, the target T2 may be a robotic device. The identification task may include identifying the state of the robotic device from temperature and images. The scope of the robotic device may be defined in the same manner as for the robotic device R1 described above. The temperature sensor S12 may be located either inside or outside the robotic device. The identification result may be fed back to the operation of the robotic device. For example, if the robotic device is identified as being in a state where cooling down is recommended, the identification device 1 may issue a command to the robotic device to perform cooling down. Alternatively, the identification device 1 may provide the identification result to the robotic device, causing the identification device 1 to determine whether or not to perform cooling down.
[0102] In another example, the object T2 may be food or drink to be delivered by a delivery robot. The identification task may include identifying the state of the food or drink from temperature and an image. The result of identifying the state of the food or drink may be fed back to the operation of the delivery robot. For example, the operation of the delivery robot may be determined according to the identification result, such as delivering hot food or drink more slowly (moving at a slower speed) compared to cold food or drink. The identification device 1 may determine an operation according to the identification result and give the delivery robot a command to execute the determined operation. Alternatively, the identification device 1 may provide the identification result to the delivery robot, thereby allowing the delivery robot (controller) to determine the operation.
[0103] According to the second example, it is possible to reduce the cost of building a classification system when identifying features using temperature data C2 and image data D1. Furthermore, when using a large-scale visual language model as the large-scale language model 6, it is possible to expect an improvement in classification accuracy by incorporating temperature data C2 (text representation 55) in the classification task in addition to image data D1.
[0104] (3) Third Case 8 is a diagram illustrating an example of a third case to which this embodiment is applied. The third case is an example of a situation in which this embodiment is applied to identify the features of an object T3 from sound data C3 and image data D1.
[0105] 8, the first data type 30 may be sound data C3. Samples of the sound data C3 (first samples 35) may be generated by a microphone S13. The microphone S13 is an example of a first sensor S1. As in the first example, the second data type 40 may be image data D1, and samples of the image data D1 (second samples 45) may be generated by an image sensor S21. The microphone S13 and the image sensor S21 may be appropriately positioned in a location where the target T3 can be observed.
[0106] The identification device 1 may construct a prompt 60 using the command sentence 50, a text representation 55 of the first sample 35 of the sound data C3, and the second sample 45 of the image data D1. In the third case, the identification device 1 may also construct the prompt 60 to further include a list 59. The identification device 1 may provide the obtained prompt 60 to the large-scale language model 6, thereby obtaining an answer 65 from the large-scale language model 6 indicating the results of identifying features in the target T3.
[0107] The identification task using sound and images may be performed for any purpose. The target T3 may be determined as appropriate depending on the embodiment. In one example, the identification task using sound and images may be performed for an inspection such as a non-destructive inspection. The target T3 may be, for example, a facility to be inspected, such as a road, building, or structure containing concrete. The identification task may include identifying the condition of the facility to be inspected (target T3). The identification result may be fed back to a user, such as an administrator.
[0108] In another example, the object T3 may be an object of observation by the robotic device. The identification task may include identifying at least one of the attributes and the state of the object of observation by the robotic device from temperature and images. The scope of the robotic device may be defined similarly to that of the robotic device R1. The identification result may be fed back to the operation of the robotic device. For example, if the robotic device is a guide robot, identifying the state of the object of observation may include identifying whether the subject is requesting guidance. The subject is an example of the object T3. If the subject is identified as being in a state of requesting guidance, the identification device 1 may move near the subject and issue a command to the guide robot to provide guidance to the subject. Alternatively, the identification device 1 may provide the identification result to the guide robot to cause the guide robot to perform an action related to guidance.
[0109] According to the third example, it is possible to reduce the cost of building a classification system when identifying features using sound data C3 and image data D1. Furthermore, when using a large-scale visual language model as the large-scale language model 6, it is possible to expect an improvement in classification accuracy by incorporating sound data C3 (text representation 55) in the classification task in addition to image data D1.
[0110] (4) Fourth Case 9 is a schematic diagram illustrating an example of a fourth case to which this embodiment is applied. The fourth case is an example of a situation in which this embodiment is applied to identify features of an object T4 from point cloud data C4 and image data D1.
[0111] 9, the first data type 30 may be point cloud data C4. Samples (first samples 35) of the point cloud data C4 may be generated by a point cloud sensor S14. The point cloud sensor S14 is an example of the first sensor S1. As in the first example, the second data type 40 may be image data D1, and samples (second samples 45) of the image data D1 may be generated by an image sensor S21. The point cloud sensor S14 and the image sensor S21 may be appropriately positioned in locations where the target T4 can be observed.
[0112] The identification device 1 may construct a prompt 60 using the command sentence 50, a text representation 55 of the first sample 35 of the point cloud data C4, and the second sample 45 of the image data D1. In the fourth example, the identification device 1 may also construct the prompt 60 to further include a list 59. The identification device 1 may provide the obtained prompt 60 to the large-scale language model 6, thereby obtaining an answer 65 from the large-scale language model 6 indicating the results of identifying features in the target T4.
[0113] The identification task using point clouds and images may be performed for any purpose. The target T4 may be determined as appropriate depending on the embodiment. In one example, the target T4 may be an observation target of the robotic device. The identification task may include identifying at least one of the attributes and the state of the observation target of the robotic device from the point cloud and images. The scope of the robotic device may be defined in the same way as the robotic device R1 described above. The identification result may be fed back to the operation of the robotic device. For example, if the robotic device is a mobile robot, the identification device 1 may determine the movement path of the mobile robot according to the identification result. For example, if it is identified that the observation target is entering or has the potential to enter the traveling direction of the mobile robot, the movement path may be determined to avoid the observation target. The identification device 1 may then instruct the mobile robot to move along the determined route. Alternatively, the identification device 1 may provide the mobile robot with the identification result, causing the mobile robot to determine a movement route according to the identification result.
[0114] According to the fourth example, it is possible to reduce the cost of building a classification system when identifying features using point cloud data C4 and image data D1. Furthermore, when using a large-scale visual language model as the large-scale language model 6, it is possible to expect improvement in classification accuracy by reflecting point cloud data C4 (text representation 55) in the classification task in addition to image data D1.
[0115] (5) Fifth Case 10 is a diagram illustrating an example of a fifth case to which this embodiment is applied. The fifth case is an example of a situation in which this embodiment is applied to identify the features of an object T5 from weight data C5 and image data D1.
[0116] 10, the first data type 30 may be weight data C5. A sample (first sample 35) of the weight data C5 may be generated by a load sensor S15. The load sensor S15 is an example of the first sensor S1. As in the first example, the second data type 40 may be image data D1, and a sample (second sample 45) of the image data D1 may be generated by an image sensor S21. The load sensor S15 and the image sensor S21 may be appropriately positioned in a location where the target T5 can be observed.
[0117] The identification device 1 may construct a prompt 60 using a command sentence 50, a text representation 55 of the first sample 35 of the weight data C5, and a second sample 45 of the image data D1. In the fifth case, the identification device 1 may also construct the prompt 60 to further include a list 59. The identification device 1 may provide the obtained prompt 60 to the large-scale language model 6 to obtain an answer 65 from the large-scale language model 6 indicating the results of identifying features in the target T5.
[0118] The identification task using weight and image may be performed for any purpose. The object T5 may be determined appropriately depending on the embodiment. In one example, the object T5 may be an item to be inspected. The item to be inspected may be, for example, a product (final product, intermediate product, etc.) on a production line, agricultural produce, etc. The identification task may include identifying the condition of the item to be inspected (object T5) from its weight and image. Identifying the condition of the item to be inspected may include identifying whether the item is defective (whether it is broken, whether it meets standards, etc.). The identification result may be fed back to a user such as a manager.
[0119] According to the fifth example, when identifying features using weight data C5 and image data D1, it is possible to reduce the cost of building a classification system. Furthermore, when using a large-scale visual language model as the large-scale language model 6, it is possible to expect improvement in classification accuracy by reflecting weight data C5 (text representation 55) in addition to image data D1 in the classification task.
[0120] (6) Sixth Case 11 is a schematic diagram illustrating an example of a sixth case to which this embodiment is applied. The sixth case is an example of a situation in which this embodiment is applied to identify the characteristics of a target T6 from odor data C6 and image data D1.
[0121] As shown in FIG. 11, the first data type 30 may be odor data C6. A sample (first sample 35) of the odor data C6 may be generated by an odor sensor S16. The odor sensor S16 is an example of the first sensor S1. As in the first example, the second data type 4 The odor sensor S16 may be image data D1, and a sample (second sample 45) of the image data D1 may be generated by the image sensor S21. The odor sensor S16 and the image sensor S21 may be appropriately disposed in a location where the target T6 can be observed.
[0122] The identification device 1 may construct a prompt 60 using a command sentence 50, a text representation 55 of the first sample 35 of the odor data C6, and a second sample 45 of the image data D1. In the sixth case, the identification device 1 may also construct the prompt 60 to further include a list 59. The identification device 1 may provide the obtained prompt 60 to the large-scale language model 6, thereby obtaining an answer 65 from the large-scale language model 6 indicating the results of identifying features in the target T6.
[0123] The identification task using smell and images may be performed for any purpose. The object T6 may be determined as appropriate depending on the embodiment. In one example, the object T6 may be an object to be inspected. The object to be inspected may include, for example, water in a facility, oil in a fryer, food and drink, etc. The facility may include water treatment equipment such as a water purification plant or a sewage treatment plant. The facility may also include natural or artificial bodies of water such as ponds and lakes. The identification task may include identifying the condition of the object (object T6) from smell and images. Identifying the condition of the object may include identifying the quality of the water in a facility, identifying the quality of oil in a fryer, identifying the quality of food and drink (e.g., whether it is fresh), etc. The identification result may be fed back to a user such as an administrator.
[0124] For example, if the target T6 is water from a sewage treatment plant and the water quality is identified as poor, the identification device 1 may generate an operation plan to increase the operation rate of the sewage treatment plant. If the water quality is identified as meeting the standard, the identification device 1 may generate an operation plan to maintain or decrease the operation rate of the sewage treatment plant. The identification device 1 may output instructions to a user or an equipment controller to operate the sewage treatment plant in accordance with the generated operation plan. Alternatively, the identification device 1 may provide the identification result to an external computer, causing the external computer to generate an operation plan. In response to this, the external computer may output instructions.
[0125] Furthermore, for example, if the target T6 is oil in a fryer, the identification device 1 may provide the results of identifying the quality of the oil to an oil manager. The quality of the oil may be defined appropriately according to any standard, including publicly known standards. If deterioration in the quality of the oil is identified, the identification device 1 may output a notification to the manager's terminal recommending that the oil be replaced.
[0126] According to the sixth example, when identifying features using scent data C6 and image data D1, it is possible to reduce the cost of building a recognition system. Furthermore, when using a large-scale visual language model as the large-scale language model 6, it is possible to expect an improvement in recognition accuracy by incorporating scent data C6 (text representation 55) in the recognition task in addition to image data D1.
[0127] (7) Seventh Case 12 is a diagram illustrating an example of a seventh case to which this embodiment is applied. The seventh case is an example of a situation in which this embodiment is applied to identify the features of an object T7 from inertial characteristic data C7 and image data D1.
[0128] As shown in FIG. 12, the first data type 30 may be inertial characteristic data C7. The inertial characteristics may include, for example, the center of gravity position, mass, moment of inertia, etc. A sample (first sample 35) of the inertial characteristic data C7 may be generated by an inertial characteristic measurement device S17. The inertial characteristic measurement device S17 is an example of a first sensor S1. As in the first example, the second data type 40 may be image data D1, and a sample (second sample 45) of the image data D1 may be generated by an image sensor S21. The inertial characteristic measurement device S17 and the image sensor S2 1 may be appropriately placed in a location where the target T7 can be observed.
[0129] The identification device 1 may construct a prompt 60 using a command sentence 50, a text representation 55 of the first sample 35 of the inertial characteristic data C7, and a second sample 45 of the image data D1. In the seventh case, the identification device 1 may also construct the prompt 60 to further include a list 59. The identification device 1 may provide the obtained prompt 60 to the large-scale language model 6 to obtain an answer 65 from the large-scale language model 6 indicating the results of identifying features in the target T7.
[0130] The identification task using inertial characteristics and images may be performed for any purpose. The object T7 may be determined as appropriate depending on the embodiment. The identification task may include identifying at least one of the attributes and state of the object T7 from the inertial characteristics and images. For example, if the object T7 is a drive target of a robotic device, identifying at least one of the attributes and state of the object T7 may include identifying an action range suitable for driving the drive target. The identification device 1 may feed back the identification result to the operation of the robotic device so that the drive target is driven within the identified action range. Furthermore, there are combinations of items that are difficult to distinguish based on appearance alone, such as a combination of a raw egg and a boiled egg, but are easily distinguishable when inertial characteristics are taken into consideration. Therefore, identifying at least one of the attributes and state of the object T7 may include identifying the type of object T7 (such as an object). The identification result may be fed back to a user or to the operation of the robotic device.
[0131] According to the seventh example, it is possible to reduce the cost of building a classification system when identifying features using inertial characteristic data C7 and image data D1. Furthermore, when using a large-scale visual language model as the large-scale language model 6, it is possible to expect improvement in classification accuracy by reflecting inertial characteristic data C7 (text representation 55) in addition to image data D1 in the classification task.
[0132] (others) The above examples may be modified as appropriate depending on the embodiment. For example, in the first to seventh examples, the image data D1 (second data type 40) may be omitted. In each example except the third example, sound data D2 may be used as the second data type 40 instead of the image data D1, and a microphone may be used as the second sensor S2. In each of the above examples, at least one of the category classification method of FIG. 3 and one form of the command statement 50 of FIG. 5 may be used.
[0133] Furthermore, the application of this embodiment is not limited to the above examples. In another example, the first data type 30 may be electromagnetic data, and the second data type 40 may be image data. The identification task may include identifying the electrical properties of the object (e.g., whether it is an insulator or a conductor). The identification result of the electrical properties may be fed back to the operation of the robot device, for example.
[0134] In one example of this embodiment, including the above cases, the first sample 35 and the second sample 45 may be measured immediately after intervention on the subject. The intervention may include, for example, physical intervention or electrical intervention. Physical intervention may include, for example, movement (shaking, rotating, etc.), application of sound waves, etc. Electromagnetic intervention may include, for example, bringing a magnet close, passing an electric current, etc. This allows the identification task to identify the effect of the intervention on the subject.
[0135] §2 Configuration example [Hardware configuration] FIG. 13 is a schematic diagram illustrating an example of the hardware configuration of the identification device 1 according to this embodiment. The identification device 1 according to this embodiment is a computer in which a control unit 11, a storage unit 12, an external interface 13, an input device 14, an output device 15, and a drive 16 are electrically connected.
[0136] The control unit 11 includes a CPU (Central Processing Unit) which is a hardware processor, The memory 12 includes RAM (Random Access Memory), ROM (Read Only Memory), etc., and is configured to execute information processing based on programs and various data. The control unit 11 (CPU) is an example of a processor resource. The storage unit 12 may be configured, for example, with a hard disk drive, a solid state drive, etc. The storage unit 12, RAM, and ROM are examples of memory resources. In this embodiment, the storage unit 12 stores various information such as the identification program 81, model data 600, and encoder data EN0.
[0137] The classification program 81 is a program for causing the classification device 1 to execute information processing (see FIG. 15 described below) related to the performance of a classification task. The classification program 81 includes a series of instructions for the information processing. The model data 600 is configured to indicate information about the large-scale language model 6. The encoder data EN0 is configured to indicate information about the trained encoder EN. If the large-scale language model 6 is deployed on an external computer, the model data 600 may be omitted. If category classification is performed on an external computer or if the trained encoder EN is not used for category classification, the encoder data EN0 may be omitted.
[0138] As long as the model data 600 can hold information for executing the calculation process of the large-scale language model 6, the configuration of the model data 600 is not particularly limited and may be determined appropriately depending on the embodiment. For example, the model data 600 may be configured to include information indicating values of calculation parameters of the large-scale language model 6 adjusted by machine learning. The model data 600 may also be configured to further include information indicating the configuration of the large-scale language model 6 (e.g., the structure of a neural network). The same applies to the encoder data EN0. For example, at least one of the model data 600 and the encoder data EN0 may be incorporated into the identification program 81.
[0139] The external interface 13 is configured to connect to an external device via a wired or wireless connection. The external interface 13 may be, for example, a USB (Universal Serial Bus) port, a dedicated port, a communication port, or the like. When the external interface 13 includes a communication port, the communication standard of the communication port may be selected arbitrarily. In this embodiment, the identification device 1 may be connected to an external device (e.g., a first sensor S1, a second sensor S2, an external computer, or the like) via the external interface 13.
[0140] The input device 14 is a device for inputting, for example, a mouse, a keyboard, etc. The output device 15 is a device for outputting, for example, a display, a speaker, etc. An operator can operate the identification device 1 by using the input device 14 and the output device 15. The input device 14 and the output device 15 may be connected via an external interface 13. The input device 14 and the output device 15 may be integrated into one device, for example, a touch panel display, etc.
[0141] The drive 16 is a device for reading various information such as programs stored in the storage medium 91. At least one of the above-mentioned identification program 81, model data 600, and encoder data EN0 may be stored in the storage medium 91 instead of or together with the storage unit 12. The storage medium 91 is configured to store various information (stored programs, etc.) by electrical, magnetic, optical, mechanical, or chemical action so that a machine such as a computer can read the information. The storage unit 12 and the storage medium 91 are non-transitory The storage medium 91 is an example of such a storage medium. The identification device 1 may acquire at least one of the identification program 81, the model data 600, and the encoder data EN0 from the storage medium 91. The storage medium 91 may be a disk-type storage medium such as a CD or a DVD, or may be a non-disk-type storage medium such as a semiconductor memory (for example, a flash memory). The type of the drive 16 may be selected appropriately depending on the type of the storage medium 91. The drive 16 may be connected via an external interface.
[0142] It should be noted that, with regard to the specific hardware configuration of the identification device 1, components can be omitted, replaced, or added as appropriate depending on the embodiment. For example, the control unit 11 may include multiple hardware processors. The hardware processor may be configured with a microprocessor, an FPGA (field-programmable gate array), a DSP (digital signal processor), a GPU (graphics processing unit), an ASIC (application specific integrated circuit), or the like. At least one of the external interface 13, the input device 14, the output device 15, and the drive 16 may be omitted. The identification device 1 may be configured with multiple computers. In this case, the hardware configuration of each computer may or may not be the same. Furthermore, the identification device 1 may be configured with an information processing device designed specifically for the service provided, as well as a general-purpose server device, a general-purpose PC (personal computer), It may be a tablet PC, a terminal device (smartphone, etc.), or the like.
[0143] At least one of the recognition program 81, the model data 600, and the encoder data EN0 may be stored in an external storage device such as a network-attached storage (NAS). A portion of the prompt 60, such as the command statement 50 and at least a portion of the list 59, may be provided by a template. The template portion of the prompt 60 may be stored in at least one of the storage unit 12 and the storage medium 91, stored in an external storage device, provided by an operator, or incorporated into the recognition program 81. The recognition device 1 may generate a portion of the prompt 60 by reading the template data. Furthermore, when using the encoder EN to classify the categories of the first samples 35, the reference value R0 of each training sample TR and the text expression associated with each category may be stored in at least one of the storage unit 12 and the storage medium 91, stored in an external storage device, or incorporated into the recognition program 81.
[0144] [Software configuration] 14 schematically shows an example of the software configuration of the identification device 1 according to this embodiment. The control unit 11 of the identification device 1 executes instructions included in the identification program 81 stored in the storage unit 12 using the CPU. As a result, the identification device 1 operates as a computer including an acquisition unit 111, a recognition unit 112, and an output processing unit 113 as software modules. That is, in this embodiment, each software module of the identification device 1 is realized by the control unit 11 (CPU).
[0145] The acquisition unit 111 is configured to acquire a text representation 55 of the first sample 35 according to a category to which the first sample 35 belongs. In one example, the acquisition unit 111 may further be configured to acquire the first sample 35 and determine a category to which the acquired first sample 35 belongs. Determining a category to which the first sample 35 belongs may include converting the first sample 35 into a feature value F1 using the trained encoder EN, and comparing the feature value F1 obtained from the first sample 35 with a reference value R0 (value of the feature F0) obtained from each training sample TR used in machine learning to extract one or more categories to which the first sample 35 may belong from among a plurality of categories respectively assigned to the training samples TR. The text representation 55 of the first sample 35 may include text representations of the extracted one or more categories. In addition, in one example, In this case, the obtaining unit 111 may be further configured to obtain second samples 45 of the second data type 40.
[0146] The identification unit 112 is configured to provide the large-scale language model 6 with a prompt 60 including a command statement 50 for instructing the large-scale language model 6 to identify features in an object and a text representation 55 of the first sample 35, thereby obtaining an answer 65 indicating a result of the feature identification from the large-scale language model 6. In one example, the prompt 60 may further include a list of labels 59. In one example, the large-scale language model 6 may be configured to accept input of a sample of the second data type 40, and the prompt 60 may further include the second sample 45. In another example, the command statement 50 may include a first sub-sentence 501 instructing the large-scale language model 6 to derive a first interim result of the feature identification from the text representation 55 of the first sample 35, a second sub-sentence 502 instructing the large-scale language model 6 to derive a second interim result of the feature identification from the second sample 45, and a third sub-sentence 503 instructing the large-scale language model 6 to derive a result of the feature identification from the first interim result and the second interim result. The output processing unit 113 is configured to output information regarding the obtained result.
[0147] In this embodiment, an example is described in which each software module of the identification device 1 is implemented by a general-purpose CPU. However, some or all of the above software modules may be implemented by one or more dedicated processors or chipsets. Each of the above modules may also be implemented as a hardware module. With regard to the software configuration of the identification device 1, modules may be omitted, replaced, or added as appropriate depending on the embodiment.
[0148] §3 Example of operation FIG. 15 is a flowchart showing an example of the processing procedure of the identification device 1 according to this embodiment. The example of FIG. 15 assumes a situation in which the embodiment of FIG. 4 is adopted. The following processing procedure is an example of an identification method executed by a computer. However, the following processing procedure is merely an example, and each step may be changed as much as possible. Furthermore, steps in the following processing procedure may be omitted, replaced, or added as appropriate depending on the embodiment.
[0149] (Step S101) In step S101, the control unit 11 operates as the acquisition unit 111 and acquires the first samples 35 of the first data type 30 and the second samples 45 of the second data type 40.
[0150] The method of acquiring each sample (35, 45) is not particularly limited and may be selected appropriately depending on the embodiment. In one example, the control unit 11 may directly or indirectly acquire the first sample 35 from the first sensor S1. The control unit 11 may directly or indirectly acquire the second sample 45 from the second sensor S2. In one example, any of the first to seventh cases above may be adopted for the combination of the first data type 30 and the second data type 40. After acquiring the first sample 35 and the second sample 45, the control unit 11 proceeds to the next step S102.
[0151] (Step S102) In step S102, the control unit 11 operates as the acquisition unit 111 and determines the category to which the first sample 35 belongs.
[0152] As described above, the method for determining the category to which the first sample 35 belongs (classifying the category of the first sample 35) may be appropriately selected depending on the embodiment. In one example, a trained encoder EN may be generated in advance by machine learning, and a reference value R0 may be generated from each training sample TR using the generated trained encoder EN. Furthermore, each category may be associated with a text expression. The control unit 11 determines the category to which the first sample 35 belongs (classifying the category of the first sample 35) by machine learning. The generated trained encoder EN may be used to convert the first sample 35 into a feature (value F1). The control unit 11 may compare the feature value F1 obtained from the first sample 35 with a reference value R0 obtained from each training sample TR. In one example, the control unit 11 may search for a neighborhood of the value F1 in a database of reference values R0. The distance used as a criterion for the neighborhood search may be defined arbitrarily. For example, a known definition such as the L1 norm or the L2 norm may be used for the distance used as a criterion for the neighborhood search. Depending on the result of this comparison, the control unit 11 may extract one or more categories that are candidates to which the first sample 35 belongs from among multiple categories assigned to the training samples TR, respectively. The extracted one or more categories correspond to a category determination result. Upon obtaining the category determination result, the control unit 11 proceeds to the next step S103.
[0153] (Step S103) In step S103, the control unit 11 operates as the acquisition unit 111 and acquires the text expression 55 of the first sample 35 according to the category to which the first sample 35 belongs.
[0154] In one example, the control unit 11 may acquire a text representation 55 of the first sample 35 according to the determination result of the category to which the first sample 35 belongs. Acquiring the text representation 55 of the first sample 35 may be configured by acquiring, as the text representation 55 of the first sample 35, a text representation associated with each of the one or more categories extracted as the determination result. In other words, the text representation 55 of the first sample may include a text representation of the one or more extracted categories. After acquiring the text representation 55, the control unit 11 proceeds to the next step S104.
[0155] (Step S104) In step S104, the control unit 11 generates a prompt 60 using the command statement 50, a text representation 55 of the acquired first sample 35, and the second sample 45. The command statement 50 may be acquired in any manner, such as by being provided by a template.
[0156] In one example, the control unit 11 may generate the prompt 60 to further include a list of labels 59. Also, in one example, the command statement 50 may be configured to include a first sub-sentence 501, a second sub-sentence 502, and a third sub-sentence 503. After generating the prompt 60, the control unit 11 proceeds to the next step S105.
[0157] (Step S105) In step S105, the control unit 11 provides the generated prompt 60 to the large-scale language model 6, thereby obtaining from the large-scale language model 6 an answer 65 indicating the result of identifying the feature.
[0158] In one example, the classification device 1 may store model data 600 and thus be equipped with a large-scale language model 6. The control unit 11 may provide a prompt 60 to the large-scale language model 6 and execute calculation processing on the large-scale language model 6, thereby obtaining an answer 65 from the large-scale language model 6. In another example, the large-scale language model 6 may be deployed on an external computer. In response, the control unit 11 may transmit the prompt 60 and a request to execute calculation to the external computer. The external computer may provide the prompt 60 received from the classification device 1 to the large-scale language model 6 and execute calculation processing on the large-scale language model 6. The external computer may return an answer 65 obtained as a result of this calculation processing to the classification device 1. The control unit 11 may obtain the answer 65 by receiving this reply. In one example, by adopting any of the first to seventh cases as the combination of the first data type 30 and the second data type 40, the control unit 11 may obtain an answer 65 indicating a classification result corresponding to any of the first to seventh cases. When the response 65 is received, the control unit 11 proceeds to the next step S106. Go ahead.
[0159] (Step S106) In step S106, the control unit 11 outputs information about the obtained result (answer 65).
[0160] The output destination and the content of the output information may be selected appropriately depending on the embodiment. In one example, the control unit 11 may directly output the acquired identification result as the response 65, such as in the case of providing feedback to the administrator. In another example, the control unit 11 may perform arbitrary information processing in accordance with the acquired identification result. The control unit 11 may output the result of the information processing as information related to the identification result. The output of the result of the information processing may include, for example, outputting a specific message in accordance with the identification result, or controlling the operation of the controlled device in accordance with the identification result. Controlling the operation of the controlled device may include, for example, feeding back the identification result to the operation of the robot device R1. The output destination may be, for example, RAM, the memory unit 12, the output device 15, another computer, the controlled device, etc.
[0161] When the output of information is completed, the control unit 11 ends the processing procedure of the identification device 1 according to this operation example. The control unit 11 may execute a series of processes from step S101 to step S106 at any timing, such as a user operation or when a condition is satisfied. The control unit 11 may execute the series of processes from step S101 to step S106 in real time, or may execute them as a post-event matching process. The control unit 11 may execute the series of processes from step S101 to step S106 each time at least one of the first sample 35 and the second sample 45 is updated.
[0162] [Features] In this embodiment, in the processes of steps S101 to S103, the first sample 35 is converted into a text representation 55 for input to the large-scale language model 6. In step S104, the text representation 55 corresponding to the category of the first sample 35 is provided to the large-scale language model 6 as a prompt 60. In step S105, when generating an answer 65 corresponding to the command statement 50, the large-scale language model 6 can derive a feature identification result from the category of the first sample 35 using acquired common sense. Therefore, according to this embodiment, by converting the first sample 35 of the first data type 30 into a text representation 55, zero-shot identification of features in a target can be achieved without fine-tuning the large-scale language model 6. This makes it possible to reduce the cost required to build an identification system for any modality.
[0163] §4 Variations Although the embodiments of the present invention have been described in detail above, the above description is merely an example of the present invention in every respect. The processes and means described in this disclosure can be freely combined and implemented as long as no technical contradiction occurs. Furthermore, various improvements or modifications may be made to the above embodiments as appropriate.
[0164] <4.1> In the processing procedure of the identification device 1 according to the above embodiment, steps may be omitted, replaced, or added. For example, the timing of acquiring the first sample 35 and the second sample 45 is not limited to the example of FIG. 15 and may be changed as desired. The second sample 45 may be acquired at any timing before step S104. The process of acquiring the second sample 45 and the process of acquiring the text representation 55 of the first sample 35 may or may not overlap at least partially. The process of obtaining 5 may be performed before or after the process of obtaining the textual representation 55 of the first sample 35.
[0165] Furthermore, for example, if the large-scale language model 6 is not configured to be able to accept input of samples of the second data type 40, the processing related to the second sample 45 may be omitted. Furthermore, for example, the processing of determining the category to which the first sample 35 belongs and deciding on the text expression 55 may be executed by an external computer other than the classification device 1. In this case, the processing related to the first sample 35 in steps S101 to S103 may be replaced by directly or indirectly acquiring, from the external computer, the text expression 55 corresponding to the category to which the first sample 35 belongs. Alternatively, the control unit 11 may acquire, from the external computer, a determination result of the category to which the first sample 35 belongs, and acquire the text expression 55 according to the acquired determination result.
[0166] §5 Working Examples The following experiments were carried out to verify the effectiveness of the above-described embodiment, but the present invention is not limited to the following examples.
[0167] (Example) Fig. 16 shows the configuration of a classification system in an embodiment. As shown in Fig. 16, tactile data (tactile sequence) was used as the first data type, and image data was used as the second data type. The classification task was to identify the type of object appearing in each sample of tactile data and image data. The method of classifying the category of the first sample was to use the autoencoder shown in Fig. 3 above. The autoencoder was the VQ-VAE proposed in Non-Patent Document 6. A two-layer bidirectional GRU (Gated Recurrent Unit) with a hidden size of 32 was used for the encoder and decoder. The number of dimensions of the embedded representation in vector quantization (VQ) was set to 4, and the number of vectors in the codebook was set to 32. Z in Fig. 16 indicates a feature, and Z q indicates the feature discretized by vector quantization. The output of VQ (Z q ) was input to a linear layer with an output size of 32 and set as the initial hidden state of the decoder. The loss function is the reconstruction error (L recon ), vector quantization error (L vq), and the commitment error (L commit ) was composed.
[0168]
number
[0169] The inference stage was configured to use a trained encoder to calculate the feature values of the tactile data and then search for neighborhoods of the feature values calculated using the L2 norm within a database of reference values. The neighborhood search was configured to extract k categories by picking out the reference values that belong to the reference values, starting with the reference value closest to the feature value of the tactile data. k was set to 3. The names of each extracted category were used as text representations of the tactile data, and a prompt was constructed from a list of commands, text representations of the tactile data, image data, and labels.
[0170] Figure 17 shows the prompt (excluding image data) in the example. The command structure is the same as that shown in Figure 5. The text representation of the tactile data (tactile information) is The {topk_refs} consists of text indicating similarities to species that fit the criteria. The names of the three categories extracted by the neighborhood search were entered. The label list {test_time_classes} is based on the labels (category names) of the dataset used for testing (described later). By adding the description of "Output Format" to the prompt, it is possible to The output format of the answer of the rule was specified. For the large-scale language model, GPT-4V proposed in Non-Patent Document 4 was adopted. The temperature of GPT-4V was set to 0.0. As described above, the classification system according to the embodiment was configured.
[0171] (Reference example) FIG. 18 shows a prompt (excluding image data) in a reference example. In the reference example, the components related to tactile data in the above-described embodiment are omitted. In other words, the reference example is configured to obtain the result of identifying the type of object from the large-scale language model by providing the large-scale language model with a prompt consisting of a command sentence, image data, and a list of labels, with the description related to tactile data omitted. In other respects, the configuration of the reference example is set to be the same as that of the embodiment.
[0172] (Dataset) (A) Training dataset Fig. 19 shows the items used to collect the training dataset used in the experiment. First, to construct the reference value database of the example, 32 items shown in Fig. 19 were prepared. The name of each item was set as the name of a category. Of the 32 categories, 27 categories were used for training, and 5 categories were used for validation.
[0173] We also prepared a commercially available robot arm equipped with a parallel gripper. A distributed triaxial tactile sensor was attached to each finger of the parallel gripper. The tactile sensor was configured to obtain 4 × 4 × 3 × 2 = 96-dimensional tactile signals per time step. The gripper was positioned so that the gripper fingers were positioned on both sides of each object. Next, the gripper was driven to close, and the gripper movement was stopped when the tactile sensor value exceeded a threshold. The threshold was set to 0.004, and the stopping time was set to 0.2 seconds. The gripper was then driven to open. A total of 1 second of sequence was recorded as one sample, from 0.4 seconds before the gripper was stopped until the gripper was stopped, the 0.2 seconds while the gripper was stopped, and the 0.4 seconds after the gripper was opened. The sampling frequency was set to 125 Hz. Next, the sequence was subsampled at 25 Hz to expand one sample to five samples. The above process was repeated 10 times for each object. This resulted in 5 × 10 = 50 training samples per category for the tactile data.
[0174] (B) First dataset FIG. 20 shows the food (real objects) and food replicas used to collect the first dataset used in the experiment. The real food and replicas are a combination of items that are difficult to distinguish from each other based on appearance. To evaluate the discrimination accuracy of the example and reference example using this combination, 11 real food items and 11 replicas of the food shown in FIG. 20 were prepared (a total of 22 categories). By omitting subsampling and observing under the same conditions as the training dataset, 10 samples of tactile data were obtained for each category. By photographing each item with a commercially available smartphone, two samples of image data were obtained for each category. By assigning one sample of image data to every five samples of tactile data, 10 sample pairs of tactile data and image data were obtained for each category.
[0175] (C) Second dataset FIG. 21 shows the actual foods used to collect the second dataset used in the experiment. For some foods, it can be difficult to distinguish between boiled and raw. To evaluate the discrimination accuracy of the examples and reference examples in each state, four foods were prepared: boiled pumpkin, raw pumpkin, boiled rice cake, and raw rice cake (four categories / two types of food in total), as shown in FIG. 21. Samples of tactile data and image data were obtained under the same conditions as those for the first dataset. That is, 10 samples of tactile data were obtained for each category. Two samples of image data were obtained for each category. One sample of image data was commonly assigned for every five samples of tactile data, resulting in 10 sample pairs of tactile data and image data for each category.
[0176] (Experimental conditions) For the example, the autoencoder of the example was implemented on a commercially available computer and trained using training samples (50 × 27 = 1350) of 27 categories out of 32 categories in the training dataset. The Adam optimizer with a learning rate of 0.01 was used for training. The batch size was set to 256. Training samples of 5 categories (50 × 5 = 250) was used to monitor the convergence of the training. This machine learning method produced a trained encoder. Then, using the trained encoder, all training samples (50 x 32 = 1600) were converted into reference values, and a database was constructed using the reference values.
[0177] Next, for each of the first and second datasets, prompts were generated from each sample of tactile data and image data for each category, and the generated prompts were fed into a large-scale language model to obtain the results of identifying the type of each object. The resulting identification results were compared with the label (category) of each object to determine whether the identification was successful. By performing this series of processes for each sample pair and category, 10 results of identification success or failure were obtained for each category. The identification accuracy was calculated for each category based on the obtained results, and the obtained identification accuracy was visualized using a confusion matrix.
[0178] In addition, for the reference example, prompts were generated from image data samples of each category for each of the first and second datasets, and the generated prompts were provided to the large-scale language model to obtain answers indicating the results of identifying the type of each item. In this way, in the reference example, the classification task was performed without processing the tactile data. Considering that the large-scale language model does not necessarily output the same answer each time, in the reference example, the classification task was repeated five times for one sample, obtaining 10 classification success / failure results for each category. Other than these conditions, the classification accuracy in the reference example was calculated for each category under the same conditions as in the example, and the obtained classification accuracy was visualized using a confusion matrix.
[0179] (result) 22 and 23 show the calculation results of the classification accuracy of the first dataset according to the reference example and the working example (FIG. 22 is the reference example, and FIG. 23 is the working example). 24 and 25 show the calculation results of the classification accuracy of the second dataset according to the reference example and the working example (FIG. 24 is the reference example, and FIG. 25 is the working example). In each figure, the vertical axis represents the true value, and the horizontal axis represents the classification result. As shown in each figure, the classification accuracy of the working example exceeded that of the reference example for both the first dataset and the second dataset. This result shows that zero-shot classification is possible by converting data not handled by a large-scale language model into a text representation and providing the obtained text representation to the large-scale language model. Therefore, it was found that the above embodiment makes it possible to build an effective classification system without fine-tuning.
[0180] This specification includes the following disclosure. [Appendix 1] A first sample (35) of a first data type (30) is detected by a first sensor (S1). obtaining a textual representation (55) of the first sample (35) generated by observing the elephant according to a category to which the first sample (35) belongs; providing a prompt (60) including an instruction statement (50) for instructing the large-scale language model (6) to identify features in the object and the text representation (55) of the first sample (35) to the large-scale language model (6), thereby obtaining an answer (65) indicating the results of identifying the features from the large-scale language model (6); outputting information about the obtained results; A control unit (11) configured to execute Identification device (1). [Appendix 2] The prompt (60) further includes a list (59) of candidate labels for the feature identification results; identifying the feature comprises selecting a label from the list (59) that corresponds to the feature; 1. The identification device (1) according to claim 1. [Appendix 3] the control unit (11) is further configured to acquire second samples (45) of a second data type (40) different from the first data type (30), the second samples (45) being generated by observing an object with a second sensor; the large-scale language model (6) is configured to receive input of samples of the second data type (40); The prompt (60) further includes the second sample (45). An identification device (1) according to appendix 1 or appendix 2. [Appendix 4] The second data type (40) is image data (D1). 1. The identification device (1) according to claim 3. [Appendix 5] The first data type (30) is tactile data (C1); 1. The identification device (1) according to claim 4. [Appendix 6] The first data type (30) is temperature data (C2). 1. The identification device (1) according to claim 4. [Appendix 7] The first data type (30) is sound data (C3). 1. The identification device (1) according to claim 4. [Appendix 8] The first data type (30) is point cloud data (C4). 1. The identification device (1) according to claim 4. [Appendix 9] The first data type (30) is weight data (C5). 1. The identification device (1) according to claim 4. [Appendix 10] The first data type (30) is odor data (C6). 1. The identification device (1) according to claim 4. [Appendix 11] The first data type (30) is inertia characteristic data (C7). 1. The identification device (1) according to claim 4. [Appendix 12] The directive (50) A second sample (35) from which the feature is identified is generated. 1. A first sub-sentence (501) instructing to derive a provisional result; a second sub-sentence (502) instructing to derive a second interim result identifying the feature from the second sample (45); and a third sub-sentence (503) instructing to derive a result of identifying the feature according to the first interim result and the second interim result; Including, An identification device (1) according to any one of appendices 3 to 11. [Appendix 13] The control unit (11) obtaining said first sample (35); and Determining the category to which the acquired first sample (35) belongs; and further configured to perform Determining the category to which the first sample (35) belongs includes: converting the first samples (35) into feature values (F1) using a trained encoder (EN) generated by machine learning; and extracting one or more categories to which the first sample (35) belongs from among a plurality of categories respectively assigned to the training samples (TR) by comparing the feature value (F1) obtained from the first sample (35) with a reference value (R0) obtained by converting each training sample (TR) used in the machine learning into the feature (F0) using the trained encoder (EN); Including, the textual representations (55) of the first sample (35) include textual representations of the one or more extracted categories; An identification device (1) according to any one of appendices 1 to 12. [Appendix 14] An identification program (81) for causing a computer (1) to execute an identification method, The identification method includes: obtaining a text representation (55) of a first sample (35) of a first data type (30) according to a category to which the first sample (35) belongs, the first sample (35) being generated by observing an object with a first sensor (S1); providing a prompt (60) including an instruction sentence (50) for instructing the large-scale language model (6) to identify features in the object and the text representation (55) of the first sample (35) to the large-scale language model (6), thereby obtaining an answer (65) indicating the results of identifying the features from the large-scale language model (6); outputting information about the obtained results; Including, Identification program (81). [Appendix 15] The prompt (60) further includes a list (59) of candidate labels for the feature identification results; identifying the feature comprises selecting a label from the list (59) that corresponds to the feature; The identification program (81) described in Appendix 14. [Appendix 16] The method further includes obtaining second samples (45) of a second data type (40) different from the first data type (30), the second samples (45) being generated by observing the object with a second sensor; the large-scale language model (6) is configured to receive input of samples of the second data type (40); The prompt (60) further includes the second sample (45). An identification program (81) according to appendix 14 or appendix 15. [Appendix 17] The directive (50) a first sub-sentence (501) instructing to derive a first interim result identifying the feature from the textual representation (55) of the first sample (35); a second sub-sentence (502) instructing to derive a second interim result identifying the feature from the second sample (45); and a third sub-sentence (503) instructing to derive a result of identifying the feature according to the first interim result and the second interim result; Including, The identification program (81) described in Appendix 16. [Appendix 18] A computer-implemented identification method comprising: The identification method includes: obtaining a text representation (55) of a first sample (35) of a first data type (30) according to a category to which the first sample (35) belongs, the first sample (35) being generated by observing an object with a first sensor (S1); providing a prompt (60) including an instruction sentence (50) for instructing the large-scale language model (6) to identify features in the object and the text representation (55) of the first sample (35) to the large-scale language model (6), thereby obtaining an answer (65) indicating the results of identifying the features from the large-scale language model (6); outputting information about the obtained results; Including, Identification method. [Appendix 19] The prompt (60) further includes a list (59) of candidate labels for the feature identification results; identifying the feature comprises selecting a label from the list (59) that corresponds to the feature; Identification method as described in Appendix 18. [Appendix 20] The method further includes obtaining second samples (45) of a second data type (40) different from the first data type (30), the second samples (45) being generated by observing the object with a second sensor; the large-scale language model (6) is configured to receive input of samples of the second data type (40); The prompt (60) further includes the second sample (45). 1. The identification method according to claim 18 or 19. [Explanation of symbols]
[0181] 1...Identification device, 11...control unit, 12...storage unit, 13...external interface, 14...input device, 15...output device, 16...drive, 81...Identification program, 91...Storage medium, 111...acquisition unit, 112...identification unit, 113...output processing unit, 30...first data type, 35...first sample, 40...second data type, 45...second sample, 50...Directives, 55...Text Expressions, 59...Lists, 60 prompts, 6 large-scale language models, S1...first sensor, S2...second sensor, D1: Image data, D2: Sound data
Claims
1. obtaining a textual representation of a first sample of a first data type according to a category to which the first sample belongs, the first sample being generated by observing an object with a first sensor; providing a prompt to a large-scale language model, the prompt including a command statement for instructing the large-scale language model to identify a feature in the object and the text representation of the first sample, to obtain an answer from the large-scale language model indicating a result of identifying the feature; and outputting information about the obtained results; a control unit configured to perform Identification device.
2. the prompt further includes a list of candidate labels for the feature identification; identifying the feature comprises selecting a label from the list that corresponds to the feature; The identification device according to claim 1 .
3. the controller is further configured to acquire second samples of a second data type different from the first data type, the second samples being generated by observing an object with a second sensor; the large-scale language model is configured to accept input of samples of the second data type; the prompt further includes the second sample. The identification device according to claim 1 .
4. the second data type is image data; The identification device according to claim 3 .
5. the first data type is haptic data; The identification device according to claim 4 .
6. the first data type is temperature data; The identification device according to claim 4 .
7. the first data type is sound data; The identification device according to claim 4 .
8. the first data type is point cloud data; The identification device according to claim 4 .
9. the first data type is weight data; The identification device according to claim 4 .
10. The first data type is odor data. The identification device according to claim 4 .
11. the first data type is inertial characteristic data; The identification device according to claim 4 .
12. The directive: a first sub-sentence directing deriving a first interim result identifying the feature from the textual representation of the first sample; a second sub-sentence instructing deriving a second interim result identifying the feature from the second sample; and a third sub-sentence indicating deriving the feature identification result in response to the first interim result and the second interim result; Including, The identification device according to claim 3 .
13. The control unit obtaining the first sample; and determining a category to which the obtained first sample belongs; and further configured to perform Determining a category to which the first sample belongs includes: converting the first samples into feature values using a trained encoder generated by machine learning; and extracting one or more categories to which the first sample belongs from among a plurality of categories assigned to each training sample by comparing the feature values obtained from the first sample with reference values obtained by converting each training sample used in the machine learning into the feature values using the trained encoder; Including, the textual representation of the first sample includes a textual representation of the one or more extracted categories; The identification device according to claim 1 .
14. An identification program for causing a computer to execute an identification method, The identification method includes: obtaining a textual representation of a first sample of a first data type according to a category to which the first sample belongs, the first sample being generated by observing an object with a first sensor; providing a prompt to a large-scale language model, the prompt including a command statement for instructing the large-scale language model to identify features in the object and the text representation of the first sample, to obtain an answer from the large-scale language model indicating the results of identifying the features; and outputting information about the obtained results; Including, Identification program.
15. the prompt further includes a list of candidate labels for the feature identification; identifying the feature comprises selecting a label from the list that corresponds to the feature; The identification program according to claim 14.
16. The method further includes acquiring a second sample of a second data type different from the first data type, the second sample being generated by observing the object with a second sensor; the large-scale language model is configured to accept input of samples of the second data type; the prompt further includes the second sample. The identification program according to claim 14.
17. The directive: a first sub-sentence directing deriving a first interim result identifying the feature from the textual representation of the first sample; a second sub-sentence instructing deriving a second interim result identifying the feature from the second sample; and a third sub-sentence indicating deriving the feature identification result in response to the first interim result and the second interim result; Including, The identification program according to claim 16.
18. 1. A computer-implemented method of identification, comprising: The identification method includes: obtaining a textual representation of a first sample of a first data type according to a category to which the first sample belongs, the first sample being generated by observing an object with a first sensor; providing a prompt to a large-scale language model, the prompt including a command statement for instructing the large-scale language model to identify features in the object and the text representation of the first sample, to obtain an answer from the large-scale language model indicating the results of identifying the features; and outputting information about the obtained results; Including, Identification method.
19. the prompt further includes a list of candidate labels for the feature identification; identifying the feature comprises selecting a label from the list that corresponds to the feature; 20. The method of claim 18.
20. The method further includes acquiring a second sample of a second data type different from the first data type, the second sample being generated by observing the object with a second sensor; the large-scale language model is configured to accept input of samples of the second data type; the prompt further includes the second sample.
20. The method of claim 18.