Identification device, identification program, and identification method

The use of large-scale language models for zero-shot classification and multiple data types addresses the high cost and complexity of building identification systems, enhancing accuracy and reducing costs.

WO2026004584A1PCT designated stage Publication Date: 2026-01-02OMRON CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/021014
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-28
Filing Date
2025-06-10
Publication Date
2026-01-02

Smart Images

  • Figure JP2025021014_02012026_PF_FP_ABST
    Figure JP2025021014_02012026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a technique for reducing the cost required for constructing an identification system in an arbitrary modality. An identification device according to one embodiment of the present invention acquires the text expression of a first sample which is of a first data type and is generated by observing an object by a first sensor, in accordance with a category to which the first sample belongs, acquires an answer indicating a result of having identified a feature from a large-scale language model by applying a prompt including the text expression of the first sample and a directive sentence for dictating identification of the feature of the object to the large-scale language model, and outputs information pertaining to the acquired result.
Need to check novelty before this filing date? Find Prior Art

Description

Identification device, identification program, and identification method

[0001] The present invention relates to an identification device, an identification program, and an identification method.

[0002] In recent years, technological development of large-scale models such as large-scale language models (LLM) and large-scale visual language models (VLM) has progressed. For example, Non-Patent Document 1 proposes preparing 43,741 combinations of tactile data, image data, and annotations, and using the prepared combinations to fine-tune the large-scale visual language model LLaMA2 using LoRA.

[0003] Letian Fu, et al. “A Touch, Vision, and Language Dataset for Multimodal Alignment”, [online], [Retrieved June 17, 2007], Internet <URL: https: / / tactile-vlm.github.io> Sai Shashank Kalakonda, et al. “Action-GPT: Leveraging Large-scale Language Models for Improved and Generalized Action Generation”, [online], [Retrieved June 17, 2007], Internet <URL: https: / / arxiv.org / abs / 2211.15603> OpenAI, “GPT-4 Technical Report”, [online], [Retrieved June 17, 2007] インターネット<URL:https: / / arxiv.org / abs / 2303.08774>Zhengyuan Yang, et al. "The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)", [online], [Reiwa June 17, 2017], インターネット<URL:https: / / arxiv.org / abs / 2309.17421>Swarup Ranjan Behera, et al. "AQUALLM: Audio Question Answering Data Generation Using Large Language Models", [online], [Reiwa 6th June 17th], インターネット<URL:https: / / arxiv.org / abs / 2312.17343v1>Aaron van den Oord, et. al. "Neural Discrete Representation Learning", [online], [Reiwa June 17, 2017], インターネット<URL:https: / / arxiv.org / abs / 1711.00937>

[0004] The present inventors have found that the above-mentioned conventional methods have the following problems. Conventional methods build large-scale models corresponding to modalities for which no large-scale models have been obtained (tactile data in the method of Non-Patent Document 1). Fine-tuning a large-scale model for each new modality requires time and effort, such as collecting training samples, which increases costs. In other words, with conventional methods, building a classification system using a large-scale model for any modality can be expensive.

[0005] In one aspect, the present invention has been made in view of the above circumstances, and an object of the present invention is to provide a technique for reducing the cost required to build an identification system in any modality.

[0006] In order to solve the above-mentioned problems, the present invention employs the following configurations. Note that the following configurations of the invention can be combined as appropriate.

[0007] According to one aspect of the present invention, an identification device includes a control unit configured to: obtain a text representation of a first sample of a first data type, the first sample being generated by observing an object with a first sensor, according to a category to which the first sample belongs; provide a prompt to a large-scale language model, the prompt including a command statement for instructing the large-scale language model to identify a feature in the object and the text representation of the first sample; and obtain an answer indicating a result of identifying the feature from the large-scale language model; and output information about the obtained result.

[0008] Some studies, such as Non-Patent Document 2, have reported that large-scale language models can acquire common sense. This configuration utilizes this property to achieve zero-shot classification. That is, a text expression corresponding to the category of a first sample is provided to the large-scale language model. The large-scale language model can derive a feature classification result from the category of the first sample in reference to common sense. Therefore, this configuration makes it possible to achieve zero-shot classification of target features without fine-tuning the large-scale language model. This reduces the cost of building a classification system for any modality.

[0009] In the identification device according to the above aspect, the prompt may further include a list of labels that are candidates for the identification result of the feature, and identifying the feature may be configured by selecting a label corresponding to the feature from the list. With this configuration, when an identification range is given, information indicating the identification range can be provided to the large-scale language model as a list, thereby narrowing the range of the identification result. This can be expected to improve identification accuracy compared to when the identification range is unlimited.

[0010] In the identification device according to the above aspect, the control unit may further be configured to acquire second samples of a second data type different from the first data type, the second samples being generated by observing an object with a second sensor. The large-scale language model may be configured to accept input of the samples of the second data type. The prompt may further include the second samples.

[0011] Large-scale language models that can accept input of data other than text, such as large-scale visual language models, already exist. This configuration can reduce the cost of building a classification system when using such large-scale language models that can accept input of data in modalities other than text. Furthermore, the use of such large-scale language models can be expected to improve classification accuracy.

[0012] In the classification device according to the above aspect, the second data type may be image data. With this configuration, it is possible to reduce the cost required to build a classification system in a situation where a large-scale visual language model is used.

[0013] In the identification device according to the above aspect, the first data type may be tactile data, and the second data type may be image data. With this configuration, it is possible to reduce the cost of building an identification system when identifying features using tactile data and image data.

[0014] In the identification device according to the above aspect, the first data type may be temperature data, and the second data type may be image data. With this configuration, when identifying features using temperature data and image data, it is possible to reduce the cost required to build an identification system.

[0015] In the identification device according to the above aspect, the first data type may be sound data, and the second data type may be image data. With this configuration, it is possible to reduce the cost required to build an identification system when identifying features using sound data and image data.

[0016] In the identification device according to the above aspect, the first data type may be point cloud data, and the second data type may be image data. With this configuration, it is possible to reduce costs involved in building an identification system when identifying features using point cloud data and image data.

[0017] In the identification device according to the above aspect, the first data type may be weight data, and the second data type may be image data. With this configuration, when identifying features using weight data and image data, it is possible to reduce the cost required to build an identification system.

[0018] In the identification device according to the above aspect, the first data type may be odor data, and the second data type may be image data. With this configuration, it is possible to reduce the cost of building an identification system when identifying features using odor data and image data.

[0019] In the identification device according to the above aspect, the first data type may be inertial characteristic data, and the second data type may be image data. With this configuration, when identifying features using the inertial characteristic data and the image data, it is possible to reduce the cost required to build an identification system.

[0020] In the identification device according to the above aspect, the command statement may include a first subsentence instructing to derive a first interim result identifying the feature from the text representation of the first sample, a second subsentence instructing to derive a second interim result identifying the feature from the second sample, and a third subsentence instructing to derive a result identifying the feature based on the first interim result and the second interim result. This configuration makes it possible to prevent the modality of either the first data type or the second data type from being ignored during the identification process. As a result, improved identification accuracy can be expected.

[0021] In the identification device according to the above aspect, the control unit may further be configured to acquire the first sample and determine a category to which the acquired first sample belongs. Determining the category to which the first sample belongs may include converting the first sample into a feature value using a trained encoder generated by machine learning, and extracting one or more categories to which the first sample may belong from among multiple categories assigned to the training samples by comparing the feature value obtained from the first sample with a reference value obtained by converting each training sample used in the machine learning into the feature using the trained encoder. The text representation of the first sample may include text representations of the one or more extracted categories. This configuration makes it possible to accurately identify the category to which the first sample belongs. This allows an accurate text representation to be provided to a large-scale language model, thereby ensuring identification accuracy.

[0022] Note that the form of the present invention is not limited to the above-described identification device (information processing device). As another mode of the identification device according to each of the above aspects, one aspect of the present invention may be an information processing method (identification method) that realizes all or part of each of the above configurations, or may be a program, or may be a storage medium readable by a machine such as a computer that stores such a program. A storage medium readable by a machine such as a computer may be a non-transitory medium that stores information such as a program by electrical, magnetic, optical, mechanical, or chemical action. The non-transitory storage medium may include a storage medium (CD, DVD, semiconductor memory, etc.), an auxiliary storage device of a computer, an external storage device connected to a computer, etc.

[0023] For example, an identification program according to an aspect of the present invention may be a program for causing a computer to execute an identification method, which may include the steps of: acquiring a text representation of a first sample of a first data type, the first sample being generated by observing an object with a first sensor, according to a category to which the first sample belongs; providing a prompt to a large-scale language model, the prompt including a command statement for instructing the large-scale language model to identify a feature in the object and the text representation of the first sample, thereby acquiring an answer indicating a result of identifying the feature from the large-scale language model; and outputting information about the acquired result.

[0024] In the identification program according to the above aspect, the prompt may further include a list of labels that are candidates for identifying the feature, and identifying the feature may be performed by selecting a label corresponding to the feature from the list.

[0025] In the identification program according to the above aspect, the identification method may further include acquiring second samples of a second data type different from the first data type, the second samples being generated by observing an object with a second sensor. The large-scale language model may be configured to accept input of the second data type samples. The prompt may further include the second samples.

[0026] In the identification program according to the above aspect, the command statement may include a first sub-sentence instructing to derive a first interim result that identifies the feature from the text representation of the first sample, a second sub-sentence instructing to derive a second interim result that identifies the feature from the second sample, and a third sub-sentence instructing to derive a result that identifies the feature based on the first interim result and the second interim result.

[0027] Furthermore, for example, an identification method according to an aspect of the present invention may be executed by a computer, and may include the steps of: acquiring a text representation of a first sample of a first data type, the first sample being generated by observing an object with a first sensor, according to a category to which the first sample belongs; providing a prompt to a large-scale language model, the prompt including a command statement for instructing the large-scale language model to identify a feature in the object and the text representation of the first sample, thereby acquiring an answer indicating a result of identifying the feature from the large-scale language model; and outputting information about the acquired result.

[0028] In the identification method according to the above aspect, the prompt may further include a list of labels that are candidates for identifying the feature, and identifying the feature may be performed by selecting a label that corresponds to the feature from the list.

[0029] The identification method according to the above aspect may further include acquiring second samples of a second data type different from the first data type, the second samples being generated by observing an object with a second sensor. The large-scale language model may be configured to accept input of the samples of the second data type. The prompt may further include the second samples.

[0030] In the identification method according to the above aspect, the command sentence may include a first sub-sentence instructing to derive a first interim result that identifies the feature from the text representation of the first sample, a second sub-sentence instructing to derive a second interim result that identifies the feature from the second sample, and a third sub-sentence instructing to derive a result that identifies the feature based on the first interim result and the second interim result.

[0031] According to the present invention, it is possible to provide a technique for reducing the cost required to build an identification system in any modality.

[0032] FIG. 1 schematically shows an example of a scenario to which the present invention is applied. FIG. 2 schematically shows an example of a prompt. FIG. 3 schematically shows an example of a method for classifying categories of a first sample. FIG. 4 schematically shows an example of a large-scale language model. FIG. 5 schematically shows an example of a command sentence. FIG. 6 schematically shows an example of a first case to which the present invention is applied. FIG. 7 schematically shows an example of a second case to which the present invention is applied. FIG. 8 schematically shows an example of a third case to which the present invention is applied. FIG. 9 schematically shows an example of a fourth case to which the present invention is applied. FIG. 10 schematically shows an example of a fifth case to which the present invention is applied. FIG. 11 schematically shows an example of a sixth case to which the present invention is applied. FIG. 12 schematically shows an example of a seventh case to which the present invention is applied. FIG. 13 schematically shows an example of the hardware configuration of a recognition device. FIG. 14 schematically shows an example of the software configuration of a recognition device. FIG. 15 is a flowchart showing an example of a processing procedure of the recognition device. FIG. 16 shows the configuration of an identification system in an example. FIG. 17 shows a prompt in an example. FIG. 18 shows a prompt in a reference example. FIG. 19 shows items used to collect the training dataset used in the experiment. FIG. 20 shows food (real objects) and food replicas used to collect the first dataset used in the experiment. FIG. 21 shows food (real objects) used to collect the second dataset used in the experiment. FIG. 22 shows the classification accuracy (results) of the first dataset according to the reference example. FIG. 23 shows the classification accuracy (results) of the first dataset according to the example. FIG. 24 shows the classification accuracy (results) of the second dataset according to the reference example. FIG. 25 shows the classification accuracy (results) of the second dataset according to the example.

[0033] An embodiment according to one aspect of the present invention will be described below with reference to the drawings. However, the embodiment described below is merely an example of the present invention in all respects. Various improvements or modifications may be made without departing from the scope of the present invention. In implementing the present invention, a specific configuration according to the embodiment may be appropriately adopted. Note that, although data appearing in this embodiment is described in natural language, more specifically, it is specified using pseudo-language, commands, parameters, machine language, etc. that can be recognized by a computer.

[0034] §1 Application Example FIG. 1 schematically illustrates an example of a scenario in which the present invention is applied. The identification device 1 according to this embodiment is one or more computers configured to perform an identification task using a large-scale language model 6. In this embodiment, a first sample 35 of a first data type 30 is generated by observing an object with a first sensor S1. The identification device 1 acquires a text representation 55 of the first sample 35 according to the category to which the first sample 35 belongs. The identification device 1 provides the large-scale language model 6 with a prompt 60 including a command statement 50 for instructing the large-scale language model 6 to identify features in the object and the text representation 55 of the first sample 35, thereby acquiring an answer 65 indicating the results of the feature identification from the large-scale language model 6. The identification device 1 outputs information related to the acquired results.

[0035] The large-scale language model 6 can acquire common sense during the machine learning process of a huge amount of data. In this embodiment, the first sample 35 is converted into a text representation 55 for input to the large-scale language model 6. The text representation 55 corresponding to the category of the first sample 35 is provided to the large-scale language model 6 as a prompt 60. When generating an answer 65 corresponding to the command statement 50, the large-scale language model 6 can derive a feature identification result from the category of the first sample 35 using the acquired common sense. Therefore, according to this embodiment, by converting the first sample 35 of the first data type 30 into the text representation 55, zero-shot feature identification of a target can be achieved without fine-tuning the large-scale language model 6. This reduces the cost of building an identification system for any modality.

[0036] [Identification Task] The identification task may include any task that classifies the features of an object. The object may be selected as appropriate depending on the embodiment from, for example, an object, a living thing (such as a person), etc. The object may be an actual entity or a virtual entity. The features may include, for example, attributes, states, etc. The attributes are static features such as type. The states are dynamically changeable features such as the presence or absence of a malfunction, the state of a chemical reaction, hardness, temperature, and posture. Identifying may include regressing the degree of match (likelihood, etc.) to the features of the object to be identified.

[0037] Prompt The prompt 60 is an input to the large-scale language model 6. The prompt 60 may be input directly to the large-scale language model 6, or may be pre-processed before being input to the large-scale language model 6. In this embodiment, the prompt 60 is configured to include the command sentence 50 and the text representation 55 of the first sample 35.

[0038] The command statement 50 is composed of a text expression that instructs the large-scale language model 6 to identify features in an object. The content and expression of the command statement 50 are not particularly limited as long as they are configured to indicate an identification command, and may be determined appropriately depending on the embodiment. In one example, the command statement 50 may be composed of text of a command statement instructing the identification of features contained in a sample.

[0039] The text expression 55 is configured to indicate the category to which the first sample 35 belongs. The content of the text expression 55 is not particularly limited as long as it corresponds to the result of categorization of the first sample 35, and may be determined depending on the embodiment. In one example, the text expression 55 may be configured with text indicating that the first sample 35 resembles a type that fits the category. Note that the number of categories extracted as the category to which the first sample 35 belongs may be one or more. In the latter case, in categorization of the first sample 35, multiple categories to which the first sample 35 is likely to belong may be extracted.

[0040] Categories may be defined as appropriate. They may be set manually or automatically using a clustering or other technique. For example, categories may be defined based on features (such as object names) that appear in samples (such as training samples) used to generate a computational model (such as an encoder, described below) used to categorize the first sample 35. The features used to set the categories and the features to be identified in the identification task may be unrelated to each other, or may be at least partially related to each other. In one example, the category set for the first sample 35 may be related to answer candidates (candidates for feature identification results) of the identification task. For example, if the identification task is to identify the type of item, the category for the first sample 35 may be set according to the type of item. However, the values ​​of the category set for the first sample 35 do not necessarily overlap with the answer candidates of the identification task. For example, if the identification task is to identify the type of food, the category set for the first sample 35 may or may not at least partially include the type of food to be identified. Furthermore, the category set for the first sample 35 may include the type of item other than food. It is assumed that the large-scale language model 6 is more likely to apply common sense to commonly used terms than to terms used only in a limited range. Therefore, when the text expression 55 is composed of text (or its tokens) including category names, it is desirable that the category names (values) be defined as general terms, including academic terms. For example, the general terms may be terms registered in any dictionary. This is expected to improve the recognition accuracy.

[0041] In this embodiment, the configuration of the prompt 60 may be determined appropriately depending on the embodiment, as long as it includes the command statement 50 and the text representation 55 of the first sample 35. The prompt 60 may be composed of only the command statement 50 and the first sample 35, or may further include any information (data) other than the command statement 50 and the text representation 55 of the first sample 35.

[0042] FIG. 2 schematically illustrates an example of a prompt 60 according to this embodiment. As illustrated in FIG. 2 , the prompt 60 may further include, in addition to the command statement 50 and the text expression 55, a list 59 of candidate labels for feature identification results. The list 59 is an example of arbitrary information. In this case, identifying a feature may be performed by selecting a label from the list 59 that corresponds to a feature included in the target. The command statement 50 may be configured to instruct the large-scale language model 6 to at least one of select a corresponding label from the list 59 and regress the degree of correspondence to each label.

[0043] According to one example of this embodiment, when an identification range is given, the range of the identification result can be narrowed by providing information indicating the identification range as a list 59 to the large-scale language model 6. This makes it possible to expect improved identification accuracy compared to when the identification range is unlimited.

[0044] (Data Format) In a typical example, the prompt 60 (such as the command statement 50, the text expression 55, and the list 59) may be composed of text data. Each category to which the first sample 35 is a candidate may be associated with text data, and the text expression 55 may be composed of text data corresponding to the classification result of the category to which the first sample 35 belongs. The text data associated with each category may be text including the name of the category (e.g., text indicating that the data resembles a type that fits the category). However, as long as the data format can be input to the large-scale language model 6, the data format of the prompt 60 is not limited to this example and may be selected appropriately depending on the embodiment. In another example, at least a portion of the prompt 60 may be composed of one or more tokens corresponding to the text data.

[0045] For example, when the text representation 55 is composed of tokens, similar to the above example, text data may be associated with each category to which the first sample 35 is a candidate. The text representation 55 may be obtained as text data according to the category classification results, and then converted into tokens by a tokenizer. In a simple example, the text data associated with each category may be text including the name of each category, and the text representation 55 may be obtained by converting text including a description of the category to which the first sample 35 belongs into tokens.

[0046] Furthermore, for example, by associating the result of category classification with a value of a token, the token (text representation 55) may be obtained directly according to the result of category classification of the first sample 35 without using text data. A token value may be assigned in advance to each category that is a candidate to which the first sample 35 belongs. The token value of each category may be set appropriately. In a simple example, the token value of each category may be obtained by converting text (such as a name) containing a description of each category into tokens using a tokenizer.

[0047] When the number of categories extracted as the category to which the first sample 35 belongs is one, the tokens obtained corresponding to the extracted category may be input to the large-scale language model 6 as the text representation 55. When the number of categories extracted as the category to which the first sample 35 belongs is multiple, statistics of the tokens obtained corresponding to each of the extracted categories may be calculated, and the calculated statistics may be input to the large-scale language model 6 as the text representation 55. The statistics may be arbitrarily selected from, for example, the mean, the median, etc. In a simple example, the statistics may be constituted by a simple average of the values ​​of the tokens in each extracted category. In another example, the values ​​of the tokens in each category may be weighted based on the degree (such as the likelihood) that the first sample 35 belongs to each category. The statistics may also be constituted by a weighted average of the values ​​of the tokens in each extracted category.

[0048] Note that the conversion into tokens is not limited to the method using a tokenizer, and may be performed by other methods, such as using a text-token correspondence table. The token values ​​may be optimized for each category for the identification task using a method such as machine learning. The same applies to the command statement 50 and the list 59, which may be composed of the above tokens.

[0049] [Category Classification] The method for classifying the category of the first sample 35 is not particularly limited and may be selected appropriately depending on the embodiment. The category classification method may employ any method, such as a machine learning method, an analytical method, or other comparison operation method. Machine learning may include, for example, supervised learning, unsupervised learning, etc. Unsupervised learning may include, for example, autoencoder learning, clustering, etc. The autoencoder may include a modified autoencoder, such as a variational autoencoder (VAE) or a vector quantized-variational autoencoder (VQ-VAE). Clustering may include, for example, spectral clustering, etc. Analytical methods may include, for example, principal component analysis, non-negative matrix factorization (NMF), independent component analysis (ICA), etc. Any computational model based on each method may be used for category classification. Depending on the method used for categorization, the categories may be referred to by other terms such as classes, clusters, etc.

[0050] In one example of this embodiment, the identification device 1 acquires a first sample 35 and determines a category to which the acquired first sample 35 belongs. Determining the category to which the first sample 35 belongs may include converting the first sample 35 into feature values ​​using a trained encoder generated by machine learning, and extracting one or more candidate categories to which the first sample 35 belongs from among multiple categories assigned to the training samples by comparing the feature values ​​obtained from the first sample 35 with reference values ​​obtained by converting each training sample used in the machine learning into a feature using the trained encoder. The text representation 55 of the first sample 35 may include text representations of the one or more extracted categories. The trained encoder is an example of a computational model used for category classification.

[0051] FIG. 3 schematically illustrates an example of a method for classifying a category to which a first sample 35 belongs using the trained encoder. In the example of FIG. 3 , an autoencoder is used as a framework for generating a trained encoder (encoder EN). The autoencoder is composed of an encoder EN and a decoder DE. The encoder EN is configured to convert input samples into features. The decoder DE is configured to reconstruct the input samples from the features obtained by the encoder EN (i.e., generate reconstructed samples). Note that the configuration of the autoencoder is not limited to this example and may be modified as appropriate depending on the type of autoencoder employed.

[0052] The encoder EN and the decoder DE may each be configured with a machine learning model. The machine learning model has one or more computational parameters that can be adjusted by machine learning. The one or more computational parameters are used for the desired inference computation. The machine learning model may be configured with, for example, a neural network, a support vector machine, a regression model, or other functional formula (computational model). When at least one of the encoder EN and the decoder DE includes a neural network, the structure of the neural network is not particularly limited and may be determined appropriately depending on the embodiment. The structure of the neural network may be specified, for example, by the number of layers from the input layer to the output layer, the type of each layer, the number of nodes (neurons) included in each layer, the connection relationships between the nodes in each layer, etc. The neural network may include any mechanism, such as a recurrent structure, a self-attention mechanism, or an autoregressive model. The neural network may include any layer, such as a fully connected layer, a convolutional layer, a pooling layer, a deconvolutional layer, an unpooling layer, a normalization layer, a dropout layer, or a long short-term memory (LSTM). The neural network may include any type of model, such as a diffusion model, a transformer model, or a generative model. The weights of the connections between the nodes included in the neural network and the thresholds of the nodes are examples of computational parameters. The machine learning method may be selected appropriately depending on the embodiment of the machine learning model to be adopted (e.g., backpropagation).

[0053] In the example of FIG. 3 , machine learning (unsupervised learning) of the autoencoder is performed in the training (learning) stage. Machine learning involves adjusting (optimizing) the values ​​of calculation parameters using training samples. Machine learning of the autoencoder may be performed as appropriate depending on the embodiment. As an example of machine learning processing, when the encoder EN and the decoder DE are configured as neural networks, training samples TR may be provided to the encoder EN, and forward calculation processing of the encoder EN may be performed. The training samples TR, like the first samples 35, are samples of the first data type 30 and are generated by observing an object using the first sensor S1. The number of training samples TR used for training may be determined as appropriate depending on the embodiment. Each training sample TR may be collected as appropriate. As a result of the calculation processing of the encoder EN, a feature F0 may be obtained from the encoder EN (i.e., the training samples TR are converted into the feature F0). The obtained feature F0 may be provided to the decoder DE, and forward calculation processing of the decoder DE may be performed. As a result of the execution of the decoder DE's computational process, a reconstructed sample RC can be obtained from the decoder DE. A loss may be calculated by calculating a reconstruction error (difference) between the reconstructed sample RC and the corresponding training sample TR. The reconstruction error is an example of a loss. The loss in the autoencoder may further include errors other than the reconstruction error, such as a regularization term, a vector quantization error, and a commitment error. Then, the values ​​of the computation parameters of the encoder EN and the decoder DE may be adjusted (optimized) so as to reduce the calculated loss. In the example of FIG. 3 , the gradient of the reconstruction error may be calculated and the calculated gradient may be backpropagated to calculate the error in the values ​​of each computation parameter of the decoder DE and the encoder EN. The values ​​of each computation parameter of the decoder DE and the encoder EN may be updated based on the calculated errors. The degree to which the values ​​of the computation parameters are updated may be adjusted by a learning rate. A series of processes from providing the training sample TR to adjusting the values ​​of the computation parameters (computation parameter adjustment process) may be repeatedly executed until a predetermined condition is met, such as the loss becoming less than a threshold or a predetermined number of repetitions.As a result of repeatedly performing this adjustment process, a trained encoder EN and a trained decoder DE can be generated.

[0054] The training samples TR may be prepared for each category. When preparing the training samples TR, multiple categories may be defined as appropriate. In one example, each category may be defined according to a feature (such as a type of object) commonly appearing in the training samples TR to which it belongs. Each category may be defined manually or automatically using a method such as clustering. Each training sample TR is assigned to at least one of the multiple categories.

[0055] After the above machine learning is completed, each training sample TR may be provided to a trained encoder EN, and the trained encoder EN may execute a calculation process to convert each training sample TR used in the machine learning into a feature F0. This allows a reference value R0 that can be used for category classification in the inference stage to be obtained. That is, the value of the feature F0 calculated from each training sample TR using the trained encoder EN may be collected as a reference value R0 to be used for category classification. Multiple reference values ​​R0 may be collected for each category, and each category (reference value R0) may be associated with a text expression such as the name of the category. This allows a database that can be used for conversion between categories (reference values ​​R0) and text to be constructed.

[0056] Note that at least a part of the pre-processing, including the machine learning process for generating the trained encoder EN and the process for collecting the reference value R0 for each category, may be executed on the classification device 1 or may be executed on an external computer other than the classification device 1. When the machine learning process is executed on an external computer, in one example, the classification device 1 may acquire the trained encoder EN directly or indirectly from the external computer at any timing and by any method. Acquiring it indirectly means acquiring it via a storage medium, another computer, etc. In another example, the trained encoder EN may be pre-installed in the classification device 1. Similarly, when the reference value R0 for each category is generated by an external computer, in one example, the classification device 1 may acquire the reference value R0 for each category directly or indirectly from the external computer. In another example, the reference value R0 for each category may be pre-installed in the classification device 1.

[0057] On the other hand, in the example of FIG. 3 , in the inference stage, the identification device 1 may acquire a first sample 35. The method of acquiring the first sample 35 may be selected appropriately depending on the embodiment. In one example, the identification device 1 may acquire the first sample 35 directly or indirectly from the first sensor S1. The first sample 35 may be acquired in response to an operator's operation or automatically. The identification device 1 may convert the acquired first sample 35 into a feature (value F1) using a trained encoder EN. The identification device 1 may compare the feature value F1 obtained from the first sample 35 with a reference value R0 obtained from each training sample TR. Depending on the result of this comparison, the identification device 1 may extract one or more candidate categories to which the first sample 35 belongs from among multiple categories assigned to each training sample TR. In one example, the identification device 1 may search for a neighborhood of the value F1 in a database (latent space) of reference values ​​R0. In this proximity search, the classification device 1 may extract k categories (k is 1 or greater) by sequentially selecting categories to which the reference values ​​R0 belong, starting with the reference value R0 closest to the value F1. The classification device 1 may acquire the extracted k categories as candidates for the category to which the first sample 35 belongs. This allows the classification device 1 to extract one or more categories from multiple categories that are candidates for the category to which the first sample 35 belongs. That is, the classification device 1 can obtain a determination (classification) result for the category to which the acquired first sample 35 belongs. The classification device 1 may construct a text representation 55 of the first sample 35 using text representations of the one or more extracted categories. According to an example of this embodiment, it is possible to accurately identify the category to which the first sample 35 belongs. This allows an accurate text representation 55 to be provided to the large-scale language model 6, which is expected to result in improved classification accuracy.

[0058] Note that the text expression may be a fixed value or a variable value among reference values ​​R0 belonging to the same category. In the latter case, the text expression may be associated with each reference value R0. Associating the text expression with a category may include associating the text expression with the reference value R0.

[0059] The trained encoder EN may also be generated by methods other than the machine learning of the autoencoder. In another example, the trained encoder EN may be generated by supervised learning. In yet another example, the trained encoder EN may be generated by an analytical method. For example, the trained encoder EN may be composed of principal component vectors obtained by principal component analysis of multiple training samples TR.

[0060] Furthermore, the classification device 1 may classify the category to which the first sample 35 belongs by directly comparing the first sample 35 with the training sample TR, rather than comparing features in a latent space. That is, the classification device 1 may compare each sample (35, TR) without compressing them into features. The classification device 1 can extract one or more categories to which the first sample 35 belongs using a method similar to the above-described feature-based method, except that the objects to be compared are replaced by each sample (35, TR) instead of each value (F1, R0). Note that when this method is adopted, each training sample TR may be read as a reference sample, etc.

[0061] Furthermore, the method for classifying the category of the first sample 35 does not need to be limited to the method based on the feature or sample comparison. In another example, a trained classifier configured to derive a category classification result from a sample may be generated. The trained classifier may be generated by machine learning, such as unsupervised learning (e.g., clustering) or supervised learning. The trained classifier is an example of a computational model used for category classification. The identification device 1 may provide the acquired first sample 35 to the trained classifier and execute computational processing of the trained classifier to obtain a category classification result for the first sample 35 from the trained classifier. The identification device 1 may generate a text representation 55 according to the acquired classification result.

[0062] Regardless of which of the above methods is adopted as the method for classifying the category of the first sample 35, at least a part of the pre-processing such as machine learning may be executed on the identification device 1 or may be executed on an external computer other than the identification device 1. When data used for classification (reference samples, trained classifiers, etc.) is generated by an external computer, the identification device 1 may obtain the data directly or indirectly from the external computer.

[0063] In one example of the above embodiment, the identification device 1 executes a series of processes from acquiring the first sample 35 to generating the text representation 55 in accordance with the classification result of the first sample 35. However, the entity that executes at least a part of this series of processes does not have to be limited to the identification device 1, and may be an external computer other than the identification device 1. In another example, the external computer may execute the series of processes from acquiring the first sample 35 to generating the text representation 55. In this case, the identification device 1 may directly or indirectly acquire the text representation 55 corresponding to the first sample 35 from the external computer.

[0064] [Large-Scale Language Model] As long as an answer can be generated from text, the configuration of the large-scale language model 6 is not particularly limited and may be determined appropriately depending on the embodiment. A known model proposed in Non-Patent Document 3 or the like may be adopted as the large-scale language model 6. Furthermore, the large-scale language model 6 may be configured to further accept input of data other than text in addition to text, such as a large-scale visual language model (Non-Patent Document 4 or the like) or an Audio Question Answering Model (Non-Patent Document 5 or the like).

[0065] FIG. 4 schematically illustrates an example of a large-scale language model 6 according to this embodiment. As illustrated in the example of FIG. 4 , the large-scale language model 6 may be configured to accept input of samples of a second data type 40 different from the first data type 30. In response to this, the identification device 1 may further acquire a second sample 45 of the second data type 40. The second sample 45 may be generated by observing an object using a second sensor S2. A method for acquiring the second sample 45 may be selected appropriately depending on the embodiment. In one example, the identification device 1 may acquire the second sample 45 directly or indirectly from the second sensor S2. The prompt 60 may further include the second sample 45 in addition to the command statement 50 and the text expression 55. The second sample 45 may be included in the prompt 60 as is, or may be included in the prompt 60 after applying any preprocessing. According to this example of the present embodiment, when a large-scale language model capable of accepting data input of a modality other than text (the second data type 40) is used as the large-scale language model 6, the cost of building an identification system can be reduced. Furthermore, by using such a large-scale language model, it is possible to expect improvement in identification accuracy. Note that in one example of this embodiment, the form shown in Figure 2 may also be adopted. That is, the prompt 60 may be configured to include the command statement 50, the text representation 55 of the first sample 35, the second sample 45, and the list 59.

[0066] (Data Type / Sensor) The first data type 30 and the second data type 40 may be appropriately selected from data types other than text. When adopting the form of FIG. 4 in which the large-scale language model 6 is configured to be able to accept samples of the second data type 40, it is desirable that the first data type 30 be selected from a minor data type for which large-scale language models that can accept samples (first samples 35) are not generally available. In particular, it is desirable that the first data type 30 be selected from a data type for which large-scale language models that can accept samples are not provided as commercial services. A minor data type may be a data type for which fine-tuning a large-scale language model is costly due to factors such as the lack of a large amount of data or the lack of standardized sensor standards.

[0067] As an example, the first data type 30 may be sensing data such as tactile data, temperature data, point cloud data, weight data, smell data, inertial property data, electromagnetic data, etc. The first sensor S1 may be configured with a tactile sensor, a temperature sensor (such as a thermometer), a point cloud sensor, a load sensor (such as a weigh scale or load meter), a smell sensor, an inertial property measuring device, an electromagnetic sensor (such as an ammeter, voltmeter, or magnetometer), etc. The point cloud sensor may include, for example, a LiDAR (light detection and ranging), an MMS (Mobile Mapping System), a depth sensor, an ultrasonic sensor, an infrared sensor, a radar, etc.

[0068] On the other hand, the second data type 40 is preferably selected from major data types for which large-scale language models expanded to accept samples (second samples 45) are widespread due to factors such as standardized sensor standards and the existence of large amounts of data. In particular, the second data type 40 is preferably selected from data types for which large-scale language models expanded to accept samples are provided as commercial services. For example, the second data type 40 may be at least one of image data D1 and sound data D2. The second sensor S2 may be composed of at least one of an image sensor (such as an RGB camera) and a microphone.

[0069] In one example, by adopting a large-scale visual language model as the large-scale language model 6, the second data type 40 may be image data D1. The image data D1 may be composed of still images or moving images. According to one example of the present embodiment, it is possible to reduce the cost of building a classification system when using a large-scale visual language model. Note that the large-scale visual language model may include a visual question answering model, an open vocabulary object detection model, an open vocabulary object segmentation model, etc.

[0070] In another example, the second data type 40 may be audio data D2 by employing an Audio Question Answering Model as the large-scale language model 6. According to this example of the present embodiment, it is possible to reduce the cost required to build a classification system when using an Audio Question Answering Model.

[0071] However, the relationship between the first data type 30 and the second data type 40 need not be limited to this example. Not only in cases where the configuration of FIG. 4 is not adopted, but also in cases where the configuration of FIG. 4 is adopted, the first data type 30 may be selected from major data types for which large-scale language models capable of accepting samples are widely available. For example, the first data type 30 may be at least one of image data and sound data. Accordingly, the first sensor S1 may be composed of at least one of an image sensor and a microphone.

[0072] In one example, in response to the selection of image data as the second data type 40, a data type other than image data, such as audio data, whether minor or major, may be selected as the first data type 30. This allows the samples of the first data type 30 (first samples 35), which are not normally acceptable, to be reflected in the classification task for the large-scale language model 6 that can accept text and samples of the second data type 40 (second samples 45). As a result, improvement in classification accuracy can be expected.

[0073] Note that any data type other than those described above may be adopted for each data type (30, 40). Each sensor (S1, S2) is any machine configured to observe an object and generate data indicating the observation results. Each sensor (S1, S2) may be configured by a computer. In one example, a sample generated by any calculation process of the computer may be treated as at least one of the first sample 35 and the second sample 45. Furthermore, each sensor (S1, S2) may be at least one of a real sensor and a virtual sensor. Observing may include simulating.

[0074] Furthermore, the number of first data types 30 does not have to be limited to one and may be two or more. The number of first samples 35 for each data type does not have to be limited to one and may be two or more. When multiple first data types 30 are set, the number of first samples 35 applied to a single classification task may be the same among the first data types 30, or may at least partially differ among the first data types 30. A text representation 55 may be obtained for each first sample 35. The text representations 55 for each first sample 35 may be integrated within a prompt 60. As long as the large-scale language model 6 can be accepted, the number of second data types 40 does not have to be limited to one and may be two or more. The number of second samples 45 for each data type does not have to be one and may be two or more.

[0075] Furthermore, the first data type 30 may be configured as a combination of multiple data types (modalities). For example, in the configuration of FIG. 3 , the encoder EN may be configured to accept input samples of each of the multiple data types and calculate features from each input sample. This allows the multiple data types to be considered as a single first data type. If the large-scale language model 6 can accept the second data type 40, the second data type 40 may also be configured as a combination of multiple data types. For example, if the large-scale language model 6 is configured to accept input of video including audio, the second data type 40 may be configured as a combination of audio data and image data.

[0076] In a typical example, data obtained from the same type of sensor may be considered to be the same data type. However, the definition of the data type is not limited to this example. In another example, data obtained from the same type of sensor under different conditions, such as different sensing methods, may be considered to be different data types. For example, when measuring tactile data using a tactile sensor provided in a gripper, a data type may be defined for each gripping method of the gripper (i.e., tactile data obtained using different gripping methods may be considered to be different data types). Accordingly, the first samples 35 of the multiple different first data types 30 may be obtained from the same sensor. In a typical example, the first sensor S1 and the second sensor S2 may be different from each other. In another example, depending on the definition of the first data type 30 and the second data type 40, the first sensor S1 and the second sensor S2 may at least partially overlap. For example, when operating to acquire data under different conditions as different data types, the first sensor S1 and the second sensor S2 may be the same sensor.

[0077] The category of the first sample 35 may be defined to have a correlation with the second data type 40, or may be defined without a correlation with the second data type 40. In one example, the category of the first sample 35 is preferably defined so as to maximize the reduction in uncertainty of the correct answer (answer candidate) of the classification task for the sample of the second data type 40 when the category of the first sample 35 is given. For example, assuming that the category of the first sample 35 is "X1," the sample of the second data type 40 is "X2," and the correct answer of the classification task is "Y," the difference "H(Y;X2) - H(Y;X1,X2)" between the relative entropy of the correct answer of the classification task for the sample of the second data type 40 and the relative entropy of the correct answer of the classification task for the category of the first sample 35 and the sample of the second data type 40 can be used as an index of the reduction in uncertainty. Relative entropy may be expressed in other ways, such as information gain or Kullback-Leibler distance. Samples of the first data type 30 and the second data type 40 may be collected, and the category of the first sample 35 may be defined such that "H(Y;X2)-H(Y;X1,X2)" exceeds a threshold value for the collected samples. The threshold value may be set as appropriate.

[0078] (Command Sentence) When the large-scale language model 6 is configured to further receive input of samples of the second data type 40, the command statement 50 may be configured to instruct the identification of features in the target from the text representation 55 of the first sample 35 and the second sample 45. If configured in this way, the content of the command statement 50 may be determined appropriately depending on the embodiment.

[0079] 5 is a schematic diagram illustrating an example of a command statement 50 according to the present embodiment. In the example of FIG. 5, a list of object types is provided as a list 59, and the command statement 50 and the list 59 are assumed to be composed of text data. As shown in the example of FIG. 5, the command statement 50 may include a presupposition subsentence 500, a first subsentence 501, a second subsentence 502, and a third subsentence 503.

[0080] The preamble 500 is configured to indicate the content of the identification task it instructs. The preamble 500 may be omitted. The first subsentence 501 may be configured to instruct deriving a first interim result of identifying features from the text representation 55 of the first sample 35. In one example, the first subsentence 501 may further include an instruction to ignore the second sample 45 in the process of deriving the first interim result from the text representation 55 of the first sample 35 ("Ignore the 2nd sample at this stage." in FIG. 5 ). The second subsentence 502 may be configured to instruct deriving a second interim result of identifying features from the second sample 45. In one example, the second subsentence 502 may further include an instruction to ignore the text representation 55 of the first sample 35 in the process of deriving the second interim result from the second sample 45 ("Ignore the 1st sample at this stage." in FIG. 5 ). The third sub-sentence 503 may be configured to instruct deriving a result (final result) of identifying features based on the first interim result and the second interim result.

[0081] According to one example of the present embodiment, the command statement 50 includes a first partial sentence 501, a second partial sentence 502, and a third partial sentence 503, thereby making it possible to prevent the modality of either the first data type 30 or the second data type 40 from being ignored during the classification process. Furthermore, the first partial sentence 501 further includes an instruction to ignore the second sample 45, and the second partial sentence 502 further includes an instruction to ignore the text representation 55 of the first sample 35, thereby making it possible to reliably reflect the modality of each of the first data type 30 and the second data type 40 in the classification process. As a result, an improvement in classification accuracy can be expected.

[0082] The content of the command statement 50 is not limited to the example shown in Figure 5 and may be modified as appropriate depending on the embodiment. For example, the command statement 50 shown in Figure 5 does not include any partial sentences other than the partial sentences 500 to 503. However, the configuration of the command statement 50 is not limited to this example and may include additional partial sentences other than the partial sentences 500 to 503. In another example, the command statement 50 may further include a partial sentence that specifies the output format of the answer 65.

[0083] 5, the command statement 50 includes, from top to bottom, a presupposition sentence 500, a first sentence 501, a second sentence 502, and a third sentence 503. However, the order in which the sentences 500 to 503 are written is not limited to this example and may be changed as appropriate depending on the embodiment. In another example, the second sentence 502 may be placed before the first sentence 501.

[0084] Also, in the first sub-sentence 501, the instruction to ignore the second sample 45 may be omitted. In the second sub-sentence 502, the instruction to ignore the text representation 55 of the first sample 35 may be omitted.

[0085] Furthermore, the content of the descriptions in each of the partial sentences 500 to 503 need not be limited to the example in FIG. 5 and may be modified as appropriate depending on the embodiment. For example, in FIG. 5, the first partial sentence 501 and the second partial sentence 502 instruct the user to evaluate the degree of belonging to each category using a score of 0-10. The score assignment results in the provisional result of the classification. However, the score range and the format of the provisional result need not be limited to this example. The score range may be set arbitrarily. The first partial sentence 501 and the second partial sentence 502 may be configured to instruct the derivation of the provisional result in a format other than a score. The list 59 may be omitted, and the description of the classification task in the command statement 50 may be modified accordingly.

[0086] (Deployment Location) The large-scale language model 6 may be deployed at any location. In one example, the large-scale language model 6 may be located in the identification device 1. In this case, the identification device 1 can provide a prompt 60 to the large-scale language model 6 and execute calculation processing on the large-scale language model 6 to obtain an answer 65 indicating the result of identifying the feature from the large-scale language model 6. In another example, the large-scale language model 6 may be deployed in an external computer other than the identification device 1. In this case, providing the prompt 60 to the large-scale language model 6 may be configured by providing the external computer with a request to execute calculation processing on the large-scale language model 6 together with the prompt 60. In response to a request received from the identification device 1, the external computer may provide the prompt 60 to the large-scale language model 6 and execute calculation processing on the large-scale language model 6. As a result, the external computer may generate an answer 65 indicating the result of identifying the feature. The identification device 1 may obtain the generated answer 65 directly or indirectly from the external computer.

[0087] [Specific Example] This embodiment is applicable to various situations in which a classification task is performed. Furthermore, the example of this embodiment shown in FIG. 4 is applicable to various situations in which a classification task is solved using two or more types of data (first data type 30, second data type 40). Each data type (30, 40) may be selected as appropriate depending on the application situation, classification task, and other aspects of the embodiment. Applications of the example of this embodiment shown in FIG. 4 may include at least one of the following first, second, third, fourth, fifth, sixth, and seventh cases. Specific application situations of the embodiment shown in FIG. 4 will be exemplified below for each case.

[0088] 6 is a schematic diagram illustrating an example of a first example to which this embodiment is applied. The first example is an example of a situation in which this embodiment is applied to identify features of an object T1 from tactile data C1 and image data D1.

[0089] 6, the first data type 30 may be tactile data C1. A sample of the tactile data C1 (first sample 35) may be generated by a tactile sensor S11. The tactile sensor S11 is an example of the first sensor S1. The second data type 40 may be image data D1. A sample of the image data D1 (second sample 45) may be generated by an image sensor S21. The image sensor S21 is an example of the second sensor S2. The tactile sensor S11 and the image sensor S21 may be appropriately disposed in a location where the target T1 can be observed.

[0090] The identification device 1 may construct a prompt 60 including a command sentence 50, a text representation 55 of the first sample 35 of the tactile data C1, and a second sample 45 of the image data D1. In the first case, the identification device 1 may construct the prompt 60 to further include a list 59. The identification device 1 may provide the obtained prompt 60 to the large-scale language model 6 to obtain an answer 65 from the large-scale language model 6 indicating the results of identifying features in the object T1.

[0091] The tactile and image-based identification task may be performed for any purpose. The target T1 may be determined appropriately depending on the embodiment. In one example, the target T1 may be a workpiece of the robot device R1, and identifying the characteristics may include identifying the type of the target T1. The tactile sensor S11 and the image sensor S21 may be appropriately positioned so as to be able to observe the workpiece (target T1) of the robot device R1. For example, the tactile sensor S11 may be attached to a robot hand such as a gripper, and the image sensor S21 may be positioned either inside or outside the robot device R1 so that the working range of the robot device R1 is within the imaging range.

[0092] The result of identifying the type of the object T1 may be used to control the robot device R1. For example, consider a scenario in which the robot device R1 is caused to perform a task of gripping the object T1 with a gripper and transporting the object T1 to a destination. In this scenario, in the initial stage of gripping the object T1 with the gripper, a first sample 35 of tactile data C1 may be obtained from the tactile sensor S11 arranged on the gripper. In this initial stage, the gripper may be controlled to grip the object T1 with a force that is too weak to lift the object T1 for transporting it, but is light enough to obtain tactile data C1 that can be used for the identification task. A second sample 45 of the image data D1 may be acquired at any timing. The identification device 1 may generate a prompt 60 from the obtained first sample 35 and second sample 45, and provide the generated prompt 60 to the large-scale language model 6 to obtain a result of identifying the type of the object T1. The identification device 1 may determine the gripping force to be used by the gripper when lifting the object T1 in accordance with the identification result of the type of the object T1, and may issue a command to the robot device R1 to grip the object T1 with the determined force and transport the object T1. Alternatively, the identification device 1 may cause the robot device R1 to determine the force when lifting the object T1 by providing the identification result of the type of the object T1 to the robot device R1.

[0093] The type of the robot device R1 is not particularly limited and may be selected appropriately depending on the embodiment. In one example, the robot device R1 may be, for example, an industrial robot used in a production line, an autonomous robot configured to operate autonomously, or a mobile robot configured to move. The industrial robot may be, for example, a vertical articulated robot, a horizontal articulated robot (SCARA robot), a parallel link robot, or an orthogonal robot. The autonomous robot may be, for example, a humanoid robot, a guide robot, an agricultural robot, a care robot, a security robot, or a transport robot (including a food delivery robot). The mobile robot may be, for example, a cleaning robot, the above-mentioned autonomous robot configured to move (including a mobile robot), a vehicle configured to be autonomously driven, an air vehicle capable of autonomous flight (such as a drone), or a ship capable of autonomous navigation (such as a ship or submarine). The robot device R1 may be operated manually.

[0094] According to the first example, it is possible to reduce the cost of building a classification system when identifying features using tactile data C1 and image data D1. Furthermore, when a large-scale visual language model is used as the large-scale language model 6, it is possible to expect improved classification accuracy by incorporating tactile data C1 (text representation 55) in the classification task in addition to image data D1.

[0095] 7 is a schematic diagram illustrating an example of a second case to which the present embodiment is applied. The second case is an example of a situation in which the present embodiment is applied to identify the characteristics of an object T2 from temperature data C2 and image data D1.

[0096] As shown in FIG. 7 , the first data type 30 may be temperature data C2. A sample of the temperature data C2 (first sample 35) may be generated by a temperature sensor S12. The temperature sensor S12 is an example of the first sensor S1. The temperature sensor S12 may be either a contact type or a non-contact type. As in the first example, the second data type 40 may be image data D1, and a sample of the image data D1 (second sample 45) may be generated by an image sensor S21. The temperature sensor S12 and the image sensor S21 may be appropriately positioned in a location where the object T2 can be observed.

[0097] The identification device 1 may construct a prompt 60 including the command sentence 50, a text representation 55 of the first sample 35 of the temperature data C2, and the second sample 45 of the image data D1. In the second case, the identification device 1 may also construct the prompt 60 to further include a list 59. The identification device 1 may provide the obtained prompt 60 to the large-scale language model 6, thereby obtaining an answer 65 from the large-scale language model 6 indicating the results of identifying the features in the target T2.

[0098] The identification task based on temperature and image may be performed for any purpose. The target T2 may be determined appropriately depending on the embodiment. In one example, the target T2 may be an object of a chemical experiment (e.g., a solution for cell culture, etc.). If the target T2 is a solution for cell culture, the identification task may include identifying a range of solutions suitable for cell culture based on temperature and image. The identification result may be fed back to the cell culture.

[0099] In another example, the target T2 may be a cooking target. The identification task may include identifying the state of the cooking process, such as whether the target T2 is in a suitable state for cooking, from the temperature and the image. When cooking is performed by a robotic device, the identification result may be fed back to the cooking operation (for example, whether to proceed to the next step may be determined depending on the identification result).

[0100] In another example, the target T2 may be a metal (such as an alloy) to be processed. The identification task may include identifying the state of the target T2 during processing from the temperature and the image. If the metal processing is performed by a robotic device, the processing operation of the robotic device may be determined according to the identification result.

[0101] In another example, the target T2 may be a robotic device. The identification task may include identifying the state of the robotic device from temperature and images. The scope of the robotic device may be defined in the same manner as the robotic device R1 described above. The temperature sensor S12 may be located either inside or outside the robotic device. The identification result may be fed back to the operation of the robotic device. For example, if the robotic device is identified as being in a state where a cool-down is recommended, the identification device 1 may issue a command to the robotic device to perform a cool-down. Alternatively, the identification device 1 may provide the identification result to the robotic device, causing the identification device 1 to determine whether or not to perform a cool-down.

[0102] In another example, the object T2 may be food or drink to be delivered by a transport robot. The identification task may include identifying the state of the food or drink from temperature and images. The result of identifying the state of the food or drink may be fed back to the operation of the transport robot. For example, the operation of the transport robot may be determined according to the identification result, such as delivering hot food or drink more slowly (moving at a slower speed) compared to cold food or drink. The identification device 1 may determine an operation according to the identification result and issue a command to the transport robot to execute the determined operation. Alternatively, the identification device 1 may provide the identification result to the transport robot, thereby allowing the transport robot (controller) to determine the operation.

[0103] According to the second example, it is possible to reduce the cost of building a classification system when identifying features using temperature data C2 and image data D1. Furthermore, when a large-scale visual language model is used as the large-scale language model 6, it is possible to expect improvement in classification accuracy by incorporating temperature data C2 (text representation 55) in the classification task in addition to image data D1.

[0104] 8 is a schematic diagram illustrating an example of a third case to which this embodiment is applied. The third case is an example of a situation in which this embodiment is applied to identify the features of an object T3 from sound data C3 and image data D1.

[0105] 8, the first data type 30 may be sound data C3. Samples of the sound data C3 (first samples 35) may be generated by a microphone S13. The microphone S13 is an example of a first sensor S1. As in the first example, the second data type 40 may be image data D1, and samples of the image data D1 (second samples 45) may be generated by an image sensor S21. The microphone S13 and the image sensor S21 may be appropriately positioned in a location where the target T3 can be observed.

[0106] The identification device 1 may construct a prompt 60 using the command sentence 50, a text representation 55 of the first sample 35 of the sound data C3, and the second sample 45 of the image data D1. In the third case, the identification device 1 may also construct the prompt 60 to further include a list 59. The identification device 1 may provide the obtained prompt 60 to the large-scale language model 6, thereby obtaining an answer 65 from the large-scale language model 6 indicating the results of identifying features in the target T3.

[0107] The identification task using sound and images may be performed for any purpose. The object T3 may be determined as appropriate depending on the embodiment. In one example, the identification task using sound and images may be performed for an inspection such as non-destructive testing. The object T3 may be, for example, an equipment to be inspected, such as a road, building, or structure containing concrete. The identification task may include identifying the condition of the equipment to be inspected (object T3). The identification result may be fed back to a user, such as an administrator.

[0108] In another example, the object T3 may be an object of observation by the robotic device. The identification task may include identifying at least one of the attributes and the state of the object of observation by the robotic device from temperature and images. The scope of the robotic device may be defined in the same manner as for the robotic device R1. The identification result may be fed back to the operation of the robotic device. For example, if the robotic device is a guide robot, identifying the state of the object of observation may include identifying whether the subject is requesting guidance. The subject is an example of the object T3. If the subject is identified as being in a state of requesting guidance, the identification device 1 may move near the subject and issue a command to the guide robot to provide guidance to the subject. Alternatively, the identification device 1 may provide the identification result to the guide robot, causing the guide robot to perform an action related to guidance.

[0109] According to the third example, it is possible to reduce the cost of building a classification system when identifying features using sound data C3 and image data D1. Furthermore, when a large-scale visual language model is used as the large-scale language model 6, it is possible to expect an improvement in classification accuracy by incorporating sound data C3 (text representation 55) in the classification task in addition to image data D1.

[0110] 9 is a schematic diagram illustrating an example of a fourth case to which the present embodiment is applied. The fourth case is an example of a situation in which the present embodiment is applied to identify features of an object T4 from point cloud data C4 and image data D1.

[0111] 9 , the first data type 30 may be point cloud data C4. A sample (first sample 35) of the point cloud data C4 may be generated by a point cloud sensor S14. The point cloud sensor S14 is an example of the first sensor S1. As in the first example, the second data type 40 may be image data D1, and a sample (second sample 45) of the image data D1 may be generated by an image sensor S21. The point cloud sensor S14 and the image sensor S21 may be appropriately positioned in a location where the object T4 can be observed.

[0112] The identification device 1 may construct a prompt 60 using the command sentence 50, a text representation 55 of the first sample 35 of the point cloud data C4, and the second sample 45 of the image data D1. In the fourth example, the identification device 1 may also construct the prompt 60 to further include a list 59. The identification device 1 may provide the obtained prompt 60 to the large-scale language model 6, thereby obtaining an answer 65 from the large-scale language model 6 indicating the results of identifying features in the object T4.

[0113] The identification task using the point cloud and image may be performed for any purpose. The target T4 may be determined appropriately depending on the embodiment. In one example, the target T4 may be an observation target of the robotic device. The identification task may include identifying at least one of the attributes and the state of the observation target of the robotic device from the point cloud and image. The scope of the robotic device may be defined similarly to the robotic device R1 described above. The identification result may be fed back to the operation of the robotic device. For example, if the robotic device is a mobile robot, the identification device 1 may determine a movement path for the mobile robot based on the identification result. For example, if it is identified that the observation target is entering or has the potential to enter the direction of travel of the mobile robot, the movement path may be determined to avoid the observation target. The identification device 1 may instruct the mobile robot to move along the determined path. Alternatively, the identification device 1 may instruct the mobile robot to move along the determined path by providing the identification result to the mobile robot, causing the mobile robot to determine a movement path based on the identification result.

[0114] According to the fourth example, it is possible to reduce the cost of building a classification system when identifying features using point cloud data C4 and image data D1. Furthermore, when a large-scale visual language model is used as the large-scale language model 6, it is possible to expect improvement in classification accuracy by incorporating point cloud data C4 (text representation 55) in the classification task in addition to image data D1.

[0115] 10 is a schematic diagram illustrating an example of a fifth case to which the present embodiment is applied. The fifth case is an example of a situation in which the present embodiment is applied to identify the features of an object T5 from weight data C5 and image data D1.

[0116] 10 , the first data type 30 may be weight data C5. A sample of the weight data C5 (first sample 35) may be generated by a load sensor S15. The load sensor S15 is an example of the first sensor S1. As in the first example, the second data type 40 may be image data D1, and a sample of the image data D1 (second sample 45) may be generated by an image sensor S21. The load sensor S15 and the image sensor S21 may be appropriately positioned in a location where the target T5 can be observed.

[0117] The identification device 1 may construct a prompt 60 including a command sentence 50, a text representation 55 of the first sample 35 of the weight data C5, and a second sample 45 of the image data D1. In the fifth case, the identification device 1 may also construct the prompt 60 to further include a list 59. The identification device 1 may provide the obtained prompt 60 to the large-scale language model 6, thereby obtaining an answer 65 from the large-scale language model 6 indicating the results of identifying features in the object T5.

[0118] The identification task using weight and image may be performed for any purpose. The object T5 may be determined appropriately depending on the embodiment. In one example, the object T5 may be an item to be inspected. The item to be inspected may be, for example, a product on a production line (final product, intermediate product, etc.), agricultural produce, etc. The identification task may include identifying the condition of the item to be inspected (object T5) from its weight and image. Identifying the condition of the item to be inspected may include identifying whether the item is defective (whether it is broken, whether it meets standards, etc.). The identification result may be fed back to a user such as an administrator.

[0119] According to the fifth example, when identifying features using weight data C5 and image data D1, it is possible to reduce the cost of building a classification system. Furthermore, when using a large-scale visual language model as the large-scale language model 6, it is possible to expect improvement in classification accuracy by incorporating weight data C5 (text representation 55) in the classification task in addition to image data D1.

[0120] 11 is a schematic diagram illustrating an example of a sixth case to which the present embodiment is applied. The sixth case is an example of a situation in which the present embodiment is applied to identify the characteristics of an object T6 from odor data C6 and image data D1.

[0121] 11 , the first data type 30 may be odor data C6. A sample (first sample 35) of the odor data C6 may be generated by an odor sensor S16. The odor sensor S16 is an example of the first sensor S1. As in the first example, the second data type 40 may be image data D1, and a sample (second sample 45) of the image data D1 may be generated by an image sensor S21. The odor sensor S16 and the image sensor S21 may be appropriately positioned in a location where the target T6 can be observed.

[0122] The identification device 1 may construct a prompt 60 using the command sentence 50, a text representation 55 of the first sample 35 of the odor data C6, and the second sample 45 of the image data D1. In the sixth case, the identification device 1 may also construct the prompt 60 to further include a list 59. The identification device 1 may provide the obtained prompt 60 to the large-scale language model 6, thereby obtaining an answer 65 from the large-scale language model 6 indicating the results of identifying the features in the target T6.

[0123] The identification task using smell and images may be performed for any purpose. The object T6 may be determined as appropriate depending on the embodiment. In one example, the object T6 may be an object to be inspected. The object to be inspected may include, for example, water in a facility, oil in a fryer, food and drink, etc. The facility may include water treatment equipment such as a water purification plant or a sewage treatment plant. The facility may also include a natural or artificial body of water such as a pond or lake. The identification task may include identifying the condition of the object (object T6) from smell and images. Identifying the condition of the object may include identifying the water quality in a facility, identifying the quality of oil in a fryer, identifying the quality of food and drink (e.g., whether it is fresh), etc. The identification result may be fed back to a user such as an administrator.

[0124] For example, if the target T6 is water from a sewage treatment plant and the water quality is identified as poor, the identification device 1 may generate an operation plan to increase the operation rate of the sewage treatment plant. If the water quality is identified as meeting the standard, the identification device 1 may generate an operation plan to maintain or decrease the operation rate of the sewage treatment plant. The identification device 1 may output instructions to a user or an equipment controller to operate the sewage treatment plant in accordance with the generated operation plan. Alternatively, the identification device 1 may provide the identification result to an external computer, causing the external computer to generate an operation plan. In response to this, the external computer may output instructions.

[0125] Furthermore, for example, if the target T6 is oil in a fryer, the identification device 1 may provide the results of identifying the quality of the oil to an oil manager. The quality of the oil may be defined appropriately according to any standard, including publicly known standards. If deterioration in the quality of the oil is identified, the identification device 1 may output a notification to the manager's terminal recommending that the oil be replaced.

[0126] According to the sixth example, when identifying features using scent data C6 and image data D1, it is possible to reduce the cost of building a recognition system. Furthermore, when using a large-scale visual language model as the large-scale language model 6, it is possible to expect improved recognition accuracy by incorporating scent data C6 (text representation 55) in the recognition task in addition to image data D1.

[0127] 12 is a schematic diagram illustrating an example of a seventh case to which the present embodiment is applied. The seventh case is an example of a situation in which the present embodiment is applied to identify the characteristics of an object T7 from inertial characteristic data C7 and image data D1.

[0128] 12 , the first data type 30 may be inertial characteristic data C7. The inertial characteristics may include, for example, center of gravity position, mass, moment of inertia, etc. A sample (first sample 35) of the inertial characteristic data C7 may be generated by an inertial characteristic measurement device S17. The inertial characteristic measurement device S17 is an example of a first sensor S1. As in the first example, the second data type 40 may be image data D1, and a sample (second sample 45) of the image data D1 may be generated by an image sensor S21. The inertial characteristic measurement device S17 and the image sensor S21 may be appropriately positioned in a location where the target T7 can be observed.

[0129] The identification device 1 may construct a prompt 60 including a command sentence 50, a text representation 55 of the first sample 35 of the inertial characteristic data C7, and a second sample 45 of the image data D1. In the seventh case, the identification device 1 may also construct the prompt 60 to further include a list 59. The identification device 1 may provide the obtained prompt 60 to the large-scale language model 6, thereby obtaining an answer 65 from the large-scale language model 6 indicating the results of identifying the features in the target T7.

[0130] The identification task using inertial characteristics and images may be performed for any purpose. The object T7 may be determined as appropriate depending on the embodiment. The identification task may include identifying at least one of the attributes and state of the object T7 from the inertial characteristics and images. For example, if the object T7 is a drive target of a robotic device, identifying at least one of the attributes and state of the object T7 may include identifying an action range suitable for driving the drive target. The identification device 1 may feed back the identification result to the operation of the robotic device so that the drive target is driven within the identified action range. Furthermore, there are combinations of items that are difficult to distinguish based on appearance alone, such as a combination of a raw egg and a boiled egg, but are easily distinguishable when inertial characteristics are taken into consideration. Therefore, identifying at least one of the attributes and state of the object T7 may include identifying the type of object T7 (e.g., an object). The identification result may be fed back to a user or to the operation of the robotic device.

[0131] According to the seventh example, it is possible to reduce the cost of building a classification system when identifying features using inertial characteristic data C7 and image data D1. Furthermore, when using a large-scale visual language model as the large-scale language model 6, it is possible to expect improvement in classification accuracy by reflecting inertial characteristic data C7 (text representation 55) in addition to image data D1 in the classification task.

[0132] (Other) The above examples may be modified as appropriate depending on the embodiment. For example, in the first to seventh examples, the image data D1 (second data type 40) may be omitted. In each example except the third example, sound data D2 may be used as the second data type 40 instead of the image data D1, and a microphone may be used as the second sensor S2. In each of the above examples, at least one of the category classification method of FIG. 3 and one form of the command statement 50 of FIG. 5 may be used.

[0133] Furthermore, the application of this embodiment is not limited to the above examples. In another example, the first data type 30 may be electromagnetic data, and the second data type 40 may be image data. The identification task may include identifying the electrical properties of the object (e.g., whether it is an insulator or a conductor). The identification result of the electrical properties may be fed back to the operation of the robotic device, for example.

[0134] In one example of this embodiment, including the above cases, the first sample 35 and the second sample 45 may be measured immediately after an intervention on the subject. The intervention may include, for example, physical intervention or electrical intervention. Physical intervention may include, for example, movement (shaking, rotating, etc.), application of sound waves, etc. Electromagnetic intervention may include, for example, bringing a magnet close, passing an electric current, etc. This allows the identification task to identify the effect of the intervention on the subject.

[0135] §2 Configuration Example [Hardware Configuration] Fig. 13 shows a schematic example of the hardware configuration of the identification device 1 according to this embodiment. The identification device 1 according to this embodiment is a computer to which a control unit 11, a storage unit 12, an external interface 13, an input device 14, an output device 15, and a drive 16 are electrically connected.

[0136] The control unit 11 includes a hardware processor such as a CPU (Central Processing Unit), RAM (Random Access Memory), and ROM (Read Only Memory), and is configured to execute information processing based on programs and various data. The control unit 11 (CPU) is an example of a processor resource. The storage unit 12 may be configured, for example, with a hard disk drive or a solid state drive. The storage unit 12, RAM, and ROM are examples of memory resources. In this embodiment, the storage unit 12 stores various information such as the identification program 81, model data 600, and encoder data EN0.

[0137] The identification program 81 is a program for causing the identification device 1 to execute information processing (see FIG. 15 described below) related to the performance of an identification task. The identification program 81 includes a series of instructions for the information processing. The model data 600 is configured to indicate information related to the large-scale language model 6. The encoder data EN0 is configured to indicate information related to the trained encoder EN. If the large-scale language model 6 is deployed on an external computer, the model data 600 may be omitted. If category classification is performed on an external computer or if the trained encoder EN is not used for category classification, the encoder data EN0 may be omitted.

[0138] As long as the model data 600 can hold information for executing the calculation process of the large-scale language model 6, the configuration of the model data 600 is not particularly limited and may be determined appropriately depending on the embodiment. For example, the model data 600 may be configured to include information indicating values ​​of calculation parameters of the large-scale language model 6 adjusted by machine learning. The model data 600 may also be configured to further include information indicating the configuration of the large-scale language model 6 (e.g., the structure of a neural network). The same applies to the encoder data EN0. For example, at least one of the model data 600 and the encoder data EN0 may be incorporated into the identification program 81.

[0139] The external interface 13 is configured to connect to an external device via a wired or wireless connection. The external interface 13 may be, for example, a Universal Serial Bus (USB) port, a dedicated port, a communication port, or the like. When the external interface 13 includes a communication port, the communication standard of the communication port may be selected arbitrarily. In this embodiment, the identification device 1 may be connected to an external device (e.g., a first sensor S1, a second sensor S2, an external computer, or the like) via the external interface 13.

[0140] The input device 14 is a device for inputting, for example, a mouse, a keyboard, etc. The output device 15 is a device for outputting, for example, a display, a speaker, etc. An operator can operate the identification device 1 by using the input device 14 and the output device 15. The input device 14 and the output device 15 may be connected via an external interface 13. The input device 14 and the output device 15 may be integrated into one device, for example, a touch panel display, etc.

[0141] The drive 16 is a device for reading various information, such as programs, stored in a storage medium 91. At least one of the identification program 81, the model data 600, and the encoder data EN0 may be stored in the storage medium 91 instead of or together with the storage unit 12. The storage medium 91 is configured to store various information (such as stored programs) by electrical, magnetic, optical, mechanical, or chemical action so that a machine such as a computer can read the information. The storage unit 12 and the storage medium 91 are examples of non-transitory storage media. The identification device 1 may acquire at least one of the identification program 81, the model data 600, and the encoder data EN0 from the storage medium 91. The storage medium 91 may be a disk-type storage medium such as a CD or DVD, or a non-disk-type storage medium such as a semiconductor memory (e.g., a flash memory). The type of the drive 16 may be selected appropriately depending on the type of the storage medium 91. The drive 16 may be connected via an external interface.

[0142] Note that, with regard to the specific hardware configuration of the identification device 1, components may be omitted, replaced, or added as appropriate depending on the embodiment. For example, the control unit 11 may include multiple hardware processors. The hardware processor may be configured with a microprocessor, a field-programmable gate array (FPGA), a digital signal processor (DSP), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), or the like. At least one of the external interface 13, the input device 14, the output device 15, and the drive 16 may be omitted. The identification device 1 may be configured with multiple computers. In this case, the hardware configurations of the computers may or may not be the same. Furthermore, the identification device 1 may be an information processing device designed specifically for the service provided, as well as a general-purpose server device, a general-purpose personal computer (PC), a tablet PC, a terminal device (smartphone, etc.), or the like.

[0143] At least one of the identification program 81, the model data 600, and the encoder data EN0 may be stored in an external storage device such as a network-attached storage (NAS). A portion of the prompt 60, such as the command statement 50 and at least a portion of the list 59, may be provided by a template. The template portion of the prompt 60 may be stored in at least one of the storage unit 12 and the storage medium 91, stored in an external storage device, provided by an operator, or incorporated into the identification program 81. The identification device 1 may generate a portion of the prompt 60 by reading the template data. Furthermore, when using the encoder EN to classify the categories of the first samples 35, the reference value R0 of each training sample TR and the text expression associated with each category may be stored in at least one of the storage unit 12 and the storage medium 91, stored in an external storage device, or incorporated into the identification program 81.

[0144] 14 schematically shows an example of the software configuration of the identification device 1 according to this embodiment. The control unit 11 of the identification device 1 executes instructions included in the identification program 81 stored in the storage unit 12 using the CPU. As a result, the identification device 1 operates as a computer including an acquisition unit 111, an identification unit 112, and an output processing unit 113 as software modules. That is, in this embodiment, each software module of the identification device 1 is realized by the control unit 11 (CPU).

[0145] The acquisition unit 111 is configured to acquire a text representation 55 of the first sample 35 according to a category to which the first sample 35 belongs. In one example, the acquisition unit 111 may further be configured to acquire the first sample 35 and determine a category to which the acquired first sample 35 belongs. Determining the category to which the first sample 35 belongs may include converting the first sample 35 into a feature value F1 using the trained encoder EN, and extracting one or more candidate categories to which the first sample 35 belongs from among multiple categories assigned to the training samples TR by comparing the feature value F1 obtained from the first sample 35 with a reference value R0 (the value of the feature F0) obtained from each training sample TR used in the machine learning. The text representation 55 of the first sample 35 may include text representations of the one or more extracted categories. In another example, the acquisition unit 111 may further be configured to acquire a second sample 45 of the second data type 40.

[0146] The identification unit 112 is configured to obtain an answer 65 indicating a result of the feature identification from the large-scale language model 6 by providing the large-scale language model 6 with a prompt 60 including a command statement 50 for instructing the large-scale language model 6 to identify features in the target and a text representation 55 of the first sample 35. In one example, the prompt 60 may further include a list of labels 59. In one example, the large-scale language model 6 may be configured to accept input of a sample of the second data type 40, and the prompt 60 may further include the second sample 45. In another example, the command statement 50 may include a first sub-sentence 501 instructing the large-scale language model 6 to derive a first interim result of the feature identification from the text representation 55 of the first sample 35, a second sub-sentence 502 instructing the large-scale language model 6 to derive a second interim result of the feature identification from the second sample 45, and a third sub-sentence 503 instructing the large-scale language model 6 to derive a result of the feature identification from the first interim result and the second interim result. The output processing unit 113 is configured to output information regarding the obtained result.

[0147] In this embodiment, an example is described in which each software module of the identification device 1 is implemented by a general-purpose CPU. However, some or all of the software modules may be implemented by one or more dedicated processors or chipsets. Each module may also be implemented as a hardware module. Regarding the software configuration of the identification device 1, modules may be omitted, replaced, or added as appropriate depending on the embodiment.

[0148] §3 Operational Example Fig. 15 is a flowchart showing an example of the processing procedure of the identification device 1 according to this embodiment. The example of Fig. 15 assumes a situation in which the configuration of Fig. 4 above is adopted. The following processing procedure is an example of an identification method executed by a computer. However, the following processing procedure is merely an example, and each step may be modified as much as possible. Furthermore, steps in the following processing procedure may be omitted, replaced, or added as appropriate depending on the embodiment.

[0149] (Step S101) In step S101, the control unit 11 operates as the acquisition unit 111 and acquires the first samples 35 of the first data type 30 and the second samples 45 of the second data type 40.

[0150] The method of acquiring each sample (35, 45) is not particularly limited and may be selected appropriately depending on the embodiment. In one example, the control unit 11 may directly or indirectly acquire the first sample 35 from the first sensor S1. The control unit 11 may directly or indirectly acquire the second sample 45 from the second sensor S2. In one example, any of the first to seventh cases described above may be adopted for the combination of the first data type 30 and the second data type 40. After acquiring the first sample 35 and the second sample 45, the control unit 11 proceeds to the next step S102.

[0151] (Step S102) In step S102, the control unit 11 operates as the acquisition unit 111 and determines the category to which the first sample 35 belongs.

[0152] As described above, the method for determining the category to which the first sample 35 belongs (classifying the category of the first sample 35) may be selected appropriately depending on the embodiment. In one example, a trained encoder EN may be generated in advance by machine learning, and a reference value R0 may be generated from each training sample TR using the generated trained encoder EN. Furthermore, each category may be associated with a text expression. The control unit 11 may convert the first sample 35 into a feature (value F1) using the trained encoder EN generated by machine learning. The control unit 11 may compare the feature value F1 obtained from the first sample 35 with the reference value R0 obtained from each training sample TR. In one example, the control unit 11 may search for a neighborhood of the value F1 within a database of reference values ​​R0. The distance serving as a criterion for the neighborhood search may be defined arbitrarily. For example, a known definition such as the L1 norm or the L2 norm may be used for the distance serving as a criterion for the neighborhood search. Depending on the result of this comparison, the control unit 11 may extract one or more categories that are candidates to which the first sample 35 belongs from among the multiple categories assigned to each training sample TR. The extracted one or more categories correspond to the category determination result. Upon obtaining the category determination result, the control unit 11 proceeds to the next step S103.

[0153] (Step S103) In step S103, the control unit 11 operates as the acquisition unit 111 and acquires the text expression 55 of the first sample 35 according to the category to which the first sample 35 belongs.

[0154] In one example, the control unit 11 may acquire a text representation 55 of the first sample 35 according to the determination result of the category to which the first sample 35 belongs. Acquiring the text representation 55 of the first sample 35 may be configured by acquiring, as the text representation 55 of the first sample 35, a text representation associated with each of the one or more categories extracted as the determination result. In other words, the text representation 55 of the first sample may include a text representation of the one or more extracted categories. After acquiring the text representation 55, the control unit 11 proceeds to the next step S104.

[0155] (Step S104) In step S104, the control unit 11 generates a prompt 60 using the command statement 50, the text representation 55 of the acquired first sample 35, and the second sample 45. The command statement 50 may be acquired by any method, such as by being provided by a template.

[0156] In one example, the control unit 11 may generate the prompt 60 so that it further includes a list of labels 59. In another example, the command statement 50 may be configured to include a first partial sentence 501, a second partial sentence 502, and a third partial sentence 503. After generating the prompt 60, the control unit 11 proceeds to the next step S105.

[0157] (Step S105) In step S105, the control unit 11 provides the generated prompt 60 to the large-scale language model 6, thereby obtaining an answer 65 from the large-scale language model 6 that indicates the result of identifying the features.

[0158] In one example, the identification device 1 may store model data 600, thereby providing a large-scale language model 6. The control unit 11 may provide a prompt 60 to the large-scale language model 6 and execute calculation processing on the large-scale language model 6, thereby obtaining an answer 65 from the large-scale language model 6. In another example, the large-scale language model 6 may be deployed on an external computer. In response, the control unit 11 may transmit the prompt 60 and a request to execute calculation to the external computer. The external computer may provide the prompt 60 received from the identification device 1 to the large-scale language model 6 and execute calculation processing on the large-scale language model 6. The external computer may return an answer 65 obtained as a result of this calculation processing to the identification device 1. The control unit 11 may obtain the answer 65 by receiving this reply. In one example, by adopting any of the first to seventh cases as the combination of the first data type 30 and the second data type 40, the control unit 11 may obtain an answer 65 indicating an identification result corresponding to any of the first to seventh cases. When the response 65 is acquired, the control unit 11 advances the process to the next step S106.

[0159] (Step S106) In step S106, the control unit 11 outputs information related to the obtained result (answer 65).

[0160] The output destination and the content of the output information may be selected appropriately depending on the embodiment. In one example, the control unit 11 may directly output the acquired identification result as the response 65, such as in the case of providing feedback to the administrator. In another example, the control unit 11 may perform arbitrary information processing in accordance with the acquired identification result. The control unit 11 may output the result of the information processing as information related to the identification result. The output of the result of the information processing may include, for example, outputting a specific message in accordance with the identification result, or controlling the operation of the controlled device in accordance with the identification result. Controlling the operation of the controlled device may include, for example, feeding back the identification result to the operation of the robot device R1. The output destination may be, for example, RAM, the memory unit 12, the output device 15, another computer, the controlled device, etc.

[0161] When the output of information is completed, the control unit 11 ends the processing procedure of the identification device 1 according to this operation example. The control unit 11 may execute the series of processes from step S101 to step S106 at any timing, such as a user operation or when a condition is satisfied. The control unit 11 may execute the series of processes from step S101 to step S106 in real time, or may execute them as a post-event matching process. The control unit 11 may execute the series of processes from step S101 to step S106 each time at least one of the first sample 35 and the second sample 45 is updated.

[0162] [Features] In this embodiment, in the processes of steps S101 to S103, the first sample 35 is converted into a text representation 55 for input to the large-scale language model 6. In step S104, the text representation 55 corresponding to the category of the first sample 35 is provided to the large-scale language model 6 as a prompt 60. In step S105, when generating an answer 65 corresponding to the command statement 50, the large-scale language model 6 can derive a feature identification result from the category of the first sample 35 using acquired common sense. Therefore, according to this embodiment, by converting the first sample 35 of the first data type 30 into the text representation 55, zero-shot identification of features in a target can be achieved without fine-tuning the large-scale language model 6. This makes it possible to reduce the cost required to build an identification system for any modality.

[0163] §4 Modifications Although the embodiments of the present invention have been described in detail above, the above description is merely an example of the present invention in every respect. The processes and means described in this disclosure can be freely combined and implemented as long as no technical contradiction occurs. Furthermore, various improvements or modifications may be made to the above embodiments as appropriate.

[0164] <4.1> Steps may be omitted, replaced, or added in the processing procedure of the identification device 1 according to the above embodiment. For example, the timing of acquiring the first sample 35 and the second sample 45 is not limited to the example of FIG. 15 and may be changed as desired. The second sample 45 may be acquired at any timing before step S104. The process of acquiring the second sample 45 and the process of acquiring the text representation 55 of the first sample 35 may or may not overlap at least partially. The process of acquiring the second sample 45 may be executed before or after the process of acquiring the text representation 55 of the first sample 35.

[0165] Furthermore, for example, if the large-scale language model 6 is not configured to accept input of samples of the second data type 40, the processing related to the second sample 45 may be omitted. Furthermore, for example, the processing of determining the category to which the first sample 35 belongs and determining the text expression 55 may be executed by an external computer other than the identification device 1. In this case, the processing related to the first sample 35 in steps S101 to S103 may be replaced by directly or indirectly acquiring, from the external computer, the text expression 55 corresponding to the category to which the first sample 35 belongs. Alternatively, the control unit 11 may acquire, from the external computer, a determination result of the category to which the first sample 35 belongs, and acquire the text expression 55 according to the acquired determination result.

[0166] §5 Examples The following experiments were carried out to verify the effectiveness of the above-described embodiment, but the present invention is not limited to the following examples.

[0167] (Example) Fig. 16 shows the configuration of a classification system in an example. As shown in Fig. 16, tactile data (tactile sequence) was used as the first data type, and image data was used as the second data type. The classification task was to identify the type of object appearing in each sample of tactile data and image data. The method of classifying the category of the first sample was to use the autoencoder shown in Fig. 3 above. The autoencoder was the VQ-VAE proposed in Non-Patent Document 6. A two-layer bidirectional GRU (Gated Recurrent Unit) with a hidden size of 32 was used for the encoder and decoder. The number of dimensions of the embedded representation in vector quantization (VQ) was set to 4, and the number of vectors in the codebook was set to 32. Z in Fig. 16 indicates a feature, and Z q indicates the feature quantity discretized by vector quantization. q ) was input to a linear layer with an output size of 32 and set as the initial hidden state of the decoder. The loss function is the reconstruction error (L recon ), vector quantization error (L vq ), and the commitment error (L commit ) was composed of:

[0168] The hyperparameter β was set to 0.25. A trained encoder was obtained by performing machine learning of VQ-VAE using the training dataset described below. Then, a reference value was calculated from each training sample used in training and validation of VQ-VAE using the trained encoder, and a database was constructed by associating each calculated reference value with the category name of the corresponding training sample. In other words, the category name was used as is for the text representation of each category.

[0169] The inference stage was configured to use a trained encoder to calculate feature values ​​for the tactile data and then search for neighborhoods of the feature values ​​calculated using the L2 norm within a database of reference values. The neighborhood search was configured to extract k categories by sequentially selecting the reference values ​​closest to the feature values ​​of the tactile data, where k was set to 3. The names of each extracted category were used as text representations of the tactile data, and a prompt was constructed from a list of commands, text representations of the tactile data, image data, and labels.

[0170] Figure 17 shows the prompt (excluding image data) in the example. The command structure used was the same as that shown in Figure 5. The text representation of the tactile data (tactile information) was composed of text indicating similarity to the type of category that fits. The names of the three categories extracted by the proximity search were entered in {topk_refs}. The label list {test_time_classes} was composed of the labels (category names) of the dataset used for testing (described below). The "Output Format" description was added to the prompt to specify the output format of the large-scale language model's response. GPT-4V, proposed in Non-Patent Document 4, was used as the large-scale language model. The temperature of GPT-4V was set to 0.0. The above steps constituted the classification system of the example.

[0171] (Reference Example) Figure 18 shows a prompt (excluding image data) in a reference example. In the reference example, the components related to tactile data in the above-described embodiment are omitted. In other words, the reference example is configured to obtain the result of identifying the type of object from the large-scale language model by providing the large-scale language model with a prompt consisting of a command sentence, image data, and a list of labels, with the description related to tactile data omitted. In all other respects, the configuration of the reference example is set to be the same as that of the embodiment.

[0172] (Dataset) (A) Training Dataset Fig. 19 shows the items used to collect the training dataset used in the experiment. First, to construct the reference value database of the example, 32 items shown in Fig. 19 were prepared. The name of each item was set as the name of a category. Of the 32 categories, 27 categories were used for training, and 5 categories were used for validation.

[0173] A commercially available robot arm equipped with a parallel gripper was also prepared. A distributed triaxial tactile sensor was attached to each finger of the parallel gripper. The tactile sensor was configured to obtain 4 × 4 × 3 × 2 = 96-dimensional tactile signals per time step. The gripper was positioned so that the gripper fingers were positioned on both sides of each object. Next, the gripper was driven to close, and the gripper movement was stopped when the tactile sensor value exceeded a threshold. The threshold was set to 0.004, and the stopping time was set to 0.2 seconds. The gripper was then driven to open. A total of 1 second of sequence was recorded as one sample, from 0.4 seconds before the gripper was stopped until the gripper was stopped, the 0.2 seconds while the gripper was stopped, and the 0.4 seconds after the gripper was opened. The sampling frequency was set to 125 Hz. Next, the sequence was subsampled at 25 Hz to expand one sample to five samples. The above process was repeated 10 times for each object. This resulted in 5 x 10 = 50 training samples per category for the tactile data.

[0174] (B) First Dataset Figure 20 shows the food (real objects) and food replicas used to collect the first dataset used in the experiment. The real food and replicas are a combination of items that are difficult to distinguish from each other based on appearance. To evaluate the identification accuracy of the Examples and Reference Examples using this combination, 11 real food items and 11 replicas of the food shown in Figure 20 were prepared (22 categories in total). By omitting subsampling and observing under the same conditions as the training dataset, 10 samples of tactile data were obtained for each category. By photographing each item with a commercially available smartphone, two samples of image data were obtained for each category. By assigning one sample of image data to every five samples of tactile data, 10 sample pairs of tactile data and image data were obtained for each category.

[0175] (C) Second Dataset Figure 21 shows the food (actual objects) used to collect the second dataset used in the experiment. For some foods, it can be difficult to distinguish between boiled and raw states. To evaluate the discrimination accuracy of the examples and reference examples in each state, four foods were prepared: boiled pumpkin, raw pumpkin, boiled rice cakes, and raw rice cakes (a total of four categories / two types of food) as shown in Figure 21. Haptic data and image data samples were obtained under the same conditions as the first dataset. That is, for tactile data, 10 samples were obtained per category. For image data, two samples were obtained per category. By assigning one image data sample to every five tactile data samples, 10 sample pairs of tactile data and image data were obtained per category.

[0176] (Experimental Conditions) For the example, the autoencoder of the example was implemented on a commercially available computer and trained using training samples (50 × 27 = 1,350 samples) of 27 of the 32 categories in the training dataset. The Adam optimizer with a learning rate of 0.01 was used for training. The batch size was set to 256. Training samples (50 × 5 = 250 samples) of five categories were used to monitor the convergence of the training. This machine learning process resulted in a trained encoder. Then, using the trained encoder, all training samples (50 × 32 = 1,600 samples) were converted into reference values, and a database was constructed using the reference values.

[0177] Next, prompts were generated from each sample of tactile data and image data for each category for each of the first and second datasets, and the generated prompts were fed into a large-scale language model to obtain results for identifying the type of each item. The resulting identification results were compared with the labels (categories) of each item to determine whether the identification was successful. This series of processes was performed for each sample pair and category, resulting in 10 results of identification success or failure for each category. The identification accuracy was calculated for each category based on the obtained results, and the obtained identification accuracy was visualized using a confusion matrix.

[0178] In addition, for the reference example, prompts were generated from image data samples of each category for each of the first and second datasets, and the generated prompts were provided to a large-scale language model to obtain answers indicating the results of identifying the type of each item. In this way, in the reference example, the classification task was performed while omitting processing related to tactile data. Considering that the large-scale language model does not necessarily output the same answer each time, in the reference example, the classification task was repeated five times for one sample, obtaining 10 classification success / failure results for each category. Other than these conditions, the classification accuracy in the reference example was calculated for each category under the same conditions as in the example, and the obtained classification accuracy was visualized using a confusion matrix.

[0179] (Results) Figures 22 and 23 show the calculation results of the classification accuracy of the first dataset by the reference example and the working example (Figure 22 is the reference example, and Figure 23 is the working example). Figures 24 and 25 show the calculation results of the classification accuracy of the second dataset by the reference example and the working example (Figure 24 is the reference example, and Figure 25 is the working example). In each figure, the vertical axis represents the true value, and the horizontal axis represents the classification result. As shown in each figure, the classification accuracy of the working example exceeded the classification accuracy of the reference example for both the first dataset and the second dataset. This result shows that zero-shot classification is possible by converting data not handled by a large-scale language model into a text representation and providing the obtained text representation to the large-scale language model. Therefore, it was found that the above embodiment makes it possible to build an effective classification system without fine-tuning.

[0180] This specification includes the following disclosure: [Supplementary Note 1] An identification device (1) including a control unit (11) configured to: acquire a text representation (55) of a first sample (35) of a first data type (30) generated by observing an object with a first sensor (S1) according to a category to which the first sample (35) belongs; provide a prompt (60) to a large-scale language model (6) including a command statement (50) for instructing the large-scale language model to identify a feature of the object and the text representation (55) of the first sample (35); and acquire an answer (65) indicating a result of identifying the feature from the large-scale language model (6); and output information about the acquired result. [Supplementary Note 2] The prompt (60) further includes a list (59) of candidate labels for the feature identification result, and identifying the feature comprises selecting a label corresponding to the feature from the list (59). [Supplementary Note 3] The identification device (1) according to Supplementary Note 1 or Supplementary Note 2, wherein the control unit (11) is further configured to acquire second samples (45) of a second data type (40) different from the first data type (30), the second samples (45) being generated by observing an object with a second sensor, the large-scale language model (6) is configured to accept input of the samples of the second data type (40), and the prompt (60) further includes the second samples (45). [Supplementary Note 4] The identification device (1) according to Supplementary Note 3, wherein the second data type (40) is image data (D1). [Supplementary Note 5] The identification device (1) according to Supplementary Note 4, wherein the first data type (30) is tactile data (C1). [Supplementary Note 6] The identification device (1) according to Supplementary Note 4, wherein the first data type (30) is temperature data (C2). [Supplementary Note 7] The identification device (1) according to Supplementary Note 4, wherein the first data type (30) is sound data (C3). [Supplementary Note 8] The identification device (1) according to Supplementary Note 4, wherein the first data type (30) is point cloud data (C4). [Supplementary Note 9] The identification device (1) according to Supplementary Note 4, wherein the first data type (30) is weight data (C5).[Supplementary Note 10] The identification device (1) according to Supplementary Note 4, wherein the first data type (30) is odor data (C6). [Supplementary Note 11] The identification device (1) according to Supplementary Note 4, wherein the first data type (30) is inertial property data (C7). [Supplementary Note 12] The identification device (1) according to any one of Supplements 3 to 11, wherein the command statement (50) includes: a first sub-sentence (501) instructing to derive a first interim result of identifying the feature from the text representation (55) of the first sample (35), a second sub-sentence (502) instructing to derive a second interim result of identifying the feature from the second sample (45), and a third sub-sentence (503) instructing to derive a result of identifying the feature according to the first interim result and the second interim result. and determining a category to which the acquired first sample belongs. The determining the category to which the first sample belongs includes: converting the first sample into a feature value (F1) using a trained encoder (EN) generated by machine learning; and extracting one or more categories to which the first sample belongs from among a plurality of categories assigned to training samples (TR) by comparing the feature value (F1) obtained from the first sample with a reference value (R0) obtained by converting each training sample (TR) used in the machine learning into the feature (F0) using the trained encoder (EN). The text representation (55) of the first sample includes a text representation of the one or more extracted categories.[Supplementary Note 14] An identification program (81) for causing a computer (1) to execute an identification method, the identification method comprising: a step of acquiring a text representation (55) of a first sample (35) of a first data type (30) generated by observing an object with a first sensor (S1) according to a category to which the first sample (35) belongs; a step of acquiring an answer (65) indicating a result of identifying the feature from a large-scale language model (6) by providing a command statement (50) for instructing the large-scale language model to identify a feature of the object and a prompt (60) including the text representation (55) of the first sample (35) to the large-scale language model (6); and a step of outputting information about the acquired result. [Supplementary Note 15] The identification program (81) according to Supplementary Note 14, wherein the prompt (60) further includes a list (59) of candidate labels for the feature identification result, and identifying the feature comprises selecting a label corresponding to the feature from the list (59). [Supplementary Note 16] The identification program (81) of Supplementary Note 14 or Supplementary Note 15, wherein the identification method further comprises the step of obtaining second samples (45) of a second data type (40) different from the first data type (30), the second samples (45) being generated by observing an object with a second sensor, the large-scale language model (6) being configured to accept input of the samples of the second data type (40), and the prompt (60) further comprising the second samples (45). [Supplementary Note 17] The identification program (81) according to Supplementary Note 16, wherein the instruction statement (50) includes: a first sub-statement (501) instructing to derive a first interim result of identifying the feature from the text representation (55) of the first sample (35); a second sub-statement (502) instructing to derive a second interim result of identifying the feature from the second sample (45); and a third sub-statement (503) instructing to derive a result of identifying the feature depending on the first interim result and the second interim result.[Supplementary Note 18] An identification method executed by a computer (1), comprising: a step of acquiring a text representation (55) of a first sample (35) of a first data type (30) generated by observing an object with a first sensor (S1) according to a category to which the first sample (35) belongs, a step of providing a prompt (60) to a large-scale language model (6) including an instruction statement (50) for instructing the large-scale language model to identify a feature of the object and the text representation (55) of the first sample (35), thereby acquiring an answer (65) indicating a result of identifying the feature from the large-scale language model (6), and a step of outputting information about the acquired result. [Supplementary Note 19] The identification method according to Supplementary Note 18, wherein the prompt (60) further includes a list (59) of candidate labels for the feature identification result, and identifying the feature comprises selecting a label corresponding to the feature from the list (59). [Supplementary Note 20] The method of claim 18 or 19, further comprising obtaining second samples (45) of a second data type (40) different from the first data type (30), the second samples (45) being generated by observing an object with a second sensor, the large-scale language model (6) being configured to accept input of the samples of the second data type (40), and the prompt (60) further comprising the second samples (45).

[0181] REFERENCE SIGNS LIST 1...Identification device, 11...Control unit, 12...Storage unit, 13...External interface, 14...Input device, 15...Output device, 16...Drive, 81...Identification program, 91...Storage medium, 111...Acquisition unit, 112...Identification unit, 113...Output processing unit, 30...First data type, 35...First sample, 40...Second data type, 45...Second sample, 50...Command statement, 55...Text expression, 59...List, 60...Prompt, 6...Large-scale language model, S1...First sensor, S2...Second sensor, D1...Image data, D2...Sound data

Claims

1. An identification device comprising: a control unit configured to perform the following: obtaining a text representation of a first sample of a first data type according to a category to which the first sample belongs, the first sample being generated by observing an object with a first sensor; obtaining a response indicating a result of identifying the feature from a large-scale language model by providing a prompt to the large-scale language model, the prompt including a command statement for instructing the large-scale language model to identify a feature in the object and the text representation of the first sample; and outputting information regarding the obtained result.

2. The identification device of claim 1, wherein the prompt further includes a list of candidate labels for identifying the feature, and identifying the feature comprises selecting a label corresponding to the feature from the list.

3. The identification device of claim 1, wherein the control unit is further configured to acquire second samples of a second data type different from the first data type, the second samples being generated by observing an object with a second sensor; the large-scale language model is configured to accept input of the samples of the second data type; and the prompt further includes the second samples.

4. The identification device according to claim 3, wherein the second data type is image data.

5. The identification device of claim 4, wherein the first data type is tactile data.

6. The identification device according to claim 4, wherein the first data type is temperature data.

7. The identification device according to claim 4, wherein the first data type is sound data.

8. The identification device according to claim 4, wherein the first data type is point cloud data.

9. The identification device according to claim 4, wherein the first data type is weight data.

10. The identification device according to claim 4, wherein the first data type is odor data.

11. The identification device of claim 4, wherein the first data type is inertial property data.

12. The identification device of claim 3, wherein the instruction sentence includes: a first sub-sentence instructing to derive a first interim result that identifies the feature from the text representation of the first sample; a second sub-sentence instructing to derive a second interim result that identifies the feature from the second sample; and a third sub-sentence instructing to derive a result that identifies the feature based on the first interim result and the second interim result.

13. The identification device of claim 1, wherein the control unit is further configured to acquire the first sample; and determine a category to which the acquired first sample belongs, wherein determining the category to which the first sample belongs includes: converting the first sample into a feature value using a trained encoder generated by machine learning; and extracting one or more categories to which the first sample may belong from among a plurality of categories assigned to training samples, by comparing the feature value obtained from the first sample with a reference value obtained by converting each training sample used in the machine learning into the feature using the trained encoder; and wherein the text representation of the first sample includes a text representation of the one or more extracted categories.

14. An identification program for causing a computer to execute an identification method, the identification method comprising: a step of obtaining a text representation of a first sample of a first data type according to a category to which the first sample belongs, the first sample being generated by observing an object with a first sensor; a step of obtaining a response indicating a result of identifying the feature from a large-scale language model by providing a prompt to the large-scale language model, the prompt including a command statement for instructing the large-scale language model to identify a feature in the object and the text representation of the first sample; and a step of outputting information regarding the obtained result.

15. The identification program of claim 14, wherein the prompt further includes a list of candidate labels for identifying the feature, and identifying the feature comprises selecting a label corresponding to the feature from the list.

16. The identification program of claim 14, wherein the identification method further includes a step of acquiring a second sample of a second data type different from the first data type, the second sample being generated by observing an object with a second sensor; the large-scale language model is configured to accept input of the sample of the second data type; and the prompt further includes the second sample.

17. The identification program of claim 16, wherein the instruction statement includes: a first sub-sentence that instructs deriving a first interim result that identifies the feature from the text representation of the first sample; a second sub-sentence that instructs deriving a second interim result that identifies the feature from the second sample; and a third sub-sentence that instructs deriving a result that identifies the feature based on the first interim result and the second interim result.

18. A computer-implemented identification method, comprising: a step of obtaining a text representation of a first sample of a first data type according to a category to which the first sample belongs, the first sample being generated by observing an object with a first sensor; a step of obtaining a response indicating a result of identifying the feature from a large-scale language model by providing a prompt to the large-scale language model, the prompt including a command statement for instructing the large-scale language model to identify a feature in the object and the text representation of the first sample; and a step of outputting information regarding the obtained result.

19. The identification method of claim 18, wherein the prompt further includes a list of candidate labels for identifying the feature, and identifying the feature comprises selecting a label corresponding to the feature from the list.

20. The method of claim 18, further comprising the step of acquiring second samples of a second data type different from the first data type, the second samples being generated by observing an object with a second sensor; the large-scale language model being configured to accept input of the samples of the second data type; and the prompt further comprising the second samples.