Information processing method, information processing device, and computer system
By employing convolutional neural networks and natural language processing models to calculate and compare features across different data fields, the method enhances the accuracy and reliability of classification tasks, addressing the limitations of existing machine learning methods in integrating diverse data types.
Patent Information
- Application Number
- JP2021188040
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-18
- Publication Date
- 2025-09-01
- Estimated Expiration
- 2041-11-18
AI Technical Summary
Existing machine learning methods struggle to improve the accuracy of classification tasks, particularly in integrating diverse data fields such as images and natural language, leading to suboptimal performance in query data categorization.
The method involves calculating features in multiple fields, using convolutional neural networks for image data and natural language processing models like BERT to extract features, and then determining similarities between these features to select appropriate answers based on high similarity scores, thereby enhancing the classification process.
This approach improves the accuracy and reliability of classification tasks by leveraging multi-field feature extraction and similarity calculations, resulting in more precise categorization of query data.
Smart Images

Figure 0007731771000001 
Figure 0007731771000002 
Figure 0007731771000003
Abstract
Description
[Technical Field]
[0001] FIELD Embodiments of the present invention relate to an information processing method, an information processing device, and a computer system. [Background technology]
[0002] Methods, devices, and systems related to machine learning have been researched and proposed. For example, various calculation methods, processing methods, system configurations, and device configurations have been researched and proposed to improve the accuracy of various machine learning tasks. Using the results of machine learning, query data, which is input data, may be classified into a certain category / class. In order to improve the accuracy of this classification, it is necessary to improve the accuracy of the machine learning task. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 4703487 specification [Patent Document 2] Patent No. 4629280 specification [Patent Document 3] Patent No. 6811645 specification [Patent Document 4] Patent No. 5121917 specification Summary of the Invention [Problem to be solved by the invention]
[0004] An information processing method, an information processing device, and a computer system are provided that improve the accuracy of machine learning tasks. [Means for solving the problem]
[0005] The information processing method of this embodiment includes receiving query data to be processed, calculating a first feature of a first field of the query data, calculating a plurality of first similarities between the first feature and each of a plurality of second feature in a first feature space of the first field, acquiring a plurality of third feature of a second field associated with one or more feature selected from the plurality of second feature based on the plurality of first similarities from a second feature space of the second field, the third feature being associated with one or more feature selected from the plurality of second feature, calculating one or more fourth feature of the second field for a plurality of options related to the query data, calculating a plurality of second similarities between the plurality of third feature and each of the one or more fourth feature, and selecting at least one answer to the query data from a plurality of answer candidates corresponding to each of the plurality of third feature based on the plurality of second similarities. [Brief explanation of the drawings]
[0006] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of a computer system according to a first embodiment. [Figure 2] FIG. 1 is a block diagram showing an example of the configuration of an information processing device according to a first embodiment. [Figure 3] FIG. 2 is a block diagram showing an example of the configuration of a part of the information processing device according to the first embodiment. [Figure 4] FIG. 4 is a block diagram showing an example of the configuration of another part of the information processing device according to the first embodiment. [Figure 5] FIG. 2 is a schematic diagram for explaining part of the concept of the information processing method according to the first embodiment. [Figure 6] FIG. 4 is a schematic diagram for explaining another part of the concept of the information processing method according to the first embodiment. [Figure 7] FIG. 4 is a schematic diagram for explaining still another part of the concept of the information processing method according to the first embodiment. [Figure 8] FIG. 4 is a schematic diagram for explaining still another part of the concept of the information processing method according to the first embodiment. [Figure 9]10 is a flowchart illustrating a preparation phase of the computer system according to the embodiment. [Figure 10] FIG. 4 is a schematic diagram for explaining a part of the advance preparation phase of the first embodiment. [Figure 11] FIG. 10 is a schematic diagram for explaining another part of the advance preparation phase of the first embodiment. [Figure 12] 10 is a flowchart for explaining a classification task phase of the computer system according to the first embodiment. [Figure 13] FIG. 4 is a schematic diagram for explaining a part of the classification task phase according to the first embodiment. [Figure 14] FIG. 10 is a schematic diagram for explaining another part of the classification task phase according to the first embodiment. [Figure 15] FIG. 10 is a schematic diagram for explaining still another part of the classification task phase according to the first embodiment. [Figure 16] FIG. 10 is a schematic diagram for explaining still another part of the classification task phase according to the first embodiment. [Figure 17] FIG. 10 is a schematic diagram for explaining still another part of the classification task phase according to the first embodiment. [Figure 18] FIG. 6 is a schematic diagram for explaining an information processing method according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0007] Hereinafter, the present embodiment will be described in detail with reference to the drawings. In the following description, elements having the same functions and configurations are designated by the same reference numerals. In addition, in each of the following embodiments, when components (e.g., circuits, wiring, various voltages and signals, etc.) that are given reference symbols with distinguishing numbers / letters at the end do not need to be distinguished from each other, descriptions (reference symbols) with the numbers / letters at the end omitted are used.
[0008] [A] First embodiment A computer system according to an embodiment, an information processing device according to an embodiment, and an information processing method according to an embodiment will be described with reference to Figures 1 to 17. The information processing method according to an embodiment may include a control method for a computer system according to an embodiment, and a control method for an information processing device according to an embodiment.
[0009] (1) Composition FIG. 1 is a schematic diagram for explaining an example of the configuration of a computer system SYS according to this embodiment.
[0010] The computer system SYS of this embodiment communicates with an information communication device 9 via a wireless or wired network NW1.
[0011] The network NW1 is, for example, the Internet or an intranet. The information communication device 9 can perform various types of information processing and data processing. The information communication device 9 is a device such as a computer device or a mobile device. An example of a computer device is a personal computer or a server computer. An example of a mobile device is a smartphone, a feature phone, or a tablet device. The information communication device 9 may be a terminal device or a host device connected to a terminal device via a network (not shown).
[0012] The computer system SYS can receive various types of information and various types of data from the information communication device 9 via the network NW1. The computer system SYS can send various types of information and various types of data to the information communication device 9 via the network NW1.
[0013] The computer system SYS can execute various types of information processing, and is equipped with, for example, knowledge-seeking artificial intelligence (AI).
[0014] The computer system SYS includes an information processing device 1 of this embodiment and a storage device 5. The information processing device 1 and the storage device 5 may be provided in the same housing (not shown) or in different housings, as long as they can communicate with each other directly or indirectly. The information processing device 1 and the storage device 5 may be provided in the same country or region, or in different countries or regions, as long as they can communicate with each other directly or indirectly.
[0015] The information processing device 1 can execute various processes and tasks based on machine learning. For example, the information processing device 1 is configured to be able to execute deep learning using supervised or unsupervised learning data. The information processing device 1 includes a computer device. The information processing device 1 is, for example, a personal computer. However, the information processing device 1 may also be a mobile device such as a smartphone or a tablet device.
[0016] The information processing device 1 includes a processor 11, a random access memory (RAM) 12, a read-only memory (ROM) 13, and a plurality of interface circuits 18 and 19.
[0017] The processor 11 performs control processing and calculation processing for executing various processes and tasks of the information processing device 1. For example, the processor 11 includes a plurality of processing units 111, 112, and 115 for various control processing and calculation processing.
[0018] The RAM 12 temporarily stores various data and software used in the information processing device 1. The RAM 12 functions as a work memory and a buffer memory in the information processing device 1. The RAM 12 can be accessed by the processor 11 to obtain data.
[0019] For example, data includes user data to be processed, setting information used for various systems and devices, parameters used for various processes, and parts of software. For example, software may include executable programs, firmware, applications, and operating systems (OS). Data and / or software may correspond to information used by various systems and devices.
[0020] The ROM 13 stores, in a substantially non-volatile manner, an operating system (OS), firmware, various software, and various data used in the information processing device 1. The ROM 13 can be accessed by the processor 11 to retrieve data.
[0021] The interface circuit 18 transfers various data and various control signals between the information processing device 1 and the information communication device 9 based on a certain interface standard.
[0022] The interface circuit 19 transfers various data and various control signals between the information processing device 1 and the storage device 5 based on a certain interface standard.
[0023] The internal configuration and functions of the information processing device 1 will be described in detail later.
[0024] In addition, the information processing device 1 may further include a display device (not shown) such as an LCD display, an audio device (not shown) such as a speaker and a microphone, a user input device (not shown) such as a keyboard and a touch panel, and / or a photographing device (not shown) such as a camera.
[0025] The storage device 5 can store various types of information and various types of data. The storage device 5 can communicate with the information processing device 1 via a wireless or wired network NW2. The storage device 5 is, for example, an SSD. When the storage device 5 is an SSD, the storage device 5 includes a controller 50 and a nonvolatile semiconductor memory device 51. When the storage device 5 is an SSD, the nonvolatile semiconductor memory device 51 is a NAND flash memory.
[0026] The controller 50 commands the nonvolatile semiconductor memory device 51 to execute various operation sequences such as a write sequence and a read sequence of the nonvolatile semiconductor memory device 51. The controller 50 manages the memory space set in the nonvolatile semiconductor memory device 51. The controller 50 controls the transfer of data between the information processing device 1 and the storage device 5. The controller 50 includes a processor 501 , a RAM 502 , a ROM 503 , and a plurality of interface circuits 508 and 509 .
[0027] The processor 501 can execute various processes such as internal processes of the storage device 5, internal processes of the controller 50, and control processes of the nonvolatile semiconductor memory device 51. For example, the processor 501 executes various processes based on commands or requests from the information processing device 1.
[0028] The RAM 502 is a memory device that temporarily stores various data used by the controller 50. The RAM 502 functions as a work memory and a buffer memory in the controller 50. The RAM 502 temporarily stores information and data from the nonvolatile semiconductor memory device 51. The RAM 502 temporarily stores information and data from the information processing device 1. The RAM 502 can be accessed by the processor 501 to retrieve data.
[0029] The ROM 503 stores, in a substantially non-volatile manner, firmware, various software, and various data used in the storage device 5. The ROM 503 can be accessed by the processor 501 to retrieve data.
[0030] The interface circuit 508 receives various types of information, various types of data, and various types of control signals from the information processing device 1 based on a certain interface standard. The interface circuit 508 sends the control signals from the information processing device 1 to the processor 501. The interface circuit 508 sends the information and data from the information processing device 1 to the RAM 502. Based on the control of the processor 501, the interface circuit 508 sends the control signals from the processor 501 and the information and data in the RAM 502 to the information processing device 1.
[0031] The interface circuit 509 communicates with the non-volatile semiconductor memory device 51 based on a certain interface standard.
[0032] The interface circuit 509 sends data in the RAM 502 to the nonvolatile semiconductor memory device 51 under the control of the processor 501. The interface circuit 509 sends commands and addresses to the nonvolatile semiconductor memory device 51 in accordance with a required operation sequence. The interface circuit 509 receives data stored in the nonvolatile semiconductor memory device 51 from the nonvolatile semiconductor memory device 51. The interface circuit 509 sends various control signals to the nonvolatile semiconductor memory device 51 under the control of the processor 501. The interface circuit 509 receives signals controlled by the nonvolatile semiconductor memory device 51. The interface circuit 509 transfers data, commands, and addresses between the controller 50 and the nonvolatile semiconductor memory device 51.
[0033] For example, if the nonvolatile semiconductor memory device 51 is a NAND flash memory, the interface standard of the interface circuit 509 complies with the Toggle DDR interface standard or the ONFi (Open NAND Flash interface) standard.
[0034] In addition to the above components, the controller 50 may further include other components such as an ECC (Error Checking and Correction) circuit. The ECC circuit is a circuit for encoding and decoding data transferred between the controller 50 and the nonvolatile semiconductor memory device 51.
[0035] The nonvolatile semiconductor memory device 51 may be a memory device other than a NAND flash memory, as long as it is capable of storing data substantially nonvolatilely. The storage device 5 may be a hard disc drive (HDD). In this case, the storage device 5 includes a magnetic disk instead of the nonvolatile semiconductor memory device 51.
[0036] 2 to 4 are schematic block diagrams for explaining the information processing device 1 of this embodiment.
[0037] For example, the information processing device 1 of this embodiment executes a classification task based on machine learning, and classifies query data, which is question data, into a certain category / class through the classification task.
[0038] As shown in FIG. 2, in the information processing device 1 of this embodiment, the processor 11 includes a first feature extraction unit 111, a second feature extraction unit 112, a similarity calculation unit 113, a judgment unit 114, a control unit 115, and a calculation unit 116.
[0039] The first feature extraction unit 111 calculates a feature related to a first field of the data to be processed based on a certain calculation model / processing model related to the first field. The feature is a vector including multiple numerical values. The first field is selected from the field of images, natural language, speech, biological signals, electrical signals, etc. The field can also be referred to as a category, type, or group.
[0040] The second feature extraction unit 112 calculates features related to a second field of the data to be processed based on a certain calculation model / processing model related to the second field. The second field is different from the first field. The second field is selected from the field of images, natural language, speech, biological signals, electrical signals, etc., excluding the field selected as the first field.
[0041] The similarity calculation unit 113 calculates the similarity between certain data and other data. For example, the similarity calculation unit 113 calculates the similarity between a feature amount related to a first field of certain data and a feature amount related to the first field of other data. For example, the similarity calculation unit 113 calculates the similarity between a feature amount related to a second field of certain data and a feature amount related to the second field of other data.
[0042] For example, the similarity is calculated based on the inner product between two feature amounts, the cosine similarity between two feature amounts, the distance between two feature amounts, etc. The distance for calculating the similarity is obtained using, for example, any one of the Euclidean distance, the Manhattan distance, and the Minkowski distance.
[0043] The determination unit 114 makes a determination on various processes executed by the processor 11. For example, the determination unit 114 determines whether or not certain data and other data (e.g., two feature amounts) are similar based on the calculation result of the similarity calculation unit 113. If the similarity calculated between the certain data and other data is equal to or greater than a certain threshold, the determination unit 114 determines that the certain data and other data are similar. If the similarity calculated between the certain data and other data is less than a certain threshold, the determination unit 114 determines that the certain data and other data are not similar.
[0044] In this way, data having a high similarity to certain data is searched for in a database DB, which will be described later, by the similarity calculation unit 113 and the determination unit 114. The database DB is stored in the storage device 5.
[0045] The control unit 115 controls various processes executed by the processor 11 . The calculation unit 116 executes various calculation processes except for the calculation processes of the feature amount and the similarity.
[0046] In the information processing device 1 of this embodiment, the processor 11 executes a classification task for the query data QR. Specifically, the processor 11 classifies the query data QR based on a feature amount in a first field related to the query data QR and a feature amount in a second field related to answer options for the query data QR.
[0047] In the following, a case where the first field is the image field and the second field is the natural language field will be described. In this case, the first feature extraction unit 111 is also called an image feature extraction unit 111, and the second feature extraction unit 112 is also called a language feature extraction unit 112. Furthermore, features related to the image field are called image features, and features in the natural language field are called language features.
[0048] FIG. 3 is a schematic diagram showing an example of the configuration of the image feature amount extraction unit 111 in the information processing device 1 of this embodiment.
[0049] 3, the image feature extraction unit 111 calculates and extracts features of image data by, for example, a convolutional neural network (CNN) 200. In this embodiment, the image data is also referred to as an image data item, an image file, or simply an image.
[0050] In the image feature extraction unit 111, the CNN 200 has an input layer 210, one or more hidden layers 220 (220A, 220B), and an output layer 230.
[0051] The input layer 210 receives all or a portion of the image data for which image features are to be calculated. The input layer 210 sends data based on the received image data to the hidden layer 220. The input layer 210 includes a plurality of processing elements 211. In FIG. 3, the processing elements 211 are indicated as "NR."
[0052] The arithmetic element 211 is also called an artificial neuron or simply a neuron. The arithmetic element 211 extracts a signal of a certain size (for example, the number of bits) based on image data containing multiple signals. The signal supplied to the hidden layer 220 may be the data extracted by the arithmetic element 211 as is, or may be data that has been subjected to any processing by the arithmetic element 211.
[0053] The hidden layer 220 performs various calculation processes on the data from the input layer 210. The hidden layer 220 has a plurality of processing elements (artificial neurons) 221 (221A, 221B).
[0054] The multiple arithmetic elements 221 are connected in a network configuration. Each arithmetic element 221 has multiple input nodes and multiple output nodes. The multiple input nodes of each arithmetic element 221 are connected to the output nodes of the multiple arithmetic elements 221 in the previous stage, respectively. The multiple output nodes of each arithmetic element 221 are connected to the input nodes of the multiple arithmetic elements 221 in the subsequent stage. Each arithmetic element 221 performs convolution processing on the supplied data using parameters. For example, the parameters used by the arithmetic elements 221 are weighting coefficients. For example, the convolution processing is a product-sum operation. For example, each arithmetic element 221 performs a product-sum operation on the supplied data using weighting coefficients that are different from each other.
[0055] For example, the hidden layer 220 is layered (multi-layered) between the input layer 210 and the output layer 230. In the example of FIG. 3, the hidden layer 220 includes two layers 220A and 220B. Each processing element 221A in the hidden layer 220A performs a calculation process on the data from the input layer 210. Each processing element 221A sends the calculation result to each processing element 221B in the hidden layer 220B. Each processing element 221B performs a predetermined calculation process on the supplied data. Each processing element 221B sends the calculation result to the output layer 230. When the hidden layer 220 has a hierarchical structure, it is possible to improve the inference, learning, and classification capabilities of the CNN 200. The number of layers in the hidden layer 220 may be three or more, or may be one.
[0056] The output layer 230 receives data from each processing element 221 in the hidden layer 220. The output layer 230 performs various processes on the received data. The output layer 230 outputs the results of the calculation processes to the subsequent layer or circuit. The output layer 230 includes multiple processing elements (artificial neurons) 231.
[0057] Each arithmetic element 231 is connected to a plurality of arithmetic elements 221. Each arithmetic element 231 executes a predetermined process on the calculation results from the plurality of arithmetic elements 221. Each arithmetic element 231 can hold and output the obtained processing results.
[0058] The CNN 200 calculates image features of the image data, thereby extracting the image features of the image data.
[0059] The configuration of the image feature extraction unit 111 is not limited to a configuration using the CNN 200. Furthermore, the image feature extraction unit 111 may be configured using a configuration other than the CNN 200 depending on the field selected for feature calculation and extraction.
[0060] FIG. 4 is a schematic diagram showing an example of the configuration of the language feature extraction unit 112 in the information processing device 1 of this embodiment.
[0061] As shown in Fig. 4, the language feature extraction unit 112 calculates and extracts features of a text label as natural language using a neural network to which a natural language processing model such as BERT (Bidirectional Encoder Representations from Transformers) is applied. A text label is data including one or more characters. In this embodiment, a text label is also referred to as a text data item, text data, a text file, or simply a label. One or more characters included in a text label are also referred to as a character string hereinafter.
[0062] The example in Fig. 4 shows the model structure of the BERT 300. As shown in Fig. 4, the language feature extraction unit 112 using the BERT 300 includes an input layer 310, transformer layers 320 (320A, 320B), and an output layer 330.
[0063] The input layer 310 tokenizes the sentence or character string included in the text label TX supplied to the linguistic feature extraction unit 112. As a result, the sentence or character string with the text label TX is converted into a token string including multiple tokens tkn. The input layer 310 sends the token string that has been subjected to various processes to the transformer layer 320.
[0064] The input layer 310 includes multiple embedders 311. For example, the embedders 311 perform token embedding, segment embedding, and / or position embedding. The embedders 311 store tokens tkn, provide information for sentence differentiation, and provide information about character positions. In FIG. 4, the embedders 311 are shown as "Em."
[0065] The input layer 310 is also called a tokenizer layer (or simply a tokenizer) or an embedder layer (or simply an embedder).
[0066] The transformer layer 320 receives a token sequence from the input layer 310. The transformer layer 320 converts each of the multiple tokens included in the received token sequence into a vector. The transformer layer 320 includes multiple arithmetic elements (hereinafter also referred to as transformer elements) 321. In FIG. 4, the transformer elements 321 are denoted as "Tm."
[0067] The multiple transformer elements 321 are connected in a network configuration. Each transformer element 321 receives data from multiple transformer elements 321 in the previous layer. Each transformer element 321 sends a processed data signal to multiple transformer elements 321 in the subsequent layer. The transformer element 321 includes an encoder 322. The encoder 322 performs vector transformation processing on the received token or signal. For example, in the BERT 300, the transformer element 321 does not include a decoder in a natural language processing model, but only includes the encoder 322. The encoder 322 is also called a transformer encoder.
[0068] For example, the transformer layer 320 is layered into two layers 320A and 320B. However, the number of layers in the transformer layer 320 may be three or more, or may be one.
[0069] The output layer 330 receives the signal from the transformer layer 320. For example, the output layer 330 conditions the signal from the transformer layer 320.
[0070] BERT300 can be pre-trained without training data, and can perform various tasks, such as classification tasks, with relatively high accuracy, even with a relatively small amount of training data.
[0071] The BERT 300 calculates the linguistic features of the text labels, thereby extracting the linguistic features of the text labels.
[0072] The configuration of the language feature extraction unit 112 is not limited to a configuration using the BERT 300. Furthermore, the language feature extraction unit 112 may have a configuration other than the BERT 300 depending on the field selected for feature calculation and extraction.
[0073] The image feature extraction unit 111 and the language feature extraction unit 112 are provided as software or firmware to the processor 11. The image feature extraction unit 111 and the language feature extraction unit 112 are stored in a storage area (not shown) of the processor 11 as computer programs written in a certain programming language such as Python. However, the image feature extraction unit 111 and the language feature extraction unit 112 may be provided as hardware inside the processor 11 or outside the processor 11.
[0074] The software for the image feature extraction unit 111 and the software for the language feature extraction unit 112 may be stored in the ROM 13 or in the storage device 5. In this case, the software is read from the ROM 13 to a storage area of the processor 11, or from the storage device 5 to a storage area of the processor 11, when processing using the image feature extraction unit 111 and the language feature extraction unit 112, which will be described later, is executed.
[0075] The software for the image feature extraction unit 111 and the language feature extraction unit 112 may be stored in RAM 12 when the image feature extraction unit 111 and the language feature extraction unit 112 are used to execute the processes described below, and the software may be executed on RAM 12 by the processor 11.
[0076] In the information processing device 1 of this embodiment, the processor 11 can calculate a plurality of types of feature amounts relating to different fields using a plurality of feature amount extraction units 111 and 112. 2 , the information processing device 1 of this embodiment performs pre-preparation for executing a classification task, such as pre-learning, using a dataset Dst supplied from the information communication device 9. The dataset Dst includes one image data IMG and one or more text labels TX associated with the image data IMG. Note that the dataset Dst may be supplied to the information processing device 1 from a device other than the information communication device 9.
[0077] The image feature extraction unit 111 described above calculates and extracts image feature values IFV of image data IMG in the data set Dst. The above-mentioned linguistic feature extraction unit 112 calculates and extracts linguistic features LFV of the text labels TX of the dataset Dst.
[0078] For example, the text label TX to be used for feature calculation in the dataset Dst may be string data indicating the file name of the image data IMG, string data in the meta information of the image data IMG, and string data associated with the image data IMG in a certain text file. The linguistic feature extraction unit 112 may calculate and extract linguistic features of string data generated for a task to be performed, such as answers to a classification task and classification options. The text label TX may also be string data indicating the folder name of a data folder containing multiple image data IMG, or string data in the meta information of this data folder.
[0079] The information processing device 1 generates a database DB related to the data set Dst by performing a calculation process on the image feature vectors IFV and the language feature vectors LFV for the supplied data set Dst.
[0080] For example, the storage device 5 stores the generated database DB. For example, the database DB includes image features IFV of the image data IMG and language features LFV of the text labels TX in each data set Dst.
[0081] The database DB is stored substantially non-volatilely in a certain area of the non-volatile semiconductor memory device 51 of the storage device 5. The area in which the database DB relating to the feature amounts IFV and LFV is stored is also called a feature amount storage area.
[0082] In this embodiment, a set of multiple features related to a first field is referred to as a first feature space, and a set of multiple features related to a second field is referred to as a second feature space. Hereinafter, a set of one or more image features IFV is referred to as an image feature space FA1. Hereinafter, a set of one or more language features LFV is referred to as a language feature space FA2.
[0083] For example, in the database DB, a common identification number (ID) is associated with an image feature IFV and one or more language feature LFVs related to common image data IMG, thereby associating one image feature IFV with one or more language feature LFVs for each image data IMG. Hereinafter, a set Fst of image features IFV and one or more language features LFV that are associated with each other will be referred to as a feature set Fst.
[0084] For example, a set of k features Fst(Fst <0> ,Fst <1> ,···,Fst <k-1>) are managed by a database DB, where k is an integer equal to or greater than 1.
[0085] Multiple feature sets Fst <0> ,Fst <1> ,···,Fst <k-1>are different identification numbers ID <0> ,ID <1> ,···,ID <k-1>The processor 11 of the information processing device 1 associates an identification number ID with the image feature IFV and the language feature LFV that are associated with each other, for each data set Dst.
[0086] For example, identification number ID <0> Feature set Fst <0> As shown above, multiple linguistic features LFV <0> However, one image feature IFV <0> On the other hand, the identification number ID <1> Feature set Fst <1> As shown above, one linguistic feature LFV <1> Only one image feature IFV <1> It may also be associated with. Note that the feature set Fst for a certain identification number stored in the database DB may include only the image feature IFV without the language feature LFV, or the feature set Fst for a certain identification number may include only the language feature LFV without the image feature IFV.
[0087] In this way, a plurality of image features IFVs and a plurality of language features LFVs are managed as a database DB so that the corresponding image features IFVs and language features LFVs are associated with each other. Mutually related image features IFVs and language features LFVs are used in pairs for classification tasks.
[0088] The image data IMG and text labels TX of the data set Dst used to calculate the features IFV and LFV may be stored in the storage device 5 as data associated with the database DB. However, as long as the image features IFV and language features LFV of each data set Dst are stored in the storage device 5 as the database DB, the image data IMG and text labels TX do not need to be stored in the storage device 5.
[0089] The information processing device 1 of this embodiment executes a classification task for query data QR using image features IFV and language features LFV of a database DB. The query data QR is data to be processed by the task. In this embodiment, the query data QR is data to be classified in the classification task.
[0090] (2) Concept The concept of processing for a task executed by the information processing device 1 in this embodiment will be described with reference to FIGS.
[0091] In the computer system SYS of this embodiment, the information processing device 1 of this embodiment executes processing for a classification task related to query data QR using the configurations of FIGS.
[0092] As shown in FIG. 5, the information processing device 1 of this embodiment executes a similarity search process on image data as query data QR.
[0093] The information processing device 1 converts the image data as the query data QR into image data IMG <0> , image data IMG <1> , ..., and image data IMG <k-1>It is determined which of the image data IMG is similar to the image data IMG.
[0094] For example, the similarity search process for the query data QR is executed by a similarity calculation process for the image feature IFVq of the query data QR and a plurality of image feature IFVs in the database DB.
[0095] Based on the result of this similarity calculation process, the information processing device 1 selects image data IMG having a high similarity to the query data QR of the classification task TK and image data IMG having a low similarity to the query data QR.
[0096] As shown in FIG. 6, the information processing device 1 of this embodiment generates options for the classification task TK based on the results of the similarity search process for the image data IMG.
[0097] The information processing device 1 selects image data IMG having a high similarity to the query data QR based on the result of the similarity search process for each image data IMG related to the query data QR. For example, the information processing device 1 selects the image data IMG (image feature IFV) having the highest similarity from among multiple results of the similarity search process for the query data QR. In the example of FIG. 6, the image feature IFV <0> Image data IMG <0> is selected as the selected image data IMG-SEL.
[0098] The information processing device 1 selects one or more options CH(CH) of the classification task TK based on the selected image data IMG-SEL. <0> ,CH <1> ,···,CH <h-1>) In this embodiment, the choice CH is represented by the text label TXq (TXq <0> ,TXq <1> ,···,TXq <h-1>) is generated and presented. That is, the option CH is character string data.
[0099] As shown in FIG. 7, the information processing device 1 of this embodiment executes a similarity search process for one or more options CH in a classification task TK for query data QR. The information processing device 1 displays one or more text labels TX (TX <0> a,TX <0> b,TX <0> c,...) to which text label it is similar. One or more text labels TX associated with the selected image data IMG-SEL are treated as answer candidates in the classification task TK.
[0100] For example, the similarity between the answer choices CH of the query data QR and the text labels TX as answer candidates is determined by the linguistic features LFVq (LFVq <0> ,LFVq <1> ,···,LFVq <h-1>) and text-level TX linguistic features LFV (LFV <0> a,LFV <0> b,FVL <0> This is performed by calculating the similarity between the two vectors (c,...).
[0101] Based on the results of this similarity calculation process, the information processing device 1 selects text labels TX that have a high similarity to each option CH in the classification task TK of the query data QR, and text labels TX that have a low similarity to each option CH in the query data QR.
[0102] As shown in Figure 8, the information processing device 1 of this embodiment selects the most appropriate answer candidate from among multiple answer candidates for multiple options CH as the answer ANS for the classification task TK based on the results of a similarity search process using a text label TX associated with image data IMG.
[0103] For example, in the example of FIG. 8, the option CH with the number "0" <0> has the string "primate" and the choice CH with number "1" <1> has the string "birds" and the choice CH with number "h-1" <h-1>has the string "mammal".
[0104] For example, the selected image data IMG-SEL (here, the image data IMG <0> ) multiple text labels associated with TX <0> a,TX <0> b,TX <0> In c,..., the text label TX <0> a has the string "mammal" and the text label TX <0> b has the string "dog" and the text label TX <0> c has the string "Labrador Retriever".
[0105] As described above, by calculating the similarity between the options CH and the text label TX, the information processing device 1 selects, as the answer ANS for the classification task TK, an answer candidate (and the corresponding option CH) that has a high similarity (e.g., the highest similarity) with a certain option CH among the text labels TX associated with the selected image data IMG-SEL from among the multiple options CH and the text labels TX of the multiple answer candidates for the query data QR. In the example of FIG. 8, the information processing device 1 selects an option CH with the text label "Mammal". <0> and text label TX as answer candidate <0> Select a as your answer ANS. As a result, the information processing device 1 obtains an answer ANS to the query data QR.
[0106] Furthermore, in the calculation result of the similarity between an option CH and a text label TX, if it is determined that multiple pairs of options CH and text label TX have high similarity based on a certain judgment criterion (threshold), multiple options CH may be selected as multiple answers ANS for the classification task TK.
[0107] As described above, in the computer system SYS of this embodiment, the information processing device 1 of this embodiment executes a task TK for query data QR based on the result of a similarity determination process for a first field (here, the image field) for query data QR, and the result of a similarity determination process for a second field (here, the natural language field) that is associated with data in the first field and different from the first field. This allows the information processing device 1 of this embodiment to improve the reliability of tasks.
[0108] (3) Information processing method An information processing method by the information processing device 1 in the computer system SYS of this embodiment will be described with reference to FIGS.
[0109] The information processing method of the embodiment may include a control method for the computer system SYS of the embodiment and a control method for the information processing device 1 of the embodiment.
[0110] (3-1) Preparation Phase 9 and 10, the process of the advance preparation phase in the information processing method by the information processing device 1 of this embodiment will be described.
[0111] In the computer system SYS, the processor 11 of the information processing device 1 of this embodiment generates image features IFV of image data IMG included in one or more datasets Dst and language features LFV of multiple text labels TX through a preparation phase using one or more datasets Dst as follows. The generated image features IFV and language features LFV are stored in the storage device 5.
[0112] For example, the advance preparation phase in this embodiment corresponds to machine learning (for example, deep learning) and advance learning of the two feature extraction units 111 and 112 of the processor 11 of the information processing device 1.
[0113] FIG. 9 is a flowchart for explaining the advance preparation phase in the information processing method of the information processing device 1 in this embodiment.
[0114] <s11> The information processing device 1 receives the data set Dst. For example, the data set Dst is supplied from the information communication device 9 to the interface circuit 18 of the information processing device 1. In the information processing device 1, the processor 11 receives the data set Dst via the interface circuit 18.
[0115] FIG. 10 is a schematic diagram for explaining various types of data used in the information processing device 1 and the computer system SYS of this embodiment.
[0116] As shown in Figure 10, each data set Dst(Dst <0> ,Dst <1> ,Dst <2> , ) includes image data IMG and one or more text labels TX. The text labels TX include one or more characters related to the content of the object in the image data IMG.
[0117] A given dataset Dst includes one image data IMG and multiple text labels TX associated with the image data IMG.
[0118] In the example of Figure 10, the data set Dst <0> The image data IMG is an image of a dog. <0> In the example, the text label TXa has the character string "mammal", the text label TXb has the character string "dog", the text label TXc has the character string "Labrador retriever", and the text label TXd has the character string "Mr. A's Labrador Retriever".
[0119] The text labels TX in the dataset Dst may be generated by a user of the information communication device 9 or the information processing device 1 based on the content of the image data IMG, or may be generated by machine learning on the image data IMG by the information processing device 1.
[0120] <s12> The processor 11 calculates the image feature amount IFV of the image data IMG of the data set Dst using the image feature amount extraction unit 111, and extracts the image feature amount IFV.
[0121] For example, the image feature extraction unit 111 performs calculation processing on the image data IMG using the CNN 200 in Figure 3. As a result, the processor 11 obtains the image feature IFV related to the image data IMG. For example, the processor 11 temporarily stores the obtained image feature IFV in the RAM 12.
[0122] As shown in Fig. 10, the image feature IFV is expressed, for example, as two-dimensional data in which multiple numerical values num are arranged in a two-dimensional space of m x n. However, the image feature IFV may also be expressed as one-dimensional data in which numerical values num are arranged in a one-dimensional space, or as multidimensional data in which numerical values num are arranged in a multidimensional space of three or more dimensions. In Fig. 10, the magnitude of each numerical value num indicating the feature is schematically indicated by a shade of color ranging from white to black. The relationship between the illustrated image data IMG and the illustrated feature IFV is merely an example, and the magnitude of the numerical value num of the feature IFV varies depending on the parameters and calculation model used in the calculation.
[0123] <s13> The processor 11 calculates the linguistic feature LFV of each text label TX in the data set Dst using the linguistic feature extraction unit 112, and extracts the linguistic feature LFV.
[0124] For example, the language feature extraction unit 112 performs a calculation process using the BERT 300 in FIG. 4 on the text label TX. As a result, the processor 11 obtains language features LFV related to the text label TX. For example, the processor 11 temporarily stores the obtained one or more language features LFV in the RAM 12.
[0125] As shown in FIG. 10, in a certain dataset Dst, multiple linguistic features LFVa, LFVb, LFVc, and LFVd are calculated and extracted so as to correspond to multiple text labels TXa, TXb, TXc, and TXd, respectively. The linguistic feature LFV is expressed, for example, as two-dimensional data in which multiple numerical values num are arranged in a two-dimensional space of i×j. However, the linguistic feature LFV may also be expressed as one-dimensional data in which numerical values num are arranged in a one-dimensional space, or as multidimensional data in which numerical values num are arranged in a multidimensional space of three or more dimensions. Note that the relationship between the illustrated text labels TX and the illustrated feature LFV is merely an example, and the magnitude of the numerical value num of the feature LFV varies depending on the parameters and calculation model used in the calculation.
[0126] In this embodiment, the image features IFV and the language features LFV generated from a certain dataset Dst do not need to be similar to each other. However, it is desirable that the language features LFV generated from the respective text labels TX of a certain dataset Dst are similar to each other. It is desirable that the computation model of the language feature extraction unit 112, the feature computation method, and / or various parameters are appropriately designed so that the multiple text labels TX of a certain dataset Dst are similar to each other.
[0127] <s14> The processor 11 associates one image feature IFV of a certain data set Dst with one or more language features. For example, the processor 11 associates a common identification number ID with the image feature IFV and the language feature LFV of a certain data set Dst.
[0128] In the example of Figure 10, the identification number ID <0> But the dataset Dst <0> and a plurality of language features LFVa, LFVb, LFVc, and LFVd.
[0129] <s15> The processor 11 stores the image features IFV and language features LFV obtained by the processes of S12, S13, and S14 in the storage device 5. The processor 11 sends the image features IFV and language features LFV associated with each other in a certain data set Dst to the storage device 5 via the interface circuit 19.
[0130] The storage device 5 receives the image feature IFV and the language feature LFV. In the storage device 5, the controller 50 writes the image feature IFV to a certain address in the nonvolatile semiconductor memory device 51. The controller 50 writes the language feature LFV to a certain address in the nonvolatile semiconductor memory device 51. Note that the image feature IFV and the language feature LFV may be written to consecutive addresses as a series of data.
[0131] For example, the identification number ID associated with the image feature IFV and the language feature LFV may be managed together with the addresses at which the image feature IFV and the language feature LFV are stored by management information of the information processing device 1 or management information of the controller 50. The identification number ID may be written to a certain address in the non-volatile semiconductor memory device 51.
[0132] As a result, the information processing device 1 of this embodiment completes the advance preparation phase for a certain data set Dst.
[0133] The information processing device 1 executes the processes from S11 to S15 for each of the plurality of data sets Dst. As a result, a database DB containing a plurality of feature sets Fst is generated.
[0134] The image feature extraction unit 111 and the language feature extraction unit 112 are trained by calculating the image feature IFV and the language feature LFV using a plurality of data sets Dst.
[0135] Here, an example is described in which the database DB is generated using a data set Dst including image data IMG and text labels TX. However, the feature calculation processes for the mutually related image data IMG and text labels TX may be performed at different times.
[0136] For example, at a certain timing, only image data IMG is supplied to the information processing device 1, and image features IFV are calculated. This forms an image feature space FA1 of the database DB. After that, at another timing, only text labels TX are supplied to the information processing device 1, and language features LFV are calculated. This forms a language feature space FA2 of the database DB. When the text labels TX are supplied or when the language features are calculated, the information processing device 1 associates the image features IFV with the language features LFV. In this way, to form the feature set Fst, the language features LFV may additionally be associated with the image features IFV.
[0137] Note that either the language feature LFV or the image feature IFV in the feature set Fst may be deleted from the feature set Fst.
[0138] In this way, the language features LFV and image features IFV in the database DB can be edited as appropriate after the advance preparation phase.
[0139] An image feature space FA1, a language feature space FA2, and a database DB may be generated by performing various processes and deep learning on certain data.
[0140] 11, the information processing device 1 may generate image data having an image different from the image in the image data IMG by performing various processes on the image data IMG of a certain data set Dst that has been supplied to it. For example, the information processing device 1 may perform inversion processing, contrast change processing, zoom processing, and the like on the certain image data IMG.
[0141] The information processing device 1 calculates an image feature amount IFVx of the image data IMGx generated by the inversion process, an image feature amount IFVy of the image data IMGy obtained by the contrast change process, and an image feature amount IFVz of the image data IMGz obtained by the zoom process.
[0142] In this case, the language features LFVa, LFVb,... associated with the image features IFVx, IFVy, IFVz of the image data IMGx, IMGy, IMGz obtained by various processes are the same as the language features LFVa, LFVb,... of the text labels TXa, TXb,... associated with the original image data IMG.
[0143] In this way, a plurality of image feature amounts IFV, IFVx, IFVy, and IFVz are obtained from one image data item IMG, thereby increasing the number of feature amount sets Fst stored in the storage device 5. As a result, the information processing device 1 can improve the recognition accuracy of the image data IMG and the text label TX in response to the query data QR.
[0144] 9 to 11, the image data IMG and the text label TX are converted into image features IFV and language features LFV, which are numerical data, respectively. The obtained features IFV and LFV are stored in the storage device 5. This allows the storage device 5 of the computer system SYS to store large amounts of data used for machine learning in the information processing device 1 more efficiently in this embodiment.
[0145] Note that the image features IFV and language features LFV of multiple data sets Dst may be written collectively in the storage device 5. Furthermore, the image features IFV and language features LFV may be stored in the storage device 5 without being associated with the image features IFV and language features LFV for each data set Dst.
[0146] (3-2) Classification task phase The processing of the classification task phase in the information processing method by the information processing device 1 of this embodiment will be described with reference to FIGS.
[0147] In the computer system SYS, the processor 11 of the information processing device 1 of this embodiment executes a classification task TK for query data QR through a two-stage similarity search process as follows: The two-stage similarity search process is a process using a plurality of image features IFV and a plurality of language features LFV in the database DB generated in the advance preparation phase.
[0148] Fig. 12 is a flowchart illustrating the classification task phase in the information processing method of the information processing device 1 according to this embodiment. Fig. 13 to Fig. 17 are schematic diagrams illustrating the classification task phase in the information processing device 1 and computer system SYS according to this embodiment.
[0149] <s20> The information processing device 1 starts the classification task TK. For example, the processor 11 of the information processing device 1 accesses the RAM 12, the ROM 13, and the storage device 5, and starts various controls and processes for executing the classification task TK.
[0150] <s21> The information processing device 1 receives the query data QR. For example, as shown in Fig. 13, the query data QR is supplied from the information communication device 9 to the interface circuit 18 of the information processing device 1. In the information processing device 1, the processor 11 receives the query data QR via the interface circuit 18. In this embodiment, the query data QR includes image data IMGq.
[0151] <s22> The information processing device 1 calculates the image feature value IFVq of the image data IMGq of the query data QR.
[0152] 13, under the control of the control unit 115, the processor 11 calculates the image feature IFVq of the image data IMGq using the image feature extraction unit 111 including the CNN 200. As a result, the image feature IFVq related to the query data QR is extracted from the image data IMGq.
[0153] For example, the image feature IFVq of the query data QR is expressed, for example, as m×n two-dimensional data, similar to the image feature IFV of the feature set Fst. Note that the image feature IFVq of the query data QR may be expressed as one-dimensional data or multi-dimensional data of three or more dimensions. The image feature IFVq includes multiple (m×n) numerical values num arranged within an m×n area. Hereinafter, the image feature IFVq of the image data IMGq included in the query data QR will also be referred to as the query image feature IFVq.
[0154] <s23> The information processing device 1 executes a first similarity search process for the image data IMGq (query image feature IFVq) serving as the query data QR. In the first similarity search process, the information processing device 1 executes a process of calculating the similarity between the image feature IFVq of the query data QR and a plurality of image feature IFVs in the database DB in order to search the image feature space FA1 for image data IMG that has a relatively high similarity to the query data QR. For example, the similarity is calculated using a calculation method such as an inner product, a cosine similarity, or an Euclidean distance.
[0155] For example, the processor 11 accesses the database DB in the storage device 5. The processor 11 reads out a plurality of image feature values IFV from the storage device 5 to the RAM 12.
[0156] For example, as shown in FIG. 14, under the control of the control unit 115, the processor 11 calculates the image feature IFVq of the query data QR and the plurality of image feature IFVs of the database DB by the similarity calculation unit 113. <0> ,IFV <1> ,···,IFV <k-1>The similarity between each of the
[0157] For example, the first similarity search process for the image features IFV and IFVq can be sped up and / or made more efficient by graphing the image features IFV and IFVq and the calculated similarities.
[0158] <s24> The information processing device 1 selects one or more image features IFV-SEL, which are deemed to be similar to the query data QR in the image data IMG, from an image feature space FA1 containing multiple image features IFV, based on the calculation results of the similarity for the image features IFVq and IFV in the first similarity search process.
[0159] For example, the processor 11 determines whether the similarity between the query image feature IFVq and the image feature IFV is equal to or greater than a threshold value using the determination unit 114. As a result, the processor 11 selects the image feature IFV-SEL having the similarity equal to or greater than the threshold value. For example, the processor 11 selects the image feature IFV-SEL having the highest similarity to the query image feature IFVq of the image data IMGq. In the example of Figure 14, ID <0> Image feature IFV with identification number <0> is treated as the selected image feature IFV-SEL.
[0160] <s25> Based on the selected image feature IFV-SEL, the information processing device 1 selects and acquires one or more language features LFV associated with the selected image feature IFV-SEL from the language feature space FA2 in the database DB.
[0161] For example, the processor 11 accesses the database DB in the storage device 5. The processor 11 reads one or more language features LFV associated with the selected image feature IFV-SEL from the storage device 5 to the RAM 12 based on the identification number of the selected image feature IFV-SEL. In this way, the processor 11 acquires the language features LFV associated with the selected image feature IFV-SEL. Note that the language features LFV may be read into the RAM 12 simultaneously with the reading of the image feature IFV-SEL.
[0162] For example, in the example of FIG. <0> Image feature IFV with identification number <0> If selected, the processor 11 <0> Multiple linguistic features LFV with identification numbers <0> a,LFV <0> b,... are selected and acquired from a language feature space FA2 that includes multiple language features LFV. In this way, based on the identification number ID, a language feature LFV having the same identification number ID as the selected image feature IFV-SEL is selected.
[0163] For example, multiple selected linguistic features LFV <0> a,LFV <0> b,··· are candidate answers in the classification task TK.
[0164] <s26> The information processing device 1 generates and obtains one or more options CH for a classification task TK for query data QR. Each option CH includes a text label TXq. 15, the processor 11 generates and acquires a plurality of text labels TXq as options CH based on the query image feature IFVq and the selected image feature IFV-SEL. The options CH and the text labels TXq may be supplied to the information processing device 1 from outside the information processing device 1, for example, from an information communication device 9. The options CH and the text labels TXq may be supplied to the information processing device 1 simultaneously with the query data QR.
[0165] The text label TXq of the option CH can also be said to be text data associated with the image data IMGq of the query data QR.
[0166] <s27> The information processing device 1 calculates the linguistic feature LFVq of the text label TXq included in each of the multiple options CH.
[0167] 15, under the control of the control unit 115, the processor 11 calculates the linguistic feature LFVq of the text label TXq of each option CH by the linguistic feature extraction unit 112 including the BERT 300. In this way, the linguistic feature LFVq related to the option CH is extracted. One or more linguistic features LFVq are obtained depending on the number of answer candidates.
[0168] For example, the language feature LFVq of the option CH is expressed as two-dimensional data of i×j, for example, like the language feature LFV of the feature set Fst. Note that the language feature LFVq of the option CH may be expressed as one-dimensional data or multi-dimensional data of three or more dimensions. The language feature LFVq includes multiple (i×j) numerical values num arranged in an i×j area.
[0169] <s28> In this embodiment, a second similarity search process is executed for the text label TXq (linguistic feature LFVq) of the option CH. In the second similarity search process, the information processing device 1 executes a process of calculating the similarity between the linguistic feature LFVq of the option CH and the acquired multiple linguistic features LFV, in order to search the linguistic feature space FA2 for a text label TX that has a relatively high similarity to the text label TXq of the option CH. As in the above example, the similarity is calculated using a calculation method such as the inner product, cosine similarity, or Euclidean distance.
[0170] For example, as shown in FIG. 16, under the control of the control unit 115, the processor 11 calculates the similarity between the linguistic features LFVqa and LFVqb and the plurality of linguistic features LFVs in the database DB by the similarity calculation unit 113. <0> a,LFV <0> b,LFV <0> c,LFV <0> The similarity between each of the d is calculated.
[0171] For example, the similarity search process for the language features LFV and LFVq can be sped up and / or made more efficient by graphing the language features LFV and LFVq and the calculated similarities.
[0172] <s29> The information processing device 1 selects one answer ANS from among the multiple options CH and multiple answer candidates based on the result of the calculation process of the similarity for the linguistic features LFV and LFVq in the second similarity search process.
[0173] For example, the processor 11 determines whether the similarity between the language feature LFVq of the option CH and the language feature LFV of the answer candidate is equal to or greater than a threshold value using the determination unit 114. The processor 11 selects the language feature LFV having the similarity equal to or greater than the threshold value. The selected linguistic feature LFV (and the linguistic feature LFVq of the corresponding option CH) becomes the answer ANS in the classification task TK.
[0174] In the example of Figure 16, of the multiple options CHa and CHb, option CHa includes a linguistic feature LFVqa corresponding to the character string "Labrador retriever," and option CHb includes a linguistic feature LFVqb corresponding to the character string "Golden retriever." Each of the multiple linguistic feature LFVs obtained as answer candidates is the linguistic feature LFV corresponding to the character string "mammal". <0> a. Linguistic feature LFV corresponding to the character string "dog" <0> b) Linguistic feature LFV corresponding to the string "Labrador retriever" <0> c, and the linguistic feature LFV corresponding to the string "Mr. A's Labrador Retriever" <0> Includes d.
[0175] In this case, the processor 11 determines the linguistic feature LFV based on the processing result of the determination unit 114. <0> The text label TX (and the choice CH of the linguistic feature LFVqa) containing the character string c is selected as the answer ANS for the classification task TK.
[0176] There may be cases where a text label TX that matches an option CH of the classification task TK (that is, a linguistic feature LFV that is the same as the linguistic feature LFVq of the option CH) does not exist in the linguistic feature space FA2 of the database DB.
[0177] For example, in the example of FIG. 17, each of the multiple options CH1, CH2, and CH3 includes a linguistic feature LFVq1 corresponding to the character string "dog" (CH1), a linguistic feature LFVq2 corresponding to the character string "cat" (CH2), and a linguistic feature LFVq3 corresponding to the character string "cat" (CH3). Each of the multiple linguistic features LFV obtained as answer candidates includes a linguistic feature LFVq1 corresponding to the character string "mammal" (CH3). <0> a. Linguistic feature LFV corresponding to the string "Labrador retriever" <0> c, and the linguistic feature LFV corresponding to the string "Mr. A's Labrador Retriever" <0> In Fig. 17, there is no linguistic feature LFV corresponding to the character string "dog". Even in this case, the information processing device 1 of this embodiment can select the text label TX corresponding to "dog" as the answer ANS based on the degree of similarity between the language feature LFVq of the option CH and the multiple language features LFV acquired as answer candidates.
[0178] As described above, in this embodiment, the language features LFV of the multiple text labels TX of each dataset Dst are calculated and extracted so that they have values that are correlated with each other during the calculation and extraction processes of the language features LFV in the advance preparation phase.
[0179] Therefore, the information processing device 1 of this embodiment can select an answer ANS based on the degree of similarity between the language feature LFVq corresponding to the option CH and the language feature LFV of the answer candidate, even if there is no answer candidate (text label TX) that exactly matches the option CH of the classification task, and / or there is an answer candidate that contains an ambiguous expression relative to the option CH.
[0180] Therefore, even if the language feature LFV corresponding to the text label TX that matches the option CH does not exist in the database DB, the text label TX that will be the answer ANS can be derived from the language feature LFV that is most similar to each option CH based on the calculation results of the similarity between the language feature LFVq of the text label TXq of the option CH and multiple language features LFV read out from the database DB.
[0181] <s30> The information processing device 1 completes the classification task TK. For example, the processor 11 classifies the query data QR into a category or class corresponding to the answer ANS based on the answer ANS of the classification task TK. The result of the classification task TK may be displayed on a display device (not shown) of the information processing device 1.
[0182] This completes the processing of the classification task by the information processing device 1 of this embodiment.
[0183] (4) Summary The information processing device 1 and the computer system SYS of this embodiment perform similarity search processing in multiple stages in multiple fields, such as a combination of images and natural languages. This allows the information processing device 1 of this embodiment to improve the accuracy of the task to be executed, compared to when an answer to a task is determined based on a similarity search process in only one field.
[0184] The information processing device 1 of this embodiment can provide diverse answers in a task for query data by obtaining multiple answer candidates through the above-described operations and processes.
[0185] Therefore, the information processing device 1 of this embodiment can select a more appropriate answer from among a plurality of answer candidates depending on the content of the question in the query data QR.
[0186] As described above, the information processing device and information processing method of this embodiment can improve the accuracy of machine learning tasks.
[0187] [B] Second embodiment An information processing method, an information processing device, and a computer system according to the second embodiment will be described with reference to FIG.
[0188] In this embodiment, the information processing device 1 can determine an answer ANS based on an inference process by a majority vote process using multiple image features IFV (IFV-SEL) that are similar to the image feature IFVq of image data IMGq as query data, and language features LFV associated with the similar image features IFV.
[0189] FIG. 18 is a schematic diagram for explaining the inference process of the information processing method by the information processing device 1 of this embodiment.
[0190] The inference process is a process of predicting and determining which option CH among a plurality of options CH (and answer candidates) the query data supplied to the information processing device 1 corresponds to.
[0191] 18, the information processing device 1 receives image data (i.e., query image data) IMGq as query data QR as described above. The information processing device 1 starts a classification task TK for the query image data IMGq.
[0192] As described above, the information processing device 1 calculates and extracts the image feature IFVq of the query image data IMGq using the image feature extraction unit 111 of the processor 11. The information processing device 1 searches for and selects, from the database DB of the storage device 5, a plurality of image feature IFVs that have a relatively high similarity to the calculated image feature IFVq through a similarity search process related to the image field.
[0193] As described above, the information processing device 1 acquires one or more linguistic features LFV associated with one or more selected image features IFV-SEL. The information processing device 1 calculates and extracts, by the linguistic feature extraction unit 112 of the processor 11, a linguistic feature LFVq for each of the text labels TXq as one or more options CH of the classification task TK.
[0194] The information processing device 1 searches and selects, from the database DB of the storage device 5, a plurality of linguistic features LFVs that have a relatively high similarity to the calculated linguistic feature LFVq of the option CH through a similarity search process related to the natural language domain.
[0195] The information processing device 1 causes the processor 11 to execute an inference process for an answer ANS to the option CH based on the calculation result of the similarity between the linguistic features LFV and LFVq in the similarity search process.
[0196] In this embodiment, during the process of inferring an answer to an option CH, the information processing device 1 selects a certain number (here, s) of language features LFVs that have relatively high similarity to the language feature LFVq of the option CH from among the multiple language features LFVs associated with the one or more selected image features IFV, in order of decreasing similarity to the language feature LFVq of the option CH, where "s" is an integer equal to or greater than 1.
[0197] The information processing device 1 counts the number of language features LFVs having substantially the same value from among the s language features LFVs. As a result, the information processing device 1 performs grouping for each set of language features LFVs having substantially the same value. Note that the number of language features LFVs belonging to a certain numerical range may be counted, not limited to language features LFVs having the same value.
[0198] Counting the number of linguistic features LFV related to a certain numerical value (or a certain numerical range) corresponds to counting the number of text labels TX having substantially the same content (e.g., character strings) for each content of the text label TX.
[0199] The information processing device 1 selects, from one or more sets including the linguistic features LFV, the set that contains the largest number of linguistic features LFV as the answer ANS for the classification task TK.
[0200] For example, as shown in FIG. 18, when a text label TXq1 of "dog" and a text label TXq2 of "cat" are presented as options CH for a classification task TK, the information processing device 1 calculates and extracts, by the linguistic feature extraction unit 112, a linguistic feature LFVq1 corresponding to the text label TXq1 of "dog" and a linguistic feature LFVq2 corresponding to the text label TXq2 of "cat".
[0201] For the similarity search process, the information processing device 1 executes a calculation process of similarity between the language feature LFVq1 and the multiple language features LFV, and a calculation process of similarity between the language feature LFVq2 and the multiple language features LFV, respectively, for the multiple language features LFV associated with each image feature IFV. As a result, the information processing device 1 acquires s language features LFV having a value equal to or greater than a certain threshold from the multiple language features LFV of the multiple feature sets Fst selected by the similarity search process of the image features IFV, IFVq, with respect to the similarity to the language feature LFVq of the option.
[0202] The information processing device 1 counts, from among the s linguistic features LFV, the number of linguistic features LFVt similar to the numerical value corresponding to "dog" and the number of linguistic features LFVu similar to the numerical value corresponding to "cat". As an example, the number of linguistic features LFV having a numerical value corresponding to "dog" is t, and the number of linguistic features LFV having a numerical value corresponding to "cat" is u. Here, "t" and "u" are each an integer greater than or equal to 0 and less than or equal to s.
[0203] If "t" is greater than "u", the information processing device 1 selects "dog" as the answer ANS from the options CH (and answer candidates) of "dog" and "cat". If "t" is less than "u", the information processing device 1 selects "cat" as the answer ANS from the options CH (and answer candidates) of "dog" and "cat". When "t" is equal to "u", the information processing device 1 selects one of the multiple options CH (and answer candidates) as the answer ANS based on a preset rule.
[0204] As described above, in this embodiment, the information processing device 1 can determine one answer ANS for a plurality of options CH (and answer candidates) in the classification task TK by majority voting on the linguistic features LFV.
[0205] As a result, the information processing device, computer system, and information processing method of this embodiment can improve task accuracy.
[0206] [C] Application example The information processing device 1 and computer system of this embodiment are applied to image recognition systems, voice recognition systems, medical systems, and the like.
[0207] When the information processing device 1 of this embodiment is applied to an image recognition system, for example, as in the above-described embodiment, an image is selected as the first field (and first feature space), and a natural language is selected as the second field (and second feature space). The image may be a person's face, fingerprint, eyeball (or iris), etc. The natural language may be a character string indicating the name of an object, a person's name, the movement of an object, etc.
[0208] In the information processing device 1 applied to an image recognition system, natural language may be selected as the first field, and images may be selected as the second field.
[0209] When the information processing device 1 of this embodiment is applied to a voice recognition system, for example, natural language may be selected as the first field and speech may be selected as the second field.
[0210] In this case, for example, a text label obtained by converting an animal's cry into a sentence is supplied as query data QR to the information processing device 1. For example, the audio data is data of an animal's cry. Hereinafter, the feature amount of the audio data will be referred to as an audio feature amount.
[0211] An information processing device 1 in a speech recognition system performs a similarity search process between a text label as query data QR and a text label in a database DB using multiple linguistic features. The information processing device 1 calculates and extracts speech features of speech data corresponding to an option CH. The information processing device 1 performs a similarity search process between the speech features of the option CH and the speech features of the speech data associated with the selected text label. Based on the results of this search, the information processing device determines speech data as an answer ANS in a classification task.
[0212] In this voice recognition system, a storage device 5 stores a plurality of feature amounts related to text labels and a plurality of feature amounts related to voice data as a database DB.
[0213] In an information processing device 1 applied to a speech recognition system, speech may be selected as the first category, and images may be selected as the second category. In an information processing device 1 applied to a speech recognition system, speech in a first language system may be selected as the first category, and a natural language in a second language system different from the first language system may be selected as the second category. In addition, the speech included in the speech data may be a sound emitted by a living thing, such as an animal cry or a human voice, or may be a sound emitted by an inanimate object, such as a machine or a structure.
[0214] When the information processing device 1 of this embodiment is applied to a medical system, for example, a biological signal may be selected as the first field and a natural language may be selected as the second field. The biological signal may include one or more of brain waves, heart rate, pulse rate, blood pressure, respiration, and sweating.
[0215] In this case, biosignal data of a certain subject is supplied to the information processing device 1 as query data QR. The information processing device 1 performs a similarity search process between the feature amounts of the biosignal data as the query data QR and the feature amounts of the biosignal data in the database DB. The information processing device 1 calculates and extracts linguistic features of natural language corresponding to the option CH. The information processing device 1 performs a similarity search process between the linguistic features of the option CH and the linguistic features associated with the selected text label. Based on the result of this process, the information processing device determines a text label as the answer ANS in the classification task.
[0216] For example, the text labels associated with the biosignal data may include the subject's condition (e.g., emotion), the name of a disease, a symptom, or the name of a medication.
[0217] In this medical system, the storage device 5 stores a plurality of feature amounts related to biosignals and a plurality of feature amounts related to text labels as a database DB.
[0218] The information processing device 1 applied to the medical system may use images as the field and feature space of the similarity search process. In this case, X-ray images, magnetic resonance images, electrocardiograms, etc. are used as images for calculating features.
[0219] The information processing device 1 of this embodiment may be applied to systems other than the system described in this application example.
[0220] A system including the information processing device 1 of this embodiment can achieve the above-mentioned effects.
[0221] [D] Other In the above-described embodiment, the information processing device 1 and the information processing method perform a classification task on query data through a two-stage similarity search process using two fields (and two feature spaces). However, the information processing device 1 and the information processing method according to the embodiment may execute a classification task for query data by a process of determining similarity at three or more levels using three or more fields (feature amount spaces).
[0222] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, and are also included in the scope of the invention and its equivalents as defined in the claims. [Explanation of symbols]
[0223] 1: information processing device, 11: processor, 111, 112: feature extraction unit, 113: similarity calculation unit, 114: determination unit, 115: control unit, 5: storage device, 9: information communication device, SYS: computer system.
Claims
1. receiving query data to be processed; Calculating a first feature of a first field of the query data; Calculating a plurality of first similarities between the first feature and each of a plurality of second feature in a first feature space of the first field; acquiring, from a second feature space of the second field, a plurality of third feature amounts of a second field associated with one or more feature amounts selected from the plurality of second feature amounts based on the plurality of first similarities; Calculating one or more fourth features in the second category for a plurality of options related to the query data; calculating a plurality of second similarities between the plurality of third feature amounts and each of the one or more fourth feature amounts; selecting at least one answer to the query data from a plurality of answer candidates corresponding to the plurality of third feature quantities, based on the plurality of second similarities; An information processing method comprising:
2. The answer is selected by majority voting among the plurality of answer candidates. The information processing method according to claim 1 .
3. generating the first feature space by performing a calculation process on features of a plurality of first data items before receiving the query data; generating the second feature space by performing a calculation process on features of a plurality of second data items associated with each of the plurality of first data items before receiving the query data; 3. The information processing method according to claim 1, further comprising:
4. receiving a third data item; generating a fourth data item by a first operation on the third data item; generating the first feature space by performing a calculation process on the features of the third and fourth data items; The information processing method according to any one of claims 1 to 3, further comprising:
5. storing information about the first feature space and the second feature space in a storage device before receiving the query data; The information processing method according to any one of claims 1 to 4, further comprising:
6. the first field is one selected from an image, a natural language, a voice, and a biological signal; The second field is one of an image, a natural language, a voice, and a biological signal, excluding the field selected as the first field.
6. The information processing method according to claim 1.
7. the first similarity is calculated based on at least one of an inner product between the first feature amount and the second feature amount, a cosine similarity between the first feature amount and the second feature amount, and a distance between the first feature amount and the second feature amount; the second similarity is calculated based on at least one of an inner product between the third feature amount and the fourth feature amount, a cosine similarity between the third feature amount and the fourth feature amount, and a distance between the third feature amount and the fourth feature amount.
7. The information processing method according to claim 1.
8. an interface circuit for receiving query data to be processed; a processor that receives the query data via the interface circuit; Equipped with The processor: Calculating a first feature of a first field of the query data; acquiring a plurality of second features from a first feature space of the first field; calculating a plurality of first similarities between the first feature amount and each of the plurality of second feature amounts; acquiring, from a second feature space relating to the second field, a plurality of third feature amounts of a second field associated with one or more feature amounts selected from the plurality of second feature amounts based on the plurality of first similarities; calculating one or more fourth features in the second field for a plurality of options related to the query data; calculating a plurality of second similarities between the plurality of third feature amounts and each of the one or more fourth feature amounts; selecting at least one answer to the query data from a plurality of answer candidates corresponding to the plurality of third feature quantities, based on the plurality of second similarities; Information processing device.
9. The information processing device of claim 8; a storage device that stores the first feature space and the second feature space; A computer system comprising:
Citation Information
Patent Citations
Sutairasusochi
JP1976021917A
Recognition device, program, and construction device
JP2021015363A
Image searching apparatus, image searching method, and program
JP2021086438A
Knowledge discovery support device and support method
JP4629280B2
Image classification method, apparatus, and program
JP4703487B2