Information processing apparatus, information processing method and program

The information processing apparatus addresses the challenge of handling complex search conditions by calculating confidence levels and applying conversion methods specific to logical operators, resulting in improved information retrieval accuracy and relevance.

JP2025093111APending Publication Date: 2025-06-23KK TOSHIBA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023208640
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-06-23

AI Technical Summary

Technical Problem

Existing information retrieval methods using expression vectors struggle to accurately handle complex search conditions involving logical operators like AND, OR, and NOT, particularly when dealing with synonyms and varying notations.

Method used

An information processing apparatus that calculates a confidence level for each search term, converts similarity scores into continuous values based on confidence levels, and applies conversion methods specific to logical operators to produce a final conversion score for each document.

Benefits of technology

This approach enables more accurate and efficient information retrieval by appropriately handling complex search conditions and synonyms, improving the relevance of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025093111000001_ABST
    Figure 2025093111000001_ABST
Patent Text Reader

Abstract

To provide an information processing apparatus, an information processing method, and a program capable of more appropriately searching for information.SOLUTION: An information processing apparatus includes a processing unit. The processing unit calculates, for each of one or more phrases included in a search criterion including the one or more phrases and one or more logical operators, a degree of certainty that represents certainty of correlation between each of pieces of target information to be searched for and the phrases and that is expressed by a discrete value. The processing unit calculates, for each of the one or more phrases, a score that is a continuous value obtained by converting degrees of similarity between the pieces of target information and the phrases such that it falls within a range determined for the degree of certainty. The processing unit converts the score calculated for each of the one or more phrases into a conversion score in accordance with a conversion method determined for the logical operator.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] In recent years, techniques for retrieving information using expression vectors have been proposed. In this technique, each search term included in the search condition and the document to be searched are represented by high-dimensional vectors called expression vectors. Various methods have been proposed for obtaining expression vectors. Basically, search terms and documents having similar meanings are learned to have mutually similar expression vectors.

[0003] If such expression vectors can be learned, for example, it can be expected that "deep learning" and "deep neural learning" have similar expression vectors. And searching for documents having expression vectors similar to the expression vector of "deep learning" means searching for documents similar to "deep neural learning" as well. Therefore, a wide range of synonyms can be handled.

[0004] The search condition may include a plurality of search terms. Also, in this case, the search condition can describe a complex search condition including logical operators such as AND, OR, and NOT. For a search using expression vectors and using a complex search condition including logical operators, it is required to be able to search for information that matches the search condition more appropriately.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0006] An object of the present invention is to provide an information processing apparatus, an information processing method, and a program that can more appropriately search for information.

Means for Solving the Problems

[0007] The information processing apparatus according to the embodiment includes a processing unit. For each of one or more phrases included in a search condition including one or more phrases and one or more logical operators, the processing unit calculates a confidence level represented by discrete values, which represents the likelihood that each of a plurality of target information to be searched is related to the phrase. The processing unit calculates a score, which is a continuous value obtained by converting the similarity between the target information and the phrase so as to be included within a range determined according to the confidence level, for each of the one or more phrases. The processing unit converts the score calculated for each of the one or more phrases into a conversion score according to a conversion method determined according to the logical operator.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

[0009] Hereinafter, preferred embodiments of the information processing apparatus according to the present invention will be described in detail with reference to the accompanying drawings.

[0010] In recent years, with the progress of IoT (Internet of Things), memory devices have been growing in scale, and an environment has been developing in which various types of document data (hereinafter simply referred to as documents) can be stored in a server. Along with this, the demand for selecting documents according to search conditions (search expressions) input by users has been increasing. In response to such a demand for document search, a method using matching with search conditions has been mainly utilized. This method is a method of selecting documents that (completely or partially) match the search conditions input by the user, and is utilized in many scenarios.

[0011] On the one hand, in such matching technologies, it was impossible to handle synonyms with different notations but similar meanings, such as "deep learning" and "deep neural network", or synonyms with the same meaning. For example, when searching for a document containing either the word "deep learning" or "deep neural network", it was necessary to enter a search condition such as "search for documents containing deep learning or deep neural network". However, in addition to the above, "deep learning" has various synonyms such as "DNN (Deep Neural Network)" and "neural network". Therefore, it is practically difficult, especially for non-experts, to enter search conditions incorporating all of those synonyms.

[0012] The search technology using the above expression vectors is a technology capable of efficiently searching for synonyms. An expression vector is a vector that represents the meaning of words, documents, etc. Expression vectors may be called distributed expression vectors, embedding expression vectors, etc.

[0013] As a search method using expression vectors and a search condition including an AND operator, for example, a method can be considered in which the sum of the expression vectors of a plurality of search terms included in the search condition is used as the expression vector of the search condition for the search. For example, documents are output as search results in descending order of the value of the inner product between the expression vector of the search condition and the expression vector of the document. This process corresponds to calculating the sum of the inner products between the expression vectors of each search term and the expression vector of the document.

[0014] Since the magnitude of the expression vector differs depending on the word, the value of the inner product with the document may differ significantly depending on the word. In such a case, a situation may occur where the sum of the expression vectors of a plurality of search terms does not correctly reflect the search condition. Also, even if the sum of the expression vectors of words corresponds to a search condition including an AND operator, it can be interpreted that it does not correspond to search conditions including other logical operators (OR operator, NOT operator, etc.). For this reason, a search method that also considers logical operators other than the AND operator is required.

[0015] In the following embodiments, by having functions such as the following, for example, it is possible to perform a search using search conditions including AND operators, OR operators, and NOT operators uniformly using expression vectors. · For each word included in the search condition, calculate a confidence level represented by a discrete value that represents the probability related to each document. · Utilize the confidence level to calculate a score, which is a continuous value obtained by converting the similarity between each document and each word included in the search condition. · Calculate a conversion score by converting the score corresponding to each word included in the search condition in a manner according to logical operators (AND, OR, NOT). · Using the conversion score, calculate a final confidence level that represents the probability that the search condition and each document are related.

[0016] Note that in the following, an example in which a word is used as one or more phrases included in the search condition will be mainly described. That is, an example in which a word is used as a unit of an operand (operand) to which a logical operator is applied will be described. The phrase used as the unit is not limited to a word, and may be, for example, a phrase including a plurality of words. In the following, a word included in the search condition may be referred to as a search term.

[0017] (First Embodiment) An information processing apparatus according to the first embodiment will be described by taking an example in which a plurality of document data are used as a plurality of pieces of information (target information) to be searched. As will be described later, the target information is not limited to document data.

[0018] FIG. 1 is a block diagram showing an example of the configuration of an information processing apparatus 100 according to the first embodiment. As shown in FIG. 1, the information processing apparatus 100 includes a reception unit 101, a similarity calculation unit 102, a confidence level calculation unit 103, a score calculation unit 104, a conversion unit 105, an output control unit 106, and a storage unit 120.

[0019] The storage unit 120 stores various information used in the information processing apparatus 100. For example, the storage unit 120 stores a document DB (database) 121, an inverted index DB 122, a word vector DB 123, a document vector DB 124, a similarity DB 125, a confidence DB 126, and a score DB 127.

[0020] Note that the storage unit 120 can be configured by any commonly used storage medium such as a flash memory, a memory card, a RAM (Random Access Memory), an HDD (Hard Disk Drive), and an optical disk.

[0021] Some or all of each data (document DB 121, inverted index DB 122, word vector DB 123, document vector DB 124, similarity DB 125, confidence DB 126, score DB 127) stored in the storage unit 120 may be stored in physically different storage media, or may be stored in different storage areas of physically the same storage medium.

[0022] FIG. 2 is a diagram showing an example of the data structure of the document DB 121. The document DB 121 is a database for storing document data to be searched. As shown in FIG. 2, the document DB 121 includes a document ID and contents. The document ID is identification information for identifying a document. The contents are data indicating the contents of the document.

[0023] The documents stored in the document DB 121 may be of any type of document, for example, the following types of documents. Any language such as Japanese, English, and other languages may be used in the document. · Presentation materials · Emails · Patent documents · Reports · Technical books · Blogs · Documents handled on SNS (Social Networking Service)

[0024] All of the plurality of documents as described above may be subject to search, or the documents selected from the plurality of documents may be subject to search. Hereinafter, it is assumed that the documents to be searched (all documents or selected documents) are stored in the document DB121. In other words, in the present embodiment, all the documents stored in the document DB121 are treated as search targets.

[0025] The document data may be preprocessed in advance. The preprocessing may be any kind of processing, but for example, it is the following kind of processing. · Leave only words whose appearance frequency is equal to or more than the specified number of times. · Remove words with weak meanings such as "desu" and "masu". · Treat a subsequence of a specified number of characters called N-gram as a word. · Cut out words from the document using functions such as a tokenizer.

[0026] Note that the result of the preprocessing as described above is used, for example, when generating an inverted index from the document data.

[0027] FIG. 3 is a diagram showing an example of the data structure of the inverted index DB122. The inverted index DB122 is a database that stores an inverted index for identifying a document containing a word. As shown in FIG. 3, the inverted index DB122 includes a word and a document ID. The inverted index is used to identify the documents in which the word is included for each word. In FIG. 3, for example, it is shown that the word "deep learning" is included in the documents with document IDs "D_A", "D_B", and "D_D".

[0028] FIG. 4 is a diagram showing an example of the data structure of the word vector DB 123. The word vector DB 123 is a database for storing the representation vectors of each word (hereinafter also referred to as word vectors). As shown in FIG. 4, the word vector DB 123 includes words and element values for each of a plurality of dimensions (dimension 1, dimension 2). Note that although FIG. 4 (and FIG. 5 described later) shows an example of a two-dimensional representation vector, generally, a representation vector including elements of a large number of dimensions is used.

[0029] FIG. 5 is a diagram showing an example of the data structure of the document vector DB 124. The document vector DB 124 is a database for storing the representation vectors of each document (hereinafter also referred to as document vectors). As shown in FIG. 5, the document vector DB 124 includes a document ID and element values for each of a plurality of dimensions (dimension 1, dimension 2).

[0030] The representation vectors (word vectors, document vectors) may be calculated by any method. For example, a method using the following techniques can be applied. · Language models and natural language processing techniques such as word2vec · Graph neural network technology

[0031] The document DB 121, the inverted index DB 122, the word vector DB 123, and the document vector DB 124 are, for example, prepared in advance and stored in the storage unit 120. On the other hand, the similarity DB 125, the confidence DB 126, and the score DB 127 correspond to databases for storing information output by the processing of each part of the information processing apparatus 100. Examples of the data structures of the similarity DB 125, the confidence DB 126, and the score DB 127 will be described later.

[0032] Return to the description of FIG. 1. The reception unit 101 receives the input of various information used in the information processing apparatus 100. For example, the reception unit 101 receives the search conditions input by a user or the like. As described above, the search conditions can include a plurality of words and logical operators. For example, when searching for documents related to both deep learning and anomaly detection but not related to both natural language and image processing, a search condition such as "(Deep Learning AND Anomaly Detection) NOT (Natural Language OR Image Processing)" is input.

[0033] The similarity calculation unit 102 calculates the similarity between each document for each search term. The similarity can be calculated by any method. For example, it can be calculated by the cosine similarity of two representation vectors (word vector, document vector), or the inner product of two representation vectors. For example, the similarity calculation unit 102 calculates, for each search term, the cosine similarity or the inner product between the word vector of the search term and the document vector of each document as the similarity. The calculated similarity is stored, for example, in the similarity DB 125.

[0034] FIG. 6 is a diagram showing an example of the data structure of the similarity DB 125. The similarity DB 125 is a database for storing the similarity between each document for each word calculated by the similarity calculation unit 102. As shown in FIG. 6, the similarity DB 125 includes a ranking (order), a document ID, and a similarity. The ranking represents the order when the documents are sorted according to the magnitude of the similarity. Note that FIG. 6 shows an example of the similarity between each full document stored in the document DB 121 for one word. When the search condition includes a plurality of search terms, the similarity is calculated for each of the plurality of search terms and stored in the similarity DB 125.

[0035] Return to the description of FIG. 1. The confidence calculation unit 103 calculates, for each search term, the confidence representing the probability that each of the plurality of documents is related to the search term. The confidence is represented by a discrete value.

[0036] For example, the confidence calculation unit 103 calculates the confidence using the ranking (ranking) based on the similarity and the determination information. The determination information is information used to determine the relevance between the search term and the document, and for example, is information indicating whether each document contains the search term. In this case, the determination information can be obtained by referring to the inverted index DB. For example, when the search term is "latest", the confidence calculation unit 103 can identify from the inverted index DB122 in FIG. 3 that the documents with document IDs "D_A" or "D_D" are documents containing "latest". The confidence calculation unit 103 sets "〇" in the determination information for documents containing the search term, and sets "×" in the determination information for documents not containing the search term.

[0037] Next, the confidence calculation unit 103 calculates the confidence for each document using the ratio (M / N) of the cumulative number (M) of documents whose determination information contains the search term to the rank (N) of the document. Note that the ratio (M / N) is an example of a value for calculating the confidence, and the confidence may be calculated using other values determined by M and N. The cumulative number of documents containing the search term corresponds to, for example, the number of documents among one or more documents ranked higher than or equal to the rank of the document for which the confidence is calculated and whose determination information indicates that they contain the search term.

[0038] The confidence calculation unit 103 calculates the confidence represented by discrete values according to, for example, the comparison result between the ratio and one or more predetermined thresholds. For example, the confidence calculation unit 103 calculates the confidence as follows using two thresholds of 0.3 and 0.6. · When the ratio is 0.6 or more and 1 or less: Confidence A · When the ratio is 0.3 or more and less than 0.6: Confidence B · When the ratio is 0 or more and less than 0.3: Confidence C

[0039] This example is an example of calculating a three-value (three-stage) confidence using two thresholds. The thresholds and the number of discrete values that the confidence can take are not limited to these. The calculated confidence is stored in, for example, the confidence DB126.

[0040] FIG. 7 is a diagram showing an example of the data structure of the confidence level DB 126. The confidence level DB 126 is a database for storing the confidence levels calculated by the confidence level calculation unit 103. As shown in FIG. 7, the confidence level calculation unit 103 includes a ranking (N), information indicating whether a document includes a search term (determination information), the cumulative number (M) of documents including the search term, a ratio (M / N), and a confidence level.

[0041] Return to the description of FIG. 1. The score calculation unit 104 calculates a score, which is a continuous value obtained by converting the similarity between each document and the search term. For example, the score calculation unit 104 calculates the score for each document with respect to each search term using the similarity stored in the similarity DB 125 and the confidence level stored in the confidence level DB 126.

[0042] More specifically, the score calculation unit 104 calculates the score by converting the similarity so as to be included within a range determined according to the confidence level for each search term and for each confidence level. At this time, the score calculation unit 104 calculates the score so that the value becomes a value maintaining the magnitude relationship of the similarities within the range.

[0043] For example, in the case of the confidence level A, a range of 0.67 or more and 1 or less is determined. The score calculation unit 104 calculates the score by converting the similarity of each document corresponding to the confidence level A so as to be included within this range. The conversion method may be any method, and for example, linear interpolation, non-linear interpolation, spline interpolation, and a method using a radial basis function can be used.

[0044] When the range of the confidence level A is 0.67 or more and 1 or less, for example, the range of the confidence level B may be 0.33 or more and less than 0.67, and the range of the confidence level C may be 0 or more and less than 0.33. This example can be interpreted as an example in which the score has a lower limit value of 0 and an upper limit value of 1, and the three confidence levels are assigned to ranges obtained by dividing the range from the lower limit value to the upper limit value into approximately three equal parts. The lower limit value and the upper limit value of the score are not limited to 0 and 1, and may be any other values. Further, the method of dividing the range from the lower limit value to the upper limit value into a plurality of parts is not limited to the method of dividing into three equal parts as described above.

[0045] The calculated score is stored in, for example, the score DB 127. FIG. 8 is a diagram showing an example of the data structure of the score DB 127. The score DB 127 is a database for storing the scores calculated by the score calculation unit 104. Note that on the right side of FIG. 8, a graph showing an example of the interpolation method when calculating the score is also described.

[0046] As shown in FIG. 8, the score DB 127 includes a document ID, a similarity, a confidence level, and a score. Note that FIG. 8 shows an example of the score for one confidence level (confidence level A) of one search term. Scores are calculated for each search term and each confidence level and stored in the score DB 127.

[0047] In the example of FIG. 8, the similarities of the five documents with the confidence level A have values in the range of 16.20 to 64.23. The score calculation unit 104 calculates a score obtained by converting these similarity values into values in the range of 0.67 or more and 1 or less determined for the confidence level A. In the example of FIG. 8, the similarity 16.20 is converted to the score 0.67, and the similarity 64.23 is converted to the score 1.

[0048] Returning to the explanation of FIG. 1. The conversion unit 105 converts the scores calculated for each search term into conversion scores according to a conversion method determined according to the logical operator.

[0049] When the logical operator is the AND operator, the conversion unit 105 calculates, as the conversion score for the two search terms, a value close to the smaller score among the two scores calculated for each of the two search terms to which the AND operator is applied. In this case, the conversion score can be interpreted as a score obtained by integrating the two scores for the two search terms to which the AND operator is applied into one.

[0050] Of the two scores, the value closer to the smaller score is, for example, the value of the smaller score among the two scores (the minimum value of the scores). The conversion score does not have to be the minimum value, and may be calculated in any way as long as it is a value closer to the smaller score among the two scores. For example, the conversion score may be a value corresponding to the first quartile of the two scores.

[0051] When the logical operator is the OR operator, the conversion unit 105 calculates, as the conversion score for the two search terms, a value closer to the larger score among the two scores calculated for each of the two search terms to which the OR operator is applied. In this case, the conversion score can be interpreted as a score that integrates the two scores for the two search terms to which the OR operator is applied into one.

[0052] Of the two scores, the value closer to the larger score is, for example, the value of the larger score among the two scores (the maximum value of the scores). The conversion score does not have to be the maximum value, and may be calculated in any way as long as it is a value closer to the larger score among the two scores. For example, the conversion score may be a value corresponding to the third quartile of the two scores.

[0053] When the logical operator is the NOT operator, the conversion unit 105 calculates, as the conversion score for the search term, a value such that the larger the value of the score calculated for the one search term to which the NOT operator is applied, the smaller the value (the smaller the value, the larger the value). For example, the conversion unit 105 calculates, as the conversion score, a value obtained by subtracting the score from the upper limit value of the score (for example, 1).

[0054] The conversion unit 105 may calculate the final confidence of each document using the conversion score. For example, the conversion unit 105 calculates the final confidence by referring to the range of values determined for each confidence level, which is used when the score calculation unit 104 calculates the score.

[0055] As in the above example, assume that the range of confidence level A is 0.67 or more and 1 or less, the range of confidence level B is 0.33 or more and less than 0.67, and the range of confidence level C is 0 or more and less than 0.33. In this case, when the conversion score is 0.7, for example, the conversion unit 105 calculates the confidence level A corresponding to the range including 0.7 as the final confidence level.

[0056] The search condition may include a plurality of logical operators. In such a case, the conversion unit 105 repeatedly performs a process (score conversion process) of converting an operand (score or conversion score) into a conversion score in accordance with the order (priority) of applying the logical operators.

[0057] For example, in the case of the above search condition of "(deep learning AND anomaly detection) NOT (natural language OR image processing)", the score conversion process may be executed in the following order. (A1) For the AND operator, calculation of the conversion score using the score of the search term "deep learning" and the score of the search term "anomaly detection" (A2) For the OR operator, calculation of the conversion score using the score of the search term "natural language" and the score of the search term "image processing" (A3) For the NOT operator, calculation of the conversion score obtained by further converting the conversion score calculated in (A2) (A4) Calculation of the conversion score obtained by performing the conversion corresponding to the AND operator on the conversion score calculated in (A1) and the conversion score calculated in (A3)

[0058] The output control unit 106 controls the output of various types of information used in the information processing apparatus 100. For example, the output control unit 106 outputs one or more documents selected according to the conversion score as search results. The output control unit 106 may output the final confidence level for each document determined based on the conversion score.

[0059] The method of outputting information by the output control unit 106 may be any method. For example, a method of displaying information on a display device such as a display, and a method of transmitting information to other devices (such as a server device) via a network can be used.

[0060] At least a part of the above-described units (reception unit 101, similarity calculation unit 102, confidence calculation unit 103, score calculation unit 104, conversion unit 105, and output control unit 106) may be realized by one or more processing units. The above-described units may be realized by, for example, one or a plurality of processors. For example, the above-described units may be realized by causing a processor such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit) to execute a program, that is, by software. The above-described units may be realized by a processor such as a dedicated IC (Integrated Circuit), that is, by hardware. The above-described units may be realized by using a combination of software and hardware. When using a plurality of processors, each processor may realize one of the units or two or more of the units.

[0061] Also, the information processing apparatus 100 may be physically configured by one device or may be physically configured by a plurality of devices. For example, the information processing apparatus 100 may be constructed in a cloud environment. Also, each unit in the information processing apparatus 100 may be provided in a distributed manner among a plurality of devices.

[0062] Next, the search process by the information processing apparatus 100 of the first embodiment will be described. FIG. 9 is a flowchart showing an example of the search process in the first embodiment.

[0063] The reception unit 101 receives the search conditions input by a user or the like (step S101). The similarity calculation unit 102 calculates the similarity between each of the plurality of documents stored in the document DB 121 and each word (search term) included in the search conditions (step S102).

[0064] The confidence calculation unit 103 calculates a confidence level, which is a discrete value, for each search term included in the search conditions using the similarity (step S103). The score calculation unit 104 calculates a score, which is a continuous value within the range corresponding to the confidence level, for each search term and for each confidence level using the similarity (step S104).

[0065] The conversion unit 105 calculates a converted score obtained by converting the score of the word according to the logical operator (AND, OR, NOT) included in the search conditions (step S105). The output control unit 106 outputs the converted score, which is the processing result (step S106), and ends the search process. The output control unit 106 may output, as the processing result, the final confidence level for each document determined using the converted score.

[0066] Next, the details of the confidence calculation process in step S103 will be described. FIG. 10 is a flowchart showing an example of the confidence calculation process.

[0067] The confidence calculation unit 103 acquires an unprocessed word among the words (search terms) included in the search conditions (step S201). The confidence calculation unit 103 determines the ranking of the documents using the similarity value calculated by the similarity calculation unit 102 (step S202). The confidence calculation unit 103 assigns information (determination information) indicating whether or not the document contains the search term to each document (step S203). For example, the confidence calculation unit 103 assigns determination information in which a value of ○ is set when the document contains the search term and a value of × is set when it does not.

[0068] The confidence calculation unit 103 calculates, for each document, a confidence level based on a value determined by M and N, such as the ratio (M / N) of the number of documents containing the search term to the document and the documents ranked higher than the document (step S204). The confidence calculation unit 103 stores the calculated confidence level in the confidence DB 126 (step S205).

[0069] The confidence calculation unit 103 determines whether all search terms have been processed (step S206). If not all search terms have been processed (step S206: No), the confidence calculation unit 103 returns to step S201 and repeats the process for the next unprocessed search term. If all search terms have been processed (step S206: Yes), the confidence calculation unit 103 ends the confidence calculation process.

[0070] Next, the details of the score calculation process in step S104 will be described. FIG. 11 is a flowchart showing an example of the score calculation process.

[0071] The score calculation unit 104 acquires an unprocessed word among the words (search terms) included in the search condition (step S301). The score calculation unit 104 acquires the unprocessed confidence for the acquired search term (step S302).

[0072] The score calculation unit 104 acquires the similarity of each document with the acquired confidence from the similarity DB125 (step S303). The score calculation unit 104 interpolates the similarity to calculate a score so that the acquired similarity falls within the set range (step S304). The score calculation unit 104 stores the calculated score in the score DB127 (step S305).

[0073] The score calculation unit 104 determines whether all confidences have been processed (step S306). If not all confidences have been processed (step S306: No), the score calculation unit 104 returns to step S302 and repeats the process for the next unprocessed confidence.

[0074] If all confidences have been processed (step S306: Yes), the score calculation unit 104 determines whether all search terms have been processed (step S307). If not all search terms have been processed (step S307: No), the score calculation unit 104 returns to step S301 and repeats the process for the next unprocessed search term. If all search terms have been processed (step S307: Yes), the score calculation unit 104 ends the score calculation process.

[0075] Next, the details of the score conversion process in step S105 will be described. FIG. 12 is a flowchart showing an example of the score conversion process.

[0076] The conversion unit 105 identifies a logical operator for a word (search term) included in the search condition (step S401). The conversion unit 105 acquires the score of the search term from the score DB 127 (step S402). The conversion unit 105 calculates a converted score obtained by converting the score by a conversion method according to the logical operator (step S403).

[0077] FIGS. 13 to 15 are diagrams showing examples of converted scores. FIGS. 13, 14, and 15 show examples of converted scores in the case of search conditions including the AND operator, the OR operator, and the NOT operator, respectively.

[0078] For example, in FIG. 13, an example of the converted score for the search condition of "deep learning AND semiconductor" is shown. For the document with document ID "D_A", the score for "deep learning" is 0.98, and the score for "semiconductor" is 0.84. Assuming that a conversion method is defined in which the smaller score value of the two scores is used as the converted score for the AND operator, the converted score is calculated as 0.84.

[0079] In FIG. 14, an example of the converted score for the search condition of "deep learning OR anomaly detection" is shown. For the document with document ID "D_A", the score for "deep learning" is 0.98, and the score for "anomaly detection" is 0.84. Assuming that a conversion method is defined in which the larger score value of the two scores is used as the converted score for the OR operator, the converted score is calculated as 0.98.

[0080] In FIG. 15, an example of the conversion score for the search condition of "NOT deep learning" is shown. For the document with the document ID "D_A", the score for "deep learning" is 0.98. Assuming that a conversion method is defined to calculate the conversion score as the value obtained by subtracting the score from the upper limit value of 1 for the NOT operator, the conversion score is calculated as 0.02.

[0081] Note that in FIGS. 13 to 15, the final confidence levels calculated using the conversion scores are also described.

[0082] As described above, in the information processing apparatus according to the first embodiment, it is possible to uniformly execute a search using a search condition including an AND operator, an OR operator, and a NOT operator using expression vectors for document data.

[0083] (Second Embodiment) In the first embodiment, information indicating whether each document includes a search term is used as determination information. The information processing apparatus according to the second embodiment uses determination information different from that of the first embodiment.

[0084] FIG. 16 is a block diagram showing an example of the configuration of the information processing apparatus 100-2 according to the second embodiment. As shown in FIG. 16, the information processing apparatus 100-2 includes a reception unit 101, a similarity calculation unit 102, a confidence level calculation unit 103-2, a score calculation unit 104, a conversion unit 105, an output control unit 106, and a storage unit 120-2.

[0085] In the second embodiment, the function of the confidence level calculation unit 103-2 and the fact that the storage unit 120-2 includes a feedback DB 128-2 instead of the inverted index DB 122 are different from those of the first embodiment. Since the other configurations and functions are the same as those in FIG. 1, which is a block diagram of the information processing apparatus 100 according to the first embodiment, the same reference numerals are given and the description thereof is omitted here.

[0086] Feedback DB 128-2 stores, for each search term, the number of times each document retrieved by the search term was selected by a user, for example, as feedback information. FIG. 17 is a diagram showing an example of the data structure of Feedback DB 128-2. As shown in FIG. 17, Feedback DB 128-2 includes a word, a document ID, and a count. The count represents the number of times a document (identified by the document ID) retrieved by a search condition including the corresponding word was selected by the user.

[0087] It can be interpreted that the higher the count, the greater the relevance between the corresponding search term and the corresponding document. Therefore, in this embodiment, such feedback information is used as determination information to determine the relevance between the search term and the document.

[0088] That is, the confidence calculation unit 103-2 calculates the confidence using the ranking based on the similarity and the determination information indicating whether the document was selected as information related to the search term.

[0089] Next, the details of the confidence calculation process in this embodiment will be described. FIG. 18 is a flowchart showing an example of the confidence calculation process.

[0090] Steps S501 to S502 are the same as steps S201 to S202 in the confidence calculation process (FIG. 10) of the first embodiment, so the description thereof is omitted.

[0091] The confidence calculation unit 103-2 assigns information (determination information) indicating whether, for example, the number of feedbacks for each document is equal to or greater than a specified number (step S503). For example, the confidence calculation unit 103-2 assigns determination information in which a value of 〇 is set when the count corresponding to the document is equal to or greater than the specified number and a value of × is set otherwise.

[0092] Steps S504 to S506 are the same as steps S204 to S206 in the confidence calculation process (FIG. 10) of the first embodiment, so the description thereof is omitted.

[0093] (Third Embodiment) Search conditions can specify complex conditions by combining logical operators (AND, NOT, OR) and symbols such as parentheses indicating the order of application of the operators (priority). On the other hand, imposing the input of overly complex search conditions on users may lead to a decrease in satisfaction. Therefore, the information processing apparatus according to the third embodiment has a function that enables easier input of search conditions including logical operators.

[0094] FIG. 19 is a block diagram showing an example of the configuration of the information processing apparatus 100-3 according to the third embodiment. As shown in FIG. 19, the information processing apparatus 100-3 includes a reception unit 101, a similarity calculation unit 102, a confidence calculation unit 103, a score calculation unit 104, a conversion unit 105, an output control unit 106, a creation unit 107-3, and a storage unit 120.

[0095] In the third embodiment, it is different from the first embodiment in that the creation unit 107-3 is added. Since the other configurations and functions are the same as those in FIG. 1 which is a block diagram of the information processing apparatus 100 according to the first embodiment, the same reference numerals are given and the description here is omitted.

[0096] The creation unit 107-3 creates search conditions from one or more words input using an input screen. The input screen includes a region RA (first region) for designating one or more words WA (first phrase) designated as words included in a document (target information) that is a search result, and a region RB (second region) for designating one or more words WB (second phrase) designated as words not included in the document that is a search result.

[0097] FIG. 20 is a diagram showing an example of the input screen 2000. The input screen 2000 includes a region 2001 corresponding to the region RA, a region 2002 corresponding to the region RB, a search button 2011, and a cancel button 2012.

[0098] When the cancel button 2012 is pressed, for example, the output control unit 106 displays the screen before the input screen is displayed. When the search button 2011 is pressed, for example, the reception unit 101 receives one or more words WA input in the area 2001 and one or more words WB input in the area 2002, and passes them to the creation unit 107-3.

[0099] The creation unit 107-3 creates a search condition using the AND, OR, and NOT operators with the passed information. For example, when a plurality of words WA are input in the area 2001, the creation unit 107-3 creates a search condition that combines the words WA using the AND operator. For example, when a plurality of words WB are input in the area 2002, the creation unit 107-3 combines the words WB using the OR operator and further creates a search condition with the NOT operator added. When search terms are input in both the areas 2001 and 2002, the creation unit 107-3 creates a search condition that combines the search conditions created for both.

[0100] For example, as shown in FIG. 20, when the word WA (words to be included) is "deep learning, anomaly detection" and the word WB (words not to be included) is "image processing, natural language", the creation unit 107-3 creates a search condition of (deep learning AND semiconductor) NOT (voice OR device).

[0101] Since the search process using the created search condition is the same as in the first embodiment, the description is omitted. Note that the creation unit 107-3 may be configured to be added to the second embodiment.

[0102] Thus, in the information processing apparatus according to the third embodiment, the search condition can be input more easily.

[0103] (Modification 1) In Modification 1, the threshold value used when calculating the confidence level is adjustable. For example, the confidence level calculation unit 103 in this modification adjusts (changes) the threshold value when calculating the confidence level according to feedback information such as the situation of document selection by the user.

[0104] For example, the threshold value may be adjusted so that the number of documents corresponding to each of the plurality of confidence levels is not biased. In the example of FIG. 7, the value of the threshold for confidence level A is changed from 0.6 to 0.8. As a result, the number of documents for which confidence level A is calculated is 4, and the number of documents for which confidence level B is calculated is 3 different.

[0105] The confidence level calculation unit 103 of this modification example may set the average value of the threshold values adjusted for the plurality of search terms as the final threshold value.

[0106] (Modification Example 2) As described above, the information to be searched (target information) is not limited to document data. As long as it is information that can be represented by a representation vector, information other than document data may be used as the target information. Examples of target information other than document data will be described below.

[0107] For example, image data and information indicating a product (hereinafter referred to as product data) can be used as target information. As a method for obtaining representation vectors for image data and product data, a method using a graph neural network and deep learning can be applied.

[0108] In the case of image data, for example, the purpose is to search for image data that includes or does not include the object described in the search condition. Therefore, the confidence level calculation unit 103 can use determination information indicating whether the image data includes the object described in the search condition.

[0109] In the case of product data, for example, the purpose is to search for product data that meets the search conditions specified by the product category and the attributes of the product (such as lemon flavor). Therefore, the confidence level calculation unit 103 can use determination information indicating whether the product data meets the specified search conditions.

[0110] As described above, according to the first to third embodiments, information can be searched more appropriately.

[0111] Next, the hardware configuration of the information processing apparatus according to the first to third embodiments will be described with reference to FIG. 21. FIG. 21 is an explanatory diagram showing a hardware configuration example of the information processing apparatus according to the first to third embodiments.

[0112] The information processing apparatus according to the first to third embodiments includes a control device such as a CPU (Central Processing Unit) 51, a storage device such as a ROM (Read Only Memory) 52 and a RAM (Random Access Memory) 53, a communication I / F 54 that connects to a network and performs communication, and a bus 61 that connects each part.

[0113] The program executed by the information processing apparatus according to the first to third embodiments is provided by being pre-embedded in the ROM 52 or the like.

[0114] The program executed by the information processing apparatus according to the first to third embodiments may be configured to be recorded on a computer-readable recording medium such as a CD-ROM (Compact Disk Read Only Memory), a flexible disk (FD), a CD-R (Compact Disk Recordable), or a DVD (Digital Versatile Disk) in an installable or executable file format and provided as a computer program product.

[0115] Furthermore, the program executed by the information processing apparatus according to the first to third embodiments may be configured to be stored on a computer connected to a network such as the Internet and downloaded via the network for providing. Also, the program executed by the information processing apparatus according to the first to third embodiments may be configured to be provided or distributed via a network such as the Internet.

[0116] The program executed by the information processing apparatus according to the first to third embodiments can cause a computer to function as each unit of the information processing apparatus described above. This computer can read a program from a computer-readable storage medium and execute it in the main storage device by the CPU 51.

[0117] A configuration example of the embodiment will be described below. (Configuration Example 1) For each of one or more of the phrases included in a search condition including one or more phrases and one or more logical operators, a probability representing the relevance between each of a plurality of target information to be searched and the phrase is calculated, and a confidence level represented by a discrete value is calculated. For each of one or more of the phrases, a score, which is a continuous value obtained by converting the similarity between the target information and the phrase so as to be included within a range determined according to the confidence level, is calculated. The scores calculated for each of one or more of the phrases are converted into conversion scores according to a conversion method determined according to the logical operator. Processing unit An information processing apparatus including the same. (Configuration Example 2) The processing unit For each of a plurality of the target information, the confidence level is calculated using the rank based on the similarity and determination information indicating whether or not the target information includes the phrase. The information processing apparatus according to Configuration Example 1. (Configuration Example 3) The processing unit Using the ratio of the number of pieces of target information among the pieces of target information ranked higher than the rank of the target information for which the confidence level is calculated, among which the determination information indicates that the target information includes the phrase, the confidence level is calculated for the rank of the target information for which the confidence level is calculated. The information processing apparatus according to Configuration Example 2. (Configuration Example 4) The processing unit The confidence level is calculated using the rank based on the similarity and determination information indicating whether or not the target information has been selected as information related to the phrase. The information processing apparatus according to Configuration Example 1. (Configuration Example 5) The processing unit calculates the score, which is a value that maintains the magnitude relationship of the similarity within the range. The information processing apparatus according to any one of Configuration Examples 1 to 4. (Configuration Example 6) The processing unit calculates the score obtained by converting the similarity by linear interpolation or non-linear interpolation. The information processing apparatus according to Configuration Example 5. (Configuration Example 7) The logical operator includes an AND operator, The processing unit calculates, as the conversion score for the two clauses, a value close to the smaller score among the two scores calculated for each of the two clauses to which the AND operator is applied. The information processing apparatus according to any one of Configuration Examples 1 to 6. (Configuration Example 8) The processing unit calculates, as the conversion score, the value of the smaller score among the two scores calculated for each of the two clauses to which the AND operator is applied. The information processing apparatus according to Configuration Example 7. (Configuration Example 9) The logical operator includes an OR operator, The processing unit calculates, as the conversion score for the two clauses, a value close to the larger score among the two scores calculated for each of the two clauses to which the OR operator is applied. The information processing apparatus according to any one of Configuration Examples 1 to 8. (Configuration Example 10) The processing unit calculates, as the conversion score, the value of the larger score among the two scores calculated for each of the two clauses to which the OR operator is applied. The information processing apparatus according to Configuration Example 9. (Configuration Example 11) The logical operator includes a NOT operator, The processing unit calculates the conversion score such that the larger the value of the score calculated for the clause to which the NOT operator is applied, the smaller the value. The information processing apparatus according to any one of Configuration Examples 1 to 10. (Configuration Example 12) The processing unit outputs one or more pieces of the target information selected according to the conversion score. The information processing apparatus according to Configuration Example 1. (Configuration Example 13) The processing unit further outputs the confidence level determined based on the conversion score. The information processing apparatus according to Configuration Example 12. (Configuration Example 14) The processing unit creates the search condition based on one or more of the first clauses included in the target information that is the search result and one or more of the second clauses not included in the target information that is the search result, using an input screen including a first area that designates one or more first clauses and a second area that designates one or more second clauses. The information processing apparatus according to any one of Configuration Examples 1 to 13. (Configuration Example 15) The processing unit when the search condition includes a plurality of the logical operators, repeatedly performs a process of converting the score into the conversion score according to the order of application of the plurality of the logical operators. The information processing apparatus according to any one of Configuration Examples 1 to 14. (Configuration Example 16) The target information is at least one of document data and image data. The information processing apparatus according to any one of Configuration Examples 1 to 15. (Configuration Example 17) The processing unit includes a confidence level calculation unit that calculates the confidence level, a score calculation unit that calculates the score, and a conversion unit that outputs the conversion score. comprising The information processing apparatus according to any one of Configuration Examples 1 to 16. (Configuration Example 18) An information processing method executed by an information processing apparatus, For each of one or more words included in a search condition including one or more words and one or more logical operators, calculating a confidence level represented by a discrete value, which represents the probability that each of a plurality of target information items to be searched and the word are related; For each of one or more words, calculating a score, which is a continuous value obtained by converting the similarity between the target information and the word so as to be included within a range determined according to the confidence level; Converting the score calculated for each of one or more words into a conversion score according to a conversion method determined according to the logical operator; An information processing method including the above. (Configuration Example 19) Causing a computer to For each of one or more words included in a search condition including one or more words and one or more logical operators, calculating a confidence level represented by a discrete value, which represents the probability that each of a plurality of target information items to be searched and the word are related; For each of one or more words, calculating a score, which is a continuous value obtained by converting the similarity between the target information and the word so as to be included within a range determined according to the confidence level; Converting the score calculated for each of one or more words into a conversion score according to a conversion method determined according to the logical operator; A program for causing the above to be executed.

[0118] Although some embodiments of the present invention have been described, these embodiments are presented by way of example and are not intended to limit the scope of the invention. These novel embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, and are included in the invention described in the claims and the equivalent scope thereof.

Explanation of Signs

[0119] 100, 100-2, 100-3 Information processing apparatus 101 Reception unit 102 Similarity calculation unit 103, 103-2 Confidence calculation unit 104 Score calculation unit 105 Conversion unit 106 Output control unit 107-3 Creation unit 120, 120-2 Storage unit 121 Document DB 122 Transposed index DB 123 Word vector DB 124 Document vector DB 125 Similarity DB 126 Confidence DB 127 Score DB 128-2 Feedback DB

Claims

1. For each of one or more of the words included in a search condition including one or more words and one or more logical operators, a confidence level representing the likelihood that each of a plurality of target information items to be searched is related to the word is calculated as a discrete value. For each of one or more of the words, a score is calculated as a continuous value obtained by converting the similarity between the target information and the word so as to be included within a range determined according to the confidence level. The scores calculated for each of one or more of the words are converted into converted scores according to a conversion method determined according to the logical operator. A processing unit An information processing apparatus comprising the same.

2. The processing unit For each of a plurality of the target information items, the confidence level is calculated using the rank based on the similarity and determination information indicating whether the target information includes the word. The information processing apparatus according to Claim 1.

3. The processing unit Using the ratio of the number of the target information items at ranks equal to or higher than the rank of the target information for which the confidence level is calculated among the target information items at ranks equal to or higher than the rank of the target information for which the confidence level is calculated, where the determination information indicates that the target information includes the word, the confidence level is calculated. The information processing apparatus according to Claim 2.

4. The processing unit Using the rank based on the similarity and determination information indicating whether the target information has been selected as information related to the word, the confidence level is calculated. The information processing apparatus according to Claim 1.

5. The processing unit Within the range, the score is calculated as a value maintaining the magnitude relationship of the similarity. The information processing apparatus according to Claim 1.

6. The processing unit Calculating the score obtained by converting the similarity by linear interpolation or non - linear interpolation, The information processing apparatus according to claim 5.

7. The logical operator includes an AND operator, The processing unit, Calculates, as the conversion score for the two statements, a value close to the smaller score among the two scores calculated for each of the two statements to which the AND operator is applied. The information processing apparatus according to claim 1.

8. The processing unit, Calculates, as the conversion score, the value of the smaller score among the two scores calculated for each of the two statements to which the AND operator is applied. The information processing apparatus according to claim 7.

9. The logical operator includes an OR operator, The processing unit, Calculates, as the conversion score for the two statements, a value close to the larger score among the two scores calculated for each of the two statements to which the OR operator is applied. The information processing apparatus according to claim 1.

10. The processing unit, Calculates, as the conversion score, the value of the larger score among the two scores calculated for each of the two statements to which the OR operator is applied. The information processing apparatus according to claim 9.

11. The logical operator includes a NOT operator, The processing unit, Calculates a conversion score such that the larger the value of the score calculated for the statement to which the NOT operator is applied, the smaller the value. The information processing apparatus according to claim 1.

12. The processing unit, Output one or more pieces of the target information selected according to the conversion score. The information processing apparatus according to claim 1.

13. The processing unit Further output the confidence level determined based on the conversion score. The information processing apparatus according to claim 12.

14. The processing unit Based on one or more of the first phrases included in the target information that is the search result and one or more of the second phrases not included in the target information that is the search result, create the search condition using an input screen including a first area for designating the one or more first phrases and a second area for designating the one or more second phrases. The information processing apparatus according to claim 1.

15. The processing unit When the search condition includes a plurality of the logical operators, repeat the process of converting the score into the conversion score according to the order of application of the plurality of the logical operators. The information processing apparatus according to claim 1.

16. The target information is at least one of document data and image data. The information processing apparatus according to claim 1.

17. The processing unit A confidence level calculation unit that calculates the confidence level A score calculation unit that calculates the score A conversion unit that outputs the conversion score Comprising The information processing apparatus according to claim 1.

18. An information processing method executed by an information processing apparatus, For each of one or more of the terms included in a search condition including one or more terms and one or more logical operators, calculating a confidence level represented by a discrete value, which represents the probability that each of a plurality of target information items to be searched and the term are related; For each of one or more of the terms, calculating a score, which is a continuous value obtained by converting the similarity between the target information and the term so as to be included within a range determined according to the confidence level; Converting the scores calculated for each of one or more of the terms into conversion scores according to a conversion method determined according to the logical operator; An information processing method including the above.

19. Causing a computer to For each of one or more of the terms included in a search condition including one or more terms and one or more logical operators, calculating a confidence level represented by a discrete value, which represents the probability that each of a plurality of target information items to be searched and the term are related; For each of one or more of the terms, calculating a score, which is a continuous value obtained by converting the similarity between the target information and the term so as to be included within a range determined according to the confidence level; Converting the scores calculated for each of one or more of the terms into conversion scores according to a conversion method determined according to the logical operator; A program for causing the above to be executed.

Citation Information

Patent Citations

  • Text retrieval device

    JP1999232303A

  • Information processing device and method therefor, storage medium, and program

    JP2006040085A

  • Multilingual document retrieval system and method using semantic vector matching

    US6006221A

  • Retrieval system, retrieval method and program

    JP2022126131A

  • Document searching system, document searching method, program, and non-transitory computer readable storage medium

    WO2019180546A1