Teacher data collection system, similarity score calculation system, literature search system, and teacher data collection program
The teacher data collection system addresses the challenge of localized searches by using feature vector generation, dimensionality reduction, and grid classification to ensure comprehensive coverage of related documents, resulting in improved similarity score calculations and search results.
Patent Information
- Application Number
- JP2021123804
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-07-29
- Publication Date
- 2025-06-05
- Estimated Expiration
- 2041-07-29
AI Technical Summary
Existing systems for collecting teacher data and calculating similarity scores often fail to cover a wide range of related documents, leading to localized searches and potential omissions in document retrieval, particularly in fields like patent search where comprehensive coverage is crucial.
A teacher data collection system that generates feature vectors for search queries and documents, applies dimensionality reduction, calculates cosine similarity, and classifies documents into grid regions to select a diverse set of teacher data, ensuring comprehensive coverage of related literature.
The system effectively collects teacher data that covers a wide range of literature, allowing for accurate and comprehensive similarity score calculations and search results presentation, thereby improving the efficiency and effectiveness of document retrieval.
Smart Images

Figure 0007688823000001 
Figure 0007688823000002 
Figure 0007688823000003
Abstract
Description
[Technical field]
[0001] The present invention relates to a teacher data collection system, a similarity score calculation system, a literature search system, and a teacher data collection program. [Background technology]
[0002] A document processing device derives document vectors of target documents, and derives cosine values (cosine similarity) of the document vectors as an index of similarity between documents (see, for example, Patent Document 1).
[0003] When collecting literature in a specific field as teacher data for machine learning, a certain teacher data collection device (a) derives feature vectors based on the frequency of occurrence of words in the reference data and the collected data, (b) derives cosine similarity between the feature vectors of the reference data and the collected data, and (c) extracts collected data whose cosine similarity falls within a specified range as teacher data (see, for example, Patent Document 2). [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 11-53396 [Patent Document 2] JP 2018-124617 A Summary of the Invention [Problem to be solved by the invention]
[0005] When searching a population of documents to find documents similar to a reference document or a search query, there are cases where you want to find many related documents, and cases where you want to find only documents with a high matching rate (highly accurate answers). In the field of patent search, searches for prior art documents are the former, and searches for patent invalidation documents are the latter.
[0006] In the former case, in order to prevent search omissions, it is preferable to cover the entire range of documents related to the reference document in the population. However, in the above-mentioned training data collection device, the training data is limited by the cosine similarity of feature vectors based on the frequency of word occurrences, and although the amount of calculation during a search is reduced, the search is localized, making it difficult to cover the entire range of documents related to the reference document in the population.
[0007] The present invention has been made in consideration of the above-mentioned problems, and aims to provide a teacher data collection system and teacher data collection program that collect teacher data that covers related literature over a wide range of a population, a similarity score calculation system that appropriately calculates similarity scores of related literature covered over a wide range of a population, and a literature search system that presents search results in an appropriate order from related literature covered over a wide range of a population. [Means for solving the problem]
[0008] The teacher data collection system according to the present invention includes a vector generation unit that derives a feature vector of a search formula as a reference feature vector and derives a feature vector of a document belonging to a population as a document feature vector; a feature extraction unit that (a) executes a dimensionality reduction process to reduce the number of dimensions of the reference feature vector and the document feature vector, and sets a dimension value obtained by the dimensionality reduction process for the reference feature vector and the document feature vector as a first feature, and (b) derives a cosine similarity between the reference feature vector and the document feature vector as a second feature; a grid division unit that classifies the documents into a predetermined first number of first partial regions obtained by dividing a feature space of the first feature, and classifies the documents into a predetermined second number of second partial regions obtained by dividing a value range of the second feature; and a teacher data extraction unit that (a) selects, for each combination of one of the first partial regions and one of the second partial regions, at least one of the documents classified into the first partial region and the second partial region, and (b) sets the documents selected for all of the combinations as teacher data.
[0009] The similarity score calculation system of the present invention comprises the above-mentioned teacher data collection system, a similarity score calculation unit that calculates similarity scores of documents in the teacher data, and a machine learning processing unit that performs machine learning on the similarity score calculation unit using the teacher data.
[0010] The document search system of the present invention comprises the above-mentioned similarity score calculation system, a search condition input unit that specifies the search formula, and a search result display unit that sorts the documents extracted as the training data in order of highest similarity score and displays combinations of the documents and their similarity scores on a display device.
[0011] The teacher data collection program of the present invention causes a computer to function as: a vector generation unit that derives a feature vector of a search formula as a reference feature vector, and derives a feature vector of a document belonging to a population as a document feature vector; (a) a feature extraction unit that executes a dimensionality reduction process to reduce the number of dimensions of the reference feature vector and the document feature vector, and sets a dimension value obtained by the dimensionality reduction process for the reference feature vector and the document feature vector as a first feature value; (b) a feature extraction unit that derives a cosine similarity between the reference feature vector and the document feature vector as a second feature value; a grid division unit that classifies the documents into a predetermined first number of first partial regions obtained by dividing a feature space of the first feature value, and classifies the documents into a predetermined second number of second partial regions obtained by dividing a value range of the second feature value; and a teacher data extraction unit that (a) selects, for each combination of one of the first partial regions and one of the second partial regions, at least one of the documents classified into the first partial region and the second partial region, and (b) sets the documents selected for all of the combinations as teacher data. Effect of the Invention
[0012] According to the present invention, there is provided a teacher data collection system and teacher data collection program that collect teacher data that covers a wide range of literature within a population, a similarity score calculation system that appropriately calculates similarity scores of related literature that is covered within a wide range of a population, and a literature search system that presents search results in an appropriate order of related literature that is covered within a wide range of a population.
[0013] The above and other objects, features and advantages of the present invention will become more apparent from the following detailed description taken in conjunction with the accompanying drawings. [Brief description of the drawings]
[0014] [Figure 1] FIG. 1 is a block diagram showing a configuration of a document search system according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a diagram showing an example of a first partial region obtained by dividing a feature space of the first feature (first and second principal components). [Diagram 3] FIG. 3 is a diagram illustrating the document search result according to the embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0016] Fig. 1 is a block diagram showing the configuration of a document search system according to an embodiment of the present invention. The document search system shown in Fig. 1 includes a processor 1 as a computer, a non-volatile memory device 2, an input device 3 that detects user operations, and a display device 4 that displays various information to the user.
[0017] The arithmetic processing device 1 is equipped with a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), etc., and functions as various processing units by loading programs stored in the RAM or the storage device 2 into the RAM and executing them with the CPU.
[0018] Here, the processor 1 functions as a document search system 11 by executing a document search program 2 a in the storage device 2 .
[0019] The document search system 11 includes a similarity score calculation system 21, a search condition input unit 22, and a search result display unit .
[0020] The similarity score calculation system 21 includes a teacher data collection system 31, a machine learning processing unit 32, and a similarity score calculation unit 33.
[0021] The teacher data collection system 31 includes a document acquisition unit 41, a morphological analysis unit 42, a vector generation unit 43, a feature extraction unit 44, a grid division unit 45, a teacher data extraction unit 46, a teacher data determination unit 47, and a teacher data balance control unit 48.
[0022] The document acquisition unit 41 performs a document search using a search formula specified by the user, and acquires documents (document data) found in the document search as a population. For example, the document acquisition unit 41 uses a communication device (not shown) to access a server on a network, causes the server to execute the above-mentioned document search, acquires the above-mentioned population from the server, and stores it in the storage device 2 or the like.
[0023] The morpheme analysis unit 42 extracts the morphemes of the above-mentioned search formula and the morphemes of each document belonging to the population by a known method. Note that as the morphemes, parts of speech (nouns only, nouns and adjectives, etc.) designated in advance are designated.
[0024] The vector generation unit 43 derives the feature vector of the above-mentioned search formula as a reference feature vector, and also derives the feature vector of the document belonging to the population (hereinafter referred to as population document) as a document feature vector. Here, for example, the vector generation unit 43 generates feature vectors of morphemes in the search formula and population document, and generates feature vectors of the search formula and population document from the feature vectors of morphemes in the search formula and population document.
[0025] For example, feature vectors are derived using count-based methods (such as TF (Term Frequency)-IDF (Inverse Document Frequency)) or distributed representation methods (such as Word2vec and BERT (Bidirectional Encoder Representations from Transformers)). In the case of count-based methods, feature vectors (of a document) are generated based on the number of times a word appears, and in the case of distributed representation methods, the sum or average of feature vectors of words in a document is derived as the feature vector (of a document).
[0026] The feature extraction unit 44 (a) performs a dimensionality reduction process to reduce the number of dimensions of the reference feature vector and the document feature vector, and sets the dimension value obtained by the dimensionality reduction process on the reference feature vector and the document feature vector as a first feature, and (b) derives the cosine similarity between the reference feature vector and the document feature vector as a second feature.
[0027] For example, dimensionality reduction processing is performed using principal component analysis (PCA), singular value decomposition (SVD), t-distributed stochastic neighbor embedding (t-SNE), or the like.
[0028] The values of the first and second features for the population document are stored, for example, in the storage device 2 in association with the population document, and are read out as necessary by downstream processing units such as the grid division unit 45 and the teacher data extraction unit 46.
[0029] The grid division unit 45 classifies the population documents into a predetermined first number (plural, for example, 16) of first partial regions obtained by dividing the feature space of the above-mentioned first feature, and also classifies the population documents into a predetermined second number (plural, for example, 3) of second partial regions obtained by dividing the value range of the above-mentioned second feature.
[0030] FIG. 2 is a diagram showing an example of first partial regions obtained by dividing the feature space of the first feature (first and second principal components). For example, in the case shown in FIG. 2, a positive boundary value, 0, and a negative boundary value are set for the first principal component, and a positive boundary value, 0, and a negative boundary value are set for the second principal component, and the first quadrant of the two-dimensional feature space is divided into four first partial regions #1-1 to #1-4, the second quadrant is divided into four first partial regions #2-1 to #2-4, the third quadrant is divided into four first partial regions #3-1 to #3-4, and the fourth quadrant is divided into four first partial regions #4-1 to #4-4. As a result, for example, in the case shown in FIG. 2, the feature space of the first feature is divided into 16 first partial regions.
[0031] In addition, the value range (range of 0 to 1) of the cosine similarity as the second feature amount is divided into three second partial regions. Here, based on the average value μ and standard deviation σ of the cosine similarity of the population documents, the second partial regions are set to the following three ranges: a range where the cosine similarity is less than (μ-σ), a range where the cosine similarity is equal to or greater than (μ-σ) and less than (μ+σ), and a range where the cosine similarity is equal to or greater than (μ+σ).
[0032] The teacher data extraction unit 46 (a) selects at least one (e.g., three) population documents classified into the first partial domain and the second partial domain for each combination of one first partial domain and one second partial domain, and (b) sets the selected population documents for all combinations of the first partial domain and the second partial domain as teacher data.
[0033] For example, as shown in Figure 2, if the feature space is divided into 16 first sub-regions, the value range is divided into three second sub-regions, and three documents are selected for each combination, 144 (=16 × 3 × 3) population documents are extracted as training data.
[0034] The teacher data judgment unit 47 performs annotation for each of the population documents of the teacher data described above, indicating whether or not the document is a desired document, and as an annotation result, a flag indicating whether or not the document is a desired document is stored in the memory device 2 or the like, and is read out as necessary by a subsequent processing unit.
[0035] Specifically, the teacher data judgment unit 47 has the user judge whether each document extracted as teacher data by the teacher data extraction unit 46 is a desired document or not, and obtains the user's judgment result (desired document or non-desired document). For example, the teacher data judgment unit 47 displays a list of documents extracted as teacher data by the teacher data extraction unit 46 on the display device 4, and identifies whether each document in the list is a desired document of the user based on a user operation detected by the input device 3, thereby obtaining the user's judgment result. When the teacher data judgment unit 47 is used, a similarity score (e.g., 1 or 0) corresponding to the document is set to a value corresponding to the judgment result.
[0036] The teacher data balance control unit 48 performs balancing processing after annotation, and in the balancing processing, thins out the documents in the teacher data so that the ratio of the number of desired documents to the number of undesired documents in the teacher data mentioned above satisfies a specified condition.
[0037] For example, if the ratio (>1) of the number of desired documents to the number of undesired documents is greater than or equal to a predetermined threshold (e.g., 1.3), the desired or undesired documents in the training data are thinned out so that at least one document remains in each of the above combinations, so that the ratio is less than the predetermined threshold.
[0038] The machine learning processing unit 32 executes machine learning in the similarity score calculation unit 33 using the above-mentioned training data (after the balancing process) (the combination of the extracted documents and their similarity scores given by the annotations).
[0039] The similarity score calculation unit 33 is a processing unit (such as a classifier) capable of machine learning, and calculates similarity scores of documents in the above-mentioned training data (after balancing processing).
[0040] This processing unit is, for example, a Support Vector Machine (SVM), a Naive Bayes classifier, a Random Forest learner, a Convolutional Neural Network, etc., and the machine learning processing unit 32 performs machine learning using teacher data, etc., with a machine learning method corresponding to the type of this processing unit.
[0041] The search condition input unit 22 uses the input device 3 and the display device 4 as a user interface to identify the above-mentioned search formula based on the detected user operation, and specifies it to the teacher data collection system 31.
[0042] The search result display unit 23 obtains the similarity scores calculated by the similarity score calculation unit 33, sorts the documents extracted as the above-mentioned teacher data in descending order of similarity score, and displays combinations of the documents and their similarity scores on the display device 4, for example as a list, in descending order of similarity score.
[0043] Furthermore, in this embodiment, the search result display unit 23 (a) detects, via the input device 3, a user operation indicating whether or not the document displayed on the display device 4 is a desired document, and (b) derives a desired document appearance rate indicating the proportion of the desired document (e.g., a moving average for the most recent predetermined number of documents) based on the user operation for the most recent predetermined number of documents displayed, and displays it on the display device 4. As the number of documents displayed increases, the desired document appearance rate decreases, so the user can refer to the current desired document appearance rate and terminate the display of the search results (e.g., when the desired document appearance rate drops to 1%).
[0044] Next, the operation of the above system will be described.
[0045] First, the search condition input unit 22 identifies a search formula desired by the user based on a user operation, and specifies it to the teacher data collection system 31. In the teacher data collection system 31, the document acquisition unit 41 acquires population documents based on the search formula specified by the user.
[0046] Next, the morpheme analysis unit 42 extracts the morphemes of the above-mentioned search formula and the morphemes of each population document. The vector generation unit 43 derives the feature vector of the above-mentioned search formula as a reference feature vector, and derives the feature vector of the population document as a document feature vector.
[0047] Then, the feature extraction unit 44 (a) performs a dimensionality reduction process to reduce the number of dimensions of the reference feature vector and the document feature vector, and sets the dimension value obtained by the dimensionality reduction process on the reference feature vector and the document feature vector as a first feature, and (b) derives the cosine similarity between the reference feature vector and the document feature vector as a second feature.
[0048] The grid division unit 45 classifies the population documents into first partial regions obtained by dividing the feature space of the above-mentioned first feature, and also classifies the population documents into second partial regions obtained by dividing the value range of the above-mentioned second feature. Then, the teacher data extraction unit 46 (a) selects at least one population document classified into the first partial region and the second partial region for each combination of one first partial region and one second partial region, and (b) sets the selected population documents for all combinations of the first partial region and the second partial region as teacher data.
[0049] The teacher data determination unit 47 performs annotation for each of the population documents of the teacher data described above, indicating whether or not the document is a desired document.
[0050] After annotation, the teacher data balance control unit 48 performs a balancing process, in which the documents in the teacher data are thinned out so that the ratio between the number of desired documents and the number of undesired documents in the teacher data satisfies a specified condition.
[0051] In this manner, training data that is widely distributed in the feature space of the first feature and in the value range of the second feature is generated.
[0052] Then, the machine learning processing unit 32 uses the teacher data to execute machine learning in the similarity score calculation unit 33. After the machine learning, the similarity score calculation unit 33 calculates a similarity score for each document in the teacher data extracted by the teacher data extraction unit 46.
[0053] In this manner, a similarity score is calculated for each document in the population, indicating the degree to which it matches the query.
[0054] Then, the search result display unit 23 acquires the similarity scores, and displays the documents extracted as teaching data on the display device 4 in descending order of similarity scores.
[0055] As described above, according to the above embodiment, the vector generation unit 43 derives the feature vector of the search formula as the reference feature vector, and derives the feature vector of the document belonging to the population as the document feature vector. The feature extraction unit 44 (a) executes a dimensionality reduction process to reduce the number of dimensions of the reference feature vector and the document feature vector, and sets the dimension value obtained by the dimensionality reduction process for the reference feature vector and the document feature vector as the first feature, and (b) derives the cosine similarity between the reference feature vector and the document feature vector as the second feature. The grid division unit 45 classifies documents into a first predetermined number of first partial regions obtained by dividing the feature space of the first feature, and classifies documents into a second predetermined number of second partial regions obtained by dividing the value range of the second feature. The teacher data extraction unit 46 (a) selects at least one document classified into the first partial region and the second partial region for each combination of one first partial region and one second partial region, and (b) sets the selected documents for all combinations as teacher data.
[0056] This allows training data that covers a wide range of literature in the population in a balanced manner to be automatically collected, and related literature that covers a wide range of the population is presented as search results in an appropriate order.
[0057] Fig. 3 is a diagram for explaining the results of a document search according to the embodiment. When verified as shown in Fig. 3, the document search according to the embodiment finds most of the desired documents from the population (35,000 cases) at an early stage, compared to the existing method.
[0058] It should be noted that various changes and modifications to the above-described embodiments will be apparent to those skilled in the art. Such changes and modifications may be made without departing from the spirit and scope of the subject matter and without diminishing its intended advantages. In other words, such changes and modifications are intended to be included within the scope of the claims.
[0059] For example, in the above embodiment, the document search program 2a may be stored in a computer-readable recording medium and installed in the storage device 2 from the recording medium. [Industrial Applicability]
[0060] The present invention is applicable, for example, to literature searching. [Explanation of symbols]
[0061] 1. Processing unit (an example of a computer) 2a Literature search program (an example of a teacher data collection program) 11 Literature Search System 21 Similarity score calculation system 23 Search results display section 31 Teacher Data Collection System 32 Machine learning processing section 33 Similarity score calculation unit 43 Vector Generation Unit 44 Feature Extraction Unit 45 Grid division section 46 Teacher Data Extraction Unit 47 Teacher Data Judgment Unit 48 Teacher Data Balance Control Section
Claims
1. a vector generation unit that derives a feature vector of the search query as a reference feature vector and derives feature vectors of documents belonging to the population as document feature vectors; (a) a feature extraction unit that performs a dimension reduction process to reduce the number of dimensions of the reference feature vector and the document feature vector, and sets a dimension value obtained by the dimension reduction process for the reference feature vector and the document feature vector as a first feature amount, and (b) derives a cosine similarity between the reference feature vector and the document feature vector as a second feature amount; a grid division unit that classifies the documents into a first predetermined number of first partial regions obtained by dividing a feature amount space of the first feature amount and classifies the documents into a second predetermined number of second partial regions obtained by dividing a value range of the second feature amount; (a) a teacher data extraction unit that selects, for each combination of one of the first partial regions and one of the second partial regions, at least one of the documents classified into the first partial region and the second partial region; and (b) sets the documents selected for all of the combinations as teacher data; A teacher data collection system comprising:
2. The teacher data collection system according to claim 1, further comprising a teacher data determination unit that performs annotation for each of the documents in the teacher data to indicate whether the document is a desired document or not.
3. The teacher data collection system described in claim 2, further comprising a teacher data balance control unit that thins out the documents so that the ratio of the number of desired documents to the number of documents that are not desired documents in the teacher data satisfies a predetermined condition.
4. A teacher data collection system according to any one of claims 1 to 3, a similarity score calculation unit for calculating a similarity score of the document of the training data; a machine learning processing unit that executes machine learning of the similarity score calculation unit using the teacher data; A similarity score calculation system comprising:
5. A similarity score calculation system according to claim 4, A search condition input section for specifying the search formula; a search result display unit that sorts the documents extracted as the teacher data in descending order of the similarity score and displays a combination of the documents and the similarity score of the documents on a display device; A document search system comprising:
6. The document search system of claim 5, wherein the search result display unit (a) detects a user operation using an input device indicating whether the document displayed on the display device is a desired document, and (b) derives a desired document appearance rate indicating the proportion of the desired document based on the user operation for a predetermined number of the most recent documents displayed, and displays the desired document appearance rate on the display device.
7. Computer, a vector generation unit that derives a feature vector of the search query as a reference feature vector and derives feature vectors of documents belonging to the population as document feature vectors; (a) a feature extraction unit that executes a dimension reduction process to reduce the number of dimensions of the reference feature vector and the document feature vector, and sets a dimension value obtained by the dimension reduction process for the reference feature vector and the document feature vector as a first feature amount, and (b) derives a cosine similarity between the reference feature vector and the document feature vector as a second feature amount; a grid division unit that classifies the documents into a first predetermined number of first partial regions obtained by dividing a feature amount space of the first feature amount and classifies the documents into a second predetermined number of second partial regions obtained by dividing a value range of the second feature amount; and (a) a training data extraction unit that selects, for each combination of one of the first partial regions and one of the second partial regions, at least one of the documents classified into the first partial region and the second partial region; and (b) sets the selected documents for all of the combinations as training data. A teacher data collection program that serves as a.
Citation Information
Patent Citations
Device and method for document processing and storage medium storing document processing program
JP1999053396A
Information retrieval unit
JP2002245067A
Teacher data collection apparatus, teacher data collection method and program
JP2018124617A
Patent information processing apparatus, patent information processing method, and program
JP2019087006A
Machine learning program, machine learning method, and machine learning apparatus
JP2020190935A