File search support method, program, and file search support system
By classifying and vector analysis of multiple files, high-precision search formulas are generated, which solves the problem of low file search accuracy in the prior art, and achieves the goal of efficiently obtaining information by users.
Patent Information
- Application Number
- CN202380079410.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-24
- Filing Date
- 2023-11-17
- Publication Date
- 2025-06-24
AI Technical Summary
It is difficult to achieve high-precision search in file retrieval, especially in the search of intellectual property-related documents. Users need to have high skills to generate appropriate search queries, and there are problems of intentional operation or guidance of search results.
By obtaining the data of multiple files, classifying them into required files and non-demand files, using vectors based on data elements and judgment labels for classifiers, extracting high-important elements to generate search formulas, and using these search formulas for file search to achieve high-precision file search.
It enables users to obtain the required information efficiently, simplifies the file retrieval process, reduces the workload and time of user-generated search, and improves the search accuracy, especially in the search of intellectual property-related documents.
Smart Images

Figure CN120202467A_ABST
Abstract
Description
Technical Field
[0001] One aspect of the present invention relates to a method for supporting document retrieval. Another aspect of the present invention relates to a program capable of supporting document retrieval. Another aspect of the present invention relates to a document retrieval support system. Background Art
[0002] Examples of patent-related operations include prior art searches, patent right establishment, and invalidation data searches. By conducting a prior art search on an invention before filing, it is possible to confirm the existence of relevant intellectual property rights. Patent documents and papers at home and abroad obtained through prior art searches can be used to confirm the novelty and inventiveness of the invention and to determine whether to apply for a patent. In addition, by conducting an invalidation data search of patent documents, it is possible to confirm whether there is a risk of invalidation of the patent rights held by oneself or whether it is possible to invalidate the patent rights held by others.
[0003] Due to the large number of patent-related operations, systems for supporting patent-related operations such as patent application document production support systems, patent information analysis systems, and patent retrieval systems have been developed in recent years. Patent Document 1 discloses a patent document retrieval technique that combines keyword retrieval and similarity retrieval. [Prior Art Document] [Patent Document]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2018-73309 Summary of the Invention Technical Problem to be Solved by the Invention
[0005] In order to retrieve a required document using text retrieval, it is necessary to select appropriate keywords. In order to improve the retrieval accuracy, it is preferable to not only use keywords related to the required document but also specify keywords for removing unnecessary documents. In addition, it is preferable to also consider synonyms (e.g., synonyms, near-synonyms) of the selected keywords. Therefore, in order to generate an appropriate retrieval query, continuous exploration is required, and the staff needs to have high skills.
[0006] In addition, the retrieval results obtained by a mechanism that can perform retrieval using a simple retrieval query such as a web search engine may be subject to intentional manipulation or guidance using algorithms such as page ranking, and there is a concern about a decrease in objectivity.
[0007] In addition, sometimes the classification mechanism is based on documents. As an example, the classification of the accompanying drawings and patent classification can be cited. Specifically, patent documents are given patent classifications such as CPC (Cooperative Patent Classification), IPC (International Patent Classification), FI (File Index), and F-term (File Forming Term). However, the number of patent classification items is very large, and sometimes it is difficult to select appropriate items to retrieve the required patents.
[0008] One of the objects of one embodiment of the present invention is to provide a user with a suitable search formula for retrieving a required document group. In addition, one of the objects of one embodiment of the present invention is to provide a document retrieval support system that is easy for users to operate. In addition, one of the objects of one embodiment of the present invention is to provide a document retrieval support method or a document retrieval support system that enables users to efficiently obtain the required information. In addition, one of the objects of one embodiment of the present invention is to achieve high-precision document retrieval with a simple input method, especially to achieve the retrieval of intellectual property-related documents.
[0009] Note that the description of these objects does not preclude the existence of other objects. One embodiment of the present invention does not need to achieve all of the above objects. Other objects than the above can be extracted from the descriptions in the specification, the drawings, and the claims. Means for Solving Technical Problems
[0010] One aspect of the present invention is a method for supporting document retrieval, comprising the following steps: obtaining data of a plurality of documents; receiving, as required documents, some of the plurality of documents; classifying the plurality of documents into a first document group including the required documents and a second document group including the remaining documents; using, as learning data, vectors based on elements in the data and determination labels indicating whether the documents are required documents in each of the plurality of documents to train a classifier; extracting, by analyzing the classifier, two or more highly important elements from the elements; classifying the extracted elements into a first group included in the first document group and a second group not included in the first document group; generating first search terms within a range of two or more and not more than twice the number of elements in the first group using the elements in the first group, and generating a first search formula such that 50% or more of the documents in the first document group include at least one first search term and 50% or more of the documents in the second document group do not include any first search term using the first search terms; generating second search terms within a range of two or more and not more than twice the number of elements in the second group using the elements in the second group, and generating a second search formula such that 50% or more of the documents in the second document group include at least one second search term using the second search terms; and outputting one or both of a third search formula generated using the first search formula and the second search formula and a result of retrieving the plurality of documents using the third search formula.
[0011] One aspect of the present invention is a method for supporting document retrieval, comprising the following steps: obtaining data of a plurality of documents; receiving, as required documents, some of the plurality of documents and receiving, as non-required documents, other parts of the documents; classifying the plurality of documents into a first document group including the required documents, a second document group including the non-required documents, and a third document group including the remaining documents; using, as learning data, vectors based on elements in the data and determination labels indicating whether the documents are required documents in each of the documents included in the first document group and the second document group to train a classifier; extracting, by analyzing the classifier, two or more highly important elements from the elements; classifying the extracted elements into a first group included in the first document group and a second group not included in the first document group; generating first search terms within a range of two or more and not more than twice the number of elements in the first group using the elements in the first group, and generating a first search formula such that 50% or more of the documents in the first document group include at least one first search term and 50% or more of the documents in the second document group do not include any first search term using the first search terms; generating second search terms within a range of two or more and not more than twice the number of elements in the second group using the elements in the second group, and generating a second search formula such that 50% or more of the documents in the second document group include at least one second search term using the second search terms; and outputting one or both of a third search formula generated using the first search formula and the second search formula and a result of retrieving the plurality of documents using the third search formula.
[0012] Preferably, the second search formula is generated after the first search formula is generated. At this time, preferably, the second search formula is generated in such a way that at least one second search term is included in more than 50% of the documents extracted from the second document group using the first search formula.
[0013] The classifier is preferably a classifier using a random forest.
[0014] The first search formula is preferably generated using a genetic algorithm.
[0015] The second search formula is preferably generated using a genetic algorithm.
[0016] The elements in the data are preferably words.
[0017] One aspect of the present invention is a program that has a function of causing a processor to execute any of the above-described file search support methods.
[0018] One aspect of the present invention is a file search support system that includes a receiving unit, a storage unit, a processing unit, and an output unit. The receiving unit has a function of receiving a selection of a required document. The storage unit includes a classifier. The processing unit has the following functions: generating vectors for a plurality of documents respectively based on elements in the data; assigning a determination label indicating whether a document is a required document to the plurality of documents according to the received selection by the receiving unit; using the vectors and the determination labels as learning data to perform learning of the classifier; extracting two or more highly important elements by analyzing the classifier; and generating a search formula using the highly important elements. The output unit has a function of outputting the search formula.
[0019] One aspect of the present invention is a document retrieval support system, which includes a receiving unit, a storage unit, a processing unit, and an output unit. The receiving unit has a function of receiving a selection of a required document. The storage unit includes a classifier. The processing unit has the following functions: classifying a plurality of documents into a first document group including the required document and a second document group including the remaining documents; generating a vector of a document based on elements in the data of the document; the receiving unit assigning a determination label of whether the document is a required document to the plurality of documents according to the received selection; using the vector and the determination label as learning data to perform learning of the classifier; extracting two or more highly important elements by analyzing the classifier; classifying the highly important elements into a first group included in the first document group and a second group not included in the first document group; generating a first search term within a range of two or more and not more than twice the number of elements in the first group using the elements in the first group, and generating a first search formula in such a way that 50% or more of the documents in the first document group include at least one first search term and 50% or more of the documents in the second document group do not include any first search term using the first search term; generating a second search term within a range of two or more and not more than twice the number of elements in the second group using the elements in the second group, and generating a second search formula in such a way that 50% or more of the documents in the second document group include at least one second search term using the second search term; and generating a third search formula using the first search formula and the second search formula, and the output unit has a function of outputting the third search formula.
[0020] One aspect of the present invention is a document retrieval support system, which includes a receiving unit, a storage unit, a processing unit, and an output unit. The receiving unit has a function of receiving the selection of required documents and non-required documents. The storage unit includes a classifier. The processing unit has the following functions: classifying a plurality of documents into a first document group including required documents, a second document group including non-required documents, and a third document group including the remaining documents; generating a vector of a document based on elements in the data of the document; the receiving unit assigns a determination label indicating whether a document is a required document to the plurality of documents according to the received selection; using the vector and the determination label as learning data to perform learning of the classifier; extracting two or more elements with high importance by analyzing the classifier; classifying the elements with high importance into a first group included in the first document group and a second group not included in the first document group; generating a first search term within a range of two or more and not more than twice the number of elements in the first group using the elements in the first group, and generating a first search formula in such a way that 50% or more of the documents in the first document group include at least one first search term and 50% or more of the documents in the second document group do not include any first search term using the first search term; generating a second search term within a range of two or more and not more than twice the number of elements in the second group using the elements in the second group, and generating a second search formula in such a way that 50% or more of the documents in the second document group include at least one second search term using the second search term; and generating a third search formula using the first search formula and the second search formula, and the output unit has a function of outputting the third search formula.
[0021] Preferably, the processing unit has a function of retrieving a plurality of documents using the third search formula. Preferably, the output unit has a function of outputting the result of the retrieved documents. Advantages of the Invention
[0022] According to one aspect of the present invention, a suitable search formula can be provided to a user to retrieve a required document group. According to one aspect of the present invention, a document retrieval support system that is easy for a user to operate can be provided. According to one aspect of the present invention, a document retrieval support method or a document retrieval support system that enables a user to efficiently obtain the required information can be provided. According to one aspect of the present invention, high-precision document retrieval can be achieved with a simple input method, and in particular, retrieval of documents related to intellectual property can be achieved.
[0023] Note that the description of these effects does not prevent the existence of other effects. One aspect of the present invention does not necessarily have all of the above effects. Effects other than the above can be extracted from the description of the specification, drawings, and claims. Brief Description of the Drawings
[0024] Figure 1 It is a diagram showing an example related to a document retrieval system. Figure 2This is a diagram showing an example of a document search support system. Figure 3 This is a diagram showing an example of a document search support method. Figure 4 This is a diagram showing an example of a document search support method. Figure 5A , Figure 5B , Figure 5C1 and Figure 5C2 This is a diagram for explaining an example of a document search support method. Figure 6 This is a diagram for explaining an example of a document search support method. Figure 7 This is a diagram showing an example of a document search support system. Figure 8 This is a diagram showing an example of a document search support system. Modes for Carrying Out the Invention
[0025] The embodiments are described in detail with reference to the accompanying drawings. Note that the present invention is not limited to the following description, and a person skilled in the art can easily understand that the mode and details can be transformed into various forms without departing from the purpose and scope of the present invention. Therefore, the present invention should not be interpreted as being limited to the contents described in the embodiments shown below.
[0026] Note that in the structure of the invention described below, the same symbols are used in common between different drawings to represent the same parts or parts with the same function, and their repeated descriptions are omitted. In addition, when representing parts with the same function, the same hatching is sometimes used without special additional symbols.
[0027] In addition, for ease of understanding, the position, size, range, etc. of each component shown in the drawings may not represent its actual position, size, range, etc. Therefore, the disclosed invention is not necessarily limited to the position, size, range, etc. disclosed in the drawings.
[0028] In this specification, etc., for the sake of convenience, ordinal numbers such as "first" and "second" are added, but they do not limit the number of components or the order of components (for example, process order or stacking order). In addition, the ordinal numbers added to components in one part of this specification may be inconsistent with the ordinal numbers added to the components in other parts of this specification or claims.
[0029] In this specification, unless otherwise specified, a document refers to a description of a phenomenon using natural language. A document is digitized and machine-readable. In addition, in this specification, a text includes one or more sentences.
[0030] (Embodiment 1) In this embodiment, with reference to Figure 1 FIG. 5, a document retrieval support system and a document retrieval support method according to one aspect of the present invention will be described.
[0031] A document retrieval support system according to one aspect of the present invention receives one or more of the documents (also referred to as required documents) included in the required document group that the user has already grasped, and uses the data of the received documents for learning a classifier. Then, by analyzing the classifier, two or more important elements are extracted to determine whether a document is a required document. Then, two or more of the extracted elements are used to generate a first search formula in such a way that the number of first search terms is small and the required documents include at least one first search term and most of the other documents do not include any first search terms, and a second search formula is formed in such a way that the number of second search terms is small and most of the other documents include at least one second search term. Then, a third search formula is generated using the two generated search formulas. As described above, a search formula for retrieving the required document group can be provided.
[0032] Specifically, a document retrieval support system according to one aspect of the present invention acquires data of a plurality of documents, receives a part of the plurality of documents as required documents, and classifies the plurality of documents into a first document group including the required documents and a second document group including the remaining documents. Alternatively, a document retrieval support system according to one aspect of the present invention may receive a part of the plurality of documents as required documents and receive another part of the documents as non-required documents. In this case, the plurality of documents are classified into a first document group including the required documents, a second document group including the non-required documents, and a third document group including the remaining documents.
[0033] Next, vectors based on elements in the data in each of the documents included in the first document group and the second document group and labels for determining whether the document is a required document are used as learning data for learning a classifier, and by analyzing the classifier, two or more highly important elements are extracted.
[0034] As the elements, one or more of the information included in the document itself and the information associated with the document can be used. For example, it is preferable to use one or both of the words included in the document text and the classification given to the document.
[0035] Next, classify the extracted elements into a first group included in the first document group and a second group not included in the first document group. Next, use the elements included in the first group to generate a first search formula in such a way that the number of first search terms is small and most of the first document group includes at least one first search term and most of the second document group does not include any first search terms. In addition, use the elements included in the second group to generate a second search formula in such a way that the number of second search terms is small and most of the second document group includes at least one second search term. And output one or both of the third search formula generated using the first search formula and the second search formula and the result of retrieving multiple documents using the third search formula.
[0036] As the first search term and the second search term, one can use an element itself and a combination of one or more logical operators (such as AND, OR, XOR, NOT, NAND, NOR) and two or more elements, respectively.
[0037] By retrieving using the first search formula, most of the required documents received by the user can be extracted. In addition, by retrieving using the second search formula, many non-required documents can be extracted. Therefore, for example, it is preferable to generate a third search formula that excludes the result of the second search formula (NOT search) from the result of the first search formula.
[0038] In addition, the document retrieval support system can also retrieve documents in the database using this search formula and output the result. In addition, the user can also input the output search formula into other search systems to perform document retrieval.
[0039] In this way, in the document retrieval support system according to one aspect of the present invention, a search formula for retrieving a required document group can be generated and provided based on a document that the user grasps as a required document. As a result, the user does not need to generate a search formula by himself, thereby reducing the workload and working time related to document retrieval. Therefore, the required information can be obtained efficiently.
[0040] As a method for providing the search formula and the search result, for example, one or both of the following methods can be performed: display on the display screen of the terminal used by the user; and output a file in CSV format or the like.
[0041] The document retrieval support system according to one aspect of the present invention can also have a document retrieval function. Note that the document retrieval support system according to one aspect of the present invention can also be a part of the function of a document retrieval system. Or, it can also be a system independent of the document retrieval system.
[0042] There are no particular restrictions on the documents received by the document retrieval support system according to one aspect of the present invention, and it can support the retrieval of various documents. As such documents, for example, patent application documents, books, magazines, newspapers, contracts, papers (including academic papers, dissertations, doctoral dissertations, short papers, journal papers, etc.), judgments, terms, product manuals, novels, publications, white papers, technical documents, and working documents can be cited. As patent application documents, more specifically, one or more of a specification, claims, and abstract of the specification can be cited.
[0043] Note that hereinafter, the patent application document is sometimes used as an example for the document for explanation.
[0044] <Document Retrieval Support System 1> Figure 1 Shows a system related to document retrieval including the document retrieval support system 100.
[0045] In Figure 1 the terminal 20 is connected to the document retrieval system 40 via the network 30a. In addition, the terminal 20 is connected to the document retrieval support system 100 via the network 30b.
[0046] The terminal 20 is an information terminal device such as a personal computer (PC) used by a user, and can also be called a client PC. In Figure 1 a notebook PC is shown as an example. In addition, as the terminal 20, a tablet PC, a desktop PC, a portable information terminal, etc. can be cited. Note that the terminal 20 is not limited to one, and there can also be multiple.
[0047] The document retrieval system 40 is a system capable of performing document retrieval. In Figure 1 a server computer capable of executing processing related to document retrieval is shown as an example. The document retrieval system 40 is not limited to a particular system, and existing systems, services, software, application programs, etc. can be used.
[0048] The document retrieval support system 100 is a system that can generate a retrieval formula using the document retrieval support method according to one aspect of the present invention. In Figure 1 a server computer capable of executing processing related to the document retrieval support method according to one aspect of the present invention is shown as an example.
[0049] As the networks 30a and 30b, computer networks such as the Internet, which is the basis of the World Wide Web (WWW), intranet, extranet, PAN (Personal Area Network), LAN (Local Area Network), CAN (Campus Area Network), MAN (Metropolitan Area Network), WAN (Wide Area Network), GAN (Global Area Network), etc. can be used. In addition, when performing wireless communication, communication standards such as the fourth-generation mobile communication system (4G), fifth-generation mobile communication system (5G), sixth-generation mobile communication system (6G), etc. or standards standardized by IEEE for communication, such as Wi-Fi (registered trademark), Bluetooth (registered trademark), etc. can be used as the communication protocol or communication technology.
[0050] Although the networks 30a and 30b are shown separately in Figure 1 they may also be the same. For example, both of the networks 30a and 30b may be the Internet. In addition, the network 30a and the network 30b may also be different from each other. For example, the file retrieval system 40 and the terminal 20 may be connected via the Internet (corresponding to the network 30a), and the file retrieval support system 100 and the terminal 20 may be connected via the company-internal LAN (corresponding to the network 30b).
[0051] For example, the user can utilize the file retrieval system 40 by using dedicated software or application programs installed on the terminal 20. Or, for example, the user can utilize the file retrieval system 40 from a web browser using the terminal 20.
[0052] Similarly, for example, the user can utilize the file retrieval support system 100 by using dedicated software or application programs installed on the terminal 20. Or, for example, the user can utilize the file retrieval support system 100 from a web browser using the terminal 20.
[0053] The user inputs information on one or more files that the user knows whether they are the required files into the file retrieval support system 100 (data) using the terminal 20. The file retrieval support system 100 generates a retrieval formula using elements included in the file data (for example, one or both of the words in the file and the classification given to the file), and outputs it to the terminal 20 (retrieval formula).
[0054] The user inputs, using the terminal 20, the search formula generated by the document search support system 100 into the document search system 40 (search formula). The document search system 40 outputs the search results obtained using this search formula to the terminal 20 (search results).
[0055] In the absence of the document search support system 100, the user needs to generate on their own the search formula to be input into the document search system 40. By using the document search support system 100, the user can reduce the work of generating the search formula, and thus can efficiently obtain the required information.
[0056] Note that Figure 1 It is shown that the document search support system 100 and the document search system 40 are different server computers, but the present invention is not limited to this. One server computer can also serve as both the document search support system 100 and the document search system 40. In addition, the document search support system according to one aspect of the present invention can also have a document search function. Therefore, the combination of the document search support system 100 and the document search system 40 can be regarded as the document search support system according to one aspect of the present invention. Further, a part or all of the functions of the document search support system 100 and the document search system 40 can also be provided in the terminal 20 used by the user.
[0057] Figure 2 A block diagram showing the document search support system 100 is presented. The document search support system 100 includes a receiving unit 110, a storage unit 120, a processing unit 130, an output unit 140, and a transmission channel 150.
[0058] In the drawings of this specification, a block diagram shows components classified according to their functions as independent boxes, but in reality, it is difficult to completely divide the actual components according to their functions, and one component may be involved in multiple functions. For example, a part of the processing unit 130 can also be used as the receiving unit 110. Further, one function may involve multiple components. For example, the processing performed in the processing unit 130 may be carried out on different servers depending on the processing.
[0059] [Receiving Unit 110] The receiving unit 110 receives while distinguishing whether the information of the document from the user is included in the required document group. For example, it can receive respectively the documents included in the required document group that the user already knows and the documents not included in the required document group.
[0060] When the document is included in a database or the like, the receiving unit 110 can receive the input of the information specifying the document.
[0061] As information on a designated document, the title of the document, the author (including the writer, the author of the manuscript, the author, etc.), and various identification numbers can be cited. In the case where the document is a patent application document, as information on the designated document, the application management number for identifying the application (including the number unique to the company), the application family management number for identifying the application family, the application number, the publication number, and the registration number can be cited.
[0062] In addition, it is also possible to receive the input of text data such as the body of the document. At this time, it is also possible to receive a document not included in the database as a required document or a non-required document.
[0063] The information on the document supplied to the receiving unit 110 is supplied to one or both of the storage unit 120 and the processing unit 130 through the transmission channel 150.
[0064] [Storage unit 120] The storage unit 120 has the function of storing the program executed by the processing unit 130. In addition, the storage unit 120 may also have the function of storing data generated by the processing unit 130 (such as calculation results, analysis results, inference results) and data input to the receiving unit 110.
[0065] Specifically, the storage unit 120 stores a classifier program that determines whether the document is a required document by inputting a vector of elements included in the data based on the document.
[0066] As algorithms that can be used for the classifier, neural networks, decision trees, Lasso regression, random forests, etc. can be cited.
[0067] In particular, when using a random forest or a decision tree, it is easy to calculate the importance of feature quantities, so it is preferred.
[0068] In addition, the storage unit 120 stores a program for generating a search formula.
[0069] In this program, combinatorial optimization processing is used to optimize the feature combination for generating the search formula. As algorithms that can be used for this program, local search methods, greedy algorithms, genetic algorithms, etc. can be cited.
[0070] In particular, in the genetic algorithm, the degree of freedom of the evaluation function is high and it is easy to define the required search formula, so it is preferred. In addition, it can be said that it is also an advantage that it is easy to prevent falling into a local optimum.
[0071] The storage unit 120 includes at least one of a volatile memory and a non-volatile memory. Examples of the volatile memory include DRAM (Dynamic Random Access Memory) and SRAM (Static Random Access Memory). Examples of the non-volatile memory include ReRAM (Resistive Random Access Memory, also known as resistive memory), PRAM (Phase change Random Access Memory), FeRAM (Ferroelectric Random Access Memory), MRAM (Magnetoresistive Random Access Memory, also known as magnetoresistive memory), and flash memory. In addition, the storage unit 120 may also include a recording medium drive. Examples of the recording medium drive include a hard disk drive (HDD) and a solid state drive (SSD).
[0072] The storage unit 120 may also include a database that stores data of files. In addition, the file retrieval support system 100 may also have a function of extracting data of files from a database existing outside the storage unit 120 or outside the system. In addition, the file retrieval support system 100 may also have a function of extracting data from both the database included in itself and the external database.
[0073] In addition, a file server may be used instead of the database. For example, when using files included in the file server, the database preferably has the paths of the files stored in the file server.
[0074] Examples of the data of the files included in the database include text data such as the body of the file, the title of the file, the author (including the writer, the author, the author, etc.), the classification, and the identification number. By comparing the data of the files included in the database with the information received by the receiving unit 110, the file can be specified.
[0075] For example, in the case where the document is a patent application document, the data of the document preferably includes at least one text data among the specification, claims, and abstract of the specification. In addition, the data of the document preferably includes, as information of the designated document, at least one among the application management number, application family management number, application number, publication number, and registration number. There is no restriction on the status of each patent application, nor on whether it is published, whether it is awaiting examination by the patent office, or whether it is registered. For example, the database may include data of at least one among applications before examination, applications under examination, and registered applications, or may include all data. In addition, the data of the document may also include data of at least one among the inventor, applicant, current owner, filing date, priority date, publication date, status, patent classification (CPC, IPC, FI, F-term, etc.). In addition, the data of the document may also include information unique to the company, such as the subject, classification, keywords, and remarks related to the document.
[0076] [Processing Unit 130] The processing unit 130 has a function of performing processes such as arithmetic operations, analysis, and inference using data supplied from one or both of the receiving unit 110 and the storage unit 120. The processing unit 130 may supply the generated data (for example, arithmetic operation results, analysis results, inference results) to one or both of the storage unit 120 and the output unit 140.
[0077] The processing unit 130 has a function of acquiring data of a plurality of documents from one or both of the storage unit 120 and the database.
[0078] In addition, the processing unit 130 has a function of classifying a plurality of documents according to the information received by the receiving unit 110.
[0079] For example, in the case where the receiving unit 110 only receives the relevant information of the documents included in the required document group, these documents are referred to as the first document group (a set of documents selected as required documents), and the remaining documents are referred to as the second document group (a set of documents not selected as required documents).
[0080] In addition, in the case where the receiving unit 110 receives the relevant information of both the documents included in the required document group and the documents not included in the required document group, the set of documents input as required documents is referred to as the first document group, the set of documents input as non-required documents is referred to as the second document group, and the remaining documents are referred to as the third document group (a set of documents for which it has not been determined whether they are required documents).
[0081] In addition, the processing unit 130 has a function of vectorizing (numericalizing) the documents according to the elements included in the data of the documents.
[0082] As described above, as elements, one or more of the information included in the document itself and the information associated with the document can be used.
[0083] As the information included in the document itself, for example, the words included in the text of the document can be cited. For example, in the case where the document is a patent application document, the words included in the text of one or more of the specification, claims, and abstract of the specification can be extracted. For example, morphological analysis and / or compound word analysis can be performed to extract the words included in the text.
[0084] In addition, specified parts of speech can also be extracted from the text. By using only the specified parts of speech, the total number of elements can be reduced, and subsequent processing can be simplified. As elements, for example, nouns are preferably used. In addition, consecutive nouns in a sentence can be regarded as one element, or can be regarded as one element by combining them into a compound noun.
[0085] As the information associated with the document, for example, the classification given to the document can be cited. For example, classification based on the classification method of the drawings or patent classification can be used. In addition, unique information associated with the document (for example, the subject, classification, keywords, and remarks about the document) can also be used as elements.
[0086] In particular, as elements, preferably one or both of words and classifications are used.
[0087] As methods for vectorizing the document using elements, various methods can be cited. For example, one-hot encoding (also called one hot vector), TF-IDF (Term Frequency-Inverse Document Frequency), and bag-of-words model can be cited.
[0088] In addition, for example, a model for converting elements into distributed representations can be learned by machine learning, and the distributed representations of the elements can be obtained through this model. A neural network is preferably used as the model.
[0089] In particular, one-hot encoding is preferably used. Generally speaking, in many cases, there are no restrictions on the number of occurrences of elements in document retrieval. By using one-hot encoding, the document can be vectorized according to whether a certain element is included in the document.
[0090] In addition, the processing unit 130 has a function of using the vector based on elements in the document and the determination label of whether it is the required document as learning data to learn a classifier. Furthermore, it has a function of extracting elements with high importance from the elements by analyzing this classifier.
[0091] In addition, the processing unit 130 has a function of generating a search formula for retrieving a required file group using elements extracted as elements with high importance. In addition, the processing unit 130 may also have a function of retrieving files in the database using the search formula. Further, the processing unit 130 has a function of outputting one or both of the search formula and the search result.
[0092] The processing unit 130 may include, for example, an arithmetic circuit. The processing unit 130 may include, for example, a central processing unit (CPU: Central Processing Unit). In addition, the processing unit 130 may include a GPU (Graphics Processing Unit: graphics processing unit).
[0093] The processing unit 130 may also include a microprocessor such as a DSP (Digital Signal Processor: digital signal processor). The microprocessor may also be implemented by a PLD (Programmable Logic Device: programmable logic device) such as an FPGA (Field Programmable Gate Array: field programmable gate array) and an FPAA (Field Programmable Analog Array: field programmable analog array). In addition, the processing unit 130 may include a quantum processor. By interpreting and executing instructions from various programs by the processor, the processing unit 130 can perform various data processing and program control. Programs executable by the processor are accommodated in at least one of the memory area included in the processor and the storage unit 120.
[0094] The processing unit 130 may also include a main memory. The main memory includes at least one of a volatile memory such as a RAM (Random Access Memory: random access memory) and a non-volatile memory such as a ROM (Read Only Memory: read only memory).
[0095] As the RAM, for example, DRAM, SRAM, etc. are used, in which a virtual storage space is allocated as the working space of the processing unit 130 and is used for the processing unit 130. The operating system, application programs, program modules, program data, look-up tables, etc. accommodated in the storage unit 120 are loaded into the RAM during execution. The processing unit 130 directly accesses and operates on these data, programs, and program modules loaded into the RAM.
[0096] The ROM can accommodate the BIOS (Basic Input / Output System) and firmware that do not need to be rewritten. Examples of ROMs include mask ROM, OTPROM (One Time Programmable Read Only Memory), and EPROM (Erasable Programmable Read Only Memory). Examples of EPROMs include UV-EPROM (Ultra-Violet Erasable Programmable Read Only Memory) that can erase stored data by ultraviolet irradiation, EEPROM (Electrically Erasable Programmable Read Only Memory), and flash memory.
[0097] At least a part of the processing of the document retrieval support system is preferably performed using artificial intelligence (AI: Artificial Intelligence).
[0098] The document retrieval support system particularly preferably uses an artificial neural network (ANN: Artificial Neural Network, hereinafter simply referred to as neural network). The neural network can be implemented by a circuit (hardware) or a program (software).
[0099] In this specification and the like, a neural network refers to all models that simulate the neural circuit network of a living being, determine the connection strength between neurons through learning, and thereby obtain the ability to solve problems. A neural network includes an input layer, an intermediate layer (hidden layer), and an output layer.
[0100] In this specification and the like, when explaining a neural network, sometimes determining the connection strength (also called the weight coefficient) between neurons based on existing information is called "learning".
[0101] In this specification and the like, sometimes constructing a neural network using the connection strength obtained through learning and deriving new conclusions from this structure is called "deduction".
[0102] For example, one or more of the vectorization of the above-mentioned document using elements, the classifier for determining whether a document is a required document, and the generation of a retrieval formula can utilize processing using AI.
[0103] [Output unit 140] The output unit 140 outputs information based on the processing result of the processing unit 130. For example, at least one of the operation result, analysis result, and inference result of the processing unit 130 can be supplied to the outside of the document retrieval support system 100. The output unit 140 can output the information to a terminal used by the user, a display, or the like.
[0104] Specifically, the output unit 140 can output the search formula generated by the processing unit 130. In addition, the result of searching for documents included in the database can be output using the search formula generated by the processing unit 130.
[0105] [Transfer channel 150] The transfer channel 150 has a function of transferring data. Data transmission and reception between the reception unit 110, the storage unit 120, the processing unit 130, and the output unit 140 can be performed through the transfer channel 150.
[0106] Refer to Figure 3 FIGS. 5 to 5 illustrate a document retrieval support method and an output method in a document retrieval support system according to an embodiment of the present invention. Note that, as an example of the output method, a display method of a display is given below. That is, the following describes a display method of the result of using the document retrieval support method according to an embodiment of the present invention.
[0107] <Document retrieval support method> The document retrieval support method of the present embodiment includes Figure 3 and Figure 4 the processes of steps S1 to S9 shown in FIGS. 5 and Figure 6 is a diagram for explaining the document retrieval support method. It can also be said that Figure 6 is an example of a graphical user interface (GUI) of the document retrieval support system according to the present embodiment. Figure 6 The windows, text boxes, etc. in FIGS. 5 to 5 are only examples, and there is no particular limitation thereto. The GUI can be configured as a web page accessed by the user through a network. Alternatively, the GUI can be configured as a screen of a program application executed on an information processing device such as a personal computer used by the user.
[0108] [Step S1] In step S1, data of a plurality of documents is acquired. For example, the processing unit 130 can acquire document data from the storage unit 120 or an external database. In addition, document data input by the user can also be acquired through the reception unit 110.
[0109] The data of the document includes at least data for specifying the document and data for generating a vector for learning by the classifier or the data of the vector.
[0110] There is no particular limitation on the number of files for which data is acquired in step S1. For example, it may be a part or all of the files that are the retrieval objects.
[0111] Step S1 can also be said to be a step of acquiring the necessary data in order to specify files using the information received in step S2.
[0112] For example, when the user's retrieval object is the entire set of Japanese patent applications, it is possible to acquire data for the entire set of Japanese patent applications or partial data. For example, in step S2, when the user only inputs data for the patent applications of their own company, it is also possible to only acquire data for the patent applications of that company in step S1.
[0113] When later entering step S31, it is possible to use the data of all the files acquired in step S1 for subsequent processing. However, the larger the number of files used, the greater the processing volume and the longer the processing time. Therefore, the number of files for which data is acquired in step S1 is preferably within a predetermined range.
[0114] Alternatively, when later entering step S32, the quantity of data for subsequent processing depends on the user's input content in step S2, so there is no particular limitation on the number of files for which data is acquired in step S1.
[0115] In addition, step S1 can also be said to be a step for acquiring the data used in step S4. Before performing step S4, the file is vectorized according to the elements in the data of the file acquired in step S1. Or, in the case where the file has already been vectorized, the data of the vector is acquired in step S1. Note that the generation or acquisition of the vector data can be performed before step S4 or after step S1.
[0116] [Step S2] In step S2, the receiving unit 110 receives the determination of the user's file through the terminal 20.
[0117] In Figure 6 the area 60 is the area where the user inputs and operates. For example, two forms 61, 62 (here text boxes) are set in the area 60, and information specifying the files included in the required file group mastered by the user is received through one, and information specifying the files not included in the required file group mastered by the user is received through the other.
[0118] As described above, as information for specifying a file, the title, author, various identification numbers, etc. of the file can be cited. In the case where the file is a patent application file, as information for specifying a file, the application management number, application family management number, application number, publication number, registration number, etc. can be cited. It is also possible to receive the input of text data such as the text of the file.
[0119] Figure 6 An example is shown in which required documents D11, D12, D13, etc. are input into form 61 using identification numbers or the like, and non-required documents UD21, UD22, UD23, etc. are input into form 62 using identification numbers or the like. The user only needs to input information specifying the required documents into form 61 at least, and form 62 can also be left blank. Then, by the user clicking or touching button 63 (Start), the information input into forms 61 and 62 is sent from terminal 20 to receiving unit 110 of file retrieval support system 100.
[0120] Here, when only form 61 is input with data, step S31 is entered, and when both form 61 and form 62 are input with data, step S32 is entered.
[0121] [Step S31] In step S31, the multiple files for which data is acquired in step S1 are classified into two types: file group D including required documents and file group UD including the remaining documents.
[0122] Processing unit 130 can determine file group D and file group UD based on the data received from terminal 20 by receiving unit 110 in step S2.
[0123] File group D includes each document (documents D11, D12, D13, etc.) input by the user into form 61 in step S2. Other documents can all be included in file group UD. When the number of documents included in file group UD is too large, the processing volume in subsequent processing increases and the processing time increases. Therefore, an upper limit can also be set for the number of documents included in file group UD. In this case, the documents constituting file group UD can be determined randomly or according to conditions such as period and field.
[0124] In the file retrieval method according to one aspect of the present invention, the user does not necessarily need to input non-required documents, so the user can easily obtain a retrieval formula for retrieving the required file group with less labor.
[0125] [Step S32] In step S32, the multiple files for which data is acquired in step S1 are classified into file group D including required documents, file group UD including non-required documents, and file group NS including the remaining documents.
[0126] Processing unit 130 can determine file group D, file group UD, and file group NS based on the data received from terminal 20 by receiving unit 110 in step S2.
[0127] The file group D includes each file (files D11, D12, D13, etc.) input by the user into form 61 in step S2. The file group UD includes each file (files UD21, UD22, UD23, etc.) input by the user into form 62 in step S2. Other files can all be included in the file group NS.
[0128] By using the relevant information of non-desired files mastered by the user, a retrieval formula with high retrieval accuracy can sometimes be generated. In particular, it is possible to suppress extracting files similar to the non-desired files in the generated retrieval formula.
[0129] Note that there is no limit to the number of files in the file group NS. The files included in the file group NS can, for example, include the retrieval objects when performing file retrieval in step S9.
[0130] [Step S4] In step S4, the vectors based on the elements in the data and the determination labels of whether the files are desired files in each file included in the file group D and the file group UD are used as learning data to perform learning of the classifier.
[0131] Specifically, the processing unit 130 uses the classifier accommodated in the storage unit 120 for processing. This classifier has the function of determining whether the file is a desired file by inputting the vector based on the elements in the data of the file.
[0132] As described above, the vector of the file can be obtained from the database or generated by the processing unit 130.
[0133] In the case where the file is a patent application file, as elements, for example, the words included in the specification, the words included in the claims, the words included in the abstract of the specification, any one or two or all of the words included in the specification, the claims, and the abstract of the specification, one or more of IPC, CPC, FI, and F-term can be used.
[0134] For example, by using the words included in the claims or the abstract of the specification to vectorize the file, a retrieval formula that can be retrieved according to the gist of the invention can be generated. In addition, when using the words included in the specification to vectorize the file, a retrieval formula that can be retrieved according to the overall content of the specification can be generated. In addition, when using IPC, CPC, FI, or F-term to vectorize the file, the deviation in the files input by the user into the form (for example, only inputting the patents of one's own company, etc.) can be reduced to generate a retrieval formula. In addition, when using IPC or CPC to vectorize the file, a retrieval formula that can perform equivalent retrieval in multiple countries can be generated.
[0135] Thus, the search expressions generated based on the selection of elements and the characteristics of the search results using such search expressions may differ. Therefore, it is also possible to vectorize documents by combining multiple elements. In addition, the user can also select the elements used to generate the search expressions according to the purpose of document search, etc.
[0136] Hereinafter, the case where words are used as elements will be described as an example.
[0137] When assigning determination labels, labels indicating required documents are assigned to the documents included in the document group D, and labels indicating non-required documents are assigned to the documents included in the document group UD.
[0138] [Step S5] In step S5, in the processing unit 130, a plurality of highly important elements are extracted by analyzing the classifier learned in step S4.
[0139] As search terms for generating the search expression, the elements extracted in step S5 are used. The more the number of search terms or elements, the higher the search accuracy of the search expression can be improved, and the fewer the number, the shorter the calculation time required to generate the search expression can be. Therefore, the number of elements extracted in step S5 is two or more, preferably 100 or more, 150 or more, or 200 or more and 2000 or less, 1500 or less, or 1000 or less. Note that the number of elements extracted in step S5 can also be less than 100 or greater than 2000.
[0140] In step S5, it is also possible to extract a specified number or a specified proportion of elements in the order of high importance. For example, it is also possible to extract elements that satisfy the extraction criteria such as having an importance higher than a specified value or being a specified value or more. In addition, for example, it is also possible to extract the top 5%, 10%, or 15% of the elements in the order of high importance. Additionally, it is also possible to calculate the average and standard deviation of the importance and extract elements greater than the average + 2σ or greater than the average + 3σ.
[0141] For example, in a classifier using a random forest, the importance of feature quantities can be calculated. Specifically, when words are used as elements, the degree of contribution of a certain word to the determination result of the classifier can be quantified as the importance. In other words, it can be said that the higher the importance of the word, the more helpful it is for determining whether the document input to the classifier is a required document.
[0142] [Step S6] In step S6, in the processing unit 130, the highly important elements are divided into two groups: group A included in the document group D and group B not included in the document group D.
[0143] For example, as Figure 5AAs shown, the words aaa, bbb included in at least one file in the file group D, and the word ccc included in both at least one file in the file group D and at least one file in the file group UD are classified into group A. In addition, the words yyy, zzz that are not included in any file in the file group D (in other words, included in at least one file in the file group UD) are classified into group B.
[0144] In addition, information about the analysis results of the classifier can also be provided to the user. Figure 6 An example of displaying information about the analysis results of the classifier in area 70 is shown.
[0145] In area 70, as a result of the processing in step S5, a list 71 of elements with high importance is displayed. In addition, as a result of the processing in step S6, a list 72 of elements in group A and a list 73 of elements in group B are displayed. Note that there is no limitation on whether to provide the lists 71, 72, 73 to the user respectively. For example, it can be arbitrarily determined to only display the list 71 of elements with high importance or to display the list 71 of elements with high importance and the list 72 of elements in group A, etc.
[0146] [Step S7] In step S7, in the processing unit 130, a search formula X is generated using the elements included in group A in such a way that the number of search terms is small, most of the files in the file group D include at least one search term, and most of the files in the file group UD do not include any search terms.
[0147] For example, one element in group A can be used as one search term. In addition, one or more logical operators (such as AND, OR, XOR, NOT, NAND, NOR) can be combined with two or more elements in group A to generate one search term.
[0148] For example, the number of search terms is two or more, preferably within a range of 2 times or less the number of elements in group A, and more preferably less than the number of elements in group A.
[0149] The search formula X is preferably generated in such a way that 50% or more of the files in the file group D include at least one search term and 50% or more of the files in the file group UD do not include any search terms. In addition, more preferably, 60% or more, 70% or more, 80% or more, 90% or more, 95% or more, 98% or more, or 100% of the files in the file group D include at least one search term. In addition, more preferably, 60% or more, 70% or more, 80% or more, 90% or more, 95% or more, or 98% or more of the files in the file group UD do not include any search terms.
[0150] The search formula X is preferably generated using a genetic algorithm. Here, the number of search terms corresponds to the number of genes, the number of search formulas as candidates for the search formula X corresponds to the population size, and the number of times of repeatedly performing the optimization calculation can be regarded as the number of generations. As the number of genes, population size, and number of generations, the more, the higher the accuracy of the search formula X can be improved, and the less, the shorter the calculation time required to generate the search formula X can be. Preferably, the number of genes, population size, and number of generations are all 50 or more, 100 or more, 150 or more, or 200 or more and 1000 or less, 700 or less, or 500 or less. In addition, the number of genes, population size, and number of generations can also all be less than 50 or greater than 1000.
[0151] The search terms for the search formula X become the search terms that make up the search formula Z generated in step S9. In particular, the search formula X has the function of extracting a plurality of files including the required files. Therefore, by generating the search formula X in a way that minimizes the number of search terms, the number of extracted files can be suppressed from being excessive.
[0152] Since the file group D includes the required files, it is preferable that all the required files can be extracted by the search formula X. Or, the more the number of extracted files, the better.
[0153] When performing step S32, since the file group UD includes non-required files, it is preferable that no non-required files are extracted by the search formula X. Or, the fewer the number of extracted files, the better.
[0154] When performing step S31, the file group UD also includes files for which it has not been determined whether they are required files, but it can also be said that they may include non-required files. Therefore, in this case, in the file group UD, the fewer the number of files extracted by the search formula X, the better.
[0155] [Step S8] In step S8, in the processing unit 130, the elements included in group B are used to generate the search formula Y in such a way that the number of search terms is small and most of the file group UD includes at least one search term.
[0156] For example, one element in group B can be used as one search term. In addition, one or more logical operators can be combined with two or more elements in group B to generate one search term.
[0157] For example, the number of search terms is two or more, preferably within a range of 2 times or less the number of elements in group B, and more preferably less than the number of elements in group B.
[0158] The search formula Y is preferably generated in such a way that it includes at least one search term in more than 50% of the documents in the document group UD. Furthermore, more preferably, more than 60%, 70%, 80%, 90%, 95% or 98% of the documents in the document group UD include at least one search term.
[0159] Similar to the generation of the search formula X, the search formula Y preferably uses a genetic algorithm.
[0160] The search terms for the search formula Y become the search terms that make up the search formula Z generated in step S9. In particular, the search formula Y has the function of removing unnecessary documents (so-called noise) from the documents extracted by the search formula X. By generating the search formula Y in such a way as to minimize the number of search terms as much as possible, noise can be appropriately removed.
[0161] Since the elements in group B are not included in the document group D, the documents included in the document group D will not be extracted by the search formula Y.
[0162] When performing step S32, since the document group UD includes unwanted documents, it is preferable to be able to extract all unwanted documents by the search formula Y. Or, the more the number of extracted documents, the better.
[0163] When performing step S31, the document group UD also includes documents for which it has not been determined whether they are wanted documents, and it can be said that they may include unwanted documents. Therefore, in this case, the more documents included in the document group UD extracted by the search formula Y, the better.
[0164] Here, either one of steps S7 and S8 can be performed first, or they can be calculated together at the same time, or they can be performed in parallel simultaneously. When performed simultaneously, parallel execution can reduce the amount of calculation, so it is preferable. Additionally, particularly preferably, step S7 is performed first, and the result obtained in step S7 is used to perform step S8. Thereby, the amount of calculation can be reduced in step S8 and the accuracy of the search formula can be improved.
[0165] Specifically, as Figure 5B shown, it can be considered that most (or all) of the documents in the document group D can be extracted by the search formula X, while a part of the document group UD (denoted as the document group YY) is extracted. Preferably, all the documents in the document group YY can be extracted by the search formula Y. Or, the more the number of extracted documents, the better. In other words, the document group YY includes the documents extracted by the search formula X from the document group UD. The search formula Y is preferably generated in such a way that more than 50% of the documents in the document group YY include at least one search term. Furthermore, more preferably, more than 60%, 70%, 80%, 90%, 95%, 98% or 100% of the documents in the document group YY include at least one search term.
[0166] As Figure 5C1 shown, when steps S7 and S8 are performed independently, whether the document is retrieved by the retrieval formula X is not considered. Therefore, the retrieval formula Y is generated to retrieve documents other than the document group YY in the large document group UD. As a result, unnecessary calculations may occur or the accuracy of the retrieval formula Y may decrease.
[0167] Therefore, as Figure 5C2 shown, it is preferable to use the elements in group B to generate a retrieval formula Y with a small number of retrieval terms and covering most of the document group YY (equivalent to the documents retrieved by the retrieval formula X in the document group UD).
[0168] [Step S9] In step S9, one or both of the retrieval formula Z using the retrieval formulas X and Y and the result of retrieving a plurality of documents using the retrieval formula Z are output.
[0169] The retrieval formula Z generated by the processing unit 130 and the result retrieved by the processing unit 130 are respectively supplied to the terminal 20 through the output unit 140.
[0170] Figure 6 An example is shown in which information about the retrieval formula is displayed in the area 80 and information about the retrieval result is displayed in the area 90.
[0171] In the area 80, a list 81 of the retrieval terms in the retrieval formula X, a list 82 of the retrieval terms in the retrieval formula Y, and the retrieval formula 83 (equivalent to the retrieval formula Z) are shown.
[0172] The list 81 shows retrieval terms including one element such as "aaa" and retrieval terms including logical operators and two elements such as "bbb AND ccc".
[0173] The list 82 shows retrieval terms including one element such as "yyy" and retrieval terms including logical operators and two elements such as "xxx AND zzz".
[0174] The area 80 preferably displays at least the retrieval formula 83. In addition, operators and the like may vary depending on the document retrieval system. Therefore, retrieval formulas can also be generated separately in the document retrieval system and multiple retrieval formulas can be displayed. As an example of the retrieval formula 83, a formula for excluding the retrieval result of the retrieval formula Y from the retrieval result of the retrieval formula X (NOT retrieval) is given.
[0175] The area 90 shows the result of retrieving documents in the database using the retrieval formula 83. The objects of this retrieval include not only the document group D and the document group UD but also the document group NS not used in steps S4 to S8. Furthermore, it may include documents not obtained in step S1.
[0176] In region 90, documents such as D11, D12, ND31, ND32, etc. are shown as document 91 retrieved by retrieval formula 83. Further, among document 91, documents such as D11, D12, etc. are shown as document 92 that matches the required documents (documents input to form 61) known to the user. Further, among document 91, documents such as ND31, ND32, etc. are shown as document 93 that is a required document not known to the user. Additionally, among the required documents (documents input to form 61) known to the user, documents such as D13, etc. are shown as document 94 that is not retrieved by retrieval formula 83.
[0177] Thus, it is preferable to display by comparing the retrieval result using retrieval formula 83 with the information received from the user in step S2. Thereby, it is easy for the user to evaluate the accuracy of retrieval formula 83. Additionally, by confirming the content of document 93, new documents in the required document group can be grasped. Additionally, in the case where a non-required document is retrieved as document 93, this document can also be added to form 62 and the retrieval formula generation can be executed again.
[0178] As described above, in the document retrieval support system according to one aspect of the present invention, by simply inputting the information related to the required documents known to the user, a retrieval formula for retrieving the required document group can be generated. Thereby, the user does not need to generate the retrieval formula by himself / herself, and thus the workload and working time related to document retrieval can be reduced. Therefore, the required information can be obtained efficiently.
[0179] This embodiment can be appropriately combined with other embodiments. Further, in this specification, in the case where a plurality of structural examples are shown in one embodiment, these structural examples can be appropriately combined.
[0180] (Embodiment 2) In this embodiment, with reference to Figure 7 and Figure 8 a document retrieval support system according to one aspect of the present invention will be described.
[0181] <Document Retrieval Support System 2> Figure 7 A block diagram showing document retrieval support system 210 is shown. Document retrieval support system 210 includes server 220 and terminal 230 (such as a personal computer). Note that the same components as those of Figure 1 the document retrieval support system 100 shown can be referred to the description of <Document Retrieval Support System 1> in Embodiment 1.
[0182] Server 220 includes communication unit 171a, transmission channel 172, storage unit 120, and processing unit 130. Although not shown in Figure 7 server 220 may further include at least one of a receiving unit, a database, an output unit, an input unit, etc.
[0183] The terminal 230 includes a communication unit 171b, a transmission channel 174, an input unit 115, a storage unit 125, a processing unit 135, and a display unit 145. As the terminal 230, various personal computers such as tablet type, notebook type, and desktop type, and various portable information terminals can be cited. In addition, the terminal 230 may be a desktop personal computer that does not include the display unit 145, and the terminal 230 may also be connected to a display or the like used as the display unit 145.
[0184] A user of the document retrieval support system 210 can input information on required documents and non-required documents from the input unit 115 of the terminal 230 to the server 220. Furthermore, document data and the like can also be input. These input contents are sent from the communication unit 171b to the communication unit 171a.
[0185] The information received by the communication unit 171a is stored in the memory or the storage unit 120 included in the processing unit 130 through the transmission channel 172. In addition, data can also be supplied to the processing unit 130 from the communication unit 171a through a receiving unit (refer to Figure 2 the receiving unit 110 shown). Or, it can also be said that the communication unit 171a is equivalent to Figure 2 the receiving unit 110 shown.
[0186] In the processing unit 130, various processes of step S4 (learning of the classifier), step S5 (analysis of the classifier), and steps S6 to S9 (generation of the retrieval formula) in the document retrieval support method described in Embodiment 1 are performed. These processes particularly require high processing capabilities, so it is preferably performed in the processing unit 130 included in the server 220. The processing capability of the processing unit 130 is preferably higher than that of the processing unit 135. Other steps are also preferably performed in the processing unit 130, and some steps can also be performed in the processing unit 135.
[0187] The processing result of the processing unit 130 is stored in the memory or the storage unit 120 included in the processing unit 130 through the transmission channel 172. Then, the processing result is output from the server 220 to the display unit 145 of the terminal 230. The processing result is sent from the communication unit 171a to the communication unit 171b. In addition, various data included in the database can also be sent from the communication unit 171a to the communication unit 171b based on the processing result of the processing unit 130. In addition, the processing result can also be supplied from the processing unit 130 to the communication unit 171a through an output unit ( Figure 2 the output unit 140 shown). Or, it can also be said that the communication unit 171a is equivalent to Figure 2 the output unit 140 shown.
[0188] [Communication unit 171a and communication unit 171b] The communication units 171a and 171b can be used to send and receive data between the server 220 and the terminal 230. As the communication units 171a and 171b, a hub, a router, a modem, etc. can be used. The sending and receiving of data can be performed wired or wirelessly (e.g., radio waves, infrared rays, etc.).
[0189] As a communication method between the communication unit 171a and the communication unit 171b, the structure shown in Embodiment 1 that can be used for the networks 30a and 30b can be used.
[0190] [Transmission channels 172 and 174] The transmission channels 172 and 174 have the function of transmitting data. The sending and receiving of data between the communication unit 171a, the storage unit 120, and the processing unit 130 can be performed through the transmission channel 172. The sending and receiving of data between the communication unit 171b, the input unit 115, the storage unit 125, the processing unit 135, and the output unit 140 can be performed through the transmission channel 174.
[0191] [Input unit 115] The input unit 115 can be used when a user designates file data, a box, etc. For example, the input unit 115 can have the function of operating the terminal 230. Specifically, a mouse, a keyboard, a touch panel, a microphone, a scanner, a camera, etc. can be cited.
[0192] The file retrieval support system 210 can also have the function of converting voice data into text data. For example, at least one of the processing unit 130 and the processing unit 135 can have this function.
[0193] The file retrieval support system 210 can also have an optical character recognition (OCR) function. Therefore, it is possible to recognize the characters included in the image data to generate text data. For example, at least one of the processing unit 130 and the processing unit 135 can have this function.
[0194] [Storage unit 125] The storage unit 125 can also store one or both of the file data and the data supplied from the server 220. In addition, the storage unit 125 can also include at least a part of the data that the storage unit 120 can include.
[0195] [Processing units 130 and 135] The processing unit 135 has the function of performing operations, etc. using the data supplied from the communication unit 171b, the storage unit 125, the input unit 115, etc. The processing unit 135 can also have the function of executing at least a part of the processing that can be performed by the processing unit 130.
[0196] Each of the processing units 130 and 135 may include one or both of a transistor (OS transistor) including a metal oxide in a channel formation region and a transistor (Si transistor) including silicon in the channel formation region.
[0197] In addition, in this specification and the like, a transistor using an oxide semiconductor or a metal oxide in a channel formation region is referred to as an Oxide Semiconductor transistor or an OS transistor. The channel formation region of the OS transistor preferably contains a metal oxide.
[0198] In this specification and the like, a metal oxide refers to an oxide of a metal in a broad sense. Metal oxides are classified into oxide insulators, oxide conductors (including transparent oxide conductors), and oxide semiconductors (Oxide Semiconductor, which may also be abbreviated as OS), etc. For example, when a metal oxide is used for a semiconductor layer of a transistor, the metal oxide is sometimes referred to as an oxide semiconductor.
[0199] The metal oxide included in the channel formation region preferably contains indium (In). When the metal oxide included in the channel formation region contains indium, the carrier mobility (electron mobility) of the OS transistor is improved. In addition, the metal oxide included in the channel formation region is preferably an oxide semiconductor containing element M. Element M is preferably at least one of aluminum (Al), gallium (Ga), and tin (Sn). Other elements that can be used as element M include boron (B), silicon (Si), titanium (Ti), iron (Fe), nickel (Ni), germanium (Ge), yttrium (Y), zirconium (Zr), molybdenum (Mo), lanthanum (La), cerium (Ce), neodymium (Nd), hafnium (Hf), tantalum (Ta), tungsten (W), etc. Note that multiple of the above elements may sometimes be combined as element M. Element M is, for example, an element with a high bond energy with oxygen. Element M is, for example, an element with a bond energy higher than that of indium with oxygen. In addition, the metal oxide included in the channel formation region preferably contains zinc (Zn). A metal oxide containing zinc is sometimes prone to crystallization.
[0200] The metal oxide included in the channel formation region is not limited to a metal oxide containing indium. The semiconductor layer may also be, for example, a metal oxide such as zinc tin oxide or gallium tin oxide that does not contain indium and contains zinc, gallium, or tin.
[0201] The processing unit 130 preferably includes OS transistors. Since the off-state current of the OS transistors is extremely small, by using the OS transistors as switches for holding the charge (data) flowing into the capacitors used as storage elements, long-term data retention periods can be ensured. By applying this characteristic to at least one of the registers and caches included in the processing unit 130, the processing unit 130 can be made to operate only when necessary, and in other cases, the previous processing information can be stored in the storage element, and the processing unit 130 can be turned off. That is to say, normally off computing can be achieved, and power consumption reduction of the file retrieval support system can be achieved.
[0202] [Display unit 145] The display unit 145 has a function of displaying the output result. Examples of the display unit 145 include a liquid crystal display device, a light-emitting display device, etc. Examples of the light-emitting element that can be used for the light-emitting display device include an LED (Light Emitting Diode), an OLED (Organic LED), a QLED (Quantum-dot LED), and a semiconductor laser. In addition, the following display devices can be used in the display unit 145: a display device using a MEMS (Micro Electro Mechanical Systems) element adopting a shutter method or an optical interference method; a display device using a display element adopting a microcapsule method, an electrophoresis method, an electrowetting method, or an electronic ink (registered trademark) method; etc.
[0203] Figure 8 A schematic diagram of the file retrieval support system according to the present embodiment is shown.
[0204] Figure 8 The shown file retrieval support system includes a server 5100 and terminals (which can also be referred to as electronic devices). Communication between the server 5100 and each terminal can be performed via the Internet line 5110.
[0205] The server 5100 can perform computations using the data input from the terminals via the Internet line 5110. The server 5100 can send the computation results to the terminals via the Internet line 5110. Therefore, the computation burden on the terminals can be reduced.
[0206] In Figure 8Among them, information terminals 5300, 5400, and 5500 are shown as terminals. Information terminal 5300 is an example of a portable information terminal such as a smartphone. Information terminal 5400 is an example of a tablet terminal. In addition, the information terminal 5400 can also be used as a notebook-type information terminal by connecting it to a housing 5450 including a keyboard. Information terminal 5500 is an example of a desktop information terminal.
[0207] By configuring in such a way, the user can access the server 5100 from information terminals 5300, 5400, 5500, etc. And the user can utilize the services provided by the administrator of the server 5100 through the communication via the Internet line 5110. As such a service, for example, a service of a file retrieval support method using one mode of the present invention can be cited. In this service, the server 5100 can also utilize artificial intelligence.
[0208] The server 5100 preferably performs file retrieval using a retrieval formula generated by a file retrieval support method using one mode of the present invention. Or, as described with reference to Figure 1 as such, file retrieval using this retrieval formula can also be performed in a server different from the server 5100.
[0209] This embodiment can be appropriately combined with other embodiments. [Reference Signs]
[0210] 20: Terminal, 30a: Network, 30b: Network, 40: File Retrieval System, 60: Area, 61: Form, 62: Form, 63: Button, 70: Area, 71: List, 72: List, 73: List, 80: Area, 81: List, 82: List, 83: Retrieval Formula, 90: Area, 91: File, 92: File, 93: File, 94: File, 100: File Retrieval Support System, 110: Receiving Unit, 115: Input Unit, 120: Storage Unit, 125: Storage Unit, 130: Processing Unit, 135: Processing Unit, 140: Output Unit, 145: Display Unit, 150: Transmission Channel, 171a: Communication Unit, 171b: Communication Unit, 172: Transmission Channel, 174: Transmission Channel, 210: File Retrieval Support System, 220: Server, 230: Terminal, 5100: Server, 5110: Internet Line, 5300: Information Terminal, 5400: Information Terminal, 5450: Housing, 5500: Information Terminal
Claims
1. A method for supporting document retrieval, comprising the following steps: Obtain data of multiple documents; Receive some of the multiple documents as required documents; Classify the multiple documents into a first document group including the required documents and a second document group including the remaining documents; Use the vectors based on the elements in the data and the determination labels indicating whether they are the required documents in each of the multiple documents as learning data to train a classifier; Extract two or more elements with high importance from the elements by analyzing the classifier; Classify the extracted elements into a first group included in the first document group and a second group not included in the first document group; Generate a first search term within a range of two or more and not more than twice the number of elements in the first group using the elements in the first group, and generate a first search formula in such a way that 50% or more of the documents in the first document group include at least one of the first search terms and 50% or more of the documents in the second document group do not include any of the first search terms; Generate a second search term within a range of two or more and not more than twice the number of elements in the second group using the elements in the second group, and generate a second search formula in such a way that 50% or more of the documents in the second document group include at least one of the second search terms; And Output one or both of the third search formula generated using the first search formula and the second search formula and the result of retrieving the multiple documents using the third search formula.
2. A method for supporting document retrieval, comprising the following steps: Obtain data of multiple documents; Receive some of the multiple documents as required documents and receive other parts of the documents as non-required documents; Classify the multiple documents into a first document group including the required documents, a second document group including the non-required documents, and a third document group including the remaining documents; Use the vectors based on the elements in the data and the determination labels indicating whether they are the required documents in each document included in the first document group and the second document group as learning data to train a classifier; Extract two or more elements with high importance from the elements by analyzing the classifier; Classify the extracted elements into a first group included in the first document group and a second group not included in the first document group; Generate a first search term within a range of two or more and not more than twice the number of elements in the first group using the elements in the first group, and generate a first search formula in such a way that 50% or more of the documents in the first document group include at least one of the first search terms and 50% or more of the documents in the second document group do not include any of the first search terms; Generate a second search term within a range of two or more and not more than twice the number of elements in the second group using the elements in the second group, and generate a second search formula in such a way that 50% or more of the documents in the second document group include at least one of the second search terms; And Output one or both of the third search formula generated using the first search formula and the second search formula and the result of retrieving the multiple files using the third search formula.
3. The file retrieval support method according to claim 1 or 2, wherein the second search formula is generated after the first search formula is generated, and the second search formula is generated in such a manner that 50% or more of the files extracted using the first search formula from the second file group include at least one of the second search terms.
4. The file retrieval support method according to claim 1 or 2, wherein the classifier is a classifier using a random forest.
5. The file retrieval support method according to claim 1 or 2, wherein the first search formula is generated using a genetic algorithm.
6. The file retrieval support method according to claim 1 or 2, wherein the second search formula is generated using a genetic algorithm.
7. The file retrieval support method according to claim 1 or 2, wherein the elements in the data are words.
8. A program having a function of causing a processor to execute the file retrieval support method according to claim 1 or 2.
9. A file retrieval support system including a receiving unit, a storage unit, a processing unit, and an output unit, Among them, wherein the receiving unit has a function of receiving a selection of a required file, the storage unit includes a classifier, and the processing unit has the following functions: generating vectors for multiple files respectively according to elements in the data; the receiving unit assigning a determination label of whether the file is the required file to the multiple files according to the received selection; using the vectors and the determination labels as learning data to perform learning of the classifier; extracting two or more elements with high importance by analyzing the classifier; and generating a search formula using the elements with high importance, and the output unit has a function of outputting the search formula.
10. A file retrieval support system including a receiving unit, a storage unit, a processing unit, and an output unit, Among them, wherein the receiving unit has a function of receiving a selection of a required file, the storage unit includes a classifier, and the processing unit has the following functions: classifying multiple files into a first file group including the required file and a second file group including the remaining files; generating a vector of the file according to elements in the data of the file; the receiving unit assigning a determination label of whether the file is the required file to the multiple files according to the received selection; using the vectors and the determination labels as learning data to perform learning of the classifier; extracting two or more elements with high importance by analyzing the classifier; classifying the elements with high importance into a first group included in the first file group and a second group not included in the first file group; generating first search terms within a range of two or more and not more than twice the number of elements in the first group using the elements in the first group, and generating a first search formula using the first search terms in such a manner that 50% or more of the files in the first file group include at least one of the first search terms and 50% or more of the files in the second file group do not include any of the first search terms; Generate a second search term within a range of more than two and less than or equal to twice the number of elements in the second group using the elements in the second group, and generate a second search formula using the second search term such that more than 50% of the documents in the second document group include at least one of the second search terms; and Generate a third search formula using the first search formula and the second search formula, and the output unit has a function of outputting the third search formula.
11. A document retrieval support system, which includes a receiving unit, a storage unit, a processing unit, and an output unit, Among them, The receiving unit has a function of receiving a selection of required documents and non-required documents, The storage unit includes a classifier, The processing unit has the following functions: Classify a plurality of documents into a first document group including the required documents, a second document group including the non-required documents, and a third document group including the remaining documents; Generate a vector of the document based on elements in the data of the document; The receiving unit assigns a determination label indicating whether the document is a required document to the plurality of documents according to the received selection; Use the vector and the determination label as learning data to perform learning of the classifier; Extract two or more highly important elements by analyzing the classifier; Classify the highly important elements into a first group included in the first document group and a second group not included in the first document group; Generate a first search term within a range of more than two and less than or equal to twice the number of elements in the first group using the elements in the first group, and generate a first search formula using the first search term such that more than 50% of the documents in the first document group include at least one of the first search terms and more than 50% of the documents in the second document group do not include any of the first search terms; Generate a second search term within a range of more than two and less than or equal to twice the number of elements in the second group using the elements in the second group, and generate a second search formula using the second search term such that more than 50% of the documents in the second document group include at least one of the second search terms; and Generate a third search formula using the first search formula and the second search formula, and the output unit has a function of outputting the third search formula.
12. The document retrieval support system according to claim 10 or 11, Wherein the processing unit has a function of retrieving the plurality of documents using the third search formula, And the output unit has a function of outputting the result of retrieving the documents.
Citation Information
Patent Citations
Document search method and system
JP2018073309A