Program, reverse cross classification device, and text classification method
Patent Information
- Application Number
- JP2023030611
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-03-01
- Publication Date
- 2026-02-13
AI Technical Summary
Existing technologies lack a reverse cross-classification technique to assign paper classifications to patents accurately, and there is uncertainty in matching codes between different classification systems, such as JST codes to IPC codes for articles and patents.
A program that utilizes a cross-classification unit to assign second type codes to first type texts, a common classification unit to assign common codes, and a similarity calculation unit to determine appropriate first type codes for second type texts based on calculated similarities, ensuring accurate code assignment.
Enables accurate and robust assignment of first type codes to second type texts, reducing errors and enhancing the reliability of code matching between different text classifications, particularly for patents and papers, facilitating industry-academia collaboration.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a program or the like that assigns a code to a text in order to classify the text. [Background technology]
[0002] Previously, there was a cross-classification technique that could assign a patent classification (e.g., IPC code), which is a different category, to a paper instead of a paper classification (e.g., JST code) (see Non-Patent Document 1). [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Hidetsugu Nanba and 3 others, "Automatic classification of academic papers into the International Patent Classification," [online], [Retrieved February 12, 2023], Internet [URL: https: / / www.japio.or.jp / 00yearbook / files / 2008book / 08_4_04.pdf] Summary of the Invention [Problem to be solved by the invention]
[0004] However, in the prior art, there was no reverse cross-classification technique that could assign a paper classification, which is a different category, to a patent.
[0005] In particular, with regard to papers and patents in a relationship subject to Article 30 of the Patent Act, in the prior art, for example, when an IPC code, which is an example of a second-type code, is assigned to a paper, which is an example of a first-type text, the IPC code must match the IPC code of the patent application for the paper, which is subject to Article 30 of the Patent Act, and when a JST code, which is an example of a first-type code, is assigned to a patent application, which is subject to Article 30 of the Patent Act, the JST code must match the JST code of the original paper, but such a match was difficult to achieve. Also, in the prior art, the accuracy of assigning a second-type code to a first-type text and the accuracy of assigning a first-type code to a second-type text were unknown, making it difficult for users to use a system that assigns the codes with confidence. [Means for solving the problem]
[0006] The program of the first invention includes a first-type storage unit storing one or more first-type texts to which one or more first-type codes have been assigned, the one or more second-type codes being assigned by a cross-classification unit that assigns a second-type code to each of the one or more first-type texts, and the one or more first-type texts to which one or more common codes have been assigned by a common classification unit that assigns a common code to each of the one or more first-type texts; and a second-type storage unit storing one or more second-type texts to which one or more second-type codes have been assigned, the one or more second-type texts to which one or more common codes have been assigned by a common classification unit that assigns a common code to each of the one or more second-type texts. and a program for causing a computer that can access the above to function as a text group acquisition unit that acquires one or more first type texts and one or more second type texts, each containing a predetermined second type code and a common code, from one or more first type texts and one or more second type texts, a similarity calculation unit that calculates a similarity between each of the one or more second type texts acquired by the text group acquisition unit and each of the one or more first type texts acquired by the text group acquisition unit, and an assignment unit that assigns, to each of the one or more second type texts acquired by the text group acquisition unit, the one or more first type codes that have been assigned to the one or more first type texts, using the similarity calculated by the similarity calculation unit.
[0007] With this configuration, the first type code added to the first type text can be appropriately assigned to the second type text.
[0008] The program of the second invention is a program for causing the computer to further function as a cross-classification unit that assigns a second-type code to each of one or more first-type texts, and a common classification unit that assigns a common code to each of one or more first-type texts and one or more second-type texts.
[0009] With this configuration, the process of reverse cross-classification can be completed.
[0010] Furthermore, in the program of the third invention, compared to the first or second invention, the number of common codes is smaller than the number of first type codes.
[0011] This configuration allows each common code to cover a wider range, resulting in less overflow.
[0012] It is desirable to set the number of common codes to 10% or less of the number of type 1 codes, i.e., to set the scope of each code to 10 times or more. Also, if the codes have a hierarchical structure, the number of codes in a given layer of the common code should be set to less than the number of codes in a given layer of the type 2 codes, desirably 10% or less.
[0013] Furthermore, the program of the fourth invention is a program for causing a computer to function in such a way that, for any one of the first to third inventions, the assignment unit acquires, for each of one or more second type texts, the top first type code of the first type text having a similarity ranking from 1st to Mth (M is a natural number) and the corresponding M similarities, acquires, for each of the one or more top first type codes, a cumulative similarity which is the sum of one or more similarities of each first type code, and assigns, to each of the one or more second type texts, the one or more first type codes corresponding to the cumulative similarity that satisfies the adoption condition.
[0014] With this configuration, one or more first-type codes added to the first-type text can be more appropriately assigned to the second-type text. More specifically, with this configuration, compared to a case where the first-type code to be assigned to the second-type text is determined only by the top code of the first-type code with the highest similarity, it is expected that a robust result can be obtained by preventing inconsistencies and errors in the logic of majority voting. The reason for this is as follows. [1] This uses the results of calculations of the similarity of not only the first but also the Mth first-class codes. However, this is simplified by assuming that only the first code is a candidate for assignment among the first-class codes. [2] Depending on whether the top codes are duplicated or not, the similarities are added up or not to obtain the cumulative similarity, and the top code of the first to Nth place in the cumulative similarity ranking is the code to be assigned, which is determined by the magnitude of the cumulative similarity. By applying this type of processing, a valid result that ensures robustness can be obtained by deciding by majority vote, and by setting N>=2, it is also possible to process multiple assignments of first-class codes.
[0015] M is often set to about 5 and N to about 3, but in cases where more caution is required, it is often appropriate to set M to about 10 and N to about 3. Note that text with low similarity may be mixed into the obtained bi-directional text group, and in extreme cases, there may be no similar text. In such cases, the disadvantage can be compensated for by determining the M rank so that text with a lower similarity than a certain level is rejected. From experience, when using COS similarity to calculate similarity, it is preferable to determine the M rank so that COS similarity of less than 0.6 is rejected.
[0016] In addition, the program of the fifth invention is a program in which, for any one of the first to fourth inventions, the first type text is a paper text, the second type text is a patent text, the first type code is a JST code, the second type code is an IPC code or FI or term or CPC code, and the common code is a JSPS common code.
[0017] With this configuration, the JST code added to the paper text can be appropriately assigned to the patent text. Once the JST code to be assigned to the patent text is determined, the paper text and the patent text will be assigned the IPC code, JSPS code, and JST code, allowing for precise display and understanding as three-dimensional coordinates, which will be of great help to matching of industry-academia collaborations, for example.
[0018] The program of the sixth invention is a program for causing a computer to function as follows: for the fifth invention, one or more second type texts are patent texts of patent applications that are subject to the exception to loss of novelty under Article 30 of the Patent Act; and one or more first type texts are paper texts for inventions that were published before the filing of a patent application for any of the patent texts among the one or more second type texts; the assignment unit assigns a JST code to each of the one or more patent texts, and for each of the one or more patent texts, determines whether the JST code assigned to the patent text matches the JST code assigned to the paper text for that patent text, obtains accuracy information regarding accuracy using the judgment result for each of the one or more patent texts, and uses the accuracy information as output.
[0019] Due to this configuration, it is inevitable that there will be a striking similarity between the patent text subject to Article 30 and the paper text. If the paper text and the patent text are fed into the above program and an extremely high similarity is obtained as expected, then the reliability of the above program can be confirmed. Effect of the Invention
[0020] According to the program of the present invention, the first type code added to the first type text can be appropriately assigned to the second type text. [Brief description of the drawings]
[0021] [Figure 1] Block diagram of a reverse cross sorting device 1 according to the first embodiment. [Diagram 2] A flowchart illustrating an example of the operation of the reverse cross sorting device 1. [Diagram 3] A flowchart illustrating an example of the cross classification process. [Figure 4] A flowchart illustrating an example of the common classification process. [Diagram 5] A flowchart illustrating an example of the reverse cross classification process. [Figure 6] A flowchart illustrating an example of the chord determination process. [Figure 7] Illustration of the calculation results of the similarity between the paper text and the patent text [Figure 8] Histogram of the calculation results of the similarity between the paper text and the patent text [Figure 9] Overview of the computer system [Figure 10] Block diagram of the computer system DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0022] Hereinafter, embodiments of a reverse cross sorting device and the like will be described with reference to the drawings. Note that components with the same reference numerals in the embodiments operate in the same manner, and therefore repeated description may be omitted.
[0023] (Embodiment 1) In this embodiment, a reverse cross-classification device that appropriately assigns a primary code added to a primary text to a secondary text using a common code will be described.
[0024] In addition, in this embodiment, a reverse cross classification device is described that calculates similarities between second type text and first type text, calculates an accumulation of the similarities for each first type code assigned to the first type text, uses the accumulation of similarities to appropriately determine a first type code to be assigned to the second type text, and assigns the first type code to the second type text.
[0025] In this embodiment, information X being associated with information Y means that information Y can be obtained from information X, or information X can be obtained from information Y, and the method of association is not important. Information X and information Y may be linked, may exist in the same buffer, information X may be included in information Y, or information Y may be included in information X, etc.
[0026] Figure 1 is a block diagram of a reverse cross classifier 1 according to this embodiment. The reverse cross classifier 1 comprises a storage unit 11, a receiving unit 12, a processing unit 13, and an output unit 14. The storage unit 11 comprises a first type storage unit 111 and a second type storage unit 112. The processing unit 13 comprises a cross classifier 131, a common classifier 132, a text group acquisition unit 133, a similarity calculation unit 134, and an assignment unit 135.
[0027] Various types of information are stored in the storage unit 11. The various types of information are, for example, a first type text, a second type text, various learning models, and adoption conditions. The storage unit 11 usually includes a first type storage unit 111 and a second type storage unit 112, but the first type storage unit 111 and the second type storage unit 112 may be included in other devices. One or more first type texts acquired from other devices may be temporarily stored in the storage unit 11 of the reverse cross classification device 1. One or more second type texts acquired from other devices may be temporarily stored in the storage unit 11.
[0028] The first type text and the second type text are texts. Here, text usually refers to the character information of a part or all of a document such as a paper or a patent. The language of the text may be Japanese or a foreign language. The text may be, for example, a text file, but the data format may be any, such as an HTML file or an XML file.
[0029] The first type storage unit 111 stores one or more first type texts. The first type text is a text of a different type from the second type text. One or more first type codes are assigned to each of the one or more first type texts. That is, one or more first type codes are associated with each of the one or more first type texts. It is preferable that the first type codes associated with the first type texts are codes assigned by an assigner who understands the contents of the first type text. The codes are information for classifying texts. The codes may be called classification codes.
[0030] A second type code, which will be described later, may be assigned in advance to each of the one or more first type texts in the first type storage unit 111. Furthermore, the cross classifying unit 131 may assign a second type code to each of the one or more first type texts. Note that when a second type code, which will be described later, is assigned in advance to a first type text, the reverse cross classifying device 1 does not need to include the cross classifying unit 131.
[0031] A common code, which will be described later, may be assigned in advance to each of the one or more first type texts in the first type storage unit 111. Alternatively, the common code may be assigned to each of the one or more first type texts by the common classification unit 132. If a common code, which will be described later, is assigned in advance to the first type texts, the reverse cross classification device 1 does not need to include the common classification unit 132.
[0032] The first type code, the second type code, and the common code are all codes. A code is a symbol or a sign that is assigned to a text in accordance with its contents. The number of codes assigned to a text is usually one, but it may be two or more. One of the codes assigned to a text is called the initial code. When the number of codes assigned to a text is two or more, the two or more codes are the initial code and the other codes. The initial code is usually the one that represents the content of the text.
[0033] A code is assigned to a text as long as the text and the code are associated with each other. For example, a code is assigned to a text when the text and the code are linked, when the code is included in the text, or when the text and the code are stored in association with each other. However, the manner in which the code is assigned to the text is not important.
[0034] The first type text is, for example, a thesis text. Thesis text is the text of a thesis. The thesis is usually an academic paper or a scientific or technical paper, but it does not matter. However, the first type text may be a type of text other than a thesis text. The first type text may be a patent text. Note that the first type text may include figures and tables.
[0035] When the first type text is a paper text, the first type code is a paper code. A paper code is a code for classifying papers. The paper code is, for example, a JST code. When the first type text is a patent text, the first type code is a patent code. The patent code may be called a patent classification code. The patent code is, for example, an IPC code, an FI, an F-term, or a CPC (Cooperative Patent Classification) code. The first type code may be, for example, an industry classification code.
[0036] The JST code is a paper classification code set by the Japan Science and Technology Agency. There are approximately 3,200 JST codes at the deepest level. The JST code may also be called the JST classification code.
[0037] The IPC code refers to the Patent Classification Code according to the Strasbourg Agreement of 1971. The number of IPC codes is approximately 80,000 at the deepest level and approximately 7,400 in the main group.
[0038] It is preferable that the primary code assigned to the primary text is a code assigned by an assigner (person). It is preferable that the primary code is a code assigned by an assigner who understands the content of the primary text.
[0039] The second type storage unit 112 stores one or more second type texts. One or more second type codes are assigned to each of the one or more second type texts. In other words, one or more second type codes are associated with each of the one or more second type texts. It is preferable that the second type codes associated with the second type texts are codes assigned by an assigner. It is preferable that the second type codes associated with the second type texts are codes assigned by an assigner who understands the contents of the second type text. Note that the second type texts may include figures and tables.
[0040] The second type of text is, for example, patent text. Patent text is the text of a patent. Patent text includes text related to patent information such as claims, specifications, and abstracts, and may include figures and tables. Patent text is, for example, text of published patent bulletins, text of patent bulletins, text of republished patent bulletins, and patent documents in preparation (e.g., one or more documents among claims, specifications, and abstracts). In other words, the status of the patent text (published, patented, pre-filed, etc.) does not matter. The second type of text may be a type of text other than patent text. The first type of text may be a paper text.
[0041] If the second kind of text is a patent text, the second kind of code is, for example, an IPC code, an FI, an F-term, or a CPC code. If the second kind of text is a thesis text, the second kind of code is, for example, a JST code.
[0042] The receiving unit 12 receives various instructions and information, such as a processing start instruction, a pre-processing start instruction, a main processing start instruction, an output instruction, and a map output instruction.
[0043] The processing start instruction is an instruction to start all processing in the processing unit 13 of the reverse cross classification device 1.
[0044] The preprocessing start instruction is an instruction to start preprocessing. Preprocessing is processing that is performed before the main processing, which will be described later. Preprocessing is cross classification processing and common classification processing performed by processing unit 13 of reverse cross classification device 1.
[0045] The instruction to start this process is an instruction to start this process. This process involves obtaining a text group that is the subject of the reverse cross-classification process described later, and performing reverse cross-classification using the text group. This process is performed by the text group obtaining unit 133, the similarity calculation unit 134, and the assignment unit 135.
[0046] The output instruction is an instruction to output information, which is a first type text to which a first type code and a second type code are assigned, or a second type text to which a first type code and a second type code are assigned.
[0047] The map output instruction is an instruction to output a map. A map is information that displays the classification of text using a code. The map is, for example, a first text map, which will be described later, or a second text map, which will be described later.
[0048] The reception unit 12 receives, for example, various instructions and information from a user. The reception unit 12 receives, for example, various instructions and information from a terminal device (not shown).
[0049] The means for inputting various instructions and information may be any means, such as a touch panel, a keyboard, a mouse, or a menu screen.
[0050] The processing unit 13 performs various types of processing. The various types of processing are, for example, processing performed by the cross classifier 131, the common classifier 132, the text group acquirer 133, the similarity calculator 134, and the assigner 135. If the reverse cross classifier 1 does not have the cross classifier 131 and the common classifier 132, the various types of processing are, for example, processing performed by the text group acquirer 133, the similarity calculator 134, and the assigner 135.
[0051] The processing unit 13, for example, uses the first type text and the second type text to which the first type code and the second type code are assigned to form a first text map, which will be described later.
[0052] The processing unit 13 constructs a second text map (to be described later) using, for example, the first type code, the second type code, and the first type text and the second type text to which the common code has been assigned.
[0053] The processing unit 13, for example, for each of one or more second type texts, determines whether or not a first type code given to the second type text matches a first type code given to a first type text for the second type text, obtains accuracy information on accuracy using a determination result for each of one or more second type texts, and outputs the accuracy information. The accuracy information is, for example, a correct answer rate, a recall rate, a matching rate, and an F-measure.
[0054] More specifically, the processing unit 13, for example, for each of one or more patent texts, determines whether or not the JST code assigned to the patent text matches the JST code assigned to the paper text for an invention published before the patent application for the patent text, obtains accuracy information regarding the accuracy using the judgment result for each of one or more patent texts, and outputs the accuracy information.
[0055] The processing unit 13, for example, determines, for each of one or more first-type texts, whether or not a second-type code assigned to the first-type text matches a second-type code assigned to a second-type text for the first-type text, obtains accuracy information regarding accuracy using the determination result for each of one or more first-type texts, and outputs the accuracy information.
[0056] More specifically, the processing unit 13, for example, determines, for each of one or more paper texts, whether or not the IPC code assigned to the paper text matches the IPC code assigned to the patent text for that paper text, obtains accuracy information regarding the accuracy using the determination result for each of one or more paper texts, and outputs the accuracy information.
[0057] The cross-classification unit 131 assigns one or more second-type codes to each of one or more first-type texts. For example, the cross-classification unit 131 assigns one or more second-type codes to each of one or more first-type texts in the first-type storage unit 111.
[0058] There is no restriction on the algorithm by which the cross classifying unit 131 assigns the second type code to the first type text. For example, the cross classifying unit 131 assigns the second type code to the first type text by the following method (1) or (2). (1) Machine learning
[0059] The cross-classification unit 131 typically acquires one or more second codes to be assigned to each of the one or more first-type texts in the first-type storage unit 111 through a machine learning prediction process using the first-type learning model and the first-type text. Next, the cross-classification unit 131 assigns the acquired one or more second codes to each of the one or more first-type texts.
[0060] More specifically, for example, the cross-classification unit 131 provides the first-type learning model and the first-type text to a machine learning prediction module for each of the one or more first-type texts in the first-type storage unit 111, executes the prediction module, and obtains one or more second-type codes. Next, the cross-classification unit 131 assigns the obtained one or more second-type codes to each of the one or more first-type texts.
[0061] The first type learning model is a learning model created by a machine learning learning process using two or more teacher data having a first type text and one or two or more second type codes. The learning model may also be called a learning device, a classifier, a classification model, or the like.
[0062] More specifically, the processing unit 13 or a learning unit not shown in the figure provides two or more teacher data having a first type of text and one or more second type codes to a machine learning learning module, executes the learning module, obtains and accumulates a first type of learning model.
[0063] In this specification, the machine learning algorithm is preferably deep learning, but may be random forest, decision tree, etc. In other words, any machine learning algorithm may be used. For machine learning, various machine learning functions such as TensorFlow library, R language random forest module, fastText, etc., and various existing libraries may be used. (2) Using document vector similarity
[0064] For example, the cross-classification unit 131 acquires a document vector of the first type text for each of one or more first type texts in the first type storage unit 111. Next, the cross-classification unit 131 may determine, for example, a document vector having the highest similarity to the document vector from the correspondence table, acquire one or more second type codes paired with the document vector from the correspondence table, and assign the one or more second type codes to the first type text. In such a case, the correspondence table has two or more pieces of correspondence information that are pairs of document vectors and one or more second type codes.
[0065] The common classification unit 132 assigns a common code to each of the one or more first type texts and each of the one or more second type texts.
[0066] The common code is a code that can be assigned to both the first type text and the second type text. It is preferable that the common code is assigned by an assigner. The number of the common codes is, for example, less than the number of the first type codes. The number of the common codes is, for example, less than the number of the second type codes.
[0067] When the first type text is a research paper text and the second type text is a patent text, and when the first type text is a patent text and the second type text is a research paper text, the common code is, for example, a JSPS common code. The common code may also be, for example, an industrial classification code.
[0068] The JSPS Common Code is the screening classification code for the Japan Society for the Promotion of Science's Grants-in-Aid for Scientific Research. Research results reports (hereafter referred to as "reports") for selected themes in the JSPS Grant-in-Aid for Scientific Research contain the JSPS Code, as well as the bibliographic information of the research paper and the bibliographic information of the patent (e.g., application number) as research results. In other words, such reports are associated with one or more paper texts. In addition, reports are associated with one or more patent texts.
[0069] The algorithm by which the common classification unit 132 assigns a common code to each of one or more first type texts and one or more second type texts is, for example, the following (1) or (2). (1) Machine learning
[0070] The common classification unit 132 typically obtains one or more common codes to be assigned to each of the one or more first-type texts in the first-type storage unit 111 and one or more second-type texts in the second-type storage unit 112 by a machine learning prediction process using a common learning model and the text. Next, the common classification unit 132 assigns the one or more acquired common codes to each of the one or more texts.
[0071] More specifically, for example, for each of one or more first type texts in the first type storage unit 111 and one or more second type texts in the second type storage unit 112, the common classification unit 132 provides a common learning model and the text to a machine learning prediction module, executes the prediction module, and obtains one or more common codes. Next, the common classification unit 132 assigns the obtained one or more common codes to each of the one or more texts.
[0072] The common learning model is a learning model created by a machine learning learning process using two or more pieces of teacher data each having a text and one or more common codes.
[0073] The common learning model is a learning model created by a machine learning learning process using, for example, training data having one or more JSPS common codes attached to one or more reports and paper text corresponding to the bibliographic information of the papers described in the report, and training data having one or more JSPS common codes attached to the report and patent text corresponding to the bibliographic information of the patents described in the report.
[0074] More specifically, the processing unit 13 or a learning unit (not shown) obtains, for each of one or more reports, one or more JSPS Common Codes corresponding to the report, obtains paper text corresponding to the bibliographic information of the paper in the report, obtains a document vector from the paper text, and obtains teacher data having the one or more JSPS Common Codes and the document vector. Also, the processing unit 13 or a learning unit (not shown) obtains, for each of one or more reports of Grants-in-Aid for Scientific Research, one or more JSPS Common Codes corresponding to the report, obtains patent text corresponding to the bibliographic information (e.g., patent application number) of the patent in the report, obtains a document vector from the patent text, and obtains teacher data having the one or more JSPS Common Codes and the document vector. Next, the processing unit 13 or a learning unit (not shown) provides two or more teacher data to a machine learning learning module, executes the learning module, and obtains and accumulates a common learning model. (2) Using document vector similarity
[0075] The common classification unit 132, for example, acquires a document vector from each of one or more first type texts in the first type storage unit 111 and one or more second type texts in the second type storage unit 112. Next, the common classification unit 132 may, for example, determine a document vector having the highest similarity to the document vector from the correspondence table, acquire one or more common codes paired with the document vector from the correspondence table, and assign the one or more common codes to the text. In such a case, the correspondence table has two or more pieces of correspondence information that are pairs of document vectors and one or more common codes.
[0076] The text group acquisition unit 133 acquires one or more first type texts and one or more second type texts including a predetermined second type code and a common code from one or more first type texts and one or more second type texts. Note that the method of giving the predetermined second type code and the common code here is not important.
[0077] The text group acquisition unit 133 typically acquires one or more first type texts and one or more second type texts that contain the same second type code and the same common code from one or more first type texts in the first type storage unit 111 and one or more second type texts in the second type storage unit 112.
[0078] For example, the text group acquisition unit 133 acquires two or more pairs of one second type code and one common code, and acquires, for each pair, one or more first type texts and one or more second type texts to which the two codes of each pair are assigned. In this case, it is preferable that the one or more first type texts and one or more second type texts to which the two codes are assigned are one or more first type texts and one or more second type texts to which each of the two codes is assigned as a first code.
[0079] The text group acquisition unit 133 acquires, for example, one or more pairs of one or more second type codes and one or more common codes provided by a user, and acquires, for each pair, one or more first type texts and one or more second type texts to which two or more codes included in each pair have been assigned.
[0080] The text group acquiring unit 133 may acquire first type texts and second type texts that include the same set of second type code and common code, regardless of the range.
[0081] The similarity calculation unit 134 calculates the similarity between each of the one or more second type texts acquired by the text group acquisition unit 133 and each of the one or more first type texts acquired by the text group acquisition unit 133 .
[0082] The similarity calculation unit 134 calculates, for each of the one or more second type texts acquired by the text group acquisition unit 133, a similarity between the second type text and each of the one or more first type texts acquired by the text group acquisition unit 133.
[0083] When obtaining the similarity between the second type text and the first type text, the similarity calculation unit 134 obtains, for example, a document vector from the second type text. The similarity calculation unit 134 also obtains, for example, a document vector from the first type text. Next, the similarity calculation unit 134 calculates, for example, the similarity between the two document vectors. Note that the technology for obtaining a document vector from text and the technology for calculating the similarity between two vectors are well known technologies, and therefore detailed explanations are omitted.
[0084] The assignment unit 135 assigns one or more first-type codes that are assigned to one or more first-type texts to each of the one or more second-type texts acquired by the text group acquisition unit 133, using the similarity calculated by the similarity calculation unit 134.
[0085] The assignment unit 135 assigns, for example, to each of one or more second-type texts acquired by the text group acquisition unit 133, one or more first-type codes corresponding to the similarity calculated by the similarity calculation unit 134 that satisfies the adoption conditions.
[0086] When the first kind of text is a thesis text and the second kind of text is a patent text, the assigning unit 135 assigns a JST code to each of one or more patent texts, for example.
[0087] The adoption condition is a condition for determining a first-type code to be assigned to a second-type text. The adoption condition is a condition related to the similarity between texts. The condition related to the similarity between texts includes a condition related to the cumulative similarity of the first-type code. The adoption condition is a condition related to the similarity between the second-type text and the first-type text. For example, the adoption condition is that the similarity between the second-type text and the first-type text is equal to or greater than a threshold value. For example, the adoption condition is that the ranking of the similarity between the second-type text and each of two or more first-type texts is equal to or greater than a threshold value. For example, the adoption condition is that the first-type code has a cumulative similarity equal to or greater than a threshold value, which will be described later. For example, the adoption condition is that the first-type code has a cumulative similarity ranking equal to or greater than a threshold value, which will be described later.
[0088] The assigning unit 135 determines the first type code to be assigned to the second type text, for example, by the following method (1) or (2). (1) Using cumulative similarity
[0089] The assigning unit 135 acquires, for example, the adoption conditions from the storage unit 11. For example, for each of one or more second kind texts, the assigning unit 135 acquires the first kind codes of the first kind texts ranked from 1st to Mth (M is a natural number) in terms of similarity and the corresponding M similarities. Next, for example, for each of one or more second kind texts, the assigning unit 135 acquires a cumulative similarity that is the sum of one or more similarities of each first kind code for each of one or more first kind codes. Note that for a first kind code that appears only once, the assigning unit 135 sets one similarity as the cumulative similarity. Next, for example, for each of one or more second kind texts, the assigning unit 135 assigns one or more first kind codes corresponding to the cumulative similarity that satisfies the adoption conditions to each of the one or more second kind texts. (2) Using document vector similarity
[0090] The assigning unit 135, for example, acquires the adoption conditions from the storage unit 11. For example, for each of the one or more second type texts acquired by the text group acquisition unit 133, the assigning unit 135 determines one or more first type texts corresponding to a similarity that satisfies the adoption conditions. Next, the assigning unit 135, for example, acquires one or more first type codes assigned to the one or more first type texts. Next, the assigning unit 135 assigns the acquired one or more first type codes to each of the one or more second type texts acquired by the text group acquisition unit 133, for example.
[0091] The output unit 14 outputs various information, such as the second type text to which the first type code is assigned, the first text map, and the second text map.
[0092] A first text map is a text map with a first type code (e.g., JST code) on the horizontal or vertical axis and a second type code (e.g., IPC code) on the vertical or horizontal axis, and is a two-dimensional map of paper texts and patent texts that have been assigned first and second type codes.
[0093] The second text map is a three-axis text map of first-type codes (e.g., JST codes), second-type codes (e.g., IPC codes), and common codes (JSPS common codes), and is a three-dimensional map of paper texts and patent texts that have been assigned first-type codes, second-type codes, and common codes.
[0094] Here, output is a concept that includes display on a display, projection using a projector, printing on a printer, sound output, transmission to an external device, storage on a recording medium, and handing over the processing results to other processing devices or other programs, etc.
[0095] The storage unit 11, the first type storage unit 111, and the second type storage unit 112 are preferably non-volatile recording media, but may also be realized as volatile recording media.
[0096] There is no restriction on the process by which information is stored in the storage unit 11, etc. For example, information may be stored in the storage unit 11, etc. via a recording medium, information transmitted via a communication line, etc. may be stored in the storage unit 11, etc., or information inputted via an input device may be stored in the storage unit 11, etc.
[0097] The reception unit 12 can be realized by a device driver for an input means such as a touch panel or a keyboard, or control software for a menu screen.
[0098] The processing unit 13, the cross classification unit 131, the common classification unit 132, the text group acquisition unit 133, the similarity calculation unit 134, and the assignment unit 135 can usually be realized by a processor, a memory, etc. The processing procedure of the processing unit 13, etc. is usually realized by software, and the software is recorded in a recording medium such as a ROM. However, it may be realized by hardware (dedicated circuit). The processor may be a CPU, MPU, GPU, etc., and the type is not important.
[0099] The output unit 14 may or may not include an output device such as a display, a speaker, etc. The output unit 14 may be realized by driver software for an output device, or a combination of driver software for an output device and an output device, etc.
[0100] Next, an example of the operation of the reverse cross sorting device 1 will be described with reference to the flowchart of FIG.
[0101] (Step S201) The processing unit 13 judges whether it is time to start pre-processing. If it is time to start pre-processing, the process proceeds to step S202, and if it is not time to start pre-processing, the process proceeds to step S204.
[0102] The timing for starting pre-processing is, for example, when the reception unit 12 receives a pre-processing start instruction or when a predetermined time arrives.
[0103] (Step S202) The cross classification unit 131 performs cross classification processing. An example of the cross classification processing will be described with reference to the flowchart of FIG.
[0104] The cross-classification process is a process of assigning one or more second-type codes (eg, IPC codes) to one or more first-type texts (eg, research paper texts).
[0105] (Step S203) The common classification unit 132 performs a common classification process. Return to step S201. An example of the common classification process will be described with reference to the flowchart of FIG.
[0106] The common classification process is a process of assigning one or more common codes (e.g., JSPS common codes) to one or more texts (e.g., academic paper texts and patent texts). The one or more texts here are one or more first-class texts and one or more second-class texts.
[0107] (Step S204) The processing unit 13 judges whether it is time to start this process. If it is time to start this process, the process proceeds to step S205, and if it is not time to start this process, the process proceeds to step S214.
[0108] The timing for starting this process may be, for example, when the accepting unit 12 accepts an instruction to start this process, when a predetermined time arrives, or when pre-processing is completed. In other words, the reverse cross classifier 1 may perform pre-processing and this process consecutively.
[0109] (Step S205) The text group acquiring unit 133 assigns 1 to a counter i.
[0110] (Step S206) The text group acquiring unit 133 judges whether or not the i-th code set exists. If the i-th code set exists, the process proceeds to step S207, and if not, the process returns to step S201.
[0111] A code set is a set of one or more second type codes and one or more common codes. It is preferable that a code set is a set of one second type code and one common code. The number of code sets is, for example, the number (X*Y) of the second type codes (X) multiplied by the number of common codes (Y). The number of code sets is, for example, the number of code sets included in the accepted instruction to start the process.
[0112] In other words, the i-th code set is, for example, the i-th code set among all combinations of one or more second type codes and one or more common codes, or the i-th code set among one or more code sets included in the accepted processing start instruction.
[0113] (Step S207) The text group acquiring unit 133 acquires, from the first type storage unit 111, one or more first type texts corresponding to one or more second type codes and one or more common codes included in the i-th code set.
[0114] (Step S208) The text group acquiring unit 133 acquires, from the second type storage unit 112, one or more second type texts corresponding to one or more second type codes and one or more common codes included in the i-th code set.
[0115] (Step S209) The similarity calculation unit 134 assigns 1 to a counter j.
[0116] (Step S210) The similarity calculation unit 134 judges whether or not the j-th type text exists among the one or more type texts acquired in step S208. If the j-th type text exists, the process proceeds to step S211, and if not, the process proceeds to step S213.
[0117] (Step S211) The similarity calculation unit 134 etc. perform reverse cross classification processing. An example of the reverse cross classification processing will be described with reference to the flowchart of FIG.
[0118] The reverse cross-classification process is a process of assigning one or more primary codes to each of one or more secondary texts.
[0119] (Step S212) The similarity calculation unit 134 increments the counter j by 1. The process returns to step S210.
[0120] (Step S213) The text group acquiring unit 133 increments the counter i by 1. The process returns to step S206.
[0121] (Step S214) The reception unit 12 judges whether or not a map output instruction has been received. If a map output instruction has been received, the process proceeds to step S215, and if not, the process returns to step S201.
[0122] (Step S215) The processing unit 13 acquires one or more first type texts from the first type storage unit 111.
[0123] (Step S216) The processing unit 13 acquires one or more second type texts from the second type storage unit 112.
[0124] (Step S217) The processing unit 13 constructs a text map, which is a two-dimensional map with the first type code and the second type code as axes, or a three-dimensional map with the first type code, the second type code, and the common code as axes, and which classifies one or more first type texts and one or more second type texts using the codes given to each text. Note that there are various types of text maps, and the technology for constructing a text map is a publicly known technology, so a detailed description of the processing will be omitted here.
[0125] (Step S218) The output unit 14 outputs the text map constructed in step S217. The process returns to step S201.
[0126] In the flowchart of FIG. 2, the process ends when the power is turned off or an interrupt occurs to end the process.
[0127] Next, an example of the cross-classification process in step S202 will be described with reference to the flowchart in FIG.
[0128] (Step S301) The cross sorting unit 131 assigns 1 to a counter i.
[0129] (Step S302) The cross-classification unit 131 judges whether or not the i-th first type text exists in the first type storage unit 111. If the i-th first type text exists, the process proceeds to step S303, and if the i-th first type text does not exist, the process returns to the upper process.
[0130] (Step S303) The cross classification unit 131 acquires the i-th primary text from the primary storage unit 111.
[0131] (Step S304) The cross classification unit 131 acquires the first type learning model from the storage unit 11.
[0132] (Step S305) The cross-classification unit 131 provides the i-th primary text and the primary learning model to the primary prediction module, and executes the primary prediction module. Then, the cross-classification unit 131 acquires one or more secondary codes that are the execution results of the primary prediction module.
[0133] (Step S306) The cross-classification unit 131 assigns one or more secondary codes acquired in step S305 to the i-th primary text.
[0134] (Step S307) The cross sorting unit 131 increments the counter i by 1. The process returns to step S302.
[0135] Next, an example of the common classification process in step S203 will be described with reference to the flowchart in FIG.
[0136] (Step S401) The common classification unit 132 assigns 1 to a counter i.
[0137] (Step S402) The common classification unit 132 judges whether or not the i-th text exists in the storage unit 11. If the i-th text exists, the process proceeds to step S403, and if the i-th text does not exist, the process returns to the upper process. Note that the i-th text is a first-class text or a second-class text.
[0138] (Step S403) The common classification unit 132 acquires the i-th text from the storage unit 11.
[0139] (Step S404) The common classification unit 132 acquires a common learning model from the storage unit 11.
[0140] (Step S405) The common classification unit 132 provides the i-th text and the common learning model to the common prediction module in the storage unit 11, and executes the common prediction module. Then, the common classification unit 132 acquires one or more common codes that are the execution results of the common prediction module. Note that the common prediction module may be the same as or different from the first type prediction module.
[0141] (Step S406) The common classification unit 132 assigns one or more common codes acquired in step S405 to the i-th text.
[0142] (Step S407) The common classification unit 132 increments the counter i by 1. The process returns to step S402.
[0143] Next, an example of the reverse cross-classification process in step S211 will be described with reference to the flowchart in FIG.
[0144] (Step S501) The similarity calculation unit 134 acquires the j-th secondary text in step S210.
[0145] (Step S502) The similarity calculation unit 134 acquires a second type vector, which is information obtained by vectorizing the second type text. The second type vector is a document vector.
[0146] (Step S503) The similarity calculation unit 134 assigns 1 to a counter i.
[0147] (Step S504) The similarity calculation unit 134 judges whether or not the i-th primary text exists among the primary texts acquired in step S207. If the i-th primary text exists, the process proceeds to step S505, and if not, the process proceeds to step S510.
[0148] (Step S505) The similarity calculation unit 134 acquires the i-th primary text.
[0149] (Step S506) The similarity calculation unit 134 acquires a primary vector, which is information obtained by vectorizing the i-th primary text. The primary vector is a document vector.
[0150] (Step S507) The similarity calculation unit 134 calculates the similarity between the second kind vector acquired in step S502 and the first kind vector acquired in step S506.
[0151] (Step S508) The similarity calculation unit 134 temporarily stores the similarity acquired in step S507 in a buffer (not shown) in association with the i-th first type text.
[0152] (Step S509) The similarity calculation unit 134 increments the counter i by 1. The process returns to step S504.
[0153] (Step S510) The assigning unit 135 determines one or more primary codes to be assigned to the secondary text acquired in step S501. An example of the code determination process will be described with reference to the flowchart of FIG.
[0154] In the code determination process, one or more first type codes to be assigned to the second type text are determined from the one or more first type codes assigned to the one or more first type texts using the similarity calculated by the similarity calculation unit 134.
[0155] (Step S511) The assigning unit 135 assigns the one or more first type codes determined in step S510 to the second type text acquired in step S501, and then returns to the upper process.
[0156] Next, an example of the chord determination process in step S510 will be described with reference to the flowchart in FIG.
[0157] (Step S601) The assignment unit 135 performs an initialization process. The initialization process is a process of substituting 0 for the accumulated similarity of all the first type codes. In addition, in the initialization process, the assignment unit 135 sorts the first type texts to which the first type codes are assigned, using the similarity in a buffer (not shown) as a key.
[0158] (Step S602) The granting unit 135 assigns 1 to a counter i.
[0159] (Step S603) The assignment unit 135 judges whether or not the i-th first type text exists among the first type texts having the 1st to Mth similarity degrees. If the i-th first type text exists, the process proceeds to step S604, and if not, the process proceeds to step S611. Note that M is a natural number. M may be the number of all first type texts, but is preferably a number smaller than the number of all first type texts (for example, 10).
[0160] (Step S604) The assignment unit 135 obtains the similarity (S) corresponding to the i-th first type text from a buffer (not shown).
[0161] (Step S605) The assigning unit 135 acquires one or more primary codes assigned to the i-th primary text. Note that it is preferable that the assigning unit 135 acquires only the top primary code. However, the assigning unit 135 may acquire all primary codes or the top N primary codes (here, N is a natural number of 2 or more).
[0162] (Step S606) The granting unit 135 assigns 1 to the counter j.
[0163] (Step S607) The adding unit 135 judges whether or not the jth first type code exists among the first type codes acquired in step S605. If the jth first type code exists, the process proceeds to step S608, and if not, the process proceeds to step S610.
[0164] (Step S608) The adding unit 135 adds the similarity (S) to the cumulative similarity of the jth primary code. If the jth primary code is not the top code, the similarity to be added to the cumulative similarity of the primary code may be a value subtracted from the similarity (S) (for example, "similarity (S) / ranking of primary code" or "similarity (S)-constant×ranking of primary code"). The constant is, for example, "0.1".
[0165] (Step S609) The granting unit 135 increments the counter j by 1. The process returns to step S607.
[0166] (Step S610) The granting unit 135 increments the counter i by 1. The process returns to step S603.
[0167] (Step S611) The assignment unit 135 acquires the adoption condition from the storage unit 11. The adoption condition is, for example, "the cumulative similarity is the maximum," "the cumulative similarity is in the top M (here, M is a natural number equal to or greater than 1)," and "the cumulative similarity is equal to or greater than a threshold."
[0168] (Step S612) The assignment unit 135 acquires one or more second type codes corresponding to the accumulated similarity that satisfies the adoption condition, and returns to the upper process.
[0169] A specific example of the operation of the reverse cross sorting device 1 according to this embodiment will now be described.
[0170] Here, the first type of text is a research paper text, the second type of text is a patent text, the first type of code is a JST code, the second type of code is an IPC code, and the common code is a JSPS common code.
[0171] A large number of research paper texts are currently stored in the first type storage unit 111 of the reverse cross-classification device 1. A JST code is assigned to each research paper text by the assigner.
[0172] Moreover, a large number of patent texts are stored in the second type storage unit 112. Each patent text is assigned an IPC code by an assigner (usually a person in charge at the Patent Office).
[0173] In this situation, the cross-classification unit 131 of the reverse cross-classification device 1 assigns one or more IPC codes to each paper text in the first type storage unit 111. The common classification unit 132 also assigns one or more JSPS common codes to each paper text. The common classification unit 132 also assigns one or more JSPS common codes to each patent text in the second type storage unit 112.
[0174] Next, the text group acquisition unit 133 acquires, from the one or more first type texts and the one or more second type texts, for each pair of the second type code and the common code, one or more first type texts and one or more second type texts that contain the same second type code and common code.
[0175] Next, the similarity calculation unit 134 calculates the similarity between one or more second-type texts acquired by the text group acquisition unit 133 and one or more first-type texts acquired by the text group acquisition unit 133. FIG. 7 shows an image of the similarity between the paper text and the patent text calculated by the similarity calculation unit 134. The first row (701) of FIG. 7 is the ID (international publication number) of the patent text, and the first column (702) of FIG. 7 is the ID of the paper text. Each cell in the second row and the second column onward of FIG. 7 is a bar graph showing the similarity between the patent text corresponding to the cell and the paper text. When the similarity here is at a maximum of "1.0", the bar graph is long enough to fill the cell. The patent text (703) in FIG. 7 is a patent text that is subject to the exception to loss of novelty under Article 30 of the Patent Act. The paper text (704) is a paper text corresponding to the patent text (703) and is a paper text published before filing. The similarity between the paper text (704) and the patent text (703) is 0.704.
[0176] Also, Figure 8 is a histogram of the similarity and number of cases between the numerous paper texts and numerous patent texts. In Figure 7, the horizontal axis is similarity and the vertical axis is number of cases. The range of the rightmost bar graph (801) of similarity on the horizontal axis is 0.697 to 0.712. In other words, the similarity between the patent text (703) and paper text (704) in Figure 7 is included in the rightmost bar graph (801) of the histogram in Figure 8.
[0177] Next, the assignment unit 135 assigns one or more JST codes assigned to one or more paper texts to each of one or more patent texts acquired by the text group acquisition unit 133 using the similarity calculated by the similarity calculation unit 134 according to the above-mentioned algorithm.
[0178] Through the above process, three types of codes, a JST code, an IPC code, and a common code, are assigned to each paper text in the first type storage unit 111 and each patent text in the second type storage unit 112.
[0179] As described above, according to the present embodiment, the first code added to the first text can be appropriately assigned to the second text. As a result, the first text and the second text to which the first code and the second code are assigned can be provided. Therefore, the first text and the second text can be seamlessly analyzed from multiple perspectives.
[0180] Furthermore, according to this embodiment, the similarity calculation unit 134 calculates only the similarity between one or more first type texts and one or more second type texts narrowed down and acquired by the text group acquisition unit 133, so the process of assigning a first type code to a second type text becomes extremely fast compared to calculating the similarities between all first type texts and all second type texts.
[0181] Furthermore, according to this embodiment, by using first type text to which an appropriate first type code has been assigned by an assigner, and first type text to which an appropriate second type code has been assigned by an assigner, it is possible to easily provide first type text and second type text to which an appropriate first type code and an appropriate second type code have been assigned.
[0182] Furthermore, according to the present embodiment, it is possible to provide a first type text and a second type text to which a first type code, a second type code, and a common code are assigned, so that the first type text and the second type text can be seamlessly analyzed from multiple perspectives.
[0183] Furthermore, according to this embodiment, the JST code added to the paper text can be appropriately assigned to the patent text.
[0184] Furthermore, according to this embodiment, patents and papers can be seamlessly classified using JST codes and IPC codes, which can contribute to the promotion of industry-academia collaboration, for example.
[0185] Furthermore, according to this embodiment, patents and papers can be seamlessly classified using JST codes, IPC codes, and JSPS common codes, which can contribute to the promotion of industry-academia collaboration, for example.
[0186] In this embodiment, it is preferable that the number of common codes is smaller than the number of first type codes. As a result, the coverage range per common code is wide, and overflow is reduced. It is preferable that the number of common codes is 10% or less of the number of second type codes, that is, the coverage range per code is 10 times or more. In addition, when the codes have a hierarchical structure, the number of codes in a given layer of the common code is preferably 10% or less so that it is smaller than the number of codes in a given layer of the second code.
[0187] Furthermore, in this embodiment, for each of one or more second type texts, the top first type code of the first type text with the highest similarity from 1st to Mth (M is a natural number) and the corresponding M similarities are obtained, and for each of one or more top first type codes, a cumulative similarity is obtained which is the sum of one or more similarities of each first type code, and one or more first type codes corresponding to the cumulative similarity that satisfies the adoption conditions are assigned to each of one or more second type texts, thereby obtaining robust results.
[0188] Furthermore, according to this embodiment, since there is inevitably a very high degree of similarity between the patent text subject to Article 30 and the paper text, if the paper text and the patent text are provided to the program and an extremely high degree of similarity is obtained as expected, the reliability of the program can be confirmed.
[0189] In the reverse cross classification device 1 of this embodiment, it is preferable that each of the one or more second-type texts is a patent text of a patent application that is subject to the exception to loss of novelty under Article 30 of the Patent Act, each of the one or more first-type texts is a paper text for an invention that was published before a patent application was filed for any of the patent texts among the one or more second-type texts, and the assignment unit 135 assigns a JST code to each of the one or more patent texts.
[0190] Furthermore, the processes of the cross classification unit 131 and the common classification unit 132 in the reverse cross classification device 1 of this embodiment are performed by another device, and the reverse cross classification device 1 is only required to acquire the results. In such a case, the reverse cross classification device 1 comprises a first type storage unit 111 which stores one or more first type texts to which one or more first type codes have been assigned, one or more second type codes have been assigned by a cross classification unit that assigns second type codes to each of the one or more first type texts, and one or more common codes have been assigned by a common classification unit 132 that assigns common codes to each of the one or more first type texts, and a second type storage unit 111 which stores one or more second type texts to which one or more second type codes have been assigned, and one or more second type texts to which one or more common codes have been assigned by a common classification unit that assigns common codes to each of the one or more second type texts. a text group acquisition unit 133 that acquires one or more first type texts and one or more second type texts, each containing a predetermined second type code and a common code, from the one or more first type texts and the one or more second type texts, a similarity calculation unit 134 that calculates a similarity between each of the one or more second type texts acquired by the text group acquisition unit 133 and each of the one or more first type texts acquired by the text group acquisition unit 133, and an assignment unit 135 that assigns one or more first type codes assigned to the one or more first type texts, to each of the one or more second type texts acquired by the text group acquisition unit 133, using the similarity calculated by the similarity calculation unit 134.
[0191] The process in this embodiment may be realized by software. This software may be distributed by software download or the like. This software may be recorded on a recording medium such as a CD-ROM and distributed. This also applies to other embodiments in this specification. The software for realizing the reverse cross classification device 1 in this embodiment is a program as follows. That is, this program accesses a first-type storage unit in which one or more first-type texts are stored, the first-type texts being one or more first-type codes assigned by a cross classification unit that assigns second-type codes to each of the one or more first-type texts, and the first-type texts being assigned one or more common codes by a common classification unit that assigns common codes to each of the one or more first-type texts, and a second-type storage unit in which one or more second-type texts are stored, the second-type texts being one or more second-type codes assigned by a common classification unit that assigns common codes to each of the one or more second-type texts. a text group acquisition unit that acquires one or more first type texts and one or more second type texts, each containing a predetermined second type code and a common code, from the one or more first type texts and the one or more second type texts; a similarity calculation unit that calculates a similarity between each of the one or more second type texts acquired by the text group acquisition unit and each of the one or more first type texts acquired by the text group acquisition unit; and an assignment unit that assigns, to each of the one or more second type texts acquired by the text group acquisition unit, the one or more first type codes assigned to the one or more first type texts, using the similarity calculated by the similarity calculation unit.
[0192] 9 shows the appearance of a computer that executes the programs described herein to realize the reverse cross classification device 1 and the like of the various embodiments described above. The above-described embodiments can be realized by computer hardware and a computer program executed thereon. FIG. 9 is an overview of this computer system 300, and FIG. 10 is a block diagram of the system 300.
[0193] In FIG. 9, computer system 300 includes a computer 301 including a CD-ROM drive, a keyboard 302, a mouse 303, and a monitor 304.
[0194] 10, computer 301 includes, in addition to CD-ROM drive 3012, MPU 3013, bus 3014 connected to CD-ROM drive 3012 etc., ROM 3015 for storing programs such as a boot-up program, RAM 3016 connected to MPU 3013 for temporarily storing instructions of application programs and providing temporary storage space, and hard disk 3017 for storing application programs, system programs, and data. Although not shown here, computer 301 may further include a network card for providing connection to a LAN.
[0195] A program for causing computer system 300 to execute the functions of reverse cross classification device 1 of the above-described embodiment may be stored on CD-ROM 3101, inserted into CD-ROM drive 3012, and then transferred to hard disk 3017. Alternatively, the program may be transmitted to computer 301 via a network (not shown) and stored on hard disk 3017. The program is loaded into RAM 3016 when executed. The program may also be loaded directly from CD-ROM 3101 or the network.
[0196] The program does not necessarily include an operating system (OS) or a third party program that causes the computer 301 to execute the functions of the reverse cross-classification device 1 of the above-described embodiment. The program need only include instructions for calling appropriate functions (modules) in a controlled manner to achieve the desired results. How the computer system 300 operates is well known, and a detailed description will be omitted.
[0197] In addition, in the above program, the steps of transmitting information and receiving information do not include processing performed by hardware, such as processing performed by a modem or interface card in the transmitting step (processing that is performed only by hardware).
[0198] The program may be executed by a single computer or a plurality of computers. That is, the program may be executed by a centralized processing or a distributed processing.
[0199] In each of the above embodiments, each process may be realized by centralized processing in a single device, or may be realized by distributed processing in a plurality of devices.
[0200] The present invention is not limited to the above-described embodiment, and various modifications are possible, and it goes without saying that these modifications are also included within the scope of the present invention. [Industrial Applicability]
[0201] As described above, the reverse cross classifier 1 according to the present invention has the effect of being able to appropriately assign a primary code added to a primary text to a secondary text, and is useful as a reverse cross classifier, etc. [Explanation of symbols]
[0202] 1 Reverse Cross Classifier 11 Storage area 12 Reception 13 Processing section 14 Output section 111 First-class storage unit 112 Second type storage unit 131 Cross Classification Section 132 Common classification section 133 Text Group Acquisition Unit 134 Similarity calculation part 135 Granting Department
Claims
1. a first type storage unit storing one or more first type texts, each of which is one or more first type codes assigned by a cross classification unit that assigns a second type code to each of the one or more first type texts, and one or more common codes assigned by a common classification unit that assigns a common code to each of the one or more first type texts; and a second type storage unit storing one or more second type texts, each of which is one or more second type codes assigned by a common classification unit that assigns a common code to each of the one or more second type texts, a text group acquisition unit that acquires one or more first type texts and one or more second type texts including a predetermined second type code and a common code from the one or more first type texts and the one or more second type texts; a similarity calculation unit that calculates a similarity between each of the one or more second type texts acquired by the text group acquisition unit and each of the one or more first type texts acquired by the text group acquisition unit; A program for functioning as an assignment unit that assigns one or more first-type codes assigned to one or more first-type texts to each of one or more second-type texts acquired by the text group acquisition unit, using the similarity calculated by the similarity calculation unit.
2. The computer, a cross-classification unit that assigns a second-type code to each of the one or more first-type texts; The program according to claim 1 , further functioning as a common classification unit that assigns a common code to each of the one or more first type texts and each of the one or more second type texts.
3. 3. The program according to claim 1, wherein the number of said common codes is smaller than the number of said first type codes.
4. The applying unit is For each of the one or more second type texts, obtain the top first type codes of the first type texts having the first to Mth similarity degrees (M is a natural number) and the corresponding M similarity degrees; For each of the one or more first primary codes, a cumulative similarity is obtained, which is the sum of one or more similarities of each of the first primary codes; 2. The program according to claim 1, for causing the computer to function as assigning one or more first type codes corresponding to a cumulative similarity that satisfies an adoption condition to each of the one or more second type texts.
5. The first text is a thesis text, and the second text is a patent text; 2. The program according to claim 1, wherein the first type code is a JST code, the second type code is an IPC code or an FI or an F-term or a CPC code, and the common code is a JSPS common code.
6. Each of the one or more second-type texts is a patent text of a patent application that is subject to the exception to loss of novelty under Article 30 of the Patent Act, Each of the one or more first-kind texts is a paper text on an invention published before a patent application for any of the patent texts among the one or more second-kind texts, The applying unit is assigning a JST code to each of said one or more patent texts; The program of claim 5, which causes the computer to function as follows: for each of the one or more patent texts, determine whether or not a JST code assigned to the patent text matches a JST code assigned to the paper text for that patent text, obtain accuracy information regarding accuracy using the judgment result for each of the one or more patent texts, and output the accuracy information.
7. a first-type storage unit for storing one or more first-type texts to which one or more first-type codes are assigned by a cross-classification unit that assigns a second-type code to each of the one or more first-type texts, and one or more first-type texts to which one or more second-type codes are assigned by a common classification unit that assigns a common code to each of the one or more first-type texts; a second-type storage unit for storing one or more second-type texts to which one or more second-type codes are assigned by a common classification unit that assigns a common code to each of the one or more second-type texts; a text group acquisition unit that acquires one or more first type texts and one or more second type texts including a predetermined second type code and a common code from the one or more first type texts and the one or more second type texts; a similarity calculation unit that calculates a similarity between each of the one or more second type texts acquired by the text group acquisition unit and each of the one or more first type texts acquired by the text group acquisition unit; an assignment unit that assigns, to each of the one or more second-type texts acquired by the text group acquisition unit, one or more first-type codes that are assigned to the one or more first-type texts, using the similarity calculated by the similarity calculation unit.
8. a first-type storage unit storing one or more first-type texts each having one or more first-type codes assigned thereto, the one or more second-type codes being assigned to each of the one or more first-type texts by a cross-classification unit that assigns a second-type code to each of the one or more first-type texts, and the one or more common codes being assigned to each of the one or more first-type texts by a common classification unit that assigns a common code to each of the one or more first-type texts; a second-type storage unit storing one or more second-type texts each having one or more second-type codes assigned thereto, the one or more second-type texts each having one or more common codes assigned thereto by a common classification unit that assigns a common code to each of the one or more second-type texts; a text group acquisition unit; a similarity calculation unit; and an assignment unit, a text group acquisition step in which the text group acquisition unit acquires one or more first type texts and one or more second type texts including a predetermined second type code and a common code from the one or more first type texts and the one or more second type texts; a similarity calculation step in which the similarity calculation unit calculates a similarity between each of one or more second type texts acquired in the text group acquisition step and each of one or more first type texts acquired in the text group acquisition step; and an assignment step in which the assignment unit assigns one or more first-type codes assigned to the one or more first-type texts to each of the one or more second-type texts acquired in the text group acquisition step, using the similarity calculated in the similarity calculation step.