Patent classification method using deep learning optimized for new growth fields
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- INHA UNIV RES & BUSINESS FOUNDATION
- Filing Date
- 2021-12-16
- Publication Date
- 2026-08-03
Smart Images

Figure 112021145755217-PAT00010_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a patent classification method using deep learning optimized for new growth fields, and more specifically, to a patent classification system and classification method using deep learning optimized for new growth fields using text mining. Background Technology
[0002] Today, patent information is utilized in such diverse ways that it influences not only research and development but also national policies regarding science and technology. Furthermore, the ability to locate desired information using well-organized documents—a key advantage of patent data—is crucial not only for understanding existing research results during R&D but also for finding evidence in appeals or litigation. Therefore, effectively utilizing existing patent information is vital for companies in preventing unnecessary R&D costs and for establishing new directions for research and development.
[0003] In order to specifically utilize the patent information found through investigation and search for technology development, the patent information must be systematically categorized and classified. However, when examining the classification process, since relevant patents must be verified and classified by humans one by one, it not only requires significant human, material, and time resources, but also frequently fails to achieve systematic classification due to errors caused by simple tasks and the lack of unified standards resulting from differences in understanding among individuals.
[0004] Registered Patent 10-0718745 "Patent search system and method using text mining" relates to a technology for deriving keywords through existing text mining and an automatic patent search technology using the same. Since it is a technology for searching for related patents for a single patent, its utility is reduced when analyzing patent databases where a large amount of patent data is collected.
[0005] Furthermore, the frequency used here has a problem in that it is based solely on the number of times words appear collectively, without considering the volume of each patent document.
[0006] Registered Patent 10-0809452 "Patent Classification Method and System Using a Computing Device" relates to a method for classifying patent data according to the hierarchical structure of technology classifications (classification tables such as IPC, UPC, ECLA, FI, F-term, Industrial Technology Classification, US Industrial Classification, and internal technology classification) by a computing device. However, since it classifies patents in a simple enumerative manner using existing IPC classifications, it has a problem in that it is insufficient for discovering new technologies within them, identifying relationships between technologies, or creating a new configuration of technology classification.
[0007] Furthermore, the IPC International Patent Classification fails to adequately represent current new technologies in its classification content, and the aforementioned classification method is merely a simple arrangement of patent documents in the patent database according to an already established patent classification system.
[0008] Therefore, there is a growing need for a new method of patent classification that allows for a more convenient understanding of the technical content within a patent database without having to check the contents of each patent individually.
[0009] Therefore, technology was needed to solve the aforementioned problems.
[0010] Meanwhile, the aforementioned background technology is technical information that the inventor possessed for the derivation of the present invention or acquired during the process of deriving the present invention, and it cannot be considered as prior art disclosed to the general public prior to the filing of the present invention. The problem to be solved
[0011] One embodiment of the present invention aims to provide a patent classification method using deep learning optimized for new growth fields. means of solving the problem
[0012] The present invention aims to solve the aforementioned problems and is characterized by comprising: a document grouping system utilizing a computer system connected via a network to a server equipped with a database storing multiple documents, wherein the system includes: a request means for requesting documents to be grouped from the server and receiving multiple documents; a vector means for parsing the received multiple documents to generate a multidimensional vector; a clustering means for clustering the multiple documents using the generated multidimensional vector and multiple document information possessed by the documents; a visualization means for producing information for visualizing the clustered multiple clusters and the multiple documents belonging to the clusters; and an output means for outputting the clusters and the multiple documents using the produced visualization information.
[0013] The present invention aims to solve the aforementioned problems and, in a document grouping method utilizing a computer system connected via a network to a server equipped with a database storing multiple documents, the method may include: a step of requesting documents to be grouped from the server and receiving multiple documents; a step of parsing the received multiple documents to generate a multidimensional vector; a step of clustering the multiple documents using the generated multidimensional vector and multiple document information possessed by the documents; a step of calculating information for visualizing the clustered multiple clusters and the multiple documents belonging to the clusters; and a step of outputting the clusters and the multiple documents using the calculated visualization information. Effects of the invention
[0014] As described above, according to the patent classification system and classification method using deep learning optimized for new growth fields according to the present invention, the effect of providing high-quality clustering results by utilizing the self-information possessed by the document in clustering is obtained. In addition, the effect of not having to review clusters one by one is obtained because the appropriateness of the clustering results can be visually judged.
[0015] In addition, according to one embodiment of the present invention, documents are automatically classified by category, so no manpower is required for document classification, thereby saving labor costs and the like.
[0016] In addition, according to one embodiment of the present invention, although there are not many experts in the field of new growth technologies, the problem of having to classify each document by category according to one embodiment of the present invention can be solved.
[0017] The effects obtainable from the present invention are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the description below. Brief explanation of the drawing
[0018] FIG. 1 is a diagram showing an example of configuring a similarity matrix between categories and keywords in a patent classification system and classification method according to one embodiment of the present invention. FIG. 2 is a diagram showing an example of configuring a similarity matrix between categories and documents in a patent classification system and classification method according to an embodiment of the present invention. FIG. 3 is a diagram showing an example of how each document is classified into a specific category in a patent classification system and classification method according to an embodiment of the present invention. Specific details for implementing the invention
[0019] Embodiments of the present invention are described below with reference to the attached drawings so that those skilled in the art can easily implement the invention. However, the present invention may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present invention in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification are denoted by similar reference numerals.
[0020] Throughout the specification, when a part is described as being “connected” to another part, this includes not only cases where they are “directly connected” but also cases where they are “electrically connected” with other components in between. Furthermore, when a part is described as “comprising” a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.
[0021] The shapes, sizes, ratios, angles, numbers, etc. disclosed in the drawings for explaining embodiments of the present invention are exemplary, and therefore the present invention is not limited to the depicted details. Throughout the specification, the same reference numerals refer to the same components. Furthermore, in describing the present invention, if it is determined that a detailed description of related known technology may unnecessarily obscure the essence of the present invention, such detailed description is omitted.
[0022] Where terms such as 'comprising,' 'having,' 'consisting of,' etc. are used in this specification, other parts may be added unless 'only' is used. Where a component is expressed in the singular, it includes cases where it is included in the plural unless specifically stated otherwise.
[0023] In interpreting the components, they are interpreted to include a margin of error even in the absence of a separate explicit statement.
[0024] In the case of describing a positional relationship, for example, when the positional relationship between two parts is described using expressions such as 'on,' 'upper,' 'lower,' or 'next to,' one or more other parts may be located between the two parts unless 'immediately' or 'directly' is used.
[0025] In the case of an explanation of a temporal relationship, for example, when a temporal sequence is explained using 'after', 'following', 'next', 'before', etc., it may include cases where the sequence is not continuous unless 'immediately' or 'directly' is used.
[0026] Although terms such as "first," "second," etc. are used to describe various components, these components are not limited by these terms. These terms are used merely to distinguish one component from another. Accordingly, the first component mentioned below may be the second component within the technical scope of the present invention.
[0027] The term “at least one” should be understood to include all combinations that can be presented from one or more related items. For example, the meaning of “at least one of the first item, the second item, and the third item” may mean not only the first item, the second item, or the third item individually, but also all combinations of items that can be presented from two or more of the first item, the second item, and the third item.
[0028] The features of each of the various embodiments of the present invention may be combined or combined with one another, either partially or wholly, and may technically enable various interlocking and operation. Each embodiment may be implemented independently of one another or may be implemented together in an associated relationship.
[0029] Below, we will specifically describe a patent classification system using deep learning optimized for new growth fields, which is an embodiment of the present invention.
[0030] FIG. 1 is a diagram showing an example of configuring a similarity matrix between a category and a keyword in a patent classification system and classification method according to an embodiment of the present invention, FIG. 2 is a diagram showing an example of configuring a similarity matrix between a category and a document in a patent classification system and classification method according to an embodiment of the present invention, and FIG. 3 is a diagram showing an example of each document being classified into a specific category in a patent classification system and classification method according to an embodiment of the present invention.
[0031] First, one embodiment of the present invention relates to a patent classification system using deep learning optimized for new growth fields, and may include: a request means for requesting documents to be grouped from a server and receiving a plurality of documents; a vector means for parsing the received plurality of documents to generate a multidimensional vector; a clustering means for clustering the plurality of documents using the generated multidimensional vector and the plurality of document information possessed by the documents; a visualization means for producing information for visualizing the clustered plurality of clusters and the plurality of documents belonging to the clusters; and an output means for outputting the clusters and the plurality of documents using the produced visualization information.
[0032] In addition, the documents applicable to the present invention are patent documents, and the term "patent document" is not limited to Korean patents or utility models, but may also include patents or utility models from other countries.
[0033] A vector means, which is an embodiment of the present invention, can first perform stop word processing. Stop word processing refers to a process of removing all words that are not necessary for the analysis of patent documents and leaving only the necessary words. For example, when classifying patents related to hydrogen, words such as hydrogen and tank may be retained, but words that are irrelevant to the patent classification, such as I and can, may be removed. Since stop word processing is a publicly disclosed technology, a detailed explanation will be omitted.
[0034] Meanwhile, the clustering means collects each word included in the aforementioned multiple documents, converts the word into a normal distribution based on the number of times the word is collected, and among the words that have undergone the normal distribution conversion, words with a z-value of 2 or more can be set as keywords.
[0035] To explain with a specific example, the clustering means can collect stopword-processed words included in multiple documents. For example, if hydrogen is included 200 times, I 400 times, tank 100 times, and can 50 times in multiple documents, the stopword-processed hydrogen 200 times and tank 100 times can be collected and stored.
[0036] The clustering means can convert collected words into canonical forms based on the number of times they were collected, and the conversion to canonical forms can be performed according to the following formula.
[0037]
[0038] z is the standard conversion value, x is the number of times each word is collected,
[0039] For example, assuming that stopword-processed words are collected an average of about 10 times, the number for "tank" will be greater than 0, and the number for "hydrogen" will be greater than that for "tank." In particular, when transformed into a normal distribution, it can be seen that the larger the value, the more frequently the word is included in the patent document.
[0040] The similarity between each category and the keyword is calculated by executing sim(A) = {simcos(a, A), simcos(b, A), simcos(c, A), ...} as shown in FIG. 1, where A is a classification category, a, b, c, ... are the collected keywords, simcos(a, A) is the cosine similarity between keyword a and category A, and sim(A) may be a similarity matrix between each keyword and category A.
[0041] For example, the similarity between a keyword and each category can be determined. This determination can be based on cosine similarity, and the formula for determining cosine similarity (simcos) is as follows.
[0042] simcos(a,A) =
[0043] and is a multidimensional vector containing statistical information of keywords a and words of category A, respectively.
[0044] Keywords and vectors are composed of multidimensional vectors, and their magnitude is determined to be between -1 and 1 according to the cosine similarity formula. That is, if the cosine similarity is determined through the inner product of each embedded keyword vector and category vector, keywords and categories with opposite meanings may show a cosine similarity close to -1, while keywords and categories with corresponding meanings may show a cosine similarity close to 1.
[0045] If we assume that keywords a, b, c, d, e, f and category A are 0.652, 0.152, -0.0342, 0.751, 0.022, and 0.096, respectively, then sim(A) can be {0.652, 0.152, -0.034, 0.751, 0.022, 0.096}.
[0046] Meanwhile, the clustering means determines whether keywords are included in each document and further determines the category of each document.
[0047] Specifically, the clustering method determines whether each document contains keywords based on the equation REF(AA) = {val(a,AA), val(b,AA}, val(c,AA),… where AA is a document that needs to be classified, val(a, AA) is whether keyword a is included in document AA, REF(AA) is an inclusion matrix representing whether each keyword is included in document AA, and the category of document AA can be determined through the inner product of REF(AA) and sim(A).
[0048] To explain more specifically, if keyword 'a' is included in document AA, val(a,AA) will be 1, and if it is not included, val(a,AA) will be 0. If AA contains a, c, and d, but lacks b, e, and f, REF(AA) will take the form of a vector {1,0,1,1,0,0}. Finally, the similarity between document AA and category A can be determined by taking the inner product of the sim(A) vector and the REF(AA) vector. In this example, the inner product of {0.652, 0.152, -0.034, 0.751, 0.022, 0.096} and {1,0,1,1,0,0} would be 1.367. Document AA can be examined for similarity with each category, such as A, B, and C, and ultimately classified as belonging to the category with the highest inner product value.
[0049] Below, we will specifically explain a patent classification method using deep learning optimized for new growth fields.
[0050] According to one embodiment of the present invention, a patent classification method using deep learning optimized for new growth fields may include: a step of requesting documents to be grouped from a server and receiving a plurality of documents; a step of parsing the received plurality of documents to generate a multidimensional vector; a step of clustering the plurality of documents using the generated multidimensional vector and the plurality of document information contained in the documents; a step of calculating information for visualizing the clustered plurality of clusters and the plurality of documents belonging to the clusters; and a step of outputting the clusters and the plurality of documents using the calculated visualization information.
[0051] One embodiment of the present invention may first perform stop word processing. Stop word processing refers to a process of removing all words that are not necessary for the analysis of patent documents and leaving only the necessary words. For example, when classifying patents related to hydrogen, words such as hydrogen and tank may be retained, but words irrelevant to the patent classification, such as I and can, may be removed. Since stop word processing is a publicly disclosed technology, a detailed explanation will be omitted.
[0052] Meanwhile, one embodiment of the present invention collects each word included in the plurality of documents, converts the word into a normal distribution based on the number of times the word is collected, and among the words that have undergone the normal distribution conversion, words with a z value of 2 or more can be set as keywords.
[0053] To explain specifically with an example, one embodiment of the present invention can collect stopword-processed words included in a number of documents. For example, if hydrogen is included 200 times, I 400 times, tank 100 times, and can 50 times in a number of documents, then 200 times hydrogen and 100 times tank that have been stopped-processed can be collected and stored.
[0054] Meanwhile, the clustering step included in one embodiment of the present invention may include: a step of collecting each word included in a plurality of documents and performing a normal distribution conversion based on the number of times the word is collected; a step of setting a word among the words that have undergone the normal distribution conversion as a keyword if the z value is 2 or higher; and a step of calculating the similarity between the two documents by executing the formula sim(A) = {simcos(a, A), simcos(b, A), simcos(c, A), ...}.
[0055] For example, the similarity between a keyword and each category can be determined. This determination can be based on cosine similarity, and the formula for determining cosine similarity (simcos) is as follows.
[0056] sincos(a,A) =
[0057] and is a multidimensional vector containing statistical information of keywords a and words of category A, respectively.
[0058] Keywords and vectors are composed of multidimensional vectors, and their magnitude is determined to be between -1 and 1 according to the cosine similarity formula. That is, if the cosine similarity is determined through the inner product of each embedded keyword vector and category vector, keywords and categories with opposite meanings may show a cosine similarity close to -1, while keywords and categories with corresponding meanings may show a cosine similarity close to 1.
[0059] If we assume that keywords a, b, c, d, e, f and category A are 0.652, 0.152, -0.0342, 0.751, 0.022, and 0.096, respectively, then sim(A) can be {0.652, 0.152, -0.034, 0.751, 0.022, 0.096}.
[0060] Meanwhile, the clustering means determines whether keywords are included in each document and further determines the category of each document.
[0061] Here, A is a classification category, a, b, c, … are the collected keywords, simcos(a, A) is the cosine similarity between keyword a and category A, and sim(A) may be a similarity matrix between each keyword and category A.
[0062] Meanwhile, the clustering step expresses whether a keyword is included in a document as a matrix using the following equation, REF(AA) = {val(a,AA), val(b,AA}, val(c,AA),… where AA is a document that needs to be classified, var(a, AA) is whether keyword a is included in document AA, REF(AA) is an inclusion matrix representing whether each keyword is included in document AA, and the category of document AA can be determined through the inner product of REF(AA) and sim(A).
[0063] Specifically, the clustering method determines whether each document contains keywords based on the equation REF(AA) = {val(a,AA), val(b,AA}, val(c,AA),… where AA is a document that needs to be classified, val(a, AA) is whether keyword a is included in document AA, REF(AA) is an inclusion matrix representing whether each keyword is included in document AA, and the category of document AA can be determined through the inner product of REF(AA) and sim(A).
[0064] To explain more specifically, if keyword 'a' is included in document AA, val(a,AA) will be 1, and if it is not included, val(a,AA) will be 0. If AA contains a, c, and d, but lacks b, e, and f, REF(AA) will take the form of a vector {1,0,1,1,0,0}. Finally, the similarity between document AA and category A can be determined by taking the inner product of the sim(A) vector and the REF(AA) vector. In this example, the inner product of {0.652, 0.152, -0.034, 0.751, 0.022, 0.096} and {1,0,1,1,0,0} would be 1.367. Document AA can be examined for similarity with each category, such as A, B, and C, and ultimately classified as belonging to the category with the highest inner product value.
[0065] The foregoing description of the present invention is for illustrative purposes only, and those skilled in the art will understand that other specific forms can be easily modified without altering the technical spirit or essential features of the present invention. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, each component described as a single unit may be implemented in a distributed manner, and components described as distributed may likewise be implemented in a combined form.
[0066] The scope of the present invention is defined by the claims set forth below rather than by the detailed description above, and all modifications or variations derived from the meaning and scope of the claims and equivalent concepts thereof should be interpreted as being included within the scope of the present invention. Explanation of the symbols delete
Claims
Claim 1 A patent classification system utilizing a computer system connected via a network to a server equipped with a database storing multiple documents, comprising: a request means for requesting documents to be grouped from the server and receiving multiple documents; a vector means for parsing the received multiple documents to generate a multidimensional vector; a clustering means for clustering the multiple documents using the generated multidimensional vector and multiple document information possessed by the documents; a visualization means for producing information for visualizing the clustered multiple clusters and the multiple documents belonging to the clusters; and an output means for outputting the clusters and the multiple documents using the produced visualization information, wherein the clustering means collects each word included in the multiple documents and performs a normal distribution transformation based on the number of times the word is collected, wherein the normal distribution transformation is performed according to the following equation. Here, z is the standard conversion value, x is the number of times each word is collected, is the average number of times collected words are collected, ε is the standard deviation, and among the words that have undergone the normal distribution transformation, words with a z-value of 2 or greater are set as keywords; for the keywords, a similarity matrix sim(A) between each keyword and category is generated using the cosine similarity between the keyword and classification category A, wherein the similarity between the multiple documents is calculated by executing sim(A) = {simcos(a, A), simcos(b, A), simcos(c, A), ...}, where A is the classification category, a, b, c, ... are the collected keywords, simcos(a, A) is the cosine similarity between the keyword and the category, sim(A) is the similarity matrix between each keyword and the category, and the formula for determining the cosine similarity is simcos(a,A) = And, and is a multidimensional vector containing statistical information of words of keyword a and category A, respectively; a category-keyword similarity matrix sim(A) is generated by determining cosine similarity through the inner product of the keyword vector and the category vector, where keywords and categories with opposite meanings show cosine similarity close to -1, and keywords and categories with corresponding meanings show cosine similarity close to 1; the clustering means generates a document-keyword inclusion matrix REF(AA) indicating whether keywords are included in each document, and determines whether keywords are included in each document based on the equation REF(AA) = {val(a,AA), val(b,AA}, val(c,AA),… where AA is a document requiring classification, var(a, AA) is whether keyword a is included in document AA, REF(AA) is an inclusion matrix representing whether each keyword is included in document AA, and through the inner product of REF(AA) and sim(A) A patent classification system using deep learning optimized for new growth fields that determines the category of AA documents. Claim 2 delete Claim 3 delete Claim 4 A patent classification method utilizing a computer system connected via a network to a server equipped with a database storing multiple documents, comprising: a step of requesting documents to be grouped from the server and receiving multiple documents; a step of parsing the received multiple documents to generate a multidimensional vector; a step of clustering the multiple documents using the generated multidimensional vector and multiple document information possessed by the documents; a step of producing information for visualizing the clustered multiple clusters and the multiple documents belonging to the clusters; and a step of outputting the clusters and the multiple documents using the produced visualization information, wherein the clustering step comprises: a step of collecting each word included in the multiple documents and performing a normal distribution conversion based on the number of times the word was collected; a step of setting words among the words that underwent the normal distribution conversion that have a z-value of 2 or more as keywords; and a step of calculating the similarity between the multiple documents by executing sim(A) = {simcos(a, A), simcos(b, A), simcos(c, A), ...}, wherein A is a classification category and a, b, c, ... is the above-mentioned collected keyword, simcos(a, A) is the cosine similarity between keyword a and category A, sim(A) is the similarity matrix between each keyword and category A, and the above-mentioned normal distribution transformation is performed by the following equation, Here, z is the standard conversion value, x is the number of times each word is collected, is the collection average count of collected words, is the standard deviation, and the formula for determining the above cosine similarity is simcos(a,A) = And, and A multidimensional vector having statistical information of words of keyword a and category A, respectively; when cosine similarity is determined through the inner product of the keyword vector and the category vector, keywords and categories with opposite meanings show a cosine similarity close to -1, and keywords and categories with corresponding meanings show a cosine similarity close to 1; the clustering step expresses whether the keyword is included in a document as a matrix through the following equation, REF(AA) = {val(a,AA), val(b,AA}, val(c,AA),… where AA is a document requiring classification, var(a, AA) is whether keyword a is included in document AA, and REF(AA) is an inclusion matrix representing whether each keyword is included in document AA; and a patent classification method using deep learning optimized for new growth fields that determines the category of each document by performing an inner product of the category-keyword similarity matrix sim(A) and the document-keyword inclusion matrix REF(AA). Claim 5 delete Claim 6 delete