Patent application pre-evaluation method based on similarity calculation, storage medium and device
By constructing a patent dataset and performing similarity calculations, the problem of low efficiency in pre-patent application evaluation is solved, enabling preliminary identification and quality assessment of patent applications and improving the efficiency and quality of patent applications.
Patent Information
- Application Number
- CN202310850633.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-11
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2043-07-11
AI Technical Summary
The lack of efficient and accurate pre-patent application evaluation methods in the existing technology leads to long patent examination cycles, low grant rates, and increased costs and risks.
By constructing a patent dataset based on similarity calculation and classifying it by patent field and IPC, the similarity of the patents to be evaluated is extracted using the LDA model and perplexity calculation, forming a similarity matrix and heat map to achieve pre-application evaluation.
It provides a preliminary identification of patent applications, determining whether they can be applied for directly or whether amendments are needed, reducing the risk of duplicate applications and improving the efficiency and quality of patent applications.
Smart Images

Figure CN116991993B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of text processing, in particular to a patent application pre-evaluation method based on similarity calculation, a storage medium and a device. BACKGROUND
[0002] According to the relevant requirements for improving the quality of university patents, conditionally qualified universities are required to accelerate the establishment of a patent pre-application evaluation system to evaluate technologies that are intended to apply for patents. At the same time, the relevant opinions on improving the evaluation mechanism of scientific and technological achievements also point out that a patent pre-application evaluation system should be established. It can be seen that patent pre-application evaluation has been increasingly mentioned.
[0003] Patent pre-application evaluation is a system arrangement and overall evaluation process for the phased evaluation of university scientific and technological achievement patent application behavior. It is also an evaluation of the invention and creation intended to be applied for a patent by a university scientific research team from the aspects of technology, market, law, etc. to determine whether to apply for a patent. Through university patent pre-application evaluation, the university scientific research management system will be restructured, which is of great significance to improving the scientific research personnel's literacy and university scientific research level, improving the quality of university patents, and promoting the landing of scientific research achievements.
[0004] In the face of the cost and risk brought by the long patent examination period and the low patent authorization rate, efficient and accurate patent pre-application evaluation is an important link to reduce repeated applications, invalid applications, save application costs, improve examination efficiency, and improve patent quality, which affects the operation efficiency of the whole chain of patent creation, use, protection, management, and service from the source. Therefore, there is an urgent need for a way to effectively evaluate patents before application. SUMMARY
[0005] In view of the defects in the prior art, the purpose of the present application is to provide a patent pre-application evaluation method based on similarity calculation, a storage medium and a device, which can provide reference support for applicants when applying for a patent.
[0006] To achieve the above purpose, the patent pre-application evaluation method based on similarity calculation provided by the present application specifically includes the following steps:
[0007] Obtain existing patent text data to construct a patent data set, and divide the patent text in the patent data set into patent fields and IPC classifications;
[0008] Based on the similarity calculation between the patent to be evaluated and the patent data set, determine the patent field and IPC classification of the patent to be evaluated;
[0009] Obtain the patent text under the IPC major group of the patent field to which the patent to be evaluated belongs, and perform similarity calculation between the obtained patent text and the patent to be evaluated to obtain a similarity calculation result;
[0010] Based on the similarity calculation result, a similarity matrix is formed, and a similarity heat map and a similarity scatter plot are obtained, so as to realize the pre-assessment of the patent application.
[0011] On the basis of the above technical scheme, the existing patent text data is acquired to construct a patent data set, and the patent texts in the patent data set are classified in the patent field and the IPC classification, and the specific steps include:
[0012] The full-field patent text data is acquired, and the abstract text and the independent claim text of the patent text are grabbed to form a patent data set;
[0013] The patent text is subjected to a natural language processing operation, and the patent data set is subjected to subject modeling in combination with an LDA model and a perplexity calculation method, so as to realize the subdivision of the patent field and obtain the patent field of the patent text in the patent data set;
[0014] Based on the patent field of the patent text, the IPC main classification of the patent text is extracted to obtain the IPC classification of the patent text in the patent data set.
[0015] On the basis of the above technical scheme, the similarity calculation between the patent to be assessed and the patent data set is performed to determine the patent field and the IPC classification of the patent to be assessed, wherein the determination of the patent field of the patent to be assessed includes:
[0016] The abstract text and the independent claim text of the patent to be assessed are extracted, and the extracted abstract text and independent claim text are subjected to a natural language processing operation to form a patent text data set to be assessed;
[0017] The patent text data set to be assessed and the patent data set are subjected to similarity calculation, the field with the highest similarity between the patent to be assessed and the patent data set is acquired, and the patent field of the patent to be assessed is determined.
[0018] On the basis of the above technical scheme, the similarity calculation between the patent to be assessed and the patent data set is performed to determine the patent field and the IPC classification of the patent to be assessed, wherein the determination of the IPC classification of the patent to be assessed includes:
[0019] Based on the patent field of the patent to be assessed, the patent text under the patent field of the patent data set is acquired;
[0020] The patent text in the acquired patent data set and the patent to be assessed are subjected to similarity calculation, and the pre-set number of patent texts in the patent data set are extracted according to the order from high to low similarity;
[0021] Based on the IPC classification of the extracted patent text, the IPC classification of the patent to be assessed is determined.
[0022] On the basis of the above technical scheme, the natural language processing operation comprises constructing a dictionary, word segmentation and stop word removal, and TF-IDF vectorization processing.
[0023] On the basis of the above technical scheme, the patent text under the IPC major group to which the to-be-evaluated patent belongs is obtained, similarity calculation is performed between the obtained patent text and the to-be-evaluated patent, and a similarity calculation result is obtained. The specific steps comprise:
[0024] In the patent data set, the abstract text and the independent claim text of the patent text under the IPC major group to which the to-be-evaluated patent belongs are obtained, and the abstract text and the independent claim text of the to-be-evaluated patent are obtained.
[0025] The abstract text of the to-be-evaluated patent is subjected to similarity calculation with the abstract text of the patent text obtained in the patent data set, and an abstract text similarity calculation result is obtained.
[0026] The independent claim text of the to-be-evaluated patent is subjected to similarity calculation with the independent claim text of the patent text obtained in the patent data set, and an independent claim text similarity calculation result is obtained.
[0027] On the basis of the above technical scheme, a similarity matrix is formed based on the similarity calculation result, and a similarity heat map and a similarity scatter plot are obtained, so as to realize pre-application evaluation of the patent. The specific steps comprise:
[0028] The abstract text similarity calculation result is set as the X axis, and the independent claim text similarity calculation result is set as the Y axis, so as to form a similarity matrix.
[0029] Based on the similarity matrix, a similarity heat map and a similarity scatter plot are generated. In combination with the similarity calculation result, the similarity heat map and the similarity scatter plot, pre-application evaluation of the patent is realized. The similarity calculation result comprises the abstract text similarity calculation result and the independent claim text similarity calculation result.
[0030] On the basis of the above technical scheme,
[0031] When the similarity calculation result is less than a first preset value, and the similarity heat map as a whole presents a light color, it is suggested that the to-be-evaluated patent is applied for a patent.
[0032] When the similarity calculation result is less than a second preset value and greater than or equal to the first preset value, and the similarity heat map as a whole presents a light color, it is suggested that the to-be-evaluated patent is modified and then applied for a patent.
[0033] When the similarity calculation result is higher than or equal to the second preset value, and the similarity heat map as a whole presents a dark color, it is suggested that the to-be-evaluated patent is not applied for.
[0034] The application provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program realizes steps of the patent pre-application evaluation method based on similarity calculation when executed by a processor.
[0035] The application provides a patent pre-application evaluation device based on similarity calculation, which comprises:
[0036] A construction module is configured to acquire existing patent text data to construct a patent data set, and divide patent texts in the patent data set according to patent fields and IPC classifications;
[0037] A determination module is configured to determine patent fields and IPC classifications of a patent to be evaluated based on similarity calculation between the patent to be evaluated and the patent data set;
[0038] A calculation module is configured to acquire patent texts under an IPC large group of a patent field to which the patent to be evaluated belongs, and perform similarity calculation between the acquired patent texts and the patent to be evaluated to obtain a similarity calculation result;
[0039] An evaluation module is configured to form a similarity matrix based on the similarity calculation result, and obtain a similarity heat map and a similarity scatter plot, so as to realize patent pre-application evaluation.
[0040] Compared with the prior art, the application has the advantages that: based on text processing, similarity calculation and the like, a patent written by an applicant but not yet applied for is preliminarily identified, functions of determining whether the patent can be directly applied for, whether the patent needs to be modified and then applied for, and whether the patent needs to be applied for are realized, the result obtained by the application can provide certain reference for the applicant before patent application, a patent pre-application evaluation process is completed, and the relevant identification result can provide reference support for the applicant when applying for a patent. BRIEF DESCRIPTION OF DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0042] Figure 1 The flowchart of the patent pre-application evaluation method based on similarity calculation in the embodiments of the present application is shown.
[0043] Figure 2 The overall flow framework diagram of the patent pre-application evaluation method based on similarity calculation is shown. DETAILED DESCRIPTION
[0044] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments of the present application.
[0045] The present application provides a patent application pre-evaluation method based on similarity calculation based on text processing, similarity calculation and the like. The patent written by the applicant but not yet applied for can be preliminarily identified to realize the functions of judging whether the patent can be directly applied for, whether the patent needs to be modified and applied for, and whether the patent needs to be applied for. The results obtained by the present application can provide certain reference for the applicant before the patent application. In the embodiments of the present application, the "artificial intelligence" field is taken as the data source of database construction, and the patent "a keyword-driven automatic testing method and a testing system thereof" is taken as the patent to be evaluated. The patent application pre-evaluation method based on similarity calculation of the present application is described in detail.
[0046] Referring to Figure 1 The patent application pre-evaluation method based on similarity calculation provided by the embodiments of the present application specifically includes the following steps:
[0047] S1: obtaining existing patent text data to construct a patent data set, and classifying the patent text in the patent data set according to the patent field and IPC (International Patent Classification, International Patent Classification);
[0048] In the present application, the existing patent text data is obtained to construct a patent data set, and the patent text in the patent data set is classified according to the patent field and IPC. The specific steps include:
[0049] S101: obtaining all-field patent text data, and grabbing the abstract text and independent claim text of the patent text to form a patent data set;
[0050] S102: performing natural language processing operation on the patent text, and combining the LDA (Latent Dirichlet Allocation, Latent Dirichlet Allocation) model and the perplexity calculation method to model the main body of the patent data set, realize the subdivision of the patent field, and obtain the patent field of the patent text in the patent data set; the natural language processing operation includes constructing a dictionary, word segmentation and stop word removal, and TF-IDF vectorization (vectorization of text data) processing.
[0051] S103: Based on the patent field of the patent text, the IPC main classification of the patent text is extracted, and the IPC classification of the patent text in the patent data set is obtained. In actual application process, the bit number of the IPC "big group" classification number is not fixed, but it can be found that the "big group" classification number is composed of a separator " / " and two "0"s, so only the characters before the separator are intercepted and connected with two "0" characters to obtain the "big group" classification number, and the IPC classification of the patent text in the patent data set is obtained.
[0052] The step S1 is described below in combination with specific examples.
[0053] (1) Data source and processing. Taking the field of "artificial intelligence" as the empirical object, the data source is the Incopat patent database, the patent type is Chinese invention authorized patent and utility model patent, the applicant is "university", the keyword search formula is ("artificial intelligence" OR "intelligent system" OR "Internet of Things" OR "human-computer interaction" OR "intelligent technology" OR "intelligent robot" OR "deep learning" OR "semantic network"), the patent legal status is "valid", the retrieval time is April 7, 2023, and the result is that 12245 "university" patents are searched, and then the patent data of "university" independent application is extracted, and the data set is de-duplicated, and finally 10396 "university" patent data sets are obtained.
[0054] (2) Technology theme modeling. Extract the "university" independent application patent "experimental data set", and perform data preprocessing, vectorization representation and technology theme extraction on the independent claim text in the patent "experimental data set". Then, by using the perplexity calculation method, the theme range is set to 1-50, the iteration number is 100, the perplexity of the LDA theme model is calculated under different theme numbers, when the theme number is 8, the perplexity is lower and the consistency index is the highest, so the number of LDA theme extraction is set to 8, the lower theme word is set to 10, the iteration number is 100, and the obtained theme and its keywords are merged, and the theme is summarized, and finally the university theme distribution is obtained, and the theme distribution is shown in Table 1.
[0055] Table 1 theme clustering
[0056]
[0057] (3) IPC major group division. In the patent document, the IPC information contained in each patent is represented by the form of "small group" classification number, so in order to obtain the corresponding major information, the patent number needs to be processed and matched in the complete IPC classification table to obtain the corresponding annotation information. In addition, the annotation information in the IPC classification table is usually concise and has many additional explanations, which is not conducive to data statistical analysis, so it needs to be processed and summarized to form a simple and easy-to-use IPC major classification table. For the processing of the IPC "major" classification number, it is found through observation that the first three digits of the fixed classification number represent it, so the first three characters of the IPC can be obtained by using the string cutting function, that is, the "major group" information of the patent, see Table 2 below.
[0058] Table 2 IPC major group classification induction (TOP10)
[0059]
[0060]
[0061] S2: determining the patent field and IPC classification of the patent to be evaluated based on the similarity calculation between the patent to be evaluated and the patent data set;
[0062] In the present application, the patent field and IPC classification of the patent to be evaluated are determined based on the similarity calculation between the patent to be evaluated and the patent data set, wherein the determination step of the patent field of the patent to be evaluated comprises:
[0063] S201: extracting the abstract text and independent claim text of the patent to be evaluated, and performing natural language processing operation on the extracted abstract text and independent claim text to form a patent text data set to be evaluated;
[0064] S202: performing similarity calculation on the patent text data set to be evaluated and the patent data set to obtain the field with the highest similarity between the patent to be evaluated and the patent data set, and determining the patent field of the patent to be evaluated.
[0065] In the present application, the patent field and IPC classification of the patent to be evaluated are determined based on the similarity calculation between the patent to be evaluated and the patent data set, wherein the determination step of the IPC classification of the patent to be evaluated comprises:
[0066] S211: obtaining the patent text under the patent field of the patent data set based on the patent field of the patent to be evaluated;
[0067] S212: performing similarity calculation on the obtained patent text in the patent data set and the patent to be evaluated, and extracting a preset number of patent texts in the patent data set according to the order from high to low similarity; the preset number can be 5.
[0068] S213: Determine the IPC classification of the patent to be evaluated based on the extracted IPC classification big groups of the patent text.
[0069] The following describes step S2 in combination with specific examples.
[0070] (1) Selection of the patent to be evaluated. The patent "A keyword-driven automatic testing method and its testing system" is selected as the patent to be evaluated.
[0071] (2) Identification of the field of the patent to be evaluated. Based on the subject clustering formed in step S1, subject identification is performed on the data in the self-built patent database, the subject type to which each patent belongs is identified, and eight subject data modules are formed. First, the abstract text and independent claim text under the eight subject modules are merged to form a subject patent data set; then, natural language processing and text vectorization operations are performed on the eight subject patent data sets to form a subject similarity calculation corpus; thirdly, the abstract text and independent claim text of the patent to be evaluated are merged and subjected to natural language processing to form a patent to be applied corpus; finally, the patent to be applied corpus is used to calculate the similarity with the subject similarity calculation corpus, and the similarity of each patent in the patent to be evaluated corpus and the subject similarity calculation corpus is obtained, and then the average similarity is calculated according to the subject to obtain the most similar subject, so as to identify the subject field of the patent to be evaluated. The identification result is shown in Table 3 below.
[0072] Table 3 Identification of the subject of the patent to be evaluated
[0073]
[0074]
[0075] It is calculated that the subject of the patent to be evaluated is Top8, and the patent data under the Top8 subject is extracted to form an IPC classification data set, which prepares data for the subsequent identification step.
[0076] (3) Identification of the IPC of the patent to be evaluated. Based on the identification of the field of the patent to be evaluated, an IPC classification data set is formed, and the similarity data of the patent text in the patent to be evaluated and the patent data set is obtained, and then the similarity values are sorted, the top 10 similarity values are extracted, and the corresponding IPC big groups are obtained, thereby completing the identification of the IPC of the patent to be evaluated. The identification result is shown in Table 4 below.
[0077] Table 4 Identification of the IPC of the patent to be evaluated
[0078] IPC major groups Similarity values G06 0.112 G06 0.105 G01 0.098 G06 0.093 A61 0.081 G06 0.077 A01 0.075 B07 0.075 H05 0.074 H04 0.071
[0079] After calculation, it is found that the IPC classification of the to-be-evaluated patent is mainly G06, G01, A61, A01, B07, H05, H04, and the patent applicant can refer to the above results when selecting the IPC of the patent application.
[0080] S3: Obtain the patent text under the IPC large group of the patent field to which the to-be-evaluated patent belongs, and perform similarity calculation between the obtained patent text and the to-be-evaluated patent to obtain a similarity calculation result.
[0081] In the present application, the patent text under the IPC large group of the patent field to which the to-be-evaluated patent belongs is obtained, and similarity calculation is performed between the obtained patent text and the to-be-evaluated patent to obtain a similarity calculation result, and the specific steps include:
[0082] S301: In the patent data set, obtain the abstract text and independent claim text of the patent text under the IPC large group of the patent field to which the to-be-evaluated patent belongs, and obtain the abstract text and independent claim text of the to-be-evaluated patent.
[0083] S302: Perform similarity calculation on the abstract text of the to-be-evaluated patent and the abstract text of the patent text obtained in the patent data set to obtain an abstract text similarity calculation result.
[0084] S303: Perform similarity calculation on the independent claim text of the to-be-evaluated patent and the independent claim text of the patent text obtained in the patent data set to obtain an independent claim text similarity calculation result.
[0085] S4: Form a similarity matrix based on the similarity calculation result, and obtain a similarity heat map and a similarity scatter plot, thereby realizing pre-patent application evaluation.
[0086] In the present application, a similarity matrix is formed based on the similarity calculation result, and a similarity heat map and a similarity scatter plot are obtained, thereby realizing pre-patent application evaluation, and the specific steps include:
[0087] S401: Set the abstract text similarity calculation result as the X-axis and the independent claim text similarity calculation result as the Y-axis to form a similarity matrix.
[0088] S402: Generate a similarity heat map and a similarity scatter plot based on the similarity matrix, and combine the similarity calculation result, the similarity heat map and the similarity scatter plot to realize pre-patent application evaluation, wherein the similarity calculation result includes the abstract text similarity calculation result and the independent claim text similarity calculation result.
[0089] In the present application, when the similarity calculation result is less than the first preset value, and the overall similarity heat map presents light color, the patent to be evaluated is recommended to apply for a patent; that is, the abstract text similarity calculation result and the independent claim text similarity calculation result are both less than the first preset value.
[0090] When the similarity calculation result is less than the second preset value, and greater than or equal to the first preset value, and the overall similarity heat map presents light color, the patent to be evaluated is recommended to be modified before applying for a patent.
[0091] When the similarity calculation result is higher than or equal to the second preset value, and the overall similarity heat map presents dark color, the patent to be evaluated is not recommended to apply for a patent.
[0092] That is, when the similarity is extremely low and the overall heat map presents light color, the patent to be evaluated is identified as a level A patent, which can directly apply for a patent; when the similarity is relatively low and the overall heat map presents relatively light color, the patent to be evaluated is identified as a level B patent, which needs to be modified before applying for a patent; when the similarity is relatively high and the overall heat map presents dark color, the patent to be evaluated is identified as a level C patent, which is not recommended to directly apply for a patent.
[0093] The steps S3 and S4 are described below in combination with examples.
[0094] Based on the subject identification result and the IPC large group identification result of step S2, it is finally determined that the subject of the patent to be evaluated is Top8, and the IPC large group is mainly G06, G01, A61, A01, B07, H05, H04. Then, based on the result, the pre-application evaluation operation is performed on the patent to be evaluated. First, based on the identified field classification and IPC "large group" of the patent to be evaluated, 403 pieces of patent data of the IPC "large group" in the field are extracted; secondly, the patent abstract and patent independent claim text data of the 403 pieces of patent data are extracted, and the abstract and independent claim text are subjected to natural language processing and text vectorization to form a pre-application evaluation corpus; next, the patent abstract and independent claim text data of the patent to be evaluated are extracted, and the natural language processing and text vectorization operation is performed to form a patent to be evaluated corpus; finally, the patent abstract, patent independent claim of the patent to be evaluated corpus, and the patent abstract and patent independent claim of the pre-application evaluation corpus are subjected to similarity calculation, the abstract similarity calculation result is set as the X axis, and the independent claim similarity calculation result is set as the Y axis, so as to form a similarity matrix, and a similarity heat map and a scatter plot are generated in combination with the matrix, and whether the patent is to be applied for a patent is identified based on the heat map and the scatter plot.
[0095] After calculation and observation of visual pictures, it can be known that the to-be-evaluated patent has low similarity with the patent application pre-evaluation corpus and high patent novelty, and can be identified as a class A patent, and the patent can be directly applied for patent.
[0096] The overall flow framework diagram of the patent application pre-evaluation method based on similarity calculation is shown in the figure Figure 2 , including patent data classification and vectorization process, to-be-evaluated patent automatic classification process, and to-be-evaluated patent application pre-evaluation process.
[0097] In a possible implementation, the embodiment of the present application also provides a non-transitory computer readable storage medium, which is located in a PLC (Programmable Logic Controller, Programmable Logic Controller) controller, and a computer program is stored on the readable storage medium, and the program is executed by a processor to realize the steps of the following patent application pre-evaluation method based on similarity calculation:
[0098] The existing patent text data is acquired to construct a patent data set, and the patent text in the patent data set is classified according to the patent field and IPC;
[0099] Based on the similarity calculation between the to-be-evaluated patent and the patent data set, the patent field and IPC classification of the to-be-evaluated patent are determined;
[0100] The patent text under the IPC large group of the patent field to which the to-be-evaluated patent belongs is acquired, the acquired patent text is subjected to similarity calculation with the to-be-evaluated patent, and a similarity calculation result is obtained;
[0101] A similarity matrix is formed based on the similarity calculation result, and a similarity heat map and a similarity scatter plot are obtained, so as to realize patent application pre-evaluation.
[0102] The storage medium can take any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus or device.
[0103] The computer-readable signal medium can include a computer-readable program code in a baseband or propagated as a carrier wave in a propagation medium. Such a propagated signal can take a wide variety of forms, including but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium that can be used to carry or store program code for use by or in connection with an instruction execution system, apparatus or device.
[0104] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0105] The embodiment of the present application provides a patent application pre-evaluation device based on similarity calculation, which comprises a construction module, a determination module, a calculation module and an evaluation module.
[0106] The construction module is used to acquire existing patent text data to construct a patent data set, and divide the patent text in the patent data set according to the patent field and IPC classification; the determination module is used to determine the patent field and IPC classification of the patent to be evaluated based on the similarity calculation between the patent to be evaluated and the patent data set; the calculation module is used to acquire the patent text under the IPC large group of the patent field to which the patent to be evaluated belongs, and perform similarity calculation between the acquired patent text and the patent to be evaluated to obtain a similarity calculation result; and the evaluation module is used to form a similarity matrix based on the similarity calculation result, and obtain a similarity heat map and a similarity scatter plot, so as to realize the pre-application evaluation of the patent.
[0107] The above only describes the specific embodiments of the present application, so that those skilled in the art can understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
[0108] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a machine that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The apparatus that implements the functions specified in one block or multiple blocks.
Claims
1. A method for pre-application patent evaluation based on similarity calculation, characterized in that, Specifically, the following steps are included: We acquire existing patent text data to construct a patent dataset, and then classify the patent texts in the dataset by patent field and IPC. Based on the similarity calculation between the patent to be evaluated and the patent dataset, the patent field and IPC classification of the patent to be evaluated are determined. Obtain the patent texts under the IPC group of the patent field to which the patent to be evaluated belongs, and calculate the similarity between the obtained patent texts and the patent to be evaluated to obtain the similarity calculation results; A similarity matrix is formed based on the similarity calculation results, and a similarity heatmap and a similarity scatter plot are obtained, thereby realizing the pre-application evaluation; The steps involved in obtaining the patent texts from the IPC group of the patent field to which the patent to be evaluated belongs, and then performing a similarity calculation between the obtained patent texts and the patent to be evaluated to obtain the similarity calculation result. These steps specifically include: In the patent dataset, obtain the abstract text and independent claim text of the patent text under the IPC group of the patent field to which the patent to be evaluated belongs; The similarity between the abstract text of the patent to be evaluated and the abstract text of the patent text obtained from the patent dataset is calculated to obtain the abstract text similarity calculation result. The similarity calculation results are obtained by comparing the independent claim text of the patent to be evaluated with the independent claim text of the patent text obtained from the patent dataset. The process of forming a similarity matrix based on the similarity calculation results, and obtaining a similarity heatmap and a similarity scatter plot to achieve pre-patent application evaluation includes the following steps: The similarity calculation results of the abstract text are set as the X-axis, and the similarity calculation results of the independent claims text are set as the Y-axis to form a similarity matrix; A similarity heatmap and a similarity scatter plot are generated based on the similarity matrix. By combining the similarity calculation results, the similarity heatmap, and the similarity scatter plot, a pre-application patent evaluation can be achieved. The similarity calculation results include the abstract text similarity calculation results and the independent claim text similarity calculation results. Among them, when the similarity calculation result is less than the first preset value and the similarity heatmap is light-colored overall, the patent to be evaluated is recommended to be applied for. If the similarity calculation result is less than the second preset value but greater than or equal to the first preset value, and the similarity heatmap is light-colored overall, then the patent application should be submitted after the patent is evaluated and the proposed modifications are made. If the similarity calculation result is higher than or equal to the second preset value, and the similarity heatmap is dark in color overall, then it is not recommended to apply for the patent to be evaluated.
2. The patent application pre-evaluation method based on similarity calculation as described in claim 1, characterized in that, The steps for acquiring existing patent text data to construct a patent dataset and classifying the patent texts in the dataset by patent field and IPC include: Acquire patent text data across all fields, and extract the abstract text and independent claim text of the patent texts to form a patent dataset; Natural language processing is performed on the patent texts, and combined with LDA model and perplexity calculation method, subject modeling is performed on the patent dataset to achieve patent field segmentation and obtain the patent field of the patent texts in the patent dataset. Based on the patent domain of the patent text, the main IPC classification of the patent text is extracted to obtain the IPC classification of the patent text in the patent dataset.
3. The patent application pre-evaluation method based on similarity calculation as described in claim 1, characterized in that, The process of determining the patent field and IPC classification of the patent to be evaluated based on the similarity calculation between the patent to be evaluated and the patent dataset includes the following steps: Extract the abstract text and independent claim text of the patent to be evaluated, perform natural language processing on the extracted abstract text and independent claim text to form a dataset of the patent text to be evaluated; The similarity between the patent text dataset to be evaluated and the patent dataset is calculated to obtain the field with the highest similarity between the patent to be evaluated and the patent dataset, thus determining the patent field of the patent to be evaluated.
4. The patent application pre-evaluation method based on similarity calculation as described in claim 3, characterized in that, The process of calculating the similarity between the patent to be evaluated and the patent dataset to determine the patent field and IPC classification of the patent to be evaluated includes the following steps: Based on the patent field of the patent to be evaluated, obtain the patent texts in that patent field from the patent dataset; The similarity between the patent texts in the acquired patent dataset and the patents to be evaluated is calculated, and a preset number of patent texts are extracted from the patent dataset according to the order of similarity from high to low. Based on the extracted patent text and IPC classification groups, the IPC classification of the patent to be evaluated is determined.
5. A pre-application patent evaluation method based on similarity calculation as described in claim 2 or 3, characterized in that: The natural language processing operations include dictionary construction, word segmentation and stop word removal, and TF-IDF vectorization.
6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the pre-application evaluation method based on similarity calculation as described in any one of claims 1 to 5.
7. A patent application pre-evaluation device based on similarity calculation, characterized in that, include: The construction module is used to acquire existing patent text data to build a patent dataset, and to classify the patent texts in the patent dataset into patent fields and IPC categories. The determination module is used to determine the patent field and IPC classification of the patent to be evaluated based on the similarity calculation between the patent to be evaluated and the patent dataset. The calculation module is used to obtain the patent text under the IPC group of the patent field to which the patent to be evaluated belongs, and to perform similarity calculation between the obtained patent text and the patent to be evaluated to obtain the similarity calculation result. The evaluation module is used to form a similarity matrix based on the similarity calculation results, and to obtain a similarity heatmap and a similarity scatter plot, thereby realizing the pre-application evaluation; The steps involved in obtaining the patent texts from the IPC group of the patent field to which the patent to be evaluated belongs, and then performing a similarity calculation between the obtained patent texts and the patent to be evaluated to obtain the similarity calculation result. These steps specifically include: In the patent dataset, obtain the abstract text and independent claim text of the patent text under the IPC group of the patent field to which the patent to be evaluated belongs; The similarity between the abstract text of the patent to be evaluated and the abstract text of the patent text obtained from the patent dataset is calculated to obtain the abstract text similarity calculation result. The similarity calculation results are obtained by comparing the independent claim text of the patent to be evaluated with the independent claim text of the patent text obtained from the patent dataset. The process of forming a similarity matrix based on the similarity calculation results, and obtaining a similarity heatmap and a similarity scatter plot to achieve pre-patent application evaluation includes the following steps: The similarity calculation results of the abstract text are set as the X-axis, and the similarity calculation results of the independent claims text are set as the Y-axis to form a similarity matrix; A similarity heatmap and a similarity scatter plot are generated based on the similarity matrix. By combining the similarity calculation results, the similarity heatmap, and the similarity scatter plot, a pre-application patent evaluation can be achieved. The similarity calculation results include the abstract text similarity calculation results and the independent claim text similarity calculation results. Among them, when the similarity calculation result is less than the first preset value and the similarity heatmap is light-colored overall, the patent to be evaluated is recommended to be applied for. If the similarity calculation result is less than the second preset value but greater than or equal to the first preset value, and the similarity heatmap is light-colored overall, then the patent application should be submitted after the patent is evaluated and the proposed modifications are made. If the similarity calculation result is higher than or equal to the second preset value, and the similarity heatmap is dark in color overall, then it is not recommended to apply for the patent to be evaluated.
Citation Information
Patent Citations
College patent personalized recommendation system
CN111259110A
Patent document evaluation task allocation method and device
CN112016830A