Technical scheme patentability evaluation method and device and storage medium
By extracting technical topic lists, feature word sets, and feature sentence data, and processing them using a vector generation model, the problem of shallow understanding of patent semantics in existing technologies is solved, more accurate technical feature extraction and sorting are achieved, and the accuracy of patentability evaluation is improved.
Patent Information
- Application Number
- CN202510746475.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-19
AI Technical Summary
The existing technology relies on keyword retrieval and text matching in patent patentability analysis, which is unable to deeply understand the semantics of patent documents, resulting in insufficient extraction of technical features and inaccurate ranking of technical similarity.
By extracting the technical subject list, technical feature word set and feature sentence data of the technical solution, using the vector generation model for vectorization processing, and combining with the preset database for similarity calculation, the patentability of the text data is determined.
It improves the depth of semantic understanding of patent documents, enhances the accuracy of technical feature extraction, and improves the accuracy of technical similarity ranking, helping users to more accurately evaluate the patentability of technical solutions.
Smart Images

Figure CN120670573A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of intelligent information processing technology, and in particular to a method, device and storage medium for evaluating the patentability of a technical solution. Background Art
[0002] Patent patentability analysis is an important task in the field of intellectual property. Its purpose is to determine whether a certain technological invention is patentable in order to decide whether to grant a patent right.
[0003] Existing patent patentability analysis systems usually rely on information retrieval and text matching technologies, which search patent documents by keywords and combine them with rule-based analysis methods to make patentability judgments.
[0004] Due to the linguistic diversity and complex descriptions of patent documents, relying solely on keywords lacks a deep understanding of the patent document's semantics, making it ineffective in processing long texts and, consequently, lacking the ability to deeply understand and extract technical features. Consequently, existing technologies are unable to accurately assess the technical similarity between patents during ranking, resulting in inaccurate ranking results. Summary of the Invention
[0005] In view of the above-mentioned solution, this application aims to propose a method, device and storage medium for evaluating the patentability of a technical solution to solve the above-mentioned technical problems.
[0006] In a first aspect, one or more embodiments of this specification provide a method for evaluating the patentability of a technical solution, including: receiving text data input by a user, wherein the text data includes a technical solution; Extracting a technical theme list, a technical feature word set, and feature sentence data of the technical solution from the text data, wherein the technical feature word set contains semantic relationships of technical features of the technical solution; Determining a first text set from a preset database according to the technical subject list and the technical feature word set; determining a second text set from the first text set based on the characteristic sentence data; Based on the second text set, the patentability of the text data is determined.
[0007] Furthermore, the patent data in the database is composed of a patent technical subject list, a patent technical feature word set, and patent feature sentence data; the patent technical subject list, the patent technical feature word set, and the patent feature sentence data correspond to the technical subject list, the technical feature word set, and the feature sentence data of the technical solution respectively; Determining a first text set from a preset database according to the technical subject list and the technical feature word set includes: After concatenating the technical theme list and the technical feature word set, vectorization is performed based on a preset vector generation model to obtain a first vector; vectorizing the patent data in the database based on the vector generation model to obtain a plurality of second vectors; determining a first similarity between the first vector and each of the second vectors respectively; Determining a second similarity between each of the patent data and the text data based on the first similarity; According to the second similarity, a first text set is determined in the database.
[0008] Furthermore, there are multiple pieces of said characteristic sentence data; Determining a second text set from the first text set according to the feature sentence data includes: For each piece of characteristic sentence data, respectively determine a third similarity between the current characteristic sentence data and each piece of the patent characteristic sentence data in the first text set; Determining text similarity between the text data and each element in the first text set based on the second similarity and the third similarity; The second text set is determined according to the text similarity.
[0009] Furthermore, the first text set includes a plurality of elements, each element corresponding to at least one third similarity; Determining text similarity between the text data and each element in the first text set based on the second similarity and the third similarity includes: For each element in the first text set, determining the similarity between the text data and the characteristic sentence of the current element according to the value range and quantity of the third similarity corresponding to the current element; The text similarity between the text data and the current element is determined according to a preset weight, the second similarity and the feature sentence similarity.
[0010] Furthermore, after receiving the text data input by the user, the method further includes: Extracting the technical effect word set and effect statement data of the technical solution from the text data; Determining a first text set from a preset database according to the technical subject list, the technical feature word set, and the technical effect word set; A second text set is determined from the first text set based on the feature sentence data and the effect sentence data.
[0011] Furthermore, the patent data in the database is composed of a patent technical subject list, a patent technical feature word set, a patent technical effect word set, patent feature statement data, and patent effect statement data; the patent technical subject list, the patent technical feature word set, the patent technical effect word set, the patent feature statement data, and the patent effect statement data correspond to the technical subject list, the technical feature word set, the technical effect word set, the feature statement data, and the effect statement data, respectively; Determining a first text set from a preset database according to the technical subject list, the technical feature word set, and the technical effect word set includes: splicing the technical subject list and the technical feature word set to obtain a first spliced set; splicing the technical subject list and the technical effect word set to obtain a second spliced set; The first spliced set and the second spliced set are combined based on a preset vector generation model to obtain a first vector; vectorizing the patent data in the database based on a preset vector generation model to obtain a plurality of second vectors; determining a first similarity between the first vector and each of the second vectors respectively; Determining a second similarity between each of the patent data and the text data based on the first similarity; According to the second similarity, a first text set is determined in the database.
[0012] Furthermore, there are a plurality of said characteristic sentence data and a plurality of said effect sentence data; Determining a second text set from the first text set according to the feature sentence data and the effect sentence data includes: For each piece of characteristic sentence data, respectively determine a third similarity between the current characteristic sentence data and each piece of the patent characteristic sentence data in the first text set; For each piece of effect statement data, respectively determining a fourth similarity between the current effect statement data and each piece of the patent effect statement data in the first text set; Determining text similarity between the text data and each element in the first text set based on the second similarity, the third similarity, and the fourth similarity; The second text set is determined according to the text similarity.
[0013] Furthermore, the first text set includes a plurality of elements, each element corresponding to at least one fourth similarity; Determining text similarity between the text data and each element in the first text set according to the second similarity, the third similarity, and the fourth similarity includes: For each element in the first text set, for each element in the first text set, determining the effect sentence similarity between the text data and the current element based on the value range and quantity of the fourth similarity corresponding to the current element; For each element in the first text set, determining the similarity between the text data and the effect statement of the current element according to the value range and quantity of the fourth similarity corresponding to the current element; The text similarity between the text data and the current element is determined according to a preset weight, the second similarity, the feature sentence similarity, and the effect sentence similarity.
[0014] In a second aspect, an embodiment of the present application provides a device for evaluating the patentability of a technical solution, comprising: A receiving module, configured to receive text data input by a user, wherein the text data includes a technical solution; an extraction module, configured to extract a technical theme list, a technical feature word set, and feature sentence data of the technical solution from the text data, wherein the technical feature word set contains semantic relationships of technical features of the technical solution; A data processing module is used to determine a first text set from a preset database based on the technical subject list and the technical feature word set; determine a second text set from the first text set based on the feature sentence data; and determine the patentability of the text data based on the second text set.
[0015] In a third aspect, an embodiment of the present application provides a storage medium for storing computer-executable instructions, characterized in that when the computer-executable instructions are executed, they implement the steps of the method for evaluating the patentability of the technical solution described in any one of the first aspects.
[0016] In a fourth aspect, an embodiment of the present application provides a method for constructing a database, comprising the following steps: Access to patent data; Based on a preset relationship extraction model, a patent technical feature word set containing a technical feature semantic relationship is extracted from the patent data; the technical feature word set contains the semantic relationship of the technical features of the technical solution; Taking the patent technology feature word set as input, a first vocabulary vector is obtained based on a preset vector generation model and stored.
[0017] Furthermore, the method further comprises: According to the patent technical feature word set, patent feature sentence data of the patent data is extracted; according to the patent feature sentence data, a feature sentence set is generated and stored; the feature sentence set is used to compare with the text input by the user to return the search results.
[0018] Furthermore, the method further comprises: According to the vector generation model, the patent feature sentence data is used to generate a first sentence vector and store it.
[0019] Furthermore, the method further comprises: Extracting technical subject keywords from the patent data; Splicing the technical subject keywords and the patent technical feature word set to form a first new set; The combination of the technical subject keywords and the patent technical feature words includes: Put the technical subject keywords in the first place and put the patent technical feature word set after the technical subject keywords.
[0020] Furthermore, the method further comprises: Extracting a set of patent technology effect words from the patent data based on a preset effect word extraction model; Taking the patent technology effect word set as input, based on the vector generation model, a second vocabulary vector is generated and stored.
[0021] Furthermore, the method further comprises: Extracting patent feature sentence data of the patent data according to the patent technical feature word set; Extracting patent effect sentence data of the patent data according to the patent technical effect word set; Generating and storing a feature statement set based on the patent feature statement data; Generating and storing an effect statement set based on the patent effect statement data; The feature statement set and the effect statement set are used to compare with the text input by the user to return the search results; Further, based on the vector generation model, the patent feature sentence data is converted into a first sentence vector and stored; Based on the vector generation model, the patent effect statement data is converted into a second statement vector and stored.
[0022] Furthermore, the method further comprises: Extracting technical subject keywords from the patent data; splicing the technical subject keywords and the patent technical feature word set to form a first new set; updating the feature word list according to the first new set; Splicing the technical subject keywords and the patent technical effect word set to form a second new set; The effect word list is updated according to the new set.
[0023] Furthermore, the technical subject keywords and the patent technical feature word set are spliced together, including: Put the technical subject keyword first and the patent technical feature word set after the technical subject keyword; The combination of the technical subject keywords and the patent technical effect word set includes: Put the technical subject keywords in the first place and put the patent technical effect word set after the technical subject keywords.
[0024] In a fifth aspect, an embodiment of the present application provides a database construction device, including: Acquisition module, used to acquire patent data; An extraction module, configured to extract a set of patent technical feature words containing semantic relationships of technical features from the patent data based on a preset relationship extraction model; the set of technical feature words contains semantic relationships of technical features of the technical solution; The data processing module takes the patent technology feature word set as input, generates a first vocabulary vector based on a preset vector generation model, and stores the first vocabulary vector.
[0025] In a sixth aspect, an embodiment of the present application provides a storage medium for storing computer-executable instructions, characterized in that the computer-executable instructions, when executed, implement the steps of the database construction method described in any one of the fourth aspects.
[0026] Compared with the existing technology, this application can at least achieve the following technical effects: This application first uses technical feature words with semantic relationships to replace existing keywords to improve the retrieval system's understanding of semantics during the retrieval process. Secondly, during the retrieval process, feature words and feature sentences are combined together to deepen the retrieval system's understanding of semantics. Finally, in order to cooperate with the above-mentioned retrieval method, the patents in the patent database are decomposed into a list of technical topics, a set of technical feature words, and a number of feature sentences to improve the retrieval accuracy of the retrieval system. Based on the above-mentioned retrieval results, users can be helped to more accurately evaluate the patentability of technical solutions. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate one or more embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] Figure 1 A flow chart of a method for evaluating the patentability of a technical solution provided for one or more embodiments of this specification; Figure 2 A schematic diagram of the structure of a patentability evaluation device for a technical solution provided in one or more embodiments of this specification; Figure 3 A flowchart of a database construction method provided in one or more embodiments of this specification; Figure 4 A schematic diagram of the structure of a database construction device provided in one or more embodiments of this specification. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this document.
[0030] Patentability is the process of determining whether an invention possesses novelty and inventiveness. Patentability evaluation involves searching prior art documents to identify the most relevant comparative documents for reference or to further evaluate the novelty and inventiveness of the relevant technical solutions based on the search results.
[0031] Therefore, the key to evaluating patentability lies in identifying the most relevant comparative documents (hereinafter referred to as comparative documents). Due to the linguistic diversity and complex descriptions of patent documents, the prior art uses keywords to create search formulas to retrieve comparative documents. However, creating search formulas requires extensive search experience, a deep understanding of patents, and numerous attempts, which are difficult for an average inventor to achieve.
[0032] Based on the above scenarios and problems, the present application embodiment provides a method for evaluating the patentability of a technical solution, such as Figure 1 As shown, including: Step 101: Receive text data input by a user.
[0033] In the embodiment of the present application, the text data is a technical solution, that is, the user can upload a technical briefing document.
[0034] Step 102: Extract the technical theme list, technical feature word set and feature sentence data of the technical solution from the text data.
[0035] In the embodiments of the present application, the technical feature word set contains the semantic relationship of the technical features of the technical solution. The semantic relationship is specifically: hierarchical relationship and whole-part relationship. For example, the briefing document records the use of metal materials, and the specific implementation method records the use of materials such as copper, iron and aluminum. The metal and copper, iron and aluminum are in a hierarchical relationship. The briefing document records that a device includes component 1 and component 2. The device, component 1 and component 2 are in a whole-part relationship. The specific process of introducing semantic relationships into the technical feature word set is: the subordinates of entity A are B and C, and the subordinate entities of B are D and E. When constructing the technical feature word set, the corresponding combination form is written into the set. The corresponding combination forms are ABC, ABDE, ABCDE, and BDE.
[0036] Step 103: Determine a first text set from a preset database according to the technical subject list and the technical feature word set.
[0037] In the examples of this application, the initial search for technical solutions is based on the semantic understanding of feature words. It should be noted that the semantic relationship of the technical feature words in this application is not the same as keyword expansion in the prior art. Keyword expansion in the prior art is based on word meaning, while the semantic relationship in this application is actually present in the briefing document. This also makes the search results of this application more accurate than keyword expansion.
[0038] Step 104: Determine a second text set from the first text set based on the feature sentence data.
[0039] In this embodiment, a group of comparative documents is first screened based on characteristic words. The correlation between these comparative documents and the technical features in the briefing document is then examined, further narrowing the search results. This results in comparative documents that are closest to the technical solution in the briefing document.
[0040] Step 105: Determine the patentability of the text data based on the second text set.
[0041] In the embodiment of the present application, the user can directly compare the briefing document with the comparative documents in the second text set. Alternatively, the customer input text and the comparative documents in the second text set can be used as part of the prompt, and similar feature long sentences and effect long sentences can be used as the other part of the prompt. According to the professional instructions in the fine-tuning process, the prompts are combined into the final large model input. The output is a reference opinion on the novelty analysis of the customer input text and the existing patent, including the similar features and similar effects of the comparative texts, the purpose achieved, and the differences that exist. The conclusion is only used as a reference for the user's novelty judgment and is not a final conclusion.
[0042] Among them, the large model can be open source models such as qwen (Qianwen Model) and chatglm (Chat Generative Language Model). After LORA fine-tuning (Low-Rank Adaptation of Large Language Models), the large model's answers are more inclined to similar feature analysis and difference discovery and differentiation analysis. The fine-tuning corpus is based on the corpus pairs after retrieval, and professional teachers construct the output comparative analysis text, including content format and professional terminology.
[0043] LoRA is a lightweight fine-tuning method designed to reduce the computational overhead and storage costs of fine-tuning large models. Its main functions include: 1. Reduce trainable parameters: Traditional full-parameter fine-tuning requires adjusting the weights of the entire model, while LoRA only adds low-rank matrices to specific layers (such as the weight matrix of the Transformer layer), so that only these small-scale parameters need to be updated during training without modifying the original model parameters.
[0044] 2. Save computing resources: Since LoRA only trains low-rank matrices, the required GPU memory and computing resources are greatly reduced, making it suitable for fine-tuning large models, especially in resource-limited environments.
[0045] 3. Maintain model generalization ability: LoRA only adjusts some parameters so that the model can adapt to new tasks while still retaining the knowledge of the original pre-trained model, avoiding the overfitting problem.
[0046] 4. Suitable for multi-task fine-tuning: When training different tasks through LoRA, you can only store different LoRA adapters instead of the entire model, which makes it easy to quickly switch tasks.
[0047] In the embodiment of the present application, in order to ensure the accuracy of the search results, the patents are stored in the database in the form of a patent technical subject list, a patent technical feature word set, and patent feature sentence data. The patent technical subject list, the patent technical feature word set, and the patent feature sentence data correspond to the technical subject list, the technical feature word set, and the feature sentence data of the technical solution, respectively.
[0048] In the embodiment of the present application, the specific process of determining the first text set is: After splicing the technical theme list and the technical feature word set, vectorization is performed based on a preset vector generation model to obtain a first vector; the patent data in the database is vectorized based on the vector generation model to obtain multiple second vectors; and the first similarity between the first vector and each second vector is determined respectively. Based on the first similarity, the second similarity between each patent data and the text data is determined respectively. Based on the second similarity, the first text set is determined in the database. Specifically, when splicing the technical theme list and the technical feature word set, the splicing is carried out in the order of the technical theme list first and the technical feature word set later. The above can improve the accuracy of calculating similarity.
[0049] For example, a first vector is input into the Faiss vector database for search. A nearest neighbor (NN) search is performed within the Faiss vector database to find the 30,000 most similar second vectors. The Euclidean distance between the first and second vectors is calculated. Based on this Euclidean distance, the similarity between the reference document and the briefing document corresponding to the second vector is calculated, i.e., the second similarity. Based on the second similarity, the corresponding reference document is selected to obtain the first text set.
[0050] The vector generation model can improve vector matching accuracy. This model is obtained by fine-tuning the general model. The specific fine-tuning process is as follows: based on the training samples, the query vector, pos (positive sample), and neg (neg sample) are constructed. The query, pos, and neg are input into the general model. The scores of pos and neg are calculated. The positive samples are regarded as the correct classification and the cross-entropy loss is calculated. The model parameters are then updated through backpropagation.
[0051] In an embodiment of the present application, there are multiple feature sentence data, so after determining the first text set, for each feature sentence data, the third similarity between the current feature sentence data and each patent feature sentence data in the first text set is determined. Based on the second similarity and the third similarity, the text similarity between the text data and each element in the first text set is determined. Based on the text similarity, the second text set is determined. The second similarity corresponds to the similarity between technical feature words, and the third similarity corresponds to the similarity between technical feature sentences. Through the retrieval process of words first and sentences later, the retrieval system advances the semantic understanding layer by layer, from shallow to deep, from easy to difficult, and from simple to complex, thereby overcoming the problem of inaccurate retrieval caused by the language diversity and description complexity of patent documents.
[0052] In the embodiments of the present application, one technical feature may correspond to multiple feature statements, which may result in the number of feature statements for some unimportant or even non-essential features being greater than the number of feature statements for necessary technical features. For example, in the process of writing the specification, the agent may introduce some non-essential features in order to facilitate the explanation of the technical solution. Therefore, these non-essential features are only for illustrating the technical solution, but cannot limit the technical solution. However, the above method will result in a relatively large number of feature statements corresponding to these non-essential features. In view of this situation, if the weight of each feature is simply set according to the number of feature statements, and then the similarity of the technical solutions in the comparative document and the disclosure book is determined according to the set weight, it may result in a large deviation between the technical solutions in the comparative document and the disclosure book that are finally retrieved.
[0053] In order to solve the above problem, for each element in the first text set, the feature sentence similarity between the text data and the current element is determined based on the value range and quantity of the third similarity corresponding to the current element; and the text similarity between the text data and the current element is determined based on the preset weight, the second similarity and the feature sentence similarity. Among them, each element in the first text set is a patent document, and each patent document contains multiple feature sentences, and therefore corresponds to multiple third similarities. When calculating the feature sentence similarity of the entire patent document, the value range and quantity of each similarity should be comprehensively considered to avoid a large deviation between the search results and the technical solution in the briefing document.
[0054] Specifically, from at least one third similarity corresponding to the current element, the third similarity that does not reach the threshold is deleted; when the number of remaining third similarities is greater than 0, the maximum similarity is selected as the feature sentence similarity; when the number of remaining third similarities is equal to 0, the feature sentence similarity is a preset value.
[0055] In an embodiment of the present application, in order to deepen the system's understanding of semantics and improve the accuracy of search results, the technical effect word set and effect statement data of the technical solution are extracted from the text data; a first text set is determined from a preset database based on the technical theme list, the technical feature word set, and the technical effect word set; and a second text set is determined from the first text set based on the feature statement data and the effect statement data. Since patentability mainly considers novelty, creativity, and practicality, when comparing novelty, substantially identical technical effects are a necessary condition. Therefore, adding a technical effect comparison can improve the accuracy of search results.
[0056] In the embodiments of the present application, after the technical effects are added, the data structure of the database will change. Specifically, the patent data in the database consists of a patent technical subject list, a patent technical feature word set, a patent technical effect word set, patent feature sentence data, and patent effect sentence data; the patent technical subject list, patent technical feature word set, patent technical effect word set, patent feature sentence data, and patent effect sentence data correspond to the technical subject list, technical feature word set, technical effect word set, feature sentence data, and effect sentence data, respectively.
[0057] At the same time, the comparison process will also have corresponding changes. Specifically, the technical subject list and the technical feature word set are spliced to obtain a first spliced set; the technical subject list and the technical effect word set are spliced to obtain a second spliced set; the first spliced set and the second spliced set are combined to obtain a first vector based on a preset vector generation model; the patent data in the database are vectorized based on the preset vector generation model to obtain multiple second vectors; the first similarity between the first vector and each second vector is determined respectively; based on the first similarity, the second similarity between each patent data and the text data is determined respectively; based on the second similarity, the first text set is determined in the database. Among them, for the splicing process, the splicing order is: first the technical subject list, then the technical feature word set; first the technical subject list, then the technical effect word set.
[0058] In the embodiment of the present application, after adding the technical effect, it is necessary to determine the overall similarity between the technical solution in the disclosure document and the comparative document based on the similarity of the technical effect words, the similarity of the technical feature words, the similarity of the technical feature sentences, and the similarity of the technical effect sentences. Specifically, for each feature sentence data, the third similarity between the current feature sentence data and the feature sentence data of each patent in the first text set is determined; for each effect sentence data, the fourth similarity between the current effect sentence data and the effect sentence data of each patent in the first text set is determined; based on the second similarity, the third similarity, and the fourth similarity, the text similarity between the text data and each element in the first text set is determined; and based on the text similarity, the second text set is determined.
[0059] In the embodiment of the present application, in the process of calculating similarity, technical effect sentences and technical feature sentences have the same problem, that is, there will be more technical effect sentences corresponding to non-core technical effects. If the similarity is simply superimposed, it will cause the final search results to deviate from the technical solutions in the briefing document. In order to solve the problem of technical effect sentences, for each element in the first text set, for each element in the first text set, according to the value range and number of the fourth similarity corresponding to the current element, the similarity between the feature sentence of the text data and the current element is determined; for each element in the first text set, according to the value range and number of the fourth similarity corresponding to the current element, the similarity between the effect sentence of the text data and the current element is determined; according to the preset weight, the second similarity, the similarity of the feature sentence and the similarity of the effect sentence, the text similarity between the text data and the current element is determined.
[0060] Specifically, the fourth similarities are classified according to the identification of the patent data in the database; wherein, the current element corresponds to at least one fourth similarity; from the at least one fourth similarity corresponding to the current element, the fourth similarities that do not reach the threshold are deleted; when the number of remaining fourth similarities is greater than 0, the maximum similarity is selected as the effect statement similarity; when the number of remaining fourth similarities is equal to 0, the effect statement similarity is a preset value.
[0061] It should be noted that after adding the technical effect, the processing method of the technical feature statement in the similarity calculation process is the same as the above method.
[0062] In the embodiment of the present application, the specific calculation process of the third similarity and the fourth similarity is: Step 1: Generate the matrix.
[0063] Convert the technical feature statement data and technical effect statement data into matrices. Specifically, ensure that the input matrix is at least two-dimensional. If the input is a one-dimensional vector (e.g., a single list), it will be converted to a two-dimensional array to facilitate matrix calculations. If the input is already a two-dimensional array, this step does not modify it.
[0064] Step 2: Normalize each row of the matrix Next, we need to normalize each row of each matrix. The purpose of normalization is to convert the vectors in each row into unit vectors, so that the modulus of each vector is 1. In this way, the calculation of cosine similarity will only focus on the direction of the vector and will not be affected by its magnitude (norm).
[0065] The specific approach is to calculate the L2 norm of each row vector (that is, the modulus of the vector), and then divide the vector of each row by its own norm to obtain a unit vector.
[0066] Step 3: Calculate the cosine similarity between the row vectors of the two matrices Computes the cosine similarity between two matrices by taking the dot product between every pair of row vectors of matrix A and matrix B.
[0067] Since each row of the matrix has been normalized, the cosine similarity is the dot product of the two vectors. The result of the dot product can be directly used as a similarity measure, ranging from 0 to 1.
[0068] Step 4: Get the similarity matrix All cosine similarities calculated by dot product will be stored in a similarity matrix. Each element of this matrix represents the similarity between the corresponding row vectors in the two matrices. The dimension of the matrix is m×n, where m is the number of rows in matrix A and n is the number of rows in matrix B.
[0069] The present application embodiment provides a device for evaluating the patentability of a technical solution, such as Figure 2 As shown, including: The receiving module 201 is configured to receive text data input by a user, wherein the text data includes a technical solution; An extraction module 202 is configured to extract a technical theme list, a technical feature word set, and feature sentence data of the technical solution from the text data, wherein the technical feature word set contains semantic relationships of technical features of the technical solution; The data processing module 202 is used to determine a first text set from a preset database based on the technical subject list and the technical feature word set; determine a second text set from the first text set based on the feature sentence data; and determine the patentability of the text data based on the second text set.
[0070] An embodiment of the present application provides a storage medium for storing computer-executable instructions, characterized in that the computer-executable instructions, when executed, implement the steps of the method for evaluating the patentability of the technical solution described in any one of the embodiments.
[0071] The present application embodiment provides a database construction method, such as Figure 3 As shown, the following steps are included: Step 301: Obtain patent data.
[0072] Step 302: Based on a preset relationship extraction model, a set of patent technical feature words containing technical feature semantic relationships is extracted from the patent data.
[0073] In the embodiment of the present application, the technical feature word set contains the semantic relationship of the technical features of the technical solution. The semantic relationship is specifically: hierarchical relationship and whole-part relationship. For example, the disclosure document records the use of metal materials, and the specific implementation method records the use of materials such as copper, iron and aluminum. Metal and copper, iron and aluminum constitute a hierarchical relationship. The disclosure document records that a device includes component 1 and component 2. The device, component 1 and component 2 are in a whole-part relationship. The specific process of introducing semantic relationships into the technical feature word set is: the subordinate entities of A are B and C, and the subordinate entities of B are D and E. When constructing the technical feature word set, the corresponding combination form is written into the set, and the corresponding combination form is ABC, ABDE, ABCDE, BDE. Among them, the patent technical feature word set contains multiple subject words, which are extracted from the patent data by the relation extraction model.
[0074] Step 303: Taking the patent technology feature word set as input, generating a model based on a preset vector, obtaining a first vocabulary vector, and storing the first vocabulary vector.
[0075] In an embodiment of the present application, the vector generation model converts patent technical feature words into vectors based on similarity optimization.
[0076] Similarity optimization involves three steps: lowering the similarity of low-similarity samples to avoid excessive clustering. Increasing the similarity of high-similarity samples to improve matching accuracy. For samples with medium similarity, based on semantic understanding, the similarity of partially related but incomplete matches is increased to improve the model's ability to distinguish keywords. Specifically, due to the lack of specialized contrastive learning, pre-trained models may still have high similarities between many unrelated texts, necessitating similarity optimization. For example, the theoretical similarity between "deep learning optimization method" and "solar cell materials" should be very low (close to 0), but a pre-trained model may yield a similarity between 0.4-0.6, resulting in insufficient discrimination. Theoretically, "convolutional neural network" and "CNN deep learning model" should have a high similarity (close to 1), but an untuned BGE may only yield a similarity between 0.7-0.8, resulting in insufficient recall.
[0077] It should be noted that the database of this application is used for patent retrieval, and the vector generation model described in step 3 is required during the search. In order to improve the accuracy of the search, this application does not simply improve the pre-trained model parameters, but adjusts the vector generation method to improve the accuracy of the search results based on the database of this application.
[0078] The specific effects after similarity optimization are: Example 1 The similarity between negative samples (dissimilar texts) is suppressed, and the corresponding similarity distribution range is 0.0-0.2, making the representation of low similarity interval more in line with actual needs. As shown in Table 1: Table 1 Comparison of fine-tuning effects of Example 1
[0079] Example 2 The similarity of positive samples (similar texts) is improved, and the corresponding similarity distribution range is between 0.9 and 1.0, making the matching more accurate. As shown in Table 2: Table 2 Comparison of fine-tuning effects of Example 2
[0080] Example 3. The similarity distribution of the pre-trained model may be too concentrated between 0.4 and 0.7, causing many samples to fall into a fuzzy range and making it difficult to effectively distinguish them. Training can be done by adding some relevant but not completely matching text. This is shown in Table 3: Table 3 Comparison of fine-tuning effects of Example 3
[0081] By using the above method, the data in the patent document is decomposed into characteristic words with semantic relationships, so that the search system can have a deeper semantic understanding of the patent document, which is conducive to users to quickly find relevant patents based on keywords. At the same time, the accuracy of the search results is improved by adjusting the vector based on semantic and text similarity. In this embodiment of the present application, to further enhance the search system's understanding of the semantics of patent documents, the patent documents are broken down into technical feature statements. Specifically, patent feature statement data is extracted from the patent data based on the patent technical feature word set. Based on the patent feature statement data, a feature statement set is generated and stored. This feature statement set is then compared with the text entered by the user to facilitate the return of search results.
[0082] In this embodiment of the present application, to improve search speed, technical subject keywords are extracted from the patent data; these technical subject keywords are then combined with the patent technical feature word set to form a first new set. The technical subject keywords can more quickly match the technical solutions in the comparative documents and the briefing document.
[0083] In the embodiment of the present application, in order to further improve the search speed, the technical subject keyword is placed in the first place, and the patent technical feature word set is placed after the technical subject keyword.
[0084] In this embodiment of the present application, a feature sentence set is converted into a first sentence vector. All first vocabulary vectors and first sentence vectors are then indexed and constructed using the Faiss vector library, including setting parameters such as the indexing method and number of clusters. This is intended to accelerate the search process and reduce hardware usage.
[0085] Specifically, it mainly involves setting the indexing method to IndexIVFPQ (Inverted File Vector Quantization Perceptual Query) and the number of clusters nlist parameter.
[0086] •nlist (number of cluster centers): determines the number of clusters in the inverted index. The recommended value is data volume / 100.
[0087] •nprobe (number of search clusters): controls the number of clusters traversed during the search. A higher value results in higher recall but lower speed.
[0088] FAISS is a vector library mainly used for: 1. Efficient similarity search: Quickly find similar vectors in massive vectors, such as semantic search and image retrieval.
[0089] 2. Data compression storage: The IVFPQ structure reduces storage requirements and is suitable for large-scale data scenarios.
[0090] 3. Approximate Nearest Neighbor (ANN) Search: The Hierarchical Navigable Small World graphs (HNSW) and IVF structures enable efficient approximate search and improve search speed.
[0091] 4. Support GPU acceleration: suitable for large-scale deep learning inference scenarios.
[0092] In this embodiment of the present application, to further enhance the retrieval system's semantic understanding of patent documents, the technical data is broken down into technical feature words and technical effect words. Specifically, based on a preset effect word extraction model, a set of patent technical effect words is extracted from the patent data. Using this set of patent technical effect words as input, a second vocabulary vector is generated and stored based on the vector generation model.
[0093] In an embodiment of the present application, in order to further enhance the search system's semantic understanding of patent documents, the patent documents are decomposed into patent feature statements and patent effect statements based on the decomposition into technical feature words and technical effect words. Specifically, based on the patent technical feature word set, patent feature statement data of the patent data is extracted; based on the patent technical effect word set, patent effect statement data of the patent data is extracted; based on the patent feature statement data, a feature statement set is generated and stored; based on the patent effect statement data, an effect statement set is generated and stored; the feature statement set and the effect statement set are used to compare with the text input by the user to facilitate the return of search results.
[0094] Preferably, the patent feature sentence data is converted into a first sentence vector based on the vector generation model and stored; and the patent effect sentence data is converted into a second sentence vector based on the vector generation model and stored. This method can improve the matching accuracy of the similarity between sentences.
[0095] In the embodiment of this application, after adding effect keywords, it is necessary to form a new set for the effect keywords and subject words to enhance the search system's semantic understanding of the patent document. Specifically, technical subject keywords are extracted from the patent data; the technical subject keywords are concatenated with the patent technical feature word set to form a first new set; the feature word list is updated based on the first new set; the technical subject keywords are concatenated with the patent technical effect word set to form a second new set; and the effect word list is updated based on the second new set.
[0096] In the embodiment of this application, the technical subject keyword is placed first, and the patent technical feature word set is placed after the technical subject keyword. The technical subject keyword is placed first, and the patent technical effect word set is placed after the technical subject keyword. This method can improve the search accuracy during the search.
[0097] In this embodiment of the present application, the feature sentence set and the effect sentence set are converted into a first sentence vector and a second sentence vector, and the effect words are converted into a second vocabulary vector. All vocabulary vectors and sentence vectors are then indexed using the Faiss vector library, including setting parameters such as the indexing method and the number of clusters. This aims to accelerate the search process and reduce hardware usage.
[0098] Specifically, it mainly involves setting the indexing method to IndexIVFPQ (Inverted File Vector Quantization Perceptual Query) and the number of clusters nlist parameter.
[0099] •nlist (number of cluster centers): determines the number of clusters in the inverted index. The recommended value is data volume / 100.
[0100] •nprobe (number of search clusters): controls the number of clusters traversed during the search. A higher value results in higher recall but lower speed.
[0101] FAISS is a vector library mainly used for: 1. Efficient similarity search: Quickly find similar vectors in massive vectors, such as semantic search and image retrieval.
[0102] 2. Data compression storage: The IVFPQ structure reduces storage requirements and is suitable for large-scale data scenarios.
[0103] 3. Approximate Nearest Neighbor (ANN) Search: The Hierarchical Navigable Small World graphs (HNSW) and IVF structures enable efficient approximate search and improve search speed.
[0104] 4. Support GPU acceleration: suitable for large-scale deep learning inference scenarios.
[0105] The present application embodiment provides a database construction device, such as Figure 4 As shown, including: Acquisition module 401, for acquiring patent data; Extraction module 402, for extracting a set of patent technical feature words containing semantic relationships of technical features from the patent data based on a preset relationship extraction model; the set of technical feature words contains semantic relationships of technical features of the technical solution; The data processing module 403 takes the patent technology feature word set as input, generates a first vocabulary vector based on a preset vector generation model, and stores the first vocabulary vector.
[0106] An embodiment of the present application provides a storage medium for storing computer-executable instructions, characterized in that the computer-executable instructions, when executed, implement the steps of the method for constructing a technical solution database in any one of the embodiments.
[0107] It should be noted that the embodiment of the storage medium in this specification and the embodiment of the blockchain-based service provision method in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding blockchain-based service provision method mentioned above, and the repeated parts will not be repeated.
[0108] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0109] In the 1930s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using physical hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly performed using software called a "logic compiler." This is similar to the software compilers used during program development. Before compilation, the original code must be written in a specific programming language, called a Hardware Description Language (HDL). There are many types of HDL, including ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that simply by programming a method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0110] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC625D, Atmel AT91SAM, Microchip PIC18F26K20, Silicone Labs C8051F320. The memory controller can also be implemented as part of the memory control logic. Those skilled in the art will also appreciate that, in addition to implementing the controller purely in computer-readable program code, the controller can also be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, an embedded microcontroller, etc. by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing the various functions included therein can also be considered as structures within the hardware component. Alternatively, the means for implementing the various functions can be considered both a software module implementing the method and a structure within the hardware component.
[0111] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0112] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0113] Those skilled in the art will appreciate that one or more embodiments of this specification may be provided as a method, system, or computer program product. Thus, one or more embodiments of this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0114] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of the processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0115] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0116] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0117] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0118] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0119] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0120] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0121] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0122] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0123] The foregoing description is merely an example of the present invention and is not intended to limit the present invention. Persons skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims herein.
Claims
1. A method for evaluating the patentability of a technical solution, characterized in that: include: receiving text data input by a user, wherein the text data includes a technical solution; Extracting a technical theme list, a technical feature word set, and feature sentence data of the technical solution from the text data, wherein the technical feature word set contains semantic relationships of technical features of the technical solution; Determining a first text set from a preset database according to the technical subject list and the technical feature word set; determining a second text set from the first text set based on the characteristic sentence data; Based on the second text set, the patentability of the text data is determined.
2. The method according to claim 1, characterized in that The patent data in the database is composed of a patent technical subject list, a patent technical feature word set and patent feature sentence data; the patent technical subject list, the patent technical feature word set and the patent feature sentence data correspond to the technical subject list, the technical feature word set and the feature sentence data of the technical solution respectively; Determining a first text set from a preset database according to the technical subject list and the technical feature word set includes: After concatenating the technical theme list and the technical feature word set, vectorization is performed based on a preset vector generation model to obtain a first vector; vectorizing the patent data in the database based on the vector generation model to obtain a plurality of second vectors; determining a first similarity between the first vector and each of the second vectors respectively; Determining a second similarity between each of the patent data and the text data based on the first similarity; According to the second similarity, a first text set is determined in the database.
3. The method according to claim 2, characterized in that There are multiple pieces of the characteristic sentence data; Determining a second text set from the first text set according to the feature sentence data includes: For each piece of characteristic sentence data, respectively determine a third similarity between the current characteristic sentence data and each piece of the patent characteristic sentence data in the first text set; Determining text similarity between the text data and each element in the first text set based on the second similarity and the third similarity; The second text set is determined according to the text similarity.
4. The method according to claim 3, characterized in that The first text set includes a plurality of elements, each element corresponding to at least one third similarity; Determining text similarity between the text data and each element in the first text set based on the second similarity and the third similarity includes: For each element in the first text set, determining the similarity between the text data and the characteristic sentence of the current element according to the value range and quantity of the third similarity corresponding to the current element; The text similarity between the text data and the current element is determined according to a preset weight, the second similarity and the feature sentence similarity.
5. The method according to claim 1, wherein After receiving text data input by the user, the method further includes: Extracting the technical effect word set and effect statement data of the technical solution from the text data; Determining a first text set from a preset database according to the technical subject list, the technical feature word set, and the technical effect word set; A second text set is determined from the first text set based on the feature sentence data and the effect sentence data.
6. The method according to claim 5, characterized in that The patent data in the database is composed of a patent technical subject list, a patent technical feature word set, a patent technical effect word set, patent feature sentence data, and patent effect sentence data; the patent technical subject list, the patent technical feature word set, the patent technical effect word set, the patent feature sentence data, and the patent effect sentence data correspond to the technical subject list, the technical feature word set, the technical effect word set, the feature sentence data, and the effect sentence data, respectively; Determining a first text set from a preset database according to the technical subject list, the technical feature word set, and the technical effect word set includes: splicing the technical subject list and the technical feature word set to obtain a first spliced set; splicing the technical subject list and the technical effect word set to obtain a second spliced set; The first spliced set and the second spliced set are combined based on a preset vector generation model to obtain a first vector; vectorizing the patent data in the database based on a preset vector generation model to obtain a plurality of second vectors; determining a first similarity between the first vector and each of the second vectors respectively; Determining a second similarity between each of the patent data and the text data based on the first similarity; According to the second similarity, a first text set is determined in the database.
7. The method according to claim 6, characterized in that There are a plurality of said feature sentence data and a plurality of said effect sentence data; Determining a second text set from the first text set according to the feature sentence data and the effect sentence data includes: For each piece of characteristic sentence data, respectively determine a third similarity between the current characteristic sentence data and each piece of the patent characteristic sentence data in the first text set; For each piece of effect statement data, respectively determining a fourth similarity between the current effect statement data and each piece of the patent effect statement data in the first text set; Determining text similarity between the text data and each element in the first text set based on the second similarity, the third similarity, and the fourth similarity; The second text set is determined according to the text similarity.
8. The method according to claim 7, characterized in that The first text set includes a plurality of elements, each element corresponding to at least one fourth similarity; Determining text similarity between the text data and each element in the first text set according to the second similarity, the third similarity, and the fourth similarity includes: For each element in the first text set, for each element in the first text set, determining the effect sentence similarity between the text data and the current element based on the value range and quantity of the fourth similarity corresponding to the current element; For each element in the first text set, determining the similarity between the text data and the effect statement of the current element according to the value range and quantity of the fourth similarity corresponding to the current element; The text similarity between the text data and the current element is determined according to a preset weight, the second similarity, the feature sentence similarity, and the effect sentence similarity.
9. A device for evaluating the patentability of a technical solution, characterized in that: include: A receiving module, configured to receive text data input by a user, wherein the text data includes a technical solution; an extraction module, configured to extract a technical theme list, a technical feature word set, and feature sentence data of the technical solution from the text data, wherein the technical feature word set contains semantic relationships of technical features of the technical solution; A data processing module is used to determine a first text set from a preset database based on the technical subject list and the technical feature word set; determine a second text set from the first text set based on the feature sentence data; and determine the patentability of the text data based on the second text set.
10. A storage medium for storing computer-executable instructions, characterized in that: When executed, the computer executable instructions implement the steps of the method for evaluating the patentability of the technical solution according to any one of claims 1 to 8.
Citation Information
Patent Citations
Patent harmful effect knowledge mining method and device, equipment and storage medium
CN112182183A
Intelligent retrieval method and device for calculating patent literature similarity based on word frequency and semantics, electronic equipment and storage medium thereof
CN112257419A
Patent duplicate checking method and device and electronic equipment
CN115905505A
Patent retrieval system combining patent images and text semantics
CN118113810A