Related Patent Recommendation Method, Device and Storage Medium Based on Semantic Understanding Model

Through the patent recommendation method based on semantic understanding model, the dependence on surface text information matching in the prior art is solved, and more accurate and comprehensive relevant patent recommendations are achieved, which improves the user's search experience and satisfaction.

CN119025665BActive Publication Date: 2025-06-17SHENZHEN ZHULIAN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410870090.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2025-06-17
Estimated Expiration
2044-06-28

AI Technical Summary

Technical Problem

The existing patent search system relies on surface text information matching and cannot effectively capture the context and deep semantic information of patent text, resulting in inaccurate and comprehensive recommendations of relevant patents.

Method used

Using relevant patent recommendation methods based on semantic understanding model, the basic patents and candidate patents are sliced ​​and semantic feature extraction, the target segments are queried and marked, and the candidate patents are sorted and recommended according to the number of target segments.

Benefits of technology

It realizes more accurate and comprehensive relevant patent recommendations, which can capture deep semantic information and complex semantic relationships, and the recommended patents are more in line with users' technical comparison and search needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119025665B_ABST
    Figure CN119025665B_ABST
Patent Text Reader

Abstract

This application relates to artificial intelligence technology, and discloses a related patent recommendation method, computer device, and storage medium based on a semantic understanding model. The method includes: screening out a patent set that matches the key information; performing document slicing on the basic patent to obtain a first text segment set; performing document slicing on the candidate patents in the patent set to obtain a second text segment set; based on the semantic understanding model, extracting first semantic features corresponding to each text segment in the first text segment set, and extracting second semantic features corresponding to each text segment in the second text segment set; querying the second semantic features that match the first semantic features, and marking the text segments associated with the queried second semantic features as target text segments; sorting the candidate patents according to the number of target text segments; and recommending the candidate patents whose sorting meets the preset positions as related patents of the basic patent. This application also discloses a control device. This application aims to more accurately and comprehensively implement the recommendation of related patents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a related patent recommendation method, a control device, a computer device, and a computer-readable storage medium based on a semantic understanding model. Background Art

[0002] Currently, in order to facilitate users to compare patents, a patent analysis system (or a patent retrieval platform) generally provides a related patent (or similar patent) recommendation service for the selected patent, so as to facilitate users to quickly query related patents and better perform patent comparison and analysis.

[0003] However, the existing retrieval and query of related patents generally use the selected patent as the basic patent, extract the keywords of the basic patent, and index and query related patents in the patent database; or calculate the similarity between the basic patent and the candidate patent through word sequence encoding to query related patents. However, these methods of querying related patents mainly rely on the matching of surface text information, ignoring the context and content of the patent text, unable to consider complex semantic relationships and deep semantic information, and can only retrieve related patents with consistent text descriptions, while ignoring those patents that, although different in text expression, are similar in concept and implementation method, thus making it difficult to recommend related patents more accurately and comprehensively.

[0004] The above content is only used to assist in understanding the technical solution of the present application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of the present application is to provide a related patent recommendation method, a control device, a computer device, and a computer-readable storage medium based on a semantic understanding model, aiming to more accurately and comprehensively realize the recommendation of related patents.

[0006] To achieve the above object, the present application provides a related patent recommendation method based on a semantic understanding model, including the following steps:

[0007] According to the key information of the basic patent, screen out a patent set that matches the key information;

[0008] Perform document slicing on the basic patent to obtain a first text segment set; and perform document slicing on the candidate patents in the patent set to obtain a second text segment set;

[0009] Based on the semantic understanding model, extract the first semantic features corresponding to each text segment in the first text segment set, and extract the second semantic features corresponding to each text segment in the second text segment set;

[0010] Query for the second semantic feature that matches the first semantic feature, and mark the passage associated with the queried second semantic feature as the target passage;

[0011] Rank the candidate patents according to the number of target passages;

[0012] Recommend the candidate patents whose ranking meets the preset position as the related patents of the basic patent.

[0013] Optionally, the semantic understanding model includes an input layer, an intermediate layer, and an output layer; among them, the input layer is constructed based on the word embedding technology, and the intermediate layer is constructed based on the BERT (Bidirectional Encoder Representations from Transformers) network;

[0014] The input layer is used to convert the passage into a word embedding vector;

[0015] The intermediate layer is used to perform forward propagation on the word embedding vector, and perform a series of non-linear transformations and feature extractions through multiple network structures;

[0016] The output layer is used to perform a feature aggregation operation on the features extracted by the intermediate layer to generate semantic features.

[0017] Optionally, the passages in the first passage set and the second passage set are divided into key passages and other passages except the key passages;

[0018] The steps of querying for the second semantic feature that matches the first semantic feature and marking the passage associated with the queried second semantic feature as the target passage include:

[0019] According to the first semantic feature corresponding to the key passage in the first passage set, query the second semantic feature corresponding to the key passage in the second passage set that matches;

[0020] Mark the key passage associated with the queried second semantic feature as the target passage;

[0021] Query whether there is a matching first semantic feature for the second semantic feature corresponding to other passages in the candidate patent to which the target passage belongs;

[0022] If so, mark the other passages as the target passages as well.

[0023] Optionally, the related patent recommendation method based on the semantic understanding model further includes:

[0024] If the similarity between the first semantic feature and the second semantic feature is greater than the preset threshold, it is determined that the first semantic feature and the second semantic feature match;

[0025] Among them, the preset threshold used when querying the key text segment is marked as the first threshold; the preset threshold used when querying other text segments is marked as the second threshold; the first threshold is greater than the second threshold.

[0026] Optionally, before the step of determining whether there is a matching first semantic feature for the second semantic feature corresponding to other text segments in the candidate patents to which the target text segment belongs, the method further includes:

[0027] Adjust the second threshold according to the number of candidate patents to which the target text segment belongs;

[0028] Among them, the number of candidate patents to which the target text segment belongs is positively correlated with the second threshold.

[0029] Optionally, the step of recommending the candidate patents whose sorting meets the preset ranking as the related patents of the basic patent includes:

[0030] Based on the document analysis function of the GPT (Generative Pre-Trained) model, analyze and summarize the basic patent and the candidate patents whose sorting meets the preset ranking respectively, and generate corresponding summary information;

[0031] Adjust the sorting results of the candidate patents within the preset ranking according to the similarity of the corresponding summary information between the basic patent and the candidate patents;

[0032] Generate a recommended list of candidate patents within the preset ranking based on the adjusted sorting results;

[0033] Perform the related patent recommendation operation of the basic patent based on the recommended list.

[0034] Optionally, after the step of querying the second semantic feature that matches the first semantic feature and marking the text segment associated with the queried second semantic feature as the target text segment, the method further includes:

[0035] If the number of target text segments among different candidate patents is the same, the higher the coherence of the target text segment in the text, the higher the corresponding ranking of the candidate patent.

[0036] To achieve the above object, the present application further provides a control device, including:

[0037] A screening module, configured to screen out a patent set that matches the key information according to the key information of the basic patent;

[0038] A slicing module, configured to perform document slicing on the basic patent to obtain a first text segment set; and perform document slicing on the candidate patents in the patent set to obtain a second text segment set;

[0039] An extraction module, configured to extract first semantic features corresponding to each passage in the first passage set and second semantic features corresponding to each passage in the second passage set based on a semantic understanding model;

[0040] A query module, configured to query second semantic features that match the first semantic features, and mark the passages associated with the queried second semantic features as target passages;

[0041] A sorting module, configured to sort candidate patents according to the number of target passages;

[0042] A recommendation module, configured to recommend candidate patents whose sorting meets a preset position as related patents of a basic patent.

[0043] To achieve the above object, the present application further provides a computer device, where the computer device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, the steps of the related patent recommendation method based on the semantic understanding model as described above are implemented.

[0044] To achieve the above object, the present application further provides a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the related patent recommendation method based on the semantic understanding model as described above are implemented.

[0045] The related patent recommendation method, control device, computer device, and computer-readable storage medium provided by the present application utilize a semantic understanding model to conduct in-depth semantic understanding and text analysis of patent texts, overcome the dependence on surface text information matching in traditional related patent search methods, and can more accurately capture the deep semantic information and relevance of patent texts. In this way, it can fully consider those patents that have similarities in concept and implementation manner although their literal expressions are different, so as to query more accurate and comprehensive related patents for recommendation, which helps users better complete patent comparison or related patent search tasks, thereby significantly improving the effect of related patent recommendation and user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a schematic diagram of the steps of the related patent recommendation method based on the semantic understanding model in an embodiment of the present application;

[0047] Figure 2 It is a schematic diagram of the control device in an embodiment of the present application;

[0048] Figure 3 It is an internal architecture schematic diagram of the computer device in an embodiment of the present application.

[0049] The realization, functional features and advantages of the present application will be further described in conjunction with embodiments with reference to the accompanying drawings. Detailed implementation manners

[0050] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.

[0051] In addition, if the description in the present application involves "first", "second", etc., it is only for descriptive purposes (such as for distinguishing the same or similar features), and should not be construed as indicating or implying its relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between the various embodiments may be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.

[0052] Referring to Figure 1 , in one embodiment, a related patent recommendation method based on a semantic understanding model includes:

[0053] Step S10: According to the key information of the basic patent, screen out a set of patents that match the key information;

[0054] Step S20: Perform document slicing on the basic patent to obtain a first text set; and perform document slicing on the candidate patents in the patent set to obtain a second text set;

[0055] Step S30: Based on the semantic understanding model, extract the first semantic features corresponding to each text in the first text set, and extract the second semantic features corresponding to each text in the second text set;

[0056] Step S40: Query the second semantic features that match the first semantic features, and mark the texts associated with the queried second semantic features as target texts;

[0057] Step S50: Sort the candidate patents according to the number of target texts;

[0058] Step S60: Recommend the candidate patents whose sorting meets the preset ranking as the related patents of the basic patent.

[0059] In this embodiment, the execution terminal of the embodiment can be a computer device, or other devices or apparatuses (such as a control device) that control the computer device.

[0060] As described in step S10, the basic patent refers to a specific patent that serves as a starting point or benchmark in a patent analysis or retrieval system. Generally, it is the specific patent that the user is researching or needs to find relevant patents for.

[0061] Optionally, the key information of the basic patent may refer to at least one of the IPC (International Patent Classification) classification number (including the main classification number and the secondary classification number), the keyword set (such as each technical term in the implementation scheme of the patent document), the applicant information, and the inventor information.

[0062] Optionally, use the advanced query function of the patent database to perform a search according to the key information conditions of the basic patent. The patent database generally allows precise queries according to multiple conditions, such as keywords, IPC classification numbers, applicants, etc.

[0063] Optionally, the patent scope can be gradually reduced in the order of the IPC classification number, the keyword set, the applicant, or the inventor information until a certain number (such as 50 - 200 pieces) of candidate patents are found and used as the patent set.

[0064] As described in step S20, document slicing is a method of splitting text into small pieces or paragraphs for further analysis or processing.

[0065] When performing document slicing on a patent document, first use natural language processing tools to parse the patent text, identify the positions or paragraphs where the background technology, the technical effects of the solution, and the embodiments of each solution are located, and then cut out one background technology, one technical effect of the solution, and one embodiment as one text paragraph from the patent document.

[0066] Optionally, refer to the above - mentioned document slicing method to perform document slicing on the basic patent, and use each text paragraph obtained after slicing as the first text paragraph set.

[0067] Optionally, refer to the above - mentioned document slicing method to perform document slicing on each candidate patent in the patent set, and use each text paragraph obtained after slicing as the second text paragraph set.

[0068] By splitting the document into smaller text paragraphs, the complexity of processing large - scale documents can be reduced. This makes it easier for algorithms to process and analyze a large amount of text information, thereby improving the analysis efficiency of patent documents.

[0069] As described in step S30, a semantic understanding model is pre-constructed and trained based on artificial intelligence technology, or an external semantic understanding model is called through an API (Application Programming Interface) interface.

[0070] Optionally, the selected semantic understanding model is used to perform semantic analysis on each passage in the first passage set, and the corresponding first semantic features of each passage are extracted.

[0071] Optionally, the selected semantic understanding model is used to perform semantic analysis on each passage in the second passage set, and the corresponding second semantic features of each passage are extracted.

[0072] It should be noted that the semantic features extracted by the semantic understanding model are represented in the form of high-dimensional vectors, which is more suitable for semantic similarity calculation and machine learning tasks.

[0073] Taking any passage as an example, the process of the semantic understanding model extracting high-dimensional vectors as semantic features is described as follows:

[0074] (1) Word segmentation: The semantic understanding model uses its built-in tokenizer to convert the passage into a token sequence. Each token represents a word or sub-word unit in the passage.

[0075] (2) Encoding: After word segmentation, the semantic understanding model converts these tokens into an input form that the model can understand. These input forms are usually sequences of vectors of a fixed length.

[0076] (3) Model processing: The vector sequence of the passage passes through the encoding layer of the model and is gradually transformed into a vector representation in the high-dimensional semantic space. These vector encodings capture the complex semantic features of the passage, including lexical context, syntactic structure, and semantic relationships.

[0077] (4) Output semantic features: Finally, the semantic understanding model outputs the semantic feature vector of the passage as the semantic feature. This vector is usually a high-dimensional numerical array, and each dimension represents the features of the passage in different semantic dimensions. This vectorized representation enables the semantic information of the passage to be compared and analyzed mathematically.

[0078] The extracted semantic features can be used for subsequent semantic similarity calculation or other natural language processing tasks. Querying for similar or identical text segments through semantic features can improve the accuracy and depth of text matching because semantic features not only consider lexical similarity but also understand the meaning and context behind the text. Therefore, they can capture the meaning and relevance of the text more accurately, rather than just the superficial textual similarity. That is, semantic features can identify the same concepts under different expressions, thus effectively overcoming the problems caused by synonyms and expression differences and improving the matching accuracy.

[0079] As described in step S40, based on the first semantic features corresponding to each text segment in the first text segment set, query for the second semantic features that match the first semantic features among the second semantic features corresponding to all text segments in the second document.

[0080] Optionally, select an appropriate similarity measurement method to compare the similarity between the two types of semantic features, such as cosine similarity, Euclidean distance, Manhattan distance, etc.

[0081] Optionally, for each text segment in the first text segment set, compare its first semantic feature with the second semantic features of all text segments in the second text segment set one by one. According to the selected similarity measurement method, calculate the similarity or distance between each pair of feature vectors.

[0082] Optionally, set a similarity threshold according to the actual requirements and task objectives. When the similarity of the semantic features between a text segment in the second text segment set and a text segment in the first text segment set exceeds this threshold, it can be considered that they are semantically matched.

[0083] Among them, if a second semantic feature that matches the first semantic feature is queried, then mark the text segment associated with the queried second semantic feature as the target text segment.

[0084] Using semantic features for querying can provide more refined and accurate matching results, capture deeper associations between texts, and thus better support the recommendation and ranking of relevant patents.

[0085] As described in step S50, the number of target text segments in each candidate patent can be counted. Then, according to the number of target text segments in each candidate patent, all candidate patents are sorted, and a corresponding sorting result is generated. Among them, the more the number of target text segments, the higher the ranking of the corresponding candidate patent.

[0086] As described in step S60, determine the number of relevant patents to be recommended or the ranking range. For example, the candidate patents ranked in the top few can be selected as relevant patents.

[0087] Optionally, according to the sorting result, select candidate patents ranked within a preset position. The larger number of target paragraphs of these patents that are similar to the basic patent indicates that they are most relevant to the basic patent in semantic matching.

[0088] Among them, the value of the preset position can be the top 5 to the top 20, preferably the top 10. For example, select candidate patents ranked within the top 10.

[0089] Optionally, generate a recommended list of candidate patents within the preset position to recommend the selected candidate patents as related patents of the basic patent.

[0090] Optionally, the recommended related patents can be used in various application scenarios, such as patent retrieval, technology competition analysis, etc.

[0091] In one embodiment, a semantic understanding model is used to conduct in-depth semantic understanding and text analysis of patent texts, overcoming the dependence on surface text information matching in traditional related patent search methods, being able to more accurately capture the deep semantic information and relevance of patent texts, so that patents with similar conceptions and implementation methods but different literal expressions can be fully considered, thereby querying more accurate and comprehensive related patents for recommendation, helping users better complete patent comparison or related patent search tasks, and thus significantly improving the effect of related patent recommendation and user satisfaction.

[0092] In one embodiment, based on the above embodiment, the semantic understanding model includes an input layer, an intermediate layer, and an output layer; wherein, the input layer is constructed based on word embedding technology, and the intermediate layer is constructed based on the BERT network;

[0093] The input layer is used to convert the text paragraph into a word embedding vector;

[0094] The intermediate layer is used to perform forward propagation on the word embedding vector, and perform a series of non-linear transformations and feature extractions through a multi-layer network structure;

[0095] The output layer is used to perform a feature aggregation operation on the features extracted by the intermediate layer to generate semantic features.

[0096] In this embodiment, a corresponding semantic understanding model is pre-constructed and trained based on artificial intelligence technology for use in the above-mentioned related patent recommendation method based on the semantic understanding model.

[0097] Optionally, the semantic understanding model includes an input layer, an intermediate layer, and an output layer.

[0098] Optionally, the main task of the input layer is to convert each word in the text passage into its corresponding word embedding vector. Word embedding techniques (such as Word2Vec, GloVe, etc.) are implemented through pre-trained models or training custom embedding layers, mapping each word to a real-valued vector of a fixed dimension. These vectors capture the semantic and syntactic relationships between words.

[0099] The form in which the input layer converts the text passage into word embedding vectors is: X i = E(W i );

[0100] where, W i represents the i-th word in the input text passage, E() represents the conversion function of the word embedding technique, and X i represents the word embedding vector of the i-th word.

[0101] Optionally, the middle layer uses the BERT network structure for more advanced semantic representation learning and feature extraction. The BERT network performs forward propagation through multiple layers of Transformer encoders, which are composed of multi-head attention mechanisms and feed-forward neural networks, and can effectively capture the complex dependencies and semantic features between words in the text passage. The BERT network has a multi-layer network structure, and the output of each layer contains rich semantic information. After multiple non-linear transformations and feature extractions, the semantic features of the text passage gradually become richer and more abstract.

[0102] Among them, the word embedding vectors obtained through the input layer serve as the input to the middle layer. The middle layer includes a multi-layer network structure. During the forward propagation process, each layer performs non-linear transformation and feature extraction on the input through a weight matrix and an activation function. The activation function can be ReLU (Rectified Linear Unit), Sigmoid, Tanh, etc., which is used to introduce non-linearity and increase the expressive power of the model. Each hidden layer in the middle layer contains multiple neurons, and each neuron performs weighted summation and activation function processing; the weight parameters of each layer are optimized through the backpropagation algorithm, enabling the model to gradually learn more expressive feature representations.

[0103] At the same time, each hidden layer is performing non-linear transformation and feature extraction: through the multiplication of the weight matrix and the action of the activation function, the input space is transformed into a higher-dimensional and more complex feature representation space; each layer learns different levels of abstract features: as the number of layers increases, the model gradually learns different levels of abstract features in the data to achieve the capture of complex patterns.

[0104] The form of feature extraction in the middle layer is: H^(j) = BERT(H^(j - 1)), H^(j - 1) = F (j-1) (X1, X2,..., X n));

[0105] Among them, H^(j) represents the hidden state of the j-th layer, BERT() represents the forward propagation function of the BERT network, and the function F() represents the forward propagation process of the BERT network, including multi-layer self-attention mechanisms and feed-forward neural networks.

[0106] Optionally, the output layer is responsible for aggregating the features extracted by the intermediate layer to generate the final semantic feature vector (i.e., the semantic feature vector). By processing the hidden state of the intermediate layer or the output of the pooling layer (such as average pooling or max pooling), the multi-layer features are integrated into a vector representation with a fixed dimension and rich semantic information (i.e., a high-dimensional vector), thereby obtaining the corresponding semantic features.

[0107] Among them, the choice of the aggregation method can be average pooling, max pooling, or an attention mechanism.

[0108] The semantic features generated by the output layer are represented as: T = G(H^(J));

[0109] J is the final layer number, G() represents the function of the feature aggregation operation, and T represents the final semantic feature representation.

[0110] In one embodiment, based on the above embodiment, the passages in the first passage set and the second passage set are divided into key passages and other passages except the key passages;

[0111] The steps of querying the second semantic feature that matches the first semantic feature and marking the passage associated with the queried second semantic feature as the target passage include:

[0112] According to the first semantic feature corresponding to the key passage in the first passage set, query the second semantic feature corresponding to the key passage in the second passage set that matches;

[0113] Mark the key passage associated with the queried second semantic feature as the target passage;

[0114] Query whether there is a matching first semantic feature for the second semantic feature corresponding to the other passages in the candidate patent to which the target passage belongs;

[0115] If so, mark the other passages as the target passages as well.

[0116] In this embodiment, the passages in the first passage set and the second passage set are divided into key passages and other passages except the key passages.

[0117] Among them, the passages corresponding to the background technology or the technical effects of the solution can be used as key passages; it should be understood that each basic patent and each candidate patent has at least one key passage.

[0118] Optionally, use the first semantic feature of the key passages in the first passage set as a query to match the second semantic features of all key passages in the second passage set. This can calculate the similarity or distance between each pair of feature vectors according to the selected similarity metric method.

[0119] If the second semantic feature of a key passage in the second passage set reaches a predetermined similarity threshold with the first semantic feature of the key passage in the first passage set, it is determined that the two match.

[0120] Optionally, mark the key passages associated with the second semantic feature as target passages. Based on this, candidate patents with target passages can be screened out, that is, candidate patents with the same or similar key passages as the basic patent are identified, and the screened candidate patents are used for subsequent recommendations.

[0121] Then, based on the candidate patents to which each target passage belongs (i.e., the screened candidate patents), further detect whether there are matching first semantic features (i.e., the first semantic features corresponding to other passages in the first passage set) for the second semantic features corresponding to other passages in these candidate patents.

[0122] Optionally, use a similarity metric method (such as cosine similarity or other similarity metrics) to calculate the similarity between the second semantic features corresponding to other passages in the screened candidate patents and the first semantic features corresponding to other passages in the first passage set.

[0123] Optionally, if the second semantic feature corresponding to other passages in the screened candidate patents reaches a predetermined similarity threshold with the first semantic feature corresponding to other passages in the first passage set, it is determined that the two match, and the corresponding other passages in the candidate patent are also marked as target passages.

[0124] In this way, while ensuring that the key passages of the candidate patents are highly relevant to the key passages of the basic patent, then marking the other passages in the candidate patents that are relevant to the basic patent as target passages for subsequent patent ranking recommendations can not only further narrow down the number of candidate patents to improve the analysis efficiency, but also more accurately match and recommend patents with strong relevance to the technical content of the basic patent, and can also improve.

[0125] In one embodiment, based on the above embodiment, the related patent recommendation method based on the semantic understanding model further includes:

[0126] If the similarity between the first semantic feature and the second semantic feature is greater than a preset threshold, it is determined that the first semantic feature and the second semantic feature match;

[0127] Among them, the preset threshold used when querying the key text segment is marked as the first threshold; the preset threshold used when querying other text segments is marked as the second threshold; the first threshold is greater than the second threshold.

[0128] In this embodiment, after using the semantic understanding model to extract the first semantic feature and the second semantic feature, the similarity between them can be calculated through similarity measurement (such as cosine similarity). If the similarity between the first semantic feature and the second semantic feature is greater than the preset threshold, it is determined that these two features match.

[0129] Optionally, in the process of querying the second semantic feature corresponding to the key text segment in the second text set that matches the first semantic feature corresponding to the key text segment in the first text set, when comparing the similarity between the first semantic feature and the second semantic feature corresponding to the key text segments in the two text sets, the preset threshold used is the first threshold, that is, if it is detected that the similarity between the first semantic feature and the second semantic feature corresponding to the key text segments in the two text sets is greater than the first threshold, it is determined that they match.

[0130] Optionally, in the process of querying whether there is a matching first semantic feature for the second semantic feature corresponding to other text segments in the candidate patents to which the target text belongs, when comparing the similarity between the first semantic feature and the second semantic feature corresponding to other text segments in the two text sets, the preset threshold used is the second threshold, that is, if it is detected that the similarity between the first semantic feature and the second semantic feature corresponding to other text segments in the two text sets is greater than the second threshold, it is determined that they match.

[0131] Among them, the first threshold is greater than the second threshold. This setting can reflect that when determining the matching of key text segments, a higher similarity threshold needs to be adopted to ensure accuracy and relevance; while when verifying other text segments, more relaxed matching conditions can be accepted to expand the recommendation scope and increase diversity.

[0132] For example, the value range of the first threshold can be 75% - 90%, and it can be optionally 80%; the value range of the second threshold can be 55% - 70%.

[0133] It should be noted that because the key text segments are often the background technology or the technical effects of the solution of the patent, which are more likely to reflect the concept of the patent and usually cover the core technology or innovation points of the patent. Therefore, when matching these parts, setting a higher similarity threshold (that is, the first threshold) can accurately match candidate patents with strong relevance to the basic patent technical content.

[0134] Other paragraphs are generally embodiments of the patent solution. Regarding these information relative to the key technical points, the matching requirements can be more relaxed. By using a relatively loose similarity threshold (i.e., the second threshold), the recommendation scope can be expanded and the diversity can be increased, enabling the recommendation system to more comprehensively cover relevant patent content.

[0135] In one embodiment, on the basis of the above embodiment, before the step of determining whether there is a matching first semantic feature for the second semantic feature corresponding to other paragraphs in the candidate patents to which the query target paragraph belongs, it further includes:

[0136] Adjust the second threshold according to the number of candidate patents to which the target paragraph belongs;

[0137] Wherein, the number of candidate patents to which the target paragraph belongs is positively correlated with the second threshold.

[0138] In this embodiment, after marking the key paragraphs that meet the requirements in the second paragraph set as target paragraphs, the number of target paragraphs in the second paragraph set at this time can roughly reflect the number of candidate patents selected. Therefore, by counting the number of candidate patents with target paragraphs currently, the number of candidate patents selected can be obtained.

[0139] At this time, the second threshold can be adjusted according to the number of candidate patents to which the target paragraph belongs (i.e., the number of candidate patents selected).

[0140] Among them, the number of candidate patents to which the target paragraph belongs is positively correlated with the second threshold. That is, when the number of candidate patents to which the target paragraph belongs is large, it indicates that the number of candidate patents selected is sufficient. At this time, the second threshold can be appropriately increased, thereby increasing the difficulty of marking other paragraphs as target paragraphs subsequently, and then improving the accuracy and precision of the recommendation; if the number of candidate patents selected is small, the second threshold can be appropriately decreased, thereby reducing the difficulty of marking other paragraphs as target paragraphs subsequently, to ensure that more other paragraphs can be marked as target paragraphs, increasing the comprehensiveness and diversity of the recommendation.

[0141] By dynamically adjusting the second threshold according to the number of candidate patents, the system can optimize the recommendation effect in different scenarios. This strategy not only maintains the accuracy of the recommendation when dealing with a large number of candidate patents, but also ensures that the comprehensiveness of the recommendation can be increased when the number of candidate patents is small, thereby enhancing the user's retrieval experience and satisfaction.

[0142] The specific way to dynamically adjust the second threshold according to the number of candidate patents can be to set a corresponding proportional function to calculate the second threshold according to the number of candidate patents. For example, a linear function or other mathematical models can be used to ensure that the second threshold maintains a positive correlation with the change in the number of candidate patents.

[0143] Alternatively, set numerical ranges for multiple candidate patent quantities, with each range associated with a corresponding second threshold, and obtain the corresponding second threshold based on the range to which the candidate patent quantity belongs. For example, several threshold levels can be defined, such as low, medium, and high, and the corresponding threshold can be selected according to the specific range in which the candidate patent quantity is located.

[0144] In one embodiment, based on the above embodiment, the step of recommending the candidate patents whose sorting meets the preset ranking as related patents of the basic patent includes:

[0145] Based on the document analysis function of the GPT model, respectively analyze and summarize the basic patent and the candidate patents whose sorting meets the preset ranking, and generate corresponding summary information;

[0146] According to the similarity of the corresponding summary information between the basic patent and the candidate patents, adjust the sorting results of the candidate patents within the preset ranking;

[0147] Based on the adjusted sorting results, generate a recommendation list of the candidate patents within the preset ranking;

[0148] Based on the recommendation list, perform the operation of recommending related patents of the basic patent.

[0149] In this embodiment, after sorting the candidate patents according to the number of target paragraphs in the candidate patents and obtaining the candidate patents whose sorting meets the preset ranking, use the document analysis function of the GPT model to respectively analyze and summarize the basic patent and the candidate patents whose sorting meets the preset ranking (such as analyzing and summarizing the full text of the patent specification), and generate corresponding summary information.

[0150] Among them, it can be to call the externally provided GPT model based on the corresponding API interface.

[0151] The reason why the document analysis function of the GPT model can generate the corresponding summary information of the patent is because of its powerful language understanding and generation capabilities, context awareness capabilities, generation diversity, and flexibility in adapting to domain-specific tasks. These characteristics make the GPT model an effective tool for processing and analyzing patent documents, and can provide accurate and comprehensive information summaries for patents.

[0152] After obtaining the summary information of the basic patent and the summary information of the candidate patents ranked within the preset ranking, calculate the similarity of the summary information between the basic patent and the candidate patents respectively. Then, according to the similarity of the corresponding summary information between the basic patent and the candidate patents, adjust the sorting results of the candidate patents within the preset ranking.

[0153] Optionally, the candidate patents within the preset ranking can be re - sorted according to the numerical value of the similarity of the summary information, so as to adjust the sorting result of the candidate patents within the preset ranking. Among them, the higher the similarity corresponding to the summary information, the higher the ranking. On the contrary, if the similarity is lower, this candidate patent will be ranked lower in the sorting.

[0154] This re - sorting based on the similarity of the summary information can effectively adjust the recommended output to ensure that the candidate patents seen by the user are more relevant to the basic patent. This method utilizes the summary information generated by the GPT model, and through quantitative similarity calculation, makes the sorting result more in line with the user's needs and expectations.

[0155] Alternatively, according to the original ranking of the candidate patents within the preset ranking and the similarity corresponding to their summary information, a weighted sum score is calculated; after obtaining the score results of each candidate patent, re - sort based on the score results, and the higher the score, the higher the ranking.

[0156] An example of the score calculation formula is as follows:

[0157] P = a * c * e + b * d * f;

[0158] Among them, P is the score value, a is the original ranking of the candidate patent, b is the similarity between the candidate patent and the summary information of the basic patent, c is the score conversion factor corresponding to the original ranking, d is the score conversion factor corresponding to the similarity, e is the first weight corresponding to the original ranking, and f is the second weight corresponding to the similarity; where e > f.

[0159] The score conversion factor refers to the coefficient or parameter used in the score calculation formula to convert the original data (such as ranking or similarity) into the final score value. These factors can be set according to specific requirements and data distribution. For example, if it is desired that the ranking has a greater impact on the score, a larger c value can be set; if the similarity has a more important impact on the score, a larger d value can be set.

[0160] In addition, the first weight is set to be greater than the second weight, so the score calculation pays more attention to the original ranking of the candidate patent. This means that the position of the candidate patent in the original ranking contributes more to the final score value. The candidate patent with a higher original ranking will have a higher starting score, thus occupying a more favorable position in the sorting.

[0161] Optionally, after adjusting the sorting of the candidate patents within the preset ranking, based on the adjusted sorting result, a recommended list of the candidate patents within the preset ranking is generated according to the ranking order, so as to display the most relevant and highest - quality patent recommendations to the user.

[0162] Optionally, the form of the recommended list can be a simple list form, such as listing the title or name of the patent, which can be accompanied by a brief description or key information; or, a card display form, that is, each patent is displayed as a card, including information such as the title and abstract of the patent, so that users can quickly browse and compare; or, a detailed view form, providing a detailed information page of the patent, including the technical field, abstract, key features, and similarity analysis with the basic patent, etc., to help users deeply understand the background and basis of the patent recommendation.

[0163] Optionally, when performing the relevant patent recommendation operation for the basic patent, the relevant patent recommendation list can be displayed to the user in an appropriate form, such as the simple list form, card display, detailed view, or chart graphical display mentioned before. Through such an operation process, a patent recommendation list related to the basic patent can be generated, helping users quickly find the patent content they are interested in among a large number of candidate patents.

[0164] In one embodiment, on the basis of the above embodiment, before the step of screening out a patent set that matches the key information according to the key information of the basic patent, it further includes:

[0165] After analyzing the patent requirements input by the user based on the GPT model, corresponding key information is extracted from the basic patent based on the analysis result.

[0166] In this embodiment, the system can receive the natural language input provided by the user, which can be a descriptive text, and can include the user's patent requirements, technical problems, or specific requirements, such as describing technical problems, requirements, or goals. For example, the user may describe the need to find a patent in a specific field or the need to solve a certain technical challenge.

[0167] Use the GPT model to perform text understanding and analysis on the user input. The GPT model can understand the key information in the context through its pre-trained language understanding ability.

[0168] During the analysis process, the model may identify content related to the IPC classification number, keyword set, applicant information, or inventor information. For example, the GPT model can extract key technical features, application scenarios, or specific technical requirements from the user input.

[0169] Based on the analysis result, the GPT model generates key information representing the user's needs. These information are usually able to clearly define the specific attributes or features of the basic patent, so that the subsequent patent screening process can accurately find the matching patent set.

[0170] Subsequently, using the key information extracted from the user input, step S10 is executed. This key information will be used as the basis for screening the patent set that matches the basic patent, and thus as the starting point or benchmark for the subsequent recommendation process.

[0171] When analyzing the user input, the GPT model can not only understand its meaning but also effectively extract the key information for patent screening and recommendation. This method utilizes natural language processing technology to automate and precision the patent search process, improving the efficiency and accuracy of the search.

[0172] Compared with extracting key information according to artificially preset rules, analyzing the patent requirements of the user input based on the GPT model and extracting key information has the following advantages:

[0173] (1) Flexibility and adaptability: The GPT model can flexibly adjust its analysis and extraction process according to different user inputs, adapting to various descriptions of patent requirements. This means that there is no need to define fixed key information extraction rules in advance, and it can handle more extensive and complex input situations.

[0174] (2) Context understanding ability: Through its deep learning architecture and the ability of large-scale pre-trained language models, the GPT model can better understand the context and meaning of the input text. Compared with static rule sets, it can more comprehensively consider the language structure and semantic information in the input text.

[0175] (3) Accuracy: Since the GPT model has been pre-trained on a large amount of data and has powerful language understanding ability, it can usually capture the core content of the user's needs more accurately when extracting key information. This helps to improve the precision and efficiency of the subsequent patent screening and recommendation process.

[0176] Generally speaking, the key information extraction method based on the GPT model is more flexible, intelligent and efficient than the traditional rule-based method, and can bring greater advantages and value to patent recommendation.

[0177] In one embodiment, on the basis of the above embodiment, after the step of querying the second semantic feature that matches the first semantic feature and marking the passage associated with the queried second semantic feature as the target passage, the following is further included:

[0178] If the number of target passages between different candidate patents is the same, the higher the coherence of the target passage in the text, the higher the ranking of the corresponding candidate patent.

[0179] In this embodiment, if the number of target paragraphs among different candidate patents is the same, then the coherence of these target paragraphs in the text is evaluated. Coherence refers to the consistency and fluency of the paragraphs in terms of content and logic, and can reflect the consistency and integrity of the patent in terms of technical implementation or innovation.

[0180] Among them, if there are no other non-target paragraphs separating the target paragraphs in the candidate patent or the number of intermediate non-target paragraphs is small, it indicates that the coherence of the target paragraphs in the text is high. Therefore, if paragraph interval analysis is needed to evaluate the coherence of the target paragraphs in the candidate patent, it can be carried out in the following way successively:

[0181] (1) Paragraph interval check: Check each target paragraph of each candidate patent one by one to confirm whether there are non-target paragraphs separating the target paragraphs or whether the number of intermediate non-target paragraphs is small.

[0182] (2) Mark non-target paragraphs: For the existing non-target paragraphs, their content can be marked or labeled to more clearly identify the relative relationship between the target paragraphs.

[0183] (3) Analyze continuity: Analyze the paragraph interval to check the degree of continuity between the target paragraphs in the candidate patent. The higher the continuity, the stronger the connection between the target paragraphs in the text in terms of content and logic, and the relatively higher the coherence.

[0184] (4) Evaluate the coherence quality: According to the results of the paragraph interval analysis, comprehensively consider the continuity situation between the target paragraphs, and evaluate the coherence quality of the target paragraphs in the candidate patent in the text.

[0185] Through the above process, the coherence degree of the target paragraphs in the candidate patent in the text can be relatively evaluated.

[0186] Optionally, according to the coherence degree, the candidate patents with the same number of target paragraphs are sorted. Among them, the candidate patent with a higher coherence degree will be ranked ahead to enhance the quality and relevance of the recommended list.

[0187] In this way, the text coherence in addition to semantic matching is emphasized, which can ensure that the recommended candidate patents are more consistent with the basic patent in terms of technology and content. By comprehensively considering semantic feature matching and text coherence, the system can provide more accurate and targeted relevant patent recommendations, meeting the user's need for in-depth understanding of technology and innovative content.

[0188] In addition, referring to Figure 2 , this embodiment of the present application also provides a control device Z10, including:

[0189] The screening module Z11 is used to screen out a set of patents that match the key information according to the key information of the basic patent;

[0190] The slicing module Z12 is used to perform document slicing on the basic patent to obtain a first text segment set; and, perform document slicing on the candidate patents in the patent set to obtain a second text segment set;

[0191] The extraction module Z13 is used to extract the first semantic features corresponding to each text segment in the first text segment set and the second semantic features corresponding to each text segment in the second text segment set based on the semantic understanding model;

[0192] The query module Z14 is used to query the second semantic features that match the first semantic features and mark the text segments associated with the queried second semantic features as target text segments;

[0193] The sorting module Z15 is used to sort the candidate patents according to the number of target text segments;

[0194] The recommendation module Z16 is used to recommend the candidate patents whose sorting meets the preset positions as the related patents of the basic patent.

[0195] Optionally, the control device Z10 can be a virtual control device (such as a virtual machine) or a physical device (such as a physical device other than a computer device that can execute the corresponding method).

[0196] In addition, an embodiment of the present application also provides a computer device, and the internal architecture of the computer device can be as Figure 3 shown, including a processor, a memory, a communication interface, and an input interface connected through a system bus. Among them, the processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database is used to store data called by the computer program. The communication interface is used to communicate with an external terminal for data. The input interface is used to receive signals input by an external device. When the computer program is executed by the processor, it is used to implement a related patent recommendation method based on a semantic understanding model as described in the above embodiments.

[0197] Those skilled in the art can understand that Figure 3 the structure shown in

[0198] In addition, the present application also provides a computer-readable storage medium, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the related patent recommendation method based on the semantic understanding model as described in the above embodiments. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0199] In summary, for the related patent recommendation method, control device, computer device, and computer-readable storage medium provided in the embodiments of the present application, by using the semantic understanding model, in-depth semantic understanding and text analysis of patent texts are carried out, overcoming the dependence on surface text information matching in traditional related patent search methods, and being able to more accurately capture the deep semantic information and relevance of patent texts. In this way, patents that are similar in concept and implementation manner but have different literal expressions can be fully considered, so as to query more accurate and comprehensive related patents for recommendation, which helps users better complete patent comparison or related patent search tasks, thus significantly improving the effect of related patent recommendation and user satisfaction.

[0200] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in the present application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0201] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, device, article or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, device, article or method including such an element.

[0202] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A method for recommending related patents based on a semantic understanding model, characterized in that: include: According to the key information of the basic patent, screen out the patent collection that matches the key information; Performing document slicing on the basic patent to obtain a first segment set; and performing document slicing on the candidate patents in the patent set to obtain a second segment set; the segments in the first segment set and the second segment set are divided into key segments and other segments except the key segments; Based on the semantic understanding model, extracting a first semantic feature corresponding to each paragraph in the first paragraph set, and extracting a second semantic feature corresponding to each paragraph in the second paragraph set; According to the first semantic feature corresponding to the key paragraph in the first paragraph set, searching for the second semantic feature corresponding to the key paragraph in the second paragraph set that matches; Marking the key text segment associated with the queried second semantic feature as the target text segment; Query the second semantic features corresponding to other text segments in the candidate patent to which the target text segment belongs to see whether there is a matching first semantic feature; if so, mark the other text segments as the target text segment; wherein, if the similarity between the first semantic feature and the second semantic feature is greater than a preset threshold, the first semantic feature is determined to match the second semantic feature; mark the preset threshold used when querying the key text segment as the first threshold; mark the preset threshold used when querying other text segments as the second threshold; the first threshold is greater than the second threshold; adjust the second threshold according to the number of candidate patents to which the target text segment belongs, and the number of candidate patents to which the target text segment belongs is positively correlated with the second threshold; Sort candidate patents according to the number of target text segments; Based on the document analysis function of the GPT model, the basic patents and candidate patents that meet the preset rankings are analyzed and summarized, and corresponding summary information is generated; According to the similarity of the corresponding summary information between the basic patent and the candidate patent, the ranking results of the candidate patents within the preset ranking are adjusted; wherein, according to the original ranking of the candidate patents within the preset ranking and the similarity corresponding to the summary information of the candidate patents, a weighted sum score is performed to obtain the score result of each candidate patent; the ranking is re-arranged based on the score result, and the higher the score, the higher the ranking; the first weight corresponding to the original ranking is greater than the second weight corresponding to the similarity corresponding to the summary information; Based on the adjusted ranking results, a recommended list of candidate patents within a preset ranking is generated; Based on the recommendation list, perform patent recommendation operations related to the basic patent.

2. The method for recommending related patents based on a semantic understanding model according to claim 1, characterized in that: The semantic understanding model includes an input layer, an intermediate layer and an output layer; wherein the input layer is constructed based on word embedding technology, and the intermediate layer is constructed based on a BERT network; The input layer is used to convert the text segment into a word embedding vector; The intermediate layer is used to forward propagate the word embedding vector, and perform a series of nonlinear transformations and feature extraction through a multi-layer network structure; The output layer is used to perform feature aggregation operation on the features extracted by the intermediate layer to generate semantic features.

3. The method for recommending related patents based on a semantic understanding model according to claim 1, characterized in that: After the step of searching for a second semantic feature matching the first semantic feature and marking the text segment associated with the searched second semantic feature as the target text segment, the method further includes: If the number of target paragraphs among different candidate patents is the same, the higher the coherence of the target paragraphs in the text, the higher the ranking of the corresponding candidate patent.

4. A control device, characterized in that: include: The screening module is used to screen out patent sets that match the key information based on the key information of the basic patent; A slicing module is used to perform document slicing on the basic patent to obtain a first segment set; and to perform document slicing on the candidate patents in the patent set to obtain a second segment set; the segments in the first segment set and the second segment set are divided into key segments and other segments except the key segments; An extraction module, configured to extract, based on the semantic understanding model, a first semantic feature corresponding to each paragraph in the first paragraph set, and a second semantic feature corresponding to each paragraph in the second paragraph set; A query module, configured to query a second semantic feature corresponding to a key paragraph in a matching second paragraph set based on a first semantic feature corresponding to a key paragraph in a first paragraph set; Mark the key text associated with the queried second semantic feature as the target text; query the second semantic features corresponding to other texts in the candidate patent to which the target text belongs to see whether there is a matching first semantic feature; if so, mark the other texts as the target text; wherein, if the similarity between the first semantic feature and the second semantic feature is greater than a preset threshold, the first semantic feature is determined to match the second semantic feature; mark the preset threshold used when querying the key text as the first threshold; mark the preset threshold used when querying other texts as the second threshold; the first threshold is greater than the second threshold; adjust the second threshold according to the number of candidate patents to which the target text belongs, and the number of candidate patents to which the target text belongs is positively correlated with the second threshold; A sorting module is used to sort the candidate patents according to the number of target text segments; The recommendation module is used for analyzing and summarizing the basic patents and the candidate patents whose rankings meet the preset rankings based on the document analysis function of the GPT model, and generating corresponding summary information; adjusting the ranking results of the candidate patents within the preset rankings according to the similarity of the corresponding summary information between the basic patents and the candidate patents; wherein, weighted summing and scoring are performed according to the original rankings of the candidate patents within the preset rankings and the similarity corresponding to the summary information of the candidate patents to obtain the scoring results of each candidate patent; re-ranking is performed based on the scoring results, and the higher the score, the higher the ranking; the first weight corresponding to the original ranking is greater than the second weight corresponding to the similarity corresponding to the summary information; based on the adjusted ranking results, a recommendation list of candidate patents within the preset rankings is generated; based on the recommendation list, the relevant patent recommendation operation of the basic patent is performed.

5. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the related patent recommendation method based on the semantic understanding model as described in any one of claims 1 to 3 are implemented.

6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the related patent recommendation method based on a semantic understanding model as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Patent recommendation method and device based on semantic link heterogeneous information network embedding

    CN110175224A

  • Text retrieval method and device, computer equipment and storage medium

    CN111444320A