Method for constructing patent depth indexing and special database based on semantic retrieval
By constructing a semantic encoding model and an indexing term semantic matching model, deep semantic vectors are generated, which solves the problem of low accuracy in patent retrieval in existing technologies and realizes the professionalism of patent indexing and the scientific nature of clustering topics.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI SANYI TECHNOLOGY CO LTD
- Filing Date
- 2026-05-08
- Publication Date
- 2026-07-31
AI Technical Summary
Existing patent indexing and database construction methods are difficult to perform deep semantic feature extraction and cluster analysis, resulting in low accuracy of patent retrieval and difficulty in uncovering core innovation points and deep technological connections.
By constructing a semantic encoding model to generate shallow and deep semantic vectors, and combining it with an indexing word semantic matching model, similarity and membership functions are calculated, and clustering and iterative optimization are performed to generate a thematic database.
It enhances the professionalism and accuracy of patent indexing, ensures the scientific nature of clustering topic division, avoids technical overlap and confusion, and achieves accurate semantic retrieval.
Smart Images

Figure CN122489741A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of patent retrieval technology, specifically to a method for deep indexing of patents and construction of thematic databases based on semantic retrieval. Background Technology
[0002] The method for deep patent indexing and thematic database construction based on semantic retrieval is a comprehensive technical system integrating semantic understanding, artificial intelligence, and patent information management technologies. The deep patent indexing method mainly consists of a deep patent indexing component and a thematic database construction component. It transforms patent text into high-dimensional semantic vectors through semantic embedding. Simultaneously, it performs hierarchical indexing of the patent's knowledge objects and knowledge elements in accordance with Chinese patent deep content indexing standards. The indexed patent data is then classified, integrated, and structured according to the needs of specific thematic domains, forming a structured thematic database.
[0003] Existing patent indexing and database construction methods struggle with deep semantic feature extraction from shallow semantic vectors and indexing vector generation. They lack clustering analysis and thematic segmentation based on new indexing vectors. Relying solely on shallow semantic vectors for patent retrieval and matching makes it difficult to uncover the core innovations and deep technological connections behind patent documents, thus limiting the professionalism, accuracy, and deep analytical capabilities of patent indexing. Most methods rely on simple patent screening based on shallow semantic similarity, lacking clustering thematic construction and iterative optimization functions based on deep semantic indexing. This makes it difficult to support the needs of accurate semantic retrieval, technical thematic segmentation, and professional patent database construction in the patent field. Summary of the Invention
[0004] In view of the shortcomings of the prior art, the purpose of this invention is to provide a method for deep indexing of patents and construction of thematic databases based on semantic retrieval, so as to solve the problems of difficulty in mining the core innovation points and deep technical connections behind patent documents, and difficulty in supporting accurate semantic retrieval in the patent field.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: The present invention provides a method for deep indexing of patents and construction of thematic databases based on semantic retrieval, the method comprising: Step S1: Obtain patent documents and define their structure; concatenate the defined text content to form complete text content; perform word segmentation on the complete text content to generate a preliminary vocabulary set; filter and remove duplicates from all technical terms in the preliminary vocabulary set to generate an effective vocabulary list. Step S2: Construct a semantic matching model for indexing terms, build a set of indexing terms and divide each indexing term in the set into different levels; concatenate and reconstruct the indexing weights to generate indexing vectors; Step S3: The user retrieves patents and generates query vectors; based on the query vectors and indexing vectors, a similarity function is constructed to calculate their similarity; a filtering function is constructed to filter out candidate patent documents and generate corresponding new indexing vectors. Step S4: Construct the average semantic distance function and introduce patent time weights to calculate the average semantic distance of the new indexing vectors; construct the membership function; construct the clustering objective function, summarize and iteratively generate a set of patent documents for clustering topics, and generate a topic set; Step S5: Construct the topical semantic mapping function, semantic index structure, and indexing vector index structure; construct the topical database.
[0006] Furthermore, the complete text content in step S1 is as follows: in, For the first The complete text of the patent document; For the first The title of the patent document; For the first Abstracts of patent documents; For the first The text of the claims in a patent document; For the first The specification text of the patent document; For patent literature indexing; For text concatenation; The preliminary vocabulary set is shown below: in, This is a preliminary vocabulary set; For the first A technical term; Total number of technical terms; The list of valid vocabulary is as follows: in, For a valid vocabulary list; Total number of valid words; For the first One effective vocabulary; For vocabulary indexing.
[0007] Furthermore, in step S1, the word frequency of each word is calculated based on the effective vocabulary, and a statistical weight is added to the word frequency in the claim text. The word frequencies are as follows: in, Frequency weighting; For the effective vocabulary list, number 1 The word, in the 1st... Total number of times it appears in the patent document; For the first The sum of the occurrences of all words in the effective vocabulary list in a patent document; The number of times the word appears in the title; The number of times a word appears in the summary; The number of times a word appears in the text of the claims. The term frequency weighting coefficients in the claim text; The number of times a word appears in the instruction manual text; A semantic encoding model is constructed to generate shallow semantic vectors. The semantic encoding model is shown below: in, This is a shallow semantic vector; For shallow semantic activation functions, It is a mask diagonal matrix; This is a mapping matrix for shallow semantics; For the first Text feature vectors in a patent document; This is a bias term for shallow semantics.
[0008] Furthermore, the different levels in step S2 are as follows: in, For the first One indexing term; A hierarchical set; As the core innovation feature level; For general technical features level; This is a feature hierarchy for an example implementation; A matching score function is constructed and combined with the weight coefficients of different levels to calculate the matching score between the patent document and the indexing term. The matching score function is as follows: in, The matching score between patent documents and index terms; This is the transpose of the semantic vector of the indexing term; , and These are the weighting coefficients for different levels; It is a deep semantic feature.
[0009] Furthermore, in step S2, a specificity index of the indexing term's level is introduced, and the indexing weight is calculated using the Softmax function, as shown below: in, For indexing weights; The matching score between patent documents and index terms; It is a natural exponential function; The sum of the exponents of all matching scores; This is a specificity index.
[0010] Furthermore, the filtering function in step S3 is as follows: in, A collection of candidate patent documents; The similarity threshold; For the first Patent documents; To determine the similarity between the query vector and the index vector; These are the shortlisted candidate patent documents. The total number of candidate patent documents.
[0011] Furthermore, in step S4, the distance to each cluster topic center is calculated based on the new indexing vector, using the following function: in, Clustering of selected candidate patent documents into specific topics semantic distance; As a weighting factor for patent popularity; For the first The center of a clustering topic, For the new indexing vector; The membership function is as follows: in, For candidate patent documents, the first Membership degree of each cluster topic; For candidate patent documents to all The sum of the reciprocals of the semantic distances of the cluster thematic centers; The candidate patent documents are divided into different cluster topics, and the patent document division rules are as follows: in, For the first A collection of patent documents for each cluster topic; This is a preset membership threshold. To select candidate patent documents; The total number of clustered topics; This is a clustering-based thematic index.
[0012] Furthermore, the clustering objective function in step S4 is as follows: in, The clustering objective function; For the first The center of a clustering topic, For the new indexing vector; These are the shortlisted candidate patent documents; The preset number of clustering topics; For the first A collection of patent documents for each cluster topic; For clustering topic indexes; A judgment function is constructed and a judgment threshold is preset. The judgment function is as follows: in, For the first The objective function value of each iteration; For the first The objective function value of each iteration; This represents the number of iterations. This is the preset judgment threshold.
[0013] Furthermore, the semantic index structure in step S5 is as follows: in, For clustering topics The corresponding patent index set; To select candidate patent documents; For clustering topic indexes; To select the target topics to which candidate patent documents belong; The indexing vector structure is shown below: in, For clustering topics The corresponding indexing vector index; For the new indexing vector; To select candidate patent documents; The thematic database is shown below: in, For specialized databases; A collection of patent documents; The updated cluster centers; For a collection of patent indexes; This is the index for the indexing vector.
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: The method of this invention improves the professionalism and accuracy of patent indexing through deep semantic feature extraction and indexing vector generation, making the indexing results more closely match the core technology of the patent. At the same time, the semantic coding model enables the generated deep semantic features to more accurately reflect the semantic information of the patent text. Finally, it can accurately select core indexing words with high semantic matching degree with the patent, making the indexing words fit the patent features of the corresponding technical field, and improving the relevance and adaptability of the indexing. The method of this invention generates a set of topics through semantic clustering, which makes the division of cluster topics closely match the actual characteristics of the technical field to which the patent belongs. This avoids the confusion caused by too many or too few clusters and ensures the scientific nature of the topic division. At the same time, it can clarify the semantic boundaries of different technical topics, effectively avoiding technical overlap and mixing between patent topics. Finally, by iteratively updating the cluster centers and measuring the clustering effect, it ensures that the division of patent cluster topics achieves the optimal effect, making the semantic similarity of patents within the same topic higher and the semantic differences between different topics more obvious. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram of the overall workflow of the method in this invention; Figure 2 This is a flowchart illustrating the specific workflow of step S1 in the method of this invention; Figure 3 This is a flowchart illustrating the specific workflow of step S2 in the method of this invention; Figure 4 This is a flowchart illustrating the specific workflow of step S3 in the method of this invention; Figure 5 This is a flowchart illustrating the specific workflow of step S4 in the method of this invention; Figure 6 This is a flowchart illustrating the specific workflow of step S5 in the method of this invention. Detailed Implementation
[0017] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to embodiments and accompanying drawings. The content mentioned in the embodiments is not intended to limit the present invention.
[0018] Reference Figure 1 As shown, the method for deep indexing of patents and construction of thematic databases based on semantic retrieval of the present invention comprises the following steps: Step S1: Obtain patent documents and define their structure; concatenate the defined text content to form complete text content; perform word segmentation on the complete text content to generate a preliminary vocabulary set; filter and remove duplicates from all technical terms in the preliminary vocabulary set to generate an effective vocabulary list; construct a semantic encoding model to generate shallow semantic vectors. First, obtain the patent document collection, which is shown below: in, A collection of patent documents; The total number of patent documents; For example, in the field of new energy vehicle power battery technology, five patent documents in related fields were selected. The patent literature covers material improvements, fast charging technology, and thermal management systems for ternary lithium batteries and lithium iron phosphate batteries. The text content of each patent document is structurally defined, and the structure of each patent document is as follows: in, For the first Patent documents; For the first The title of the patent document; For the first Abstracts of patent documents; For the first The text of the claims in a patent document; For the first The specification text of the patent document; For patent literature indexing; Secondly, the defined text content is concatenated to form the complete text content: in, For the first The complete text of the patent document; For text concatenation; Based on the complete text content, word segmentation is performed to generate a preliminary vocabulary set: in, This is a preliminary vocabulary set; For the first A technical term; Total number of technical terms; For example, a patent titled "A Fast Charging Heat Dissipation Structure for a Ternary Lithium Battery" yields a preliminary vocabulary set after word segmentation: in, ; After filtering and deduplicating all technical terms in the initial vocabulary set, a valid vocabulary list is generated, as shown below: in, For a valid vocabulary list; Total number of valid words; For the first One effective vocabulary; For a vocabulary index, for example, for the total number of valid words. 10 valid words were randomly selected; After deduplication and word segmentation of the titles, abstracts, claims, and descriptions of five patent documents, an effective vocabulary was constructed, as shown below: Based on an effective vocabulary, the frequency of each word is calculated. Targeting the high density of technical terms in patent texts, core technical terms in individual patents are selected for filtering stop words in the patent field after word segmentation. This improves the efficiency and accuracy of feature extraction and increases the statistical weight of word frequencies in the claims. The word frequencies are shown below: in, Frequency weighting; For the effective vocabulary list, number 1 The word, in the 1st... Total number of times it appears in the patent document; For the first The sum of the occurrences of all words in the effective vocabulary list in a patent document; The number of times the word appears in the title; The number of times a word appears in the summary; The number of times a word appears in the text of the claims. The term frequency weighting coefficients in the claims text are adjusted by experts in the technical research and development field based on the core requirements of the patent indexing, and the most accurate values are selected as the final term frequency weighting coefficients. For example, when When a technical term appears once in the text of the claims, it is counted as two occurrences. The number of times a word appears in the instruction manual text; Taking ternary lithium battery fast charging technology as an example, the data is substituted into the formula to calculate the word frequency of each word. The word frequencies are shown in the table below: Based on an effective vocabulary list, the number of patent documents containing the term is counted. This aligns with the core needs of patent retrieval, allowing for the selection of unique technical terms from the patent collection. This avoids interference from generic technical terms on patent characteristics. The statistical function is shown below: in, Inverse patent document frequency; For including the first Number of patent documents per word; The total number of patent documents; The frequency of statistical terms in patent literature is calculated by substituting the data into a formula. The inverse patent literature frequency is shown in the table below: Based on term frequency weights and inverse patent document frequency, a weight calculation function is constructed, as shown below: in, For the final weight; Finally, the final weights are combined to form a text feature matrix, as shown below: in, For the first Text feature vectors in a patent document; Based on text feature vectors, a semantic encoding model is constructed to generate shallow semantic vectors. The semantic encoding model is shown below: in, This is a shallow semantic vector; is the activation function for shallow semantics; This is a mapping matrix for shallow semantics; This is a bias term for shallow semantics; A lexical masking factor is introduced into the activation function, as shown below: in, It is a mask diagonal matrix; For the first The mask factor of each word, when These terms belong to core technology terms. ,otherwise, ; Constructor of a diagonal matrix; Substituting the data into the formula yields the shallow semantic vector: Step S2: Construct a semantic matching model for indexing terms, build a set of indexing terms and divide each indexing term in the set into different levels; concatenate and reconstruct the indexing weights to generate indexing vectors; First, deep semantic features are extracted from the shallow semantic vectors to avoid indexing merely remaining at the level of surface technical terms. This enables indexing of the core innovative points of the patent, improving the professionalism and accuracy of patent indexing. The extraction function is shown below: in, For deep semantic features; This is the activation function for deep semantics, which is consistent with the activation function for shallow semantics; This is a mapping matrix for deep semantics; This is a bias term for deep semantics; Substituting the shallow semantic vector data into the formula yields the deep semantic features: The patent document set was divided into a training set and a validation set in a 7:3 ratio. The text feature vectors in the training set were labeled. Then, the mapping matrices (the mapping matrices for shallow semantics and the mapping matrices for deep semantics) were initialized with random normality. The bias terms (the bias terms for shallow semantics and the bias terms for deep semantics) were initialized with zero. The matching error between the semantic vectors and the technical subdomain labels was used as the loss function. The semantic coding model was trained using the gradient descent method. The mapping matrix and bias terms were updated using backpropagation. Finally, the semantic coding accuracy of the semantic coding model was tested on the validation set. If the accuracy did not meet the preset requirements, the model learning rate was adjusted and training continued until the accuracy met the requirements, thus obtaining the final mapping matrix and bias terms. Secondly, construct a semantic matching model for index terms: Define a set of indexing terms, and generate a corresponding semantic vector for each indexing term in the set. The set of indexing terms is shown below: in, For indexing terms; For the first One indexing term; The total number of indexed terms; for example, selecting 10 core technical terms in the field of new energy vehicle power battery technology. ; For the first The semantic vector of each indexing term; for A dimensional real vector; Each index term in the index term set is divided into different levels: in, A hierarchical set; As the core innovation feature level; For general technical features level; This is a feature hierarchy for an example implementation; The indexing term set is defined based on an effective vocabulary, and a corresponding semantic vector is generated for each indexing term in the set. The indexing term set and the indexing term semantic vectors are shown in the table below: A matching score function is constructed and weighted coefficients at different levels are used to calculate the matching score between patent documents and indexing terms. For a patent set in a specific technical field, indexing terms with high matching scores are selected as core indexing terms for that field. Simultaneously, weighted coefficients at different levels ensure that the matching results between indexing terms and patents align with the technological development priorities of the corresponding field, making the indexing terminology library relevant to the patent technology characteristics of that field and improving the targeting of indexing. The matching score function is shown below: in, The matching score between patent documents and index terms; This is the transpose of the semantic vector of the indexing term; , and Weighting coefficients for different levels are assigned based on the hierarchical attributes of the indexing terms, with initial weights allocated from high to low technical importance. Core innovative features, being the core of the patent indexing, are assigned the highest initial value, while implementation features, being auxiliary, are assigned the lowest initial value. The judgment matrix is then adjusted by industry technical experts, taking into account the R&D priorities of the target technology field, to align with industry R&D trends. For example... , , ; Substitute the data into the formula to obtain the matching score, as shown in the table below: A domain-specific index is introduced, and the indexing weight is calculated using the Softmax function to reflect the importance of the indexing term to the patent semantics. The Softmax function is shown below: in, For indexing weights; The matching score between patent documents and index terms; It is a natural exponential function; The sum of the exponents of all matching scores; As a specificity index, the higher the correlation between the index term and the IPC classification, the closer the specificity index is to 1. For example, core innovative features... General technical features , Features of the embodiment Based on the core IPC classification number of the corresponding technical field, the index is quantitatively scored from two dimensions: whether the index term is a core technical term of the IPC classification and whether the index term is a characteristic term of the domain IPC classification, and then a specificity index is obtained. Substituting the matching score into the formula yields the indexing weights, as shown in the table below: Finally, the indexing weights are concatenated and reconstructed in order to generate the indexing vector. The reconstruction function is shown below: in, For indexing vectors; For the first The indexing weight of each indexing term; Generate indexing vectors based on indexing weight data: Step S3: The user retrieves patents and generates query vectors; based on the query vectors and indexing vectors, a similarity function is constructed to calculate their similarity; a filtering function is constructed to filter out candidate patent documents and generate corresponding new indexing vectors. First, the user retrieves a patent and generates a query vector, as shown below: in, For query vector; To generate the corresponding first... Individual index weights; Taking the user search query "Optimization of fast charging technology and thermal management for power batteries of new energy vehicles" as an example, the following query index weights are generated: Secondly, based on the query vector and index vector, a similarity function is constructed to calculate the similarity between the query vector and the index vector. The similarity function is shown below: in, The similarity between the query vector and the index vector is such that when the similarity approaches 1, it means that the query weight set is more similar to the index vector. Substitute the data into the formula to obtain the similarity score: Finally, a similarity threshold is preset, and a filtering function is constructed to filter candidate patent documents based on the similarity threshold, eliminating a large number of irrelevant patents, significantly reducing computational costs, and improving the efficiency of patent retrieval and analysis. The filtering function is as follows: in, A collection of candidate patent documents; To determine the similarity threshold, a subset of patent documents is randomly selected from the candidate patent document set as a test set. This simulates different types of user search needs and generates corresponding query vectors. The similarity distribution interval generated under the simulation is calculated, and the lower quartile of the interval is taken as the initial value of the similarity threshold. Based on the initial value, the similarity threshold is adjusted, and data under different similarity thresholds is calculated. The optimal value is taken as the final similarity threshold. For example, the similarity threshold... ; These are the shortlisted candidate patent documents. The total number of candidate patent documents; This is a filtering function that selects patent documents from all patent documents whose similarity between the query vector and the index vector is greater than a similarity threshold. The patent documents were screened to obtain a set of candidate patent documents: Based on the selected candidate patent documents, corresponding new indexing vectors are generated, as shown below: in, For the new indexing vector; For the first A new index weight.
[0019] Step S4: Construct an average semantic distance function and introduce patent time weights to calculate the average semantic distance of the new indexing vectors; construct a membership function to calculate the membership degree of candidate patent documents to the clustering topics and divide the candidate patent documents into different clustering topics; construct a clustering objective function, summarize and iteratively generate the set of patent documents for the clustering topics and generate a topic set; First, based on the new indexing vector, an average semantic distance function is constructed and patent time weights are introduced to calculate the average semantic distance. The average semantic distance function is shown below: in, The average semantic distance, To determine the preset number of cluster topics, an initial value for the number of cluster topics is determined based on the technical field classification standards of the candidate patent document set and the main classification number. The average semantic distance corresponding to different cluster topic counts is calculated. A curve is plotted with the number of cluster topics on the x-axis and the average semantic distance on the y-axis. The inflection point where the curve transitions from a rapid decline to a gentle flattening is considered the optimal number of cluster topics. For example, the preset number of cluster topics... Based on IPC classification, the core topics in the field of power batteries are "fast charging technology" and "thermal management system". Let the squared distance from the new index vector to the nearest cluster topic center be . For the first The center of a clustering topic; For patent time-weighted applications, the more recent the patent document filing date, the higher the weighting. The closer to 1, the more quantitative weights are assigned to patents with different application dates, using the latest application date in the candidate patent literature set as the baseline time and a linear decay method. Substituting the data into the formula, we obtain the average semantic distance: Secondly, based on the new indexing vector, the distance to the center of each cluster topic is calculated to clarify the semantic boundaries of different technology cluster topics and avoid technical overlap and mixing between patent topics. The function is as follows: in, Clustering of selected candidate patent documents into specific topics semantic distance; As a weighting factor for patent popularity, the higher the technology's popularity, The closer it is to 1, for example, the higher the patent popularity weight of fast charging technology. =0.95, Patent Heat Weight of Thermal Management System =0.84, the degree of hotspot for each technology cluster topic was determined by statistical analysis of patent application volume and industry R&D dynamics; by For example, substituting the data into the formula, we obtain the semantic distance: Based on the distance from the new indexing vector to the center of each cluster topic, the membership degree of candidate patent documents to the cluster topics is calculated. The membership degree calculation function is as follows: in, For candidate patent documents, the first Membership degree of each cluster topic; For candidate patent documents to all The sum of the reciprocals of the semantic distances of the cluster thematic centers; Based on membership, the selected candidate patent documents are divided into different cluster topics. The patent document division rules are as follows: in, For the first A collection of patent documents for a specific cluster topic, representing all patent documents whose semantics belong to that topic; Using the new index vectors of candidate patent documents as clustering samples, and different membership thresholds, cluster analysis is performed for each threshold. The cluster profile coefficient is calculated for each threshold, and the threshold with the largest profile coefficient is selected as the initial value. Finally, the patent documents in each cluster are examined to ensure that the core technical features of patents within the same cluster are consistent and that there is no significant technical overlap between different clusters. If overlap exists, the threshold is appropriately increased until the semantic boundary requirements are met. For example... ,when At the time of establishment, the candidate patent documents belonged to the [number missing] category. A clustering topic; The selected candidate patent documents are categorized into cluster topics: After a clustering operation is completed, the cluster centers are updated using the following update function: in, These are the updated cluster centers, used for calculation in the next iteration; For the first Each cluster topic contains the total number of patent documents; Substituting the clustered data into the formula, we obtain the updated cluster centers: Finally, a clustering objective function is constructed to measure the overall effectiveness of patent clustering. By determining whether the clustering converges, the optimality of the clustering effect is ensured. The clustering objective function is shown below: in, The clustering objective function; A judgment function is constructed and a preset judgment threshold is set to determine whether the clustering iteration has stopped. The judgment function is as follows: in, For the first The objective function value of each iteration; For the first The objective function value of each iteration; This represents the number of iterations. Using a preset judgment threshold and a set of candidate patent documents as clustering samples, with a preset initial number of clusters, the change in the clustering objective function value is counted in each iteration, and the value is taken when the change is less than the threshold for the first time. The value is used as the initial value for the judgment threshold, for example, the preset judgment threshold. ; After two rounds of iteration: The patent documents generated from the iterative clustering are summarized and a topic set is generated, as shown below: in, This is a collection of thematic works; Indicates the first to the second A collection of patent documents for each cluster topic; The thematic collection is as follows: Step S5: Construct the topical semantic mapping function, semantic index structure, and indexing vector index structure; construct the topical database; First, a topic semantic mapping function is constructed for fast database retrieval. The topic semantic mapping function is shown below: in, To select the target topics to which candidate patent documents belong; For candidate patent documents, the first Membership degree of each cluster topic; This is a preset membership threshold. For clustering topic indexes; This represents the total number of clustered topics. Substitute the data into the formula to filter out the target topics to which the candidate patent documents belong: Secondly, a semantic index structure is constructed to improve the retrieval efficiency of the database. The semantic index structure is shown below: in, For clustering topics The corresponding patent index set; To select candidate patent documents; For clustering topic indexes; Substituting the data into the formula yields the semantic index structure: Construct an indexing vector index structure for nearest neighbor search in semantic retrieval. The indexing vector index structure is shown below: in, For clustering topics The corresponding indexing vector index; For the new indexing vector; To select candidate patent documents; Substituting the data into the formula, we obtain the indexing vector index: Finally, based on the patent document set of the clustering topics, the updated cluster centers, the patent index set, and the indexing vector index, a thematic database is constructed, as shown below: in, For specialized databases; A collection of patent documents for clustering specific topics; The updated cluster centers; For a collection of patent indexes; For indexing vectors; Based on the above data, a thematic database is constructed, as shown in the table below: The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for deep indexing of patents and construction of thematic databases based on semantic retrieval, characterized in that, The steps are as follows: Step S1: Obtain patent documents and define their structure; concatenate the defined text content to form complete text content; perform word segmentation on the complete text content to generate a preliminary vocabulary set; filter and remove duplicates from all technical terms in the preliminary vocabulary set to generate an effective vocabulary list. Step S2: Construct a semantic matching model for indexing terms, build a set of indexing terms, and divide each indexing term in the set into different levels; The indexing weights are concatenated and reconstructed to generate the indexing vector; Step S3: The user retrieves patents and generates query vectors; based on the query vectors and indexing vectors, a similarity function is constructed to calculate their similarity; a filtering function is constructed to filter out candidate patent documents and generate corresponding new indexing vectors. Step S4: Construct the average semantic distance function and introduce patent time weights to calculate the average semantic distance of the new indexing vectors; construct the membership function; construct the clustering objective function, summarize and iteratively generate a set of patent documents for clustering topics, and generate a topic set; Step S5: Construct the topical semantic mapping function, semantic index structure, and indexing vector index structure; construct the topical database.
2. The method according to claim 1, characterized in that, The complete text content in step S1 is as follows: in, For the first The complete text of the patent document; For the first The title of the patent document; For the first Abstracts of patent documents; For the first The text of the claims in a patent document; For the first The specification text of the patent document; For patent literature indexing; For text concatenation; The preliminary vocabulary set is shown below: in, This is a preliminary vocabulary set; For the first A technical term; Total number of technical terms; The list of valid vocabulary is as follows: in, For a valid vocabulary list; Total number of valid words; For the first One effective vocabulary; For vocabulary indexing.
3. The method according to claim 1, characterized in that, In step S1, the frequency of each word is calculated based on the effective vocabulary, and a statistical weight is added to the frequency of words in the claim text. The word frequencies are as follows: in, Frequency weighting; For the vocabulary list, number The word, in the 1st... Total number of times it appears in the patent document; For the first The sum of the occurrences of all words in the effective vocabulary list in a patent document; The number of times the word appears in the title; The number of times a word appears in the summary; The number of times a word appears in the text of the claims. The term frequency weighting coefficients in the claim text; The number of times a word appears in the instruction manual text; A semantic encoding model is constructed to generate shallow semantic vectors. The semantic encoding model is shown below: in, This is a shallow semantic vector; For shallow semantic activation functions, It is a mask diagonal matrix; This is a mapping matrix for shallow semantics; For the first Text feature vectors in a patent document; This is a bias term for shallow semantics.
4. The method according to claim 1, characterized in that, The different levels in step S2 are shown below: in, For the first One indexing term; A hierarchical set; As the core innovation feature level; For general technical features level; This is a feature hierarchy for an example implementation; A matching score function is constructed and combined with the weight coefficients of different levels to calculate the matching score between the patent document and the indexing term. The matching score function is as follows: in, The matching score between patent documents and index terms; This is the transpose of the semantic vector of the indexing term; , and These are the weighting coefficients for different levels; It is a deep semantic feature.
5. The method according to claim 1, characterized in that, In step S2, a specificity index of the indexing term's level is introduced, and the indexing weight is calculated using the Softmax function, which is shown below: in, For indexing weights; The matching score between patent documents and index terms; It is a natural exponential function; The sum of the exponents of all matching scores; This is a specificity index.
6. The method according to claim 1, characterized in that, The filtering function in step S3 is as follows: in, A collection of candidate patent documents; The similarity threshold; For the first Patent documents; To determine the similarity between the query vector and the index vector; These are the shortlisted candidate patent documents. The total number of candidate patent documents.
7. The method according to claim 1, characterized in that, In step S4, the distance to each cluster topic center is calculated based on the new indexing vector, using the following function: in, To cluster the selected candidate patent documents into specific topics semantic distance; As a weighting factor for patent popularity; For the first The center of a clustering topic, For the new indexing vector; The membership function is as follows: in, For candidate patent documents, the first Membership degree of each cluster topic; For candidate patent documents to all The sum of the reciprocals of the semantic distances of the cluster thematic centers; The candidate patent documents are divided into different cluster topics, and the patent document division rules are as follows: in, For the first A collection of patent documents for each cluster topic; This is a preset membership threshold. To select candidate patent documents; The total number of clustered topics; This is a clustering-based thematic index.
8. The method according to claim 1, characterized in that, The clustering objective function in step S4 is as follows: in, The clustering objective function; For the first The center of a clustering topic, For the new indexing vector; These are the shortlisted candidate patent documents; The preset number of clustering topics; For the first A collection of patent documents for each cluster topic; For clustering topic indexes; A judgment function is constructed and a judgment threshold is preset. The judgment function is as follows: in, For the first The objective function value of each iteration; For the first The objective function value of each iteration; This represents the number of iterations. This is the preset judgment threshold.
9. The method according to claim 1, characterized in that, The semantic index structure in step S5 is as follows: in, For clustering topics The corresponding patent index set; To select candidate patent documents; For clustering topic indexes; To select the target topics to which candidate patent documents belong; The indexing vector structure is shown below: in, For the topic The corresponding indexing vector index; For the new indexing vector; To select candidate patent documents; The thematic database is shown below: in, For specialized databases; A collection of patent documents; The updated cluster centers; For a collection of patent indexes; This is the index for the indexing vector.