Patent Classification Retrieval Method and System Based on Multivariate Data Fusion

By constructing and updating the domain knowledge base in real time, combining intelligent analysis technology, the problem of inconsistency between patent classification and search is solved, and efficient and accurate classification and search of patent documents is achieved, and it is suitable for patent information management system for multivariate data fusion.

CN119739847BActive Publication Date: 2025-07-22GUANGXI SOUTH CHINA TECHNOLOGY EXCHANGE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411919692.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-07-22
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

The existing patent classification and search technology have shortcomings in accuracy and efficiency, especially in the terms in cross-domain or multi-field literature, resulting in inaccurate classification and uncorrelated search results, and lack of dynamic updates and intelligent analysis of the technical context and meaning of terminology.

Method used

Build a domain knowledge base, update term definitions in real time, and evaluate the meaning and semantic accuracy of terminology in patent literature through multivariate data fusion and intelligent analysis technology, establish a semantic accurate identification mechanism, and improve the accuracy and efficiency of patent classification and retrieval.

Benefits of technology

It achieves consistency between the terms of patent literature and the standard terms in the knowledge base, improves the accuracy of patent classification and the relevance of search results, ensures efficient use of patent information, and can quickly locate patent documents that meet technical needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119739847B_ABST
    Figure CN119739847B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of patent classification retrieval, and specifically discloses a patent classification retrieval method and system based on multi-source data fusion, including real-time updating of the domain knowledge base, dynamically adjusting the definition and application of terms to ensure the consistency and accuracy of terms in different technical fields; the present invention combines intelligent analysis technology to deeply analyze the terms and their contexts in patent documents, accurately understand the meanings of terms, and evaluate the semantic accuracy of patent texts, avoiding understanding deviations caused by term ambiguity or unclear context; by analyzing the technical problems and solutions in patent texts and combining the technical context and the technical background information of the domain knowledge base, the precision of patent text analysis is further improved, thereby enhancing the precision of patent classification and the relevance of retrieval results; the present invention optimizes the classification and retrieval processes of patent documents by establishing an accurate semantic recognition mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of patent classification and retrieval, and specifically relates to a patent classification and retrieval method and system based on multi-source data fusion. Background Art

[0002] With the rapid development of global technological innovation, the accumulation of patent literature in various technical fields has become increasingly huge. How to efficiently and accurately classify, retrieve, and analyze patents has become an important issue in intellectual property management. Currently, the classification and retrieval of patent literature mostly rely on the standardization of terms and the understanding of technical contexts. However, in the prior art, the definitions and applications of terms often vary, especially in cross-domain or multi-domain technical documents, where the differences in the use of terms among different patents are significant, resulting in inaccurate patent classification, irrelevant retrieval results, and even affecting the efficiency and quality of patent examination. The existing patent classification and retrieval methods fail to effectively integrate the domain knowledge base, lacking dynamic updates and intelligent analysis of technical contexts and term meanings, often leading to term ambiguities and unclear contexts.

[0003] The prior art has the following deficiencies:

[0004] The existing patent classification and retrieval technologies have many deficiencies in terms of accuracy and efficiency, mainly manifested in the inability to adjust the definitions and applications of terms in real time and dynamically, resulting in low classification accuracy and poor relevance of retrieval results. Especially when the technical contexts of patent literature change significantly, the prior art lacks in-depth analysis of technical content and semantic accurate recognition mechanisms, and cannot effectively solve the diversity and complexity of term meanings. Therefore, in view of these problems in the prior art, the present invention proposes a dynamic update mechanism based on the domain knowledge base, combined with intelligent analysis means, to accurately identify the meanings of terms in patent literature, and through in-depth analysis of technical contexts, to improve the accuracy and efficiency of patent classification and retrieval, thereby effectively solving the problems of term inconsistency and inaccurate classification and retrieval in the prior art. Summary of the Invention

[0005] The purpose of the present invention is to provide a patent classification and retrieval method and system based on multi-source data fusion to solve the problems in the above background.

[0006] The purpose of the present invention can be achieved by the following technical solutions:

[0007] A patent classification and retrieval method based on multi-source data fusion includes the following steps:

[0008] S1: Construct a domain knowledge base for the technical fields related to patents, including: terms and technical background information, to provide a precise semantic benchmark for patent texts and ensure semantic consistency in classification and retrieval;

[0009] Among them, the terms include: technical vocabulary, professional terms, and abbreviations; the technical background information includes: the application scenarios of the terms, the development history of the relevant technical fields, the methods of technical implementation, technical problems and solutions;

[0010] S2: According to the constructed domain knowledge base, the knowledge base is updated in real time, and the definitions and applications of the terms are dynamically adjusted according to different technical fields. The technical terms in all patent documents are comprehensively analyzed with the standard terms in the knowledge base to evaluate the accuracy of patent classification and retrieval;

[0011] S3: By intelligently analyzing the terms in the patent document and the technical content in the context, analyze the meaning of the terms to evaluate the semantic accuracy of the text information;

[0012] S4: According to the accurate patent classification and retrieval and the accurate semantics of the text information, and combined with the technical context and domain knowledge of the patent text, establish a semantic accurate recognition mechanism in the classification and retrieval process.

[0013] As a further solution of the present invention: The construction of the domain knowledge base related to the patent technology fields specifically includes:

[0014] For each technical field, obtain the terms and technical background information of the corresponding technical field, analyze the terms and technical background information of the patent documents, generate a comprehensive feature index according to the analysis results, and evaluate the semantic consistency of the text in classification and retrieval according to the comprehensive feature index.

[0015] As a further solution of the present invention: The process of obtaining the comprehensive feature index is as follows:

[0016] Convert the terms and their background information in the patent documents into vector representations;

[0017] Convert each term and technical background into a vector representation of a fixed dimension;

[0018] For each term T in the patent document i , obtain the word vector representation through the word embedding method

[0019] Obtain the technical background information B j , and convert the technical background information into a background vector by the method based on document embedding

[0020] For each term T in each patent document i and the corresponding technical background information B j , calculate the semantic matching degree between the term and the background through the cosine similarity;

[0021] Among them, the calculation expression of the cosine similarity is as follows:

[0022]

[0023] In the formula, represents the similarity between the terms in the patent document and the technical background information, represents the dot product of the term vector and the background vector, represents the Euclidean norm of the term vector, represents the Euclidean norm of the background vector, i represents the i-th term in the patent document, and j represents the j-th item in the technical background information;

[0024] For each term in the patent document, calculate the similarity with all technical background information and obtain a similarity vector, and calculate the comprehensive feature index through weighted average;

[0025] Among them, the calculation expression of the comprehensive feature index is as follows:

[0026]

[0027] In the formula, C i represents the comprehensive feature index, n represents the total number of technical background information, and w j represents the weight coefficient of the technical background information B j .

[0028] As a further solution of the present invention: comprehensively analyze the technical terms in all patent documents and the standard terms in the knowledge base to evaluate the accuracy of patent classification and retrieval, specifically including:

[0029] Obtain the technical terms in the patent document, calculate the embedding representation on the text and image in the patent document, and generate a technical term accuracy coefficient according to the calculation result;

[0030] Obtain the standard terms in the knowledge base of the patent document, analyze the connection degree and semantic consistency of the term nodes, and calculate the standard term accuracy coefficient according to the analysis result;

[0031] Comprehensively calculate and analyze the technical term accuracy coefficient and the standard term accuracy coefficient, and calculate the accuracy coefficient according to the analysis result to evaluate the accuracy of patent classification and retrieval.

[0032] As a further solution of the present invention: the process of obtaining the technical term accuracy coefficient is as follows:

[0033] Represent the text part of the patent document as an embedding vector and extract the image features associated with the technical terms in the patent document;

[0034] Fuse text embeddings and image embeddings to form multimodal embedding vectors for terms

[0035] According to the multimodal embedding vectors, calculate the text-image consistency score of technical terms through cosine similarity calculation technology;

[0036] Calculate the semantic weight of technical terms according to the occurrence frequency of technical terms in the literature and the importance weight in the technical field knowledge base. The calculation expression is:

[0037] W = log(1 + f) * w1;

[0038] In the formula, W represents the semantic weight, w1 represents the importance weight, and f represents the occurrence frequency of technical terms in patent literature;

[0039] Comprehensively process the consistency score and the semantic weight to calculate the accuracy coefficient of technical terms. The calculation expression is:

[0040]

[0041] In the formula, TTAC represents the accuracy coefficient of technical terms, λ represents the adjustment parameter, F represents the consistency score, and max(W) represents the maximum semantic weight in the term set.

[0042] As a further solution of the present invention: the process of obtaining the standard term accuracy coefficient is:

[0043] Structuralize the term definitions, domain backgrounds, synonyms, and hyponymy relationships in the knowledge base into a graph;

[0044] The graph nodes represent terms, and the graph edges represent the semantic relationships between terms;

[0045] Normalize the frequency of standard terms in the literature within the corresponding technical field to obtain the domain weight parameter;

[0046] Sum up all the domain weight parameters in the patent literature to obtain the connectivity G of the language nodes;

[0047] Calculate the semantic similarity between the standard term definition and its hyponym and hypernym nodes. The calculation expression is:

[0048]

[0049] In the formula, T a represents the standard term, a and b represent the number of standard terms in the patent literature, sin represents the similarity between the standard term vectors, N represents the total number of standard terms, and S(T a ) represents the semantic similarity between the standard term definition and its hyponym and hypernym nodes;

[0050] Calculate the ratio of the connection degree G of the language node to the semantic similarity S(T a ) to obtain the standard term accuracy coefficient, denoted as TTAB.

[0051] As a further solution of the present invention: Analyze the meaning of the terms and evaluate the semantic accuracy of the text information, specifically including:

[0052] Among them, the terms include technical terms and standard terms;

[0053] Model the context content of the terms in the patent document as a graph;

[0054] Use the graph convolutional network to aggregate the neighborhood features of the nodes layer by layer, update the semantic representation of the nodes, and the calculation expression is:

[0055]

[0056] In the formula, represents the feature vector of node q at the l+1 layer, q represents the number of nodes, z represents the number of neighbor nodes, represents the set of neighbor nodes of node q, d q represents the connection number of node q, d z represents the connection number of neighbor node z, W (l) represents the weight matrix of the l layer, l is the number of graph convolutional layers, b (l) represents the bias vector of the l layer, and σ represents the activation function;

[0057] Calculate the semantic similarity between the term and the context through the cosine similarity of the updated feature vectors of the term node and the context node;

[0058] Calculate the semantic accuracy coefficient of the term by aggregating the similarity scores between the term node and all its neighbor nodes, and the calculation expression is:

[0059]

[0060] In the formula, Y q represents the semantic accuracy coefficient, represents the set of neighbor nodes, x(q) represents the semantic similarity score between the term and the neighbor nodes, and q represents the number of nodes.

[0061] As a further solution of the present invention: Establish a semantic accurate recognition mechanism in the classification and retrieval process, specifically including:

[0062] Normalize the semantic accuracy coefficient of patent documents and the accuracy coefficient of patent document classification and retrieval, calculate the comprehensive recognition score, and based on the comprehensive recognition score, divide the semantic accuracy of patent classification and retrieval, and divide the semantics of document classification and retrieval into qualified semantics and unqualified semantics.

[0063] A patent classification and retrieval system based on multi-source data fusion, comprising:

[0064] A domain knowledge base construction module, which constructs a domain knowledge base for the technical field related to patents, including: terms and technical background information, to provide an accurate semantic benchmark for patent texts and ensure semantic consistency in classification and retrieval;

[0065] A domain knowledge base update and term standardization module, which updates the domain knowledge base in real time, dynamically adjusts the definition and application of terms according to different technical fields, analyzes the matching degree between the terms in patent documents and the standard terms in the knowledge base, and evaluates the accuracy of patent classification and retrieval;

[0066] A term semantic analysis and evaluation module, which evaluates the semantic accuracy of terms by intelligently analyzing the terms and their contexts in patent documents;

[0067] A semantic accurate recognition and classification retrieval module, which, based on accurate patent classification and retrieval, combines the technical context and domain knowledge of patent texts to establish an effective semantic accurate recognition mechanism, and improves the accuracy of patent text classification and retrieval.

[0068] The beneficial effects of the present invention:

[0069] (1) By constructing and continuously updating a domain knowledge base and combining advanced semantic processing technologies, the present invention realizes the consistency between the terms in patent documents and the standard terms in the knowledge base, and successfully solves the accuracy problem caused by term inconsistency in patent classification and retrieval; this knowledge base not only contains various professional terms, but also covers the application scenarios of terms, technical background information, and the evolution process of related technical fields, thus providing in-depth technical context support; through the dynamic update and real-time adjustment of terms, the present invention can adapt to the changing technological progress, ensuring that patent documents in each technical field can be classified under the latest and most accurate semantic framework; this method improves the accuracy of patent document classification, avoids the possible errors and biases in the traditional classification system, and ensures a high correlation between the retrieval results and user needs, making patent retrieval more accurate and efficient, enabling users to quickly locate the patent documents that best meet their technical needs, and effectively improving the utilization efficiency and retrieval effect of patent information;

[0070] (2) Through intelligent analysis means, the present invention deeply excavates the terms and their context information in patent documents, realizing accurate interpretation of the meanings of terms and efficient evaluation of semantic accuracy; by comprehensively analyzing the usage of terms in specific technical contexts and combining the rich technical background information and development context in the domain knowledge base, the present invention can effectively eliminate potential understanding deviations caused by term ambiguity or unclear context, ensuring comprehensive and accurate understanding of various technical contents in patent texts; especially in identifying technical problems and solutions in patent documents, the present invention precisely locates and analyzes the innovation points and technical elements in patent documents through in-depth association of terms with technical backgrounds, thereby improving the analysis accuracy and semantic accuracy of patent texts; this innovative method provides more reliable support for processes such as patent examination, retrieval, and technical evaluation, ensuring the accuracy of patent classification and retrieval, and further promoting technological innovation and precise information management in the field of intellectual property rights. Brief Description of the Drawings

[0071] The present invention will be further described below with reference to the accompanying drawings.

[0072] Figure 1 is a flowchart of the specific steps of the patent classification and retrieval method based on multi-source data fusion of the present invention;

[0073] Figure 2 is a flowchart of the patent classification and retrieval system based on multi-source data fusion in the present invention. Detailed Embodiments

[0074] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0075] Please refer to Figure 1 as shown, the present invention is a patent classification and retrieval method based on multi-source data fusion, including the following steps:

[0076] S1: Construct a domain knowledge base for the technical field related to patents, including: terms and technical background information, which are used to provide an accurate semantic benchmark for patent texts and ensure semantic consistency in classification and retrieval;

[0077] Among them, the terms include: technical vocabulary, professional terms, abbreviations; the technical background information includes: application scenarios of terms, development history of related technical fields, methods of technical implementation, technical problems and solutions;

[0078] S2: According to the constructed domain knowledge base, the knowledge base is updated in real time, and the definitions and applications of terms are dynamically adjusted according to different technical fields. The technical terms in all patent documents are comprehensively analyzed with the standard terms in the knowledge base to evaluate the accuracy of patent classification and retrieval;

[0079] S3: By intelligently analyzing the terms and technical content in the patent document, analyze the meaning of the terms to evaluate the semantic accuracy of the text information;

[0080] S4: According to the accurate patent classification and retrieval and the accurate semantics of the text information, combined with the technical context and domain knowledge of the patent text, establish a semantic accurate recognition mechanism in the classification and retrieval process to improve the retrieval accuracy and efficiency.

[0081] In S1, construct a domain knowledge base for the technical fields related to patents, including: terms and technical background information, which are used to provide a precise semantic benchmark for patent texts to ensure semantic consistency in classification and retrieval. Specifically, it includes:

[0082] For each technical field, obtain the terms and technical background information of the corresponding technical field, analyze the terms and technical background information of the patent document, generate a comprehensive feature index according to the analysis results, and evaluate the semantic consistency of the text in classification and retrieval according to the comprehensive feature index;

[0083] Among them, the process of obtaining the comprehensive feature index is as follows:

[0084] When constructing the domain knowledge base for the technical fields related to patents, use cosine similarity to calculate the comprehensive feature index. First, the terms and technical background information in the patent document need to be vectorized, and then the similarity between the text and the standard terms in the domain knowledge base is calculated, and then a comprehensive feature index is generated to evaluate the semantic consistency of the text. Specifically, it includes:

[0085] Convert the terms and their background information in the patent document into vector representations;

[0086] Convert each term or technical background into a vector representation of a fixed dimension;

[0087] For each term T in the patent document i , obtain the word vector representation through the word embedding method

[0088] Obtain the technical background information B j , and convert the technical background information into a background vector using the method of document-based embedding

[0089] For each term T in each patent document i and the corresponding technical background information Bj , the semantic matching degree between the term and the background is calculated by cosine similarity;

[0090] Among them, the calculation expression of the cosine similarity is:

[0091]

[0092] In the formula, represents the similarity between the term in the patent document and the technical background information, represents the dot product of the term vector and the background vector, represents the Euclidean norm of the term vector, represents the Euclidean norm of the background vector, i represents the i-th term in the patent document, and j represents the j-th item in the technical background information;

[0093] For each term in the patent document, calculate the similarity with all technical background information, and obtain a similarity vector, and calculate the comprehensive feature index through weighted average;

[0094] Among them, the calculation expression of the comprehensive feature index is:

[0095]

[0096] In the formula, C i represents the comprehensive feature index, n represents the total number of technical background information, w j represents the weight coefficient of the technical background information B j ;

[0097] Compare the comprehensive feature index with a preset threshold;

[0098] If the comprehensive feature index is greater than or equal to the preset threshold, it indicates that the semantics of the corresponding patent are consistent in classification and retrieval;

[0099] If the comprehensive feature index is less than the preset threshold, it indicates that the semantics of the corresponding patent are inconsistent in classification and retrieval;

[0100] It should be noted that: the comprehensive feature index reflects whether the semantics of the patent are consistent in the classification and retrieval process, and when the value of the comprehensive feature index is larger, the degree of semantic consistency in the corresponding patent classification and retrieval process is higher.

[0101] In S2, according to the constructed domain knowledge base, the knowledge base is updated in real time, and the definitions and applications of terms are dynamically adjusted according to different technical fields. The technical terms in all patent documents are comprehensively analyzed with the standard terms in the knowledge base to evaluate the accuracy of patent classification and retrieval, specifically including:

[0102] Obtain technical terms in patent documents, calculate the embedding representations on the text and images in patent documents, and generate a technical term accuracy coefficient based on the calculation results;

[0103] Obtain standard terms in the knowledge base of patent documents, analyze the connection degree and semantic consistency of term nodes, and calculate a standard term accuracy coefficient based on the analysis results;

[0104] Perform comprehensive calculation and analysis on the technical term accuracy coefficient and the standard term accuracy coefficient, and calculate an accuracy coefficient based on the analysis results for evaluating the accuracy of patent classification and retrieval, specifically including:

[0105] Among them, the process of obtaining the technical term accuracy coefficient is as follows:

[0106] Represent the text part of the patent document as an embedding vector, and extract the image features associated with technical terms in the patent document;

[0107] Fuse the text embedding and the image embedding to form a multi-modal embedding vector of the term

[0108] According to the multi-modal embedding vector, calculate the text and image consistency score of the technical term through cosine similarity;

[0109] According to the occurrence frequency of the technical term in the document and the importance weight in the technical field knowledge base, calculate the semantic weight of the technical term. The calculation formula is:

[0110] W = log(1 + f) * w1;

[0111] In the formula, W represents the semantic weight, w1 represents the importance weight, and f represents the occurrence frequency of the technical term in the patent document;

[0112] Perform comprehensive processing on the consistency score and the semantic weight, and calculate the technical term accuracy coefficient. The calculation formula is:

[0113]

[0114] In the formula, TTAC represents the technical term accuracy coefficient, λ represents the adjustment parameter, F represents the consistency score, and max(W) represents the maximum semantic weight in the term set;

[0115] Among them, the process of obtaining the standard term accuracy coefficient is as follows:

[0116] Structuralize the term definitions, domain backgrounds, synonyms, hyponymy relationships, etc. in the knowledge base into a graph;

[0117] The graph nodes represent terms, and the graph edges represent the semantic relationships between terms;

[0118] Normalize the frequency of standard terms in the literature within the corresponding technical field to obtain the domain weight parameter;

[0119] Sum up all the domain weight parameters in the patent literature to obtain the connection degree G of the language node;

[0120] Calculate the semantic similarity between the definition of the standard term and its upper and lower nodes. The calculation expression is:

[0121]

[0122] In the formula, T a represents the standard term, a and b represent the number of standard terms in the patent literature, sin represents the similarity between the standard term vectors, N represents the total number of standard terms, and S(T a ) represents the semantic similarity between the definition of the standard term and its upper and lower nodes;

[0123] Calculate the ratio of the connection degree G of the language node to the semantic similarity S(T a ) to obtain the standard term accuracy coefficient, denoted as TTAB;

[0124] Among them, the calculation expression of the accuracy coefficient is:

[0125]

[0126] In the formula, Z d represents the accuracy coefficient, o1 and o2 represent preset proportionality coefficients, and both o1 and o2 are greater than 0.

[0127] In S3, by intelligently analyzing the terms in the patent literature and the technical content in the context, analyze the meaning of the terms and evaluate the semantic accuracy of the text information, specifically including:

[0128] Among them, the terms include technical terms and standard terms;

[0129] Model the context content of the terms in the patent literature as a graph;

[0130] Use the graph convolutional network to layer by layer aggregate the neighborhood features of the nodes and update the semantic representation of the nodes. The calculation expression is:

[0131]

[0132] In the formula, represents the feature vector of node q at the l + 1 layer, q represents the number of nodes, z represents the number of neighbor nodes, represents the set of neighbor nodes of node q, d q represents the connection number of node q, d zThe number of connections of the node representing the neighbor node z, W (l) Represents the weight matrix of the l-th layer, l is the number of graph convolutional layers, b (l) Represents the bias vector of the l-th layer, and σ represents the activation function;

[0133] The feature vectors of the updated term nodes and context nodes are used to calculate the semantic similarity between the term and the context through cosine similarity;

[0134] By aggregating the similarity scores between the term node and all its neighbor nodes, the semantic accuracy coefficient of the term is calculated, and the calculation expression is:

[0135]

[0136] In the formula, Y q Represents the semantic accuracy coefficient, Represents the set of neighbor nodes, x(q) represents the semantic similarity score between the term and the neighbor nodes, and q represents the number of nodes;

[0137] Compare the semantic accuracy coefficient with a preset threshold;

[0138] If the semantic accuracy coefficient is greater than or equal to the preset threshold, it indicates that the semantics of the corresponding patent document term is accurate;

[0139] If the semantic accuracy coefficient is less than the preset threshold, it indicates that the semantics of the corresponding patent document term is inaccurate.

[0140] In S4, according to the accurate patent classification and retrieval of the semantics of accurate text information, and combined with the technical context and domain knowledge of the patent text, a semantic accurate recognition mechanism in the classification and retrieval process is established to improve the retrieval accuracy and efficiency, specifically including:

[0141] Normalize the semantic accuracy coefficient of the patent document and the accuracy coefficient of the patent document classification and retrieval, calculate the comprehensive recognition score, and divide the semantic accuracy of the patent classification and retrieval according to the comprehensive recognition score, and divide the semantics of the document classification and retrieval into qualified semantics and unqualified semantics;

[0142] Among them, the calculation expression of the comprehensive recognition score is:

[0143]

[0144] In the formula, E z Represents the comprehensive recognition score;

[0145] Perform semantic accuracy recognition on each patent document group during classification and retrieval, including:

[0146] Compare the comprehensive recognition score with a preset threshold;

[0147] If the comprehensive recognition score is greater than or equal to the preset threshold, it indicates that the corresponding patent document group has high semantic accuracy in classification and retrieval, and is recorded as qualified semantics;

[0148] If the comprehensive recognition score is less than the preset threshold, it indicates that the corresponding patent document group has low semantic accuracy in classification and retrieval, and is recorded as unqualified semantics.

[0149] Please refer to Figure 2 as shown, the patent classification and retrieval system based on multi-source data fusion includes:

[0150] A domain knowledge base construction module, which constructs a domain knowledge base for the technical field related to patents, including: terms and technical background information, to provide an accurate semantic benchmark for patent texts and ensure semantic consistency in classification and retrieval;

[0151] A domain knowledge base update and term standardization module, which updates the domain knowledge base in real time, dynamically adjusts the definition and application of terms according to different technical fields, analyzes the matching degree between the terms in patent documents and the standard terms in the knowledge base, and evaluates the accuracy of patent classification and retrieval;

[0152] A term semantic analysis and evaluation module, which evaluates the semantic accuracy of terms by intelligently analyzing the terms and their contexts in patent documents;

[0153] A semantic accurate recognition and classification retrieval module, which, based on accurate patent classification and retrieval, combines the technical context and domain knowledge of patent texts to establish an effective semantic accurate recognition mechanism and improve the accuracy of patent text classification and retrieval.

[0154] Working principle of the present invention: By constructing a domain knowledge base in the technical field related to patents, including technical terms, background information, application scenarios, technical development history, etc., to provide an accurate semantic benchmark for patent texts and ensure semantic consistency in classification and retrieval; on this basis, calculate the semantic matching degree between terms and technical background information through cosine similarity to generate a comprehensive feature index, thereby evaluating the semantic consistency of patent documents in classification and retrieval; secondly, the method intelligently analyzes the terms and context in patent documents, uses a graph neural network to aggregate the features of term nodes and neighbor nodes, calculates the semantic accuracy coefficient of terms, and further analyzes the semantic precision of patent documents. In addition, the accuracy coefficients of technical terms and standard terms in patent documents are calculated through multi-modal embedding vectors, and the connection degree and semantic similarity of term nodes are analyzed in combination with the graph spectrum to evaluate the accuracy of patent classification and retrieval; finally, based on the accuracy coefficient and semantic consistency, a semantic accurate recognition mechanism in the classification and retrieval process is established, and patent documents are semantically accurately divided through a comprehensive recognition score to optimize the retrieval accuracy and efficiency. This method can dynamically adjust term definitions, improve the classification and retrieval effect of patent documents, and is applicable to patent information management and retrieval systems in various technical fields.

[0155] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula that is closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0156] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0157] It should be understood that the term "and / or" in this text is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this text generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, and the specific meaning can be understood by referring to the context before and after.

[0158] It should be understood that in various embodiments of this application, the magnitudes of the serial numbers of the above processes do not imply the sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0159] The above has described in detail one embodiment of the present invention, but the content described is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.

Claims

1. A patent classification retrieval method based on multi-source data fusion, characterized in that It includes the following steps: S1: Construct a domain knowledge base for the technical field related to patents, including: standard terms and technical background information, which are used to provide an accurate semantic benchmark for patent texts and ensure semantic consistency in classification and retrieval. Specifically, it includes: For each technical field, obtain the standard terms and technical background information of the corresponding technical field, analyze the technical terms and technical background information in patent documents, generate a comprehensive feature index according to the analysis results, and evaluate the semantic consistency of the text in classification and retrieval according to the comprehensive feature index; S2: According to the constructed domain knowledge base, update the knowledge base in real time, dynamically adjust the definitions and applications of technical terms according to different technical fields, comprehensively analyze the technical terms in all patent documents and the standard terms in the knowledge base, and evaluate the accuracy of patent classification and retrieval. Specifically, it includes: Obtain the technical terms in the patent document, calculate the embedding representations on the text and images in the patent document, and generate a technical term accuracy coefficient according to the calculation results; Obtain the standard terms in the knowledge base, analyze the connection degree and semantic consistency of the term nodes, and calculate the standard term accuracy coefficient according to the analysis results; Comprehensively calculate and analyze the technical term accuracy coefficient and the standard term accuracy coefficient, and calculate the accuracy coefficient according to the analysis results to evaluate the accuracy of patent classification and retrieval; S3: Analyze the meaning of technical terms by intelligently analyzing the technical terms and technical content in the context of the patent document, and evaluate the semantic accuracy of the text information. Specifically, it includes: Model the context content of the technical terms in the patent document as a graph; Use the graph convolutional network to aggregate the neighborhood features of the nodes layer by layer, update the semantic representation of the nodes, and the calculation expression is: In the formula, represents the feature vector of node q at the (l + 1)-th layer, q represents the q-th node, and z represents the z-th neighbor node. represents the set of neighbor nodes of node q, and d q represents the number of connections of node q, and d z represents the number of connections of neighbor node z, and W (l) represents the weight matrix of the l-th layer, l represents the l-th convolutional layer, and b (l) represents the bias vector of the l-th layer, and σ represents the activation function. The feature vectors of the updated technical term nodes and context nodes are used to calculate the semantic similarity between the term and the context through cosine similarity; By aggregating the similarity scores of the technical term nodes and all their neighbor nodes, calculate the semantic accuracy coefficient of the technical terms, and the calculation expression is: Where Y q represents the semantic accuracy coefficient, represents the set of neighbor nodes, represents the set the number of elements in, and x(q, z) represents the semantic similarity score between the technical term and the neighbor nodes; S4: Establish a semantic accurate recognition mechanism in the classification and retrieval process according to the accurate patent classification and retrieval and the semantics of accurate text information, and combine the technical context and domain knowledge of the patent text.

2. The method for patent classification retrieval based on multi-source data fusion according to claim 1, wherein The process of obtaining the comprehensive feature index is: Convert the technical terms and their background information in the patent document into vector representations; Convert each technical term and technical background into a vector representation of a fixed dimension; For each technical term T in the patent literature i , a word vector representation is obtained through the word embedding method Obtain the technical background information B j , and convert the technical background information into a background vector using a document-based embedding method For each technical term T in a patent document i and the corresponding technical background information B j , the semantic matching degree between the technical term and the background is calculated by cosine similarity; Among them, the calculation expression of the cosine similarity is: In the formula, represents the similarity between the technical terms and the technical background information in the patent document, represents the dot product of the technical term vector and the background vector, represents the Euclidean norm of the technical term vector, represents the Euclidean norm of the background vector, i represents the i-th technical term in the patent document, and j represents the j-th item in the technical background information; For the technical terms in each patent document, calculate the similarity with all technical background information, obtain a similarity vector, and calculate the comprehensive feature index through weighted average; Among them, the calculation expression of the comprehensive feature index is: Where C i represents the comprehensive feature index, n represents the total number of technical background information, and w j represents the weight coefficient of the technical background information B j .

3. The method for patent classification retrieval based on multi-source data fusion according to claim 1, wherein, The process of obtaining the technical term accuracy coefficient is: Represent the text part of the patent document as an embedding vector, and extract the image features associated with the technical terms in the patent document; Fuse text embeddings and image embeddings to form multimodal embedding vectors for terms According to the multi-modal embedding vector, calculate the text and image consistency score of the technical terms through cosine similarity; Calculate the semantic weight of technical terms according to the frequency of occurrence of technical terms in the literature and the importance weight in the technical field knowledge base. The calculation formula is as follows: W = log(1 + f) * w1; In the formula, W represents the semantic weight, w1 represents the importance weight, and f represents the frequency of occurrence of technical terms in patent literature; Comprehensively process the consistency score and the semantic weight to calculate the accuracy coefficient of technical terms. The calculation formula is as follows: In the formula, TTAC represents the accuracy coefficient of technical terms, λ represents the adjustment parameter, F represents the consistency score, and max(W) represents the maximum semantic weight in the set of technical terms.

4. The method for patent classification retrieval based on multi-source data fusion according to claim 1, wherein The process of obtaining the standard term accuracy coefficient is as follows: Structurize the standard term definitions, domain backgrounds, synonyms, and hyponymy relationships in the knowledge base into a graph; The graph nodes represent terms, and the graph edges represent the semantic relationships between terms; Normalize the frequency of standard terms in the literature in the corresponding technical field to obtain the domain weight parameter; Sum up all the domain weight parameters in the patent literature to obtain the connection degree G of the semantic nodes; Calculate the semantic similarity between the standard term and its hyponym and hypernym nodes. The calculation formula is as follows: Wherein, T a represents a standard term, T b represents the upper and lower nodes related to T a , sin represents the similarity between standard term vectors, N(T a ) represents the set of all direct upper and lower nodes of T a , |N(T a )| represents the number of upper and lower nodes of T a , T b belongs to an element in this set, and S represents the semantic similarity between the standard term and its upper and lower nodes; Calculate the ratio of the connection degree G of the semantic node to the semantic similarity S to obtain the standard term accuracy coefficient, denoted as TTAB.

5. The method for patent classification retrieval based on multi-source data fusion according to claim 1, wherein, The establishment of the semantic accurate recognition mechanism in the classification and retrieval process specifically includes: Normalize the semantic accuracy coefficient of the patent literature and the accuracy coefficient of patent literature classification and retrieval, calculate the comprehensive recognition score, and divide the semantic accuracy of patent classification and retrieval according to the comprehensive recognition score. Divide the semantics of literature classification and retrieval into qualified semantics and unqualified semantics.

6. A patent classification retrieval system based on multi-source data fusion, characterized in that, For the patent classification and retrieval method based on multi-source data fusion according to any one of claims 1-5, it includes: A domain knowledge base construction module, which constructs a domain knowledge base related to the technology field of patents, including standard terms and technical background information, to provide an accurate semantic benchmark for patent texts and ensure semantic consistency in classification and retrieval; A domain knowledge base update and term standardization module, which updates the domain knowledge base in real time, dynamically adjusts the definitions and applications of terms according to different technical fields, analyzes the matching degree between the terms in the patent literature and the standard terms in the knowledge base, and evaluates the accuracy of patent classification and retrieval; A term semantic analysis and evaluation module, which evaluates the semantic accuracy of technical terms by intelligently analyzing the technical terms and their contexts in the patent literature; A semantic accurate recognition and classification retrieval module, which establishes an effective semantic accurate recognition mechanism based on accurate patent classification and retrieval, combines the technical context and domain knowledge of the patent text, and improves the accuracy of patent text classification and retrieval.

Citation Information

Patent Citations

  • Semantic representation method for patent text vectors

    CN104199809A

  • Construction method and system of Chinese patent key information corpus and computer equipment

    CN115617989A