Big data-based patent retrieval method, device and equipment, and storage medium

By expanding keywords and constructing a relevance score matrix, a set of highly relevant patents is identified and technology trend prediction is performed. This solves the problem of insufficient relevance and trend capture in existing patent search methods, and achieves more accurate and comprehensive big data patent search.

CN119537564BActive Publication Date: 2025-12-19SHENZHEN HUAXIA TAIHE INTELLECTUAL PROPERTY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411569066.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-05
Publication Date
2025-12-19
Estimated Expiration
2044-11-05

AI Technical Summary

Technical Problem

Existing patent search methods rely on keyword matching and classification number retrieval, which cannot effectively capture the correlation between patents and technological development trends. They are also easily affected by synonyms and polysemous words, resulting in insufficient comprehensiveness and accuracy in the search process.

Method used

By acquiring and expanding the keywords input by users, a relevance score matrix is ​​constructed to identify a set of highly relevant patents. In conjunction with the application date time series, technology trend prediction is performed to generate a set of search results.

Benefits of technology

It improves the comprehensiveness and accuracy of patent searches, can identify potential technological intersections and innovations, provides insights into future technological developments, and helps users grasp industry trends and market demands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119537564B_ABST
    Figure CN119537564B_ABST
Patent Text Reader

Abstract

The application relates to a patent retrieval method and device based on big data, equipment and a storage medium, and relates to the technical field of computers.The method comprises the following steps: obtaining and expanding a keyword input by a user to obtain an expanded keyword set, determining a preliminary retrieval result set according to the keyword set, determining a relevance score matrix according to a text vector representation of each patent in the preliminary retrieval result set, determining a high-relevance patent set according to the relevance score matrix and the application date time sequence of each patent in the high-relevance patent set, determining a technology trend prediction result of the high-relevance patent set according to the high-relevance patent set, the corresponding application date time sequence and the relevance score matrix, and generating a retrieval result set for the user according to the high-relevance patent set, the corresponding technology trend prediction result and the relevance score matrix.The application can improve the accuracy and comprehensiveness of retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of big data, and particularly relates to a patent retrieval method and device based on big data, equipment and a storage medium. BACKGROUND

[0002] At present, the existing patent retrieval method mainly depends on keyword matching and classification number retrieval, and a user finds relevant patents by inputting specific terms.

[0003] However, the existing retrieval method cannot effectively capture the relevance between patents and the development trend of technology, keyword retrieval is easily affected by synonyms and polysemous words, and the user may miss important relevant patents. Meanwhile, the existing retrieval method lacks the use of big data analysis, so that the retrieval process fails to fully tap potential technological innovation and market demand, reducing the comprehensiveness and accuracy of retrieval. SUMMARY

[0004] The present application aims at the technical problems in the prior art, and provides a patent retrieval method and device based on big data, equipment and a storage medium, which can improve the comprehensiveness and accuracy of patent retrieval.

[0005] In a first aspect, the present application provides a patent retrieval method based on big data, which comprises:

[0006] Obtaining and expanding the keywords input by a user to obtain an expanded keyword set;

[0007] Determining a preliminary retrieval result set according to the keyword set;

[0008] Determining a relevance score matrix according to the text vector representation of each patent in the preliminary retrieval result set; the elements in the relevance score matrix represent the relevance between each patent and other patents;

[0009] Determining a high-relevance patent set and the application date time sequence of each patent in the high-relevance patent set according to the relevance score matrix;

[0010] Determining the technology trend prediction result of the high-relevance patent set according to the high-relevance patent set and the corresponding application date time sequence and the relevance score matrix;

[0011] Generating a retrieval result set for the user according to the high-relevance patent set and the corresponding technology trend prediction result and the relevance score matrix.

[0012] In a second aspect, the present application further provides a patent retrieval device based on big data, which comprises:

[0013] a vocabulary expansion module configured to obtain and expand the keyword input by the user to obtain an expanded keyword set;

[0014] a preliminary search module configured to determine a preliminary search result set based on the keyword set;

[0015] a relevance construction module configured to determine a relevance score matrix based on the text vector representation of each patent in the preliminary search result set; the elements in the relevance score matrix represent the relevance between each patent and other patents;

[0016] a relevance screening module configured to determine a high-relevance patent set and the application date time sequence of each patent in the high-relevance patent set based on the relevance score matrix;

[0017] a technology trend module configured to determine a technology trend prediction result of the high-relevance patent set based on the high-relevance patent set, the corresponding application date time sequence, and the relevance score matrix;

[0018] a target search module configured to generate a search result set for the user based on the high-relevance patent set, the corresponding technology trend prediction result, and the relevance score matrix.

[0019] In a third aspect, the present application further provides an electronic device, comprising: a memory configured to store a computer software program; and a processor configured to read and execute the computer software program, thereby implementing the above-described patent search method based on big data.

[0020] In a fourth aspect, the present application further provides a non-transitory computer readable storage medium, wherein the storage medium stores the computer software program, and the computer software program is executed by a processor to implement the above-described patent search method based on big data.

[0021] Compared with the prior art, the above technical solutions provided by the embodiments of the present application have the following advantages:

[0022] The present application comprehensively considers the citation relationship and content similarity of patents, making the relevance score between patents more comprehensive and enabling the identification of potential technical intersections and innovation points. Combined with the application time sequence of high-relevance patents and the adjusted trend prediction, it can provide insights into future technology development, helping users grasp industry dynamics and market demand. Through big data analysis capabilities, the patent search results can be updated and adjusted in real time to adapt to the rapidly changing market and technology environment, ensuring that users obtain the latest technical information and improving the comprehensiveness and accuracy of patent search. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1A flow chart of a patent retrieval method based on big data provided in the present application;

[0024] Figure 2 A structural schematic diagram of a patent retrieval device based on big data provided in the present application;

[0025] Figure 3 A hardware structural schematic diagram of a possible electronic device provided in the present application. DETAILED DESCRIPTION

[0026] The technical solutions in the embodiments of the present application will be clearly and completely described in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0027] In the description of the present application, the terms "first", "second" are used only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.

[0028] In the description of the present application, the term "for example" is used to indicate "as an example, illustration or description". Any embodiment described as "for example" in the present application is not necessarily interpreted as more preferred or more advantageous than other embodiments. In order to enable any person skilled in the art to implement and use the present application, the following description is given. In the following description, details are listed for the purpose of explanation. It should be understood that those skilled in the art can realize the present application without using these specific details. In other examples, well-known structures and processes will not be described in detail to avoid unnecessary details making the description of the present application obscure. Therefore, the present application is not intended to be limited to the shown embodiments, but is consistent with the broadest scope of principles and features disclosed in the present application.

[0029] Please refer to Figure 1 , a flow chart of a patent retrieval method based on big data provided in the present application is provided, including the following steps:

[0030] Step 101, obtaining and expanding the keywords input by the user to obtain an expanded keyword set.

[0031] The user inputted keyword refers to the initial search terms provided by the user when conducting a patent search. These terms are usually the technical fields, products, problems or other related concepts that the user is interested in. Assuming that the user inputted keyword is "smart home", it is a broad term involving multiple technical fields.

[0032] The main purpose of keyword expansion is to capture all possible terms and concepts that are more relevant to the user's intention, so as to improve the comprehensiveness and accuracy of the search. In some embodiments, synonyms of the user inputted keyword can be found. For example, synonyms of "smart home" can include "smart home", "smart home appliance", etc. Terms related to the user's keyword can be introduced, which can come from technical literature, industry standards or professional vocabulary. For example, terms related to "smart home" include "Internet of Things", "automation system", etc. The context of the user inputted keyword can be analyzed, and potential associated keywords can be identified using natural language processing techniques. In this way, other relevant terms can be intelligently recommended based on the background of the user's query.

[0033] After the expansion process, a more comprehensive keyword set is formed, which includes the user's initial input and all related expanded terms. For example, for the input "smart home", the expanded keyword set can include "smart home", "smart home", "smart home appliance", etc.

[0034] Step 102: Determine the preliminary search result set based on the keyword set.

[0035] The expanded keyword set can include the user inputted keyword and its synonyms, related terms, etc. This set is the basis of the search process. Assuming that the expanded keyword set is "smart home", "Internet of Things", "home automation", etc.

[0036] In some embodiments, the system uses the keyword set to search for patent documents, usually through the following ways: keyword matching: find all patent documents containing keywords in the set in the patent database. It can include the title, abstract, body, etc. of the patent. Classification number retrieval: according to the mapping of keywords to related international patent classification (IPC) or other classification systems, to obtain related patents.

[0037] In some embodiments, the search system can match the keywords to the patent documents that meet the conditions, and collect these documents as the preliminary search result set. After obtaining the preliminary results, the system can perform deduplication processing to ensure that each patent document appears only once. At the same time, the results are preliminarily sorted according to the matching degree or other standards.

[0038] The final set of preliminary search results contains patent documents highly relevant to the user's input keywords, and can include detailed information of multiple patents. For example, for the keyword set "smart home", the preliminary search result set can contain multiple patents related to this topic, such as: Patent A: technology related to smart home control system. Patent B: application of Internet of Things in home automation. Patent C: patent describing smart home appliance interconnection.

[0039] In summary, by expanding the keyword set, the preliminary search result set can more comprehensively cover the technical field of interest to the user, improving relevance. The preliminary result set provides basic data for subsequent content similarity calculation and relevance analysis, ensuring the effectiveness and accuracy of subsequent steps. It can effectively convert the user's needs into actual patent information, laying a good foundation for further analysis and optimization.

[0040] Step 103, determining a relevance score matrix according to the text vector representation of each patent in the preliminary search result set.

[0041] Wherein, the elements in the relevance score matrix represent the relevance between each patent and other patents.

[0042] In some embodiments, step 103 can include:

[0043] According to the text vector representation of each patent, calculate the content similarity between each patent and other patents, and obtain the content similarity score between each patent and other patents;

[0044] According to the citation relationship between each two patents, obtain the citation relationship score between each patent and other patents;

[0045] According to the content similarity score and the citation relationship score of each patent, determine the relevance score between each patent and other patents;

[0046] According to each of the relevance scores, construct the relevance score matrix.

[0047] In some embodiments, the content similarity score can be determined according to the following way:

[0048] Obtain the norm of the first vector representation of the patent;

[0049] Obtain the norm of the second vector representation of the other patent;

[0050] According to the first product of the first vector and the second vector, and the second product of the norm of the first vector representation and the norm of the second vector representation, determine the content similarity score between each patent and other patents.

[0051] In some embodiments, the content similarity score can be represented as:

[0052]

[0053] wherein, is the content similarity score between the i-th patent and the j-th patent, V i is the first vector representation corresponding to the i-th patent, V j is the second vector representation corresponding to the j-th patent, ||V i ||, ||V j || are the norms of the first and second vector representations, respectively.

[0054] In a specific implementation, The higher the score, the more similar the content of the two patents. V i and V j are usually generated by natural language processing techniques (such as TF-IDF, Word2Vec, or BERT, etc.), which can capture semantic information in patent documents. V i ·V j is the dot product of the two vectors, which reflects their similarity in semantic space. If the two vectors point in the same direction (content similarity), the value of the dot product will be larger. Normalization is performed using the norms, ||V i ||, ||V j The product of the norms can eliminate the influence of different vector lengths, ensuring that the similarity score only reflects the similarity in content, not the difference in vector length.

[0055] The value of is usually between -1 and 1, with a score of 1 indicating that the two patents are identical (pointing in the same direction in vector space). A score of 0 indicates that the two patents have no similarity (orthogonal vectors). A score of -1 indicates that the two patents are completely opposite in content. By calculating the content similarity score, patents with similar technical content can be effectively identified, providing important evidence for subsequent relevance analysis. Patents with high similarity scores can help users find related technical fields and innovation points, improving the accuracy and effectiveness of patent retrieval. During the retrieval process, the system can recommend other patents with similar content to the user's interest through the content similarity score, thereby expanding the user's perspective and promoting in-depth understanding and research of technology.

[0056] In summary, the content similarity score calculation method of the present application can effectively quantify the similarity between patents, providing important data support for subsequent relevance analysis and retrieval result optimization.

[0057] In some embodiments, the citation relationship score can be represented as:

[0058]

[0059] wherein C ij is the patent citation data, count(C ij ) represents the number of times the i-th patent cites the j-th patent.

[0060] In specific implementations, represents the score of the i-th patent citing the j-th patent, reflecting the strength of the citation relationship between the two patents. C ij Generally includes the number of patents, citation time, citation type and other information. count(C ij ) value is higher, indicating that the i-th patent has greater dependence and citation strength on the j-th patent.

[0061] If indicates that there is a technical association between the two. If indicates that the direct relationship between the two is weak or does not exist. By calculating the citation relationship score, patents that are technically dependent on each other can be effectively identified, which is of great significance for understanding the context of technological development and the evolution of innovation. Patents with high citation scores are generally considered to have more influence in a certain technical field, which can help users quickly identify key and representative patents.

[0062] In the patent retrieval process, the citation relationship score can provide users with information about the technology innovation chain, helping them identify important patents and their technical background in related fields, and enhance their understanding of technological development.

[0063] Through this process, the calculation of the citation relationship score can provide important data support for subsequent relevance analysis and patent recommendation, improving the comprehensiveness and accuracy of patent retrieval.

[0064] In some embodiments, the relevance score of each said patent to other patents is represented as:

[0065]

[0066] wherein R ij is the relevance score of the i-th patent to the j-th patent, is the content similarity score of the i-th patent to the i-th patent, is the citation relationship score of the i-th patent to the i-th patent, and α is the first weight and β is the second weight.

[0067] In a specific implementation, the settings of the first weight and the second weight can be based on domain knowledge, user demand, or derived through historical data analysis. Depending on the specific search target and technical background, users or researchers can adjust these two parameters to better reflect the actual situation of technical association.

[0068] High score R ij A high score R indicates a strong association between the i-th patent and the j-th patent, which means technical complementarity, close citation relationship, or high content similarity. A low score indicates a weak association between the two, meaning they have no direct relationship in terms of technical content or application field.

[0069] By combining content similarity and citation relationship, the association score provides a comprehensive perspective to understand the technical connection between patents. Patents with high association scores can be recommended first in subsequent search results, helping users quickly find relevant technical innovations and solutions. In patent search and analysis, the association score can provide users with the context of technical development and the evolution of innovation, supporting users to make more informed technical decisions.

[0070] In summary, the calculation of the association score in this application provides important data support for subsequent search optimization and patent recommendation, enhancing the accuracy and effectiveness of patent search.

[0071] Step 104, according to the association score matrix, determine a high association patent set, and a time sequence of application dates of each patent in the high association patent set.

[0072] Among them, the time sequence of application dates of each patent in the high association patent set refers to the collection of all application dates related to a specific patent, usually arranged in chronological order. This sequence reflects the application activities and development history of patents in a certain technical field.

[0073] In some embodiments, step 104 can include:

[0074] Sorting all association scores in the association score matrix to obtain a corresponding first sorting result;

[0075] According to the first sorting result and a first threshold, determine a target association score;

[0076] Determine all patents corresponding to the target association score as the high association patent set.

[0077] Among them, the association score matrix is a two-dimensional matrix, where each element R ij represents the association score between the i-th patent and the j-th patent. Each row and each column in the matrix corresponds to a different patent.

[0078] In some embodiments, all scores in the matrix can be extracted to form a one-dimensional list. This list can be sorted according to the values of the scores, from high to low. This process can use standard sorting algorithms (such as quicksort, mergesort, etc.). After sorting, a new list can be obtained, called the first sorting result. The elements in this list are arranged in descending order of relevance scores.

[0079] In some embodiments, a first threshold value can be set according to the actual application requirements or historical data analysis. This threshold value is used to filter out patents with high relevance scores, usually determined based on experience or data analysis. The relevance scores in the first sorting result can be compared with the threshold value. All patents with scores higher than or equal to the threshold value are considered target patents. According to the target relevance score, all patents corresponding to the target score are determined. These patents constitute the high-relevance patent set.

[0080] For example, the first sorting result is [0.95, 0.85, 0.80, 0.75, 0.70], and the first threshold value is set to 0.8, then the high-relevance patent set will include patents with scores of 0.95, 0.85 and 0.80.

[0081] Through this process, users can quickly filter out patents highly relevant to their query or research topic, improving the accuracy of retrieval. The high-relevance patent set provides important basic data for subsequent technical analysis, market evaluation and R&D decision-making, helping users grasp technology trends and market opportunities. In the fields of patent retrieval, technical analysis, competitor research, etc., this process can help researchers and enterprises effectively identify key and representative patents, providing support for technology innovation and market strategy.

[0082] In summary, the sorting of the relevance score matrix and the determination of the high-relevance patent set in this application provide users with more accurate and relevant patent information, enhancing the practicality and effectiveness of patent retrieval.

[0083] Step 105, according to the high-relevance patent set and its corresponding application date time sequence, and the relevance score matrix, determine the technology trend prediction result of the high-relevance patent set.

[0084] In some embodiments, step 105 can include:

[0085] According to the relevance score matrix and the high-relevance patent set, construct a weight function for determining the technology trend prediction result;

[0086] According to the application date time sequence and the weight function, determine the technology trend prediction result of the high-relevance patent set.

[0087] In some embodiments, the technology trend prediction results of highly related patent sets can be expressed as:

[0088]

[0089] in, It is a technology trend forecast result, R a It is a highly relevant patent set, T is a highly relevant patent set R a The patent application time series, f(A,R) a ) is the weight function.

[0090] In the specific implementation, This is a technology trend forecast. The result reflects technological development trends based on a highly correlated patent set. T is the patent application time series of the highly correlated patent set, including the application dates of all related patents. This series is used to capture the temporal characteristics of technological development. R a This is a highly relevant patent set, meaning a collection of patents highly relevant to user needs selected through the previous steps. f(A,R) a () is a weighting function used to adjust and weight technology trends. It dynamically adjusts the impact of technology trends based on the characteristics of the correlation score matrix and the highly correlated patent set.

[0091] In some embodiments, the contribution of each patent to a technological trend can be determined by analyzing the content similarity and citation relationships among patents in a highly correlated patent set. The weighting function can be linear or non-linear, depending on the specific application requirements.

[0092] Specifically, the application dates of all patents in the highly correlated patent set are first extracted to form an application time series. A weighting function can then be applied to the application time series to adjust the technology trend prediction results. This function involves statistical analysis of the time series, weighted averaging, and other processing methods.

[0093] It provides predictions about the future development of specific technology fields, helping users understand potential trends in technological evolution. Technology trend forecasts can provide a basis for corporate R&D strategies, market layout, and investment decisions, enhancing the foresight of future technological developments. In fields such as technology research, market analysis, and intellectual property management, technology trend forecasts can help users grasp industry dynamics, identify technological opportunities, and support more effective decision-making.

[0094] In summary, the technology trend prediction results of the highly relevant patent set in this application can provide users with in-depth technology insights and trend analysis, enhancing the effectiveness of patent search and technological innovation.

[0095] Step 106: Based on the high-correlation patent set and its corresponding technology trend prediction results, as well as the correlation score matrix, a search result set is generated for the user.

[0096] In some embodiments, step 106 can include:

[0097] According to the correlation scores of each patent in the high-correlation patent set, the patents are sorted to obtain a corresponding second sorting result;

[0098] According to the inverse correlation score matrix and the technology trend prediction results, the second sorting result is adjusted to obtain a corresponding third sorting result;

[0099] According to the third sorting result and a second threshold, a plurality of target patents are determined from the high-correlation patent set to obtain the search result set including the plurality of target patents.

[0100] In some embodiments, all patents in the high-correlation patent set can be extracted into a list according to the correlation scores of each patent. The list is sorted from high to low according to the scores to form a second sorting result. This process can use standard sorting algorithms.

[0101] It can be understood that the inverse correlation score matrix is a reverse processing of the original correlation score matrix. It is obtained by calculation or other mathematical methods, and reflects the inverse correlation between patents. The values in the inverse correlation score matrix and the technology trend prediction results can be combined to dynamically adjust the scores of the patents. The specific method can be through weighted average or direct multiplication, etc. to ensure that the sorting result can reflect the current technology trend and inverse correlation.

[0102] In some embodiments, a second threshold can be set according to specific application requirements or historical data analysis. This threshold is used to filter patents with higher scores. The scores of patents in the third sorting result can be compared with the second threshold. All patents with scores higher than or equal to the threshold are considered target patents. According to the target patents, a search result set is determined. This result set will include all target patents that meet the conditions and is provided to the user as the final search result.

[0103] Through this process, the user can quickly identify patents that are highly relevant to their research or business needs, improving the accuracy and effectiveness of the search. The search result set provides important basic data for subsequent technology analysis, market evaluation, and R&D decision-making, helping users grasp technology trends and market opportunities.

[0104] In the fields of patent search, technology analysis, and competitor research, this process can help researchers and enterprises effectively identify key and representative patents, supporting technology innovation and market strategy.

[0105] In summary, the present application can help researchers and enterprises effectively identify key and representative patents in patent retrieval, technical analysis, and competitor research, supporting technology innovation and market strategy.

[0106] Please refer to Figure 2 , Figure 2 The structure diagram of a patent retrieval device based on big data provided by the present application is shown.

[0107] As Figure 2 shown, the patent retrieval device based on big data provided by the present application includes:

[0108] The vocabulary expansion module 201 is configured to obtain and expand the keywords input by the user to obtain an expanded keyword set;

[0109] The preliminary retrieval module 202 is configured to determine a preliminary retrieval result set according to the keyword set;

[0110] The association construction module 203 is configured to determine an association score matrix according to the text vector representation of each patent in the preliminary retrieval result set; the elements in the association score matrix represent the association between each patent and other patents;

[0111] The association screening module 204 is configured to determine a high-association patent set and the application date time sequence of each patent in the high-association patent set according to the association score matrix;

[0112] The technology trend module 205 is configured to determine a technology trend prediction result of the high-association patent set according to the high-association patent set, the corresponding application date time sequence, and the association score matrix;

[0113] The target retrieval module 206 is configured to generate a retrieval result set for the user according to the high-association patent set, the corresponding technology trend prediction result, and the association score matrix.

[0114] Please refer to Figure 3 , Figure 3 The embodiment of the electronic device provided by the present application is shown. As Figure 3 shown, the present application provides an electronic device, which includes a processor 301, a communication interface 302, a memory 303, and a communication bus 304, wherein the processor 301, the communication interface 302, and the memory 303 communicate with each other through the communication bus 304,

[0115] The memory 303 is configured to store a computer program;

[0116] In an embodiment of the present application, the processor 301, when executing the program stored in the memory 303, implements the big data-based patent retrieval method provided by any one of the foregoing method embodiments, including:

[0117] obtaining and expanding the keyword input by the user to obtain an expanded keyword set;

[0118] determining a preliminary retrieval result set according to the keyword set;

[0119] determining a relevance score matrix according to the text vector representation of each patent in the preliminary retrieval result set; an element in the relevance score matrix represents the relevance between each patent and other patents;

[0120] determining a high-relevance patent set and the application date time sequence of each patent in the high-relevance patent set according to the relevance score matrix;

[0121] determining a technology trend prediction result of the high-relevance patent set according to the high-relevance patent set, the corresponding application date time sequence, and the relevance score matrix;

[0122] generating a retrieval result set for the user according to the high-relevance patent set, the corresponding technology trend prediction result, and the relevance score matrix.

[0123] The embodiment of the present application also provides a computer readable storage medium, which has a computer program stored thereon, and the computer program is executed by a processor to implement the following steps:

[0124] obtaining and expanding the keyword input by the user to obtain an expanded keyword set;

[0125] determining a preliminary retrieval result set according to the keyword set;

[0126] determining a relevance score matrix according to the text vector representation of each patent in the preliminary retrieval result set; an element in the relevance score matrix represents the relevance between each patent and other patents;

[0127] determining a high-relevance patent set and the application date time sequence of each patent in the high-relevance patent set according to the relevance score matrix;

[0128] determining a technology trend prediction result of the high-relevance patent set according to the high-relevance patent set, the corresponding application date time sequence, and the relevance score matrix;

[0129] According to the high-correlation patent set and the corresponding technical trend prediction result thereof, and the correlation score matrix, a search result set is generated for the user.

[0130] It should be noted that in the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0131] Those skilled in the art understand that the embodiments of the present application can be provided as a method, device, or computer program product. Therefore, the present application can adopt a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0132] The present application is described with reference to flowcharts and / or block diagrams according to the method, device (apparatus), and computer program product of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in one or more blocks.

[0133] These computer program instructions can also be stored in a computer readable storage medium that can direct the computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer readable storage medium produce a product including instruction apparatus, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in one or more blocks.

[0134] These computer program instructions can also be loaded into a computer or other programmable data processing device, so that a series of operation steps are performed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in one or more blocks.

[0135] While the preferred embodiments of the application have been described, additional variations and modifications can be made to these embodiments by those skilled in the art once they have the benefit of the present disclosure without departing from the spirit and scope of the application. Accordingly, it is intended that the appended claims include all such modifications and variations as fall within the scope of the present application.

[0136] Obviously, many modifications and variations of the present application are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.

Claims

1. A patent retrieval method based on big data, characterized in that, The method includes: Obtain and expand the keywords input by the user to obtain an expanded set of keywords; Based on the aforementioned keyword set, a preliminary search result set is determined; Based on the text vector representation of each patent in the preliminary search result set, a relevance score matrix is ​​determined; the elements in the relevance score matrix represent the relevance between each patent and other patents. Based on the correlation score matrix, a set of highly correlated patents is determined, along with the application date time sequence of each patent in the set of highly correlated patents. Based on the highly correlated patent set and its corresponding application date time series, as well as the correlation score matrix, the technology trend prediction result of the highly correlated patent set is determined, including: constructing a weight function for determining the technology trend prediction result based on the correlation score matrix and the highly correlated patent set, and determining the technology trend prediction result of the highly correlated patent set based on the application date time series and the weight function; Based on the highly relevant patent set and its corresponding technology trend prediction results, as well as the relevance score matrix, a search result set is generated for the user, including: sorting each patent according to the relevance score of each patent in the highly relevant patent set to obtain a corresponding second sorting result; adjusting the second sorting result according to the inverse relevance score matrix and the technology trend prediction results to obtain a corresponding third sorting result; and determining multiple target patents from the highly relevant patent set according to the third sorting result and a second threshold to obtain the search result set including the multiple target patents. The step of determining the relevance score matrix based on the text vector representation of each patent in the preliminary search result set includes: Based on the text vector representation of each patent, the content similarity of each patent with other patents is calculated to obtain the content similarity score of each patent with other patents; Based on the citation relationship between each pair of patents, a citation relationship score for each patent with other patents is obtained; Based on the content similarity score and citation relationship score of each patent, determine the relevance score of each patent to other patents; Construct the correlation score matrix based on each correlation score; The citation relationship score is represented as follows: in, It is patent citation data. This represents the number of times the i-th patent cites the j-th patent.

2. The patent retrieval method based on big data according to claim 1, characterized in that, The step of calculating the content similarity between each patent and other patents based on the text vector representation of each patent to obtain a content similarity score for each patent and other patents includes: Obtain the modulus of the first vector representation of the patent; Obtain the modulus of the second vector representation of other patents; Based on the first product of the first vector and the second vector, and the second product of the modulus represented by the first vector and the modulus represented by the second vector, a content similarity score for each patent with other patents is determined. The content similarity score is represented as follows: in, It is the content similarity score between the i-th patent and the j-th patent. It is the first vector representation corresponding to the i-th patent. It is the second vector representation corresponding to the j-th patent. These are the modulus represented by the first vector and the modulus represented by the second vector, respectively.

3. The patent retrieval method based on big data according to claim 2, characterized in that, The relevance score of each patent to other patents is expressed as follows: in, It is the correlation score between the i-th patent and the j-th patent. It is the content similarity score between the i-th patent and the j-th patent. It is the score of the citation relationship between the i-th patent and the j-th patent. It is the first weight. It is the second weight.

4. The patent retrieval method based on big data according to claim 3, characterized in that, The step of determining the highly relevant patent set based on the correlation score matrix includes: Sort all correlation scores in the correlation score matrix to obtain the corresponding first sorting result; Based on the first ranking result and the first threshold, the target relevance score is determined; All patents corresponding to the target relevance score are identified as the highly relevant patent set.

5. A patent search device based on big data, characterized in that, The device includes: The vocabulary expansion module is used to obtain and expand the keywords input by the user, resulting in an expanded set of keywords; The preliminary search module is used to determine a preliminary search result set based on the keyword set; The association construction module is used to determine an association score matrix based on the text vector representation of each patent in the preliminary search result set; the elements in the association score matrix represent the association between each patent and other patents; The association filtering module is used to determine the highly associated patent set and the application date time sequence of each patent in the highly associated patent set based on the association score matrix. The technology trend module is used to determine the technology trend prediction result of the highly correlated patent set based on the highly correlated patent set and its corresponding application date time series, as well as the correlation score matrix. The module includes: constructing a weight function for determining the technology trend prediction result based on the correlation score matrix and the highly correlated patent set; and determining the technology trend prediction result of the highly correlated patent set based on the application date time series and the weight function. The target retrieval module is used to generate a retrieval result set for the user based on the highly relevant patent set and its corresponding technology trend prediction results, as well as the relevance score matrix. The module includes: sorting each patent in the highly relevant patent set according to its relevance score to obtain a corresponding second sorting result; adjusting the second sorting result according to the inverse relevance score matrix and the technology trend prediction results to obtain a corresponding third sorting result; and determining multiple target patents from the highly relevant patent set based on the third sorting result and a second threshold to obtain the retrieval result set including the multiple target patents. The step of determining the relevance score matrix based on the text vector representation of each patent in the preliminary search result set includes: Based on the text vector representation of each patent, the content similarity of each patent with other patents is calculated to obtain the content similarity score of each patent with other patents; Based on the citation relationship between each pair of patents, a citation relationship score for each patent with other patents is obtained; Based on the content similarity score and citation relationship score of each patent, determine the relevance score of each patent to other patents; Construct the correlation score matrix based on each correlation score; The citation relationship score is represented as follows: in, It is patent citation data. This represents the number of times the i-th patent cites the j-th patent.

6. An electronic device, characterized in that, include: Memory, used to store computer software programs; A processor is configured to read and execute the computer software program to implement a patent retrieval method based on big data as described in any one of claims 1 to 4.

7. A non-transitory computer-readable storage medium, wherein the storage medium stores the computer software program, characterized in that, When the computer software program is executed by the processor, it implements a patent retrieval method based on big data as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Patent introduction predicted value calculation method

    CN103208038A

  • Patent retrieval method, device and system, computer equipment and storage medium

    CN111291159A