Generation method of patent text vectorization large model for patent relevancy calculation

By segmenting and training the patents to identify invention points, a large-scale vectorized model of the patent text, encompassing both overall and partial invention points, is generated. This solves the problem of local relevance in patent relevance calculation, achieving higher computational accuracy and innovation efficiency.

CN120910243APending Publication Date: 2025-11-07TRS INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511076384.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In existing patent relevance calculations, locally related patents are prone to ranking bias, and a single vector is insufficient to capture the correlation of detailed technical features, resulting in insufficient accuracy in relevance calculations.

Method used

By segmenting patents according to their inventive points, generating overall and partial inventive points, and combining patent bias, a general text vectorization model is trained to improve vector representation capabilities and relevance calculation accuracy.

Benefits of technology

It improves the accuracy of patent relevance calculation, reduces legal risks and R&D costs for enterprises, and enhances innovation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910243A_ABST
    Figure CN120910243A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of patent information processing, and provides a method for generating a patent text vectorization large model for patent relevancy calculation, which comprises the following steps of: cutting each training patent text, and extracting an overall invention point and a local invention point of each training patent text; and vectorizing the overall invention points and the local invention points of each training patent text through a general text vectorization large model to obtain an overall feature vector and a local feature vector of each training patent text, and calculating the patent relevancy of the training patent pairs. And based on the patent relevancy of the training patent pair, feeding back and adjusting parameters of the general text vectorization large model until a vectorization result output by the general text vectorization large model reaches a set expectation, obtaining a patent text vectorization large model, vectorizing the overall invention points and the local invention points through the patent text vectorization large model, and obtaining a patent text vectorization result. The patent text vectorization quality is improved, and the accuracy of patent relevancy calculation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of patent information processing, and particularly relates to a generation method of a patent text vectorization large model for patent relevance calculation. BACKGROUND

[0002] In the field of patent information processing, patent relevance calculation as a core technology has been widely applied in patent retrieval, infringement judgment, layout navigation and value assessment scenarios: in patent retrieval, the semantic engine technology is used to optimize the result display through relevance sorting, and high-relevance patents are preferentially presented, which effectively improves the retrieval efficiency of the overall relevant (more than 70%) patents; in infringement judgment, the text and technical features of the analyzed technical scheme and the existing patents are compared through the relevance, which provides a technical reference for infringement boundary judgment; in patent layout and navigation, the relevant degree distribution can be used to identify technical blank areas and dense areas, and to assist in formulating accurate R&D strategies; in value assessment, the relevance can reflect the positioning of the target patent in the technical system, and provide a basis for judging its market value and technical influence.

[0003] However, the current patent relevance calculation still has significant limitations: on the one hand, for locally relevant patents, the overall similarity is low, which easily causes ranking deviation and leads to missed detection, seriously affecting the credibility of the retrieval results; on the other hand, the existing technology mostly uses a single vector to represent the patent text, and a piece of invention often contains multiple invention points, so when two patents have local relevance, the single vector cannot capture the relevance of the subdivided technical features, resulting in insufficient accuracy of the relevance calculation. At the same time, the reliability of the relevance result is highly dependent on the quality of the vector, which further affects the accuracy of subsequent decisions such as patent inventiveness judgment based on relevance. Therefore, how to train a high-quality patent text vectorization large model to improve the vector representation ability and the accuracy of relevance calculation has become a key problem to break through the current technical bottleneck.

[0004] Based on this, the present application provides a generation method of a patent text vectorization large model for patent relevance calculation. SUMMARY

[0005] To solve the technical problem that the existing patent relevance calculation cannot accurately calculate the patent relevance, especially the local relevance, due to the low quality of the text vectorization, the present application provides a generation method of a patent text vectorization large model for patent relevance calculation, which cuts the patent according to the invention points to generate overall invention points and local invention points, and then trains a general text vectorization large model in combination with the corresponding bias of the patent, so as to improve the quality of the patent text vectorization and finally improve the accuracy of the patent relevance calculation.

[0006] The specification provides a generation method of a patent text vectorization large model for patent relevance calculation, comprising: S1: Assemble training patent pairs and extract overall and local invention points: Assemble training patent pairs, a set of the training patent pairs includes two training patent texts, cut each training patent text to obtain the overall invention point and the local invention point corresponding to each training patent text; S2: Invention point vectorization: Based on the general text vectorization large model, the overall invention point and the local invention point of each training patent text are vectorized to obtain the overall feature vector and the local feature vector of each training patent text; S3: Calculate the patent relevance: According to each overall feature vector and each local feature vector, the patent relevance of the training patent pair is calculated; S4: Adjust the general text vectorization large model based on the patent relevance: Based on the patent relevance of the training patent pair, adjust the parameters of the general text vectorization large model until the vectorization result output by the general text vectorization large model reaches the set expectation, and obtain a patent text vectorization large model.

[0007] Optionally, the patent text vectorization large model obtained in S4 specifically comprises: Redetermine the training patent pair, and repeat steps S2-S4 until the cycle number of feedback adjustment reaches a preset threshold, and the general text vectorization large model is used as a patent text vectorization large model.

[0008] Optionally, in S4, the method for adjusting the parameters of the general text vectorization large model based on the patent relevance is: According to the patent relevance and the bias of the training patent pair, the loss function of the general text vectorization large model constructed based on the patent relevance is used to calculate the loss function value, and the parameters of the general text vectorization large model are adjusted according to the loss function value; wherein, when the two training patent texts in the training patent pair are related, the bias of the training patent pair is positive; when the two training patent texts in the training patent pair are not related, the bias of the training patent pair is negative.

[0009] Optionally, the loss function in S4 is: (1) In the formula, S is the patent relevance.

[0010] Optionally, the evaluation and deployment method of the patent text vectorization large model in S4 comprises: obtaining a test patent pair, and cutting each test patent text in the test patent pair respectively to obtain an overall invention point and a local invention point corresponding to each test patent text respectively; inputting the overall invention point and the local invention point of each test patent text into the general text vectorization large model respectively to obtain an overall feature vector and a local feature vector of each test patent text; calculating a patent relevance of the test patent pair according to the overall feature vector and the local feature vector of each test patent text; determining a bias of the test patent pair, evaluating the general text vectorization large model according to the determined bias and the calculated patent relevance, and deploying the general text vectorization large model as a patent text vectorization large model after the evaluation passes.

[0011] Optionally, the general text vectorization large model is a Transformer model based on an attention mechanism.

[0012] Optionally, the patent text vectorization large model is used to calculate the patent relevance of two patents, and the specific application method is: obtaining two pieces of to-be-calculated patent texts, and cutting each to-be-calculated patent text respectively to obtain an overall invention point and a local invention point corresponding to each to-be-calculated patent text respectively; using the patent text vectorization large model deployed on the target device to vectorize the overall invention point and the local invention point of each to-be-calculated patent text respectively to obtain an overall feature vector and a local feature vector of each to-be-calculated patent text; calculating a patent relevance of the two pieces of to-be-calculated patent texts according to the overall feature vector and the local feature vector of each to-be-calculated patent text.

[0013] Optionally, the cutting of each training patent text in S1 to obtain the overall invention point and the local invention point corresponding to each training patent text specifically includes: S11: Extracting key terms: extracting key terms matching a pre-constructed patent knowledge graph from each training patent text based on the patent knowledge graph; S12: Generating an overall invention point: concatenating each training patent text and the extracted key terms of each training patent text by using a first prompt word template, and inputting into a first generative large model to obtain an overall invention point corresponding to each training patent text; S13: generating a local invention point, determining the number of times that the key term of each training patent text appears in each paragraph of the training patent text, and determining the local invention point corresponding to each training patent text according to the number of times corresponding to each paragraph and the weight corresponding to the key term of each training patent text.

[0014] Optionally, the determination of the local invention point corresponding to each training patent text according to the number of times corresponding to each paragraph and the weight corresponding to the key term of each training patent text in S13 specifically comprises: S131: initializing the processing state of each paragraph of the training patent text as a first state; S132: when the processing state of each paragraph of the training patent text is the first state, determining the importance of each paragraph of the training patent text according to the weight corresponding to the key term of the training patent text and the number of times that the key term of the training patent text appears in each paragraph of the training patent text; S133: combining adjacent paragraphs according to the importance of each paragraph of the training patent text, determining the shortest paragraph group of the training patent text, and taking the shortest paragraph group as a patent segment, wherein the generation condition of the patent segment is that the importance of the patent segment of the training patent text is greater than half of the sum of the weights of all key terms of the training patent text, and the importance of the patent segment is generated by superimposing the importance corresponding to each paragraph in the patent segment; S134: setting the processing state of the paragraph in the shortest paragraph group of the training patent text to a second state, and adjusting the weight value of the key term hit in the shortest paragraph group of the training patent text to half of the current weight; S135: repeating the steps of S132-S134, recalculating the importance corresponding to each paragraph and determining the patent segment of the current period based on the adjusted weight value of the key term in S134, until the execution times reach the set threshold or no new patent segment can be generated, and based on each patent segment of the training patent text, using a second generative large model to determine each local invention point of the training patent text.

[0015] Optionally, S3 specifically comprises: The two training patent texts in the training patent pair are taken as a first training patent text and a second training patent text respectively, and the global feature vector and the local feature vector of the first training patent text are taken as each first feature vector, and the global feature vector and the local feature vector of the second training patent text are taken as each second feature vector; combining each first feature vector with each second feature vector respectively to obtain each matching pair, and calculating the feature similarity between the first feature vector and the second feature vector in each matching pair respectively; determining the similarity weight corresponding to each matching pair according to the weight corresponding to the first feature vector and the second feature vector in the matching pair respectively; determining the maximum feature similarity in the feature similarities of the matching pairs, and determining the similarity weight corresponding to the maximum feature similarity as the maximum similarity weight; determining the matching pairs satisfying the target condition from the matching pairs as target matching pairs; the first feature vector and the second feature vector included in each target matching pair are not the same, and the sum of the feature similarities of the target matching pairs is maximum; calculating the patent relevance of the training patent pair according to the maximum feature similarity, the maximum similarity weight, the feature similarity of each target matching pair, and the similarity weight of each target matching pair.

[0016] The above at least one technical solution adopted in the specification can achieve the following beneficial effects: The present application provides a generation method of a patent text vectorization large model for patent relevance calculation, which realizes the consideration of local invention points in the calculation of patent relevance, and the improvement of the quality of text vectorization. Specifically, the overall invention points and local invention points of each training patent text in the assembled training patents can be extracted by cutting each training patent text, and the overall invention points and local invention points of each training patent text can be vectorized by a general text vectorization large model to obtain the overall feature vector and local feature vector of each training patent text. Then, the patent relevance of the training patent pair is calculated based on the overall feature vector and the local feature vector, and the parameters of the general text vectorization large model are adjusted based on the patent relevance of the training patent pair until the vectorization result output by the general text vectorization large model reaches the set expectation, and the patent text vectorization large model is obtained. Therefore, the overall invention points and local invention points extracted are vectorized by the trained patent text vectorization large model, the quality of patent text vectorization is improved, the accuracy of patent relevance calculation is improved, and the accuracy of subsequent decisions such as patent inventiveness judgment based on patent relevance is improved.

[0017] The application can adopt the determined training sample pair to train the general text vectorization large model for multiple times in a cycle until the cycle number of feedback adjustment reaches the preset threshold, and the general text vectorization large model at this time is taken as the patent text vectorization large model, so as to ensure that the general text vectorization large model can fully learn knowledge, so that the quality of the output vector in subsequent application is high. Moreover, different loss function values are determined through the set loss function for different bias, so that the loss function values of different loss functions can be calculated based on the loss function according to different bias when the general text vectorization large model is trained, and different feedbacks are given to the general text vectorization large model, so as to realize the conversion of the general large model (i.e. the general text vectorization large model) to the patent large model (i.e. the patent text vectorization large model), so that the patent text vectorization large model obtained by training is more in line with the characteristics of the patent text, and the quality of the output vector is higher.

[0018] In order to ensure the quality of the output vector of the trained patent text vectorization large model, the application can use test patents to evaluate the general text vectorization large model, and only after the evaluation is passed, the general text vectorization large model is taken as the patent text vectorization large model. Moreover, after obtaining the trained patent text vectorization large model, the patent text vectorization large model can be used to vectorize the overall invention points and local invention points of each to-be-calculated patent text respectively, to obtain the overall feature vector and local feature vector of each to-be-calculated patent text, and calculate the patent relevance of two to-be-calculated patent texts according to the overall feature vector and local feature vector, so as to improve the quality of the overall feature vector and local feature vector through the patent text vectorization large model, so that the patent relevance calculated based on the overall feature vector and local feature vector is more accurate.

[0019] When the overall invention points and local invention points are obtained, the key terms matched with the patent knowledge graph can be extracted from each training patent text through the patent knowledge graph. Since the patent knowledge graph contains a large amount of search information input when searching for patents, that is, a large amount of keywords (i.e. key terms) used when searching for patents, the key terms of each training patent text can be extracted through the patent knowledge graph, which helps to improve the accuracy of the subsequently determined overall invention points and local invention points.

[0020] When the overall invention point is obtained, each training patent text and the key term of each training patent text are spliced through the first prompt word template and input into the first generative large model to obtain the overall invention point corresponding to each training patent text, so that the first generative large model extracts the overall invention point from each training patent text through the first prompt word template and the key term, thereby improving the accuracy of determining the overall invention point, and the overall invention point is the original text in each training patent text, thereby better representing the invention point of the training patent text.

[0021] When the local invention point is obtained, the shortest paragraph group of each training patent text is calculated as a patent segment based on the weight of the key term of each training patent text and the number of times the key term appears in each paragraph of each training patent text, so as to cover the key term more and improve the accuracy of the subsequently generated local invention point, and the local invention point of each training patent text is determined based on the patent segment of each training patent text, and the local invention point is also derived from the original text in each training patent text, thereby better representing the invention point of the training patent text.

[0022] When the patent relevance is calculated, not only the feature similarity between each feature vector is considered, but also the importance of each feature vector, that is, the patent relevance of two training patent texts is determined according to the similarity weight of each matching pair and the feature similarity of each matching pair, so as to improve the accuracy of the determined patent relevance, solve the problem of local relevance between patents, improve the accuracy of patent relevance calculation, and realize the upgrade from "overall fuzzy matching" to "fine-grained feature association". This upgrade has practical application value in patent retrieval, infringement judgment, layout navigation and value evaluation in the scene of patent information retrieval, can reduce the legal risk and research and development cost of enterprises, and improve the innovation efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 A flowchart of a generation method of a patent text vectorization large model for patent relevance calculation provided in the specification; Figure 2 A schematic diagram of a method for calculating patent relevance based on a patent text vectorization large model provided in the specification; Figure 3 A schematic diagram of a training process of a patent text vectorization large model for patent relevance calculation provided in the specification; Figure 4 A schematic diagram of a process of calculating patent relevance by using the first comparison method provided in the specification; Figure 5A schematic diagram of a process for calculating patent relevance using a second control method is provided in the present specification. DETAILED DESCRIPTION

[0024] The present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0025] The present specification provides a generation method of a patent text vectorization large model for patent relevance calculation, specifically as shown in Figure 1 Figure 1 A flowchart of a generation method of a patent text vectorization large model for patent relevance calculation is provided in the present specification, specifically including the following steps: S1: Assemble training patent pairs and extract overall and local invention points: Assemble training patent pairs, a set of the training patent pairs includes two training patent texts, cut each training patent text to obtain the overall invention point and the local invention point corresponding to each training patent text.

[0026] In the present specification, the device for generating a patent text vectorization large model can assemble training patent pairs and extract overall and local invention points, that is, assemble training patent pairs, a set of training patent pairs includes two training patent texts, specifically, training patent pairs can be selected from a patent database. Then cut each training patent text to obtain the overall invention point and the local invention point corresponding to each training patent text. The device for generating a patent text vectorization large model can be a server, but also a system, or an electronic device such as a desktop computer, a notebook computer, etc. For ease of description, the following will take the server as the execution subject to explain the generation method of a patent text vectorization large model for patent relevance calculation provided in the present specification. The type of each training patent text in the above training patent pair can be invention, but also invention, which is not limited in the present specification.

[0027] The above patent database can be a pre-constructed database, which can include a large number of patent texts, and the types of the included patent texts can be invention, but also invention. The above training patent pair can be a patent pair selected from the patent database, which can include two training patent texts, which can be related or unrelated, that is, when the two training patent texts in the above training patent pair are related, the bias of the training patent pair is positive, and when the two training patent texts in the training patent pair are not related, the bias of the training patent pair is negative.

[0028] ​In the present specification, the above-mentioned training patent pairs can be multiple, that is, the server can screen a target number of training patent pairs from the patent database, and the target number can be a pre-set value, such as 6 million. The specific value is not limited in the present specification. Moreover, the number of positive training patent pairs in each of the above-mentioned training patent pairs can be a first number, and the number of negative training patent pairs can be a second number. The first number and the second number can also be pre-set values, such as 1 million for the first number and 5 million for the second number, and the sum of the first number and the second number is the target number. The ratio of the number of positive training patent pairs to the number of negative training patent pairs satisfies a first ratio, and the first ratio can be pre-set, such as 1:5. The above-mentioned positive training patent pairs are training patent pairs with a positive bias, and the above-mentioned negative training patent pairs are training patent pairs with a negative bias. The two training patent texts in each of the above-mentioned positive training patent pairs are related, and the two training patent texts in each of the above-mentioned negative training patent pairs are not the same. Moreover, each of the above-mentioned negative training pairs can be obtained by randomly combining two patents in the above-mentioned patent database, and the two patents randomly combined are not related. When screening the training patent pairs from the patent database, a pre-set algorithm or model can be used to screen the training patent pairs from the patent database.

[0029] Currently, the self-attention mechanism needs to pay attention to all other positions at each position in the attention mechanism-based model, so long text sequences will bring large computation and storage costs. Therefore, there is a certain limit to the length of the text, which is generally 512, although it can be solved by modifying the configuration and source code, but the cost performance is very low, such as greatly reducing the training efficiency. When processing long text sequences, the commonly used strategies are blocking, truncation, sliding window, etc., but they will all lose the context information of the sequence. Assuming that the sequence length supported by the model completely meets the requirements of the length of the patent text, such as 100,000 words, but since the correlation between patent texts is mostly local rather than global, the whole patent is not considered to be vectorized, but needs to be split according to the characteristics of the patent text, so as to fundamentally solve the problem of local correlation and make the length of the text input into the large model meet the requirements. Based on this, the server can first filter the training patent pairs from the patent database, and based on the patent invention point extraction algorithm, cut each training patent text in the training sample pair to obtain the corresponding overall invention point and local invention point of each training patent text. The above patent invention point extraction algorithm is constructed based on the generative large model. Specifically, when the patent invention point extraction algorithm is used to cut each training patent text in the training sample pair to obtain the corresponding overall invention point and local invention point of each training patent text, or when each training patent text is cut in S1 to obtain the corresponding overall invention point and local invention point of each training patent text, the server can perform the following steps: S11: Extract key terms: extract key terms matching the patent knowledge graph from each training patent text based on the pre-constructed patent knowledge graph.

[0030] S12: Generate overall invention points: splice each training patent text and the extracted key terms of each training patent text through the first prompt word template, and input into the first generative large model to obtain the corresponding overall invention points of each training patent text.

[0031] S13: Generate local invention points: determine the number of times each key term of each training patent text appears in each paragraph of each training patent text, and determine the corresponding local invention points of each training patent text according to the determined number of times of each paragraph and the corresponding weight of each key term of each training patent text.

[0032] The patent knowledge graph in S11 above is constructed based on the search log information of the patent search system. The nodes in the patent knowledge graph represent key terms, and the edges represent the relationships between the key terms. The patent search system accumulates a large amount of search log information. In addition, to avoid loss of search log information, the search log information in the patent search system can be regularly transferred every month. The search log information contains a large amount of valuable information, i.e., the search log information includes a large amount of search information used when searching for patents, i.e., patent search formulas, such as (IC='F21V21 / 30' OR IC='F21V21 / 14') AND ('support' AND / SEN 'rod' AND / SEN 'base'), IC represents the international patent classification number, 'F21V21 / 30' and 'F21V21 / 14' represent the classification numbers, and ((CPC='B28B11 / 245')) AND (('temperature control' OR 'temperature sensor') AND ('air cooler' OR 'water cooler') AND 'cooling water pipe') AND (USE='reduce concrete structure cracks'), CPC represents the cooperative patent classification, 'B28B11 / 245' represents the classification number, and USE represents the use or purpose of the technical solution described in the patent. The server can construct a patent knowledge graph based on the search log information, and can extract key terms from the training patent text based on the patent knowledge graph. Specifically, in S11 above, the server can extract key terms matching the nodes in the pre-constructed patent knowledge graph from each training patent text based on the pre-constructed patent knowledge graph. The number of key terms extracted from each training patent text that best represent each training patent text is not more than a preset number, which can be pre-set, such as 30. The generative large model at least includes the first generative large model described above, and can also include the second generative large model described below.

[0033] In the above S12, the server can splice the training patent text and the extracted key terms of the training patent text through the first prompt word template for each training patent text, and input into the first generative large model to obtain the overall invention point corresponding to the training patent text. Wherein, the above first prompt word template can include first role information, first input requirement, first output requirement, card slot corresponding to first input content and card slot corresponding to key terms, the first role information can be pre-set, the first role information can be "now you are a professional patent examiner, you need to use your professional knowledge to read and understand the following long text patent, and extract 3-6 invention points and implementation steps that can best represent the novelty and creativity of the patent". The above first input requirement and first output requirement can also be pre-set, the first input requirement can be "the input content is the patent text", and the first output requirement can be "filtering known theories and technologies, the output content must come from the input content, and the output content should as much as possible revolve around the given key terms". Therefore, the above first prompt word template can be "now you are a professional patent examiner, you need to use your professional knowledge to read and understand the following long text patent, and extract 3-6 invention points and implementation steps that can best represent the novelty and creativity of the patent, but you must pay attention to the following four points: (1) filtering known theories and technologies; (2) the output content must come from the input content; (3) the output content should as much as possible revolve around the given key terms, the given key terms are:

key terms

patent text

key terms

patent text

[0034] In the above S13, when determining the number of times the key terms of each training patent text appear in each paragraph of each training patent text, the server can determine the number of times the key terms appear in each paragraph of each training patent text. Wherein, the paragraphs corresponding to each training patent text can be the paragraphs corresponding to the specification of each training patent text.

[0035] In the step S13, the server determines the local invention points of each training patent text according to the number of times each paragraph corresponds to and the weight of each key term of each training patent text. S131: initialize the processing state of each paragraph of each training patent text to the first state.

[0036] S132: when the processing state of each paragraph of each training patent text is the first state, determine the importance of each paragraph of each training patent text according to the weight of each key term of each training patent text and the number of times each key term of each training patent text appears in each paragraph of each training patent text.

[0037] S133: combine adjacent paragraphs according to the importance of each paragraph of each training patent text, determine the shortest paragraph group of each training patent text, and take the patent segment as the condition for generating the patent segment: the importance of the patent segment of each training patent text is greater than half of the sum of the weights of all key terms of each training patent text; the importance of the patent segment is generated by superimposing the importance of each paragraph in the patent segment.

[0038] S134: set the processing state of the paragraph in the shortest paragraph group of each training patent text to the second state, and adjust the weight value of the hit key term in the shortest paragraph group of each training patent text to half of the current weight.

[0039] S135: repeat the steps S132-S134, recalculate the importance of each paragraph based on the adjusted weight value of the key term in S134 and determine the patent segment of the current period, until the execution times reach the set threshold or no new patent segment can be generated, and based on each patent segment of each training patent text, determine each local invention point of each training patent text using the second generative large model.

[0040] The first state in the above S131 can be pre-set, which can be represented by 0, and can be specifically represented as S i =0, i=1,2,…,p, p represents the number of paragraphs, S i represents the processing state of the paragraph. Specifically, the server can initialize the processing state of each paragraph of each training patent text by setting the processing state of each paragraph of the training patent text to the first state.

[0041] In S132, the server determines the importance of each paragraph of each training patent text according to the weight of the key term corresponding to the key term and the number of occurrences of the key term in the paragraph when the processing state of the paragraph is the first state. Then in S133, the server combines adjacent paragraphs of each training patent text according to the importance of each paragraph of the training patent text to determine the shortest paragraph group of the training patent text as a patent segment. The weight of each key term of each training patent text is pre-set (i.e., the initial weight), i.e., W j =1, j = 1, 2, …, q, q represents the number of key terms, and W j represents the weight of the key term j.

[0042] The shortest paragraph group is composed of consecutive paragraphs, and the server can use any pre-set algorithm to determine the consecutive shortest paragraph group of the training patent text according to the importance of each paragraph of the training patent text. Moreover, the condition for generating the patent segment is that the importance of the patent segment of each training patent text is greater than half of the sum of the weights of all key terms of each training patent text, and the importance of the patent segment is generated by superimposing the importance of each paragraph in the patent segment.

[0043] In S134, the server can set the processing state of the paragraphs in the shortest paragraph group of each training patent text to a second state, and adjust the weight value of the key terms hit in the shortest paragraph group of the training patent text to half of the current weight value after determining the shortest paragraph group of the training patent text. Then, in S135, the server can repeatedly perform S132-S134 to recalculate the importance of each paragraph and determine the patent segments of the current period based on the adjusted weight value of the key terms in S134 until the execution number reaches the set threshold value or no new patent segment can be generated. Based on the patent segments of the training patent text, the local invention points of the training patent text are determined. The second state can be pre-set, which can be represented by 1. The server can update the processing state of each paragraph in the shortest paragraph group of the training patent text from the first state to the second state. In addition, the server also adjusts the weight value of the key terms hit in the shortest paragraph group of the training patent text to half of the current weight value, and the current weight value is the initial weight value when the shortest paragraph group is the first generated shortest paragraph group. The recalculated importance of each paragraph refers to the paragraphs included in each training patent text, which is performed according to the process of S132. The determination of the patent segments of the current period is performed according to the process of S133, and the current period refers to the current execution round. One period (i.e., one execution) represents the process of generating a patent segment, i.e., the process of performing S132-S134. The set threshold value is pre-set. In addition, the server cannot generate a new patent segment refers to the importance of the combination of all adjacent paragraphs being less than half of the sum of the weights of the key terms.

[0044] In S135, based on the patent segments of each training patent text, the server can determine the local invention points of each training patent text using the second generative large model. For each training patent text, the server can use the patent segments of the training patent text as the local invention points of the training patent text.

[0045] In addition, in order to avoid the length of the generated local invention points exceeding the input length of the subsequent general text vectorization large model, when the patent segments of the training patent text exceed the limited length, the server can splice the second prompt word template with the patent segments (i.e., the patent segments exceeding the limited length) and input them into the second generative large model to reduce the length of the patent segments through the second generative large model to obtain the local invention points corresponding to the patent segments. When the patent segments of the training patent text do not exceed the limited length, the patent segments are used as the local invention points.

[0046] Specifically, the server can splice the second prompt word template with each patent segment of each training patent text when the patent segment exceeds the limited length, and input to the second generative large model to reduce the length of the patent segment through the second generative large model to obtain the corresponding local invention point of the patent segment. When the patent segment does not exceed the limited length, the patent segment is directly taken as the local invention point. The limited length can be pre-set, and the limited length can be 512 characters. The second generative large model can be the first generative large model, or other general large models, or large models fine-tuned from general large models, which are not limited in the specification. The second prompt word template is pre-set, and the second prompt word template includes second role information, second output requirements, card slots corresponding to second input content, and card slots corresponding to key terms. The second role information can be pre-set, and the second role information can be "Now you are a professional patent examiner, you need to use your professional knowledge to read and understand the following long text patent, and reduce the text to within 500 words". The second output requirement can also be pre-set, and the second output requirement can be "The output content must come from the input content, and the output content should as much as possible contain the given key terms". In addition, the second output requirement can also include the number of words of the output content, such as "The output content must be within 500 words". Therefore, the second prompt word template can be "Now you are a professional patent examiner, you need to use your professional knowledge to read and understand the following long text patent, and reduce the text to within 500 words, but please pay attention to the following two points: (1) The output content must come from the input content; (2) The output content should as much as possible contain the given key terms, the given key terms are:

key terms

patent segment

key terms

patent segment

[0047] S2: Invention point vectorization: vectorizing the overall invention point and the local invention point of each training patent text based on a general text vectorization large model to obtain the overall feature vector and the local feature vector of each training patent text.

[0048] In the specification, the server can invent point vectorization, that is, based on the general text vectorization large model, the overall invention points and the local invention points of each training patent text are vectorized to obtain the overall feature vector and the local feature vector of each training patent text. Specifically, the server can input the overall invention points and the local invention points of each training patent text into the general text vectorization large model respectively to obtain the overall feature vector and the local feature vector of each training patent text. Wherein, the general text vectorization large model is a Transformer model based on attention mechanism, and the above general text vectorization large model is a large model. Of course, the server can input the overall invention points of each training patent text into the general text vectorization large model to obtain the overall feature vector, and input the local invention points of each training patent text into the general text vectorization large model to obtain the local feature vector.

[0049] S3: Calculate the patent relevance: according to each overall feature vector and each local feature vector, calculate the patent relevance of the training patent pair.

[0050] In the specification, the server can calculate the patent relevance, that is, according to each overall feature vector and each local feature vector, the patent relevance of the training patent pair is calculated. Specifically, the server can use a multi-vector matching algorithm to extract one feature vector from the overall feature vector and the local feature vector corresponding to the two training patent texts to form a matching pair, calculate the feature similarity of all matching pairs of the two training patents, construct a feature similarity matrix, and determine the patent relevance of the training patent pair according to the feature similarity matrix. Wherein, all the above matching pairs are the overall feature vector and the local feature vector of any one of the two training patent texts and the overall feature vector and the local feature vector of the other training patent text except the training patent text.

[0051] In addition, the server can also perform the following steps: S31: The two training patent texts in the training patent pair are taken as the first training patent text and the second training patent text respectively, and the overall feature vector and the local feature vector of the first training patent text are taken as each first feature vector, and the overall feature vector and the local feature vector of the second training patent text are taken as each second feature vector.

[0052] S32: Each first feature vector is combined with each second feature vector respectively to obtain each matching pair, and the feature similarity between the first feature vector and the second feature vector in each matching pair is calculated respectively.

[0053] The above-described multi-vector matching algorithm extracts one feature vector from the overall feature vector and one from the local feature vectors corresponding to the two training patent texts and combines them into matching pairs. The feature similarity of all matching pairs between the two training patent texts is calculated, and the process of constructing the feature similarity matrix is ​​similar to steps S31-S32 above, and will not be repeated here. Furthermore, each of the above-described first feature vectors includes the overall feature vector and local feature vectors of the first training patent text, and each of the above-described second feature vectors includes the overall feature vector and local feature vectors of the second training patent text. Each matching pair includes one first feature vector and one second feature vector. When obtaining each matching pair, the server can combine each first feature vector with each second feature vector to obtain each matching pair. For example, suppose there are 5 first feature vectors, i.e., V... w V m V o1 V o2 V o3 and 4 second eigenvectors, namely U w U m U o1 U o2 Therefore, there are 20 matching pairs, i.e., V. w and U w V w and U m V w and U o1 V w and U o2 V m and U w V m and U m V m and U o1 V m and U o2 V o1 and U w V o1 and U m V o1 and U o1 V o1 and U o2 V o2 and U w V o2 and U m V o2 and U o1 V o2 and U o2 V o3 and U w V o3 and U m Vo3 and U o1 , V o3 and U o2 . The feature similarity can be a cosine similarity between the first feature vector and the second feature vector, and the server can calculate the feature similarity of each matching pair through a cosine similarity algorithm. The local invention point is also taken into account in the similarity calculation, i.e., the local similarity has an impact on the degree of association of the patents. The feature similarity matrix is composed of the feature similarities of the matching pairs.

[0054] S33: Determine the similarity weight corresponding to each matching pair according to the weights corresponding to the first feature vector and the second feature vector in each matching pair.

[0055] S34: Determine the maximum feature similarity in the feature similarities of the matching pairs, and determine the similarity weight corresponding to the maximum feature similarity as the maximum similarity weight.

[0056] S35: From the matching pairs, determine the matching pairs that meet the target condition as target matching pairs; the first feature vector and the second feature vector included in each target matching pair are not the same, and the sum of the feature similarities of the target matching pairs is the largest.

[0057] S36: Calculate the patent correlation of the training patent pair according to the maximum feature similarity, the maximum similarity weight, the feature similarity of each target matching pair, and the similarity weight of each target matching pair.

[0058] The process of determining the patent correlation of the training patent pair according to the feature similarity matrix is similar to the processes of steps S33-S36, and will not be described here. The weights corresponding to the first feature vector and the second feature vector in each matching pair are pre-set, such as the weight of each first feature vector can be W w , W m , W o1 , W o2 , W o3 , and the weight of each second feature vector can be W w , W m , W o1 , W o2 , wherein W w =1.0, W m =0.9, W o1 = W o2 = W o3 =0.8. The feature similarity of each matching pair can be S1, S2, S3, …, S 20 , and the corresponding similarity weight can be W1, W2, W3, …, W 20The similarity weight is calculated by multiplying the weight of the first feature vector and the weight of the second feature vector in each matching pair. w and U m For example, the similarity weight of the matching pair including V .

[0059] The similarity weight of each matching pair is the product of the weight of the first feature vector and the weight of the second feature vector in each matching pair. Specifically, in S33, the server can multiply the weight of the first feature vector and the weight of the second feature vector in each matching pair as the similarity weight of the matching pair. The maximum feature similarity is the maximum of the feature similarities of each matching pair, and the maximum similarity weight is the similarity weight of the matching pair corresponding to the maximum feature similarity. The maximum feature similarity can be represented by , and the maximum similarity weight can be represented by .

[0060] The target condition in S35 can be pre-set, and the target condition can be that the first feature vector and the second feature vector included in each target matching pair are different, and the feature similarity of each target matching pair is the largest. That is, the server can first determine the first number of the first feature vector of the first training patent text, and determine the second number of the second feature vector of the second training patent text, compare the first number and the second number, and take the smaller number as the target number. Based on each matching pair, a combination group of matching pairs including the target number is constructed, and the first feature vector and the second feature vector of each matching pair in the constructed combination group are different. The sum of the feature similarities of each matching pair in each combination group is calculated as the total feature similarity, and the combination group corresponding to the maximum total feature similarity is taken as the target combination group, and each matching pair in the target combination group is taken as each target matching pair satisfying the target condition.

[0061] In S36, the server can first calculate the first weighted similarity of the training patent pair according to the feature similarity of each target matching pair and the similarity weight of each target matching pair.

[0062] Then, the product of the maximum feature similarity and the maximum similarity weight is calculated as the second weighted similarity, and the patent relevance of the training patent pair is calculated according to the first weighted similarity and the second weighted similarity.

[0063] S4: Adjusting the general text vectorization large model based on the patent relevance: adjusting the parameters of the general text vectorization large model based on the patent relevance of the training patent pair until the vectorization result output by the general text vectorization large model reaches the set expectation, obtaining a patent text vectorization large model.

[0064] In the specification, the server can adjust the general text vectorization large model based on the patent relevance, that is, adjust the parameters of the general text vectorization large model based on the patent relevance feedback of the training patent pair until the vectorization result output by the general text vectorization large model reaches the set expectation to obtain the patent text vectorization large model. When adjusting the parameters of the general text vectorization large model based on the patent relevance feedback of the training patent pair, the server can calculate the loss function value by using the loss function of the general text vectorization large model constructed based on the patent relevance according to the patent relevance and the bias of the training patent pair, and feed back and adjust the parameters of the general text vectorization large model according to the loss function value. When the two training patent texts in the training patent pair are related, the bias of the training patent pair is positive. When the two training patent texts in the training patent pair are not related, the bias of the training patent pair is negative. The above loss function is constructed based on the bias and the patent relevance in advance, different biases have different loss function values calculated based on the loss function, and different feedbacks to the general text vectorization large model. The above trained patent text vectorization large model is used for vectorizing the input text. The above vectorization result is the output result of the general text vectorization large model, and the vectorization result reaching the set expectation can represent that the number of feedback adjustment cycles reaches the preset threshold, or that the difference between the patent relevance and the bias determined based on the vectorization result is less than the specified threshold, or that the loss function value is less than the target threshold, which is not limited in the specification. The above preset threshold, specified threshold and target threshold are all set in advance.

[0065] When calculating the loss function value by using the loss function of the general text vectorization large model constructed based on the patent relevance according to the bias and the patent relevance of the training patent pair, the server can calculate the loss function value of the training patent pair based on the loss function constructed based on the bias and the patent relevance according to the bias and the patent relevance of the training patent pair. The above loss function can be formula (1) as follows: (1) Wherein, loss represents the loss function value calculated by the loss function, and S represents the patent relevance.

[0066] The above loss function is an indicator for adjusting the model parameters. The smaller the loss function value determined based on the loss function, the more rewards the model parameters will be given, and the larger the loss function value determined based on the loss function, the more penalties the model parameters will be given. If the bias of the training patent pair is positive, the larger the patent relevance, the smaller the difference, and the smaller the loss function value. On the contrary, the smaller the patent relevance, the larger the difference, and the larger the loss function value. If the bias of the training patent pair is negative, the larger the patent relevance, the larger the difference, and the larger the loss function value. On the contrary, the smaller the patent relevance, the smaller the difference, and the smaller the loss function value.

[0067] In addition, when the bias is positive, the server can also calculate a first loss function value according to the first loss function pre-constructed according to the patent relevance, and train the general text vectorization large model according to the first loss function value. When the bias is negative, a second loss function value is calculated according to the second loss function pre-constructed according to the patent relevance, and the general text vectorization large model is trained according to the second loss function value. Wherein, the above-mentioned first loss function can be the following formula (2): (2) Wherein, the above-mentioned formula (2) is a first loss function pre-set, S represents the patent relevance.

[0068] And the second loss function can be the following formula (3): (3) Wherein, the above-mentioned formula (3) is a second loss function pre-set, S represents the patent relevance.

[0069] In some embodiments of the present specification, when the importance of the paragraph is determined according to the weight of the key term corresponding to the training patent text and the number of times the key term appears in the paragraph in the above S132, the following formula (4) can be used for calculation: (4) Wherein, M i represents the importance of paragraph i, represents the number of times the key term j appears in paragraph i.

[0070] In some embodiments of the present specification, when the first weighted similarity of the training patent pair is calculated in the above S36, the following formula (5) can be used for calculation: (5) Wherein, represents the first weighted similarity, represents the feature similarity corresponding to the target matching pair, WW k represents the similarity weight corresponding to the target matching pair, represents the sum of the similarity weights of the k target matching pairs, represents the sum of the product of the feature similarity and the similarity weight of the k target matching pairs, and k represents the target number.

[0071] When the patent relevance of the training patent pair is calculated according to the first weighted similarity and the second weighted similarity, the following formula (6) can be used for calculation: (6) wherein S represents a patent similarity, represents a first weighted similarity, represents a maximum feature similarity, represents a maximum similarity weight, represents a second weighted similarity, W1 represents a weight corresponding to the first weighted similarity, W2 represents a weight corresponding to the second weighted similarity, W1 and W2 are preset, W1 can be 0.3, and W2 can be 0.7.

[0072] In some embodiments of the present disclosure, when the patent text vectorization large model is obtained in the step S4, the server can redetermine the training patent pairs and repeatedly perform the steps S2-S4 until the number of cycles of the feedback adjustment reaches a preset threshold, and then the general text vectorization large model is used as the patent text vectorization large model. The preset threshold can be preset, and the number of cycles of the feedback adjustment can be used to represent the number of training times. The redetermination of the training patent pairs can be that the training patent pairs are reselected from the patent database, that is, the training patent pairs are reassembled. In addition, since the training patent pairs can be multiple, the redetermination of the training patent pairs can be performed on the multiple training patent pairs.

[0073] In addition, for one training in logic, that is, one cycle, the general text vectorization large model is trained based on a batch of training patent pairs, and each training patent pair in the batch is processed in parallel. For ease of description, the training patent pairs are described in the present disclosure. The number of training patent pairs in each batch can be preset, for example, 16.

[0074] In some embodiments of the present disclosure, the server can test and evaluate the general text vectorization large model, and use the general text vectorization large model as the patent text vectorization large model after the evaluation is passed. Based on this, the evaluation and deployment method of the patent text vectorization large model in the step S4 is that the server can obtain a test patent pair, and cut each test patent text in the test patent pair to obtain the overall invention point and the local invention point corresponding to each test patent text, respectively. The specific process is similar to that of the step S1, and will not be described here.

[0075] Then, the server can input the overall invention point and the local invention point of each test patent text into the general text vectorization large model to obtain the overall feature vector and the local feature vector of each test patent text. According to the overall feature vector and the local feature vector of each test patent text, the patent relevance of the test patent pair is calculated. The specific process is similar to that of the steps S2-S3, and will not be described here.

[0076] Afterwards, the server can determine the bias of the test patent pair, evaluate the general text vectorization large model according to the determined bias and the calculated patent relevance, and use the general text vectorization large model as the patent text vectorization large model after the evaluation is passed. The evaluation index can be accuracy, and other evaluation indexes are also possible, which are not limited in the specification.

[0077] In addition, after obtaining the patent text vectorization large model, the patent text vectorization large model can be deployed to a target device, which is a system, program, software or hardware device that applies the patent text vectorization large model, and the target device can be pre-set, and the specific device is not limited in the specification.

[0078] In some embodiments of the specification, the first local invention point generated in the process of S13 is the main invention point, and the local invention points other than the main invention point are the secondary invention points. For an invention patent, the number of secondary invention points is not more than three, and for an invention patent, the number of secondary invention points is zero. The overall feature vector corresponding to the overall invention point can be represented by V w Or U w , the feature vector corresponding to the main invention point can be represented by V m Or U m , and the feature vector corresponding to the secondary invention point can be represented by V o1 , V o2 , V o3 Or U o1 , U o2 .

[0079] In some embodiments of the specification, after obtaining the patent text vectorization large model, the patent text vectorization large model can be used to calculate the patent relevance of two patents. The specific application method is that the server can obtain two patent texts to be calculated, and cut each patent text to be calculated to obtain the overall invention point and the local invention point corresponding to each patent text to be calculated, the specific process being similar to the process of step S1, which will not be repeated here. The patent text vectorization large model deployed on the target device is used to vectorize the overall invention point and the local invention point of each patent text to be calculated to obtain the overall feature vector and the local feature vector of each patent text to be calculated, and the specific process is similar to the process of step S2, which will not be repeated here. According to the overall feature vector and the local feature vector of each patent text to be calculated, the patent relevance of the two patent texts to be calculated is calculated, and the specific process is similar to the process of step S3, which will not be repeated here.

[0080] Specifically, as shown in Figure 2 , the patent text vectorization large model is used to vectorize the overall invention point and the local invention point of each patent text to be calculated to obtain the overall feature vector and the local feature vector of each patent text to be calculated, and the specific process is similar to the process of step S2, which will not be repeated here. According to the overall feature vector and the local feature vector of each patent text to be calculated, the patent relevance of the two patent texts to be calculated is calculated, and the specific process is similar to the process of step S3, which will not be repeated here. Figure 2For the schematic diagram of a method for calculating patent relevance based on a large patent text vectorization model provided in the specification, the above two patent texts to be calculated (i.e. the first text to be calculated and the second text to be calculated) are the patent text 1 to be calculated (i.e. the first patent text to be calculated) and the patent text 2 to be calculated (i.e. the second patent text to be calculated) in Figure 2 . The server can determine the overall invention points and local invention points of the patent text 1 to be calculated and the patent text 2 to be calculated based on the patent invention point extraction algorithm, and determine the first feature vectors (i.e. the first feature vectors 1~n) of the patent text 1 to be calculated and the second feature vectors (i.e. the second feature vectors 1~m) of the patent text 2 to be calculated based on the large patent text vectorization model, and then determine the matching pairs according to the first feature vectors and the second feature vectors, and calculate the feature similarity of each matching pair. Figure 2 In the above formula, D.1-1~D.1-m, D.2-1~D.2-m, D.3-1~D.3-m …… D.n-1~D.n-m respectively represent the feature similarity corresponding to each matching pair. D.n-m represents the feature similarity between the nth first feature vector and the mth second feature vector. Figure 2 The "calculate the best matching" in the above formula is to determine the target matching pairs that meet the target conditions from the matching pairs, Figure 2 The feature similarity of the matching pair marked in gray in the above formula is D.1-2, D.2-1, D.3-m …… D.n-3, that is, Figure 2 The feature similarity of the matching pair marked in gray in the above formula is D.1-2, D.2-1, D.3-m …… D.n-3, that is, Figure 2 The "calculate the patent relevance" in the above formula is to calculate the patent relevance between the patent text 1 to be calculated and the patent text 2 to be calculated, which can be specifically implemented as the above steps S33~S36.

[0081] In the specification, as shown in Figure 3 , Figure 3 is a schematic diagram of the training process of a large patent text vectorization model for patent relevance calculation provided in the specification, Figure 3 , only one training patent pair is taken as an example in the above formula, that is, Figure 3 the training patent text 1 and the training patent text 2 in the above formula, the server can cut the training patent text 1 and the training patent text 2 respectively by using the patent invention point extraction algorithm to obtain the overall invention points and local invention points of the training patent text 1 and the training patent text 2, and then determine the first feature vectors (i.e. the first feature vectors 1~N) of the training patent text 1 and the second feature vectors (i.e. the second feature vectors 1~N) of the training patent text 2 based on the general text vectorization large model. Figure 3 Figure 3 ​the second feature vectors 1~M in the second feature vector set). Then, according to the first feature vectors and the second feature vectors, the matching pairs are determined, and the feature similarities of each matching pair are calculated. Figure 3 D.1-1~D.1-M, D.2-1~D.2-M, D.3-1~D.3-M …… D.N-1~D.N-M in the table respectively represent the feature similarities corresponding to each matching pair. D.N-M represents the feature similarity between the Nth first feature vector and the Mth second feature vector. Figure 3 “Calculate the best matching” in the table is to determine the target matching pairs satisfying the target condition from the matching pairs, Figure 3 The feature similarities corresponding to the target matching pairs shown in the table are D.1-2, D.2-1, D.3-M …… D.N-3, that is, Figure 3 The feature similarities of the matching pairs marked in gray in the table are calculated. Then, the server can calculate the patent relevance between the training patent text 1 and the training patent text 2, that is, Figure 3 “Calculate the patent relevance” in the table can be specifically as described above in steps S33~S36. Then, according to the bias of the training patent pair and the patent relevance, the pre-set loss function is used to feedback (i.e., train) the general text vectorization large model, so as to obtain the patent text vectorization large model, that is, Figure 3 The patent text vectorization large model represented by the dashed line in the table.

[0082] In some embodiments of the present specification, an evaluation patent pair can be obtained, and the evaluation patent pair includes a first patent text and a second patent text related to the first patent text. The first patent text is respectively subjected to the above Figure 2 The method shown in the table and the two control methods perform patent retrieval based on a patent database to respectively obtain the top 400 third patent texts related to the first patent text. The first control method is to use a general text vectorization large model without training, to use a sliding window method to vectorize the patent text, and to use cosine similarity to calculate the patent relevance, which is specifically as shown in Figure 4 , Figure 4 is a schematic diagram of a process of calculating the patent relevance by using the first control method provided in the present specification, Figure 4 In the table, the patent text 1 is the first patent text, and the patent text 2 is any patent text in the patent database. The patent text 1 and the patent text 2 are respectively input into the general text vectorization large model to obtain the feature vectors corresponding to the patent text 1 and the patent text 2, that is, the feature vector corresponding to the patent text 1 is composed of T1~ T A (i.e. Figure 4 The feature vector corresponding to the patent text 2 is composed of T1~ T B (i.e. Figure 4the eigenvectors in the vector 1~B) in the patent text 1, wherein T1~ T A the eigenvectors in the vector 1~B) in the patent text 1, wherein T1~ T B the eigenvectors in the vector 1~B) in the patent text 1, wherein T1~ T Figure 4 the eigenvectors in the vector 1~B) in the patent text 1, wherein T1~ T

[0083] The second comparison method is to use the patent invention point extraction algorithm, use the general text vectorization large model which has not been trained to vectorize the patent text, and use the multi-vector matching algorithm to calculate the patent correlation, as shown in Figure 5 Figure 5 is a schematic diagram of a process of calculating the patent correlation using the second comparison method provided in the specification, Figure 5 In the patent text 1 is the first patent text, and the patent text 2 is any patent text in the patent database. Figure 5 The process of calculating the patent correlation between the first patent text and any patent text in the patent database is shown. First, use the patent invention point extraction algorithm to determine the overall invention point and the local invention point corresponding to the patent text 1 and the overall invention point and the local invention point corresponding to the patent text 2, respectively. Then use the general text vectorization large model to determine the first eigenvector of the patent text 1 (i.e. Figure 5 the first eigenvector 1~Q) in the vector 1~Q) and the second eigenvector of the patent text 2 (i.e. Figure 5 the second eigenvector 1~W) in the vector 1~W). Then determine each matching pair according to each first eigenvector and each second eigenvector, and calculate the feature similarity of each matching pair. Figure 5 D.1-1~D.1-W, D.2-1~D.2-W, D.3-1~D.3-W…D.Q-1~D.Q-W in the D.1-1~D.1-W, D.2-1~D.2-W, D.3-1~D.3-W…D.Q-1~D.Q-W respectively represent the feature similarity corresponding to each matching pair. D.Q-W represents the feature similarity between the Qth first eigenvector and the Wth second eigenvector. Figure 5 “Calculate the best match” in the D.1-1~D.1-W, D.2-1~D.2-W, D.3-1~D.3-W…D.Q-1~D.Q-W is to determine each target matching pair that meets the target condition from each matching pair, Figure 5 The feature similarity corresponding to each target matching pair is shown as D.1-2, D.2-1, D.3-W…D.Q-3, i.e. Figure 5 ​The feature similarity of the matched pairs marked in gray is then calculated. The server can then calculate the patent relevance between patent text 1 and patent text 2, i.e. Figure 5 The "calculation of patent relevance" can be performed as described in steps S33-S36 above. Based on the patent relevance between patent text 1 and each patent text 2, the third patent text is determined from each patent text 2.

[0084] and Figure 2 The method shown involves using a patent invention point extraction algorithm, vectorizing the patent text using a large-scale patent text vectorization model, and calculating the patent relevance using a multi-vector matching algorithm. It should be noted that the two patent texts to be calculated can be either the first patent text or any patent text from the patent database. When the patent relevance of the two patent texts exceeds a threshold, it indicates that the two patent texts are related, and therefore the other patent text among the two patent texts is related to the first patent text. Conversely, when the patent relevance of the two patent texts does not exceed the threshold, it indicates that the two patent texts are not related, and therefore the other patent text among the two patent texts is not related to the first patent text. Based on Figure 2 The method shown identifies the third patent text as a patent text whose patent relevance to the first patent text exceeds a threshold.

[0085] Next, the proportion of the third patent text that matches the second patent text, determined based on the first comparison method, is determined as the first proportion, which is approximately 50%. The proportion of the third patent text that matches the second patent text, determined based on the second comparison method, is determined as the second proportion, which is approximately 56%. And the proportion based on... Figure 2 The method shown determines the proportion of the third patent text that matches the second patent text, and this third proportion is approximately 65%. In summary, Figure 2 The evaluation results of the method shown (i.e., the third proportion) are better than the evaluation results of the second control method (i.e., the second proportion), and the evaluation results of the second control method are better than the evaluation results of the first control method (i.e., the first proportion). Therefore, it is shown that... Figure 2 The patent invention point extraction algorithm, the large model of patent text vectorization, and the multi-vector matching algorithm improve the accuracy of patent relevance calculation, especially the large model of patent text vectorization.

[0086] It should be noted that the above detailed description of the specific embodiments of the present application is not intended to limit the present application in any way. Thus, while the present application has been described in detail with reference to specific embodiments thereof, it will be apparent to those skilled in the art that various modifications and changes can be made thereto without departing from the spirit and scope of the present application.

Claims

1. A method for generating a large patent text vectorization model for patent relevance calculation, characterized in that, Comprise: S1: Assemble training patent pairs and extract overall and local invention points: Assemble training patent pairs, a set of said training patent pairs comprises two training patent texts, cut each training patent text to obtain the overall invention point and the local invention point corresponding to each training patent text; S2: Invention point vectorization: based on a general text vectorization large model, the overall invention point and the local invention point of each training patent text are vectorized to obtain the overall feature vector and the local feature vector of each training patent text; S3: Calculate patent relevance: according to each overall feature vector and each local feature vector, calculate the patent relevance of the training patent pair; S4: Adjust the general text vectorization large model based on the patent relevance: adjust the parameters of the general text vectorization large model based on the patent relevance of the training patent pair until the vectorization result output by the general text vectorization large model reaches the set expectation, and obtain a patent text vectorization large model.

2. The method of claim 1, wherein the method further comprises: determining a number of patents to be used for training the model; and selecting the number of patents from the patent database based on the number of patents to be used for training the model. The patent text vectorization large model obtained in S4 specifically comprises: Redetermine the training patent pair, and repeat steps S2-S4 until the number of feedback adjustment cycles reaches a preset threshold, and the general text vectorization large model is used as a patent text vectorization large model.

3. The method of claim 1, wherein the method further comprises: determining a number of patents to be used for training the model; and selecting the number of patents from the patent database based on the number of patents to be used for training the model. In S4, the method for adjusting the parameters of the general text vectorization large model based on the patent relevance is: According to the patent relevance and the bias of the training patent pair, use the loss function of the general text vectorization large model constructed based on the patent relevance to calculate the loss function value, and according to the loss function value, feedback adjust the parameters of the general text vectorization large model; wherein, when the two training patent texts in the training patent pair are related, the bias of the training patent pair is positive; when the two training patent texts in the training patent pair are not related, the bias of the training patent pair is negative.

4. The method of claim 3, wherein the method further comprises: determining a weight of each of the plurality of words in the patent text; and determining a weight of each of the plurality of words in the patent text based on a frequency of each of the plurality of words in the patent text. The loss function in S4 is: (1) In the formula, S is the patent relevance.

5. The method of claim 1, wherein the method further comprises: determining a number of patents to be used for training the model; and selecting the number of patents from the patent database based on the number of patents to be used for training the model. The evaluation and deployment method of the patent text vectorization large model in S4 comprises: Obtain a test patent pair, and cut each test patent text in the test patent pair respectively to obtain the overall invention point and the local invention point corresponding to each test patent text respectively; Input the overall invention point and the local invention point of each test patent text into the general text vectorization large model respectively to obtain the overall feature vector and the local feature vector of each test patent text; According to the overall feature vector and the local feature vector of each test patent text, calculate the patent relevance of the test patent pair; Determine the bias of the test patent pair, and according to the determined bias and the calculated patent relevance, evaluate the general text vectorization large model, and after the evaluation is passed, use the general text vectorization large model as a patent text vectorization large model, and deploy the patent text vectorization large model to a target device.

6. The method of claim 1, wherein the method further comprises: determining a number of patents to be used for training the model; and selecting the number of patents from the patent database based on the number of patents to be used for training the model. The general text vectorization large model is a Transformer model based on an attention mechanism.

7. The method of claim 1, wherein the method further comprises: determining a number of patents to be used for training the model; and selecting the number of patents from the patent database based on the number of patents to be used for training the model. The patent text vectorization large model is used to calculate the patent relevance of two patents, and the specific application method is: Obtain two patent texts to be calculated, and cut each patent text to be calculated to obtain the overall invention points and local invention points corresponding to each patent text to be calculated respectively; Use the patent text vectorization large model deployed on the target device to vectorize the overall invention points and local invention points of each patent text to be calculated to obtain the overall feature vectors and local feature vectors of each patent text to be calculated; According to the overall feature vectors and local feature vectors of each patent text to be calculated, the patent relevance of the two patent texts to be calculated is calculated.

8. The method of claim 1, wherein the method further comprises: determining a number of patents to be used for training the model; and selecting the number of patents from the patent database based on the number of patents to be used for training the model. In S1, each training patent text is cut to obtain the overall invention points and local invention points corresponding to each training patent text, which specifically includes: S11: Extract key terms: based on a pre-constructed patent knowledge graph, extract key terms matching the patent knowledge graph from each training patent text; S12: Generate overall invention points: splice the first prompt word template, the each training patent text and the key terms extracted from the each training patent text, and input into the first generative large model to obtain the overall invention points corresponding to the each training patent text; S13: Generate local invention points: determine the number of times the key terms of the each training patent text appear in each paragraph of the each training patent text, and determine the local invention points corresponding to the each training patent text according to the number of times of each paragraph and the weight corresponding to the key terms of the each training patent text.

9. The method of claim 7, wherein the method further comprises: determining a weight of each of the plurality of words in the patent text; and determining a weight of each of the plurality of words in the patent text based on a frequency of each of the plurality of words in the patent text. In S13, the local invention points corresponding to the each training patent text are determined according to the number of times of each paragraph and the weight corresponding to the key terms of the each training patent text, which specifically includes: S131: initialize the processing state of each paragraph of the each training patent text to a first state; S132: when the processing state of each paragraph of the each training patent text is the first state, determine the importance of each paragraph of the each training patent text according to the weight corresponding to the key terms of the each training patent text and the number of times of the key terms of the each training patent text appearing in each paragraph of the each training patent text; S133: combine adjacent paragraphs according to the importance of each paragraph of the each training patent text to determine the shortest paragraph group of the each training patent text, and the patent segment is generated, the conditions for generating the patent segment are: the importance of the patent segment of the each training patent text is greater than half of the sum of the weights of all key terms of the each training patent text; the importance of the patent segment is generated by superimposing the importance of each paragraph in the patent segment; S134: set the processing state of the paragraph in the shortest paragraph group of each training patent text to a second state, and adjust the weight value of the hit key term in the shortest paragraph group of each training patent text to half of the current weight; S135: repeat steps S132-S134, based on the adjusted weight value of the key term in S134, recalculate the importance of each paragraph and determine the patent segment of the current period, until the execution times reach the set threshold or no new patent segment can be generated, based on each patent segment of the training patent text, using a second generative large model to determine each local invention point of the training patent text.

10. The method of claim 1, wherein the method further comprises: determining a number of patents to be used for training the model; and selecting the number of patents from the patent database based on the number of patents to be used for training the model. The S3 specifically comprises: The two training patent texts in the training patent pair are taken as a first training patent text and a second training patent text respectively, and the global feature vector and the local feature vector of the first training patent text are taken as each first feature vector, and the global feature vector and the local feature vector of the second training patent text are taken as each second feature vector; Each first feature vector is combined with each second feature vector to obtain each matching pair, and the feature similarity between the first feature vector and the second feature vector in each matching pair is calculated respectively; According to the weight corresponding to the first feature vector and the second feature vector in each matching pair, the similarity weight corresponding to each matching pair is determined; Determine the maximum feature similarity in the feature similarity of each matching pair, and determine the similarity weight corresponding to the maximum feature similarity, and take it as the maximum similarity weight; From the matching pairs, determine each matching pair that meets the target condition, and take it as each target matching pair; each target matching pair includes different first feature vectors and second feature vectors, and the sum of the feature similarities of the target matching pairs is maximum; According to the maximum feature similarity, the maximum similarity weight, the feature similarity of each target matching pair, and the similarity weight of each target matching pair, the patent relevance of the training patent pair is calculated.