A method and device for evaluating cross-border data

By establishing a structured patent knowledge base and patent generation model, similarity search and evaluation of cross-border data technical solutions is solved, and the rapid evaluation problem in cross-border data patent applications is improved, and the application efficiency and quality are improved.

CN120105240BActive Publication Date: 2025-08-29FU JIUE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510601214.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-29
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

The technical briefing for enterprises’ cross-border data lacks the ability to make quick judgments when applying for patents, resulting in repeated modifications and affecting application efficiency.

Method used

Establish a structured patent knowledge base, search and evaluate the similarity of the technical solutions to be evaluated related to cross-border data through a pre-trained patent generation model, and generate a recommendation application index.

Benefits of technology

It has achieved rapid and accurate assessment of corporate cross-border technology and improved the efficiency and quality of patent applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105240B_ABST
    Figure CN120105240B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for evaluating cross-border data, comprising the following steps: collecting multiple authorized and publicly available patents associated with the cross-border data, establishing a structured patent knowledge base based on the structured patent content of the publicly available patents; optimizing a pre-trained patent generation model based on the structured patent content; obtaining technical solutions to be evaluated associated with the cross-border data, extracting technical keywords and technical structure data from the technical solutions to be evaluated, inputting these into the patent knowledge base for similarity search, and obtaining similar patent structure paragraphs; generating a technical text to be evaluated based on the technical keywords and technical structure data using the optimized patent generation model, generating similar patent texts based on the technical keywords and similar patent structure paragraphs, and generating a recommended application index for the technical solution to be evaluated based on the similarity scores between the technical text to be evaluated and the similar patent texts. The present invention can rapidly evaluate the patentability of a company's cross-border technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data evaluation technology, and in particular to a method and device for evaluating cross-border data. Background Art

[0002] Cross-border data refers to data transmitted, stored, or processed between different countries or regions. Therefore, the compliance and security of cross-border data are crucial. Consequently, technical solutions related to cross-border data require patent applications, which necessitates technical personnel to provide a technical briefing document. However, current company technical personnel lack the ability to quickly determine whether a technical briefing document is patentable. Furthermore, cross-border data involves application regulations in different countries or regions, making it difficult for technical briefing documents to quickly meet corresponding quality requirements. Consequently, repeated revisions are required, impacting application efficiency.

[0003] Therefore, it is necessary to provide a preliminary and rapid evaluation technology for whether a company's cross-border technology can be patented. Summary of the Invention

[0004] In order to solve the above-mentioned problems of the prior art, the present invention provides a method and device for evaluating cross-border data, which can quickly evaluate the patentability of an enterprise's cross-border technology.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is:

[0006] In a first aspect, the present invention provides a method for evaluating cross-border data, comprising the steps of:

[0007] S1. Collect multiple authorized and published patents associated with cross-border data, and establish a structured patent knowledge base based on the structured patent content of the published patents;

[0008] S2. Optimizing a pre-trained patent generation model based on the structured patent content;

[0009] S3. Obtain the technical solutions to be evaluated that are associated with the cross-border data, extract technical keywords and technical structure data from the technical solutions to be evaluated, input them into the patent knowledge base for similarity search, and obtain similar patent structure paragraphs;

[0010] S4. The optimized patent generation model generates a technical text to be evaluated based on the technical keywords and the technical structure data, and generates similar patent texts based on the technical keywords and the similar patent structure paragraphs, and generates a recommended application index for the technical solution to be evaluated based on the similarity scores between the technical text to be evaluated and the similar patent texts.

[0011] The beneficial effects of the present invention are as follows: structured patent content is established based on existing public patents, and the pre-trained patent generation model is optimized to ensure that the text generated by the optimized patent generation model is more in line with patent writing specifications; then, similar paragraphs are retrieved for the enterprise's cross-border technical solutions, thereby generating a technical text to be evaluated based on the enterprise's cross-border technical solutions and a similar patent text based on similar patent structure paragraphs, and finally generating a recommended application index for the technical solution to be evaluated based on the degree of similarity between the two, so that the technical solution to be evaluated is compared with the authorized public patent in terms of patent structure, thereby enabling a rapid evaluation of the patentability of the enterprise's cross-border technology.

[0012] Optionally, step S1 includes the following steps:

[0013] S11. Collect multiple authorized patent publications associated with cross-border data, and obtain patent keywords for the patent publications based on the reviewers' areas of interest when reviewing the patent publications;

[0014] S12. Arrange the disclosed patent into patent structure paragraphs corresponding to the technical means and technical effects according to each claim in the claims, and obtain structured patent content including patent keywords and patent structure paragraphs. The technical means and technical effects respectively include the original content corresponding to the claims in the specification;

[0015] S13. Establish a structured patent knowledge base based on the structured patent content of the disclosed patent.

[0016] According to the above description, semantics is ensured by patent keywords, and logic is ensured by patent structure paragraphs corresponding to technical means and technical effects, thereby ensuring that the structured patent content stored in the patent knowledge base is semantic and logical.

[0017] Optionally, the patent keywords include the content keywords of the paragraphs where the specification and claims are located, the abstract keywords of the paragraph where the specification abstract is located, and the figure keywords of the area where the specification figures are located. Then, the structured patent content including the patent keywords and the patent structure paragraphs obtained in step S12 is:

[0018] The drawing data including the drawing keywords, the abstract data including the original content of the specification abstract and the abstract keywords, and the structured patent content of the patent structure paragraph corresponding to each claim are obtained. The patent structure paragraph includes the corresponding technical means and technical effects. The technical means include the original technical content of the current claim, the corresponding original technical content in the specification, and the content keywords of the two parts of the original technical content. The technical effect includes the corresponding original effect content in the specification and the content keywords of the original effect content.

[0019] Optionally, step S11 includes the following steps:

[0020] S111. Collect multiple authorized patents associated with cross-border data, and obtain the attention weight of each word in the current paragraph based on the area of ​​interest of the reviewer when reviewing the patent. If it is a figure in the specification of the patent, identify the text content in each figure, and treat the text content of each figure as a paragraph. The attention weight W eye The formula for (w) is:

[0021] ;

[0022] Where, TFD(w) is the fixation duration of word w, TFD max is the maximum fixation duration of a single word in the current paragraph, FC(w) is the number of fixations on word w, RC(w) is the number of returns to look at word w in the subsequent region, and RC avg is the average number of times all words in the current paragraph are looked back;

[0023] S112. Obtain content keywords, abstract keywords, and figure keywords of the disclosed patent according to the location of the current paragraph and the attention weight of each word in the current paragraph.

[0024] According to the above description, the objectivity of keyword screening is ensured by converting the reviewer's gaze duration, number of gazes and return gaze behavior into calculable weight values.

[0025] Optionally, the step S111 further includes the following steps:

[0026] If the fixation duration of the sentence containing word w is greater than the average fixation duration of all sentences in the current paragraph by a preset multiple standard deviation, the fixation weight is accumulated by a first preset coefficient to obtain a comprehensive word weight, and both the preset multiple and the first preset coefficient are greater than 1;

[0027] And / or if the number of times the sentence containing word w is looked back is greater than a preset proportion of the entire sentence in the patent disclosure, then the gaze weight is accumulated by a second preset coefficient to obtain a comprehensive word weight, the preset proportion is greater than 40%, and the second preset coefficient is greater than 1;

[0028] And / or identifying the weight of each word in the patent disclosure according to the word weight model, and accumulating the weight with the attention weight to form a comprehensive word weight.

[0029] According to the above description, the coefficients are multiplied for abnormally long gazes and intensive return glances to ensure that the final word weights are more in line with the true intentions of the reviewers.

[0030] Optionally, the patent generation model includes a text generator and a paragraph generator, and step S2 includes the following steps:

[0031] S21. Based on the structured patent content, the paragraph generator generates a patent comparison paragraph according to the content keywords of each paragraph, and the text generator generates a patent comparison text according to the patent comparison paragraph, abstract keywords, and figure keywords;

[0032] S22. Input the average similarity between all the patent comparison paragraphs and the corresponding patent structure paragraphs in the structured patent content, as well as the average similarity between all the patent comparison texts and the corresponding published patents, into the loss function to update and optimize the text generator and the paragraph generator according to the loss function.

[0033] Optionally, the loss function L is:

[0034] L = α × L Paragraph +β×L Article ;

[0035] Where α and β are hyperparameters, L Paragraph is the average similarity between the patent comparison paragraph and the corresponding patent structure paragraph, L Article It is the average similarity between the patent comparison text and the corresponding published patent.

[0036] Optionally, step S4 includes the following steps:

[0037] S41, generating a technical text to be evaluated based on the technical keywords and the technical structure data by the optimized text generator and paragraph generator;

[0038] S42, generating a similar patent text based on the technical keywords and the similar patent structure paragraphs by the optimized text generator;

[0039] S43. Generate a recommended application index for the technical solution to be evaluated based on the similarity scores between the technical text to be evaluated and the similar patent text.

[0040] Optionally, step S43 includes the following steps:

[0041] Generate an overall similarity S based on the similarity scores of the technical text to be evaluated and the similar patent text total ;

[0042] Calculate the paragraph similarity (s1, s2, ..., s n );

[0043] According to the overall similarity S total and paragraph similarity (s1,s2,…,s n ) Generate a recommendation application index for the technical solution to be evaluated, wherein the formula for the recommendation application index InnovationScore is:

[0044] InnovationScore=(1-S total )×(1-max(s1,s2,…,s n ));

[0045] Where, max(s1,s2,…,s n ) is the paragraph similarity (s1,s2,…,s n ) is the maximum value among .

[0046] In a second aspect, the present invention provides a device for evaluating cross-border data, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a method for evaluating cross-border data according to the first aspect is implemented.

[0047] Among them, the technical effects corresponding to the cross-border data evaluation device provided by the second aspect refer to the relevant description of the cross-border data evaluation method provided by the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 A schematic diagram of the main process of a cross-border data evaluation method according to an embodiment of the present invention;

[0049] Figure 2 A schematic diagram of a specific process of a cross-border data evaluation method according to an embodiment of the present invention;

[0050] Figure 3 A schematic diagram of storage of structured patent content involved in an embodiment of the present invention;

[0051] Figure 4 A schematic diagram of the framework of a cross-border data evaluation device according to an embodiment of the present invention.

[0052] Description of reference numerals:

[0053] 1. A device for evaluating cross-border data;

[0054] 2. Processor;

[0055] 3. Memory. DETAILED DESCRIPTION

[0056] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0057] Example 1

[0058] This embodiment is suitable for scenarios requiring patentability assessment of a company's cross-border technology solutions. Existing techniques rely on technicians making independent judgments, which can be slow and inaccurate. This embodiment establishes a structured patent knowledge base and uses it to conduct similarity assessments of a company's cross-border technology solutions, enabling rapid patentability assessments. See below for details.

[0059] It should be noted that the cross-border data evaluation method provided in this embodiment is intended to provide technical personnel with a preliminary evaluation of whether the technical solution that an enterprise is about to apply for a patent can be applied for. The evaluation result is for technical personnel to understand their own technical solution or to modify the technical solution accordingly, rather than to stamp a final conclusion on the technical solution.

[0060] Please refer to Figures 1 to 3 , a cross-border data evaluation method, comprising the steps of:

[0061] S1. Collect multiple authorized public patents associated with cross-border data, and establish a structured patent knowledge base based on the structured patent content of the public patents.

[0062] Among them, by collecting and downloading multiple authorized public patents from existing public websites, this application mainly involves cross-border data, so public patents related to cross-border data are downloaded. In other equivalent embodiments, in order to solve patent solutions in other fields, public patents corresponding to other fields can also be downloaded.

[0063] At this point, after obtaining the original data set of the public patent, existing data cleaning methods can be used to perform data deduplication, format standardization, and other processing.

[0064] In this embodiment, referring to Figure 2 It can be seen that step S1 includes the following steps:

[0065] S11. Collect multiple authorized public patents associated with cross-border data, and obtain patent keywords of the public patents based on the reviewers' areas of interest when reviewing the public patents.

[0066] The reviewer is not limited to patent examiners but can also be a patent agent or corporate IPR, etc. By using eye tracking equipment, the areas that the reviewer focuses on during the review are tracked, which are called areas of interest. Keywords of the published patents are then obtained based on these areas.

[0067] Among them, patent keywords include content keywords in the paragraphs where the specification and claims are located, abstract keywords in the paragraph where the specification abstract is located, and drawing keywords in the area where the specification drawings are located.

[0068] The text content of the published patent is divided into multiple areas of interest (AOIs), including:

[0069] (1) Vocabulary-level AOI: single word bounding box;

[0070] (2) Sentence-level AOI: the area of ​​each sentence;

[0071] (3) Paragraph-level AOI: the content range of the entire paragraph;

[0072] In this way, the auditors' gaze duration, number of gazes, and number of return glances in different areas of interest can be collected. The gaze duration refers to the cumulative length of time the gaze point stays in the area of ​​interest, the number of gazes refers to the total number of times the area of ​​interest is entered, and the number of return glances refers to the number of times the subsequent area returns to the area of ​​interest.

[0073] Specifically, in this embodiment, step S11 includes the following steps:

[0074] S111. Collect multiple authorized patents that are related to cross-border data, and obtain the attention weight of each word in the current paragraph based on the reviewer's interest area when reviewing the patent. If it is a drawing in the specification of the patent, identify the text content in each drawing, and treat the text content of each drawing as a paragraph. The attention weight W eye The formula for (w) is:

[0075] ;

[0076] Where, TFD(w) is the fixation duration of word w, TFD max is the maximum fixation duration of a single word in the current paragraph, FC(w) is the number of fixations on word w, RC(w) is the number of returns to look at word w in the subsequent region, and RC avg is the average number of times all words in the current paragraph are looked back.

[0077] In this embodiment, step S111 further includes the following steps:

[0078] If the fixation duration of the sentence containing word w is greater than the average fixation duration of all sentences in the current paragraph by a preset multiple standard deviation, the fixation weight is accumulated by a first preset coefficient to obtain a comprehensive word weight, and both the preset multiple and the first preset coefficient are greater than 1;

[0079] And / or if the number of times the sentence containing word w is looked back is greater than a preset ratio of the entire sentence in the published patent, then the gaze weight is accumulated by a second preset coefficient to obtain a comprehensive word weight, the preset ratio is greater than 40%, and the second preset coefficient is greater than 1;

[0080] And / or identifying the weight of each word in the disclosed patent according to the word weight model, and accumulating it with the attention weight to form a comprehensive word weight.

[0081] In this embodiment, the weights of the above-mentioned abnormally long gazes, intensive return gazes and model evaluation will all be used to calculate the gaze weights. Specifically, the gaze weights are accumulated by the preset coefficients of abnormally long gazes and intensive return gazes, and then added to the weights of the model evaluation to obtain the comprehensive word weights.

[0082] In this embodiment, the preset multiple is 1.5 times, the preset ratio is 75%, and the first preset coefficient and the second preset coefficient are 1.2. The word weight model can adopt the existing TextRank, large language model or deep learning model.

[0083] S112. Obtain content keywords, abstract keywords, and figure keywords of the published patent according to the location of the current paragraph and the attention weight of each word in the current paragraph.

[0084] In this embodiment, for the paragraphs of the specification and claims, the specification abstract, and the specification drawings, the words with the top m% comprehensive word weights are taken as the specification paragraph keywords, abstract keywords, and flowchart keywords, respectively.

[0085] S12. Organize the published patents into patent structure paragraphs corresponding to the technical means and technical effects according to each claim in the claims, and obtain structured patent content including patent keywords and patent structure paragraphs. The technical means and technical effects respectively include the original content corresponding to the claims in the specification.

[0086] In this embodiment, referring to Figure 3 It can be seen that the structured patent content including patent keywords and patent structure paragraphs obtained in step S12 is:

[0087] The drawing data including the keywords of the drawing, the abstract data including the original content of the specification abstract and the keywords of the abstract, and the structured patent content of the patent structure paragraph corresponding to each claim are obtained. The patent structure paragraph includes the corresponding technical means and technical effects. The technical means include the original technical content of the current claim, the corresponding original technical content in the specification, and the content keywords of the two parts of the original technical content. The technical effect includes the corresponding original effect content in the specification and the content keywords of the original effect content.

[0088] Among them, the original technical content of the claims, the original technical content of the specification, and the original effect content of the specification all refer to the content originally recorded in the application documents, and are named based on their location and content type.

[0089] S13. Establish a structured patent knowledge base based on the structured patent content of published patents.

[0090] Therefore, semantics is ensured through patent keywords, and logic is ensured through patent structure paragraphs corresponding to technical means and technical effects, thereby ensuring that the structured patent content stored in the patent knowledge base is semantic and logical.

[0091] S2. Optimize the pre-trained patent generation model based on structured patent content.

[0092] In this embodiment, the patent generation model includes a text generator and a paragraph generator. Figure 2 It can be seen that step S2 includes the following steps:

[0093] S21. Based on the structured patent content, the paragraph generator generates patent comparison paragraphs according to the content keywords of each paragraph, and the text generator generates patent comparison text according to the patent comparison paragraphs, abstract keywords and figure keywords.

[0094] In this embodiment, the paragraph generator uses lightweight pre-trained Transformer, BART and other models, and the text generator uses pre-trained large models such as GPT-3 and PaLM.

[0095] S22. The average similarity between all patent comparison paragraphs and the corresponding patent structure paragraphs in the structured patent content and the average similarity between all patent comparison texts and the corresponding published patents are input into the loss function to update and optimize the text generator and the paragraph generator according to the loss function.

[0096] Among them, the loss function L is:

[0097] L = α × L Paragraph +β×L Article ;

[0098] Where α and β are hyperparameters used to adjust L Article , L Paragraph The weight, L Paragraph is the average similarity between the patent comparison paragraph and the corresponding patent structure paragraph, L Article It is the average similarity between the patent comparison text and the corresponding published patent.

[0099] In this embodiment, the similarity between two paragraphs or two texts can be achieved using TF-IDF, cosine similarity, Jaccard similarity, etc.

[0100] S3. Obtain the technical solutions to be evaluated that are associated with the cross-border data, extract technical keywords and technical structure data from the technical solutions to be evaluated, input them into the patent knowledge base for similarity search, and obtain similar patent structure paragraphs.

[0101] Among them, the data scenario of this embodiment is cross-border data, and the cross-border related technical solutions are uniformly transmitted to the cross-border IDC for analysis. The technicians can choose whether to perform patent analysis. If so, the product technical description information is extracted through common data mining algorithms such as deep learning / regular expressions, and the product technical description information is divided into multiple groups of technical means-technical effects paragraphs, and the technical keywords of the product technical description information are obtained. The technical keywords and patent keywords also include content keywords, abstract keywords and figure keywords. Use search enhancement generation to first retrieve similar patent structure paragraphs from the patent knowledge base.

[0102] S4. The optimized patent generation model generates a technical text to be evaluated based on technical keywords and technical structure data, and generates similar patent texts based on technical keywords and similar patent structure paragraphs. The recommended application index of the technical solution to be evaluated is generated based on the similarity scores between the technical text to be evaluated and the similar patent texts.

[0103] In this embodiment, step S4 includes the following steps:

[0104] S41. The optimized text generator and paragraph generator generate the technical text to be evaluated based on the technical keywords and technical structure data.

[0105] At this time, the technical text B to be evaluated is generated based on the technical keywords and technical structure data of the technical solution A to be evaluated.

[0106] S42. The optimized text generator generates similar patent texts based on technical keywords and similar patent structure paragraphs.

[0107] At this time, a similar patent text C is generated based on the technical keywords of the technical solution A to be evaluated and the similar patent structure paragraphs.

[0108] S43. Generate a recommended application index for the technical solution to be evaluated based on the similarity scores between the technical text to be evaluated and similar patent texts.

[0109] In this embodiment, step S43 includes the following steps:

[0110] S431. Generate an overall similarity S based on the similarity scores of the technical text to be evaluated and the similar patent texts. total .

[0111] At this time, calculate the similarity between the technical text B to be evaluated and the similar patent text C, and get the overall similarity S total .

[0112] S432, calculate the paragraph similarity (s1, s2, ..., s n ).

[0113] At this time, the technical solution A to be evaluated has n paragraphs, and the paragraph similarity (s1, s2, ..., s n ).

[0114] S433, according to the overall similarity S total and paragraph similarity (s1,s2,…,s n ) Generate the recommended application index of the technical solution to be evaluated. The formula of the recommended application index InnovationScore is:

[0115] InnovationScore=(1-S total )×(1-max(s1,s2,…,s n ));

[0116] Where, max(s1,s2,…,s n ) is the paragraph similarity (s1,s2,…,s n ) is the maximum value among .

[0117] Therefore, this embodiment can provide a solution for rapid patentability evaluation of an enterprise's cross-border technical solutions for enterprise cross-border technical personnel.

[0118] Example 2

[0119] Please refer to Figure 4 A cross-border data evaluation device 1 includes a memory 3, a processor 2, and a computer program stored in the memory 3 and executable on the processor 2. When the processor 2 executes the computer program, the steps in the above-mentioned embodiment 1 are implemented.

[0120] Since the systems / devices described in the above embodiments of the present invention are systems / devices used to implement the methods of the above embodiments of the present invention, those skilled in the art will be able to understand the specific structures and variations of these systems / devices based on the methods described in the above embodiments of the present invention, and thus will not be described in detail here. All systems / devices used in the methods of the above embodiments of the present invention are within the scope of protection of the present invention.

[0121] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0122] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions.

[0123] It should be noted that, in the claims, any reference signs placed between brackets shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention may be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In claims enumerating several means, several of these means may be embodied by one and the same hardware. The use of the words first, second, third etc. is for convenience only and does not indicate any order. These words may be understood as part of the component name.

[0124] In addition, it should be noted that, in the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.

[0125] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments after learning the basic creative concept. Therefore, the claims should be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0126] Obviously, those skilled in the art may make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if such modifications and variations fall within the scope of the claims and their equivalents, the present invention shall also include such modifications and variations.

Claims

1. A method for evaluating cross-border data, characterized in that: Including steps: S1. Collect multiple authorized and published patents associated with cross-border data, and establish a structured patent knowledge base based on the structured patent content of the published patents; S2. Optimizing a pre-trained patent generation model based on the structured patent content; S3. Obtain the technical solutions to be evaluated that are associated with the cross-border data, extract technical keywords and technical structure data from the technical solutions to be evaluated, input them into the patent knowledge base for similarity search, and obtain similar patent structure paragraphs; S4. The optimized patent generation model generates a technical text to be evaluated based on the technical keywords and the technical structure data, generates similar patent texts based on the technical keywords and the similar patent structure paragraphs, and generates a recommended application index for the technical solution to be evaluated based on the similarity scores between the technical text to be evaluated and the similar patent texts; The step S1 comprises the following steps: S11. Collect multiple authorized patent publications associated with cross-border data, and obtain patent keywords for the patent publications based on the reviewers' areas of interest when reviewing the patent publications; S12. Arrange the disclosed patent into patent structure paragraphs corresponding to the technical means and technical effects according to each claim in the claims, and obtain structured patent content including patent keywords and patent structure paragraphs. The technical means and technical effects respectively include the original content corresponding to the claims in the specification; S13. Establishing a structured patent knowledge base based on the structured patent content of the disclosed patent; The patent keywords include the content keywords of the paragraphs where the specification and claims are located, the abstract keywords of the paragraph where the specification abstract is located, and the figure keywords of the area where the specification figures are located. The patent generation model includes a text generator and a paragraph generator, and the step S2 includes the following steps: S21. Based on the structured patent content, the paragraph generator generates a patent comparison paragraph according to the content keywords of each paragraph, and the text generator generates a patent comparison text according to the patent comparison paragraph, abstract keywords, and figure keywords; S22. Input the average similarity between all the patent comparison paragraphs and the corresponding patent structure paragraphs in the structured patent content, as well as the average similarity between all the patent comparison texts and the corresponding published patents, into the loss function to update and optimize the text generator and the paragraph generator according to the loss function.

2. A cross-border data evaluation method according to claim 1, characterized in that: The structured patent content including patent keywords and patent structure paragraphs obtained in step S12 is: The drawing data including the drawing keywords, the abstract data including the original content of the specification abstract and the abstract keywords, and the structured patent content of the patent structure paragraph corresponding to each claim are obtained. The patent structure paragraph includes the corresponding technical means and technical effects. The technical means include the original technical content of the current claim, the corresponding original technical content in the specification, and the content keywords of the two parts of the original technical content. The technical effect includes the corresponding original effect content in the specification and the content keywords of the original effect content.

3. A cross-border data evaluation method according to claim 2, characterized in that: The step S11 includes the following steps: S111. Collect multiple authorized patents associated with cross-border data, and obtain the attention weight of each word in the current paragraph based on the area of ​​interest of the reviewer when reviewing the patent. If it is a figure in the specification of the patent, identify the text content in each figure, and treat the text content of each figure as a paragraph. The attention weight W eye The formula for (w) is: ; Where, TFD(w) is the fixation duration of word w, TFD max is the maximum fixation duration of a single word in the current paragraph, FC(w) is the number of fixations on word w, RC(w) is the number of returns to look at word w in the subsequent region, and RC avg is the average number of times all words in the current paragraph are looked back; S112. Obtain content keywords, abstract keywords, and figure keywords of the disclosed patent according to the location of the current paragraph and the attention weight of each word in the current paragraph.

4. A cross-border data evaluation method according to claim 3, characterized in that: The step S111 further includes the following steps: If the fixation duration of the sentence containing word w is greater than the average fixation duration of all sentences in the current paragraph by a preset multiple standard deviation, the fixation weight is accumulated by a first preset coefficient to obtain a comprehensive word weight, and both the preset multiple and the first preset coefficient are greater than 1; And / or if the number of times the sentence containing word w is looked back is greater than a preset proportion of the entire sentence in the patent disclosure, then the gaze weight is accumulated by a second preset coefficient to obtain a comprehensive word weight, the preset proportion is greater than 40%, and the second preset coefficient is greater than 1; And / or identifying the weight of each word in the patent disclosure according to the word weight model, and accumulating the weight with the attention weight to form a comprehensive word weight.

5. The cross-border data evaluation method according to claim 1, characterized in that: The loss function L is: L=α×L Paragraph +β×L Article ; Where α and β are hyperparameters, L Paragraph is the average similarity between the patent comparison paragraph and the corresponding patent structure paragraph, L Article It is the average similarity between the patent comparison text and the corresponding published patent.

6. The method for evaluating cross-border data according to claim 1, characterized in that: The step S4 comprises the following steps: S41, generating a technical text to be evaluated based on the technical keywords and the technical structure data by the optimized text generator and paragraph generator; S42, generating a similar patent text based on the technical keywords and the similar patent structure paragraphs by the optimized text generator; S43. Generate a recommended application index for the technical solution to be evaluated based on the similarity scores between the technical text to be evaluated and the similar patent text.

7. A cross-border data evaluation method according to claim 6, characterized in that: The step S43 includes the following steps: Generate an overall similarity S based on the similarity scores of the technical text to be evaluated and the similar patent text total ; Calculate the paragraph similarity (s1, s2, ..., s n ); According to the overall similarity S total and paragraph similarity (s1,s2,…,s n ) Generate a recommendation application index for the technical solution to be evaluated, wherein the formula for the recommendation application index InnovationScore is: InnovationScore=(1-S total )×(1-max(s1,s2,…,s n )); Where, max(s1,s2,…,s n ) is the paragraph similarity (s1, s2, ..., s n ) is the maximum value among .

8. A device for evaluating cross-border data, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements a method for evaluating cross-border data as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Program text generation method and system

    CN107133210A

  • Personalized text abstract generation method and system fusing eye movement data

    CN115098669A

  • Patent evaluation determination method, patent evaluation determination device, and patent evaluation determination program

    JP2020021455A