Method and System for Automatically Generating a Sci-Tech Novelty Search Report Based on a Database and a Large Model
Through the automatic generation method of scientific and technological new report based on database and large models, the problem of innovative evaluation results in the existing technology is solved, and efficient new search and report generation is achieved with simultaneous analysis of innovativeness of multiple modules, improving the comprehensiveness and accuracy of evaluation results.
Patent Information
- Application Number
- CN202510265824.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-07
AI Technical Summary
The existing method of generating new scientific and technological reports ignores the writing characteristics and content characteristics of different modules of the text, resulting in the one-sided results of innovative evaluations, affecting comprehensiveness and accuracy.
The automatic generation method of scientific and technological new search report based on database and large models is adopted. By inputting multiple content modules into the big model to extract keywords, combining the word expansion model and keyword comparison model, multiple result chains are generated, and a prompt word model is used to generate new search report.
The efficiency of new search is improved, and the innovation of each module is analyzed in multiple dimensions, the comprehensiveness and accuracy of evaluation results are enhanced, and the contents of various parts of the report are in an orderly manner. The generated reports follow scientific and technological writing standards, which improves quality and authority.
Smart Images

Figure CN119782507B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of scientific and technological novelty search, and particularly to a method and system for automatically generating scientific and technological novelty search reports based on a database and a large model. Background Art
[0002] Traditional scientific and technological novelty search report methods mainly rely on manual operations and have many limitations. Some current automatic scientific and technological novelty search methods use a single model for text innovation evaluation, without considering the writing characteristics and differences of different modules. During the text comparison process, the evaluation factors for innovation are single, which will result in a relatively one-sided evaluation result and affect the comprehensiveness and accuracy of the innovation evaluation result.
[0003] For example, Chinese Patent Application No. CN118427303A discloses a method, device, electronic device, and storage medium for generating a scientific and technological novelty search report, including: representing a text to be queried and a text to be matched through a trained large language model to determine a reference text related to the text to be queried in the text to be matched; extracting key information of the reference text through the trained large language model and the chain of thought multi-step prompting method according to the reference text; and generating a novelty search report for the text to be queried through the trained large language model and the chain of thought multi-step prompting method according to the text to be queried and the key information.
[0004] The above existing technologies have the problems raised in this background art: evaluating innovation from the overall text, ignoring the writing characteristics and content features of different modules of the text; during the text comparison process, the evaluation factors for innovation are single, which will result in a relatively one-sided evaluation result and affect the comprehensiveness and accuracy of the innovation evaluation result; to solve at least one of the above problems, the present invention proposes a method and system for automatically generating scientific and technological novelty search reports based on a database and a large model. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the main purpose of the present invention is to provide a method and system for automatically generating scientific and technological novelty search reports based on a database and a large model, which can effectively solve the problems in the background art. The specific technical solutions of the present invention are as follows:
[0006] A method for automatically generating a scientific and technological novelty search report based on a database and a large model includes:
[0007] Inputting multiple content modules into a corresponding large model respectively, where the corresponding large model outputs a keyword set of the content module, and the multiple content modules belong to a text to be searched;
[0008] Inputting all the keyword sets into a preset word expansion model, and the word expansion model outputs a corresponding expanded keyword set;
[0009] Based on all the extended keyword sets and a preset database, obtain the relevant technical texts for each content module;
[0010] Based on the relevant technical texts for each content module and a preset keyword comparison model, obtain the comparison results corresponding to each content module;
[0011] Combining the causal associations between each content module, use directed edges to connect the comparison results in each content module to obtain multiple result chains, where the directed edge points from the cause node to the result node;
[0012] Based on the multiple result chains, combine with a preset prompt word model to generate a novelty search report.
[0013] Specifically, input all the keyword sets into a preset word expansion model, and the word expansion model outputs the corresponding extended keyword sets. The word expansion model includes a synonym replacement plugin, a near-synonym rewriting plugin, and a word generalization plugin, including:
[0014] Through a preset synonym replacement plugin, perform synonym replacement on all keywords in the keyword set to obtain a synonym set;
[0015] Through a preset near-synonym rewriting plugin, perform near-synonym rewriting on all keywords in the keyword set to obtain a near-synonym set;
[0016] Through a preset word generalization plugin, perform word upper-level generalization on all keywords in the keyword set to obtain an upper-level word set;
[0017] Combine the keyword set, synonym set, near-synonym set, and upper-level word set to obtain an extended keyword set.
[0018] Specifically, the obtaining of the relevant technical texts for each content module based on all the extended keyword sets and a preset database includes:
[0019] Retrieve in the database according to the extended keyword set to obtain retrieval texts;
[0020] Use a preset evaluation model to screen the retrieval texts for each content module respectively to obtain the relevant technical texts for each content module.
[0021] Specifically, the using of a preset evaluation model to screen the retrieval texts for each content module respectively to obtain the relevant technical texts for each content module, where the evaluation model includes a similarity evaluation model and a matching degree evaluation model, including:
[0022] For each content module, respectively use a preset similarity evaluation model to perform similarity screening on the retrieved text to obtain similar texts for each module;
[0023] For each content module, respectively use a preset matching degree evaluation model to perform matching degree screening on the similar texts to obtain matching texts for each module;
[0024] Combine the matching texts of each module to obtain the related technical texts of each module.
[0025] Specifically, for each content module, respectively use a preset similarity evaluation model to perform similarity screening on the retrieved text to obtain similar texts for each module. Among them, the similarity evaluation model includes a text vectorization model and a similarity calculation plugin, including:
[0026] According to the keyword set, respectively vectorize the text to be searched and the retrieved text through a preset text vectorization model to obtain a text vector to be searched and multiple retrieved text vectors;
[0027] Through a preset similarity calculation plugin, respectively calculate the similarity between the text vector to be searched and each retrieved text vector to obtain multiple similarity values;
[0028] Combine the corresponding retrieved texts with similarity values greater than the preset similarity threshold to obtain similar texts.
[0029] Specifically, for each content module, respectively use a preset matching degree evaluation model to perform matching degree screening on the similar texts to obtain matching texts for each module. Among them, the matching degree evaluation model includes a graph construction model and a matching degree calculation plugin, including:
[0030] According to the keyword set, respectively construct graphs for the text to be searched and the similar texts through a preset graph construction model to obtain a graph of the text to be searched and multiple graphs of similar texts;
[0031] Through a preset matching degree calculation plugin, respectively calculate the matching degree between the graph of the text to be searched and each graph of similar texts to obtain multiple matching degree values;
[0032] Combine the corresponding similar texts with matching degree values greater than the preset matching degree threshold to obtain matching texts.
[0033] Specifically, according to the related technical texts of each content module and a preset keyword comparison model, obtain the comparison result corresponding to each content module. Among them, the content module includes a background content module, a solution content module, and an effect content module, including:
[0034] Extract keywords from relevant technical texts respectively to obtain a keyword set for each content module of the relevant technology. The keyword set includes keywords and keyword connection relationships;
[0035] Split the keywords and keyword connection relationships in the keyword set of each content module of the relevant technology and put them into the background library, solution library, and effect library respectively;
[0036] Use a preset keyword recombination model to recombine the keywords in the background library, solution library, and effect library to obtain multiple recombined texts;
[0037] Compare the text to be searched with each recombined text respectively through a preset similarity comparison model to obtain the comparison results of each content module.
[0038] Specifically, generate a novelty search report according to the multiple result chains in combination with a preset prompt word model, including:
[0039] Screen the multiple result chains through preset causal constraints to obtain the target result chain;
[0040] According to the target result chain, use a preset prompt word model to logically connect the comparison results of each module in the target result chain to obtain a combined result;
[0041] Extract the feature vector of the combined result through a preset feature vector extraction model to obtain the result feature vector;
[0042] Input the result feature vector into a preset text generation model to generate the report content;
[0043] Adjust the format of the report content through a preset format adjustment model to obtain the novelty search report.
[0044] A scientific and technological novelty search report automatic generation system based on a database and a large model, used to implement the scientific and technological novelty search report automatic generation method based on a database and a large model, includes:
[0045] A keyword extraction module that inputs multiple content modules into a corresponding large model respectively, where the corresponding large model outputs a keyword set of the content module, and the multiple content modules belong to a text to be searched;
[0046] A keyword expansion module that inputs all keyword sets into a preset word expansion model, and the word expansion model outputs a corresponding expanded keyword set;
[0047] A novelty search module that generates a novelty search report according to all expanded keyword sets and a preset database.
[0048] Compared with the prior art, the present invention has the following beneficial effects:
[0049] By dividing the text into modules and analyzing the innovation of multiple modules simultaneously, the novelty search efficiency can be improved. According to the writing characteristics and content features of each module, the innovation analysis of each module is carried out separately, avoiding inaccurate innovation results due to a single standard. Through similarity evaluation and matching degree evaluation, the recall rate and precision rate of the innovation evaluation of the report can be improved. Analyzing the innovation of each module from multiple dimensions can avoid the one-sidedness of single-factor evaluation and improve the comprehensiveness and accuracy of the innovation evaluation results. By automatically generating and formatting the report, the orderly connection of each part of the novelty search report can be ensured. Automatically generating replaces manual writing, greatly shortening the report generation time, and the generated content follows the scientific and technological writing norms, which can improve the quality and authority of the report. Brief Description of the Drawings
[0050] Figure 1 It is a flowchart of the method for automatically generating a scientific and technological novelty search report based on a database and a large model in Embodiment 1 of the present invention;
[0051] Figure 2 It is a schematic diagram of text module division in Embodiment 1 of the present invention;
[0052] Figure 3 It is a flowchart of the keyword comparison model in Embodiment 1 of the present invention;
[0053] Figure 4 It is a schematic diagram of the result combination model in Embodiment 1 of the present invention;
[0054] Figure 5 It is a schematic diagram of the structure of the system for automatically generating a scientific and technological novelty search report based on a database and a large model in Embodiment 2 of the present invention. Detailed Embodiments
[0055] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given in conjunction with the accompanying drawings of the specification.
[0056] Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0057] Second, the "one embodiment" or "embodiment" referred to herein means a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that are mutually exclusive of other embodiments.
[0058] Embodiment 1:
[0059] This embodiment provides an automatic generation method for scientific and technological novelty search reports based on a database and a large model. As Figure 1 shown, the automatic generation method for scientific and technological novelty search reports based on a database and a large model includes:
[0060] S101. Input each of a plurality of content modules into a corresponding large model, where the corresponding large model outputs a keyword set of the content module, and the plurality of content modules belong to a text to be searched;
[0061] S102. Input all the keyword sets into a preset word expansion model, and the word expansion model outputs a corresponding expanded keyword set;
[0062] S103. Obtain relevant technical texts of each content module according to all the expanded keyword sets and a preset database;
[0063] S104. Obtain a comparison result corresponding to each content module according to the relevant technical text of each content module and a preset keyword comparison model;
[0064] S105. Combine the causal associations between each content module, and use directed edges to connect the comparison results in each content module to obtain a plurality of result chains, where the directed edges point from the cause nodes to the result nodes;
[0065] S106. Generate a novelty search report according to the plurality of result chains and in combination with a preset prompt word model.
[0066] Since the current scientific and technological report content is written in blocks, including modules such as background technology, technical solutions, and technical effects, using a single innovation evaluation standard will ignore the writing differences between different modules and affect the accuracy of the innovation evaluation results. The present invention provides an automatic generation method for scientific and technological novelty search reports based on a database and a large model, which conducts innovation evaluation according to the report content in blocks, taking into account the writing habits of different modules and the degree of influence on innovation. As Figure 2 shown, the text is divided into three modules according to the title, and innovation evaluations are respectively conducted on the first module, the second module, and the third module.
[0067] In this embodiment, each text is divided into multiple sections. For example, the background technology section usually includes an elaboration of the industry status quo and an analysis of existing problems, and the solution content module will describe specific technical means and work processes. The focuses and writing methods of different modules are different. Therefore, according to the text content and title, the text to be searched is divided into multiple modules, and the innovation of each content module is evaluated separately. By dividing the modules and analyzing the innovation of multiple modules simultaneously, the novelty search efficiency can be improved. According to the writing characteristics and content features of each content module, the innovation analysis of each module is carried out separately to avoid inaccurate innovation results due to a single standard. The background section contributes less to the innovation evaluation of the report. When synthesizing the innovation evaluation results of each module, a smaller weight can be assigned to this section to achieve dynamic adjustment of the comprehensive innovation evaluation results and adapt to the needs of different text contents.
[0068] Specifically, the multiple content modules include a background content module, a solution content module, and an effect content module. Through the first large model, the problem keywords and the logical connection relationships between the problem keywords in the background content module are extracted to obtain a set of problem keywords. By deeply analyzing the relevant expressions of the core problems in the text of the background content module, sorting out the key problem keywords among them, and analyzing the logical connection relationships between the keywords, the core problems around which the background technology focuses can be accurately grasped. Through the second large model, the solution keywords and the logical connection relationships between the solution keywords in the solution content module are extracted to obtain a set of solution keywords. Identify the key expressions that can reflect specific technical means, technical implementation paths, etc., extract the solution keywords, and analyze the logical associations between them, so as to clearly present the core constituent elements and mutual relationships of the technical solution. Through the third large model, the effect keywords and the logical connection relationships between the effect keywords in the effect content module are extracted to obtain a set of effect keywords, obtain the effect keywords and their mutual logical connections, and clearly show the actual effect of the technology.
[0069] Specifically, by using a preset word expansion model, based on the existing keyword sets of problems, solutions, and effects, and leveraging mechanisms such as internal semantic association expansion, synonym replacement, and related concept extension within it, the original keywords are expanded to enrich the coverage of the keywords, thereby improving the comprehensiveness of subsequent retrievals. Using the expanded keyword set as the retrieval basis, a matching query is performed in a large amount of text in the database to find texts that contain these keywords or are semantically related to them, so as to obtain retrieval texts related to the text to be searched. Each content module's trained evaluation model is obtained by training with a large number of training texts of this module. This evaluation model is trained based on a large model. According to the extracted set of feature words, it is compared with the retrieved text, and evaluations are carried out using similarity and matching degree respectively. Through similarity evaluation, texts with a certain degree of relevance can be screened out, and through matching degree evaluation, highly relevant texts can be further screened out from the relevant texts. Highly relevant texts are of great value to the innovation of the evaluation report, which can improve the recall rate and precision rate of the innovation evaluation of the report. Analyzing the innovation of each content module from multiple dimensions can avoid the one-sidedness of single-factor evaluation and improve the comprehensiveness and accuracy of the innovation evaluation results.
[0070] According to the relevant technical texts of each content module and a preset keyword comparison model, the corresponding comparison results of each content module are obtained. In this embodiment, each content module includes multiple comparison results, and there is a causal relationship between the comparison results of different modules. The causal association between modules can be determined according to the theme and core content of each content module, which can be determined by referring to relevant literature, expert experience, and understanding of the business process, and a connection relationship between modules is constructed. For example, the technical problems in the background will guide the formulation of technical solutions. By constructing the connection relationship, a clear logical context can be formed among the various modules of the novelty search report, avoiding looseness and confusion of the content. Regarding the comparison results of each content module as nodes, according to the constructed connection relationship, a directed edge pointing from the cause node to the result node is set to connect the results in each content module, obtaining multiple result chains.
[0071] Specifically, by combining multiple result chains and a prompt model, a novelty search report is generated. After calculating the innovation evaluation results of each content module separately, based on the general logical structure and standard format of a scientific and technological novelty search report, the innovation evaluation results of each module are concatenated in an orderly manner. Through a preset text combination model, the relationships and coherence among the innovation evaluation results of each module are analyzed and combined into a relatively complete and coherent result. Since there may be some cases of incoherence or incorrect proper nouns among the directly combined results, a key feature vector extraction model is used to extract the key feature vectors in the concatenated report. Around the key feature vectors, a preset text generation model is used to generate the report content and adjust the format issues in the report content. By automatically generating and formatting the report, it is possible to ensure the orderly connection of each part of the novelty search report, automatically generate it to replace manual writing, greatly shorten the report generation time, and the generated content follows scientific and technological writing norms, which can improve the quality and authority of the report.
[0072] Through module division, the present invention can simultaneously analyze the innovation of multiple modules, improve the novelty search efficiency, and separately conduct the innovation analysis of each module according to the writing characteristics and content features of each module, avoiding inaccurate innovation results due to a single standard. Through similarity evaluation and matching degree evaluation, the recall rate and precision rate of the innovation evaluation of the report can be improved, the innovation of each module can be analyzed from multiple dimensions, avoiding the one-sidedness of single-factor evaluation, and improving the comprehensiveness and accuracy of the innovation evaluation results. By automatically generating and formatting the report, it is possible to ensure the orderly connection of each part of the novelty search report, automatically generate it to replace manual writing, greatly shorten the report generation time, and the generated content follows scientific and technological writing norms, which can improve the quality and authority of the report.
[0073] Further, all the keyword sets are input into a preset word expansion model, and the word expansion model outputs the corresponding expanded keyword sets. Among them, the word expansion model includes a synonym replacement plug-in, a near-synonym rewriting plug-in, and a word generalization plug-in, including:
[0074] S201. Through a preset synonym replacement plug-in, all the keywords in the keyword set are replaced with synonyms to obtain a synonym set;
[0075] S202. Through a preset near-synonym rewriting plug-in, all the keywords in the keyword set are rewritten with near-synonyms to obtain a near-synonym set;
[0076] S203. Through a preset word generalization plug-in, all the keywords in the keyword set are generalized to obtain a hypernym set;
[0077] S204. Combine the keyword set, synonym set, near-synonym set, and hypernym set to obtain an extended keyword set.
[0078] In this embodiment, the word expansion model includes a synonym replacement plugin, a near-synonym rewriting plugin, and a word generalization plugin, which perform synonym replacement, near-synonym rewriting, and word hypernym generalization on the keywords of each content module respectively. For example, if the keyword is "improve", the plugin will find synonyms such as "enhance" and "promote" in the dictionary to generate a synonym set; word hypernym generalization can analyze the position of each keyword in the lexical semantic network, determine its concept category, and then find a more general hypernym. For example, for the keywords "smartphone" and "tablet computer", their hypernym is "mobile terminal device". By synonym replacement, near-synonym rewriting, and word hypernym generalization, the retrieval range can be expanded, which helps to discover more extensive relevant literature and avoid missing searches. Combine the original problem keyword set, solution keyword set, and effect keyword set with the synonym set, near-synonym set, and hypernym set generated through the previous steps. The integrated keyword set makes full use of various lexical variants and concept expansions generated in the previous steps, and can comprehensively utilize these lexical information during subsequent database retrieval, performing retrieval from multiple angles and levels to maximize the possibility of retrieving relevant literature.
[0079] Further, obtaining the relevant technical texts of each content module according to all the extended keyword sets and a preset database includes:
[0080] S301. Retrieve in the database according to the extended keyword set to obtain retrieval texts.
[0081] S302. Use a preset evaluation model to screen the retrieval texts of each content module respectively to obtain the relevant technical texts of each content module.
[0082] Further, when using a preset evaluation model to screen the retrieval texts of each content module respectively to obtain the relevant technical texts of each content module, the evaluation model includes a similarity evaluation model and a matching degree evaluation model, and includes:
[0083] S401. For each content module, use the preset similarity evaluation model to perform similarity screening on the retrieval texts to obtain the similar texts of each module.
[0084] S402. For each content module, use the preset matching degree evaluation model to perform matching degree screening on the similar texts to obtain the matching texts of each module.
[0085] S403. Combine the matching texts of each module to obtain the relevant technical texts of each module.
[0086] In this embodiment, innovative evaluations are respectively carried out on each content module. The innovative evaluation is divided into similarity evaluation and matching degree evaluation. The evaluation model includes a similarity evaluation model and a matching degree evaluation model. For similarity evaluation, through a pre-trained similarity evaluation model, the similarity between the text to be searched and the database text is calculated. For example, if the similarity value between the text to be searched and a certain database text is 0.7, it indicates a high correlation between the two in terms of semantics and technical themes; if the similarity value with another database text is 0.3, then the contribution of this database text to evaluating the text to be searched is not significant. By calculating the similarity between the text to be searched and the database text, the degree of proximity between the text to be searched and the existing knowledge can be accurately displayed, breaking the ambiguity of traditional qualitative judgments, and providing an intuitive and reliable quantitative index for innovative definition; it can quickly locate the text similar to the text to be searched in a large number of database texts, improving the comparison efficiency.
[0087] After the similarity evaluation, for the database texts with high similarity, the matching degree between the text to be searched and the text with high similarity is further calculated to evaluate the overlap of the existing achievements between the text to be searched and the database text, and then the innovation of the text to be searched is judged. Through a pre-trained matching degree evaluation model, the matching degree between the text to be searched and the database text with high similarity is calculated. For example, if the matching degree value between the text to be searched and a database text is 0.8, it means a high degree of matching between the two in terms of technical architecture and function implementation, indicating that the technical solution of the text to be searched is very similar to the existing technology to a certain extent. Therefore, the innovation level of the technical solution of this text to be searched is not high. By calculating the matching degree, the matching degree between the technology of the text to be searched and the existing technology can be deeply analyzed from aspects such as technical architecture and technical effects, so as to accurately judge the technical innovation of the text to be searched. Screening out the matching texts highly relevant to the text to be searched from a large number of retrieved texts to obtain the relevant technical texts of each module, reducing the subsequent calculation amount and improving the efficiency of innovative analysis.
[0088] Furthermore, for each content module, a pre-set similarity evaluation model is respectively used to perform similarity screening on the retrieved text to obtain the similar texts of each module. Among them, the similarity evaluation model includes a text vectorization model and a similarity calculation plugin, including:
[0089] S501. According to the keyword set, through a pre-set text vectorization model, the text to be searched and the retrieved texts are respectively vectorized to obtain a text vector to be searched and multiple retrieved text vectors;
[0090] S502. Through a pre-set similarity calculation plugin, the similarity between the text vector to be searched and each retrieved text vector is respectively calculated to obtain multiple similarity values;
[0091] S503. Combine the corresponding retrieved texts with similarity values greater than the preset similarity threshold to obtain similar texts.
[0092] In this embodiment, first, the similarity between the text to be searched and the retrieved texts is evaluated. The similarity evaluation model includes a text vectorization model and a similarity calculation plugin. According to the keyword set, the texts are vectorized to facilitate the quantitative calculation of similarity. Through the pre-trained text vectorization model, the texts are transformed into numerical vector forms that can be efficiently processed by a computer. When the text to be searched and the retrieved texts are input, the model maps them into vectors of a fixed length respectively. These vectors can represent the semantic features of the texts in a high-dimensional space, and the vector dimension can be set according to actual needs. Through text vectorization, texts of different forms and lengths can be uniformly transformed into fixed-length vectors, which is convenient for calculating similarity.
[0093] Specifically, by calculating the similarity between vectors, the similarity degree between the corresponding texts is reflected. Using the preset similarity calculation plugin, the cosine similarity between vectors is calculated, and the value range is between -1 and 1. When the calculated cosine similarity value is larger, the corresponding texts are more similar. Through similarity calculation, the vague similarity concept between texts can be transformed into specific numerical values, so that the similarity degree between the text to be searched and the retrieved texts can be quantitatively analyzed. It can quickly calculate the pairwise similarity between a large number of database text vectors and the text vector to be searched, greatly improving the comparison efficiency. By setting the similarity threshold, texts highly relevant to the text to be searched are selected from numerous retrieved texts. The similarity threshold can be set according to actual calculation needs. The texts with similarity values greater than the similarity threshold are selected as similar texts. These texts have a high overlap in semantics with the innovative information and are used as key comparison texts for subsequent in-depth analysis, which helps to focus on key comparison materials and improve the novelty search efficiency.
[0094] Furthermore, for each content module, the preset matching degree evaluation model is respectively used to screen the matching degree of the similar texts to obtain the matching texts of each module. Among them, the matching degree evaluation model includes a graph construction model and a matching degree calculation plugin, including:
[0095] S601. According to the keyword set, through the preset graph construction model, the graph of the text to be searched and the similar texts are respectively constructed to obtain the graph of the text to be searched and multiple graphs of similar texts;
[0096] S602. Through the preset matching degree calculation plugin, calculate the matching degree between the graph of the text to be searched and each graph of similar texts respectively to obtain multiple matching degree values;
[0097] S603. Combine the corresponding similar texts with matching degree values greater than the preset matching degree threshold to obtain the matching texts.
[0098] In this embodiment, after obtaining the similar texts, the matching degree of the texts is further calculated. The matching degree evaluation model includes a graph construction model and a matching degree calculation plug-in. Through the preset graph construction model, the graph construction is carried out for the text to be searched and the similar texts respectively. The graph construction model identifies the professional terms and key concepts therein as nodes through natural language processing technology, and then determines the logical relationships between the nodes, such as causal, compositional, and application relationships as edges, to construct a graph reflecting the technical architecture and function implementation of the text; through graph construction, the disordered technical information of the text is transformed into a graph form with clear organization and strict structure, which is convenient for in-depth analysis of the innovation of the text.
[0099] Specifically, after constructing the graphs of the text to be searched and the similar texts, the matching degree between the graphs is quantitatively calculated. Through the pre-trained matching degree calculation plug-in, the degree of fit between the two graphs is quantified. The model calculates the matching degree based on a combination of the graph edit distance algorithm and the subgraph isomorphism algorithm to take into account both the overall structural similarity of the graph and the matching of local key technologies. For example, for the graph of the text to be searched of the smart home security system and the graph of the similar text in the constructed database, the model first calculates the preliminary edit distance value according to the graph edit distance algorithm, and then combines the subgraph isomorphism algorithm to judge the matching of the key subgraphs, and comprehensively obtains a matching degree value of 0.6. Repeat this operation for other similar text graphs in the database, and calculate the matching degree values with the graph of the text to be searched in turn; through the quantitative calculation of the matching degree, the matching degree of the innovation information and the existing knowledge can be quantified from the underlying details such as the technical architecture and the function implementation process, and the innovation uniqueness can be more accurately reflected.
[0100] After calculating the matching degree value, set a matching degree threshold, and select the database texts with the matching degree value greater than the matching degree threshold. For example, select 3 highly matching texts from many database texts, and use these texts as the matching texts. The matching texts cover the core technical solutions, key function implementations, etc. that are similar to the innovation points of the text to be searched, which is crucial for accurately judging the innovation.
[0101] Further, according to the relevant technical texts of each content module and the preset keyword comparison model, the comparison result corresponding to each content module is obtained, where the content module includes a background content module, a solution content module, and an effect content module, including:
[0102] S701. Extract keywords from the relevant technical texts respectively to obtain the keyword set of each content module of the relevant technology. The keyword set includes keywords and keyword connection relationships;
[0103] S702. Split the keywords and keyword connection relationships in the keyword set of each content module of the related technology, and put them into the background library, solution library, and effect library respectively;
[0104] S703. Use a preset keyword recombination model to recombine the keywords in the background library, solution library, and effect library to obtain multiple recombined texts;
[0105] S704. Through a preset similarity comparison model, compare the text to be searched with each recombined text respectively to obtain the comparison results of each content module.
[0106] In this embodiment, as Figure 3 , for the related technology texts obtained after screening, extract the problem keyword set, solution keyword set, and effect keyword set in each related technology text respectively. The keyword set includes keywords / sentences and the logical connection relationships (and, or, causality) between keywords / sentences. For example, for the solution content module, dig out the key expressions related to technical implementation means, process flows, etc. as the solution keywords and the corresponding logical connection relationships; split the keywords and connection relationships in the extracted related technology text keyword set to obtain independent keywords, and put them into the corresponding background library, solution library, and effect library. By splitting the keywords and their connection relationships, put the keywords related to technical problems into the background library, those related to technical solutions into the solution library, and those reflecting technical effects into the effect library. In this way, the information can be sorted according to different technical dimensions, facilitating subsequent recombination and comparison operations.
[0107] Specifically, there is a certain correlation between the solutions or effects of different technical documents. Randomly recombining the keywords of different technical documents will result in new technical documents. Through a preset keyword recombination model, the keywords stored in the background library, solution library, and effect library are recombined and matched respectively to generate multiple recombined texts. It will consider the reasonable semantic coherence and technical relevance between the keywords in different libraries, try to construct various text expressions, and simulate different technical scenarios and situations; by generating multiple recombined texts, more diverse text forms can be created to simulate various technical situations, increasing the number of samples for comparison with the text to be searched, so as to more comprehensively detect the similarity between the text to be searched and the existing related technology under different technical expressions and scenarios, and avoid missing important similarity situations due to a single comparison sample.
[0108] According to the reorganized text, use a preset similarity comparison model to analyze the similarity between the text to be searched and each reorganized text in terms of semantics, vocabulary usage, logical structure, etc., calculate the similarity degree value between them, set a threshold, and when the similarity degree value is greater than the threshold, the existing technology has a huge impact on the technical composition of the text to be searched. Compare the text to be searched with each reorganized text one by one, analyze the similarity between the text to be searched and the existing related technologies from multiple perspectives and various technical scenarios, respectively obtain the comparison results of each reorganized text in each content module, and comprehensively understand the innovation of the technology to be searched in the entire technical field, which helps to accurately generate a novelty search report.
[0109] Further, according to the multiple result chains, combined with a preset prompt word model, a novelty search report is generated, including:
[0110] S801. Through preset causal constraints, screen in the multiple result chains to obtain a target result chain;
[0111] S802. According to the target result chain, use a preset prompt word model to logically connect the comparison results of each module in the target result chain to obtain a combined result;
[0112] S803. Extract feature vectors from the combined result through a preset feature vector extraction model to obtain result feature vectors;
[0113] S804. Input the result feature vectors into a preset text generation model to generate report content;
[0114] S805. Through a preset format adjustment model, adjust the format of the report content to obtain a novelty search report.
[0115] In this embodiment, as Figure 4 , each content module includes multiple comparison results, and there is a causal relationship between the comparison results of different modules. The causal association between modules can be determined according to the theme and core content of each content module, and can be determined by referring to relevant literature, expert experience, and understanding of business processes, and a connection relationship between modules is constructed. For example, the technical problems in the background will guide the formulation of technical solutions; by constructing the connection relationship, a clear logical context can be formed between the various modules of the novelty search report, avoiding looseness and chaos of the content.
[0116] Specifically, consider the comparison results of each content module as nodes. According to the constructed connection relationships, set directed edges pointing from cause nodes to result nodes, and connect the results in each content module to obtain multiple result chains. Among them, the logical relationships and results between some result chains cannot comprehensively reflect the innovation of the text to be searched. Screen the multiple result chains, and select the most logical and representative result chain from them. According to the preset causal constraints, for example, stipulate that the causal relationship in the result chain has corresponding credibility; through the preset causal constraint rules, evaluate each result chain, and select one target result chain that meets all constraint conditions and is the most representative. By screening out the target result chain that conforms to the causal constraints, the logical rigor and accuracy of the novelty search report content can be ensured, and the inclusion of unreasonable or invalid information can be avoided.
[0117] Specifically, according to the target result chain and the logical relationships between the result chains, use the preset prompt model to generate logical keywords. Use the logical keywords to combine the comparison results in the target result chain, and insert the logical keywords into appropriate positions so that each result is combined into a complete sentence in logical order, which can transform the discrete module comparison results into coherent and smooth text content, enabling readers to clearly understand the logical derivation process of the report content.
[0118] In this embodiment, according to the comparison results between the text to be searched and each combined text analyzed, the novelty search report is written. The comparison result of each content module includes the comparison results with each combined text. In this embodiment, according to the causal relationship between the modules, the most logical and representative comparison results are selected from multiple comparison results, and the scattered comparison results are organized according to certain rules to form a coherent and logical overall result.
[0119] Furthermore, according to the overall result, through the preset text generation model, the report content is automatically generated. The text generation model is based on natural language generation (NLG) technology. By means of the language knowledge, expression paradigms, and logical structures accumulated during the pre-training stage on a large number of scientific and technological texts and novelty search report examples, it deeply understands the innovation associations of each module through the internal multi-head attention mechanism, imitates the professional writing style, and transforms the structured innovation key points into natural language text paragraphs that are fluent, professional, and in line with industry norms to generate the report content; through the automatic generation of the report content, the structured innovation evaluation can be quickly transformed into a complete report, and the generated text follows the scientific and technological writing norms, improving the quality and readability of the report.
[0120] Specifically, with the result feature vector as the core, a preset text generation model is used to automatically write the report content and generate the report content. The text generation model is based on natural language generation (NLG) technology. Through the internal multi-head attention mechanism, it comprehensively captures the innovative associations in the feature vector, selects accurate words and reasonable sentence patterns from a rich corpus, and transforms the structured feature vector into a coherent, professional, and highly readable natural language paragraph. The result feature vector is input into the text generation model, and at the same time, necessary metadata is attached according to the text attributes, such as the text name "Research and Development of New High-Energy-Density New Energy Vehicle Batteries", the field "Interdisciplinary Field of New Energy and Materials Science", and the novelty search requirement keywords "Application of High-Nickel Ternary Materials", etc., so that the generated content closely follows the theme, meets professional requirements, and enhances the professionalism and standardization of the content.
[0121] Specifically, there are some cases of content repetition or grammar errors in the directly generated report content, and the report content needs to be corrected and adjusted. Through a preset format adjustment model, the unsmooth and semantically repetitive positions in the report content are automatically identified, and their formats are adjusted or rewritten, so that the formatted novelty search report is reasonably typeset, standardized in format, and smooth in sentences, thereby increasing the standardization and readability of the novelty search report.
[0122] Further, adjusting the report content through the preset format adjustment model to obtain a novelty search report includes:
[0123] S901. According to the report content, perform a syntax structure analysis through a preset syntax analysis model to obtain a text syntax tree;
[0124] S902. According to the text syntax tree, detect the positions with incorrect report content formats through a preset format matching model to obtain text difference points;
[0125] S903. Through a preset language adjustment model, adjust the content of the text difference points in terms of language and format to obtain a novelty search report.
[0126] In this embodiment, there are cases where the directly generated report content has ungrammatical logic or repetitive content. The logical correction and adjustment of the report content are carried out to improve the readability and standardization of the novelty search report. First, the grammar analysis of the report content is performed. Through a preset grammar analysis model, the part-of-speech in the text content (such as nouns, verbs, adjectives, etc.) is identified, the sentence components (subject, predicate, object, attributive, adverbial, complement) are determined, and the dependency relationship between each component is established to construct a text grammar tree that reflects the grammar logic of the text. For example, for the sentence "The application of the new high-nickel ternary material in the battery significantly improves the energy density, which benefits from its unique crystal structure.", the model identifies "the new high-nickel ternary material" as a noun phrase serving as the subject, "the application in the battery" as a prepositional phrase serving as a postpositive attributive modifying the subject, "significantly improves" as a verb phrase serving as the predicate, "the energy density" as the object, and "which benefits from its unique crystal structure" as a reason adverbial clause, and then constructs a text grammar tree containing these nodes and dependency relationships. By constructing the text grammar tree, the internal logical organization of the text can be reflected, and when identifying format errors, it can accurately locate the positions at the sentence and phrase levels instead of blindly searching, improving the adjustment efficiency and accuracy.
[0127] Specifically, according to the text grammar tree, the incorrect positions in the text content are detected. Through a preset format matching model, based on a pre-set standard format template, the template covers detailed specifications such as font, font size, paragraph format (indentation, line spacing, paragraph spacing), numbering rules, title hierarchy styles, etc. The format matching model traverses each part of the report text according to each node of the text grammar tree, compares the actual text presentation style with the template requirements one by one, and uses a rule-based matching algorithm. Once it is found that the actual format does not conform to the standard, such as the repetition of the text content in the body paragraph, incorrect font size of the title, etc., the position is immediately marked as a text difference point. By comparing through the structural guidance of the grammar tree and the template, problems such as non-standard content in the text can be quickly and accurately identified.
[0128] Specifically, after identifying the text differences, the incorrect content is corrected. Through a preset language adjustment model, the content of the differences is adjusted in terms of language and format. The language adjustment model integrates rule-based text processing and machine learning-assisted optimization capabilities. For the text content involved in format differences, on the one hand, preset rule sets are used, such as term unification rules (ensuring consistent terminology for the same concept, e.g., unifying "artificial intelligence algorithm" and "AI algorithm"), typesetting rules (resetting paragraph indents, line spacing, etc. according to standards); on the other hand, machine learning components are introduced to judge the spelling correctness of professional terms based on a classifier and fine-tune the sentence smoothness using a language model. For example, inputting the text fragment of the difference point into the language model and using the model to predict a better expression way, such as optimizing the awkward sentence "By implementing a special processing technology on this material, the purpose of improving battery performance is achieved" to "By adopting a special processing technology on this material, the battery performance is improved", thus enhancing the readability of the text. By adjusting the language and format of the text content, the report can be made more professional and standardized, improving the readability and quality of the report.
[0129] Embodiment 2:
[0130] In this embodiment, as Figure 5 , a system for automatically generating a science and technology novelty search report based on a database and a large model is provided, which is used to implement the method for automatically generating a science and technology novelty search report based on a database and a large model, and includes:
[0131] A keyword extraction module inputs each of multiple content modules into a corresponding large model, where the corresponding large model outputs a keyword set of the content module, and the multiple content modules belong to a text to be searched;
[0132] A keyword expansion module inputs all the keyword sets into a preset word expansion model, and the word expansion model outputs a corresponding expanded keyword set;
[0133] A novelty search module generates a novelty search report according to all the expanded keyword sets and a preset database.
[0134] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for automatically generating a scientific and technological novelty search report based on a database and a large model, characterized in that: include: Inputting each of the plurality of content modules into a corresponding large model, wherein the corresponding large model outputs a keyword set of the content module, wherein the plurality of content modules belong to a text to be searched; Input all keyword sets into a preset word expansion model, and the word expansion model outputs the corresponding expanded keyword set; According to all the expanded keyword sets and the preset database, the relevant technical texts of each content module are obtained; According to the relevant technical text of each content module and the preset keyword comparison model, the comparison result corresponding to each content module is obtained; Combined with the causal relationship between each content module, the comparison results in each content module are connected by directed edges to obtain multiple result chains, wherein the nodes in the result chain are the comparison results in each content module, and the directed edges point from the cause node to the result node; According to the multiple result chains, combined with the preset prompt word model, a novelty search report is generated, including: By using preset causal constraints, the multiple result chains are screened to obtain a target result chain; According to the target result chain, using a preset prompt word model, logically connect the comparison results of each module in the target result chain to obtain a combined result; Extracting feature vectors from the combination result using a preset feature vector extraction model to obtain a result feature vector; Inputting the result feature vector into a preset text generation model to generate report content; The format of the report content is adjusted through a preset format adjustment model to obtain a novelty search report.
2. The method for automatically generating a scientific and technological novelty search report based on a database and a large model according to claim 1 is characterized in that: All keyword sets are input into a preset word expansion model, and the word expansion model outputs a corresponding expanded keyword set, wherein the word expansion model includes a synonym replacement plug-in, a synonym rewriting plug-in and a word summarization plug-in, including: Using a preset synonym replacement plug-in, all keywords in the keyword set are replaced with synonyms to obtain a synonym set; Through the preset synonym rewriting plug-in, all keywords in the keyword set are rewritten into synonyms to obtain a synonym set; Through the preset word summarization plug-in, all keywords in the keyword set are summarized into hypernyms to obtain a hypernym set; The keyword set, synonym set, near-synonym set and hypernym set are combined to obtain an expanded keyword set.
3. The method for automatically generating a scientific and technological novelty search report based on a database and a large model according to claim 1 is characterized in that: The relevant technical text of each content module is obtained according to all the expanded keyword sets and the preset database, including: Search in the database according to the expanded keyword set to obtain the search text; Using the preset evaluation model, the search text of each content module is screened respectively to obtain the relevant technical text of each content module.
4. The method for automatically generating a scientific and technological novelty search report based on a database and a large model according to claim 3 is characterized in that: The preset evaluation model is used to screen the search text of each content module to obtain the relevant technical text of each content module, wherein the evaluation model includes a similarity evaluation model and a matching evaluation model, including: For each content module, the preset similarity evaluation model is used to screen the search texts for similarity, and similar texts of each module are obtained; For each content module, a preset matching evaluation model is used to screen the similar texts for matching, so as to obtain matching texts for each module; The matching texts of the modules are combined to obtain relevant technical texts of the modules.
5. The method for automatically generating a scientific and technological novelty search report based on a database and a large model according to claim 4 is characterized in that: For each content module, a preset similarity evaluation model is used to perform similarity screening on the search text to obtain similar texts of each module, wherein the similarity evaluation model includes a text vectorization model and a similarity calculation plug-in, including: According to the keyword set, the to-be-queried text and the search text are respectively vectorized by a preset text vectorization model to obtain a to-be-queried text vector and multiple search text vectors; By using a preset similarity calculation plug-in, the similarity between the text vector to be searched and each search text vector is calculated respectively to obtain multiple similarity values; The corresponding search texts whose similarity values are greater than a preset similarity threshold are combined to obtain similar texts.
6. The method for automatically generating a scientific and technological novelty search report based on a database and a large model according to claim 4 is characterized in that: For each content module, a preset matching evaluation model is used to perform matching screening on the similar texts to obtain matching texts of each module, wherein the matching evaluation model includes a graph construction model and a matching calculation plug-in, including: According to the keyword set, the preset graph construction model is used to construct graphs for the text to be searched and similar texts respectively, so as to obtain a graph of the text to be searched and multiple graphs of similar texts; The matching degree between the to-be-queried text graph and each similar text graph is calculated respectively by a preset matching degree calculation plug-in to obtain a plurality of matching degree values; The corresponding similar texts whose matching values are greater than a preset matching threshold are combined to obtain a matching text.
7. The method for automatically generating a scientific and technological novelty search report based on a database and a large model according to claim 1 is characterized in that: The method obtains the comparison result corresponding to each content module according to the relevant technical text of each content module and the preset keyword comparison model, wherein the content module includes a background content module, a solution content module and an effect content module, including: Extract keywords from the relevant technical texts respectively to obtain a keyword set for each content module of the relevant technology, wherein the keyword set includes keywords and keyword connection relationships; Split the keywords and keyword connection relationships in the keyword set of each content module of the related technology, and put them into the background library, solution library and effect library respectively; Using the preset keyword reorganization model, the keywords in the background library, solution library and effect library are reorganized to obtain multiple reorganized texts; The text to be checked is compared with each reorganized text respectively through a preset similarity comparison model to obtain a comparison result of each content module.
8. The automatic generation system of scientific and technological novelty search reports based on database and large model is characterized by: A method for automatically generating a scientific and technological novelty search report based on a database and a large model as claimed in any one of claims 1 to 7, comprising: A keyword extraction module, which inputs each of the multiple content modules into a corresponding large model, wherein the corresponding large model outputs a keyword set of the content modules, wherein the multiple content modules belong to a text to be searched; The keyword expansion module inputs all keyword sets into a preset word expansion model, and the word expansion model outputs the corresponding expanded keyword set; The novelty search module generates novelty search reports based on all expanded keyword sets and preset databases.
Citation Information
Patent Citations
Patent literature similarity measurement method based on ontology
CN107247780A
Comprehensive retrieval method and system for patent and periodical literature
CN115794743A
Short text information retrieval system based on big data application
CN116881538A
Scientific and technological novelty search report generation method and device, electronic equipment and storage medium
CN118427303A
Generation method and device of patent technology route and computer equipment
CN118427356A