Water and electricity science and technology document duplicate checking method and device based on knowledge graph
Through the knowledge graph-based method, hydropower technology documents are cleaned and structured disassembled, and keyword extraction and content summary are used to use natural language processing algorithms and large language models to extract and content summary, the problems of low efficiency and low accuracy of hydropower technology documents in the existing technology are solved, and fast and efficient document weight checking is achieved.
Patent Information
- Application Number
- CN202510299440.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, the plagiarism checking method of hydropower technology documents has problems such as poor text similarity calculation effect, low efficiency, and slow calculation.
Using a knowledge graph-based method, data cleaning and structural disassembly are carried out by obtaining historical and technological documents, natural language processing algorithms and large language models are used to extract keywords and summarize contents, and document plagiarism-checking knowledge graphs are constructed.
It improves the efficiency and accuracy of information extraction, provides deeper content understanding and decision-making support, and achieves rapid and efficient scientific and technological documents checking in the water and power industry.
Smart Images

Figure CN120218206A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method and device for checking the duplication of hydropower science and technology documents based on a knowledge graph. Background Art
[0002] A knowledge graph is a structured semantic knowledge base that organizes entities and their relationships in the form of a graph and is widely used in fields such as search engines, recommendation systems, and natural language processing. The key technologies for constructing a knowledge graph include entity recognition, relationship extraction, knowledge fusion, and knowledge reasoning. These technologies are the basis for constructing a high-quality knowledge graph.
[0003] Knowledge graphs are usually stored in graph databases, which support efficient graph query and analysis operations. For example, Neo4j is a popular graph database system suitable for storing and querying knowledge graphs. Knowledge graphs need to be continuously updated to reflect changes in the real world. Automated update mechanisms and incremental learning algorithms are current research hotspots. In terms of applications, knowledge graphs have broad application prospects in professional fields such as healthcare, finance, and law. For example, in the healthcare field, knowledge graphs can help integrate and analyze medical data to support clinical decision-making. Existing document duplication checking methods have problems such as poor text similarity calculation results, low efficiency, and slow operation.
[0004] In summary, how to design a method for checking the duplication of hydropower science and technology documents based on a knowledge graph with high extraction efficiency and fast operation speed is an urgent problem to be solved at present. Summary of the Invention
[0005] This application aims to solve at least one of the technical problems in the related technologies to some extent.
[0006] To this end, the first object of this application is to propose a method for checking the duplication of hydropower science and technology documents based on a knowledge graph to solve the problems of large limitations of existing technical means, poor text similarity calculation results, low efficiency, and slow operation.
[0007] The second object of this application is to propose a device.
[0008] The third object of this application is to propose an electronic device.
[0009] The fourth object of this application is to propose a computer-readable storage medium.
[0010] To achieve the above object, the first aspect embodiment of this application proposes a method for checking the duplication of hydropower science and technology documents based on a knowledge graph, including:
[0011] Obtaining historical science and technology document information;
[0012] Read the historical scientific and technological document information content and convert the historical scientific and technological document into a text file;
[0013] Clean the data of the text file and perform annotation processing to obtain a preprocessed text;
[0014] Analyze and process based on the structure of the preprocessed text, use regular expressions to disassemble the structure, and obtain the disassembled text;
[0015] Use natural language processing algorithms and large language models to extract keywords and summarize the content of the disassembled text to obtain a knowledge graph dataset;
[0016] Input the knowledge graph dataset into a knowledge graph integration tool to obtain a document duplicate checking knowledge graph.
[0017] Preferably, the step of reading the historical scientific and technological document information content and converting the historical scientific and technological document into a text file includes:
[0018] Use the built-in algorithms of Python to read the historical scientific and technological document, traverse all block-level elements in the document. The block-level elements include paragraphs, tables, and pictures, parse the current block-level element type, and convert the parsed content into a text file.
[0019] Preferably, the step of cleaning the data of the text file and performing annotation processing to obtain a preprocessed text includes:
[0020] Construct a stop word dictionary, perform word segmentation on the text content extracted from each element block, remove the useless headers, footers, and symbols in the text file through regular expressions, and convert the text file into a specified format to obtain a preprocessed text.
[0021] Preferably, the step of analyzing and processing based on the structure of the preprocessed text, using regular expressions to disassemble the structure, and obtaining the disassembled text includes:
[0022] Preset regular expressions based on the structure of the preprocessed text, split the content of the preprocessed text into information text blocks based on the titles in the preprocessed text, and construct the disassembled text based on all the information text blocks.
[0023] Preferably, the step of using natural language processing algorithms and large language models to extract keywords and summarize the content of the disassembled text to obtain a knowledge graph dataset includes:
[0024] Analyze each information text block in the disassembled text, construct and train a large language model, and use the trained large language model to extract knowledge and summarize the content of the text of each text block to obtain a knowledge graph dataset.
[0025] Preferably, building and training the large language model includes: obtaining a training data set by using algorithms, annotation, and manual verification methods, and training the large language model with the training data set to obtain a trained large language model.
[0026] Preferably, inputting the knowledge graph data set into a knowledge graph integration tool to obtain a document duplicate check knowledge graph includes:
[0027] Summarizing the knowledge graph data set, converting it into a specified format file, and then inputting it into the knowledge graph integration tool to obtain a document duplicate check knowledge graph.
[0028] To achieve the above object, an embodiment of the second aspect of the present application proposes a hydroelectric science and technology document duplicate check device based on a knowledge graph, including:
[0029] A data acquisition module for acquiring historical science and technology document information;
[0030] A document conversion module for reading the content of the historical science and technology document information and converting the historical science and technology document into a text file;
[0031] A data cleaning module for cleaning the text file and performing annotation processing to obtain preprocessed text;
[0032] A disassembling module for analyzing and processing based on the structure of the preprocessed text, and using regular expressions to disassemble the structure to obtain disassembled text;
[0033] An extraction module for extracting keywords and summarizing the content of the disassembled text by using natural language processing algorithms and a large language model to obtain a knowledge graph data set;
[0034] A summarizing module for inputting the knowledge graph data set into a knowledge graph integration tool to obtain a document duplicate check knowledge graph.
[0035] To achieve the above object, an embodiment of the third aspect of the present application proposes an electronic device, including: a processor and a memory communicatively connected to the processor;
[0036] The memory stores computer execution instructions;
[0037] The processor executes the computer execution instructions stored in the memory to implement the method described in any one of the above.
[0038] To achieve the above object, an embodiment of the fourth aspect of the present application proposes a computer-readable storage medium, including computer execution instructions stored in the computer-readable storage medium, and the computer execution instructions are used to implement the method described in any one of the above when executed by a processor.
[0039] A method for detecting duplicate hydroelectric science and technology documents based on a knowledge graph provided by this application constructs an information extraction prompt engineering for large models, which can not only improve the efficiency and accuracy of information extraction, but also provide users with deeper content understanding and decision-making support. Automated information extraction can greatly improve the efficiency of processing text and save time. Through structured information extraction, it helps users better understand the core content of the text. The extracted key information and summary can provide a basis for decision-making, especially in fields such as policy formulation and market analysis. Through the extraction of keywords, it is more convenient to conduct information retrieval and knowledge sharing. By combining the characteristics of the document and designing appropriate prompt templates, the authenticity and effectiveness of the extracted information can be ensured. It realizes the detection of duplicate hydroelectric industry science and technology documents using the knowledge graph.
[0040] Additional aspects and advantages of this application will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of this application. Brief Description of the Drawings
[0041] The above and / or additional aspects and advantages of this application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0042] Figure 1 is a flowchart of the first specific embodiment of a method for detecting duplicate hydroelectric science and technology documents based on a knowledge graph provided by the present invention;
[0043] Figure 2 is a flowchart of another method for detecting duplicate hydroelectric science and technology documents based on a knowledge graph provided by the present invention;
[0044] Figure 3 is a structural block diagram of a device for detecting duplicate hydroelectric science and technology documents based on a knowledge graph provided by an embodiment of the present invention. Detailed Description of the Embodiments
[0045] The core of the present invention is to provide a method and device for detecting duplicate hydroelectric science and technology documents based on a knowledge graph, which constructs an information extraction prompt engineering for large models, can not only improve the efficiency and accuracy of information extraction, but also provide users with deeper content understanding and decision-making support.
[0046] In order to enable those skilled in the art of this technology to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the drawings and specific embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts belong to the scope of protection of the present invention.
[0047] Please refer to Figure 1 , Figure 1 which is a flowchart of the first specific embodiment of a method for checking the duplication of hydropower science and technology documents based on a knowledge graph provided by the present invention; the specific operation steps are as follows:
[0048] Step S101: Obtain historical science and technology document information;
[0049] Step S102: Read the content of the historical science and technology document information and convert the historical science and technology document into a text file;
[0050] Use the built-in algorithm of Python to read the historical science and technology document, traverse all block-level elements in the document, where the block-level elements include paragraphs, tables, and pictures, parse the current block-level element type, and convert the parsed content into a text file.
[0051] Step S103: Clean the data of the text file and perform annotation processing to obtain preprocessed text;
[0052] Construct a stop word dictionary, perform word segmentation on the text content extracted from each element block, and remove useless headers, footers, and symbols in the text file through regular expressions, and then convert the text file into a specified format to obtain preprocessed text.
[0053] Step S104: Analyze and process based on the structure of the preprocessed text, use regular expressions for structure decomposition, and obtain the decomposed text;
[0054] Preset regular expressions based on the structure of the preprocessed text, split the content of the preprocessed text into information text blocks based on the titles in the preprocessed text, and construct the decomposed text based on all information text blocks.
[0055] Step S105: Use natural language processing algorithms and large language models to extract keywords and summarize the content of the decomposed text to obtain a knowledge graph dataset;
[0056] Analyze each information text block in the decomposed text, construct a large language model and train it, and use the trained large language model to perform knowledge extraction and content summary on the text content of each text block to obtain a knowledge graph dataset.
[0057] In one embodiment, constructing a large language model and training it includes: using algorithms, annotation, and manual verification methods to obtain a training dataset, using the training dataset to train the large language model, and obtaining a trained large language model.
[0058] Step S106: Input the knowledge graph dataset into a knowledge graph integration tool to obtain a document duplication checking knowledge graph.
[0059] Summarize the knowledge graph dataset, convert it into a specified format file, and then input it into the knowledge graph integration tool to obtain a knowledge graph for document duplicate checking.
[0060] This embodiment provides a method for checking the duplication of hydropower science and technology documents based on a knowledge graph, constructing an information extraction prompt project for large models, which can not only improve the efficiency and accuracy of information extraction, but also provide users with deeper content understanding and decision-making support. Automated information extraction can greatly improve the efficiency of processing text and save time. Through structured information extraction, it helps users better understand the core content of the text. The extracted key information and summary can provide a basis for decision-making, especially in fields such as policy formulation and market analysis. Through the extraction of keywords, it is more convenient to conduct information retrieval and knowledge sharing. By combining the characteristics of the document and designing a suitable prompt template, the authenticity and effectiveness of the extracted information can be ensured. It realizes the duplication checking of science and technology documents in the hydropower industry using a knowledge graph.
[0061] Based on the above embodiment, this embodiment describes the method for checking the duplication of hydropower science and technology documents based on a knowledge graph, as Figure 2 shown below:
[0062] S1: Organize the original materials
[0063] Organize historical science and technology documents, all of which are word documents and belong to unstructured text materials. The key information in the application cannot be directly extracted, and text preprocessing and document disassembly of the application are required.
[0064] S2: Read the data
[0065] Use the built-in Python algorithm Python-docx to read the content of the docx document. The specific operation is to use the self-written "docx2txt" algorithm to convert the read docx file into a txt text file.
[0066] Specifically, it includes:
[0067] The implementation steps of the docx2txt algorithm are as follows: 1. Install the python-docx toolkit and ensure that the python version is consistent with the python-docx version; 2. Use python-docx to read the document at the specified path doc = docx.Document(docx_path); 3. Traverse all block-level elements in the doc document (including paragraphs, tables, and images), for child in doc.element.body.iterchildren(); 4. Develop a block-level element recognition tool, including image, table, and paragraph recognition. The python-docx library represents Word documents using XML elements. By parsing XML, the type of the current element block is identified respectively; 5. Traverse all element blocks in the document, identify all element block types, and extract the content for each element block type. The extracted content is saved in a python list for further processing.
[0068] S3: Data cleaning
[0069] The extracted txt file contains useless content such as headers and footers, special symbols such as commas and periods, and stop words such as "xxx is crucial", etc. The process of removing these contents is called "data cleaning". After cleaning, corresponding more concise text data can be obtained.
[0070] The implementation of this process depends on the support of a self-written algorithm, which contains regular expressions, etc., as follows:
[0071] Build a stop word dictionary, perform word segmentation on the text content extracted from each element block, and use the stop word dictionary for filtering during the word segmentation process to remove stop words irrelevant to the text theme. Use regular expressions to build a special symbol representation: r'[,。,。?!;:"”‘’‘’@#$%^&*()_+{}|<>\[\]\\-=`~\'“”‘’“”…]', and match and clean special characters through regular expressions.
[0072] S4: Standardization
[0073] After S3 data cleaning, the text data format is GBK or Unicode format, and it is changed to the commonly used utf-8 format through a format conversion tool.
[0074] S5: Structural disassembly
[0075] Technical documents contain different contents separated by headings such as "Project background" and "Project content", and will also separately identify tables in the document such as "Project personnel information table" and "Data statistics tables in each chapter", etc.
[0076] In this step, the text is mainly disassembled into information text blocks by identifying the obvious "structural segmentation" headings mentioned above. At the same time, table information will be extracted separately.
[0077] Here is a supplementary description of the representative detailed self-written algorithm code and writing logic:
[0078] Use regular expressions and the xml structure information of the document for disassembling, analyze the structural characteristics of the document, design regular expressions for document structure markers, and identify the start and end positions of text blocks respectively. Based on the xml style information in text blocks without special marks, divide the text, and finally obtain text blocks with complete semantic content. Table information is read through the xml style and the read table content is saved.
[0079] S6: Content Disassembly and Analysis
[0080] After the structural disassembly in S5, there are still cases of mixed content in each information text block and table. For example, in the table, there are different information such as personnel names, personnel titles, contact information, etc. In the project background text block, not only background information is included, but also some technical route descriptions exist.
[0081] Facing the above situation, it is necessary to disassemble and analyze the content of text blocks and tables. In this step, natural language processing algorithms and large language model technologies are used to extract knowledge from the text content after structural disassembly, and realize content summary and keyword extraction. Both content summary and keyword extraction are based on algorithms. Especially in the keyword extraction link, this patent independently researches and develops a data annotation tool suitable for algorithm training, which can upload the text content of the application form, and obtain the data for algorithm model training through the method of algorithm pre-annotation + manual verification. The annotation tool supports uploading of multiple types of documents, including various data formats such as pictures, texts, and audios.
[0082] Here is a supplementary description of the representative detailed self-written algorithm code and writing logic:
[0083] Content extraction and keyword extraction adopt a large model generative algorithm. By constructing a prompt engineering for extracting information from the large model, semantic understanding of the input Chinese text block is carried out, and then content summary information and keyword information are generated. The construction of the project needs to be combined with the characteristics of the document to ensure the authenticity of the extracted information, and extraction examples are added to ensure the information extraction effect of the large model.
[0084] S7: Formatting
[0085] Summarize the basic information of the application form and the knowledge graph dataset obtained through all the above steps into files in formats such as xlxs, json, and dict, and input them into the knowledge graph integration tool to obtain the final knowledge graph that can assist in checking the duplication of scientific and technological projects for hydropower projects.
[0086] A method for checking the duplication of hydropower science and technology documents based on a knowledge graph provided by an embodiment of the present invention constructs an information extraction prompt project for a large model, which can not only improve the efficiency and accuracy of information extraction, but also provide users with deeper content understanding and decision-making support. Automated information extraction can greatly improve the efficiency of processing text and save time. Through structured information extraction, it helps users better understand the core content of the text. The extracted key information and summary can provide a basis for decision-making, especially in fields such as policy formulation and market analysis. Through the extraction of keywords, it is more convenient to conduct information retrieval and knowledge sharing. By combining the characteristics of the document and designing appropriate prompt templates, the authenticity and effectiveness of the extracted information can be ensured. It realizes the duplication checking of hydropower industry science and technology documents using the knowledge graph.
[0087] Please refer to Figure 3 , Figure 3 is the structural block diagram of a device for checking the duplication of hydropower science and technology documents provided by an embodiment of the present invention; the specific device may include:
[0088] A data acquisition module 100 to acquire historical science and technology document information;
[0089] A document conversion module 200 to read the content of the historical science and technology document information and convert the historical science and technology document into a text file;
[0090] A data cleaning module 300 to clean the text file and perform annotation processing to obtain a preprocessed text;
[0091] A disassembling module 400 to analyze and process based on the structure of the preprocessed text, and use regular expressions to disassemble the structure to obtain the disassembled text;
[0092] An extraction module 500 to extract keywords and summarize the content of the disassembled text using natural language processing algorithms and large language models to obtain a knowledge graph dataset;
[0093] A summarizing module 600 to input the knowledge graph dataset into a knowledge graph integration tool to obtain a document duplication checking knowledge graph.
[0094] A hydroelectric power technology document duplicate checking device based on a knowledge graph in this embodiment is used to implement the aforementioned hydroelectric power technology document duplicate checking method based on a knowledge graph. Therefore, the specific implementation manners in a hydroelectric power technology document duplicate checking device based on a knowledge graph can be seen in the embodiment part of the aforementioned hydroelectric power technology document duplicate checking method based on a knowledge graph. For example, a data acquisition module 100, a document conversion module 200, a data cleaning module 300, a disassembling module 400, an extraction module 500, and a summarizing module 600 are respectively used to implement steps S101, S102, S103, S104, S105, and S106 in the aforementioned hydroelectric power technology document duplicate checking method based on a knowledge graph. Therefore, the specific implementation manners can refer to the descriptions of the corresponding various part embodiments and will not be elaborated herein.
[0095] To implement the above embodiments, the present application also proposes an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided in the foregoing embodiments.
[0096] To implement the above embodiments, the present application also proposes a computer-readable storage medium storing computer-executable instructions, and the computer-executable instructions are used to implement the method provided in the foregoing embodiments when executed by a processor.
[0097] To implement the above embodiments, the present application also proposes a computer program product including a computer program, and the computer program implements the method provided in the foregoing embodiments when executed by a processor.
[0098] The collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved in the present application all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0099] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of these legal uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the users, including but not limited to notifying the users to read the user agreement / user notice and signing an agreement / authorization including authorizing the relevant user information before the users use this function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.
[0100] This application is expected to provide an implementation scheme for users to selectively block the use or access of personal information data. That is, the present disclosure is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of users.
[0101] In the description of the foregoing embodiments, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0102] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0103] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0104] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by an instruction execution system, apparatus, or device or in combination with these instruction execution systems, apparatuses, or devices. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.
[0105] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well-known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0106] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0107] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0108] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.
Claims
1. A method for checking duplicate content of hydropower technology documents based on knowledge graph, characterized in that: include: Obtain historical scientific and technological document information; Read the information content of the historical scientific and technological documents and convert the historical scientific and technological documents into text files; Performing data cleaning and annotation processing on the text file to obtain a preprocessed text; Analyze and process the structure of the preprocessed text, and use regular expressions to decompose the structure to obtain the decomposed text; Using a natural language processing algorithm and a large language model to extract keywords and summarize the content of the disassembled text to obtain a knowledge graph dataset; The knowledge graph dataset is input into the knowledge graph integration tool to obtain the document duplication checking knowledge graph.
2. The method for checking duplicate hydropower technology documents based on knowledge graph according to claim 1 is characterized in that: The reading of the information content of the historical scientific and technological document and converting the historical scientific and technological document into a text file comprises: The historical scientific and technological document is read using Python's built-in algorithm, all block-level elements in the document are traversed, including paragraphs, tables, and pictures, and the current block-level element type is parsed, and the parsed content is converted into a text file.
3. The method for checking duplicate hydropower technology documents based on knowledge graph according to claim 1 is characterized in that: The data cleaning and annotation processing of the text file to obtain the preprocessed text includes: A stop word dictionary is constructed, and the text content extracted from each element block is segmented. After the useless headers, footers and symbols in the text file are removed through regular expressions, the text file is converted into a specified format to obtain a preprocessed text.
4. The method for checking duplicate hydropower technology documents based on knowledge graph according to claim 1 is characterized in that: The analysis and processing based on the structure of the preprocessed text, using regular expressions to decompose the structure, and obtaining the decomposed text includes: A regular expression is preset based on the structure of the preprocessed text, the preprocessed text content is decomposed into information text blocks based on the titles in the preprocessed text, and the decomposed text is constructed based on all the information text blocks.
5. The method for checking duplicate hydropower technology documents based on knowledge graph according to claim 4 is characterized in that: The natural language processing algorithm and the large language model are used to extract keywords and summarize the content of the disassembled text to obtain a knowledge graph dataset including: Each information text block in the disassembled text is analyzed, a large language model is constructed and trained, and the trained large language model is used to extract knowledge and summarize the text content of each text block to obtain a knowledge graph data set.
6. The method for checking duplicate hydropower technology documents based on knowledge graph according to claim 5 is characterized in that: The constructing and training of a large language model includes: obtaining a training data set by using an algorithm, a labeling method, and a manual verification method, and training the large language model by using the training data set to obtain a trained large language model.
7. The method for checking duplicate hydropower technology documents based on knowledge graph according to claim 1 is characterized in that: The step of inputting the knowledge graph dataset into the knowledge graph integration tool to obtain the document duplicate checking knowledge graph comprises: The knowledge graph dataset is aggregated, converted into a file in a specified format, and then input into a knowledge graph integration tool to obtain a document duplication checking knowledge graph.
8. A hydropower technology document duplication checking device based on knowledge graph, characterized in that: include: Data acquisition module, to obtain historical scientific and technological document information; A document conversion module, which reads the information content of the historical scientific and technological document and converts the historical scientific and technological document into a text file; A data cleaning module performs data cleaning on the text file and performs annotation processing to obtain a preprocessed text; A disassembly module performs analysis and processing based on the structure of the preprocessed text, uses regular expressions to perform structural disassembly, and obtains the disassembled text; An extraction module uses a natural language processing algorithm and a large language model to extract keywords and summarize the content of the disassembled text to obtain a knowledge graph data set; The summary module inputs the knowledge graph dataset into the knowledge graph integration tool to obtain the document duplication checking knowledge graph.
9. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.