Document duplicate checking method and computing equipment
By combining the surface and deep feature plagiarism checking methods, using the knowledge graph to analyze key information in the document, the problems of low efficiency of plagiarism checking and in-depth semantic understanding are solved, and efficient and accurate document plagiarism checking is achieved.
Patent Information
- Application Number
- CN202510398555.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-08-26
AI Technical Summary
The existing document diversion checking methods are inefficient and cannot deeply understand the text semantics, resulting in high false positive rates.
Combining surface feature comparison and deep feature mutation check, a knowledge graph is constructed by extracting key information, and a joint analysis of text semantic features and deep features of the graph are carried out.
It improves the efficiency and accuracy of document plagiarism checking, reduces the false positive rate, and allows a deeper understanding of the text semantics in the document.
Smart Images

Figure CN120542406A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a document duplication checking method and computing device. Background Art
[0002] With the rapid development of information technology, document generation and dissemination are becoming increasingly common. Traditional methods for checking for duplicate content have numerous limitations. For example, they focus solely on surface-level similarity, resulting in low duplication detection efficiency and an inability to effectively understand complex semantics.
[0003] Therefore, it is particularly important to develop an efficient and accurate document duplication checking method. Summary of the Invention
[0004] The embodiments of the present application provide a document duplication checking method and computing device, which combine surface feature comparison and deep feature duplication checking to achieve joint analysis of text semantic features and deep graph features, thereby performing efficient and accurate document duplication checking and reducing the false alarm rate.
[0005] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:
[0006] In a first aspect, a document duplication checking method is provided, the method comprising:
[0007] Determining a target text based on a first document to be checked for duplicates; the target text is used to represent key information in the first document;
[0008] Determining a second document that satisfies a similarity condition with the target text from a plurality of preset documents;
[0009] Performing knowledge graph matching based on a first knowledge graph corresponding to the first document and a second knowledge graph corresponding to the second document to obtain a first similarity feature between the first document and the second document;
[0010] Determining semantically similar features between the first document and the second document as second similar features between the first document and the second document;
[0011] Based on the first similarity feature and the second similarity feature, a duplicate checking result for the first document is determined.
[0012] As can be seen, the target text representing key information in the first document is first extracted, and similar text is quickly matched based on the target text. The semantic similarity features of the first document and similar documents are then compared, and deep duplication detection is performed in combination with knowledge graph comparison. This combines surface feature comparison with deep feature duplication detection, achieving a joint analysis of text semantic features and deep graph features, leading to efficient and accurate document duplication detection and reducing false positive rates.
[0013] In a possible implementation, determining the target text based on the first document to be checked for duplicates includes:
[0014] Get the original text under the specified directory node of the first document;
[0015] The original text and the first prompt word are input into the first model, and the text output by the first model is obtained as the target text; wherein the first model is used to extract key information from the original text according to the first prompt word.
[0016] As can be seen, using advanced large-scale model technology to conduct in-depth analysis of the original text under a specified directory node and extract key features and key content helps further focus on important information points, laying a solid foundation for subsequent vectorization processing and knowledge graph construction. Therefore, for document duplication checks based on fixed templates, this helps improve the efficiency and accuracy of duplicate checks, effectively reducing the error parsing rate, and further facilitating a deeper understanding of the textual semantics within the document.
[0017] In a possible implementation, the designated directory node includes a first directory node, and the target text includes a first target text extracted from an original text under the first directory node;
[0018] The method also includes: constructing a first knowledge graph based on the first target text.
[0019] Constructing a first knowledge graph according to the first target text, including:
[0020] Inputting the first target text and the second prompt word into the second model to obtain triple information output by the second model;
[0021] Constructing a first knowledge graph based on triple information;
[0022] Among them, the second model is used to determine the entities, entity relationships and entity attribute information in the first target text according to the instructions of the second prompt word to output triple information; the first similarity feature includes triple information whose similarity between the first knowledge graph and the second knowledge graph meets the conditions.
[0023] It can be seen that in the embodiment of the present application, a directory node suitable for checking for duplicate content by means of knowledge graph comparison can be pre-specified, i.e., a first directory node. The first directory node can be selected according to the requirements and targets of checking for duplicate content. During the document checking for duplicate content, a knowledge graph is constructed based only on the key information extracted under the first directory node, and then a knowledge graph comparison method is adopted to check for duplicate content. For other directory nodes that are not suitable for checking for duplicate content by knowledge graph comparison, other methods of checking for duplicate content can be adopted. Thus, processing the texts under different directory nodes by multiple checking engines can help provide a multi-level perspective of checking for duplicate content, and carry out refined checking for duplicate content analysis on documents from different angles. This helps to improve the efficiency and accuracy of checking for duplicate content, and to gain an in-depth understanding of the text semantics in the document. For the checking for duplicate content of fixed-structure documents, it is more capable of accurate positioning and refined processing by chapter, which greatly improves the efficiency and accuracy of checking for duplicate content.
[0024] In one possible implementation, the first target text includes a plurality of subtexts corresponding to different content themes; the first knowledge graph includes a first subgraph created based on each subtext; and the second knowledge graph includes a plurality of second subgraphs created based on texts of different content themes.
[0025] Performing knowledge graph matching based on a first knowledge graph corresponding to the first document and a second knowledge graph corresponding to the second document to obtain a first similarity feature between the first document and the second document includes:
[0026] Determine a first sub-graph and a second sub-graph that match the content theme;
[0027] Determine the graph similarity between the first sub-graph and the second sub-graph that have matching content themes, and obtain a preset number of groups of sub-graphs with the highest graph similarity;
[0028] The third prompt word and each group of sub-graphs are input into the third model in sequence, and similar features between each group of sub-graphs output by the third model are obtained as first similar features; wherein the third model is used to determine similar content between each group of sub-graphs according to the indication of the third prompt word.
[0029] It can be seen that in the embodiment of the present application, a more refined division is performed based on the first target text, sub-texts of different content themes are determined, and a first sub-graph is constructed based on the sub-texts. By analyzing the graph similarity between the first sub-graph and the second sub-graph that match the content themes, multiple groups of sub-graphs with higher graph similarity are screened out. This removes content themes with lower similarity and reduces the burden of duplicate checking on the large model. Moreover, by further analyzing the similarity between sub-graphs with higher similarity through the large model, a refined and in-depth graph analysis can be provided, which helps to improve the precision and accuracy of duplicate checking.
[0030] In a possible implementation, the target text further includes a second target text extracted from the text under the second directory node specified by the first document.
[0031] In one possible implementation, determining semantically similar features between the first document and the second document as second similar features between the first document and the second document includes:
[0032] Determining a third directory node belonging to a basic section in the first document;
[0033] The fourth prompt word, the text under each third directory node in the first document, and the text under the corresponding directory node in the second document are input into the fourth model to obtain a second similarity feature; the fourth model is used to determine the semantically similar features between the input text of the first document and the input text of the second document based on the indication of the fourth prompt word.
[0034] As can be seen, each directory node within the basic section of the first document is compared for similar semantic features based on the large semantic model. This combines semantic similarity features with deep graph features for comprehensive duplication detection. This helps provide a multi-layered duplication detection perspective, enabling refined duplication analysis of documents from different angles. This improves duplication detection efficiency and accuracy, and provides a deeper understanding of the textual semantics within the document.
[0035] In a possible implementation, the key information represents text information related to the duplicate checking target;
[0036] When the target of duplicate checking is a construction project, the entities of the first knowledge graph include at least one of the construction function point name, construction location, technical indicators, and implementation stage; the attribute information of the first knowledge graph includes the construction number and / or construction funds of each construction function point.
[0037] As can be seen, when the duplicate check target is a construction project, by extracting knowledge graphs from the text, it is possible to extract various entities such as the construction function point name, construction location, technical indicators, and implementation phase, and identify attribute information such as the number of construction function points and construction funding. In the subsequent comparison stage, it is possible to accurately analyze the relevant information related to the construction function points, thereby detecting potential duplicate construction projects and avoiding the waste of resources caused by duplicate construction.
[0038] In one possible implementation, determining a second document that satisfies a similarity condition with the target text from a plurality of preset documents includes:
[0039] Convert the target text into a target text vector;
[0040] Retrieving candidate vectors that meet similarity conditions from a vector library; the vector library includes text vectors converted based on key information of each preset document;
[0041] A second document is determined from the preset documents corresponding to the candidate vector.
[0042] As can be seen, the embodiments of this application, by pre-establishing a vector library, quickly locate similar texts, laying the foundation for subsequent semantic feature duplication checking and deep graph feature duplication checking, facilitating efficient and accurate duplication checking. Furthermore, the system supports the addition of a large number of documents from different fields to expand the size of the vector library and text library to accommodate diverse application scenarios.
[0043] In the second aspect, a document duplication checking device is provided, which includes: a functional unit for executing any one of the methods provided in the first aspect, and the actions performed by each functional unit are implemented by hardware or by hardware executing corresponding software. For example, the document duplication checking device may include: a first determination unit for determining a target text based on a first document to be checked for duplication; the target text is used to represent key information in the first document; a second determination unit for determining a second document that meets a similarity condition with the target text from a plurality of preset documents; a matching unit for performing knowledge graph matching based on a first knowledge graph corresponding to the first document and a second knowledge graph corresponding to the second document to obtain a first similarity feature between the first document and the second document; a third determination unit for determining a semantically similar feature between the first document and the second document as a second similarity feature between the first document and the second document; and a fourth determination unit for determining a duplication checking result for the first document based on the first similarity feature and the second similarity feature.
[0044] In a third aspect, a computing device is provided, comprising: a processor and a memory; the processor is coupled to the memory; the memory is used for computer program instructions; and the processor is used to call the computer program instructions in the memory so that the computing device executes any one of the methods provided in the first aspect or the second aspect.
[0045] In a fourth aspect, a computer-readable storage medium is provided, which stores computer execution instructions. When the computer execution instructions are executed on a computing device, the computing device executes any one of the methods provided in the first aspect.
[0046] In a fifth aspect, a computer program product is provided, comprising: computer execution instructions, which, when executed on a computing device, cause the computing device to execute any one of the methods provided in the first aspect.
[0047] Among them, the technical effects brought about by any implementation method in the second to fifth aspects can refer to the technical effects brought about by different implementation methods in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A schematic diagram of the system architecture of an exemplary application scenario of the document duplication checking method provided in an embodiment of the present application;
[0049] Figure 2 A schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0050] Figure 3 A flowchart of a document duplication checking method provided in an embodiment of the present application;
[0051] Figure 4 A schematic diagram of a process for comparing knowledge graph similarities provided in an embodiment of the present application;
[0052] Figure 5 A schematic diagram of a process for determining similar documents provided in an embodiment of the present application;
[0053] Figure 6 A schematic diagram of a document duplication checking method provided in an embodiment of the present application;
[0054] Figure 7 A structural diagram of a document duplication checking device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0056] In the description of this application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship, for example, A / B can represent A or B; "and / or" in this application is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural.
[0057] Furthermore, in the description of this application, unless otherwise specified, "plurality" means two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.
[0058] In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit differences. At the same time, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way for easy understanding.
[0059] The following is a brief introduction to the relevant terms involved in the embodiments of this application.
[0060] Artificial intelligence (AI) is a discipline that studies and develops theories, methods, techniques, and application systems for simulating, extending, and expanding human intelligence. Currently, AI technology has been widely used due to its high degree of automation, high precision, and low cost.
[0061] A knowledge graph (KG) is a semantic network used to describe the relationships between entities. It is a representation method for semi-structured data, used to describe entities, attributes, and relationships between entities. The core idea of a knowledge graph is to transform real-world information into a graph, where nodes represent entities and edges represent relationships between entities. A knowledge graph is not just a graphical knowledge base; it also contains semantic descriptions of entities and relationships that can be understood and processed by computers.
[0062] Deep learning (DL) is a new research direction in machine learning (ML). It learns the inherent patterns and representational hierarchies of sample data. The information gained from this learning process is highly helpful for interpreting data such as text, images, and sound. Its ultimate goal is to enable machines to acquire human-like analytical and learning capabilities, enabling them to recognize text, images, and sound. Specifically, this research focuses on neural network systems based on convolutional operations, known as convolutional neural networks; autoencoder neural networks based on multi-layer neurons; and deep belief networks, which use multi-layer autoencoder neural networks for pre-training and then incorporate discriminant information to further optimize the neural network weights. Deep learning has achieved significant results in search technology, data mining, machine learning, machine translation, natural language processing, multimedia learning, speech, recommendation and personalization technologies, and other related fields.
[0063] Natural language processing (NLP): NLP is a branch of artificial intelligence that focuses on the interaction between computers and human natural language, aiming to enable computers to understand and generate human language. NLP involves processing text data, speech recognition, machine translation, sentiment analysis, question-answering systems, and other aspects.
[0064] Large language models (LLMs): Deep learning models with powerful natural language processing capabilities that can understand and generate complex natural language text. Examples include bidirectional encoder representations from transformers (BERT), autoregressive blank-filling general language model pretraining (GLM), and generative pre-trained transformer (GPT).
[0065] The following is an illustrative introduction to the application scenarios of the embodiments of the present application.
[0066] For example, the document duplication checking method provided in the embodiments of the present application can be applied to application scenarios such as duplication checking of academic works, duplication checking of internal enterprise documents (such as technical documents, research reports and training materials), and duplication checking of policy documents issued by the government.
[0067] For example, checking for plagiarism in academic works can detect plagiarism and duplication. Checking for plagiarism in internal corporate documents can prevent duplicate submissions and ensure their originality and consistency. Checking for plagiarism in government-issued policy documents can ensure their originality and authority and prevent duplicate content. Furthermore, checking for plagiarism in construction project documents can identify potential duplicate projects and avoid the resulting waste of resources.
[0068] Figure 1 A schematic diagram of the system architecture of an exemplary application scenario of the document duplication checking method provided in an embodiment of the present application.
[0069] like Figure 1 As shown, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables. The terminal device 101 includes but is not limited to desktop computers, portable computers, smart phones, tablet computers, etc. It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be configured as needed. For example, server 103 may be a server cluster consisting of multiple servers.
[0070] As a possible application scenario of the embodiment of the present application, the document duplication checking method is executed by the server 103. The document duplication checking method provided by the embodiment of the present application requires the use of a model, so the model can also be pre-deployed in the server 103. Thus, the user accesses the server 103 through the terminal device 101, and the user can interact through the interactive interface displayed on the terminal device 101. The terminal device 101 sends the information input by the user to the server 103, and the server 103 executes the method provided by the embodiment of the present application to check for document duplication, and feeds back the results to the terminal device 101, which then displays the duplication checking results.
[0071] As another possible application scenario of the embodiment of the present application, the document duplication checking method is executed by the terminal 101. The model used in the duplication checking process can also be pre-deployed in the terminal device 101. The user can interact through the interactive interface displayed on the terminal device 101. The terminal device 101 executes the document duplication checking method provided in the embodiment of the present application based on the document content input by the user and displays the duplication checking results.
[0072] Currently, methods for checking for duplicate content in documents have low efficiency and a lack of in-depth understanding of semantics.
[0073] In view of this, the present invention provides a document duplication detection method. The inventive concept is to extract key information from the document to be detected, retrieve similar documents based on this key information, and then deeply analyze the similar features between the document to be detected and similar documents by matching knowledge graphs. This method can solve the technical problem of the limited semantic understanding of existing duplication detection methods while achieving high duplication detection efficiency.
[0074] The following is an exemplary introduction to the system architecture of the embodiment of the present application.
[0075] Figure 2 A schematic diagram of the structure of a computing device provided in an embodiment of the present application.
[0076] Need to explain, Figure 2 The system architecture shown is merely an example and does not constitute a limitation on the system architecture of the computing device provided in the embodiments of the present application.
[0077] In the embodiments of the present application, the computing device may specifically be a network device. The network device may include a server, etc. The server may be a single physical server, or two or more physical servers sharing different responsibilities and cooperating to implement various server functions.
[0078] For example, the server may be a blade server, a high-density server, a rack server, or a tower server, etc. The terminal device may include a personal digital assistant (PDA), an ultra-mobile personal computer (UMPC), a notebook computer, a netbook, a desktop computer, an all-in-one computer, etc.
[0079] The hardware of the computing device includes a processor, a basic input / output system (BIOS) chip, an out-of-band controller, and memory, while the software mainly includes the BIOS, an out-of-band management module, and an operating system (OS). Figure 2 shown.
[0080] A processor may include a central processing unit (CPU), which includes one or more CPU cores. The CPU's data processing operations are all performed by the CPU cores. The more CPU cores a CPU includes, the faster it can process data.
[0081] The BIOS chip is a chip installed on the motherboard that initializes and detects various hardware components during the computer's startup process. The BIOS chip includes a flash memory area.
[0082] The out-of-band management module is located in the out-of-band controller, and the operating system is located in the processor.
[0083] An out-of-band management module can be a management unit for non-business modules. For example, an out-of-band management module can remotely maintain and manage a computing device through a dedicated data channel. This out-of-band management module is completely independent of the computing device's operating system and can communicate with the BIOS and operating system through the computing device's out-of-band management interface.
[0084] Exemplarily, the out-of-band management module may include a management unit for the computing device's operating status, a management system in a management chip, a baseboard management controller (BMC) for the computing device, a system management module (SMM), etc. It should be noted that the embodiments of the present application do not limit the specific form of the out-of-band management module, and the above description is merely exemplary.
[0085] The operating system (OS) is a computer program that manages and controls the hardware and software resources of a computing device. Any other software must be supported by the OS to run. After a computing device is powered on, the BIOS first performs a series of operations, including self-tests and initialization, and then boots the OS, allowing the user to use the computing device normally.
[0086] BIOS is a set of programs embedded in the BIOS chip on the motherboard of a computing device. The main function of BIOS is to provide the lowest-level and most direct hardware settings and control for the computing device.
[0087] Memory, also known as internal storage or main memory, is installed in memory slots on the motherboard of a computing device.
[0088] It should be noted that the system architecture and application scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0089] For ease of understanding, the document duplication checking method provided in the embodiment of the present application is exemplarily introduced below in combination with the above system architecture and accompanying drawings.
[0090] Figure 3The flowchart of a document duplicate checking method according to an exemplary embodiment is shown. Exemplarily, the document duplicate checking method can be applied to a computing device, and the method steps include the following S301-S304.
[0091] S301: Determine a target text based on a first document to be checked for duplicates; the target text is used to represent key information in the first document.
[0092] In the embodiment of the present application, the first document may be an academic paper, a project document published by an enterprise, training materials, a technical report, or a policy document issued by the government, etc., without specific limitation.
[0093] In this step, a target text is determined from the first document, and the target text is used to represent key information in the first document. As an example, the target text can be text related to the duplicate check target of this document duplicate check.
[0094] Among them, the duplicate checking target can be understood as the information that may constitute duplication and is of greatest concern in this document duplicate checking, and may be related to the type of document.
[0095] For example, documents related to construction projects often contain information about planned construction projects, such as plans to build "xx" charging stations in "xx" area. The purpose of checking for duplicate content in such documents is to identify similar construction plans that appear in similar documents, thereby detecting potential duplicate construction projects and avoiding the waste of resources caused by duplicate construction. The corresponding duplication check target could be the construction project, meaning that the content in the document related to the construction project is considered key information.
[0096] As a possible implementation of the embodiment of the present application, key information can be extracted from the first document using a large language model, where a large language model refers to a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of spoken text.
[0097] Specifically, when using large language models, prompt engineering techniques can be used. Prompt engineering techniques involve providing appropriate guidance and instructions to help the model generate the desired output.
[0098] Specifically, a corresponding prompt word can be input into the large language model. The prompt word can be a vocabulary word or a short sentence. The prompt word is used to instruct the large language model to perform the corresponding task according to the prompt word.
[0099] For example, when using a large language model to extract key text from a first document, a corresponding prompt word may be input, such as "extract key sentences from the document." Under the guidance of the prompt word, the large language model understands the text semantics in the document and outputs the key text.
[0100] As a possible implementation method of an embodiment of the present application, considering that the key information in a document is usually concentrated in a specific chapter, the specified directory node can be located according to the directory structure of the first document, and then the key text can be extracted based on the original text under the specified directory node.
[0101] The above step of determining the target text based on the first document to be checked for duplicates may specifically include obtaining the original text under a specified directory node of the first document, inputting the original text and the first prompt word into the first model, and obtaining the text output by the first model as the target text. The first model is configured to extract key information from the original text based on the first prompt word.
[0102] Specifically, if the first document is a fixed template document and its directory structure matches the preset directory structure, the directory structure of the first document can be directly identified to locate the original text under the specified directory node. The large speech model is then used to refine the original text under the specified directory node and extract key text.
[0103] For example, if the first document to be checked for duplicate content has a fixed template, sequentially containing directory nodes such as "Project Background," "Construction Objectives," "Construction Content Introduction," "Facility Unit," and "Construction Function Module," one or more of these directory nodes can be pre-determined as designated directory nodes, such as "Construction Content Introduction" and "Construction Function Module." By identifying the directory structure of the first document and locating these two directory nodes, only the original text under the designated directory nodes will be considered when extracting key text.
[0104] It can be seen that for documents with fixed templates, designated directory nodes containing key information of the document can be pre-specified, which is equivalent to establishing a document structure fingerprint. In the process of document duplication checking, since the document template is fixed, the original text under the designated directory node can be quickly located to extract key information. In addition, advanced large model technology is used to conduct in-depth analysis of the original text under the designated directory node to extract key features and key content, which helps to further focus on important information points and lay a solid foundation for subsequent vectorization processing and knowledge graph construction. Therefore, for document duplication checking with fixed templates, it helps to improve the efficiency and accuracy of duplication checking and effectively reduce the error parsing rate. It also helps to deeply understand the text semantics in the document.
[0105] S302: Determine a second document from a plurality of preset documents that satisfies a similarity condition with the target text.
[0106] In the embodiment of the present application, the preset document may be a document in a document library, and the document library may be pre-established.
[0107] To facilitate understanding, the establishment of the document library and vector library is explained.
[0108] The preset documents are the documents selected for storage, and can be selected based on specific needs. In the embodiment of the present application, a large number of documents from different fields can be added to expand the size of the vector library and text library to adapt to different application scenarios, such as academic paper duplication detection, legal document comparison, and corporate document management.
[0109] Alternatively, in an embodiment of the present application, for a single application scenario or technical field, documents related to the application scenario or technical field may be added to make the established document library more professional.
[0110] For each preset document, key information is extracted from the document and converted into a text vector. The step of extracting key information can be seen in step S301. Alternatively, a large language model can be used to extract key information from the preset document, or a specific directory node in the preset document can be first located and key information can be extracted based on the original text within the specified directory node.
[0111] Subsequently, an embedding model can be used to convert the key information of the preset document into a vector form. Embedding models can include, but are not limited to, the Word2Vec model and the bidirectional encoder representations from transformers (BERT) embedding model. Embedding models are used to convert text into numerical vector representations, allowing similarity calculations to be performed on the text in a mathematical space.
[0112] Subsequently, the vectors corresponding to the key information of the preset document are stored in a vector database, such as the Faiss vector database and the Annoy vector database, to support efficient similarity search.
[0113] In addition, you can also store preset documents in a document library, such as a MySQL database, to preserve the complete document content, making it easier to find the original text and conduct detailed duplicate checks.
[0114] In the embodiment of the present application, after determining the key information of the first document, a second document with a higher semantic similarity can be searched from the document library based on the key information.
[0115] Exemplarily, the key information of the first document is converted into a vector form, a vector with high similarity is matched in a vector library, and then the original document is searched from the document library based on the matched vector as a second document with high semantic similarity.
[0116] The above examples are only possible implementations. For other specific implementations of determining the second document, please refer to the following. Figure 5 The embodiment shown.
[0117] S303: Perform knowledge graph matching based on the first knowledge graph corresponding to the first document and the second knowledge graph corresponding to the second document to obtain a first similarity feature between the first document and the second document.
[0118] In an embodiment of the present application, a first knowledge graph can be constructed for a specified portion of the first document.
[0119] The designated portion may be a portion of the target text, the entire target text, or other text in the first document except the target text, which is not limited in this embodiment of the present application.
[0120] As a possible implementation of an embodiment of the present application, after locating a directory node for a first document, the first directory node can be located from a designated directory node. The original text under the designated directory node is used to extract key information from the first document. The first directory node is contained within the designated directory node, and accordingly, the target text includes the first target text extracted from the original text under the first directory node.
[0121] As a possible implementation method of an embodiment of the present application, a first knowledge graph can be constructed based on the first target text, that is, the first knowledge graph can be constructed based on the key text extracted from the first directory node.
[0122] Exemplarily, when the target of document duplication check is a construction project, the designated directory nodes may include a directory node for describing the construction project profile and a directory node for describing at least one construction function point in the construction project. Considering that these two directory nodes contain information for describing the construction project, both are used as designated directory nodes to extract target text representing key information. Among them, the construction function point can be understood as a more detailed construction entity under the construction project. The construction entity can be a physical construction content, such as "charging piles", "parking spaces", "cameras", etc., or it can be a non-physical construction content, such as "appointment registration platform", "medical reimbursement platform", etc. A construction project mentioned in a construction target document usually contains multiple construction function points.
[0123] For the construction of knowledge graphs, the directory node used to describe the introduction of the construction project only contains the introduction of the construction project, and it is difficult to extract an effective knowledge graph. Therefore, only the directory node used to describe at least one construction function point in the construction project can be used as the first directory node, and then the first knowledge graph can be constructed based on the first target text extracted from the original text under the first directory node.
[0124] It can be seen that in the embodiment of the present application, a directory node suitable for checking for duplicate content by means of knowledge graph comparison can be pre-specified, i.e., a first directory node. The first directory node can be selected according to the requirements and targets of checking for duplicate content. During the document checking for duplicate content, a knowledge graph is constructed based only on the key information extracted under the first directory node, and then a knowledge graph comparison method is adopted to check for duplicate content. For other directory nodes that are not suitable for checking for duplicate content by knowledge graph comparison, other methods of checking for duplicate content can be adopted. Thus, processing the texts under different directory nodes by multiple checking engines can help provide a multi-level perspective of checking for duplicate content, and carry out refined checking for duplicate content analysis on documents from different angles. This helps to improve the efficiency and accuracy of checking for duplicate content, and to gain an in-depth understanding of the text semantics in the document. For the checking for duplicate content of fixed-structure documents, it is more capable of accurate positioning and refined processing by chapter, which greatly improves the efficiency and accuracy of checking for duplicate content.
[0125] As a possible implementation of an embodiment of the present application, constructing a first knowledge graph based on a first target text may specifically include: inputting the first target text and a second prompt word into a second model to obtain triple information output by the second model; and constructing the first knowledge graph based on the triple information. The second model is configured to determine entities, entity relationships, and entity attribute information in the first target text based on the indication of the second prompt word, and output triple information.
[0126] For example, the second prompt word can be "build a knowledge graph based on the input text and output triple information of the knowledge graph." Under the guidance of the prompt word, the large language model analyzes the text semantics and extracts entities, entity attribute information, and relationships between entities.
[0127] Entities are basic units, and relationships represent the associations between entities, which can be direct or indirect.
[0128] As a possible implementation method of an embodiment of the present application, the key information represents text information related to the duplicate checking target; when the duplicate checking target is a construction project, the entities of the first knowledge graph include at least one of the construction function point name, construction location, technical indicators, and implementation stage; the attribute information of the first knowledge graph includes the construction number and / or construction funds of each construction function point.
[0129] For example, in the case where the above-mentioned construction function point is a charging pile, the entities may include "charging pile", "location area", the relationship between "charging pile" and "location area" may be "located in", and the attribute information of "charging pile" may include quantity, price, etc.
[0130] In an embodiment of the present application, for preset documents in the document library, a knowledge graph is also pre-constructed based on the text under the same directory node and stored in the knowledge graph library corresponding to the document library.
[0131] Furthermore, knowledge graph matching can be performed based on the first knowledge graph corresponding to the first document and the second knowledge graph corresponding to the second document to determine the first similarity feature.
[0132] When performing knowledge graph matching, the comprehensive similarity between the first knowledge graph and the second knowledge graph can be compared. The comprehensive similarity includes at least one of entity similarity, relationship similarity, attribute information similarity, and topological structure similarity.
[0133] As an example, the matching degree between knowledge graphs can be analyzed based on a related algorithm, such as the GraphSAGE algorithm. The first knowledge graph and the second knowledge graph are input into the related algorithm, and the algorithm can output the similarity between the two in terms of triples and the similarity in terms of overall topology.
[0134] In an embodiment of the present application, the first similarity feature between the first document and the second document may include: triple information whose similarity between the first knowledge graph and the second knowledge graph meets the conditions obtained by comparing the first knowledge graph and the second knowledge graph.
[0135] For example, the similarity condition may be that the similarity evaluation value is greater than a preset threshold, wherein the similarity evaluation value may be obtained by weighted evaluation based on entity similarity, association relationship similarity, and attribute information similarity, and the weighting coefficient may be predetermined.
[0136] S304: Determine semantically similar features between the first document and the second document as second similar features between the first document and the second document.
[0137] Specifically, the first similarity feature is determined based on the knowledge graph, which belongs to the deep graph check. In order to further improve the accuracy and completeness of the check, the shallow semantic check and the deep graph check can be combined.
[0138] The semantically similar features between the first document and the second document may be determined by chapters.
[0139] As mentioned above, we check for duplicate content using knowledge graph comparison for the original text under the first directory node. We can then identify semantically similar features for other directory nodes.
[0140] As a possible implementation method of an embodiment of the present application, the above-mentioned determination of semantically similar features between the first document and the second document may specifically include: for other directory nodes in the first document except the first directory node, sequentially determining the similar semantic features of the first document and the second document under the matching directory nodes.
[0141] As a possible implementation method of an embodiment of the present application, the above-mentioned determination of semantically similar features between the first document and the second document may specifically include: determining the third directory node belonging to the basic section in the first document; inputting the fourth prompt word, the text under each third directory node in the first document, and the text under the corresponding directory node in the second document into the fourth model to obtain the second similar feature; the fourth model is used to determine the semantically similar features between the input text of the first document and the input text of the second document according to the indication of the fourth prompt word.
[0142] Specifically, the first document may include a basic section, and the definition of the basic section may be set as required. Exemplarily, the basic section includes other directory nodes in addition to the above-mentioned first directory node, or the basic section may include other directory nodes in addition to the above-mentioned first directory node and second directory node. In an embodiment of the present application, the directory node included in the basic section may be recorded as a third directory node, and the basic section may include multiple third directory nodes.
[0143] For example, if the first document to be checked for duplicates sequentially contains directory nodes such as "Project Background," "Construction Objectives," "Construction Content Introduction," "Facility Unit," and "Construction Function Module," among which "Construction Function Module" is the first directory node, it can be determined that the directory nodes "Project Background," "Construction Objectives," "Construction Content Introduction," and "Facility Unit" other than the first directory node are all third directory nodes of the basic section. Then, a comparison of similar semantic features is performed for each third directory node in turn.
[0144] Specifically, the key text of the original text of the first document in the third directory node and the key text of the original text of the second document in the corresponding directory node can be extracted respectively, and then input into the large language model respectively, along with the prompt words, to instruct the large language model to evaluate the similar semantic features of the third directory node.
[0145] The large language model can evaluate the similarity overview, similarity score, and original text comparison results of the similar parts under each third directory node between the first document and the second document.
[0146] For example, when using a large language model, corresponding prompt words can be input to instruct the large language model to identify similar parts and / or the degree of similarity. For example, the prompt words can be "identify plagiarized content" or "identify similar content." The large language model identifies similar parts and / or the degree of similarity under the guidance of the prompt words.
[0147] As can be seen, each directory node within the basic section of the first document is compared for similar semantic features based on the large semantic model. This combines semantic similarity features with deep graph features for comprehensive duplication detection. This helps provide a multi-layered duplication detection perspective, enabling refined duplication analysis of documents from different angles. This improves duplication detection efficiency and accuracy, and provides a deeper understanding of the textual semantics within the document.
[0148] S305: Determine a duplicate checking result for the first document based on the first similarity feature and the second similarity feature.
[0149] As a possible implementation method of an embodiment of the present application, the first similarity feature is obtained by comparing the text under the first directory node in the first document with the text under the corresponding directory node in the second document through a knowledge graph, and the second similarity feature is obtained by comparing the semantic similarity features of the text under the third directory node in the first document with the text under the corresponding directory node in the second document. The first similarity feature and the second similarity feature are aggregated to obtain the final duplicate check result.
[0150] Furthermore, in order to facilitate user viewing, a duplicate checking report can be generated in a preset format, and the duplicate checking report includes the above-mentioned duplicate checking results, thereby providing users with clear and detailed duplicate checking results.
[0151] It can be seen that in the embodiment of the present application, the powerful text comprehension ability of the large language model is used to analyze the similar parts and similarity levels of the first document and the second document under a specific directory node. When the large language model analyzes similar texts, it will automatically analyze them in combination with the understanding of the context, thereby improving the understanding of the text semantics and the judgment of similarity, improving the accuracy of the duplicate checking results, and reducing the probability of missed detection and false detection. In addition, for the specified directory node, deep duplicate checking is achieved based on the comparison of the knowledge graph. Thus, the semantic feature comparison and the graph feature comparison are combined to realize the joint analysis of the text semantic features and the deep features of the graph, so as to perform efficient and accurate document duplicate checking and reduce the false alarm rate.
[0152] As a possible implementation method of an embodiment of the present application, the duplication check report includes one or more of the following: an overview of similar content between the first document and the second document, an overall similarity score, a text comparison result of similar content under the matching directory nodes, and a similarity score of the content under the matching directory nodes; wherein the overall similarity score is obtained by performing a weighted operation based on the similarity scores under each matching directory node.
[0153] Specifically, for the first directory node, a knowledge graph comparison method is used to determine the original text comparison results and similarity scores for the similar content of the first and second documents under that directory node. For each directory node included in the basic section, semantic analysis can be performed based on the large model to obtain the original text comparison results and similarity scores for the similar content of the first and second documents under that directory node.
[0154] The large language model can also generate an overview of similar content based on similar situations under different directory nodes, making it easier for users to quickly perceive similar content.
[0155] In an embodiment of the present application, the weighted weight corresponding to each directory node can be pre-set, and the similarity scores under each directory node are weighted according to the weighted weight corresponding to each directory node to obtain an overall similarity score.
[0156] For example, for documents related to a construction project, the directory nodes include: "Project Background," "Construction Objectives," "Construction Content Introduction," "Facility Unit," and "Construction Function Module." For more important directory nodes, such as "Construction Function Module" and "Construction Content Introduction," a larger weight can be preset.
[0157] By integrating and analyzing the similarities of each directory node, the overall similarity is evaluated, an overview of similar content is obtained, and an overall similarity score is given.
[0158] Furthermore, in order to facilitate user viewing, a duplicate checking report can be generated in a preset format, and the duplicate checking report includes the above-mentioned duplicate checking results, thereby providing users with clear and detailed duplicate checking results.
[0159] For example, the duplicate check report first displays the overall similarity score and an overview of similar content. It then displays the original text comparison results and similarity scores for each directory node.
[0160] As a possible implementation method of the present application, a report template can be customized based on prompt engineering technology. Specifically, a report can be generated by a large language model. After the text to be compared belonging to the first document and the second document under each directory node is input into the large language model, a specific prompt word is input to instruct the large language model to integrate the overall duplicate checking results and generate a duplicate checking report according to the report template.
[0161] Among them, the report template can be related to the user role. For example, different report templates can be pre-customized according to the differentiated needs of auditors, operators and other personnel. The weighting coefficients and layout methods of each directory node in different report templates may be different.
[0162] It can be seen that in the embodiment of the present application, multiple duplicate checking engines are used to process the texts under different directory nodes, combining semantic feature comparison and deep graph feature comparison, and integrating the text comparison results and deep graph comparison results to generate a multi-dimensional evidence chain, which is presented in the form of a duplicate checking report. The duplicate checking report can include duplicate checking comparison results in multiple dimensions and can support the customization of different duplicate checking report templates to meet the differentiated needs of different user roles and reduce the cost of understanding the duplicate checking report.
[0163] In an embodiment of the present application, the first target text may include multiple sub-texts corresponding to different content topics. Accordingly, the first knowledge graph includes a first sub-graph created based on each sub-text, and the second knowledge graph includes multiple second sub-graphs created based on texts of different content topics.
[0164] The content theme can be understood as a detailed classification entity under the duplicate checking target. For example, for documents related to construction projects, the duplicate checking target is the construction project, then the content theme can be used to characterize the content related to the construction function point, that is, the text related to each construction function point can be understood as the text of a single content theme. For example, the text related to "charging piles", that is, the text with the content theme of charging piles, is used to describe the area, location, price and other information of the charging piles planned to be built. The text related to "parking spaces", that is, the text with the content theme of parking spaces.
[0165] In an embodiment of the present application, for the first document, knowledge graphs can be constructed based on the text corresponding to different content themes in the first target text. In order to distinguish it from the first knowledge graph, the knowledge graph constructed based on the text corresponding to a single content theme is recorded as the first subgraph, that is, the first knowledge graph can include first subgraphs corresponding to multiple content themes. Correspondingly, the second knowledge graph can include second subgraphs corresponding to multiple content bodies.
[0166] It should be noted that when the large language model is instructed to extract a knowledge graph from the first target text through prompt words, if a single knowledge graph is required, the large language model can be instructed to output a single knowledge graph, that is, a single knowledge graph integrates the contents of different content topics; if multiple knowledge graphs are required, the large language model can also be instructed to output multiple knowledge graphs based on different content topics, that is, a single knowledge graph only involves one of the content topics.
[0167] For example, for a document on construction functions, the content topic may represent a construction function point. In other words, the entities contained in the first subgraph corresponding to the content topic may include one or more of the following: construction function point name, construction location, technical indicators, and implementation phase. Attribute information for the entity may include the number of construction function points constructed, construction funding, and so on.
[0168] See also Figure 4 The above step S303: performing knowledge graph matching based on the first knowledge graph corresponding to the first document and the second knowledge graph corresponding to the second document to obtain the first similarity feature between the first document and the second document may specifically include:
[0169] S401: Determine a first sub-graph and a second sub-graph that have matching content themes.
[0170] The matching may be identical or similar. For example, “parking space” and “parking lot” are similar content topics and can be regarded as matching content topics.
[0171] In this step, if the second document also contains the same or similar content topics as the first document, it is determined that the two content topics match.
[0172] S402: Determine the graph similarity between the first sub-graph and the second sub-graph that have matching content themes, and obtain a preset number of groups of sub-graphs with the highest graph similarity.
[0173] In the embodiment of the present application, there may be multiple matching content topics. If similar features are analyzed based on a large model for each group of matching content topics, it will not be conducive to improving the analysis effect.
[0174] Therefore, for the first sub-graph and the second sub-graph that match the content themes, a graph similarity analysis may be performed first to obtain a preset number of groups of sub-graphs with the highest graph similarity.
[0175] Among them, graph similarity analysis can include analyzing the similarity of triples in the knowledge graph and the similarity of the overall topological structure, which will not be elaborated here.
[0176] After the above screening, a preset number of groups of sub-graphs with the highest graph similarity are determined, and subsequent duplication analysis is performed only on these sub-graphs.
[0177] For example, for documents related to construction projects, the graph similarity between sub-graphs can represent the degree of construction similarity between construction function points. For a specific construction function point, such as "charging piles," if the location, number, and construction funding of the charging piles represented by the first sub-graph differ significantly from those represented by the second sub-graph, it can be determined that the graph similarity between the two is low, meaning that the construction similarity of the charging piles is also low, and there is no potential risk of duplicate construction. Therefore, there is no need for subsequent duplication checking based on the large model.
[0178] S403: Input the third prompt word and each group of sub-graphs into the third model in sequence, and obtain the similarity features between each group of sub-graphs output by the third model as the first similarity features. The third model is used to determine the similar content between each group of sub-graphs based on the indication of the third prompt word.
[0179] In an embodiment of the present application, each group of sub-graphs and prompt words can be input into the large language model in sequence, instructing the large language model to generate duplicate checking results based on similar knowledge graphs.
[0180] For example, the prompt word could be "compare the similarities and differences between two knowledge graph triples and output the analysis results."
[0181] Under the guidance of the prompt word, the large language model compares and analyzes the first sub-graph and the second sub-graph with matching subject content, analyzes the similarities and dissimilarities, and then outputs the analysis results.
[0182] It can be seen that in the embodiment of the present application, a more refined division is performed based on the first target text, sub-texts of different content themes are determined, and a first sub-graph is constructed based on the sub-texts. By analyzing the graph similarity between the first sub-graph and the second sub-graph that match the content themes, multiple groups of sub-graphs with higher graph similarity are screened out. This removes content themes with lower similarity and reduces the burden of duplicate checking on the large model. Moreover, by further analyzing the similarity between sub-graphs with higher similarity through the large model, a refined and in-depth graph analysis can be provided, which helps to improve the precision and accuracy of duplicate checking.
[0183] The following describes how to determine a second document whose semantic similarity with the target text meets the conditions.
[0184] See also Figure 5 As a possible implementation of the embodiment of the present application, the above step of determining a second document whose semantic similarity with the target text satisfies the conditions from multiple preset documents may specifically include:
[0185] S501: Convert the target text into a target text vector.
[0186] Specifically, the target text can be processed based on the above embedding model to obtain a target text vector.
[0187] S502: Retrieve candidate vectors that meet similarity conditions from a vector library; the vector library includes text vectors converted according to key information of each preset document.
[0188] The vector library stores text vectors corresponding to multiple pre-stored preset texts. These text vectors are converted based on the key information of the preset documents. The steps for extracting key information from the preset texts can be found above and will not be repeated here.
[0189] In this step, candidate vectors whose similarity with the target text vector meets the preset conditions are retrieved from the vector library.
[0190] Exemplarily, the preset condition may be that the similarity between vectors is greater than a preset threshold.
[0191] After this retrieval step, multiple candidate vectors and the pre-stored preset documents to which each candidate vector belongs can be obtained.
[0192] S503: Determine a second document from the preset documents corresponding to the candidate vector.
[0193] In the embodiment of the present application, the preset text corresponding to the candidate vector may be used as a candidate document, and the second document may be determined from the candidate documents.
[0194] As a possible implementation of the embodiment of the present application, the candidate documents corresponding to the candidate vectors whose similarity is greater than a preset threshold can be determined as the second document. As a possible implementation of the embodiment of the present application, a preset number of candidate documents can be selected as the second document in descending order of similarity.
[0195] In an embodiment of the present application, there may be multiple second documents determined. When there are multiple second documents, a duplicate check may be performed on the similarity between the first document and each second document.
[0196] As can be seen, the embodiments of this application, by pre-establishing a vector library, quickly locate similar texts, laying the foundation for subsequent semantic feature duplication checking and deep graph feature duplication checking, facilitating efficient and accurate duplication checking. Furthermore, the system supports the addition of a large number of documents from different fields to expand the size of the vector library and text library to accommodate diverse application scenarios.
[0197] See below Figure 6 , combined with a specific example, the document duplication checking method provided in the embodiment of the present application is further explained.
[0198] Exemplarily, the document duplication checking method provided in the embodiment of the present application can be run on a duplication checking server. In addition, as mentioned above, the document duplication checking method provided in the embodiment of the present application can also be run on a terminal device.
[0199] like Figure 6 As shown in the example, a user uploads a document about a construction project as the first document to be checked for duplicate content. The duplicate checking server needs to efficiently and accurately detect duplicate content between this document and existing documents in the database. If it detects potential duplicate content, it notifies the user to avoid wasting resources.
[0200] As shown in the diagram, the first document to be checked for duplicate content is first analyzed for its structure, followed by the location of directory nodes. Targeted directory nodes, such as "Construction Content Introduction" and "Construction Function Module," facilitate large-scale model refinement and extraction of key text. This is then routed through the first routing module.
[0201] The key text extracted from the specified directory node is vectorized and similar documents are retrieved from the vector library. The second routing module then performs branching processing.
[0202] On the other hand, a first knowledge graph is constructed based on the key text of the first directory node (eg, construction function module) in the specified directory node.
[0203] For basic sections (or basic chapters), for example, other directory nodes except the construction function module, a large model is refined and similarities are analyzed based on the large model.
[0204] On the other hand, a knowledge graph comparison is performed based on the first knowledge graph and the second knowledge graph corresponding to the similar document, which may specifically include but is not limited to attribute comparison, structure comparison, and relationship comparison to obtain a knowledge graph comparison result.
[0205] Subsequently, the similarities analyzed based on the large model and the knowledge graph comparison results are summarized and processed, and then a duplicate checking report is output.
[0206] As can be seen, the target text representing key information in the first document is first extracted, similar text is quickly matched based on the target text, and then deep duplication detection is achieved by combining surface feature comparison with deep feature duplication detection, achieving a joint analysis of text semantic features and deep graph features, thereby performing efficient and accurate document duplication detection and reducing the false positive rate.
[0207] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the method. In order to realize the above functions, the training device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0208] In the embodiment of the present application, the functional modules of the document duplication checking device can be divided according to the above method. For example, the document duplication checking device can include functional modules corresponding to the functional divisions, or two or more functions can be integrated into one processing module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation.
[0209] For example, Figure 7 A possible schematic diagram of a document duplication checking device involved in the above embodiment is shown. The document duplication checking device 700 may include a first determining unit 701, a second determining unit 702, a matching unit 703, a third determining unit 704, and a fourth determining unit 705. The first determining unit 701 is used to determine a target text based on a first document to be checked for duplication; the target text is used to represent key information in the first document; the second determining unit 702 is used to determine a second document that meets a similarity condition with the target text from a plurality of preset documents; the matching unit 703 is used to perform knowledge graph matching based on a first knowledge graph corresponding to the first document and a second knowledge graph corresponding to the second document to obtain a first similarity feature between the first document and the second document; the third determining unit 704 is used to determine a semantically similar feature between the first document and the second document as a second similarity feature between the first document and the second document; and the fourth determining unit 705 is used to determine a duplication checking result for the first document based on the first similarity feature and the second similarity feature.
[0210] As can be seen, the target text representing key information in the first document is first extracted, and similar text is quickly matched based on the target text. The semantic similarity features of the first document and similar documents are then compared, and deep duplication detection is performed in combination with knowledge graph comparison. This combines surface feature comparison with deep feature duplication detection, achieving a joint analysis of text semantic features and deep graph features, leading to efficient and accurate document duplication detection and reducing false positive rates.
[0211] Optionally, determining the target text based on the first document to be checked for duplicates includes:
[0212] Get the original text under the specified directory node of the first document;
[0213] The original text and the first prompt word are input into the first model, and the text output by the first model is obtained as the target text; wherein the first model is used to extract key information from the original text according to the first prompt word.
[0214] As can be seen, using advanced large-scale model technology to conduct in-depth analysis of the original text under a specified directory node and extract key features and key content helps further focus on important information points, laying a solid foundation for subsequent vectorization processing and knowledge graph construction. Therefore, for document duplication checks based on fixed templates, this helps improve the efficiency and accuracy of duplicate checks, effectively reducing the error parsing rate, and further facilitating a deeper understanding of the textual semantics within the document.
[0215] Optionally, the designated directory node includes a first directory node, and the target text includes a first target text extracted from the original text under the first directory node;
[0216] The method also includes: constructing a first knowledge graph based on the first target text.
[0217] Constructing a first knowledge graph according to the first target text, including:
[0218] Inputting the first target text and the second prompt word into the second model to obtain triple information output by the second model;
[0219] Constructing a first knowledge graph based on triple information;
[0220] Among them, the second model is used to determine the entities, entity relationships and entity attribute information in the first target text according to the instructions of the second prompt word to output triple information; the first similarity feature includes triple information whose similarity between the first knowledge graph and the second knowledge graph meets the conditions.
[0221] It can be seen that in the embodiment of the present application, a directory node suitable for checking for duplicate content by means of knowledge graph comparison can be pre-specified, i.e., a first directory node. The first directory node can be selected according to the requirements and targets of checking for duplicate content. During the document checking for duplicate content, a knowledge graph is constructed based only on the key information extracted under the first directory node, and then a knowledge graph comparison method is adopted to check for duplicate content. For other directory nodes that are not suitable for checking for duplicate content by knowledge graph comparison, other methods of checking for duplicate content can be adopted. Thus, processing the texts under different directory nodes by multiple checking engines can help provide a multi-level perspective of checking for duplicate content, and carry out refined checking for duplicate content analysis on documents from different angles. This helps to improve the efficiency and accuracy of checking for duplicate content, and to gain an in-depth understanding of the text semantics in the document. For the checking for duplicate content of fixed-structure documents, it is more capable of accurate positioning and refined processing by chapter, which greatly improves the efficiency and accuracy of checking for duplicate content.
[0222] Optionally, the first target text includes a plurality of subtexts corresponding to different content themes; the first knowledge graph includes a first subgraph created based on each subtext; and the second knowledge graph includes a plurality of second subgraphs created based on texts of different content themes.
[0223] Performing knowledge graph matching based on a first knowledge graph corresponding to the first document and a second knowledge graph corresponding to the second document to obtain a first similarity feature between the first document and the second document includes:
[0224] Determine a first sub-graph and a second sub-graph that match the content theme;
[0225] Determine the graph similarity between the first sub-graph and the second sub-graph that have matching content themes, and obtain a preset number of groups of sub-graphs with the highest graph similarity;
[0226] The third prompt word and each group of sub-graphs are input into the third model in sequence, and similar features between each group of sub-graphs output by the third model are obtained as first similar features; wherein the third model is used to determine similar content between each group of sub-graphs according to the indication of the third prompt word.
[0227] It can be seen that in the embodiment of the present application, a more refined division is performed based on the first target text, sub-texts of different content themes are determined, and a first sub-graph is constructed based on the sub-texts. By analyzing the graph similarity between the first sub-graph and the second sub-graph that match the content themes, multiple groups of sub-graphs with higher graph similarity are screened out. This removes content themes with lower similarity and reduces the burden of duplicate checking on the large model. Moreover, by further analyzing the similarity between sub-graphs with higher similarity through the large model, a refined and in-depth graph analysis can be provided, which helps to improve the precision and accuracy of duplicate checking.
[0228] Optionally, the target text further includes a second target text extracted from the text under the second directory node specified by the first document.
[0229] In one possible implementation, determining semantically similar features between the first document and the second document as second similar features between the first document and the second document includes:
[0230] Determining a third directory node belonging to a basic section in the first document;
[0231] The fourth prompt word, the text under each third directory node in the first document, and the text under the corresponding directory node in the second document are input into the fourth model to obtain a second similarity feature; the fourth model is used to determine the semantically similar features between the input text of the first document and the input text of the second document based on the indication of the fourth prompt word.
[0232] As can be seen, each directory node within the basic section of the first document is compared for similar semantic features based on the large semantic model. This combines semantic similarity features with deep graph features for comprehensive duplication detection. This helps provide a multi-layered duplication detection perspective, enabling refined duplication analysis of documents from different angles. This improves duplication detection efficiency and accuracy, and provides a deeper understanding of the textual semantics within the document.
[0233] Optionally, key information represents text information related to the duplicate checking target;
[0234] When the target of duplicate checking is a construction project, the entities of the first knowledge graph include at least one of the construction function point name, construction location, technical indicators, and implementation stage; the attribute information of the first knowledge graph includes the construction number and / or construction funds of each construction function point.
[0235] As can be seen, when the duplicate check target is a construction project, by extracting knowledge graphs from the text, it is possible to extract various entities such as the construction function point name, construction location, technical indicators, and implementation phase, and identify attribute information such as the number of construction function points and construction funding. In the subsequent comparison stage, it is possible to accurately analyze the relevant information related to the construction function points, thereby detecting potential duplicate construction projects and avoiding the waste of resources caused by duplicate construction.
[0236] Optionally, determining a second document that satisfies a similarity condition with the target text from a plurality of preset documents includes:
[0237] Convert the target text into a target text vector;
[0238] Retrieving candidate vectors that meet similarity conditions from a vector library; the vector library includes text vectors converted based on key information of each preset document;
[0239] A second document is determined from the preset documents corresponding to the candidate vector.
[0240] As can be seen, the embodiments of this application, by pre-establishing a vector library, quickly locate similar texts, laying the foundation for subsequent semantic feature duplication checking and deep graph feature duplication checking, facilitating efficient and accurate duplication checking. Furthermore, the system supports the addition of a large number of documents from different fields to expand the size of the vector library and text library to accommodate diverse application scenarios.
[0241] An embodiment of the present application also provides a computing device, including: a processor and a memory; the processor is coupled to the memory; the memory is used for computer program instructions; the processor is used to call the computer program instructions in the memory so that the computing device executes any one of the methods in the above embodiments.
[0242] An embodiment of the present application further provides a computer-readable storage medium storing computer execution instructions. When the computer execution instructions are executed on a computing device, the computing device executes any one of the methods in the above embodiments.
[0243] For explanations of the relevant contents and descriptions of the beneficial effects of any of the computer-readable storage media provided above, reference may be made to the corresponding embodiments described above, and no further details will be given here.
[0244] The embodiment of the present application also provides a computer program product comprising instructions, which, when run on a computing device, causes the computing device to perform any one of the methods in the above embodiments. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. It should be noted that the above-mentioned devices for storing computer instructions or computer programs provided in the embodiment of the present application, such as but not limited to the above-mentioned memory, computer-readable storage medium and communication chip, etc., are all non-transitory.
[0245] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using a software program, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.
[0246] Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. A computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. Available media can be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
[0247] Although the present application has been described with reference to specific features and embodiments thereof, it is apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are merely illustrative of the present application as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art may make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, the present application is intended to include such modifications and variations as fall within the scope of the claims of the present application and their equivalents.
Claims
1. A document duplication checking method, characterized in that: The method comprises: Determining a target text based on a first document to be checked for duplicates; the target text is used to represent text information in the first document that is related to the duplicate checking target; Determining, from a document library, a second document that satisfies a similarity condition with the target text; Performing knowledge graph matching based on a first knowledge graph corresponding to the first document and a second knowledge graph corresponding to the second document to obtain a first similarity feature between the first document and the second document; determining a semantically similar feature between the first document and the second document as a second similar feature between the first document and the second document; Based on the first similarity feature and the second similarity feature, a duplicate checking result of the first document is determined.
2. The method according to claim 1, characterized in that The step of determining the target text based on the first document to be checked for duplicates includes: Obtaining the original text under the specified directory node of the first document; The original text and the first prompt word are input into a first model to obtain a text output by the first model as the target text; wherein the first model is used to extract key information from the original text according to the first prompt word.
3. The method according to claim 2, characterized in that The designated directory node includes a first directory node, the target text includes a first target text extracted from the original text under the first directory node, and the first target text includes a plurality of subtexts corresponding to different content themes; The method also includes: constructing the first knowledge graph based on multiple sub-texts in the first target text.
4. The method according to claim 3, characterized in that The constructing the first knowledge graph according to the first target text includes: Inputting the first target text and the second prompt word into the second model to obtain triple information output by the second model; Constructing the first knowledge graph based on the triple information; The second model is used to determine the entities, entity relationships and entity attribute information in the first target text according to the indication of the second prompt word, so as to output the triple information.
5. The method according to claim 4, characterized in that The first similarity feature includes triple information whose similarity between the first knowledge graph and the second knowledge graph meets the conditions.
6. The method according to claim 3 or 4, characterized in that The first knowledge graph includes a first sub-graph created based on each sub-text; the second knowledge graph includes a plurality of second sub-graphs created based on texts of different content themes; The performing knowledge graph matching based on the first knowledge graph corresponding to the first document and the second knowledge graph corresponding to the second document to obtain a first similarity feature between the first document and the second document includes: Determining a first sub-graph and a second sub-graph that match the content theme; Determine the graph similarity between the first sub-graph and the second sub-graph that match the content theme, and obtain a preset number of groups of sub-graphs with the highest graph similarity; The third prompt word and each group of sub-graphs are sequentially input into the third model to obtain similar features between each group of sub-graphs output by the third model as the first similar features; wherein the third model is used to determine similar content between each group of sub-graphs according to the indication of the third prompt word.
7. The method according to claim 1, characterized in that The determining of the semantically similar feature between the first document and the second document as the second similar feature between the first document and the second document includes: Determining a third directory node in the first document; A fourth prompt word, the text under the third directory node in the first document, and the text under the corresponding directory node in the second document are input into a fourth model to obtain the second similarity feature; the fourth model is used to determine the semantically similar features between the input text of the first document and the input text of the second document based on the indication of the fourth prompt word.
8. The method according to any one of claims 1 to 7, characterized in that When the target of the duplicate checking is a construction project, the entities of the first knowledge graph include at least one of the construction function point name, construction location, technical indicators, and implementation stage; the attribute information of the first knowledge graph includes the construction number and / or construction funds of each of the construction function points.
9. The method according to claim 1, characterized in that The step of determining, from a plurality of preset documents, a second document that satisfies a similarity condition with the target text includes: Converting the target text into a target text vector; Retrieving candidate vectors that meet the similarity condition from a vector library; the vector library includes text vectors converted according to key information of each of the preset documents; The second document is determined from preset documents corresponding to the candidate vectors.
10. A computing device, characterized in that comprising a processor and a memory; the processor being coupled to the memory; The memory is used for computer program instructions; The processor is configured to call computer program instructions in the memory to enable the computing device to execute the method according to any one of claims 1 to 9.
Citation Information
Cited By
Document duplicate checking method and device and medium
CN121211029A