Document duplicate checking method and computing equipment

By performing structured analysis and matching of functional points on documents, and using large models and vector technology to generate dilution check results, the problem of low dilution checking of existing documents is solved, and efficient and automated document dilution checking is achieved.

CN120542407APending Publication Date: 2025-08-26HENAN QINWEI DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510400602.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing document plagiarism checking methods require manual intervention, which consumes high labor costs and is inefficient.

Method used

By analyzing the styles of the content elements of the document, structured data are generated, functional points are determined, and similarity calculations are performed using large-scale models and vector matching techniques to generate dilution check results.

Benefits of technology

It improves the efficiency of document plagiarism checking, reduces manual intervention, and saves time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120542407A_ABST
    Figure CN120542407A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a document duplicate checking method and computing equipment. The method comprises the steps that a document to be subjected to duplicate checking is analyzed by means of styles of document content elements, structured data corresponding to the document to be subjected to duplicate checking are obtained, the structured data are used for indicating a data structure corresponding to document content of the document to be subjected to duplicate checking, and the structured data comprise hierarchical titles in the document to be subjected to duplicate checking; determining function points of the document to be subjected to duplicate checking based on the structured data; and performing similarity matching on the function points of the document to be subjected to duplicate checking by utilizing the function points of the historical document to obtain a duplicate checking result of the document to be subjected to duplicate checking. According to the method, the duplicate checking efficiency of duplicate checking of the document to be subjected to duplicate checking can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computing technology, and in particular to a document duplication checking method and computing device. Background Art

[0002] In today's era of information explosion, document processing and knowledge management have become indispensable across all industries. With the acceleration of digitization, the number and variety of documents has increased dramatically. Whether academic papers, business reports, patent applications, news reports, or policy documents, they all carry a wealth of information and innovative ideas. However, this wealth of information also presents challenges, the most prominent of which is verifying document originality and authenticity.

[0003] Therefore, to verify the originality and authenticity of a document, a duplicate check is necessary to determine whether duplicate or similar documents already exist. Related technologies for duplicate checking include manual part-of-speech tagging, keyword extraction and segmentation, building a synonym database, and rule matching to determine duplicate check results. However, this method not only requires manual intervention and is extremely labor-intensive, but also offers limited efficiency. Summary of the Invention

[0004] The embodiments of the present application provide a document duplication checking method and computing device for improving the efficiency of document duplication checking.

[0005] In the first aspect, an embodiment of the present application provides a document duplication checking method, comprising: parsing the document to be checked for duplicates using the style of the document content elements to obtain structured data corresponding to the document to be checked for duplicates, the structured data being used to indicate the data structure corresponding to the document content of the document to be checked for duplicates, the structured data including hierarchical titles in the document to be checked for duplicates; determining the functional points of the document to be checked for duplicates based on the structured data; the functional points being used to represent the content summary of the text content under the title; performing similarity matching on the functional points of the document to be checked for duplicates using the functional points of historical documents to obtain a duplication checking result of the document to be checked for duplicates, the duplication checking result being used to indicate the degree of similarity between the functional points of the document to be checked for duplicates and the functional points of historical documents.

[0006] The document duplicate checking method that the embodiment of the application provides is parsed to the document to be checked for duplicates by means of the style of the document content element, obtains the structured data corresponding to the document to be checked for duplicates.Should be understood that this structured data can be used to indicate the data structure corresponding to the document content of the document to be checked for duplicates, and the structured data comprises the level title in the document to be checked for duplicates.Because the document content corresponding to each level title in the document normally launches around a specific function point, it has certain independence and can be read and understood separately.Therefore, based on the structured data, the function point of the document to be checked for duplicates is determined, a plurality of function points in the document to be checked for duplicates can be determined more quickly.And then the function point of the historical document is utilized to carry out similar matching to the function point of the document to be checked for duplicates, determine the duplicate checking result of the degree of similarity between the function point of the document to be checked for duplicates and the function point of the historical document.

[0007] It can be seen that the solution provided in the embodiment of the present application can quickly break down the document content into quantifiable functional points by parsing the document to be checked for duplicates, and then can perform duplicate checking based on the functional points when checking for duplicates, without the need to manually compare the entire document, thereby greatly saving the time for document duplicate checking and improving the efficiency of document duplicate checking.

[0008] In one possible implementation, determining functional points of a document to be checked for duplicate content based on structured data includes: determining a target title within a hierarchy of headings based on the structured data; wherein the target title has a level higher than a preset level and no subheadings below the target title; or, alternatively, the target title has a level equal to a preset level; and determining the functional points corresponding to the target title based on its subheadings or the textual content corresponding to the target title. Lower-level headings typically correspond to very small units of content within a document, specifically addressing a specific detail while ignoring their relevance to higher-level content. Determining functional points based on these headings may result in an excessive number of functional points and excessive fragmentation. Therefore, when determining the target title, it is necessary to ensure that the target title has a level higher than or equal to a preset level. In other words, a title within the hierarchy of headings at a preset level is used as the target title. The preset level is pre-designed based on requirements or the structural characteristics of the document, and its corresponding hierarchy is relatively reasonable. The determined target title is positioned appropriately within the document's hierarchy, summarizing the main content without being overly broad or specific. Considering that some content in the document to be checked for duplicate content may not have a title at the preset level, the corresponding lowest-level title may have a level higher than the preset level. If the title at the preset level is used as the target title, functional points in the document to be checked for duplicate content may be missed. Therefore, the target title has a level higher than the preset level, and no subtitles exist under the target title.

[0009] In one possible implementation, the functional points corresponding to the target title are determined based on the subtitles of the target title or the text content corresponding to the target title, including: if there are subtitles under the target title, the functional points of the target title are determined based on the subtitles under the target title; if there are no subtitles under the target title, the functional points corresponding to the target title are determined based on the text content corresponding to the target title. It can be understood that subtitles are usually titles at the next level or lower level of the target title, and the content of subtitles is usually closely related to the target title, usually further refining and expanding the target title. In addition, subtitles are often descriptions or summaries of the specific content under the target title. Each subtitle may correspond to a functional point of the target title. By summarizing and analyzing these subtitles, the specific functions, characteristics, or aspects contained in the target title can be understood. Therefore, if there are subtitles under the target title, the functional points of the target title can be determined based on all subtitles under the target title.

[0010] In one possible implementation, determining the function points corresponding to a target title based on the text content corresponding to the target title involves inputting the text content corresponding to the target title in the document to be checked for duplicates into a large model, and then extracting function points from the text content corresponding to the target title using the large model to obtain the function points of the target title. The large model analyzes the input text content and identifies key information and functional descriptions related to the target title. It then organizes this information into function points for output.

[0011] In one possible implementation method, the main text content corresponding to the target title in the document to be checked for duplicates is input into the big model, including: inputting the main text content corresponding to the target title in the document to be checked for duplicates and prompt information into the big model; the prompt information is used to instruct the big model to summarize the content of the input text to obtain the functional points of the target title.

[0012] In one possible implementation, the functional points of a document to be checked for duplicate content are determined based on structured data, including: determining target paragraph content related to a duplication check index in the document to be checked for duplicate content based on the structured data; the duplication check index is used to represent the content of interest to the user; and determining the functional points of the document to be checked for duplicate content based on the hierarchical title corresponding to the target paragraph content. Considering that the document to be checked for duplicate content may contain a lot of content unrelated to the actual function, if the entire document is checked for duplicate content, the amount of data required to be processed is large. Therefore, during the duplication check, it is possible to only check the key content that the user is concerned about.

[0013] In one possible implementation, the structured data also includes the page number range corresponding to each level of the hierarchical title; the method also includes: determining the page number range corresponding to each functional point of the document to be checked for duplicates based on the page number range corresponding to each level of the hierarchical title; determining the text content corresponding to the functional points of the document to be checked for duplicates based on the page number range corresponding to the functional points of the document to be checked for duplicates; generating a duplicate checking report based on the text content corresponding to the functional points of the document to be checked for duplicates and the duplicate checking results of the document to be checked for duplicates.

[0014] In one possible implementation, a duplication checking report is generated based on the text content corresponding to the functional points of the document to be checked for duplication and the duplication checking results of the document to be checked for duplication, including: retrieving the target text content related to the functional points of the document to be checked for duplication from the text content corresponding to the functional points of the document to be checked for duplication; generating a duplication checking report, and in the duplication checking report, displaying the duplication checking results corresponding to the functional points in association with the target text content related to the functional points.

[0015] In one possible implementation, the duplication check report includes: the functional points of the document to be checked for duplication, the target text content corresponding to the functional points of the document to be checked for duplication, similar functional points and the degree of similarity corresponding to the similar functional points; similar functional points are functional points of historical documents whose degree of similarity to the functional points is higher than a preset similarity threshold.

[0016] In one possible implementation, the functional points of the historical documents are used to perform similarity matching on the functional points of the document to be checked for duplicates, and the duplicate checking result of the document to be checked for duplicates is obtained, including: determining the first vector corresponding to the functional points of the document to be checked for duplicates, and the second vector corresponding to the functional points of the historical documents; among the second vectors corresponding to the functional points of the historical documents, determining the M second vectors with the highest semantic similarity to the first vector; based on the functional points of the historical documents corresponding to the M second vectors and the semantic similarity, determining the duplicate checking result corresponding to the functional points of the document to be checked for duplicates. When the functional points of the historical documents are used to perform similarity matching on the functional points of the document to be checked for duplicates, since the text data is unstructured, the computer cannot process it directly. Therefore, it is necessary to convert the text into a numerical vector through vectorization so that the computer can use mathematical and statistical methods to calculate the similarity.

[0017] In a second aspect, an embodiment of the present application provides a document duplication checking device, which is used to execute any one of the document duplication checking methods provided in the first aspect above.

[0018] In a third aspect, an embodiment of the present application provides a computing device comprising a processor and a memory; the processor is coupled to the memory; the memory is used to store computer instructions, which are loaded and executed by the processor to enable the computing device to implement the method of the first or second aspect above.

[0019] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes: computer software instructions; when the computer software instructions are executed in a computing device, the computing device implements the method of the first aspect or the second aspect above.

[0020] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on a computing device, the computing device executes the steps of the relevant method described in the first or second aspect above to implement the method of the first or second aspect above.

[0021] The beneficial effects of the second to fifth aspects mentioned above can be referred to the corresponding description of the first aspect and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 A schematic diagram of an application scenario provided in an embodiment of the present application;

[0023] Figure 2 A schematic diagram of the system architecture of a server provided in an embodiment of the present application;

[0024] Figure 3 A flowchart of a document duplication checking method provided in an embodiment of the present application;

[0025] Figure 4 A flowchart of another document duplication checking method provided in an embodiment of the present application;

[0026] Figure 5 A schematic diagram of a duplicate check report provided in an embodiment of the present application;

[0027] Figure 6 A flowchart of another document duplication checking method provided in an embodiment of the present application;

[0028] Figure 7 A flowchart of another document duplication checking method provided in an embodiment of the present application;

[0029] Figure 8 A schematic diagram of another duplicate check report provided in an embodiment of the present application;

[0030] Figure 9 A flowchart of another document duplication checking method provided in an embodiment of the present application;

[0031] Figure 10 A schematic diagram of a complete solution process provided in an embodiment of the present application;

[0032] Figure 11 A schematic diagram of a document duplication checking device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0034] It should be noted that in the embodiments of this application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described in the embodiments of this application as "exemplarily" or "for example" should not be interpreted as being more preferred or advantageous than other embodiments or designs. Rather, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a concrete manner.

[0035] In order to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order.

[0036] The following is a brief description of the professional terms involved in the embodiments of this application:

[0037] 1. Large language model (LLM): It can be referred to as a large model, which refers to a machine learning model with large-scale parameters and computing power. These models are usually built with deep neural networks and have billions or even hundreds of billions of parameters. The large model is trained by inputting a large amount of corpus, so that the computer can acquire human-like "thinking" ability, understand text, pictures, voice and other content, and perform text generation, image generation, reasoning question and answer, scientific prediction and other tasks. In the embodiment of the present application, the large language model is mainly used to analyze the degree of semantic similarity between different texts.

[0038] 2. Vector Database: A vector database is a database system specifically designed for storing and processing vector data. In mathematics, vector data is represented as quantities with magnitude and direction. It can be represented by a line segment with an arrow. The direction of the arrow represents the direction of the vector, and the length of the line segment represents the magnitude of the vector.

[0039] In today's era of information explosion, document processing and knowledge management have become indispensable across all industries. With the acceleration of digitization, the number and variety of documents has increased dramatically. Whether academic papers, business reports, patent applications, news reports, or policy documents, they all carry a wealth of information and innovative ideas. However, this wealth of information also presents challenges, the most prominent of which is verifying document originality and authenticity.

[0040] The embodiment of the present application provides a method for checking for duplicate documents. First, the document to be checked for duplicate documents is parsed by means of the style of the document content elements, and the structured data corresponding to the document to be checked for duplicate documents is obtained. It should be understood that the structured data can be used to indicate the data structure corresponding to the document content of the document to be checked for duplicate documents, and the structured data includes the hierarchical titles in the document to be checked for duplicate documents. It is understandable that documents are usually divided based on hierarchical titles when they are written. The hierarchical titles can divide the document content into multiple levels, so that the document structure is clear and the logic is clear. Each hierarchical title usually corresponds to a part of the content of the document, such as a chapter, subchapter or paragraph. The content under each title usually revolves around a specific functional point, has a certain independence, and can be read and understood separately.

[0041] Secondly, by determining the functional points of the document to be checked for duplicates based on structured data including hierarchical titles, the document to be checked for duplicates can be divided into multiple parts more quickly, each part corresponding to a functional point. In other words, by determining the functional points of the document to be checked for duplicates based on structured data, multiple functional points in the document to be checked for duplicates can be determined more quickly. Furthermore, the functional points of the historical documents are used to perform similarity matching on the functional points of the document to be checked for duplicates, and a duplicate checking result is determined that indicates the degree of similarity between the functional points of the document to be checked for duplicates and the functional points of the historical documents.

[0042] It can be seen that the solution provided in the embodiment of the present application can quickly break down the document content into quantifiable functional points by parsing the document to be checked for duplicates, and then can perform duplicate checking based on the functional points when checking for duplicates, without the need to manually compare the entire document, thereby greatly saving the time for document duplicate checking and improving the efficiency of document duplicate checking.

[0043] Figure 1 A schematic diagram of a scenario provided by an embodiment of the present application is shown. Figure 1 As shown, the embodiment of the present application can be executed by a computing device 100, which can be a server running a document duplication checking system.

[0044] The computing device 100 can interact with the user-side display device 110. The display page of the document duplication checking system can be displayed on the user-side display device 110.

[0045] Among them, the display device 110 on the user side can be a device with an interface display function. For example, the display device 110 on the user side can display the display page of the document duplication checking system through a browser.

[0046] Optionally, the display device 110 on the user side may be a smart phone, a tablet computer, a personal portable computer, etc., which is not limited here.

[0047] For example, if the document duplication checking system is run by the computing device 100, when the display device 110 on the user side accesses the document duplication checking system, the display device 110 may display a first interface for uploading the document to be checked for duplication. Furthermore, in response to the user's first operation, the display device 110 sends the document to be checked for duplication to the computing device 100. The computing device 100 runs the document duplication checking system to check for duplication, obtains the duplication check results, and returns them to the display device 110. The display device 110 then displays the duplication check results on the interface so that the user can understand the similarity between the document to be checked for duplication and the historical documents.

[0048] The following uses the computing device as an example to introduce its system architecture. Figure 2 This is a schematic diagram of the server system architecture, such as Figure 2 As shown, the server's hardware includes a processor, an out-of-band controller, an external storage device, and a memory. The software includes an out-of-band management module and an operating system (OS).

[0049] The out-of-band management module runs in the out-of-band controller, and the OS runs on the processor (such as Figure 2 shown).

[0050] The out-of-band management module can be a management unit for non-business modules. For example, the out-of-band management module can remotely maintain and manage the server through a dedicated data channel. The out-of-band management module is completely independent of the server's operating system and can communicate with the basic input and output system (BIOS) and OS through the server's out-of-band management interface.

[0051] Exemplarily, the out-of-band management module may include a management unit for managing the server's operating status, a management system in a management chip, a system management module (SMM), etc. It should be noted that the embodiments of the present application do not limit the specific form of the out-of-band management module, and the above description is merely exemplary.

[0052] Memory, also known as internal memory or main memory, is installed in the memory slots on the server's motherboard. The memory communicates with the memory controller through a memory channel.

[0053] The external memory may be a hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the server, or an external storage device such as a USB flash drive.

[0054] The document duplication checking method provided in the embodiment of the present application can be applied to Figure 2 processor shown.

[0055] It should be noted that the system architecture and application scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0056] Figure 3 This is a flowchart of a document duplicate checking method provided in an embodiment of the present application. For example, the document duplicate checking method provided in an embodiment of the present application can be applied to Figure 1 or Figure 2 In the computing device shown, in other words, the method can be executed by the computing device, specifically, the method can be executed by a processor of the computing device.

[0057] like Figure 3 As shown, the document duplication checking method provided in the embodiment of the present application may include the following steps:

[0058] S301. parse the document to be checked for duplicates using the style of the document content elements to obtain structured data corresponding to the document to be checked for duplicates.

[0059] The structured data is used to indicate the data structure corresponding to the document content of the document to be checked for duplicates, and the structured data includes the hierarchical titles in the document to be checked for duplicates.

[0060] Document content elements (such as titles, body text, tables, and images) typically have specific style attributes, such as font, size, color, alignment, and indentation. This style information can be used to identify the structure and content of a document. By analyzing the styles of document content elements, unstructured document content can be converted into structured data, facilitating subsequent processing.

[0061] It should be understood that structured data refers to data organized in a specific format, such as extensible markup language (XML), JavaScript object notation (JSON) or a database table, and the embodiments of the present application are not limited to this.

[0062] Specifically, as a feasible implementation method, the content elements and their style information (such as bold font of the title, paragraph indentation of the text, etc.) in the document to be checked for duplicates (such as Word, PDF, etc.) can be extracted; the document structure corresponding to some content (such as title level, paragraph, list, table, etc.) can be identified based on the style information. The content elements are then classified into different structural units (such as title, text, picture, table, etc.). The identified structural units are converted into a structured data format (such as XML or JSON). Metadata (such as title level, paragraph number, table name, etc.) is added to each structural unit to generate structured data, and stored in a database or file to facilitate subsequent further processing of the structured data (such as duplicate checking, analysis, retrieval, etc.).

[0063] It should be understood that the document to be checked for duplicates may include a table of contents, which typically records content such as chapter numbers, titles, and page numbers. Therefore, as another feasible implementation method, the table of contents can be identified by identifying the style corresponding to the table of contents, and then parsing the table of contents into content such as chapter numbers, titles, and page numbers, and integrating the parsed style content to generate structured data.

[0064] For example, if the document to be checked for duplicates is in docx format, the Python docx library can be used to determine the table of contents style as "TOC," the first-level heading style as "Heading 1," the second-level heading style as "Heading 2," the third-level heading style as "Heading 3," ... the body content style as "normal," and the table style as "tbl." Structured data can then be generated based on the identified styles.

[0065] In the embodiments of the present application, the structured data includes at least the hierarchical headings in the document to be checked for duplicate content. It should be understood that the hierarchical headings in the document can divide the document content into multiple levels, making the document structure clear and logically distinct. Through the hierarchical display of the hierarchical headings, readers can more easily understand the overall structure and key content of the document.

[0066] It should be understood that each level of heading typically corresponds to a portion of the document, such as a chapter, subchapter, or paragraph. The hierarchy of the level of headings reflects the hierarchy of the content, for example, level one headings correspond to chapters, level two headings correspond to subchapter, and level three headings correspond to subsections. The content under each heading is generally independent and can be read and understood independently.

[0067] The common structure of hierarchical headings generally includes: a first-level heading (Title / Heading 1), which is usually used for the main title or chapter title of a document; a second-level heading (Heading 2), which is used to divide sub-chapter within a chapter; and a third-level heading (Heading 3), which is used to further subdivide the content and is usually used for the title of a section or paragraph. It should be understood that hierarchical headings can also include fourth-level headings or sub-headings below the third-level heading, and this embodiment of the application does not limit this.

[0068] It should be understood that the structured data may also include the text content, page numbers, image data, etc. in the document to be checked for duplicates, and the embodiments of the present application do not impose any restrictions on this.

[0069] S302: Determine the functional points of the document to be checked for duplicates based on the structured data.

[0070] It should be understood that a function point refers to a specific function, feature, or operation described in a document. For example, in a technical document, a function point might be an API interface, a software module, or an operation step; in a business document, a function point might be a product feature, a marketing strategy, or a business process.

[0071] Because structured data reflects the hierarchical structure of document content, particularly the hierarchical headings within structured data, the content under each hierarchical heading typically describes one or more corresponding functional points and possesses a certain degree of independence. In other words, each heading typically corresponds to one or more corresponding functional points or themes. Therefore, the functional points of the documents to be checked for duplicate content can be determined based on structured data.

[0072] In some embodiments, the functions of titles at different levels in the hierarchical title are different. Titles at higher levels usually have a larger scope of overview. For example, first-level titles are usually used for the main title or chapter title of a document. If the functional points of the document to be checked for duplicates are determined based on the first-level titles, the division of the determined functional points is likely to be too general and cannot accurately reflect the detailed content of the document. Titles at lower levels usually correspond to very small content units in the document (such as a specific operation step, parameter setting, etc.). If functional points are determined based on these titles, it may result in too many functional points and they may be too trivial. Therefore, when determining functional points based on hierarchical titles, it is necessary to reasonably determine the target title in the hierarchical title.

[0073] As a possible implementation, see Figure 4 , S302 can be specifically implemented as follows:

[0074] S3021. Determine a target title in the hierarchical titles based on the structured data.

[0075] The level of the target title is higher than the preset level, and there is no subtitle under the target title; or the level of the target title is the preset level.

[0076] It should be noted that the preset level is pre-set and can be determined according to the needs or the structural characteristics of the document itself during actual application. This embodiment of the present application does not limit this. For example, as a feasible implementation method, the preset level can be three levels.

[0077] It should be understood that lower-level headings often correspond to very small units of content within a document, specifically, they focus on a specific detail while neglecting their relevance to higher-level content. Determining function points based on these headings can result in an excessive number of trivial function points. Therefore, when determining target headings, ensure that the target heading's level is higher than or equal to the preset level.

[0078] Specifically, as an implementation manner, the level of the target title is a preset level.

[0079] In other words, the target title is chosen as the title at a preset level in the hierarchy. The preset levels are pre-designed based on requirements or the document's structural characteristics, and their corresponding hierarchies are relatively reasonable. The target title is positioned appropriately within the document's hierarchy, summarizing the main content without being too broad or too specific.

[0080] However, since some content in the document to be checked for duplicates may not have a title of the preset level, the level of the corresponding lowest-level title may be higher than the preset level. If the title of the preset level is used as the target title, the functional points in the document to be checked for duplicates may be missed.

[0081] Therefore, as another implementation manner, the level of the target title is higher than the preset level, and there is no subtitle under the target title.

[0082] The level of the target title is higher than the preset level, that is, in the hierarchical titles, the level (or importance, hierarchy) of the target title is higher than a preset level. For example, if the preset level is "level three title", then the target title determined according to this implementation method is a "level two title" or "level one title". Furthermore, among the multiple titles higher than the preset level, it is necessary to determine a title that does not have a subtitle as the target title, that is, there is no further subdivision or subtitle under the target title. In other words, this title is a "terminal" title, and there are no more specific sub-items or detail titles under it. Exemplarily, assuming that the "level two title" in the above example is a subtitle of the "level one title", and there is no "level three title" under the "level two title", then the "level two title" is the target title.

[0083] For example, assuming the preset level is the third-level title, the hierarchical title structure of the document to be checked for duplicates is as follows:

[0084] "First level title: Project Overview

[0085] Secondary heading: Project background

[0086] Level 3 heading: Market trends

[0087] Level 3 title: Industry analysis

[0088] Secondary heading: Project objectives

[0089] First level heading: Project scope

[0090] Based on the above "the level of the target title is the preset level", the third-level titles "Market Trends" and "Industry Analysis" can be used as target titles; based on the above "the level of the target title is higher than the preset level, and there is no sub-title under the target title", the second-level title "Project Goals" and the first-level title "Project Scope" can be used as target titles.

[0091] S3022: Determine the function point corresponding to the target title based on the subtitle of the target title or the text content corresponding to the target title.

[0092] After obtaining the target title, the functional point corresponding to the target title may be determined based on the subtitle of the target title or the text content corresponding to the target title.

[0093] Specifically, as a feasible implementation method, when there are subtitles under the target title, the functional points of the target title are determined based on the subtitles under the target title.

[0094] It's understandable that subheadings are typically titles at the next level or lower than the target title. The content of subheadings is often closely related to the target title, typically further refining and expanding upon it. Furthermore, subheadings often describe or summarize the specific content within the target title. Each subheading may correspond to a functional point within the target title. By summarizing and analyzing these subheadings, we can understand the specific functions, features, or aspects encompassed by the target title. Therefore, if subheadings exist within a target title, the functional points of the target title can be determined based on all of its subheadings.

[0095] As an implementation method, since subtitles are often descriptions or summaries of the specific content under the target title, all subtitles under the target title can be directly integrated as functional points of the target title.

[0096] As another feasible implementation method, when there is no subtitle under the target title, the function point corresponding to the target title is determined based on the text content corresponding to the target title.

[0097] Functional points refer to the functions, features, or effects represented by the target title. If there are no subtitles under the target title, the functional points are usually implied in the description of the main content. Therefore, to determine the functional points corresponding to the target title, it is necessary to conduct an in-depth analysis of the main content corresponding to the target title to determine its functional points.

[0098] Since large models have high semantic understanding capabilities, they are able to process and analyze large amounts of text data and extract meaningful information from them. Therefore, large models can serve as a powerful tool when determining the functional points corresponding to the target title.

[0099] Specifically, as an implementation method, the text content corresponding to the target title in the document to be checked for duplicates is input into the big model, and the function points of the text content corresponding to the target title are extracted through the big model to obtain the function points of the target title.

[0100] The target title's corresponding text content is fed into the large model so it can fully understand its semantics. The large model analyzes the input text content, identifying key information and functional descriptions related to the target title. It then organizes this information into functional points for output.

[0101] It should be understood that since large models generally have a character limit for input content, in actual applications, the text content can be divided into multiple segments and input into the large model separately, depending on the actual situation of the large model. For example, the text content can be truncated at intervals of 4,000 characters to obtain multiple text segments.

[0102] Furthermore, when dividing the main text, try to maintain the integrity and coherence of each segment so that the large model can accurately understand its semantics. Specifically, consider dividing the segments according to natural boundaries such as paragraphs and functional descriptions, which helps maintain semantic consistency within the segments.

[0103] As an implementation method, the text content corresponding to the target title in the document to be checked for duplicates and prompt information are input into the large model. The prompt information is used to instruct the large model to summarize the content of the input text to obtain the functional points of the target title.

[0104] It's important to understand that prompts for large models play a crucial role in interacting with them. They not only guide the model in generating desired responses but also directly impact the quality, accuracy, and relevance of those responses. Carefully designed prompts can clearly inform the model of the tasks it needs to perform and guide it to focus on specific aspects or angles when generating responses, ensuring that responses are closely aligned with user needs.

[0105] Therefore, when determining the functional points of the target title based on the large model, you can set prompt information and input it into the large model synchronously, so that the large model can summarize the content based on the input text to obtain the functional points of the target title. For example, the prompt information can be: "Please analyze the following text content and extract relevant functional points."

[0106] It should be understood that although large models have high semantic understanding capabilities, their accuracy may still be affected by various factors, such as the diversity of training data and the complexity of context, and their output may not be completely accurate or comprehensive. Therefore, as an implementation method, after obtaining the output of the large model, it can be displayed to the user through a display interface to facilitate manual review of the function points through the display interface. Through manual review, function points that may be omitted or misinterpreted by the large model can be discovered and corrected, thereby ensuring the accuracy and comprehensiveness of the function points.

[0107] It is understandable that the operation of a large model requires a large amount of computing resources, including CPU, GPU, and memory. Therefore, in this embodiment, when there are subtitles under the target title, directly determining the function points based on the subtitles can reduce the frequency of use of the large model, thereby avoiding the need to process the full text of the large model every time, thereby significantly reducing the consumption of computing resources. In addition, the large model is relatively slow when processing long texts or texts with complex structures. Subtitles are often summaries and refinements of the main text content. Determining the function points directly based on the subtitles can focus on the core content of the document more accurately, avoid the processing of irrelevant information, and can greatly increase the processing speed and improve overall efficiency.

[0108] When the target title doesn't contain subheadings, the large model leverages its powerful semantic understanding capabilities to deeply analyze the text content, accurately identifying and extracting functional points, thereby ensuring accurate extraction. Furthermore, the large model automates information extraction, reducing manual intervention. This level of automation is crucial for improving the accuracy and consistency of information processing when subheadings are absent.

[0109] S303: Use the function points of the historical documents to perform similarity matching with the function points of the document to be checked for duplicates, and obtain the duplicate checking result of the document to be checked for duplicates.

[0110] It should be understood that the method for obtaining the functional points of historical documents can refer to the method for obtaining the functional points of the document to be checked for duplicates in the above embodiment, and the embodiments of this application will not be repeated here.

[0111] Furthermore, as a feasible implementation method, after extracting the function points of a document, they can be stored in a database. This allows users to quickly understand the content of the document based on the function points. When using historical documents for duplicate checking, the database can be searched to see if the function points of the historical documents exist, thus avoiding the phenomenon of repeated extraction of function points.

[0112] After obtaining the function points of the historical document and the document to be checked for duplicates, we can determine whether there are similar function points in the historical document for each function point in the document to be checked for duplicates. As a feasible implementation method, we can establish a structured feature library and use semantic analysis and vector matching technology to calculate function point similarity. Finally, we can generate a duplicate check report that includes similarity scores and detailed comparisons.

[0113] The specific method of similarity matching of functional points of the document to be checked for duplicates can be found in the following embodiments, and this application will not go into details here.

[0114] It should be understood that the embodiment of the present application does not limit the specific content of the duplicate checking results. The duplicate checking results of the document to be checked for duplicates may include: the similarity between the functional points of the document to be checked for duplicates and the functional points of the historical documents, the overall similarity between the document to be checked for duplicates and the historical documents, a detailed comparative analysis report, and the location of similar sources. This allows users to compare the similarities between the document to be checked for duplicates and the historical documents.

[0115] As a feasible implementation method, the duplicate check results of the document to be checked can also be structured or visualized to ensure the clarity and interpretability of the results. For example, charts (such as bar charts and heat maps) can be used on the display interface of the computing device to intuitively display the similarity of functional points and overall similarity.

[0116] As an implementation method, a function point association network diagram can be provided on a display interface of a computing device to display the association relationship between similar function points.

[0117] Furthermore, users can click on similar function points on the display interface of the computing device to view detailed matching criteria and comparison content. In addition, the function of exporting duplicate check reports can be provided, supporting multiple formats (such as PDF and Excel) for user reference.

[0118] As can be seen, the document duplicate checking method provided by the embodiment of the present application, is parsed to the document to be checked for duplicates by means of the style of the document content element, obtains the structured data corresponding to the document to be checked for duplicates.Should be understood that this structured data can be used to indicate the data structure corresponding to the document content of the document to be checked for duplicates, and the structured data comprises the level title in the document to be checked for duplicates.Because the document content corresponding to each level title in the document normally launches around a specific function point, it has certain independence and can be read and understood separately.Therefore, based on structured data, determine the function point of the document to be checked for duplicates, can more quickly determine a plurality of function points in the document to be checked for duplicates.And then utilize the function point of historical documents to carry out similar matching to the function point of the document to be checked for duplicates, determine the duplicate checking result of the degree of similarity between the function point of the document to be checked for duplicates and the function point of historical documents.

[0119] That is to say, the solution provided in the embodiment of the present application can quickly break down the document content into quantifiable functional points by parsing the document to be checked for duplicates, and then can perform duplicate checking based on the functional points when checking for duplicates, without the need to manually compare the entire document, thereby greatly saving the time of document duplicate checking and improving the efficiency of document duplicate checking.

[0120] In some embodiments, the document to be checked for duplicate content may contain a lot of content that is irrelevant to the actual function. If the entire document is checked for duplicate content, the amount of data required to be processed is large. Therefore, when checking for duplicate content, the user can focus on the key content to be checked for duplicate content.

[0121] As a feasible implementation method, S302 can be specifically implemented as follows:

[0122] S401. Based on the structured data, determine the target paragraph content related to the duplication check index in the document to be checked for duplication.

[0123] Among them, the duplication check index is used to indicate the content that users are concerned about.

[0124] It should be understood that the duplication check index can include multiple user-preset tags, which are used to represent the content that the user is interested in. For example, the tags can be: functional point description, that is, the specific functions or features described in the document that the user is interested in; technical parameters: such as performance indicators, technical specifications, etc.; business logic, that is, the business process or logic described in the document; key terms, that is, professional terms or industry terms that the user considers important. This embodiment of the application does not limit this.

[0125] For example, when the document to be checked for duplicates is a feasibility study report document, the corresponding tag words may include: construction unit, construction goal, construction content, and construction function.

[0126] As a feasible implementation method, after determining the plagiarism check index, you can traverse the text content and title of each paragraph in the structured data of the document to be checked for plagiarism based on the label words of the plagiarism check index to find whether it contains these label words, and then use the paragraph content related to the label words as the target paragraph content.

[0127] For example, if a paragraph contains a tag word, the paragraph content is marked as the target paragraph content related to the duplicate check index. If a certain level of title contains a tag word, all the text content under the title can be marked as the target paragraph content.

[0128] As an implementation method, since first-level headings usually correspond to chapters in a document and can divide the document into multiple sections, tag words can be determined based on the first-level headings of the document to be checked for duplicate content. Specifically, the first-level headings can be displayed to the user through a display interface, allowing the user to select a tag word from multiple first-level headings, and then the content under the first-level heading corresponding to the selected tag word is used as the target paragraph content.

[0129] As another feasible implementation method, a regular expression (regex or regexp for short) can be constructed based on multiple label words of the plagiarism check index, and then the target paragraph content related to the plagiarism check index can be filtered out through the regular expression.

[0130] It should be understood that a regular expression is a pattern used to match character combinations in a string. It uses specific characters and symbols to define search patterns, which can efficiently search, edit, or manipulate text.

[0131] For example, if the duplicate check indicator is a paragraph containing a specific keyword, you can use \b (word boundary) and | (or) to construct a keyword matching pattern. Then use the constructed regular expression to match the document content and extract the target paragraph content that meets the duplicate check indicator.

[0132] As another feasible implementation method, the semantic relevance between each paragraph and each label word in the duplicate check index can be determined first, and then the paragraphs with semantic relevance higher than a preset threshold are used as target paragraphs corresponding to the label word.

[0133] Specifically, as an implementation method, the text understanding capabilities of the large model can be leveraged to identify target paragraph content related to plagiarism detection indicators. Specifically, the paragraph and tag words can be input into the large model to extract their semantic representations, and then the cosine similarity of the semantic representations can be calculated as semantic relevance.

[0134] As another implementation, the tag words and the words in the paragraph can be mapped into a high-dimensional vector space to capture the semantic relationship between the words. The semantic relevance between the paragraph and the tag words can then be measured by calculating the average similarity between the word vectors of the paragraph and the tag words.

[0135] S402: Determine the functional points of the document to be checked for duplicates based on the hierarchical titles corresponding to the target paragraph content.

[0136] The method of determining the functional points of the document to be checked for duplicates based on the hierarchical title corresponding to the target paragraph content can refer to the above S3021-S3022, and this application will not go into details here.

[0137] It should be noted that when checking for plagiarism based on the target paragraph content to be checked for plagiarism, the target paragraph content corresponding to each label word of the plagiarism check index can be checked for plagiarism separately, so that each label word corresponds to a similarity, which is convenient for users to check.

[0138] For example, see Figure 5 , in the case where the label words include construction unit, construction target, construction content and construction function, their proportions in the documents to be checked for duplicates can be determined respectively (e.g. Figure 5 5%, 10%, 10%, 70% as shown), and the corresponding similarity (as shown Figure 5 50%, 87%, 64%, and 61% as shown). Based on the similarity corresponding to each tag word, the overall similarity is determined to be 66%.

[0139] It should be understood that, as an implementation method, the duplicate check report can also generate a similarity radar chart based on the similarities corresponding to multiple tag words. This chart displays the similarity of documents across different tag words or dimensions in the form of a radar chart. Using the similarity radar chart, users can quickly identify the dimensions in which a document is similar to other documents, thereby gaining a more comprehensive understanding of the similarities and differences between documents.

[0140] It can be seen that the scheme provided by the present embodiment filters out the target paragraph content related to the duplicate checking index in the document to be checked for duplicates, and then performs a document duplicate checking based on the target paragraph content. Because the document to be checked for duplicates often contains a large amount of information that is irrelevant to the content of the user's attention, such as background introduction, cited documents, etc. These information may cause interference when checking for duplicates, causing the duplicate checking result to be inaccurate. Therefore, the scheme provided by the present embodiment checks the part that the user pays attention to, which can avoid this interference and improve the accuracy of the duplicate checking. Moreover, by focusing on the part that the user pays attention to, the computing resources of processing other parts can be more effectively utilized, avoiding wasting time and computing power on unnecessary content, and can also significantly reduce the amount of data that needs to be processed, thereby improving the duplicate checking speed.

[0141] In some embodiments, when using the functional points of historical documents to perform similarity matching with the functional points of the document to be checked for duplicates, computers cannot directly process text data because it is unstructured. Therefore, it is necessary to convert the text into a numerical vector through vectorization, so that the computer can use mathematical and statistical methods to calculate similarity.

[0142] Specifically, as a feasible implementation method, please refer to Figure 6 , S303 can be specifically implemented as follows:

[0143] S3031. Determine a first vector corresponding to a functional point of the document to be checked for duplicates, and a second vector corresponding to a functional point of the historical document.

[0144] That is to say, through vectorization, text data is converted into numerical vectors that can be processed by computers. It should be understood that vectorization is the basis for similarity calculation and duplicate checking.

[0145] When determining the first vector and the second vector, the text data may be converted into vectors using a vector conversion model. The vector conversion model may include an embedding model, a bag-of-words model (BOW), a TF-IDF (Term Frequency-Inverse Document Frequency) model, or a large model, which is not limited in this embodiment of the present application.

[0146] Exemplarily, as an implementation, an embedding (bge-m3) model may be used to convert the function points of the document to be checked for duplicates and the function points of the historical documents into a first vector and a second vector.

[0147] S3032. Determine, among the second vectors corresponding to the functional points of the historical documents, M second vectors having the highest semantic similarity to the first vector.

[0148] That is, using the similarity recall method, based on the vectorization result, documents or paragraphs similar to the second vector are retrieved from the second vector corresponding to the historical documents.

[0149] Specifically, the similarity between the first vector and the second vector can be determined by semantic similarity calculation. It should be understood that the calculation of semantic similarity can be carried out in one of the following ways: cosine similarity, which measures the cosine value of the angle between two vectors, where the closer the value is to 1, the more similar they are; Euclidean distance, which calculates the Euclidean distance between two vectors, where the smaller the value, the more similar they are; Jaccard similarity, which is applicable to set-type data and measures the ratio of the intersection to the union of two sets. The embodiments of the present application are not limited to this.

[0150] Then, the M function points with the highest semantic similarity are recalled (M is set according to the requirements, such as 3), that is, the function points of the M historical documents that are most similar to the function points of the document to be checked for duplicates.

[0151] As an implementation, you can input the vectors of all function points into the FAISS vectorization tool to construct a vector index. Then, use the KMeans clustering algorithm provided by FAISS (or implement KMeans using other libraries such as scikit-learn, and then input the cluster centers into FAISS for subsequent operations) to cluster the function point vectors and recall the similarities in the M most similar historical documents.

[0152] It should be understood that, as an implementation, a semantic similarity threshold may be pre-set, and function points with semantic similarity exceeding the threshold may be recalled to filter out function points with lower semantic similarity to avoid interference to users.

[0153] S3033. Based on the functional points and semantic similarities of the historical documents corresponding to the M second vectors, determine the duplicate checking results corresponding to the functional points of the document to be checked for duplicates.

[0154] After obtaining the functional points and semantic similarities of the historical documents corresponding to the M second vectors, they can be sorted into duplicate checking results and output for easy reference by users.

[0155] It should be understood that the duplicate checking results can be organized into an easy-to-understand format, such as using a table style to list each functional point of the document to be checked for duplicates and its similar historical document functional points (similar functional points), semantic similarity scores and other information.

[0156] As an implementation method, after obtaining the semantic similarities corresponding to the M second vectors, the functional points of the historical documents corresponding to the second vectors can be sorted based on the semantic similarities, so that users can more clearly understand which functional points have a higher risk of duplication.

[0157] In other words, the M second vectors initially recalled can be further sorted by re-sorting to improve the accuracy and relevance of the duplicate checking results.

[0158] Specifically, the sorting criteria may include similarity scores and other features. That is, the sorting is performed based on the similarity calculation results, and other features such as the release time, authority, and number of citations of historical documents are used as auxiliary sorting criteria.

[0159] As can be seen, the solution provided by this embodiment, by converting function points into numerical vectors through vectorization technology, can accurately capture the semantic information in the function points, facilitate subsequent calculations and storage, and improve processing efficiency. Furthermore, by using vector similarity calculations (such as cosine similarity and Euclidean distance), function points in historical documents that are similar to those in the document to be checked for duplicates can be quickly screened, greatly narrowing the scope of the duplicate check and improving the efficiency of the duplicate check.

[0160] In some embodiments, after obtaining the duplicate check results, in order to facilitate the user to review the actual similarities between the document to be checked for duplicates and the historical documents, the similar text content of the document to be checked for duplicates and the historical documents can be obtained and output. This not only enhances the transparency and interpretability of the duplicate check results, but also provides users with a more intuitive and specific basis for comparison, making it easier for users to review.

[0161] In order to facilitate the acquisition of the text content corresponding to each functional point, the page number can be used as the parsed content when parsing the document to be checked for duplicates. Therefore, the obtained structured data also includes the page number range corresponding to each level of the hierarchical title.

[0162] As a possible implementation, see Figure 7 The document duplication checking method provided in the embodiment of the present application further includes the following steps:

[0163] S501. Determine the page number range corresponding to each functional point of the document to be checked for duplicates based on the page number range corresponding to each level of the hierarchical title.

[0164] As can be seen from the above embodiment, function points are determined based on titles. Therefore, by analyzing the correspondence between titles and page numbers, a mapping relationship between function points and page ranges can be established. For example, a first-level title (such as "2. Project Overview") may correspond to pages 15-18; a second-level title (such as "2.1. Project Name") may correspond to pages 16-17. This mapping relationship provides a basic framework for function point location.

[0165] S502: Determine the text content corresponding to the functional points of the document to be checked for duplicates based on the page number range corresponding to the functional points of the document to be checked for duplicates.

[0166] After obtaining the page number range corresponding to the function point, the text content corresponding to the function point can be determined in the document to be checked for duplicates based on the page number.

[0167] Alternatively, as another implementation, when the structured data includes text content, the text content corresponding to the title corresponding to each functional point can be directly determined based on the structured data to obtain the text content corresponding to the functional point.

[0168] S503: Generate a duplicate checking report based on the text content corresponding to the functional points of the document to be checked for duplicates and the duplicate checking results of the document to be checked for duplicates.

[0169] Since the text content corresponding to the function points is large in size, displaying all the text content to users will inevitably affect their browsing experience. Therefore, based on the summarized function points, we can search the text content corresponding to the function points to find the paragraph content that is most relevant to the function points (the similarity is higher than the threshold), and then display the duplicate checking results of the paragraph content and the function points in the duplicate checking report.

[0170] Specifically, as a feasible implementation method, S503 can be specifically as follows: in the text content corresponding to the functional point of the document to be checked for duplicates, retrieve the target text content related to the functional point of the document to be checked for duplicates; generate a duplicate checking report, and in the duplicate checking report, associate and display the duplicate checking results corresponding to the functional point with the target text content related to the functional point.

[0171] Functional points in the document to be checked for duplicate content are identified using NLP techniques (such as entity recognition and semantic search) to locate relevant description paragraphs within the main text, which serve as the target content. The duplicate check report then accurately associates the functional points with the main text, allowing the report to not only display the duplication ratio but also directly identify the specific duplicate paragraphs. This allows users to compare implementation differences between different documents without requiring a secondary search.

[0172] For example, the large model can be used to calculate the semantic similarity between the function points and the paragraphs in the text content corresponding to the function points, and then the paragraph with the highest similarity is selected as the target text content and displayed in the duplicate checking report.

[0173] It should be understood that Figure 8 As shown, the duplicate check report can be displayed in a table, list or text form to ensure that the information is clear and easy to read. The embodiment of the present application does not limit the specific form of the duplicate check report. Moreover, the embodiment of the present application does not limit the specific content of the duplicate check report.

[0174] As an implementation method, the duplication check report includes: the functional points of the document to be checked for duplicates, the target text content corresponding to the functional points of the document to be checked for duplicates, similar functional points and the degree of similarity corresponding to the similar functional points; similar functional points are functional points of historical documents whose degree of similarity with the functional points of the document to be checked for duplicates is higher than a preset similarity threshold.

[0175] It is understandable that the duplicate check report includes the functional points of the document to be checked for duplicates, which can help users understand the functions of the document to be checked for duplicates, provide a benchmark for subsequent similarity comparisons, and ensure that the analysis direction is correct. Displaying the target text content in the duplicate check report can help users associate the functional points with specific description paragraphs, and users can understand the context without secondary search. Displaying similar functional points in the duplicate check report helps users quickly find paragraphs with similar functions in historical documents. Displaying the degree of similarity corresponding to similar functional points in the duplicate check report, that is, the similarity index (such as 85%, 92%), can intuitively display the degree of duplication, allowing users to quickly evaluate the scope of impact.

[0176] It should be understood that the duplicate check report may include more or less content, such as the text content corresponding to similar functional points, to facilitate user review. The embodiment of the present application does not limit the specific content of the duplicate check report.

[0177] In some embodiments, the duplicate checking results of all functional points can be integrated to determine a historical document that is most similar to the document to be checked for duplicates, so that users can perform comparative analysis.

[0178] As an implementation method, for each historical document, the similarity scores of all functional points with the document to be checked for duplicates can be accumulated to reflect the overall similarity between the historical document and the document to be checked for duplicates, and then the multiple historical documents with the highest scores (historical similar documents) can be regarded as the documents most similar to the document to be checked for duplicates.

[0179] As another implementation, in real applications, different function points may have different importance. Therefore, each function point can be assigned a weight and these weights can be taken into account when accumulating the similarity score. This step can make the similarity calculation more in line with actual needs.

[0180] As another implementation method, in addition to the similarity score, other factors (such as the number and distribution of function points) can also be considered to comprehensively determine the most similar historical documents, making the determination result more comprehensive and accurate.

[0181] For example, Figure 5 As shown, the duplicate check report may include: document information of historical similar documents. The document information may include: Figure 5It should be understood that the document information may also include other content, such as document author, document time, etc., which is not limited in this application.

[0182] In some embodiments, in order to facilitate users to identify differences, the original contents corresponding to similar functional points of the document to be checked for duplicates and historical similar documents can be displayed in the duplicate checking report through the above S501-S503 methods for user reference.

[0183] As a way to implement Figure 8 As shown, it is also possible to provide a side-by-side comparison of the original text content of the historically similar document and the original text content of the document to be checked for duplicates, so that users can easily view the differences. Figure 8 As shown, similar texts or words in the original content of the document to be checked for duplicates and historical similar documents can also be displayed in the same way ( Figure 8 The bold or underlined characters in the figure are displayed for the convenience of users to refer to the corresponding characters.

[0184] As an implementation method, when displaying the original content corresponding to the similar functional points of the document to be checked for duplicates and the historical similar documents, such as Figure 8 As shown, in the duplicate checking report, the chapters or parent titles corresponding to the original content of the document to be checked for duplicates can be displayed accordingly, so that the user can determine its source in the document.

[0185] In some embodiments, the duplicate checking results of each function point can also be summarized and counted, such as determining the repetition rate and number of duplicate sources corresponding to the function point based on all similar historical function points corresponding to the function point.

[0186] In some embodiments, the document to be checked for duplicates usually also includes image data, which is also part of the text, but this part of data is usually not obtained when parsing the document to be checked for duplicates. Therefore, when the document to be checked for duplicates includes image data, it can be identified by the following identification method.

[0187] Specifically, as a feasible implementation method, the structured data corresponding to the document to be checked for duplicates in the embodiment of the present application also includes the text content of the document to be checked for duplicates. Figure 9 In the document duplication checking method provided in the embodiment of the present application, S301 can be specifically implemented as follows:

[0188] S601: Identify image data in the document to be checked for duplicates.

[0189] Images in documents usually exist in a specific form, so image data in a document can be identified based on the style or format corresponding to the image.

[0190] For example, when the document to be checked for duplicates is a ".docx" file, the ".docx" file is actually a ZIP compressed package, which contains multiple files and folders. The pictures are stored as independent files in a specific folder within this compressed package, and the association between the pictures and the document content is defined by an XML file, especially the document.xml file. In this file, the pictures are defined by specific XML elements (such as <w:drawing>or <wp:inline>, depending on how the image was inserted and the Office version). Therefore, by parsing the XML structure of the document to be checked for duplicates, you can locate elements containing images. These elements may have specific tags, such as "pict" or "drawing."

[0191] Since the image data in the document to be checked for duplicates may contain tables or other types of image content, such as diagrams, charts, photos, etc., in order to effectively process these image data and perform duplicate checking, tables can be distinguished from other types of content.

[0192] As a feasible implementation method, a table location detection and recognition algorithm can be used to determine whether the image content is a table or ordinary text. The table location detection and recognition algorithm can automatically identify and locate the position of the table by analyzing the visual features in the image, and then distinguish whether the image content is a table or ordinary text.

[0193] Furthermore, optical character recognition (OCR) technology can be used to convert text in images into editable text. OCR technology can recognize text in images and convert it into a computer-processable format, facilitating subsequent duplicate checking and analysis.

[0194] S602: When the image data is a table, the image data is converted into table data based on the table format of the image data, and the converted table data is used as the main text content of the document to be checked for duplicates.

[0195] When the image data is a table, the text information recognized by OCR can be sorted according to the rows and columns of the recognized table, and the sorted data can be converted into a computer-processable table format, such as CSV, Excel, or JSON, to form structured data.

[0196] For example, as an implementation method, tab characters may be used between columns, line breaks may be used between rows, and then characters may be concatenated to obtain data content in a table format.

[0197] Finally, the converted table data is used as the main content of the document to be checked for duplicate content, so as to facilitate document duplication checking.

[0198] S603: When the image data is ordinary text, the ordinary text corresponding to the image data is used as the main text content of the document to be checked for duplicates.

[0199] In the case where the image data is ordinary text or other types of images, the ordinary text extracted from the image can be used as the main text content of the document to be checked for duplicate content, so as to facilitate document duplicate checking.

[0200] It can be seen that the solution provided in this embodiment converts images into tables or text information to facilitate subsequent duplicate checking. It is understandable that if duplicate checking is performed directly in the form of images, it is often difficult to accurately identify the content therein. After conversion to a table or text, the duplicate checking system can perform a comparison based on the precise text content, thereby improving the accuracy of the duplicate checking. In addition, text and table data are easier to quickly compare than image data. The duplicate checking system can quickly scan and compare large amounts of text or table data, thereby significantly improving the efficiency of duplicate checking.

[0201] It can be seen that the above mainly introduces the solution provided by the embodiment of the present application from the perspective of method. In order to achieve the above functions, the embodiment of the present application provides hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily appreciate that, in combination with the modules and algorithm steps of each example described in the embodiment disclosed herein, the embodiment of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific documents and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific document, but such implementation should not be considered to be beyond the scope of this application.

[0202] In some embodiments, see Figure 10 , Figure 10 A schematic diagram of a document duplication check provided in an embodiment of the present application. Figure 10 As shown, for the document to be checked for duplicate content, the structured data is first extracted by analyzing the style of the document content elements (element style or directory style), and then the target paragraph content (that is, the paragraph content that needs to be checked for duplicate content) can be determined by text extraction or large model reasoning. Among them, text extraction is used to retrieve the target paragraph content related to the tag word in the title or article content based on the tag word, and large model reasoning is used to determine the target paragraph content based on the text comprehension ability of the large model. The specific implementation method can refer to the above embodiment.

[0203] Then, for the target paragraph content, function points are extracted through rule matching and / or a hybrid of the large model. Rule matching directly uses the subheadings of the target title with subheadings as function points. The method for extracting function points using the large model can refer to the above-mentioned embodiment. It should be understood that the function point extraction process for historical documents can refer to the function point extraction process for documents to be checked for duplicates.

[0204] After obtaining the function points of the document to be checked for duplicates and historical documents, they can be converted into vectors using a vector database. Data recall can then be used to identify M function points in historical documents that are similar to those in the document to be checked for duplicates. To facilitate user review, a large model can be used to generate similarities and conclusions based on these similar function points. A duplicate checking report is then generated based on the output of the large model and the document to be checked for duplicates.

[0205] In an exemplary embodiment, the present application further provides a document duplication checking device. The document duplication checking device may be the aforementioned computing device, or a processor within the computing device. The document duplication checking device may include one or more functional modules for implementing the document duplication checking method of the above method embodiment.

[0206] For example, Figure 11 This is a schematic diagram of a document duplicate checking device provided in an embodiment of the present application. Figure 11 As shown, the document duplicate checking device 110 includes: a parsing module 111 , a determination module 112 and a matching module 113 .

[0207] Among them, the parsing module 111 is used to parse the document to be checked for duplicates using the style of the document content elements to obtain structured data corresponding to the document to be checked for duplicates. The structured data is used to indicate the data structure corresponding to the document content of the document to be checked for duplicates. The structured data includes the hierarchical titles in the document to be checked for duplicates.

[0208] The determination module 112 is used to determine the functional points of the document to be checked for duplicates based on the structured data; the functional points are used to represent the content summary of the text content under the title. The matching module 113 is used to perform similarity matching on the functional points of the document to be checked for duplicates using the functional points of the historical documents, thereby obtaining a duplicate checking result for the document to be checked for duplicates, which is used to indicate the degree of similarity between the functional points of the document to be checked for duplicates and the functional points of the historical documents.

[0209] In one possible implementation, the determination module 112 is specifically used to determine the target title in the hierarchical title based on the structured data; wherein the level of the target title is higher than the preset level, and there is no subtitle under the target title; or, the level of the target title is the preset level; based on the subtitle of the target title or the text content corresponding to the target title, determine the functional point corresponding to the target title.

[0210] In one possible implementation, the determination module 112 is specifically used to, when there is a subtitle under the target title, determine the functional point of the target title based on the subtitle under the target title; when there is no subtitle under the target title, determine the functional point corresponding to the target title based on the text content corresponding to the target title.

[0211] In one possible implementation, the determination module 112 is specifically configured to input the text content corresponding to the target title in the document to be checked for duplicates into the large model, extract function points from the text content corresponding to the target title through the large model, and obtain function points of the target title.

[0212] In one possible implementation, the determination module 112 is specifically used to input the text content corresponding to the target title in the document to be checked for duplicates, as well as prompt information, into the large model; the prompt information is used to instruct the large model to summarize the content of the input text to obtain the functional points of the target title.

[0213] In one possible implementation, the determination module 112 is specifically used to determine the target paragraph content related to the duplication checking index in the document to be checked for duplication based on structured data; the duplication checking index is used to represent the user's attention content; and based on the hierarchical title corresponding to the target paragraph content, the functional points of the document to be checked for duplication are determined.

[0214] In one possible implementation, the structured data also includes the page number range corresponding to each level of the hierarchical title; the determination module 112 is also used to determine the page number range corresponding to each functional point of the document to be checked for duplicates based on the page number range corresponding to each level of the hierarchical title; determine the text content corresponding to the functional point of the document to be checked for duplicates based on the page number range corresponding to the functional point of the document to be checked for duplicates; generate a duplicate checking report based on the text content corresponding to the functional point of the document to be checked for duplicates and the duplicate checking results of the document to be checked for duplicates.

[0215] In one possible implementation, the determination module 112 is specifically used to retrieve the target text content related to the functional point of the document to be checked for duplicates from the text content corresponding to the functional point of the document to be checked for duplicates; generate a duplicate checking report, and in the duplicate checking report, associate and display the duplicate checking results corresponding to the functional point with the target text content related to the functional point.

[0216] In one possible implementation, the duplication check report includes: the functional points of the document to be checked for duplication, the target text content corresponding to the functional points of the document to be checked for duplication, similar functional points and the degree of similarity corresponding to the similar functional points; similar functional points are functional points of historical documents whose degree of similarity to the functional points is higher than a preset similarity threshold.

[0217] In one possible implementation, the structured data also includes the main text content of the document to be checked for duplicates; the parsing module 111 is specifically used to identify image data in the document to be checked for duplicates; when the image data is a table, the image data is converted into table data based on the table format of the image data, and the converted table data is used as the main text content of the document to be checked for duplicates; when the image data is ordinary text, the ordinary text corresponding to the image data is used as the main text content of the document to be checked for duplicates.

[0218] In one possible implementation, the matching module 113 is specifically used to determine the first vector corresponding to the functional point of the document to be checked for duplicates, and the second vector corresponding to the functional point of the historical document; among the second vectors corresponding to the functional points of the historical document, determine the M second vectors with the highest semantic similarity to the first vector; based on the functional points of the historical document corresponding to the M second vectors and the semantic similarity, determine the duplicate checking results corresponding to the functional points of the document to be checked for duplicates.

[0219] For the detailed description of the above optional methods, please refer to the above method embodiments, which will not be repeated here. In addition, the explanation of any of the above-mentioned document duplicate checking devices and the description of their beneficial effects can refer to the above-mentioned corresponding method embodiments, which will not be repeated here.

[0220] The embodiment of the present application also provides a computer-readable storage medium. All or part of the processes in the above method embodiments can be completed by computer instructions to instruct related hardware. Exemplarily, the related hardware can be a processor of a computing device. The program instructions can be stored in the above computer-readable storage medium. When the program instructions are executed, the processes of the above method embodiments can be implemented. The computer-readable storage medium can be a memory. The above computer-readable storage medium can also be an external storage device, such as a hard disk, a smart memory card (smart media card, SMC), a secure digital (secure digital, SD) card, a flash card (flash card), etc. Further, the above computer-readable storage medium can also include both a memory and an external storage device. The above computer-readable storage medium is used to store the above computer program instructions and other programs and data required for the above document duplicate checking method.

[0221] An embodiment of the present application also provides a computer program product, which includes a computer program. When the computer program product is run on a computing device, the computing device executes any one of the document duplication checking methods provided in the above embodiments.

[0222] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple situations. A single processor or other unit may implement several functions listed in the claims. The fact that certain measures are recorded in different dependent claims does not mean that these measures cannot be combined to produce good results.

[0223] Although the present application has been described with reference to specific features and embodiments thereof, it is apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are merely illustrative of the present application as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art may make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, the present application is intended to include such modifications and variations as fall within the scope of the claims of the present application and their equivalents.

[0224] The above are only specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.< / wp:inline> < / w:drawing>

Claims

1. A document duplication checking method, characterized in that: include: Parsing the document to be checked for duplicates using the style of the document content elements to obtain structured data corresponding to the document to be checked for duplicates, wherein the structured data is used to indicate the data structure corresponding to the document content of the document to be checked for duplicates, and the structured data includes hierarchical titles in the document to be checked for duplicates; Determining the functional points of the document to be checked for duplicates based on the structured data; the functional points are used to represent the content summary of the text content under the title; The function points of the historical document are used to perform similarity matching on the function points of the document to be checked for duplicates, and a duplicate checking result of the document to be checked for duplicates is obtained. The duplicate checking result is used to indicate the degree of similarity between the function points of the document to be checked for duplicates and the function points of the historical document.

2. The method according to claim 1, characterized in that The determining of the functional points of the document to be checked for duplicates based on the structured data includes: Determining a target title in the hierarchical titles based on the structured data; wherein the level of the target title is higher than a preset level and there is no subtitle under the target title; or, the level of the target title is the preset level; Based on the subtitle of the target title or the main text content corresponding to the target title, a functional point corresponding to the target title is determined.

3. The method according to claim 2, characterized in that The determining of the function point corresponding to the target title based on the subtitle of the target title or the text content corresponding to the target title includes: In the case where there is a subtitle under the target title, determining the function point of the target title based on the subtitle under the target title; In the case that there is no subtitle under the target title, the function point corresponding to the target title is determined based on the text content corresponding to the target title.

4. The method according to claim 3, characterized in that The determining of the function point corresponding to the target title based on the text content corresponding to the target title includes: The text content corresponding to the target title in the document to be checked for duplicates is input into a large model, and the function points of the text content corresponding to the target title are extracted using the large model to obtain the function points of the target title.

5. The method according to claim 4, characterized in that The step of inputting the text content corresponding to the target title in the document to be checked for duplicates into the macro model comprises: The main text content corresponding to the target title in the document to be checked for duplicates and the prompt information are input into the large model; the prompt information is used to instruct the large model to summarize the content of the input text to obtain the functional points of the target title.

6. The method according to claim 1, characterized in that The determining of the functional points of the document to be checked for duplicates based on the structured data includes: Based on the structured data, determining target paragraph content related to a duplication check index in the document to be checked for duplicate content; the duplication check index is used to represent the user's focus; Based on the hierarchical title corresponding to the target paragraph content, the functional points of the document to be checked for duplicates are determined.

7. The method according to claim 1, characterized in that The structured data also includes a page number range corresponding to each level of the hierarchical title; the method further includes: Determine the page number range corresponding to each functional point of the document to be checked for duplicates based on the page number range corresponding to each level of the hierarchical title; Determining the text content corresponding to the functional point of the document to be checked for duplicates based on the page number range corresponding to the functional point of the document to be checked for duplicates; A duplicate checking report is generated based on the text content corresponding to the functional points of the document to be checked for duplicates and the duplicate checking results of the document to be checked for duplicates.

8. The method according to claim 7, characterized in that The generating of a duplicate checking report based on the text content corresponding to the functional points of the document to be checked for duplicates and the duplicate checking result of the document to be checked for duplicates includes: Retrieving target text content related to the functional point of the document to be checked for duplicates from the text content corresponding to the functional point of the document to be checked for duplicates; Generate the duplicate checking report, and in the duplicate checking report, associate the duplicate checking results corresponding to the function points with the target text content related to the function points and display them.

9. The method according to claim 8, characterized in that The duplication check report includes: the functional points of the document to be checked for duplication, the target text content corresponding to the functional points of the document to be checked for duplication, similar functional points and the degree of similarity corresponding to the similar functional points; the similar functional points are functional points of historical documents whose degree of similarity to the functional points is higher than a preset similarity threshold.

10. The method according to claim 1, characterized in that The method of performing similarity matching on the function points of the document to be checked for duplicates by using the function points of the historical document to obtain the duplicate checking result of the document to be checked for duplicates includes: Determine a first vector corresponding to a functional point of the document to be checked for duplicates, and a second vector corresponding to a functional point of the historical document; Determining, among the second vectors corresponding to the functional points of the historical documents, M second vectors having the highest semantic similarity to the first vector; Based on the functional points and semantic similarities of the historical documents corresponding to the M second vectors, a duplicate checking result corresponding to the functional points of the document to be duplicate checked is determined.

11. A computing device, characterized in that The computing device includes a processor and a memory; the processor is coupled to the memory; The memory is used to store computer instructions; The computer instructions are loaded and executed by the processor to enable the computing device to implement the method according to any one of claims 1 to 10.