Text duplicate checking method and device, electronic equipment and storage medium
By employing multiple language processing models and vector matching algorithms, this technology addresses the issue of low accuracy in identifying synonyms and semantic similarity in existing text plagiarism detection methods. It achieves multi-dimensional text matching, thereby improving the accuracy and efficiency of plagiarism detection results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-04-07
AI Technical Summary
Existing text plagiarism detection technologies cannot effectively handle synonyms and semantic similarity, resulting in low accuracy in recognizing generative texts that have undergone synonym rewriting, word order adjustment, and cross-language rewriting.
By using multiple language processing models to perform feature processing on the text to be checked and the text to be compared, multi-dimensional feature vectors are generated. Pre-set vector matching algorithms are used for matching, and algorithms such as cosine similarity and Manhattan distance are combined to determine the matching relationship of text segments, thus achieving multi-dimensional similarity matching.
It improves the matching accuracy of text plagiarism detection, avoids omissions, and ensures the accuracy and comprehensiveness of plagiarism detection results, making it suitable for efficient processing of long texts.
Smart Images

Figure CN121809445A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of text duplication checking, and particularly relates to a text duplication checking method, a text duplication checking device, an electronic device, and a computer readable storage medium. BACKGROUND
[0002] The current mainstream text duplication checking technology is mainly based on statistical characteristics. This kind of method mainly analyzes the statistical characteristics such as word frequency, part of speech, and syntactic structure of the text to calculate the similarity between texts. However, this method cannot effectively handle synonyms and semantic similarity, and has low recognition accuracy for generated texts that have undergone synonym rewriting, syntax adjustment, cross-language rewriting, etc. SUMMARY
[0003] The present application aims to at least solve one of the technical problems existing in the prior art. To this end, the present application provides a text duplication checking method, a text duplication checking device, an electronic device, and a computer readable storage medium, which can improve matching accuracy and avoid omission, thereby improving the accuracy of the duplication checking result determined based on the matching result.
[0004] In a first aspect, the present application provides a text duplication checking method, which comprises: processing features of the to-be-checked text and the comparison text respectively through various different language processing models to output a first feature vector corresponding to the to-be-checked text and a second feature vector corresponding to each of the comparison texts; matching the first feature vector and the second feature vector through a preset vector matching algorithm to obtain a second text segment matched by the first text segment of the to-be-checked text in the comparison text; determining a duplication checking result of the to-be-checked text based on the second text segment matched by the first text segment, wherein the duplication checking result comprises a target text matched by the to-be-checked text in the comparison text.
[0005] In a second aspect, the present application provides a text duplication checking device, which comprises: a feature processing module configured to process features of the to-be-checked text and the comparison text respectively through various different language processing models to output a first feature vector corresponding to the to-be-checked text and a second feature vector corresponding to each of the comparison texts; a matching module configured to match the first feature vector and the second feature vector based on a preset vector matching algorithm to obtain a second text segment matched by the first text segment of the to-be-checked text in the comparison text; The determining module is used to determine the plagiarism detection result of the text to be checked based on the second text segment matched by the first text segment, wherein the plagiarism detection result includes the target text matched by the text to be checked in the comparison text.
[0006] Thirdly, this application provides an electronic device, comprising: Memory and processor; The memory is used to store one or more computer instructions; The processor is used to execute one or more computer instructions to implement the text plagiarism detection method described above.
[0007] Fourthly, this application provides a computer-readable storage medium storing one or more computer instructions that, when executed by a processor, implement the above-described text deduplication method.
[0008] The text plagiarism detection method, device, electronic device, and computer-readable storage medium provided in this application extract feature vectors from the text to be checked and the comparison text using multiple language processing models, achieving multi-dimensional feature extraction. Then, for each model's feature vector, vector matching is performed to achieve multi-dimensional similarity matching, improving matching accuracy and avoiding omissions. Finally, based on the accurate and complete matching of the first and second text segments, plagiarism detection of the text to be checked is achieved, thereby improving the accuracy of the subsequent plagiarism detection results determined based on the matching results.
[0009] Additional aspects and advantages of embodiments of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of embodiments of this application. Attached Figure Description
[0010] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a schematic diagram illustrating an application scenario of the text plagiarism detection method provided in this application embodiment; Figure 2 This is a schematic diagram of the first process of the text plagiarism detection method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the second process of the text plagiarism detection method provided in the embodiments of this application; Figure 4 This is a schematic diagram of the modules of the text plagiarism detection device provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application; and Figure 6This is a schematic diagram of the hardware structure of the terminal device provided in the embodiments of this application. Detailed Implementation
[0011] Numerous specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.
[0012] It should be noted that the terms "first," "second," "third," etc., in the claims, specification, and drawings of this application are used to distinguish similar objects and are not used to describe a specific order or sequence. Such data are interchangeable where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown or described herein. Furthermore, the terms "comprising," "having," and their variations are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0013] It should be understood that in the embodiments of this application, "at least one" means one or more, and "more than one" means two or more. "And / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. The character " / " generally indicates that the related objects before and after it are in an "or" relationship. "Contains A, B and / or C" means containing any one, two, or three of A, B, and C.
[0014] It should be understood that in the embodiments of this application, "B corresponding to A", "B corresponding to A", "A corresponds to B" or "B corresponds to A" means that B is associated with A, and B can be determined based on A. Determining B based on A does not mean that B is determined solely based on A; B can also be determined based on A and / or other information.
[0015] This application provides a text deduplication method, apparatus, electronic device, and computer-readable storage medium.
[0016] The text plagiarism detection method provided in this application can be executed by an electronic device, which can be a terminal device or a server, etc. The terminal device can be a smartphone, tablet computer, laptop computer, etc.
[0017] A server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0018] In one alternative embodiment, when the text plagiarism detection method is run on a terminal device, the terminal device may include a display screen and a processor, the display screen being used to present the plagiarism detection results.
[0019] There are several ways in which a terminal device can provide a graphical user interface to a user. For example, it can be rendered and displayed on the terminal device's screen, or the graphical user interface can be presented through holographic projection.
[0020] It should be noted that, in this embodiment of the application, the entity executing the text deduplication method can be a terminal device or a server. This embodiment of the application does not limit the type of the entity executing the method.
[0021] For example, in conjunction with the above description, Figure 1 This application illustrates a text plagiarism detection system 1000 based on an embodiment of the present application. The plagiarism detection system 1000 may include at least one terminal device 100, at least one server 200, at least one database 300, and a network.
[0022] The user's terminal device 100 can connect to different servers via a network. The terminal device is any device with computing hardware capable of supporting and executing the software application tools corresponding to the plagiarism detection system 1000.
[0023] In the aforementioned plagiarism detection system 1000, the terminal device 100 is used to install and run the plagiarism detection application. In some cases, the plagiarism detection application may not need to be installed on the terminal device 100 in advance, and users can directly access the plagiarism detection system through a browser or other client.
[0024] Users log in to the plagiarism detection application using their registered accounts. When logging in, terminal device 100 sends a login request to server 200. Server 200 verifies the user's account, and if the verification is successful, it returns a login success notification to terminal device 100.
[0025] Furthermore, when the plagiarism detection system 1000 includes multiple terminal devices, multiple servers, and multiple networks, different terminal devices can connect to each other through different networks and different servers. The network can be a wireless network or a wired network; for example, wireless networks include Wireless Local Area Network (WLAN), Local Area Network (LAN), cellular network, 2G network, 3G network, 4G network, 5G network, etc.
[0026] In addition, different terminal devices can also use their own Bluetooth network or hotspot network to connect to other terminal devices or to servers.
[0027] In addition, the system 100 may include multiple databases, which are coupled to different servers for storing answer information.
[0028] It should be noted that, Figure 1 The schematic diagram of the plagiarism detection system shown is merely an example. The plagiarism detection system 1000 described in this application embodiment is intended to more clearly illustrate the technical solutions of this application embodiment and does not constitute a limitation on the technical solutions provided in this application embodiment. As those skilled in the art will know, with the evolution of plagiarism detection systems and the emergence of new business scenarios, the technical solutions provided in this application embodiment are also applicable to similar technical problems.
[0029] The technical solution of this application will be described in detail below through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0030] Please see Figure 2 The text plagiarism detection method of this application may include the following steps 011 to 013, which are described in detail below.
[0031] Step 011: Using various language processing models, perform feature processing on the text to be checked and the comparison text respectively, so as to output the first feature vector corresponding to the text to be checked and the second feature vector corresponding to each comparison text.
[0032] Among them, the language processing model is a sentence embedding model that can map texts of different languages to a unified vector space, enabling cross-language semantic alignment and unified semantic representation.
[0033] In some embodiments, multiple pre-trained multilingual models are used to perform feature processing on the text to be checked and the comparison text, respectively, so as to output the first dense feature vector corresponding to the text to be checked and the second dense feature vector corresponding to each comparison text.
[0034] Among them, the first dense feature vector and the second dense feature vector are low-dimensional dense vectors.
[0035] For example, a multilingual model could be a multilingual sentence embedding model within the Sentence-transformer framework. Various language processing models include the paraphrase-multilingual-MiniLM-L12-v2 model and the paraphrase-multilingual-mpnet-base-v2 model.
[0036] In some embodiments, a preset bag-of-words model is used to perform feature processing on the text to be checked and the comparison text, respectively, so as to output the first sparse feature vector corresponding to the text to be checked and the second sparse feature vector corresponding to each comparison text.
[0037] Among them, the first sparse feature vector and the second sparse feature vector are high-dimensional sparse vectors.
[0038] For example, a language processing model could be the Bag-of-Words (BoW) model, a simple text representation method that treats text as a "set of words," ignoring grammar and word order, and only counting the frequency of word occurrences to represent the semantics of the text.
[0039] The text to be checked for plagiarism refers to the text that needs to be detected for duplicates or plagiarism; it is the core object of the plagiarism detection task. The comparison text refers to the reference text used as the basis for duplicate detection; it is the "standard library / source library" for determining whether the text to be checked for plagiarism is duplicated.
[0040] Once the various language processing models are determined, each model can be used to perform feature processing on the text to be checked and the comparison text, mapping them to a unified vector space. This yields the first feature vector for the text to be checked and the second feature vectors for each comparison text. In this way, each language processing model can obtain its corresponding first and second feature vectors.
[0041] It is understandable that different language processing models can perform vector mapping from different dimensions, thereby extracting feature vectors of the text to be checked and the text to be compared from different dimensions, thus achieving multi-dimensional feature extraction.
[0042] For example, the paraphrase-multilingual-MiniLM-L12-v2 model and the paraphrase-multilingual-mpnet-base-v2 model differ in embedding dimension, semantic accuracy, and applicable scenarios.
[0043] For example, the multilingual sentence embedding model and the bag-of-words model under the Sentence-transformer framework differ in semantic understanding (the bag-of-words model does not need to understand semantics, while the multilingual sentence embedding model can), word order information (the bag-of-words model ignores word order, while the multilingual sentence embedding model needs to consider word order), vector properties (the bag-of-words model's vectors are high-dimensional sparse vectors, while the multilingual sentence embedding model's vectors are low-dimensional dense vectors), and representation ability (the bag-of-words model has weak representation ability and cannot distinguish between sentences that are semantically similar but have different word orders, while the multilingual sentence embedding model has strong representation ability and can distinguish between sentences that are semantically similar but have different word orders).
[0044] Thus, in the multi-model vector representation stage, multiple pre-trained Sentence-transformer multilingual models (such as the paraphrase-multilingual-MiniLM-L12-v2 model or the paraphrase-multilingual-mpnet-base-v2 model) are used to generate various dense vector representations to capture the semantic information of the text. These models, trained on a large-scale multilingual corpus, are capable of deeply understanding the semantic connotations of texts in different languages, effectively avoiding the loss of semantic information due to language differences. Simultaneously, to further enhance the richness and accuracy of the vector representation, traditional bag-of-words models (such as the TF-IDF method) are used to generate sparse vector representations to capture word frequency-based lexical features.
[0045] In one optional embodiment, the feature vectors corresponding to historical text (i.e., text that has been deduplicated) or comparison text can be added to a vector database (such as Milvus or Faiss) for storage and indexing, so as to facilitate rapid retrieval and comparison in the future and improve the efficiency of subsequent vector matching.
[0046] Step 012: Using a preset vector matching algorithm, match the first feature vector and the second feature vector to obtain the second text segment that matches the first text segment of the text to be checked for plagiarism in the comparison text.
[0047] Vector matching algorithms are used to calculate the similarity / distance between two or more vectors.
[0048] In one alternative embodiment, different vector matching algorithms are used to match vectors of different types. Alternatively, all language processing models can use the same vector matching algorithm, which can improve matching efficiency.
[0049] For example, for dense vectors (such as vectors output by multilingual models), cosine similarity matching algorithm and approximate matching algorithm (Hierarchical Navigable Small World, HNSW is currently the mainstream and high-performance approximate nearest neighbor (ANN) vector retrieval algorithm); for sparse vectors (such as vectors output by bag-of-words models), Manhattan distance matching algorithm and LSH (Locality-Sensitive Hashing) algorithm are preferred.
[0050] For example, all language processing models use cosine similarity matching algorithms for vector matching.
[0051] The first text segment is a part of the text to be checked for plagiarism, such as a clause or a paragraph of the text to be checked; the second text segment is a part of the text to be compared, such as a clause or a paragraph of the text to be compared.
[0052] Specifically, for each language processing model, a corresponding preset vector matching algorithm can be used to match the first feature vector of each first text segment with the second feature vector of the second text segment, thereby obtaining the second text segment matched by the first text segment of the text to be deduplicated in the comparison text.
[0053] For example, taking text fragments as clauses, a multi-level matching algorithm was designed in the similarity matching stage. Given two sets of clauses in the text to be checked for plagiarism and the text to be compared, for each clause in the text to be checked for plagiarism, vector indexing techniques are used, such as using the Milvus vector database to implement approximate nearest neighbor search, to quickly retrieve the second text fragments that match each first text fragment under different vector matching methods.
[0054] Please see Figure 3 In one optional embodiment, step 012 includes: Step 0121: Calculate the cosine similarity between the first feature vector of the first text segment and the second feature vector of the second text segment; Step 0122: Based on the similarity threshold and cosine similarity corresponding to each language processing model, determine the second text segment that matches the first text segment. Set the corresponding similarity threshold for each language processing model.
[0055] Specifically, the feature vectors generated after processing text differ for each language processing model. Therefore, due to these differences in feature vectors, even identical text segments may have different actual similarity scores during vector matching. Thus, for each language processing model, a reasonable matching threshold can be set to ensure the accuracy of vector matching across different models.
[0056] For each language processing model, after calculating the cosine similarity between the first feature vector of the first text segment and the second feature vector of the second text segment, it is possible to determine whether the first feature vector and the second feature vector match based on the similarity threshold corresponding to the model (e.g., whether the cosine similarity between the first feature vector and the second feature vector is greater than the similarity threshold), thereby accurately determining whether the first text segment and the second text segment match.
[0057] In an optional embodiment, the matching priorities of the multiple language processing models are different, and step 0122 includes: Step 01221: For the first text segment that is currently matched, determine the target language processing model corresponding to the current matching priority from high to low, and calculate the cosine similarity corresponding to the target language processing model; Step 01222: Determine the target cosine similarity among the cosine similarities corresponding to the first text segment currently matched, which is greater than the similarity threshold corresponding to the target language processing model, and determine whether the second text segment corresponding to the target cosine similarity matches the first text segment currently matched. Step 01223: If the cosine similarity of each of the currently matched first text segments is less than the similarity threshold corresponding to the target language processing model, then the current matching priority is reduced by one level, and the process re-enters the step of determining the target language processing model corresponding to the current matching first text segment according to the matching priority from high to low, and calculating the cosine similarity corresponding to the target language processing model.
[0058] It's understandable that if every first text segment were processed using all language processing models for feature processing and vector matching, the workload of feature processing and vector matching would be excessive, severely reducing plagiarism detection efficiency. However, without fully utilizing various language processing models, it's difficult to guarantee the accuracy and comprehensiveness of plagiarism detection. Therefore, a method to achieve a balance between plagiarism detection efficiency and accuracy is urgently needed.
[0059] Specifically, for the first text segment currently matched, the target language processing model corresponding to the current matching priority is determined from high to low according to the matching priority for feature processing, thereby extracting the corresponding first feature vector. Then, based on the first feature vector, the cosine similarity corresponding to the target language processing model is calculated, that is, the cosine similarity between the first feature vector and the second feature vector corresponding to the target language processing model in the compared text is calculated.
[0060] Then, determine the target cosine similarity among the cosine similarities corresponding to the first text segment currently matched, which is greater than the similarity threshold corresponding to the target language processing model. If a target cosine similarity exists, it can be determined that the second text segment corresponding to the target cosine similarity matches the first text segment currently matched. If there is no target cosine similarity, that is, the cosine similarity of each text segment matched at the moment is less than the similarity threshold of the target language processing model, then the current matching priority is reduced by one level, the lower-level language processing model is determined as the target language processing model, and based on the newly determined target language processing model, steps 01221 and 01222 are executed again until a second text segment matching the current first text segment is found, or all language processing models have performed vector matching.
[0061] The embodiments of this application prioritize different language processing models for matching, and then perform vector matching based on these priorities. Feature processing is performed sequentially from highest to lowest priority to generate corresponding feature vectors, and vector matching is then performed until a second text segment matching the first text segment is found. At this point, the feature processing and vector matching processes of subsequent language processing models based on matching priorities are stopped, thereby improving matching efficiency. Before a second text segment matching the first text segment is found, multiple language processing models can be fully utilized to achieve vector matching. This ensures accurate and complete matching of each first text segment while reducing unnecessary feature processing and vector matching processes, thus balancing plagiarism detection efficiency, accuracy, and comprehensiveness.
[0062] In an optional embodiment, step 0122 further includes: Step 01224: If all language processing models have completed matching, and the first text segment currently matched has no matching second text segment, then calculate the weighted average of the cosine similarity corresponding to each language processing model, and set the corresponding weight for each language processing model. Step 01225: Determine the target weighted average value among the weighted average values corresponding to the first text segment currently matched, which is greater than the corresponding weighted similarity threshold, and determine whether the second text segment corresponding to the target weighted average value matches the first text segment currently matched.
[0063] It is understandable that for the first text segment being matched, all language processing models may have already completed their matching process without finding a matching second text segment. Therefore, in order to maximize the matching success rate of the first text segment, [further steps are needed].
[0064] When all language processing models have completed matching and there is no matching second text segment for the first matched text segment, the weighted average of the cosine similarity of each language processing model can be calculated. Each language processing model is assigned a corresponding weight (e.g., the higher the matching priority, the greater the weight).
[0065] Finally, among the weighted averages corresponding to the currently matched first text segment, the target weighted average is determined to be greater than the corresponding weighted similarity threshold. If a target weighted average exists, the second text segment corresponding to the target weighted average is determined to match the currently matched first text segment. If no target weighted average exists, the currently matched first text segment is determined to have no matching second text segment.
[0066] In this way, by flexibly setting similarity thresholds and weight ratios, accurate matching of similar clauses can be achieved, and the similarity matching results of each clause in the text to be checked for plagiarism can be obtained in the comparison text.
[0067] Step 013: Based on the second text segment matched by the first text segment, determine the plagiarism detection result of the text to be checked. The plagiarism detection result includes the target text matched by the text to be checked in the comparison text.
[0068] Specifically, after determining that there is a matching second text fragment in the text to be checked for plagiarism, the plagiarism result of the text to be checked can be determined. It can be determined whether there is a matching target text in the text to be checked. If there is, the text to be checked can be determined to be plagiarized text. If there is no matching target text, the text to be checked can be determined to be original text.
[0069] Please continue reading. Figure 3 In one optional embodiment, step 013 includes: Step 0131: Based on the second text segment matched by the first text segment, determine the matching results between the text to be checked for plagiarism and each comparison text; Specifically, based on the matching results between the text to be checked and each comparison text (such as the first and second matching text segments), it can be determined whether the text to be checked matches each comparison text. If they match, the matching comparison text is the target text that the text to be checked plagiarized. If there is no matching comparison text, the text to be checked is the original text.
[0070] In an optional embodiment, the matching result includes at least one of text segment matching ratio, paragraph matching ratio, matching length ratio, and average similarity. Step 0131: Based on the second text segment matched by the first text segment, determine the matching result between the text to be checked for plagiarism and each comparison text, including: Step 01311: Determine the proportion of the number of first text fragments that have matching second text fragments in the target comparison text to the total number of text fragments. This proportion is the text fragment matching ratio corresponding to the target comparison text. The target comparison text is any comparison text, and the total number of text fragments is the larger of the total number of first text fragments and the total number of second text fragments. Step 01312: Determine the proportion of the first paragraph of the text to be checked that has a matching second paragraph in the target comparison text to the total number of paragraphs. This is the paragraph matching ratio corresponding to the target comparison text. Among the matching first paragraph and second paragraph, the matching ratio of the first text fragment of the first paragraph and the second text fragment of the second paragraph is greater than the preset ratio. The total number of paragraphs is the larger of the total number of the first paragraph and the total number of the second paragraph. Step 01313: Determine the average ratio of the first proportion of the total number of words in all matched first text segments to the total number of words in all first text segments and the second proportion of the total number of words in all matched second text segments to the total number of words in all second text segments. This ratio is the matching length ratio corresponding to the target text. Step 01314: Determine the average cosine similarity of all matching first and second text segments as the average similarity of the target text.
[0071] The text fragment matching ratio is used to characterize the proportion of matching first and second text fragments in the text to be checked for plagiarism and the comparison text. For example, the proportion of first text fragments that have matching second text fragments in the target comparison text (any comparison text) to the total number of text fragments can be determined as the text fragment matching ratio corresponding to the target comparison text.
[0072] It is understandable that, in order to avoid an excessively high matching ratio, the larger of the total number of the first text fragments and the total number of the second text fragments can be used as the total number of text fragments.
[0073] The paragraph matching ratio is used to characterize the proportion of matching first and second paragraphs in the text to be checked for plagiarism and the comparison text. For example, the proportion of the first paragraph of the text to be checked that has a matching second paragraph in the target comparison text to the total number of paragraphs can be determined as the paragraph matching ratio of the target comparison text.
[0074] Specifically, in the matching first paragraph and second paragraph, the matching ratio of the first text segment of the first paragraph and the second text segment of the second paragraph is greater than a preset ratio (such as 80%, 90%, etc.). Matching the first paragraph and the second paragraph means that the first paragraph and the second paragraph are basically the same.
[0075] The matching length ratio is the ratio of the number of words matched in the text to be checked for plagiarism to the number of words matched in the comparison text. For example, the average ratio of the total number of words in all matched first text segments to the total number of words in all first text segments, and the average ratio of the total number of words in all matched second text segments to the total number of words in all second text segments, is the matching length ratio corresponding to the target comparison text. Thus, using the average ratio of the first ratio of the text to be checked and the second ratio of the comparison text as the matching length ratio improves the accuracy of the matching length ratio.
[0076] The average similarity is the average of the cosine similarities of the first and second text segments in each pair of matches, which can approximate the text similarity between the text to be checked and the comparison text. The higher the average similarity, the more similar the text to be checked and the comparison text are.
[0077] Currently, deep learning-based methods utilize neural network models for feature extraction and similarity calculation of text, which improves the accuracy of plagiarism detection to some extent. However, these models have limited processing capabilities for long texts. Directly inputting long texts for calculation often results in high computational resource consumption and low processing efficiency. Furthermore, directly encoding long texts into a single vector can easily lose semantic details, leading to a lack of interpretability in the overall text similarity calculation. Therefore, by employing a multi-dimensional similarity calculation method (i.e., multiple matching results), this approach overcomes the limitations of traditional plagiarism detection methods that rely solely on single-dimensional similarity calculations. Simultaneously, it provides richer information to the plagiarism detection results, enhancing their interpretability.
[0078] Step 0132: Determine the plagiarism detection result based on the matching results between the text to be checked and each comparison text.
[0079] After determining the matching results between the text to be checked and each comparison text, the plagiarism detection result can be determined. This can be based on factors such as the matching ratio of text segments, paragraph matching ratio, matching length ratio, and average similarity for each comparison text.
[0080] In an optional embodiment, step 0132: Step 01321: Calculate the weighted matching values of various matching results for the text to be checked for plagiarism and each comparison text, and set the corresponding weights for each matching result; Step 01322: Determine the comparison text corresponding to the largest weighted matching value as the target text for matching the text to be checked for plagiarism; or, determine the comparison text corresponding to the weighted matching value that is greater than the preset matching value threshold as the target text for matching the text to be checked for plagiarism.
[0081] Specifically, different weights can be assigned to each matching result, and the weight settings can be flexibly adjusted according to the actual plagiarism detection scenario and needs.
[0082] Then, based on the weighted matching value of various matching results corresponding to each comparison text, the weighted matching value of each comparison text is compared. The larger the weighted matching value, the more similar the corresponding comparison text and the text to be checked for plagiarism are.
[0083] When determining the most similar target text, the text corresponding to the highest weighted matching value can be identified as the target text for the text to be checked for plagiarism. In this way, the plagiarism check results will at least output the most similar comparison text as the target text. Then, the target text and the text to be checked for plagiarism are compared manually to finally determine whether the text to be checked for plagiarism is suspected of plagiarism. This can reduce the burden on plagiarism checkers and improve the efficiency and accuracy of plagiarism checks.
[0084] Alternatively, the text corresponding to a weighted matching value greater than a preset matching threshold (such as 70%, 80%, etc.) can be identified as the target text for the text to be checked for plagiarism. In this way, one or more target texts that are relatively close to the text to be checked can be selected. Subsequently, the target text and the text to be checked can be compared manually to finally determine whether the text to be checked is suspected of plagiarism. This can reduce the burden on plagiarism checkers and improve the efficiency and accuracy of plagiarism checking.
[0085] The text plagiarism detection method in this application extracts feature vectors from both the text to be checked and the comparison text using multiple language processing models, achieving multi-dimensional feature extraction. Then, for each model's feature vector, vector matching is performed to achieve multi-dimensional similarity matching, improving matching accuracy and avoiding omissions. Finally, based on the accurate and complete matching of the first and second text segments, plagiarism detection of the text to be checked is achieved, thereby improving the accuracy of subsequent plagiarism detection results determined based on the matching results.
[0086] In some embodiments, to improve matching accuracy, the text to be checked and the text to be compared can be preprocessed and segmented into clauses.
[0087] It is understandable that for long texts, i.e. long document content, existing technologies usually adopt a block-segmentation scheme, which divides the text into fixed-length blocks, encodes them separately, and then performs brute-force comparison. However, the effect is limited by the block segmentation method, and the computational complexity of brute-force comparison is high, resulting in low efficiency.
[0088] In the text preprocessing and intelligent segmentation stages, this application first performs text preprocessing on the text to be checked or compared, including noise removal and text formatting to ensure the text quality meets the requirements of subsequent processing. For example, for Markdown formatted text, formatting marks, redundant blank lines, and whitespace characters are removed, retaining only valid text content. Subsequently, intelligent segmentation is performed, adaptively dividing long texts into several logically coherent paragraph units based on the document structure. During this process, the text's structural organization (such as headings) and the use of delimiters and punctuation are comprehensively considered to ensure that each paragraph unit contains relatively complete information, while avoiding grouping content from different topics into the same unit.
[0089] Next, for each paragraph unit, clause segmentation is further implemented using clause segmentation techniques (such as the pySBD rule-based sentence boundary detection library in Python). Invalid clauses, such as those containing only punctuation marks or lacking actual semantic meaning, are identified and filtered using natural language processing techniques (such as part-of-speech tagging or using a trained text classification model), ensuring that subsequent processing targets only valid text content.
[0090] During this process, the system retains the original paragraph information for each clause. Ultimately, each given text will generate a set of valid clauses and their corresponding paragraphs, and a word segmenter (such as the jieba library in Python) will be used to count the number of words in each clause. This facilitates subsequent calculations of the matching length ratio.
[0091] This application employs a multi-stage processing workflow, including four stages: text preprocessing and intelligent segmentation, multi-model vector representation, multi-level similarity matching, and result integration and output. Through this multi-stage workflow, the present invention can comprehensively and accurately assess the similarity between texts while maintaining efficient processing capabilities for long texts.
[0092] Based on the method described in the above embodiments, this application also provides a text plagiarism detection device 300 for answer sheets, used to perform the steps in the above text plagiarism detection method. Please refer to... Figure 4 , Figure 4 This is a schematic diagram of the modules of the text plagiarism detection device 300 provided in this application embodiment. The text plagiarism detection device 300 includes: The feature processing module 301 is used to perform feature processing on the text to be checked and the comparison text through various language processing models, so as to output the first feature vector corresponding to the text to be checked and the second feature vector corresponding to each comparison text. The matching module 302 is used to match the first feature vector and the second feature vector respectively based on a preset vector matching algorithm to obtain the second text segment matched in the comparison text by the first text segment of the text to be deduplicated. The determination module 303 is used to determine the plagiarism detection result of the text to be checked based on the second text segment matched by the first text segment. The plagiarism detection result includes the target text matched by the text to be checked in the comparison text.
[0093] In the embodiments of this application, the terms "module" or "unit" refer to computer instructions or a portion of computer instructions that have a predetermined function and work together with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0094] In some embodiments, the text plagiarism detection device in this application can be implemented in hardware, such as a terminal device or a component in the terminal device, such as an integrated circuit or a chip; the text plagiarism detection device can also be implemented in software, such as as an application installed in the terminal device.
[0095] In some embodiments, please refer to Figure 5 , Figure 5 This is a schematic diagram of the terminal device provided in the embodiments of this application. The terminal device 500 includes a processor 501, a memory 502, and a display screen 503. The memory 502 stores computer instructions 504 that can be executed by the processor 501. When the processor 501 executes the instructions 504, it implements the various processes of the above-described text plagiarism detection method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0096] Please see Figure 6 , Figure 6 This is a schematic diagram of the hardware structure of a terminal device provided in an embodiment of this application. The terminal device can be a terminal or a server. Exemplarily, the terminal device 700 includes a central processing unit (CPU) 701, a system memory 704 including random access memory (RAM) 702 and read-only memory (ROM) 703, and a system bus 705 connecting the system memory 704 and the central processing unit 701.
[0097] In some embodiments, the terminal device 700 may further include a basic input / output system 706 that helps transmit information between various devices within the computer, and a large-capacity storage device 707 for storing the operating system 713, the client 714, and other program modules 715.
[0098] In some embodiments, the basic input / output system 706 includes a display 708 for displaying a graphical user interface and input devices 709 for user input, such as touch panels and other input devices. A touch panel is also called a touchscreen. A touch panel may include both a touch device and a touch controller. Other input devices may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described further here.
[0099] Both the display 708 and the input device 709 are connected to the central processing unit 701 via an input / output controller 710 connected to the system bus 705. The basic input / output system 706 may also include the input / output controller 710 for receiving and processing input from touch panels, other input devices, etc. Similarly, the input / output system 706 also includes output devices such as displays, printers, or other types of output devices.
[0100] Mass storage device 707 is connected to central processing unit 701 via a mass storage controller (not shown) connected to system bus 705. Mass storage device 707 and its associated computer-readable media provide non-volatile storage for terminal device 700. That is, mass storage device 707 may include computer-readable media (not shown) such as hard disk or compact disc read-only memory (CD-ROM) drive.
[0101] According to various embodiments of this application, the terminal device 700 can also be connected to a remote computer on a network, such as the Internet. That is, the terminal device 700 can be connected to the network 717 via the network interface unit 716 connected to the system bus 705, or the network interface unit 716 can be used to connect to other types of networks or remote computer systems (not shown).
[0102] This application also provides a non-transitory computer-readable storage medium storing computer instructions. When these computer instructions are executed by a processor, they implement the various processes of the above-described text plagiarism detection method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0103] The processor can be the processor in the terminal device described in the above embodiments. The computer-readable storage medium can be a computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc.
[0104] Computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types.
[0105] This application also provides a computer instruction product, including computer instructions that, when executed by a processor, implement the aforementioned text plagiarism detection method. The processor may be the processor in the terminal device described in the above embodiments. When executed by the processor, the computer instructions implement various processes of the embodiments of the aforementioned text plagiarism detection method and achieve the same technical effects; therefore, to avoid repetition, they will not be described again here.
[0106] It is understood that in the specific implementation of this application, data related to user identity or characteristics is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
Claims
1. A text plagiarism detection method, characterized in that, include: Using various language processing models, feature processing is performed on the text to be checked and the comparison text to output the first feature vector corresponding to the text to be checked and the second feature vector corresponding to each of the comparison texts. The first feature vector and the second feature vector are matched using a preset vector matching algorithm to obtain the second text segment that matches the first text segment of the text to be checked for plagiarism in the comparison text. Based on the second text segment matched by the first text segment, the plagiarism detection result of the text to be checked is determined, and the plagiarism detection result includes the target text matched by the text to be checked in the comparison text.
2. The plagiarism detection method according to claim 1, characterized in that, The process involves performing feature processing on the text to be checked and the comparison text using various language processing models to output a first feature vector corresponding to the text to be checked and a second feature vector corresponding to each of the comparison texts, including: Using multiple pre-trained multilingual models, feature processing is performed on the text to be checked and the comparison text respectively, so as to output the first dense feature vector corresponding to the text to be checked and the second dense feature vector corresponding to each of the comparison texts; Using a pre-defined bag-of-words model, feature processing is performed on the text to be checked and the comparison texts to be compared, so as to output the first sparse feature vector corresponding to the text to be checked and the second sparse feature vector corresponding to each of the comparison texts.
3. The plagiarism detection method according to claim 1 or 2, characterized in that, The text to be checked for plagiarism includes multiple first text segments, and the comparison text includes multiple second text segments. The step of matching the first feature vectors and second feature vectors using a preset vector matching algorithm to obtain the second text segments in the comparison text that match the first text segments of the text to be checked for plagiarism includes: Calculate the cosine similarity between the first feature vector of the first text segment and the second feature vector of the second text segment to determine the second text segment that matches the first text segment.
4. The plagiarism detection method according to claim 3, characterized in that, The step of calculating the cosine similarity between the first feature vector of the first text segment and the second feature vector of the second text segment to determine the second text segment that matches the first text segment includes: For each of the aforementioned language processing models, the cosine similarity between the first feature vector of the first text segment and the second feature vector of the second text segment is calculated; Based on the similarity threshold corresponding to each of the language processing models and the cosine similarity, the second text segment that matches the first text segment is determined, and the multiple language processing models are respectively set with corresponding similarity thresholds.
5. The plagiarism detection method according to claim 3, characterized in that, The matching priorities of the multiple language processing models are different. The step of calculating the cosine similarity between the first feature vector of the first text segment and the second feature vector of the second text segment to determine the second text segment that matches the first text segment includes: For the first text segment that is currently matched, the target language processing model corresponding to the current matching priority is determined according to the matching priority from high to low, and the cosine similarity corresponding to the target language processing model is calculated. Determine the target cosine similarity among the cosine similarities corresponding to the first text segment currently matched, which is greater than the similarity threshold corresponding to the target language processing model, and determine that the second text segment corresponding to the target cosine similarity matches the first text segment currently matched. If the cosine similarity of each of the currently matched first text segments is less than the similarity threshold corresponding to the target language processing model, then the current matching priority is reduced by one level, and the process re-enters the steps of determining the target language processing model corresponding to the current matching priority according to the matching priority from high to low for the currently matched first text segment, and calculating the cosine similarity corresponding to the target language processing model.
6. The plagiarism detection method according to claim 5, characterized in that, The step of calculating the cosine similarity between the first feature vector of the first text segment and the second feature vector of the second text segment to determine the second text segment that matches the first text segment further includes: If all the language processing models have completed the matching, and the currently matched first text segment has no matching second text segment, then the weighted average of the cosine similarity corresponding to each language processing model is calculated, and each language processing model is assigned a corresponding weight. Among the weighted average values corresponding to the first text segment currently matched, determine the target weighted average value that is greater than the corresponding weighted similarity threshold, and determine that the second text segment corresponding to the target weighted average value matches the first text segment currently matched.
7. The plagiarism detection method according to claim 1, characterized in that, The step of determining the plagiarism detection result of the text to be checked based on the second text segment matched with the first text segment includes: Based on the second text segment matched by the first text segment, the matching results of the text to be deduplicated and each of the comparison texts are determined; The plagiarism detection result is determined based on the matching results of the text to be checked and each of the comparison texts.
8. The plagiarism detection method according to claim 7, characterized in that, The matching result includes at least one of text segment matching ratio, paragraph matching ratio, matching length ratio, and average similarity. The step of determining the matching result between the text to be deduplicated and each of the comparison texts based on the second text segment matched by the first text segment includes: The proportion of the number of first text segments that have a matching second text segment in the target comparison text to the total number of text segments is determined as the text segment matching ratio corresponding to the target comparison text. The target comparison text is any of the comparison texts. The total number of text segments is the larger of the total number of first text segments and the total number of second text segments. The proportion of the first paragraph of the text to be deduplicated, in which a matching second paragraph exists in the target comparison text, is determined to be the paragraph matching ratio corresponding to the target comparison text. Among the matched first paragraph and second paragraph, the matching ratio of the first text segment of the first paragraph and the second text segment of the second paragraph is greater than a preset ratio. The total number of paragraphs is the larger of the total number of the first paragraph and the total number of the second paragraph. The average ratio of the total number of words contained in all matched first text segments to the total number of words in all first text segments and the average ratio of the total number of words contained in all matched second text segments to the total number of words in all second text segments is determined as the matching length ratio corresponding to the target comparison text. The average cosine similarity of all matching first and second text segments is determined as the average similarity corresponding to the target text.
9. The plagiarism detection method according to claim 7, characterized in that, The matching result includes at least one of text segment matching ratio, paragraph matching ratio, matching length ratio, and average similarity. Determining the plagiarism detection result based on the matching result between the text to be checked and each of the compared texts includes: Calculate the weighted matching values of various matching results for the text to be checked for plagiarism and each of the comparison texts, and set corresponding weights for each of the various matching results; The comparison text corresponding to the largest weighted matching value is determined as the target text that matches the text to be deduplicated; or, the comparison text corresponding to the weighted matching value that is greater than a preset matching value threshold is determined as the target text that matches the text to be deduplicated.
10. A text plagiarism detection device, characterized in that, The device includes: The feature processing module is used to perform feature processing on the text to be checked and the comparison text through various language processing models, so as to output the first feature vector corresponding to the text to be checked and the second feature vector corresponding to each of the comparison texts. The matching module is used to match the first feature vector and the second feature vector based on a preset vector matching algorithm, so as to obtain the second text segment that matches the first text segment of the text to be deduplicated in the comparison text. The determining module is used to determine the plagiarism detection result of the text to be checked based on the second text segment matched by the first text segment, wherein the plagiarism detection result includes the target text matched by the text to be checked in the comparison text.
11. An electronic device, characterized in that, include: Memory and processor; The memory is used to store one or more computer instructions; The processor is configured to execute one or more computer instructions to implement the method as described in any one of claims 1-9.
12. A computer-readable storage medium storing one or more computer instructions thereon, characterized in that, The instruction is executed by the processor to implement the method as described in any one of claims 1-9.