Text detection method and device, electronic equipment and computer readable storage medium
By dividing the text to be detected into fragments and using the search tree and type prediction model, the problem that the prior art cannot detect deep processing of text plagiarism is solved, and the precise distinction between the types of text plagiarism is achieved.
Patent Information
- Application Number
- CN202510466943.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art cannot effectively detect plagiarism of deeply processed texts, and cannot accurately distinguish the types of plagiarism of texts.
By dividing the text to be detected into fragments to be detected, its vector representation is determined, and a reference fragment with the highest semantic similarity is found based on the pre-constructed search tree. Then, based on the semantic similarity and TF-IDF value, the target similarity is determined, and the fragment with the target similarity greater than the threshold is selected as the target detection fragment, and input it into the type prediction model to obtain the plagiarism type.
It realizes the precise distinction between the types of plagiarism treated by detecting text, and solves the problem that plagiarists cannot be detected by deeply processing the original text.
Smart Images

Figure CN120012754A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of natural language processing, and in particular to a text detection method, device, electronic device and computer-readable storage medium. Background Art
[0002] Plagiarism is a more hidden and complex form of academic misconduct. Its core lies in the fact that the plagiarist deeply processes the original content through semantic rewriting tools. This behavior is superficially different from direct text copying, but it retains the core ideas and logical structure of the original content, making it difficult to be identified by traditional text matching detection systems.
[0003] With the rapid development of generative artificial intelligence technology, large-scale language models represented by GPT-4 have high-level semantic understanding and generation capabilities. These models can rewrite input texts, not only replacing words and phrases, but also reorganizing sentence structures and adjusting expression logic, thereby generating texts with high fluency and coherence.
[0004] Online rewriting tools also have similar functions to a certain extent, using simple vocabulary replacement or syntactic adjustment to process the input text. Although the operation of such tools is relatively rudimentary, they are able to bypass some duplicate checking systems based on string matching. In contrast, the rewritten text generated by the large model based on deep learning is more semantically consistent and closer to human creation in terms of expression form, which significantly improves the concealment and disguise of plagiarism.
[0005] Therefore, related text detection technologies have the problem of being unable to detect plagiarism in deeply processed texts, and being unable to accurately distinguish the types of plagiarism in texts. Summary of the invention
[0006] The embodiments of the present application provide a text detection method, device, electronic device and computer-readable storage medium, which are used to solve the technical problems of being unable to detect plagiarism of deeply processed texts and being unable to accurately distinguish the types of plagiarism of texts.
[0007] According to a first aspect of an embodiment of the present application, a text detection method is provided, the method comprising: dividing a text to be detected into at least one segment to be detected, determining a vector representation of the segment to be detected for each segment to be detected, and determining a reference segment corresponding to the segment to be detected from the retrieval tree based on the vector representation and a pre-constructed retrieval tree; the reference segment is a segment in the retrieval tree with the highest semantic similarity to the segment to be detected, and the retrieval tree is composed of vector representations of text segments of all texts in a corpus; For each segment to be detected, the target similarity of the segment to be detected is determined according to the semantic similarity between the segment to be detected and the corresponding reference segment and the TF-IDF value of the segment to be detected; the TF-IDF value is used to characterize the importance of the words in the segment to be detected in the reference text, and the target similarity is used to characterize the similarity between the segment to be detected and the corresponding reference segment; From each to-be-detected segment, a to-be-detected segment whose target similarity is greater than a preset similarity threshold is selected as a target detection segment; For each target detection segment, input the target detection segment and a reference segment corresponding to the target detection segment into a type prediction model, obtain a type of plagiarism corresponding to the target detection segment output by the type prediction model, and generate a detection result, the detection result including the plagiarism type corresponding to each target detection segment; Among them, the type prediction model is trained with the sample segment and a type of plagiarized segment corresponding to the sample segment as training samples, and the plagiarism type of the plagiarized segment as training labels.
[0008] In a possible implementation, after obtaining a type of plagiarism corresponding to the target detection segment output by the type prediction model, for each target detection segment, according to the target detection segment and the corresponding reference segment, the context matching degree between the context of the target detection segment and the context of the reference segment is determined; the context matching degree is used to characterize the matching degree between the context of the target detection segment and the reference segment; For each target detection segment, the semantic similarity between each semantic unit in the target detection segment and each semantic unit in the corresponding reference segment is calculated, and a comparison matrix corresponding to the target detection segment is constructed; the elements in the comparison matrix are the semantic similarities between the semantic units of the target detection segment and the semantic units of the reference text; For each target detection segment, a plagiarism report is generated according to the plagiarism type corresponding to the target detection segment, the contextual matching degree between the target detection segment and the reference segment, and the comparison matrix corresponding to the target detection segment.
[0009] In another possible implementation, the target detection segment is divided into a plurality of first windows in a sliding window manner according to a predefined window size, and the words in each first window are converted into vector representations; For each first window of the target detection segment, determining a semantic similarity between the first window and a corresponding reference segment according to a vector representation of the first window; The context matching degree between the target detection segment and the reference segment is determined according to the semantic similarity corresponding to each first window of the target detection segment.
[0010] In another possible implementation, the plagiarism report includes the overall plagiarism situation and the partial plagiarism situation of the text to be detected, the partial plagiarism situation refers to the proportion of plagiarized text in each target detection segment, the overall plagiarism situation refers to the proportion of plagiarized text in the full text of the text to be detected, and the plagiarized text refers to the text obtained after the corresponding plagiarism type in the target detection segment plagiarizes the reference segment; For each target detection segment, obtain the target elements greater than the element similarity threshold from the comparison matrix corresponding to the target detection segment; For each target element, the semantic unit in the target detection segment corresponding to the target element is taken as the plagiarized text; Determine the first character number of the plagiarized text and the second character number of the target detection segment, and determine the local plagiarism situation according to the first character number and the second character number; The overall plagiarism situation is determined based on the number of first characters of the plagiarized text and the number of third characters of the text to be detected in each target detection segment.
[0011] In another possible implementation, the text to be detected is segmented to obtain at least one keyword contained in the text to be detected; Get the reference text corresponding to the reference fragment; For each keyword, determine the frequency of the keyword in the segment to be detected, and determine the number of segments in the reference text that contain the keyword; The TF-IDF value of the keyword in the reference text is determined based on the frequency, the total number of fragments in the reference text, and the number of fragments in the reference text that contain the keyword.
[0012] In yet another possible implementation, for each text in the corpus, the text is segmented into a plurality of text segments, and for each text segment, a vector representation of the text segment is determined; A leaf node of the retrieval tree is formed based on the vector representation of each text fragment of each text; Perform multiple rounds of recursion on all leaf nodes until any recursive stopping condition is met; Each round of recursion includes: Get each element of this round of recursion. The element of the first round of recursion is the text fragment; Clustering all elements through a clustering algorithm to obtain at least one cluster; For each cluster, determine the vector representation of the cluster according to the vector representation of the elements contained in the cluster; The clusters are used as the elements of the next round of recursion, and the elements obtained in this round of recursion are used as nodes of a new layer in the retrieval tree. The nodes corresponding to the elements obtained in this round of recursion are located in the upper layer of the corresponding nodes in the previous round in the retrieval tree.
[0013] In yet another possible implementation, the recursive stop condition includes any of the following: The retrieval tree reaches the preset number of layers; The number of clusters is lower than the preset threshold; The similarity between any two clusters is lower than the preset cluster similarity threshold.
[0014] In yet another possible implementation, the plagiarism type output by the type prediction model includes at least one of the following: copy; semantic substitution; Restructuring; The plagiarized segment with the plagiarism type of copying is obtained by copying the sample segment; The plagiarism segment with the plagiarism type of semantic replacement is obtained by rewriting the sample segment through the synonym replacement function; The plagiarized fragment with structural recombination type is obtained by recombining the sequence of the sample fragment.
[0015] According to a second aspect of an embodiment of the present application, a text detection device is provided, the device comprising: A division module is used to divide the text to be detected into at least one segment to be detected, determine the vector representation of each segment to be detected, and determine the reference segment corresponding to the segment to be detected from the retrieval tree based on the vector representation and a pre-built retrieval tree; the reference segment is the segment with the highest semantic similarity with the segment to be detected in the retrieval tree, and the retrieval tree is composed of the vector representations of the text segments of all the texts in the corpus; A determination module is used to determine, for each segment to be detected, a target similarity of the segment to be detected according to the semantic similarity between the segment to be detected and the corresponding reference segment and the TF-IDF value of the segment to be detected; the TF-IDF value is used to characterize the importance of the words in the segment to be detected in the reference text, and the target similarity is used to characterize the similarity between the segment to be detected and the corresponding reference segment; A selection module, used for selecting, from among the segments to be detected, segments to be detected whose target similarity is greater than a preset similarity threshold as target detection segments; An input module is used to input the target detection segment and the reference segment corresponding to the target detection segment into a type prediction model for each target detection segment, obtain a type of plagiarism corresponding to the target detection segment output by the type prediction model, and generate a detection result, wherein the detection result includes the plagiarism type corresponding to each target detection segment; Among them, the type prediction model is trained with the sample segment and a type of plagiarized segment corresponding to the sample segment as training samples, and the plagiarism type of the plagiarized segment as training labels.
[0016] According to a third aspect of an embodiment of the present application, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory, and when the processor executes the program, the steps of the method provided in the first aspect are implemented.
[0017] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method provided in the first aspect are implemented.
[0018] According to the fifth aspect of the embodiment of the present application, a computer program product is provided, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. When a processor of a computer device reads the computer instructions from the computer-readable storage medium, the processor executes the computer instructions, so that the computer device executes the steps of implementing the method provided in the first aspect.
[0019] The beneficial effects of the technical solution provided by the embodiment of the present application are: The text detection method provided by the embodiment of the present application divides the to-be-detected text into at least one to-be-detected segment, determines the reference segment with the highest semantic similarity to each to-be-detected segment from the retrieval tree of the pre-constructed corpus according to the vector representation of each to-be-detected segment, determines the target similarity characterizing the degree of similarity between the to-be-detected segment and the corresponding reference segment according to the semantic similarity between the to-be-detected segment and the corresponding reference segment and the TF-IDF value of the to-be-detected segment, selects the to-be-detected segment with the target similarity greater than a preset similarity threshold as the target detection segment, and for each target detection segment, inputs the target detection segment into a pre-trained type prediction model to obtain the plagiarism type corresponding to the target detection segment output by the type prediction model, thereby obtaining a detection result containing the plagiarism type corresponding to each target detection segment. Since the type prediction model is trained with sample text and plagiarized text as training samples and the plagiarism type corresponding to the plagiarized text as training labels, accurate distinction of the plagiarism type of the to-be-detected text is achieved, and the problem that the plagiarist cannot be detected for plagiarism by deep processing of the original text is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in describing the embodiments of the present application are briefly introduced below.
[0021] Figure 1 A schematic diagram of a system architecture for implementing a text detection method provided in an embodiment of the present application; Figure 2 A flowchart of a text detection method provided in an embodiment of the present application; Figure 3A schematic diagram of a process for generating a plagiarism report in a text detection method provided in an embodiment of the present application; Figure 4 A flowchart of a method for determining a context matching degree in a text detection method provided in an embodiment of the present application; Figure 5 A flowchart of a method for obtaining a TF-IDF value in a text detection method provided in an embodiment of the present application; Figure 6 A schematic diagram of a process for constructing a retrieval tree in a text detection method provided in an embodiment of the present application; Figure 7 A schematic diagram of each round of recursion in a method for constructing a retrieval tree provided in an embodiment of the present application; Figure 8 A flowchart of a method for determining local plagiarism and overall plagiarism in a text detection method provided in an embodiment of the present application; Fig. 9 A structural schematic diagram of a text detection device provided in an embodiment of the present application; Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] The embodiments of the present application are described below in conjunction with the drawings in the present application. It should be understood that the implementation methods described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0023] It will be understood by those skilled in the art that, unless specifically stated, the singular forms "one", "an" and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application refer to that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation as other features, information, data, steps, operations, elements, components and / or combinations thereof supported by the technical field. It should be understood that when we call an element "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can refer to that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here may include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or as "B", or as "A and B".
[0024] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0025] The following is an explanation of the relevant technology: The current mainstream text detection method mainly relies on string matching algorithms. For example, by comparing the string overlap rate of the text to be detected with the existing documents in the database, the text similarity is judged. This method has a good detection effect on plagiarism such as direct copying or slight rewriting; alternatively, the comparison is performed in units of sentences or short texts of fixed length.
[0026] However, currently, large models such as GPT-4 can be used to semantically rewrite and reconstruct texts. Even if the words or phrases in the original text are significantly changed, the semantics of sentences or even paragraphs can remain consistent. Therefore, since the relevant text detection methods rely on literal matching, when the sentence structure and vocabulary of the text are completely changed, the relevant technology cannot effectively perceive semantic plagiarism, and there is a problem of being unable to detect plagiarism in deeply processed texts, and being unable to accurately distinguish the types of plagiarism in texts.
[0027] In view of at least one of the above-mentioned technical problems or areas that need improvement in the related art, the present application proposes a text detection method, which divides the text to be detected into at least one segment to be detected, determines the reference segment with the highest semantic similarity to each segment to be detected from the retrieval tree of the pre-constructed corpus according to the vector representation of each segment to be detected, and determines the target similarity representing the degree of similarity between the segment to be detected and the corresponding reference segment according to the semantic similarity between the segment to be detected and the corresponding reference segment and the TF-IDF value of the segment to be detected, selects the segment to be detected with the target similarity greater than the preset similarity threshold as the target detection segment, and for each target detection segment, inputs the target detection segment into a pre-trained type prediction model to obtain the plagiarism type corresponding to the target detection segment output by the type prediction model, thereby obtaining a detection result containing the plagiarism type corresponding to each target detection segment. Since the type prediction model is trained with sample text and plagiarized text as training samples and the plagiarism type corresponding to the plagiarized text as training labels, it realizes the accurate distinction of the plagiarism type of the text to be detected, and solves the problem that the plagiarist cannot be detected for plagiarism by deep processing of the original text.
[0028] The following describes several exemplary embodiments to illustrate the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application. It should be noted that the following embodiments can refer to, draw on or combine with each other, and the same terms, similar features and similar implementation steps in different embodiments will not be described repeatedly.
[0029] Figure 1A schematic diagram of a system architecture for implementing a text detection method provided in an embodiment of the present application, wherein the system architecture includes: a terminal 120 and a server 140.
[0030] The terminal 120 installs and runs an application program having a text detection method, and the terminal 120 is used to determine the detection result of the text to be detected.
[0031] The terminal 120 is connected to the server 140 via a wireless network or a wired network.
[0032] The server 140 includes at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Schematically, the server 140 includes a processor 144 and a memory 142, and the memory 142 includes a display module 1421, a control module 1422, and a receiving module 1423. The server 140 is used to provide background services for the application of the method. Optionally, the server 140 undertakes the main computing work and the terminal 120 undertakes the secondary computing work; or, the server 140 undertakes the secondary computing work and the terminal 120 undertakes the main computing work; or, the server 140 and the terminal 120 adopt a distributed computing architecture for collaborative computing.
[0033] Optionally, the device types of the terminal include: at least one of a smart phone, a tablet computer, an e-book reader, a Moving Picture Experts Group Audio Layer III (MP3) player, a Moving Picture Experts Group Audio Layer IV (MP4) player, a laptop computer and a desktop computer.
[0034] Those skilled in the art will appreciate that the number of the above terminals may be more or less. For example, the above terminal may be only one, or the above terminals may be dozens or hundreds, or more. The embodiment of the present application does not limit the number of terminals and device types.
[0035] The present application provides a text detection method, such as Figure 2 As shown, the method includes: S101, divide the text to be detected into at least one segment to be detected, determine the vector representation of each segment to be detected, and based on the vector representation and a pre-constructed retrieval tree, determine the reference segment corresponding to the segment to be detected from the retrieval tree.
[0036] In an embodiment of the present application, the text to be detected refers to the text to be detected for plagiarism. Since the text to be detected is usually a long text, in order to accurately detect whether there is plagiarism in each part of the text to be detected, when performing plagiarism detection on the text to be detected, the text to be detected will be divided into at least one segment to be detected, and then the vector representation of each segment to be detected will be determined. The vector representation of the segment to be detected can be obtained through a pre-trained large-scale language model. The large-scale language model can be a bidirectional Transformer pre-trained language model, a Word2Vec model, or a GPT series model.
[0037] In one example, the text to be detected is divided to obtain a set of segments to be detected of the text to be detected. , through the model , generate text vector representation , the specific formula is as follows:
[0038] in, is a fragment of text T The embedded vector representation of .
[0039] In an embodiment of the present application, the reference fragment is the fragment with the highest semantic similarity with the fragment to be detected in the retrieval tree. The retrieval tree is composed of vector representations of text fragments of all texts in the corpus. That is to say, the retrieval tree of the corpus is a tree that stores vector representations of each text fragment of each text in the corpus. The retrieval tree can significantly accelerate the retrieval process of the reference fragment by organizing the text vectors of each text in the corpus into a tree structure. By determining the semantic similarity between the fragment to be detected and the text fragments stored in the retrieval tree, the reference fragment with the highest semantic similarity with the fragment to be detected can be found from the retrieval tree to further analyze the plagiarism of the article to be detected.
[0040] In an embodiment of the present application, the semantic similarity between the segment to be detected and the text segment in the search tree can be determined by calculating the cosine similarity between the vector representation of the segment to be detected and the vector representation of the text segment.
[0041] In an embodiment of the present application, after determining the vector representation of each segment to be detected, for each segment to be detected, the semantic similarity between the segment to be detected and the retrieved text segment is calculated through the vector representation of the segment to be detected and the vector representation of the text segment in the retrieval tree, and the semantic similarities are sorted in order from high to low similarity, and the text segment with the highest semantic similarity with the segment to be detected is selected as the reference segment.
[0042] S102 : For each segment to be detected, determine a target similarity of the segment to be detected according to a semantic similarity between the segment to be detected and a corresponding reference segment and a TF-IDF value of the segment to be detected.
[0043] In the embodiment of the present application, the target similarity is used to characterize the similarity between the segment to be detected and the corresponding reference segment. The target similarity between the segment to be detected and the corresponding reference segment is calculated from multiple dimensions.
[0044] In the embodiment of the present application, the TF-IDF value is used to characterize the importance of the words in the segment to be detected in the reference text. The target similarity is obtained by weighted combination of the semantic similarity and the TF-IDF value. The specific formula is as follows:
[0045] in, represents the target similarity, Represents the semantic similarity between the text to be detected and the reference text, Indicates the TF-IDF value of the text to be detected. is a weight parameter used to adjust the importance of semantic matching and keyword matching between the text to be detected and the reference text.
[0046] S103 : Selecting, from each to-be-detected segment, a to-be-detected segment whose target similarity is greater than a preset similarity threshold as a target detection segment.
[0047] In an embodiment of the present application, after calculating the target similarity between each text to be detected and the corresponding reference segment, the segment to be detected whose target similarity is greater than a preset similarity threshold is selected as the target detection segment, that is, the segment to be detected with a low probability of plagiarism is first screened out through the preset similarity threshold, and then the remaining segment to be detected is used as the target detection segment to further analyze the plagiarism situation of the text to be detected.
[0048] S104, for each target detection segment, input the target detection segment and the reference segment corresponding to the target detection segment into a type prediction model, obtain a type of plagiarism corresponding to the target detection segment output by the type prediction model, and generate a detection result.
[0049] In an embodiment of the present application, the type prediction model is used to determine the plagiarism type of the fragment to be detected based on the fragment to be detected and the reference fragment. The plagiarism type refers to the plagiarism method of the fragment to be detected from the corresponding reference fragment, such as directly copying the reference fragment as the plagiarism method of the fragment to be detected, or semantically replacing part of the text of the reference fragment to obtain the plagiarism method of the fragment to be detected, or reorganizing the structure of the reference text to obtain the plagiarism method of the fragment to be detected.
[0050] In an embodiment of the present application, the detection result includes the plagiarism type corresponding to each target detection segment. That is to say, since the target detection segments input into the type prediction model are all segments to be detected with a higher probability of plagiarism, the type prediction model will output a class of plagiarism types based on the segments to be detected and the corresponding reference segments, and the detection result includes the plagiarism type corresponding to each target detection segment.
[0051] In the embodiment of the present application, the type prediction model is trained using the sample segment and a type of plagiarized segment corresponding to the sample segment as training samples, and using the plagiarism type of the plagiarized segment as training labels.
[0052] In an embodiment of the present application, the type of plagiarism is determined, and then based on the type of plagiarism, the sample segment is converted into a plagiarism segment corresponding to the plagiarism type, the sample segment and the plagiarism segment are used as training samples, and the plagiarism type is used as a training label to train the type prediction model. The type prediction model is trained by generating plagiarism segments of different plagiarism types, so that the type prediction model can identify various types of plagiarism.
[0053] In the embodiment of the present application, the type prediction model can be trained by means of Few-shot learning, which means training an efficient and generalizable type prediction model with only a small amount of labeled data. Its goal is to enable the model to understand new tasks and make accurate predictions based on the few training samples and training labels provided.
[0054] The text detection method provided by the embodiment of the present application divides the to-be-detected text into at least one to-be-detected segment, determines the reference segment with the highest semantic similarity to each to-be-detected segment from the retrieval tree of the pre-constructed corpus according to the vector representation of each to-be-detected segment, determines the target similarity characterizing the degree of similarity between the to-be-detected segment and the corresponding reference segment according to the semantic similarity between the to-be-detected segment and the corresponding reference segment and the TF-IDF value of the to-be-detected segment, selects the to-be-detected segment with the target similarity greater than a preset similarity threshold as the target detection segment, and for each target detection segment, inputs the target detection segment into a pre-trained type prediction model to obtain the plagiarism type corresponding to the target detection segment output by the type prediction model, thereby obtaining a detection result containing the plagiarism type corresponding to each target detection segment. Since the type prediction model is trained with sample text and plagiarized text as training samples and the plagiarism type corresponding to the plagiarized text as training labels, accurate distinction of the plagiarism type of the to-be-detected text is achieved, and the problem that the plagiarist cannot be detected for plagiarism by deep processing of the original text is solved.
[0055] Based on the above embodiments, as an optional embodiment, after obtaining a type of plagiarism type corresponding to the target detection segment output by the type prediction model, a method for generating a plagiarism report is as follows: Figure 3 The specific contents are as follows: S201, for each target detection segment, determining a context matching degree between a context of the target detection segment and a context of the reference segment according to the target detection segment and a corresponding reference segment; S202, for each target detection segment, calculating the semantic similarity between each semantic unit in the target detection segment and each semantic unit in the corresponding reference segment, and constructing a comparison matrix corresponding to the target detection segment; S203 , for each target detection segment, generating a plagiarism report according to the plagiarism type corresponding to the target detection segment, the contextual matching degree between the target detection segment and the reference segment, and the comparison matrix corresponding to the target detection segment.
[0056] In S201 of the embodiment of the present application, the context matching degree is used to characterize the degree of context matching between the target detection segment and the reference segment. By comparing the degree of context matching between the target detection segment and the reference segment, the similarity in expression between the target detection segment and the reference segment can be determined, and the plagiarism of the text to be detected can be deeply analyzed and detected, so as to more accurately evaluate the innovation of the text to be detected, which helps to improve the accuracy of plagiarism detection and avoid misjudgment caused by simple surface matching.
[0057] In the embodiment of the present application, the semantic unit of a fragment refers to a basic unit that can express a complete meaning in the fragment. The semantic unit can be a single word, multiple dimensions, a phrase, or a sentence.
[0058] In S202 of the embodiment of the present application, the elements in the generated comparison matrix are the semantic similarities between the semantic units of the target detection segment and the semantic units of the reference text. Therefore, for each target detection segment, it is necessary to calculate the semantic similarities between each semantic unit and each semantic unit of the reference segment one by one, and then obtain the comparison matrix corresponding to the current target detection segment.
[0059] In one example, semantic extraction is performed on a target detection segment and a corresponding reference segment respectively to obtain at least one semantic unit of the target detection segment and at least one semantic unit of the reference segment, each semantic unit of the target detection segment and each semantic unit of the reference segment are converted into a vector representation, and for each semantic unit of the target detection segment, the semantic similarity between the semantic unit and any semantic unit of the reference segment is determined; based on the semantic similarity between each semantic unit of the target detection segment and any semantic unit of the reference segment, a comparison matrix corresponding to the target detection segment is constructed.
[0060] In S203 of the embodiment of the present application, for each target detection segment, the plagiarism type corresponding to the target detection segment, the context matching degree between the target detection segment and the reference segment, and the comparison matrix corresponding to the target detection segment are obtained, so as to obtain a plagiarism report characterizing the plagiarism situation of the text to be detected.
[0061] In the above scheme, by calculating the matching degree between each target detection segment and the reference segment, the comparison matrix between each target detection segment and the corresponding reference segment, and combining the plagiarism type corresponding to the target detection segment, the plagiarism of the text to be detected is considered from multiple dimensions, and a plagiarism report showing the overall plagiarism situation and the local plagiarism situation is generated, thereby improving the accuracy of plagiarism detection. By gradually comparing the semantic units of the target detection segment and the reference segment, the model can more accurately capture local similarities. This can identify subtle differences that may be overlooked in a large range.
[0062] Based on the above embodiments, as an optional embodiment, a method for determining context matching is provided, such as Figure 4 The specific contents include: S301, dividing the target detection segment into a plurality of first windows in a sliding window manner according to a predefined window size, and converting words in each first window into vector representation; S302, for each first window of the target detection segment, determining the semantic similarity between the first window and the corresponding reference segment according to the vector representation of the first window; S303: Determine a contextual matching degree between the target detection segment and the reference segment according to the semantic similarities corresponding to each first window of the target detection segment.
[0063] In the embodiment of the present application, the sliding window represents a text segment of a fixed length, and when sliding, the content in the window will change as the position changes.
[0064] In S301 of the embodiment of the present application, the target detection segment is divided into multiple first windows according to a predefined window size, the reference segment corresponding to the target detection segment is divided into multiple second windows, and the words in each first window are converted into vector representations, that is, for each target detection segment, there are multiple first windows, and the text content in each first window of the target detection segment is inconsistent.
[0065] In one example, the target detection segment is "the sliding window is a fixed-length text segment", and the preset length of the sliding window is 5, which means that five words are analyzed each time, and the multiple first windows obtained by splitting include: First window: ["sliding", "window", "is", "one", "fixed"] Second first window: ["window", "is", "one", "fixed", "length"] The third first window: ["is", "a", "fixed", "length", "of"] Fourth first window: ["Fixed", "Length", "Of", "Text", "Fragment"] In S302 of the embodiment of the present application, for each first window, the cosine similarity between the vector representation of the first window and the vector representation of the corresponding reference segment is calculated, thereby obtaining the semantic similarity between the first window and the reference segment.
[0066] In S303 of the embodiment of the present application, the average value of the semantic similarities between each first window and the reference segment can be calculated as the contextual matching degree between the target detection segment and the reference segment; the contextual matching degree between the target detection segment and the reference segment can also be obtained by weighted semantic similarity, such as assigning different weights according to the semantic importance of different first windows, and then taking the weighted average of each semantic similarity to obtain the contextual matching degree; the maximum semantic similarity can also be selected from the calculated semantic similarities as the contextual matching degree.
[0067] In the above scheme, the target detection segment is divided into windows of preset length by sliding windows, and the semantic similarity between each first window and the reference text is calculated, so that the contextual matching degree of the local context can be obtained, thereby analyzing the similarity between the target detection segment and the reference text in different parts, and improving the accuracy of the contextual matching degree. In addition, the contextual matching degree between the target detection segment and the reference segment is calculated by sliding windows, which can dynamically adjust the window size to find a suitable comparison range in contexts of different lengths, avoiding missing important contextual information in long texts. The accuracy of the contextual matching degree calculation is improved.
[0068] Based on the above embodiments, as an optional embodiment, the present application embodiment further provides a method for obtaining the TF-IDF value of a to-be-detected segment, such as Figure 5 As shown, the content is as follows: S401, performing word segmentation on the segment to be detected, and obtaining at least one keyword contained in the segment to be detected; S402, obtaining a reference text corresponding to the reference segment; S403, for each keyword, determining the frequency of the keyword in the text to be detected, and determining the number of segments containing the keyword in the reference text; S404, determining the TF-IDF value of the keyword in the reference text according to the frequency, the total number of segments of the reference text, and the number of segments containing the keyword in the reference text.
[0069] In S401 of the embodiment of the present application, keywords refer to words that can represent the main content of the text to be detected and are highly relevant to the text subject. The fragment to be detected is segmented, the text is divided into independent words, and multiple keywords that are highly relevant to the text subject are obtained.
[0070] In S402 of the embodiment of the present application, the full text to which the reference segment belongs, ie, the reference text, is determined according to the reference segment corresponding to the segment to be detected.
[0071] In S403 of the embodiment of the present application, for each keyword, the frequency of occurrence of the keyword in the text to be detected is determined, and the number of segments containing the keyword in the reference text is determined.
[0072] In S404 of the embodiment of the present application, the TF-IDF value of the keyword in the reference text is determined according to the frequency, the total number of segments of the reference text, and the number of segments containing the keyword in the reference text. The specific formula is as follows:
[0073] in, Indicates keywords Clip to be detected The frequency of occurrence in To contain keywords The number of fragments, is the total number of fragments of the reference text.
[0074] In the above scheme, by capturing the TF-IDF value of the keyword, the accuracy of the similarity calculation is improved while avoiding the influence of common words. In addition, it can be applied to different text fields and text lengths, with strong applicability and high accuracy.
[0075] Based on the above embodiments, as an optional embodiment, a method for constructing a search tree is provided, such as Figure 6 The specific contents are as follows: S501, for each text in the corpus, segment the text into multiple text segments, and for each text segment, determine a vector representation of the text segment; S502, constructing a leaf node of a retrieval tree based on the vector representation of each text segment of each text; S503, performing multiple rounds of recursion according to all leaf nodes until any recursion stopping condition is met.
[0076] In S501 of the embodiment of the present application, when constructing a retrieval tree of a corpus, each text in the corpus is first segmented into multiple text segments of predefined lengths, and for each text segment, a vector representation of the text segment is generated based on a large language model.
[0077] In S502 of the embodiment of the present application, in the process of constructing the retrieval tree, the vector representation of each text segment of each text constitutes a leaf node of the retrieval tree.
[0078] In S503 of the embodiment of the present application, performing multiple rounds of iterations based on all leaf nodes means that text segments with high semantic similarity are clustered multiple times through a clustering algorithm to obtain abstract semantic vector representations of multiple rounds of high-level nodes until any recursive stopping condition is met.
[0079] refer to Figure 7 As shown, the process of each round of recursion includes: S601, obtaining each element of this round of recursion, where the element of the first round of recursion is a text segment; S602, clustering all elements by using a clustering algorithm to obtain at least one cluster; S603, for each cluster, determining a vector representation of the cluster according to the vector representations of the elements included in the cluster; S604, taking the cluster as the element of the next round of recursion, and taking the elements obtained in this round of recursion as the nodes of a new layer in the search tree, the nodes corresponding to the elements obtained in this round of recursion are located in the upper layer of the corresponding nodes in the previous round in the search tree.
[0080] In S601 of the embodiment of the present application, in the first round of recursion, the elements participating in the recursion are text fragments, and in the non-first round of recursion, the elements participating in the recursion are clusters generated in the previous round of recursion.
[0081] In S602 of the embodiment of the present application, all elements are clustered by a clustering algorithm to obtain at least one cluster, and the similarity between the vector representations of any two elements in each cluster is not less than a preset cluster similarity threshold. In the first round of recursion, the vector representation of each text segment is determined, and the similarity between the vector representations of any two text segments is calculated. Then, the text segments are clustered by the clustering algorithm and the similarity between the vector representations of any two text segments to obtain at least one cluster, and each cluster contains at least one text segment; in the non-first round of recursion, the vector representation of each element is determined, and the similarity between the vector representations of any two elements is calculated. According to the clustering algorithm and the similarity between the vector representations of any two elements, the elements are clustered to obtain at least one cluster, and each cluster contains at least one element. The clustering algorithm can be a Gaussian mixture model-based clustering algorithm.
[0082] In S603 of the embodiment of the present application, for each cluster, the vector representation of the current cluster can be determined according to the vector representation of the elements contained in the cluster. The specific formula is as follows:
[0083] in, is the set of elements contained in the cluster, is the vector representation of the elements, is the vector representation of the cluster.
[0084] In the first round of recursion, for each cluster, the text segments contained in the cluster are determined, and the vector representation of the cluster is determined based on the vector representation of the contained text segments; in the non-first round of recursion, the elements contained in the cluster are determined, and the vector representation of the cluster is determined based on the vector representation of the contained elements.
[0085] In S604 of an embodiment of the present application, after determining the vector representation of the cluster, the cluster obtained in this round of recursion is used as an element participating in the next round of recursion, and the cluster obtained in this round of recursion is used as a node of a new layer in the retrieval tree. The node corresponding to the element obtained in this round of recursion is located in the upper layer of the corresponding node in the previous round in the retrieval tree.
[0086] In one example, in the first round of recursion, clusters A, B, and C are obtained. Cluster A contains text segments 1 and 2, cluster B contains text segments 3 and 4, and cluster C contains text segments 5 and 6. Then in the retrieval tree, the leaf nodes corresponding to text segments 1 and 2 are child nodes of the node corresponding to cluster A, the leaf nodes corresponding to text segments 3 and 4 are child nodes of the node corresponding to cluster B, and the leaf nodes corresponding to text segments 5 and 6 are child nodes of the node corresponding to cluster C.
[0087] In the above scheme, by constructing a retrieval tree of the corpus, the text fragments are aggregated based on the similarity between each text fragment, and the text data of the corpus is efficiently organized and indexed, so that there is no need to traverse the entire corpus when retrieving reference fragments, which greatly reduces the time consumed in the search.
[0088] In the embodiment of the present application, the recursive stop condition includes any of the following: The retrieval tree reaches the preset number of layers; The number of clusters is lower than the preset threshold; The similarity between any two clusters is lower than the preset cluster similarity threshold.
[0089] In an embodiment of the present application, the number of layers of the retrieval tree may be predefined, and when the retrieval tree constructed in the recursive process reaches the predefined number of layers, the recursion is stopped to obtain the retrieval tree.
[0090] In an embodiment of the present application, a preset number of clusters can also be pre-set. When the number of clusters obtained by clustering during the recursive process is lower than the preset number threshold, it means that the elements do not need to be clustered anymore, and clustering is stopped to obtain a retrieval tree based on the current recursive results.
[0091] In an embodiment of the present application, a clustering similarity threshold may be predefined. When the similarity threshold between any two clustering clusters is lower than the preset clustering similarity threshold, it indicates that the current similarity between the clustering clusters is low and is not suitable for further clustering. Therefore, the recursion is stopped to obtain a retrieval tree.
[0092] In the embodiment of the present application, when searching for reference segments of the to-be-detected segments in the search tree, a dynamic path selection strategy is adopted to maximize the node similarity through recursive search. The specific formula is as follows:
[0093] in, is the child node with the largest similarity in the output cluster. is the vector representation of the segment to be detected, is the vector representation of the cluster.
[0094] Based on the above embodiments, as an optional embodiment, the plagiarism type output by the type prediction model includes at least one of the following: copy; semantic substitution; Restructuring; The plagiarized segment with the plagiarism type of copying is obtained by copying the sample segment; The plagiarism segment with the plagiarism type of semantic replacement is obtained by rewriting the sample segment through the synonym replacement function; The plagiarized fragment with structural recombination type is obtained by recombining the sequence of the sample fragment.
[0095] In the embodiment of the present application, the copied plagiarized segment is obtained by copying the sample segment, and the formula is as follows:
[0096] in, For sample fragments, This is a plagiarized fragment whose plagiarism type is copying.
[0097] In the embodiment of the present application, the plagiarized segment with the plagiarism type of semantic replacement is obtained by rewriting the sample segment through the synonym replacement function, and the formula is as follows:
[0098] in, For plagiarized fragments whose plagiarism type is semantic replacement, is a synonym replacement function, For sample fragments, These are the words that are semantically replaced in the sample fragment.
[0099] In the embodiment of the present application, the plagiarized segment with the plagiarism type of structural reorganization is obtained by reorganizing the sequence of the sample segment, and the specific formula is as follows:
[0100] in, is a function for rearranging sentence sequences, For plagiarized fragments whose plagiarism type is structural reorganization, is a sample fragment, and the sentence sequence set of the sample fragment is .
[0101] In the embodiment of the present application, the training samples generated by the above method are combined with a small amount of manually annotated training labels to construct a training set. , as the training input for Few-shot learning. Model training is optimized through the loss function, the formula is as follows:
[0102] in, is the actual training label, Output probability for the model.
[0103] In the embodiment of the present application, the F1 score is used as an indicator to evaluate the model during the training process. The higher the F1 score, the better the performance of the model.
[0104] In the above scheme, the sample fragments are processed by using operations corresponding to the plagiarism types, so as to train a type prediction model that can identify plagiarism types such as copying, semantic replacement, and structural reorganization. The model is trained with a small number of manually annotated cases, reducing the time and labor costs required for model training.
[0105] On the basis of the above embodiments, as an optional embodiment, a method for determining partial plagiarism and overall plagiarism is provided, such as Figure 8 The specific contents are as follows: S701, for each target detection segment, obtaining a target element greater than an element similarity threshold from a comparison matrix corresponding to the target detection segment; S702, for each target element, taking a semantic unit in a target detection segment corresponding to the target element as plagiarized text; S703, determining the first character number of the plagiarized text and the second character number of the target detection segment, and determining the local plagiarism situation according to the first character number and the second character number; S704: Determine the overall plagiarism situation according to the number of first characters of the plagiarized text of each target detection segment and the number of third characters of the text to be detected.
[0106] In an embodiment of the present application, the plagiarism report is used to display the overall plagiarism situation and local plagiarism situation of the text to be detected. The local plagiarism situation refers to the proportion of plagiarized text in each target detection segment in the text to be detected, and the overall plagiarism situation refers to the proportion of plagiarized text in the full text of the text to be detected. The plagiarized text refers to the text obtained by plagiarizing the reference segment in the segment to be detected through the corresponding plagiarism type.
[0107] In S701 of the embodiment of the present application, a target element greater than an element similarity threshold is obtained from the comparison matrix of the target detection segment in order to screen out semantic units with a high degree of similarity with the reference segment from the target detection segment. The elements can be sorted in descending order, and then elements greater than the element similarity threshold are obtained as target elements.
[0108] In S702 of the embodiment of the present application, each target element has a corresponding semantic unit of the target detection segment. Since the semantic similarity represented by the target elements is relatively large, the semantic unit of the target detection segment corresponding to the target element is used as the plagiarized text in the target detection segment.
[0109] In S703 of the embodiment of the present application, the proportion of the plagiarized text in the target detection segment can be determined based on the number of words in the plagiarized text and the number of words in the target detection segment. The first number of characters is used to characterize the number of characters in the plagiarized text, and the second number of characters is used to characterize the number of characters in the target detection segment. Therefore, the first number of characters in the plagiarized text and the second number of characters in the target detection segment are determined, and the ratio of the first number of characters to the second number of characters is used as the proportion of the plagiarized text in the target detection segment to obtain the local plagiarism situation.
[0110] In S704 of the embodiment of the present application, the third number of characters is used to represent the number of characters in the text to be detected. After determining the number of words in all plagiarized texts and the number of all words in the text to be detected, the proportion of all plagiarized texts in the full text to be detected can be determined. Therefore, the proportion of plagiarized texts in the full text is determined according to the first number of characters of the plagiarized texts in each target detection segment and the third number of characters in the text to be detected to obtain the overall plagiarism situation.
[0111] In an embodiment of the present application, the plagiarism report may include the plagiarism type of each target detection segment in the text to be detected, the local plagiarism situation and the overall plagiarism situation of the text to be detected, and the contextual matching degree between each target detection segment in the text to be detected and the reference text.
[0112] In the embodiment of the present application, the plagiarism report can be presented in natural language generation technology, which is an artificial intelligence technology used by large models, which can convert data into natural language text that humans can understand. In the context of plagiarism report, natural language generation technology is used to convert the analysis results into easy-to-understand reports. In natural language generation technology, generating text is usually regarded as a sequence generation problem, that is, building sentences one word by one, and optimizing the output of plagiarism report through the following formula Process:
[0113] Among them, the sequence generation probability Indicates that in the generated word sequence Based on the next word Probability of occurrence.
[0114] In an embodiment of the present application, before processing the text to be detected, a natural language processing tool is used to remove noise and redundant data.
[0115] In an embodiment of the present application, when dividing the text to be detected, a long document is divided into logical paragraphs, and a rule-based algorithm is used to ensure the integrity of the paragraphs and avoid semantic loss.
[0116] In an embodiment of the present application, when determining the text to be detected, the text to be detected can be automatically subject-categorized so as to quickly match the subject-related corpus in the retrieval stage.
[0117] The present application embodiment provides a text detection device, such as Fig. 9 As shown, the text detection device 90 may include: a division module 901 , a determination module 902 , a selection module 903 and an input module 904 .
[0118] Specifically, the division module 901 is used to divide the text to be detected into at least one segment to be detected, determine the vector representation of each segment to be detected, and determine the reference segment corresponding to the segment to be detected from the retrieval tree based on the vector representation and a pre-built retrieval tree; the reference segment is the segment with the highest semantic similarity with the segment to be detected in the retrieval tree, and the retrieval tree is composed of the vector representations of the text segments of all the texts in the corpus; A determination module 902 is used to determine, for each segment to be detected, a target similarity of the segment to be detected according to the semantic similarity between the segment to be detected and the corresponding reference segment and the TF-IDF value of the segment to be detected; the TF-IDF value is used to characterize the importance of the words in the segment to be detected in the reference text, and the target similarity is used to characterize the similarity between the segment to be detected and the corresponding reference segment; A selection module 903 is used to select, from among the segments to be detected, segments to be detected whose target similarity is greater than a preset similarity threshold as target detection segments; An input module 904 is used to input the target detection segment and the reference segment corresponding to the target detection segment into a type prediction model for each target detection segment, obtain a type of plagiarism corresponding to the target detection segment output by the type prediction model, and generate a detection result, wherein the detection result includes the plagiarism type corresponding to each target detection segment; Among them, the type prediction model is trained with the sample segment and a type of plagiarized segment corresponding to the sample segment as training samples, and the plagiarism type of the plagiarized segment as training labels.
[0119] The device of the embodiments of the present application can execute the method provided by the embodiments of the present application, and the implementation principles are similar. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed functional description of each module of the device, please refer to the description in the corresponding method shown in the previous text, which will not be repeated here.
[0120] The text detection device provided by the embodiment of the present application divides the to-be-detected text into at least one to-be-detected segment, determines the reference segment with the highest semantic similarity to each to-be-detected segment from the retrieval tree of the pre-constructed corpus according to the vector representation of each to-be-detected segment, determines the target similarity representing the degree of similarity between the to-be-detected segment and the corresponding reference segment according to the semantic similarity between the to-be-detected segment and the corresponding reference segment and the TF-IDF value of the to-be-detected segment, selects the to-be-detected segment with the target similarity greater than a preset similarity threshold as the target detection segment, and for each target detection segment, inputs the target detection segment into a pre-trained type prediction model to obtain the plagiarism type corresponding to the target detection segment output by the type prediction model, thereby obtaining a detection result containing the plagiarism type corresponding to each target detection segment. Since the type prediction model is trained with sample text and plagiarized text as training samples and the plagiarism type corresponding to the plagiarized text as training labels, accurate distinction of the plagiarism type of the to-be-detected text is achieved, and the problem that the plagiarist cannot be detected for plagiarism by deeply processing the original text is solved.
[0121] In an embodiment of the present application, an electronic device (computer device / equipment / system) is provided, including a memory, a processor and a computer program stored in the memory, and the processor executes the above-mentioned computer program to implement the steps of the text detection method. Compared with the related art, it can be achieved: by dividing the text to be detected into at least one segment to be detected, according to the vector representation of each segment to be detected, determining the reference segment with the highest semantic similarity with each segment to be detected from the retrieval tree of the pre-constructed corpus, and according to the semantic similarity between the segment to be detected and the corresponding reference segment and the TF-IDF value of the segment to be detected, determining the similarity between the segment to be detected and the corresponding reference segment. The target similarity of the target is measured, and the target similarity to be detected is greater than the preset similarity threshold, and the target detection segment whose target similarity is greater than the preset similarity threshold is selected as the target detection segment. For each target detection segment, the target detection segment is input into the pre-trained type prediction model to obtain the plagiarism type corresponding to the target detection segment output by the type prediction model, thereby obtaining a detection result containing the plagiarism type corresponding to each target detection segment. Since the type prediction model is trained with sample text and plagiarized text as training samples and the plagiarism type corresponding to the plagiarized text as training label, it realizes the accurate distinction of the plagiarism type of the text to be detected, and solves the problem that the plagiarist cannot be detected for plagiarism through deep processing of the original text.
[0122] In an alternative embodiment, an electronic device is provided, such as Fig.10 As shown, Fig.10The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may also include a transceiver 4004, which may be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.
[0123] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0124] The bus 4002 may include a path to transmit information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.10 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0125] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation herein.
[0126] The memory 4003 is used to store the computer program for executing the embodiment of the present application, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the above method embodiment.
[0127] Among them, the electronic equipment package may include but is not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Fig.10 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0128] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented. Compared with the prior art, the present invention can achieve the following: by dividing the text to be detected into at least one segment to be detected, according to the vector representation of each segment to be detected, determining the reference segment with the highest semantic similarity with each segment to be detected from the retrieval tree of the pre-constructed corpus, and according to the semantic similarity between the segment to be detected and the corresponding reference segment and the TF-IDF value of the segment to be detected, determining the target similarity characterizing the degree of similarity between the segment to be detected and the corresponding reference segment, selecting the segment to be detected with the target similarity greater than a preset similarity threshold as the target detection segment, for each target detection segment, inputting the target detection segment into a pre-trained type prediction model, obtaining the plagiarism type corresponding to the target detection segment output by the type prediction model, thereby obtaining a detection result containing the plagiarism type corresponding to each target detection segment, and since the type prediction model is trained with sample text and plagiarized text as training samples and the plagiarism type corresponding to the plagiarized text as training labels, accurate distinction of the plagiarism type of the text to be detected is achieved, and the problem that the plagiarist cannot be detected for plagiarism by deep processing of the original text is solved.
[0129] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0130] The present application also provides a computer program product, including a computer program, which can implement the steps and corresponding contents of the above method embodiments when executed by a processor. Compared with the prior art, it can achieve: The method divides the text to be detected into at least one segment to be detected, determines the reference segment with the highest semantic similarity with each segment to be detected from the retrieval tree of the pre-constructed corpus according to the vector representation of each segment to be detected, determines the target similarity representing the degree of similarity between the segment to be detected and the corresponding reference segment according to the semantic similarity between the segment to be detected and the corresponding reference segment and the TF-IDF value of the segment to be detected, selects the segment to be detected with the target similarity greater than a preset similarity threshold as the target detection segment, and for each target detection segment, inputs the target detection segment into a pre-trained type prediction model to obtain the plagiarism type corresponding to the target detection segment output by the type prediction model, thereby obtaining a detection result containing the plagiarism type corresponding to each target detection segment. Since the type prediction model is trained with sample text and plagiarized text as training samples and the plagiarism type corresponding to the plagiarized text as training labels, the accurate distinction of the plagiarism type of the text to be detected is achieved, and the problem that the plagiarist cannot be detected for plagiarism due to deep processing of the original text is solved.
[0131] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that shown or described in the drawings.
[0132] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the implementation order of these steps is not limited to the order indicated by the arrows. Unless clearly stated herein, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders according to demand. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on actual implementation scenarios. Some or all of these sub-steps or stages may be executed at the same time, and each sub-step or stage in these sub-steps or stages may also be executed at different times respectively. In different scenarios of execution time, the execution order of these sub-steps or stages may be flexibly configured according to demand, and the embodiment of the present application does not limit this.
[0133] The above are only optional implementation methods for some implementation scenarios of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the scheme of the present application, other similar implementation methods based on the technical ideas of the present application are also within the protection scope of the embodiments of the present application.
Claims
1. A text detection method, characterized in that: include: Divide the text to be detected into at least one segment to be detected, determine a vector representation of each segment to be detected, and determine a reference segment corresponding to the segment to be detected from the retrieval tree based on the vector representation and a pre-constructed retrieval tree; the reference segment is a segment with the highest semantic similarity with the segment to be detected in the retrieval tree, and the retrieval tree is composed of vector representations of text segments of all texts in a corpus; For each segment to be detected, the target similarity of the segment to be detected is determined according to the semantic similarity between the segment to be detected and the corresponding reference segment and the TF-IDF value of the segment to be detected; the TF-IDF value is used to characterize the importance of the words in the segment to be detected in the reference text, and the target similarity is used to characterize the similarity between the segment to be detected and the corresponding reference segment; From each to-be-detected segment, select a to-be-detected segment whose target similarity is greater than a preset similarity threshold as a target detection segment; For each target detection segment, input the target detection segment and a reference segment corresponding to the target detection segment into a type prediction model, obtain a type of plagiarism corresponding to the target detection segment output by the type prediction model, and generate a detection result, wherein the detection result includes the type of plagiarism corresponding to each target detection segment; The type prediction model is trained by taking the sample segment and a type of plagiarized segment corresponding to the sample segment as training samples, and taking the plagiarism type of the plagiarized segment as training labels.
2. The method according to claim 1, characterized in that The step of obtaining a type of plagiarism corresponding to the target detection segment output by the type prediction model further includes: For each target detection segment, determining a context matching degree between a context of the target detection segment and a context of the reference segment according to the target detection segment and a corresponding reference segment; the context matching degree is used to characterize a matching degree of the context between the target detection segment and the reference segment; For each target detection segment, the semantic similarity between each semantic unit in the target detection segment and each semantic unit in the corresponding reference segment is calculated, and a comparison matrix corresponding to the target detection segment is constructed; the elements in the comparison matrix are the semantic similarities between the semantic units of the target detection segment and the semantic units of the reference text; For each target detection segment, a plagiarism report is generated according to the plagiarism type corresponding to the target detection segment, the context matching degree between the target detection segment and the reference segment, and the comparison matrix corresponding to the target detection segment.
3. The method according to claim 2, characterized in that The determining of the context matching degree between the context of the target detection segment and the context of the reference segment includes: According to a predefined window size, the target detection segment is divided into a plurality of first windows in a sliding window manner, and words in each first window are converted into vector representations; For each first window of the target detection segment, determining a semantic similarity between the first window and a corresponding reference segment according to a vector representation of the first window; A context matching degree between the target detection segment and the reference segment is determined according to the semantic similarities corresponding to the first windows of the target detection segment.
4. The method according to claim 2, characterized in that: The plagiarism report includes the overall plagiarism situation and partial plagiarism situation of the text to be detected, the partial plagiarism situation refers to the proportion of plagiarized text in each target detection segment, the overall plagiarism situation refers to the proportion of plagiarized text in the full text of the text to be detected, and the plagiarized text refers to the text obtained after the corresponding plagiarism type plagiarizes the reference segment in the target detection segment; The steps of obtaining the partial plagiarism situation and the overall plagiarism situation include: For each target detection segment, obtaining target elements greater than an element similarity threshold from a comparison matrix corresponding to the target detection segment; For each target element, a semantic unit in the target detection segment corresponding to the target element is used as plagiarized text; Determine a first character number of the plagiarized text and a second character number of the target detection segment, and determine the local plagiarism situation according to the first character number and the second character number; The overall plagiarism situation is determined according to the number of first characters of the plagiarized text of each target detection segment and the number of third characters of the text to be detected.
5. The method according to claim 1, characterized in that The step of obtaining the TF-IDF value of the segment to be detected includes: Performing word segmentation on the text to be detected to obtain at least one keyword contained in the text to be detected; Get the reference text corresponding to the reference fragment; For each keyword, determining the frequency of occurrence of the keyword in the segment to be detected, and determining the number of segments in the reference text that contain the keyword; The TF-IDF value of the keyword in the reference text is determined according to the frequency, the total number of segments of the reference text, and the number of segments of the reference text containing the keyword.
6. The method according to claim 1, characterized in that The method for constructing the retrieval tree comprises: For each text in the corpus, the text is segmented into a plurality of text segments, and for each text segment, a vector representation of the text segment is determined; A leaf node of the retrieval tree is formed based on the vector representation of each text segment of each text; Perform multiple rounds of recursion on all leaf nodes until any recursive stopping condition is met; Each round of recursion includes: Get each element of this round of recursion. The element of the first round of recursion is the text fragment; Clustering all elements through a clustering algorithm to obtain at least one cluster; For each cluster, determine the vector representation of the cluster according to the vector representation of the elements contained in the cluster; The cluster is used as the element of the next round of recursion, and the elements obtained in this round of recursion are used as nodes of a new layer in the retrieval tree. The nodes corresponding to the elements obtained in this round of recursion are located in the upper layer of the corresponding nodes in the previous round in the retrieval tree.
7. The method according to claim 6, characterized in that The recursive stop condition includes any of the following: The retrieval tree reaches a preset number of layers; The number of clusters is lower than a preset number threshold; The similarity between any two clusters is lower than the preset cluster similarity threshold.
8. The method according to claim 1, characterized in that The plagiarism type output by the type prediction model includes at least one of the following: copy; semantic substitution; Restructuring; The plagiarized segment whose plagiarism type is copying is obtained by copying the sample segment; The plagiarism segment whose plagiarism type is semantic replacement is obtained by rewriting the sample segment through a synonym replacement function; The plagiarism fragment whose plagiarism type is structural recombination is obtained by recombining the sequence of the sample fragment.
9. A text detection device, characterized in that: include: A division module, used for dividing the to-be-detected text into at least one to-be-detected segment, determining a vector representation of each to-be-detected segment, and determining a reference segment corresponding to the to-be-detected segment from the retrieval tree based on the vector representation and a pre-constructed retrieval tree; the reference segment is a segment in the retrieval tree with the highest semantic similarity to the to-be-detected segment, and the retrieval tree is composed of vector representations of text segments of all texts in a corpus; A determination module is used to determine, for each segment to be detected, a target similarity of the segment to be detected according to the semantic similarity between the segment to be detected and the corresponding reference segment and the TF-IDF value of the segment to be detected; the TF-IDF value is used to characterize the importance of the words in the segment to be detected in the reference text, and the target similarity is used to characterize the similarity between the segment to be detected and the corresponding reference segment; A selection module, configured to select, from among the segments to be detected, a segment to be detected whose target similarity is greater than a preset similarity threshold as a target detection segment; An input module is used to input, for each target detection segment, the target detection segment and a reference segment corresponding to the target detection segment into a type prediction model, obtain a type of plagiarism corresponding to the target detection segment output by the type prediction model, and generate a detection result, wherein the detection result includes the plagiarism type corresponding to each target detection segment; The type prediction model is trained by taking the sample segment and a type of plagiarized segment corresponding to the sample segment as training samples, and taking the plagiarism type of the plagiarized segment as training labels.
10. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Electronic homework plagiarism preventing system and method based on paragraph plagiarism detection
CN103678528A
Content plagiarism identification method, apparatus and device, and storage medium
CN112214984A
Similar text retrieval method and device, electronic equipment and storage medium
CN113407738A
Detection method, device, equipment, medium and product
CN119203991A
Systems, methods and computer program products for a snippet based proximal search
US20110252030A1
Cited By
Text sharing compliance detection method and device, electronic equipment and program product
CN120873619A