An intelligent duplicate checking and detection system and method based on multi-feature mimicry of target entities
The method leverages natural language processing and graph neural networks to extract and fuse text and semantic features, dynamically adjusting similarity scores based on text structure, improving the accuracy and robustness of text similarity detection.
Patent Information
- Application Number
- CN202510265333.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-07
AI Technical Summary
The existing plagiarism checking methods cannot fully combine the specific content of the target text, and lack analysis of the multi-dimensional and multi-level characteristics of the target entity, resulting in inaccurate similarity calculation.
An intelligent plagiarism detection method based on multi-eigen morphology of target entities is adopted. Through natural language processing word segmentation, text features and semantic features are extracted in combination with bag-of-word model and BERT model, and fused into plagiarism features. When the plagiarism similarity is lower than the threshold, structural features are extracted using graph neural network to adjust.
It improves the intelligence and accuracy of the plagiarism checking system, can more comprehensively understand and capture the potential similarity of the text, reduce misjudgment and misjudgment, and enhances the accuracy and robustness of plagiarism checking.
Smart Images

Figure CN119761339B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data detection, and more particularly, to an intelligent duplicate checking and detection system and method based on multi-feature mimicry of target entities. Background Art
[0002] With the rapid development of information technology, the generation and dissemination volume of text data increase year by year. Especially in the fields of academic research, document management, and intellectual property, the importance of duplicate checking technology has become increasingly prominent. Current duplicate checking is usually based on direct comparison of texts, using string matching algorithms or text similarity calculation methods based on word frequency statistics. These methods can effectively identify duplicate content in some scenarios, but with the development of semantic analysis technology, traditional duplicate checking methods have gradually revealed some deficiencies.
[0003] Currently, duplicate checking methods based on natural language processing (NLP) technology have gradually attracted attention, and the accuracy of duplicate checking is improved by introducing semantic understanding and context analysis. However, existing methods still have certain limitations. For example, the extraction of semantic features often depends on pre-trained models and cannot fully combine the specific content of the target text; existing duplicate checking methods mostly rely on a single feature (such as text or semantic features) for comparison, lacking a comprehensive analysis of multi-dimensional and multi-level features of the target entity; the mining of text structure features is lacking, and dynamic adjustment cannot be performed according to the differences in text structures, resulting in inaccurate calculation of duplicate checking similarity.
[0004] Therefore, it is necessary to design an intelligent duplicate checking and detection system and method based on multi-feature mimicry of target entities to solve the problems existing in the current technology. Summary of the Invention
[0005] In view of this, the present invention proposes an intelligent duplicate checking and detection system and method based on multi-feature mimicry of target entities, aiming to solve the problems that current duplicate checking methods cannot fully combine the specific content of the target text and lack multi-dimensional analysis of the target entity, resulting in inaccurate similarity calculation.
[0006] On the one hand, the present invention proposes an intelligent duplicate checking and detection method based on multi-feature mimicry of target entities, including:
[0007] Collect data of the target entity to be detected, and segment the data of the target entity to be detected based on natural language processing;
[0008] Process the segmented data based on the bag-of-words model to extract text features, and process the segmented data based on BERT (Bidirectional Encoder Representations from Transformers pre-trained language model) to extract semantic features;
[0009] Fuse the text features and semantic features to obtain mimetic features, compare the mimetic features with the historical database to determine duplicate files and determine the duplicate similarity;
[0010] Compare the duplicate similarity with the similarity threshold, and judge whether to adjust the duplicate similarity according to the comparison result; when the duplicate similarity is less than the similarity threshold, it is determined to adjust the duplicate similarity, and the data after word segmentation is processed based on a graph neural network to extract structural features, and the structural features are compared with the historical structural features in the historical database, and an adjustment coefficient is determined according to the comparison result to adjust the duplicate similarity.
[0011] Further, when performing word segmentation on the target entity data to be detected based on natural language processing, it includes:
[0012] Perform data cleaning on the target entity data to be detected. The data cleaning includes removing stop words and punctuation marks, and converting English text to lowercase letters;
[0013] Count the number of words in the cleaned data, including the number of Chinese characters and the number of English words, and determine the word segmentation tool according to the ratio of the number of Chinese characters to the number of English words;
[0014] When the ratio of the number of Chinese characters to the number of English words is greater than 1, the word segmentation tools include Jieba, THULAC, and HanLP;
[0015] When the ratio of the number of Chinese characters to the number of English words is less than or equal to 1, the word segmentation tools include NLTK, spaCy, and StanfordNLP.
[0016] Further, when processing the data after word segmentation based on the bag-of-words model to extract text features, it includes:
[0017] Pre-construct a vocabulary V = {w1, w2,..., wn}, where Wn represents the nth word in the vocabulary;
[0018] Calculate the number of occurrences of each word in the data after word segmentation in the vocabulary;
[0019] Represent the number of occurrences of all the data after word segmentation as a feature vector to obtain the text features.
[0020] Further, when fusing the text features and semantic features to obtain mimetic features, it includes:
[0021] Use zero-padding to adjust the text features and semantic features to the same dimension;
[0022] The text features and semantic features are fused using weighted fusion to obtain the mimicry features;
[0023] ;
[0024] Among them, V represents the mimicry feature, Vw represents the text feature, Vy represents the semantic feature, a1 and a2 represent weight coefficients, and a1 + a2 = 1.
[0025] Further, when comparing the mimicry feature with the historical database to determine the duplicate-check file and the duplicate-check similarity, it includes:
[0026] The historical database includes the historical mimicry features of several historical detection target entity data;
[0027] Calculate the duplicate-check similarity between the mimicry feature and each historical mimicry feature;
[0028] ;
[0029] Among them, Si represents the duplicate-check similarity between the mimicry feature and the i-th historical mimicry feature, V represents the mimicry feature, Vsi represents the i-th historical mimicry feature, represents the modulus of the vector.
[0030] Further, when judging whether to adjust the duplicate-check similarity according to the comparison result, it includes:
[0031] When the duplicate-check similarity is greater than or equal to the similarity threshold, it is determined not to adjust the duplicate-check similarity, and the duplicate-check file and the duplicate-check similarity are output;
[0032] When the duplicate-check similarity is less than the similarity threshold, it is determined to adjust the duplicate-check similarity.
[0033] Further, when it is determined to adjust the duplicate-check similarity and the structure features are extracted by processing the segmented data based on the graph neural network, it includes:
[0034] Convert the segmented data into a graph structure, where the nodes in the graph structure represent basic elements, and the basic elements include words or sentences, and the edges in the graph structure represent the relationships between elements;
[0035] Use Word2Vec to map each basic element to a vector of a fixed dimension;
[0036] Update the representation of the nodes in each layer based on the convolution operation;
[0037] ;
[0038] Among them, Denote the representation of node v at the (k + 1)-th layer; Denote the set of neighbor nodes of node v; Denote the weight matrix at the k-th layer; Denote the bias term, Denote the activation function, Denote the representation of node u at the k-th layer;
[0039] Obtain the structural feature by performing average pooling on the representations of all nodes.
[0040] Further, when determining the adjustment coefficient according to the comparison result to adjust the duplicate check similarity, it includes:
[0041] Select the historical mimic feature corresponding to the maximum duplicate check similarity between the mimic feature and the historical mimic feature, obtain the historical structural feature of the historical mimic feature, obtain the modulus difference between the structural feature and the historical structural feature, and determine the adjustment coefficient according to the modulus difference to adjust the duplicate check similarity;
[0042] Compare the modulus difference with a preset first preset modulus difference and a second preset modulus difference respectively, and determine the adjustment coefficient according to the comparison result to adjust the duplicate check similarity; the first preset modulus difference is less than the second preset modulus difference;
[0043] When the modulus difference is less than or equal to the first preset modulus difference, determine the first adjustment coefficient to adjust the duplicate check similarity; when the modulus difference is greater than the first preset modulus difference and less than or equal to the second preset modulus difference, determine the second adjustment coefficient to adjust the duplicate check similarity; when the modulus difference is greater than the second preset modulus difference, determine the third adjustment coefficient to adjust the duplicate check similarity; the first adjustment coefficient is greater than the second adjustment coefficient, the second adjustment coefficient is greater than the third adjustment coefficient, and the value range of the adjustment coefficient is (0, 1), and the adjusted duplicate check similarity is the product of the duplicate check similarity and the adjustment coefficient.
[0044] Compared with the prior art, the beneficial effects of the present invention are as follows: By introducing the method of multi-feature fusion and combining technologies such as natural language processing, deep learning, and graph neural networks, the intelligence and accuracy of the plagiarism detection system are improved. Based on the word segmentation and bag-of-words model of natural language processing, text features are extracted, and then combined with the BERT model to extract semantic features. The text of the target entity is analyzed from multiple levels and dimensions, avoiding the limitations of single-feature comparison in traditional methods, and being able to more comprehensively understand and capture the potential similarity of the text. The text features and semantic features are fused to generate mimetic features, and then through comparison with the historical database, the plagiarism detection similarity is accurately calculated, thereby realizing more efficient plagiarism detection. When the plagiarism detection similarity is lower than the set threshold, the graph neural network is further introduced to mine and compare the text structure features, and the similarity is adjusted to cope with the structural differences of the text, avoiding the plagiarism detection error caused by structural differences. The accuracy of plagiarism detection is improved, the similarity of the target text is comprehensively evaluated, and the precision and robustness of plagiarism detection are improved.
[0045] On the other hand, the present application also provides an intelligent plagiarism detection system based on multi-feature mimesis of a target entity, which is used to apply the above-mentioned intelligent plagiarism detection method based on multi-feature mimesis of a target entity, including:
[0046] An acquisition unit, configured to acquire data of a target entity to be detected, and perform word segmentation on the data of the target entity to be detected based on natural language processing;
[0047] A processing unit, configured to process the segmented data based on the bag-of-words model to extract text features, and process the segmented data based on BERT to extract semantic features;
[0048] A judgment unit, configured to fuse the text features and semantic features to obtain mimetic features, compare the mimetic features with the historical database, determine the plagiarized file, and determine the plagiarism detection similarity;
[0049] An adjustment unit, configured to compare the plagiarism detection similarity with a similarity threshold, and judge whether to adjust the plagiarism detection similarity according to the comparison result; when the plagiarism detection similarity is less than the similarity threshold, it is determined that the plagiarism detection similarity is adjusted, and the segmented data is processed based on the graph neural network to extract structural features, and the structural features are compared with the historical structural features in the historical database, and an adjustment coefficient is determined according to the comparison result to adjust the plagiarism detection similarity.
[0050] Furthermore, when the adjustment unit determines an adjustment coefficient according to the comparison result to adjust the plagiarism detection similarity, it further includes:
[0051] The adjustment unit selects the historical mimic feature corresponding to the maximum duplicate check similarity between the mimic feature and the historical mimic feature, obtains the historical structural feature of the historical mimic feature, obtains the modulus difference between the structural feature and the historical structural feature, and determines an adjustment coefficient according to the modulus difference to adjust the duplicate check similarity.
[0052] It can be understood that the above intelligent duplicate check detection system and method based on multi-feature mimicry of the target entity have the same beneficial effects and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0054] Figure 1 is a flowchart of an intelligent duplicate check detection method based on multi-feature mimicry of the target entity provided by an embodiment of the present invention;
[0055] Figure 2 is a functional block diagram of an intelligent duplicate check detection system based on multi-feature mimicry of the target entity provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.
[0057] In some embodiments of the present application, referring to Figure 1 shown, an intelligent duplicate check detection method based on multi-feature mimicry of the target entity includes:
[0058] S100: Collect data of the target entity to be detected, and perform word segmentation on the data of the target entity to be detected based on natural language processing.
[0059] S200: Process the segmented data based on the bag-of-words model to extract text features, and process the segmented data based on BERT to extract semantic features.
[0060] S300: Integrate the text features and semantic features to obtain the mimetic features, compare the mimetic features with the historical database to determine the duplicate files and the duplicate similarity.
[0061] S400: Compare the duplicate similarity with the similarity threshold, and judge whether to adjust the duplicate similarity according to the comparison result. When the duplicate similarity is less than the similarity threshold, it is determined to adjust the duplicate similarity, and the data after word segmentation is processed based on the graph neural network to extract the structural features, and the structural features are compared with the historical structural features in the historical database, and the adjustment coefficient is determined according to the comparison result to adjust the duplicate similarity.
[0062] Specifically, in S100, the target text data to be detected is collected. Through natural language processing technology, the text data is segmented to divide basic language units (such as words, phrases, etc.), providing a basis for subsequent feature extraction and analysis. The text information is converted into structured data that can be understood by machines. In S200, on the data after word segmentation, the bag-of-words model is used to extract text features. The bag-of-words model helps capture the basic information in the text, such as the repetition and distribution patterns of vocabulary, by counting the occurrence frequencies of each word in the text. The BERT model is used to extract the semantic features of the text. BERT is a deep learning language model that can better understand the context and thus capture the deep semantic information in the text. This embodiment can not only identify the literal similarity of the text, but also understand the meaning and context of the text, thereby improving the accuracy of duplicate checking. In S300, the text features extracted by the bag-of-words model and the semantic features extracted by BERT are integrated to generate a "mimetic feature". The mimetic feature combines the surface information and deep semantic information of the text, representing the multi-dimensional features of the target entity. The mimetic feature is compared with the content in the historical database to identify the existing similar texts. By calculating the duplicate similarity, the similarity between the text and the historical data is initially determined. In S400, on the basis of the preliminary duplicate check, the duplicate similarity is compared with the set similarity threshold. If the duplicate similarity is lower than the threshold, further adjustment is made. Based on the graph neural network (GNN Graph Neural Network), the structural features of the text are extracted to analyze the structure and organization form of the text. The potential similarity in the structure of the text (such as paragraph, sentence structure, etc.) is captured through the graph neural network and compared with the historical structural features in the historical database.
[0063] According to the comparison result, the duplicate similarity is adjusted according to the similarity of the text structure, so as to more accurately evaluate the degree of text repetition. The adjusted similarity can more effectively identify the true similarity of the text, avoiding misjudgment and missed judgment.
[0064] It is understandable that by integrating text features, semantic features, and structural features, and adopting advanced technologies such as natural language processing, BERT models, and graph neural networks, the accuracy and robustness of duplicate checking detection have been improved. It can simultaneously consider the surface features, semantic levels, and differences in text structures of the text, avoiding the limitations of traditional methods that rely on single features, and can more comprehensively identify the similarities of complex texts; based on the semantic feature extraction of deep learning and the structural feature analysis of graph neural networks, the system can dynamically adjust the duplicate checking similarity, reducing misjudgments and missed judgments caused by factors such as text rewriting and synonym replacement; the multi-dimensional feature fusion and dynamic adjustment mechanism enable duplicate checking to more intelligently handle different types of texts, with higher accuracy and flexibility compared to traditional duplicate checking methods.
[0065] In some embodiments of the present application, when tokenizing the target entity data to be detected based on natural language processing, it includes: cleaning the target entity data to be detected, and the data cleaning includes removing stop words and punctuation marks, and converting English text to lowercase letters.
[0066] Count the number of words in the cleaned data, including the number of Chinese characters and the number of English words, and determine the tokenization tool according to the ratio of the number of Chinese characters to the number of English words.
[0067] When the ratio of the number of Chinese characters to the number of English words is greater than 1, the tokenization tools include Jieba, THULAC, and HanLP.
[0068] When the ratio of the number of Chinese characters to the number of English words is less than or equal to 1, the tokenization tools include NLTK, spaCy, and StanfordNLP.
[0069] It is understandable that stop words refer to words that frequently appear in the text but contribute little to the meaning of the text (such as "de", "he", "shi", etc.). Removing these stop words helps reduce noise and improve the effectiveness of subsequent analysis. Although punctuation marks are important for the transmission of word meanings, they are usually not treated as independent units when tokenizing, so they need to be removed. By counting the number of words in the target entity data and analyzing the language ratio, the most suitable tokenization tool is selected, thus ensuring the efficiency and accuracy of tokenization in different language environments. By flexibly adjusting the tokenization tool used according to the language characteristics of the text, the limitations in the application of traditional single tokenization tools are effectively avoided, and complex texts containing mixed languages (such as Chinese-English mixtures) can be handled. The adaptability and accuracy of text preprocessing are improved, and for texts containing a large amount of English or Chinese-English mixtures, they can be processed more accurately, thereby improving the overall effect of subsequent text feature extraction, semantic analysis, and duplicate checking calculation. The applicability and accuracy of the duplicate checking detection system in multi-language and multi-text type environments are enhanced.
[0070] In some embodiments of the present application, when processing the segmented data based on the bag-of-words model to extract text features, it includes: pre-constructing a vocabulary V = {w1, w2,..., wn}, where Wn represents the nth word in the vocabulary. Calculating the number of occurrences of each word in the segmented data in the vocabulary. Representing all the occurrences of the segmented data as a feature vector to obtain text features.
[0071] In some embodiments of the present application, when fusing the text features and semantic features to obtain mimetic features, it includes:
[0072] Using the zero-padding method to adjust the text features and semantic features to the same dimension.
[0073] ;
[0074] where, V represents the mimetic feature, Vw represents the text feature, Vy represents the semantic feature, a1, a2 represent weight coefficients, and a1 + a2 = 1.
[0075] It can be understood that through feature extraction and fusion, the ability of the duplicate check detection system is improved. By extracting features from the text through the bag-of-words model, the lexical distribution in the text can be fully captured, providing basic information for subsequent duplicate checks. Using the zero-padding method ensures the dimensional consistency of the text features and semantic features, enabling the two to be smoothly fused. Through the weighted fusion method, the relative importance of the text features and semantic features can be flexibly adjusted according to specific requirements, so as to obtain the optimal duplicate check effect in different application scenarios. It can optimize the duplicate check accuracy at multiple feature levels (text, semantics), improve the adaptability and accuracy, and maintain high processing ability when facing different types of texts (for example, documents with rich semantic information or texts with complex language structures).
[0076] In some embodiments of the present application, when comparing the mimetic feature with the historical database to determine the duplicate check file and the duplicate check similarity, it includes:
[0077] The historical database includes the historical mimetic features of several historical detection target entity data.
[0078] Calculating the duplicate check similarity between the mimetic feature and each historical mimetic feature.
[0079] ;
[0080] where, Si represents the duplicate check similarity between the mimetic feature and the ith historical mimetic feature, V represents the mimetic feature, Vsi represents the ith historical mimetic feature, represents the modulus of the vector.
[0081] In some embodiments of the present application, when determining whether to adjust the duplicate check similarity based on the comparison result, it includes: when the duplicate check similarity is greater than or equal to the similarity threshold, it is determined that the duplicate check similarity is not adjusted, and the duplicate check file and the duplicate check similarity are output. When the duplicate check similarity is less than the similarity threshold, it is determined that the duplicate check similarity is adjusted.
[0082] It can be understood that, assuming an article to be detected, and the mimicry features of several articles that have been detected are stored in the historical database. First, calculate the mimicry features of the article to be detected, and calculate the similarity with the mimicry features of each article in the historical database. The similarity between the mimicry features of the article to be detected and the mimicry features of the 5th article in the historical database is 0.85, while the similarity with the 10th article is 0.65. The set similarity threshold is 0.8. It is determined that the similarity between the 5th article and the article to be detected is higher than the threshold, and the duplicate check result is directly output and marked as high similarity; while the similarity of the 10th article is lower than the threshold, the similarity will be adjusted according to further analysis, and more text structure features will be considered to optimize or adjust the similarity.
[0083] By calculating the similarity between the mimicry features and the historical mimicry features in the historical database, texts with high similarity can be identified, avoiding the limitations of simple literal matching or duplicate check methods based on surface features. The introduction of the similarity threshold judgment and the dynamic adjustment mechanism improves the flexibility and accuracy of the duplicate check system. When facing texts with low similarity or potential subtle similarities, the duplicate check results are optimized through further adjustment to avoid missed judgments or misjudgments. The multi-stage duplicate check and similarity adjustment process can handle complex texts and provide more accurate and comprehensive duplicate check detections.
[0084] In some embodiments of the present application, when it is determined that the duplicate check similarity is adjusted and the structure features are extracted by processing the tokenized data based on a graph neural network, it includes:
[0085] Convert the tokenized data into a graph structure, where the nodes in the graph structure represent basic elements, and the basic elements include words or sentences, and the edges in the graph structure represent the relationships between elements.
[0086] Use Word2Vec to map each basic element to a vector of a fixed dimension.
[0087] Update the representation of the nodes in each layer based on convolutional operations.
[0088] ;
[0089] where, represents the representation of node v in the (k + 1)-th layer; represents the set of neighbor nodes of node v; denotes the weight matrix of the k-th layer; denotes the bias term, denotes the activation function, denotes the representation of node u at the k-th layer.
[0090] Average pooling is performed on the representations of all nodes to obtain the structural features.
[0091] In some embodiments of the present application, when determining the adjustment coefficient to adjust the duplicate checking similarity according to the comparison result, it includes: selecting the historical mimetic feature corresponding to the maximum duplicate checking similarity between the mimetic feature and the historical mimetic feature, obtaining the historical structural feature of the historical mimetic feature, obtaining the modulus difference between the structural feature and the historical structural feature, and determining the adjustment coefficient according to the modulus difference to adjust the duplicate checking similarity. The modulus difference is compared with a preset first preset modulus difference and a second preset modulus difference respectively, and the adjustment coefficient is determined according to the comparison result to adjust the duplicate checking similarity. The first preset modulus difference is less than the second preset modulus difference.
[0092] Specifically, when the modulus difference is less than or equal to the first preset modulus difference, determine the first adjustment coefficient to adjust the duplicate checking similarity. When the modulus difference is greater than the first preset modulus difference and less than or equal to the second preset modulus difference, determine the second adjustment coefficient to adjust the duplicate checking similarity. When the modulus difference is greater than the second preset modulus difference, determine the third adjustment coefficient to adjust the duplicate checking similarity. The first adjustment coefficient is greater than the second adjustment coefficient, the second adjustment coefficient is greater than the third adjustment coefficient, and the value range of the adjustment coefficient is (0, 1). The adjusted duplicate checking similarity is the product of the duplicate checking similarity and the adjustment coefficient.
[0093] Specifically, the tokenized data is converted into a graph structure. The nodes in the graph represent the basic elements in the text, which can be words, phrases, sentences, etc. The edges in the graph represent the relationships between these elements, usually defined based on grammar, semantics, or context information. For example, an edge can represent the dependency relationship between words or the semantic connection between sentences. The Word2Vec model is used to map each basic element (such as a word or a sentence) into a vector space of a fixed dimension. These vectors can preserve the semantic similarity between elements. Therefore, the representation of nodes can reflect the semantic content of the elements in the text. The convolutional operation in the graph neural network (Graph Convolutional Network, GCN) is used to update the node representation at each layer. At each layer, the representation of a node is affected by its neighboring nodes. After updating the node representation through multiple layers of convolutional operations, the representations of all nodes are aggregated into a global structural feature through average pooling, which reflects the overall structural information of the text. These structural features provide additional context information during the duplicate checking process, helping to distinguish texts that are superficially similar but structurally different.
[0094] It can be understood that by converting the text into a graph structure and using graph neural networks to capture the grammar, semantics, and structural information in the text, the deep similarities and differences between texts can be more accurately identified during duplicate checking. Through the mechanism of modulus difference and adjustment coefficient, the duplicate checking similarity is dynamically adjusted according to the structural differences of the text, making the final duplicate checking result more accurate. The similarity adjustment not only improves the accuracy of duplicate checking but also avoids over-reliance on surface features, enhancing the robustness when dealing with texts with complex structures and diverse grammars.
[0095] In the above embodiments, by introducing the method of multi-feature fusion and combining technologies such as natural language processing, deep learning, and graph neural networks, the intelligence and accuracy of the duplicate checking system are improved. The text features are extracted based on word segmentation and the bag-of-words model in natural language processing, and then combined with the BERT model to extract semantic features. The text of the target entity is analyzed from multiple levels and dimensions, avoiding the limitations of single-feature comparison in traditional methods and being able to comprehensively understand and capture the potential similarities of the text. The text features and semantic features are fused to generate mimetic features, and then through comparison with the historical database, the duplicate checking similarity is accurately calculated, thus achieving more efficient duplicate checking. When the duplicate checking similarity is lower than the set threshold, the graph neural network is further introduced to mine and compare the text structural features, and the similarity is adjusted to cope with the structural differences of the text, avoiding the duplicate checking errors caused by structural differences. The accuracy of duplicate checking is improved, the similarity of the target text is comprehensively evaluated, and the precision and robustness of duplicate checking are improved.
[0096] In another preferred manner based on the above embodiments, refer toFigure 2 As shown, this embodiment provides an intelligent duplicate check detection system based on multi-feature mimicry of target entities, which is used to apply the above-mentioned intelligent duplicate check detection method based on multi-feature mimicry of target entities, including:
[0097] An acquisition unit, configured to acquire data of the target entity to be detected, and perform word segmentation on the data of the target entity to be detected based on natural language processing.
[0098] A processing unit, configured to process the segmented data based on the bag-of-words model to extract text features, and process the segmented data based on BERT to extract semantic features.
[0099] A judgment unit, configured to fuse the text features and semantic features to obtain mimicry features, compare the mimicry features with the historical database, determine duplicate check files and determine the duplicate check similarity.
[0100] An adjustment unit, configured to compare the duplicate check similarity with a similarity threshold, and judge whether to adjust the duplicate check similarity according to the comparison result. When the duplicate check similarity is less than the similarity threshold, it is determined to adjust the duplicate check similarity, and the segmented data is processed based on a graph neural network to extract structural features, and the structural features are compared with the historical structural features in the historical database, and an adjustment coefficient is determined according to the comparison result to adjust the duplicate check similarity.
[0101] Furthermore, when the adjustment unit determines an adjustment coefficient to adjust the duplicate check similarity according to the comparison result, it further includes:
[0102] The adjustment unit selects the historical mimicry feature corresponding to the maximum value of the duplicate check similarity between the mimicry feature and the historical mimicry feature, obtains the historical structural feature of the historical mimicry feature, obtains the modulus difference between the structural feature and the historical structural feature, and determines an adjustment coefficient according to the modulus difference to adjust the duplicate check similarity.
[0103] It is understandable that by introducing the method of multi-feature fusion and combining technologies such as natural language processing, deep learning, and graph neural networks, the intelligence and accuracy of the duplicate checking system have been improved. Based on the word segmentation and bag-of-words model of natural language processing to extract text features, and then combined with the BERT model to extract semantic features, the text of the target entity is analyzed from multiple levels and dimensions, avoiding the limitations of single-feature comparison in traditional methods, and being able to more comprehensively understand and capture the potential similarity of the text. The text features and semantic features are fused to generate mimic features, and then through comparison with the historical database, the duplicate checking similarity is accurately calculated, so as to achieve more efficient duplicate checking. When the duplicate checking similarity is lower than the set threshold, the graph neural network is further introduced to mine and compare the text structure features, and the similarity is adjusted to cope with the structural differences of the text, avoiding the duplicate checking errors caused by structural differences. The accuracy of duplicate checking is improved, the similarity of the target text is comprehensively evaluated, and the accuracy and robustness of duplicate checking are improved.
[0104] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0105] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 a block or multiple blocks.
[0106] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 a block or multiple blocks.
[0107] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the steps specified in one process or a plurality of processes and / or blocks Figure 1 in one block or a plurality of blocks Figure 1 of the functions specified in the flow.
[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. An intelligent duplicate check and detection method based on multi - feature mimicry of target entities, characterized in that, Including: Collect the data of the target entity to be detected, and perform word segmentation on the data of the target entity to be detected based on natural language processing; Process the segmented data based on the bag-of-words model to extract text features, and process the segmented data based on BERT to extract semantic features; Fuse the text features and semantic features to obtain mimic features, compare the mimic features with the historical database to determine duplicate files and determine the duplicate similarity; Compare the duplicate similarity with the similarity threshold, and judge whether to adjust the duplicate similarity according to the comparison result; when the duplicate similarity is less than the similarity threshold, it is determined that the duplicate similarity is adjusted, and the segmented data is processed based on the graph neural network to extract structural features, compare the structural features with the historical structural features in the historical database, and determine the adjustment coefficient according to the comparison result to adjust the duplicate similarity; When determining the adjustment coefficient according to the comparison result to adjust the duplicate similarity, it includes: Select the historical mimic feature corresponding to the maximum duplicate similarity between the mimic feature and the historical mimic feature, obtain the historical structural feature of the historical mimic feature, obtain the modulus difference between the structural feature and the historical structural feature, and determine the adjustment coefficient according to the modulus difference to adjust the duplicate similarity; Compare the modulus difference with the preset first preset modulus difference and the second preset modulus difference respectively, and determine the adjustment coefficient according to the comparison result to adjust the duplicate similarity; the first preset modulus difference is less than the second preset modulus difference; When the modulus difference is less than or equal to the first preset modulus difference, determine the first adjustment coefficient to adjust the duplicate similarity; when the modulus difference is greater than the first preset modulus difference and less than or equal to the second preset modulus difference, determine the second adjustment coefficient to adjust the duplicate similarity; when the modulus difference is greater than the second preset modulus difference, determine the third adjustment coefficient to adjust the duplicate similarity; the first adjustment coefficient is greater than the second adjustment coefficient, the second adjustment coefficient is greater than the third adjustment coefficient, and the value range of the adjustment coefficient is (0, 1), and the adjusted duplicate similarity is the product of the duplicate similarity and the adjustment coefficient.
2. The intelligent duplicate checking and detection method based on multi - feature mimicry of a target entity according to claim 1, wherein, When performing word segmentation on the data of the target entity to be detected based on natural language processing, it includes: Perform data cleaning on the data of the target entity to be detected, and the data cleaning includes removing stop words and punctuation marks and converting English text to lowercase letters; Count the number of words in the cleaned data, including the number of Chinese characters and the number of English words, and determine the word segmentation tool according to the ratio of the number of Chinese characters to the number of English words; When the ratio of the number of Chinese characters to the number of English words is greater than 1, the word segmentation tools include Jieba, THULAC, and HanLP; When the ratio of the number of Chinese characters to the number of English words is less than or equal to 1, the word segmentation tools include NLTK, spaCy, and StanfordNLP.
3. The intelligent duplicate checking and detecting method based on multi-feature mimicry of a target entity according to claim 2, wherein When processing the segmented data based on the bag-of-words model to extract text features, it includes: Pre - construct a vocabulary V = {w1, w2,..., wn}, where Wn represents the nth word in the vocabulary; Calculate the number of occurrences of each word in the segmented data in the vocabulary; Represent all the occurrences of the segmented data as a feature vector to obtain the text features.
4. The intelligent duplicate check and detection method based on multi-feature mimicry of a target entity according to claim 3, wherein, When fusing the text features and semantic features to obtain mimic features, it includes: Use zero padding to adjust the text features and semantic features to the same dimension; Fuse the text features and semantic features using weighted fusion to obtain the mimic features; ; Among them, V represents the mimic features, Vw represents the text features, Vy represents the semantic features, a1 and a2 represent weight coefficients, and a1 + a2 = 1.
5. The intelligent duplicate checking and detection method based on multi-feature mimicry of a target entity according to claim 4, wherein When comparing the mimic features with the historical database to determine duplicate files and determine the duplicate similarity, it includes: The historical database includes the historical mimic features of several historical detection target entity data; Calculate the duplicate similarity between the mimic features and each historical mimic feature; ; Among them, Si represents the similarity of duplicate checking between the mimicry feature and the i-th historical mimicry feature, V represents the mimicry feature, and Vsi represents the i-th historical mimicry feature. Represents the norm of the vector.
6. The intelligent duplicate check and detection method based on multi-feature mimicry of a target entity according to claim 5, wherein, When judging whether to adjust the duplicate similarity according to the comparison result, it includes: When the duplicate similarity is greater than or equal to the similarity threshold, it is determined not to adjust the duplicate similarity, and output the duplicate file and the duplicate similarity; When the duplicate similarity is less than the similarity threshold, it is determined to adjust the duplicate similarity.
7. The intelligent duplicate check and detection method based on multi-feature mimicry of a target entity according to claim 6, wherein When it is determined to adjust the duplicate similarity and process the segmented data based on a graph neural network to extract structural features, it includes: Convert the segmented data into a graph structure, where the nodes in the graph structure represent basic elements, and the basic elements include words or sentences, and the edges in the graph structure represent the relationships between elements; Use Word2Vec to map each basic element to a vector of a fixed dimension; Update the representation of nodes in each layer based on convolutional operations; ; Among them, represents the representation of node v at the (k + 1)-th layer; represents the set of neighbor nodes of node v; represents the weight matrix at the k-th layer; represents the bias term, σ represents the activation function, represents the representation of node u at the k-th layer; Perform average pooling on the representations of all nodes to obtain the structural features.
8. An intelligent duplicate detection system based on multi-feature mimicry of target entities, which is used to apply the intelligent duplicate detection method based on multi-feature mimicry of target entities according to any one of claims 1-7, and is characterized in that, It includes: A collection unit configured to collect data of the target entity to be detected, and segment the data of the target entity to be detected based on natural language processing; A processing unit configured to process the segmented data based on the bag - of - words model to extract text features, and process the segmented data based on BERT to extract semantic features; A judgment unit configured to fuse the text features and semantic features to obtain mimic features, compare the mimic features with the historical database to determine duplicate files and determine the duplicate similarity; An adjustment unit configured to compare the duplicate similarity with the similarity threshold, judge whether to adjust the duplicate similarity according to the comparison result; when the duplicate similarity is less than the similarity threshold, it is determined to adjust the duplicate similarity, and process the segmented data based on a graph neural network to extract structural features, compare the structural features with the historical structural features in the historical database, and determine an adjustment coefficient according to the comparison result to adjust the duplicate similarity.
9. The intelligent duplicate check and detection system based on multi-feature mimicry of a target entity according to claim 8, wherein When the adjustment unit determines an adjustment coefficient according to the comparison result to adjust the duplicate similarity, it further includes: The adjustment unit selects the historical mimicry feature corresponding to the maximum value of the duplicate check similarity between the mimicry feature and the historical mimicry feature, obtains the historical structural feature of the historical mimicry feature, obtains the modulus difference between the structural feature and the historical structural feature, and determines an adjustment coefficient according to the modulus difference to adjust the duplicate check similarity.
Citation Information
Patent Citations
Similarity detection method and system based on power information system code file
CN110471835A
Electronic case duplicate checking method and device based on word segmentation text and computer equipment
CN111814447A