Text duplicate checking method and device, electronic equipment and storage medium

By obtaining the target keyword clusters of the text and using the pre-trained domain similarity model to calculate domain similarity, the problem of low accuracy in existing text duplication detection technology is solved, and more accurate text duplication detection is achieved.

CN120670570APending Publication Date: 2025-09-19CHINA MOBILE GROUP ZHEJIANG +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510621221.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing text duplication detection technology has the problem of low accuracy, especially when the semantic similarity is high but the domain similarity is low, it is easy to make misjudgments.

Method used

By obtaining multiple target keyword clusters of the text to be checked for duplicates and the reference text, the domain similarity is calculated using a pre-trained domain similarity model to determine the duplicate checking results.

Benefits of technology

It improves the accuracy of text duplication checking, avoids misjudgment when the semantic similarity is high but the domain similarity is low, and ensures the accuracy of the duplication checking results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670570A_ABST
    Figure CN120670570A_ABST
Patent Text Reader

Abstract

The invention discloses a text duplicate checking method and device, electronic equipment and a storage medium, belongs to the technical field of information processing, and is used for solving the problem that a related text duplicate checking technology is low in duplicate checking accuracy. The method comprises the following steps: acquiring a first text and a second text; wherein the first text is a text to be subjected to duplicate checking; the second text is used for performing duplicate checking on the first text; obtaining a plurality of target keyword clusters of the historical text; wherein a target keyword in the target keyword cluster represents a domain feature of the historical text; through a pre-trained field similarity model, according to the multiple target keyword clusters, the first text and the second text, performing field similarity calculation to obtain a field similarity result of the first text and the second text; and determining a duplicate checking result based on the domain similarity result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of information processing technology, and specifically relates to a text duplication checking method, device, electronic device and storage medium. Background Art

[0002] In recent years, with the development of information technology and the continuous emergence of new technologies, the number and scale of various scientific and technological research projects have increased significantly, whether in enterprises, universities, or scientific research institutions. Generally, when carrying out project approval work, institutions at all levels will check the project description text for duplicates to prevent duplicate project approvals.

[0003] Current technology for text duplication detection uses keyword and semantic similarity to determine similarity, but this method is prone to misjudgment. For example, if two projects, Project A and Project B, have 90% technical similarity, and Project A belongs to the large-scale model technology route and Project B belongs to the traditional artificial intelligence (AI) route, and duplication is only determined based on semantic similarity, Project A and Project B will be judged as duplicates, but in fact, Project A and Project B are completely different.

[0004] In other words, the relevant text duplication detection technology has the problem of low accuracy. Summary of the Invention

[0005] The embodiments of the present application provide a text duplication checking method, device, electronic device and storage medium, which can solve the problem of low accuracy of duplication checking in related text duplication checking technologies.

[0006] In a first aspect, an embodiment of the present application provides a method for checking for duplicate text, the method comprising: obtaining a first text and a second text; wherein the first text is a text to be checked for duplicate text; the second text is a text used to check for duplicate text on the first text; obtaining multiple target keyword clusters of a historical text; wherein the target keywords in the target keyword cluster represent the domain characteristics of the historical text; performing domain similarity calculation based on the multiple target keyword clusters, the first text, and the second text through a pre-trained domain similarity model to obtain a domain similarity result between the first text and the second text; and determining a duplicate checking result based on the domain similarity result.

[0007] In the second aspect, an embodiment of the present application provides a text duplication checking device, which includes: a first acquisition module, used to obtain a first text and a second text; wherein the first text is the text to be checked for duplication; the second text is the text used to check for duplication of the first text; a second acquisition module, used to obtain multiple target keyword clusters of historical texts; wherein the target keywords in the target keyword clusters represent the domain characteristics of the historical text; a calculation module, used to perform domain similarity calculation based on the multiple target keyword clusters, the first text and the second text through a pre-trained domain similarity model, to obtain the domain similarity result of the first text and the second text; a determination module, used to determine the duplication checking result based on the domain similarity result.

[0008] In a third aspect, an embodiment of the present application provides an electronic device comprising: a processor; and a memory arranged to store computer-executable instructions, wherein the executable instructions are configured to be executed by the processor, and the executable instructions include instructions for executing the text duplication checking method as described in the first aspect.

[0009] In a fourth aspect, an embodiment of the present application provides a storage medium for storing computer-executable instructions, wherein the computer-executable instructions enable a computer to execute the text duplication checking method as described in the first aspect.

[0010] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the text duplication checking method as described in the first aspect.

[0011] In a sixth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the text duplication checking method as described in the first aspect.

[0012] In an embodiment of the present application, a first text and a second text are obtained; wherein the first text is a text to be checked for duplicates; the second text is a text used to check for duplicates of the first text; a plurality of target keyword clusters of a historical text are obtained; wherein the target keywords in the target keyword clusters represent the domain characteristics of the historical text; a domain similarity calculation is performed based on the plurality of target keyword clusters, the first text, and the second text using a pre-trained domain similarity model to obtain a domain similarity result between the first text and the second text; and a duplicate checking result is determined based on the domain similarity result. This solution, through the pre-trained domain similarity model, can automatically perform domain similarity calculation based on the plurality of target keyword clusters, the first text, and the second text to obtain a domain similarity result between the first text and the second text, thereby determining a duplicate checking result based on the domain similarity result, thereby avoiding a false positive when the semantic similarity between the first text and the second text is high but the domain similarity is low. Compared with the related text duplicate checking technology, which only performs repetitive judgment based on the semantic similarity when judging, the duplicate checking result determined by this solution is more accurate, solving the problem of low accuracy of the related text duplicate checking technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 This is a flowchart of a text duplication checking method provided in an embodiment of the present application; Figure 2 This is a schematic diagram of a domain similarity calculation using a domain similarity model provided in an embodiment of the present application; Figure 3 This is a schematic diagram of a process for constructing a keyword library provided by an embodiment of the present application; Figure 4 This is a schematic diagram of generating duplicate checking results provided by an embodiment of the present application; Figure 5 This is a flowchart of another text duplication checking method provided in an embodiment of the present application; Figure 6 This is a structural diagram of a text duplication checking device provided in an embodiment of the present application; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0014] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0015] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0016] The text duplication checking method, device, electronic device and storage medium provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.

[0017] Figure 1 A method for checking for duplicate text provided by an embodiment of the present invention is shown. The method can be performed by an electronic device, which may include: a server and / or a terminal device, wherein the terminal device may be, for example, a vehicle-mounted terminal or a mobile phone terminal. In other words, the method can be performed by software or hardware installed in the electronic device, and the method includes the following steps: S102: Acquire a first text and a second text.

[0018] The first text is the text to be checked for plagiarism; the second text is the text used to check for plagiarism on the first text.

[0019] In practical applications, this application can be applied to at least the following two application scenarios: the first application scenario is for enterprises, colleges, or scientific research institutions, etc., which need to check for project duplication when carrying out project establishment work, and can check for duplication of the project description text, wherein the project can be a project of various scenarios, and there is no specific limitation on this; the second application scenario is for educational institutions, publishers, scientific research institutions, and any organization that needs to ensure the originality of the content, which needs to verify the originality of the text, and can check for duplication of the text of the project itself, such as checking for duplication of the paper. It should be noted that this application is not limited to the above two application scenarios, and this application can determine whether there is duplication between two or more texts (or projects).

[0020] Accordingly, in the first application scenario, the first text can be the descriptive text of the first project, and the second text can be the descriptive text of the second project. The descriptive text of the first project can at least represent the project content of the first project, and may also represent the project background, requirements, and implementation objectives (project goals) of the first project, without specific limitations. Similarly, the descriptive text of the second project can at least represent the project content of the second project, and may also represent the project background, requirements, and implementation objectives (project goals) of the second project, without specific limitations. In the second application scenario, the first text can be the text of the first project itself, such as a paper. Similarly, the second text can also be the text of the second project itself, such as a paper. Both the first and second projects can be one or more. The first project can be the project to be checked for plagiarism. The second project can also be the project to be checked for plagiarism and compared with the first project. The second project can also be the project to be checked for plagiarism and compared with the first project. Similarly, the first and second texts can be one or more. The second text can be the text to be checked for plagiarism and compared with the first text. The second text can also be the text to be checked for plagiarism and compared with the first text. Therefore, there are no specific limitations on the second text.

[0021] S104: Acquire multiple target keyword clusters of the historical text.

[0022] Among them, the target keywords in the target keyword cluster represent the domain characteristics of the historical text.

[0023] The second text may be one or more of the historical texts, or may not be a historical text. The historical text may be a description of a historical project or the text of the historical project itself, such as a description of an approved project, a published paper, etc. Of course, the historical text may also be a text that has been checked for plagiarism and has a plagiarism check result.

[0024] Specifically, the historical text can be clustered to determine multiple target keyword clusters of the historical text, or the multiple target keyword clusters of the historical text can be directly obtained from the keyword library below; of course, it is also possible to flexibly select a suitable professional vocabulary according to the needs of different scenarios and obtain multiple target keyword clusters of the historical text from the professional vocabulary, and there is no specific limitation on this.

[0025] For example, multiple target keyword clusters ,in, is the number of target keywords in each cluster.

[0026] S106: Using a pre-trained domain similarity model, domain similarity calculation is performed based on the multiple target keyword clusters, the first text, and the second text to obtain a domain similarity result between the first text and the second text.

[0027] S108: Determine the duplicate checking result based on the domain similarity result.

[0028] Specifically, the duplicate checking results may include domain similarity results; they may also include a specific analysis of the reasons for the similarity comparison between the first text and the second text, thereby increasing the transparency of the system, helping users better understand why the first text and the second text are considered similar or different, thereby enhancing the user's sense of trust and facilitating subsequent manual review work.

[0029] The present invention provides a method for checking for duplicate text, comprising obtaining a first text and a second text; wherein the first text is a text to be checked for duplicate text; and the second text is a text used to check for duplicate text on the first text; obtaining multiple target keyword clusters of a historical text; wherein the target keywords in the target keyword clusters represent domain characteristics of the historical text; performing domain similarity calculation based on the multiple target keyword clusters, the first text, and the second text using a pre-trained domain similarity model to obtain domain similarity results between the first text and the second text; and determining a duplicate text checking result based on the domain similarity results. The present invention uses a pre-trained domain similarity model to automatically perform domain similarity calculation based on the multiple target keyword clusters, the first text, and the second text to obtain domain similarity results between the first text and the second text, thereby determining a duplicate text checking result based on the domain similarity results. Furthermore, when the semantic similarity between the first text and the second text is high but the domain similarity is low, a false duplicate text check is not made. Compared to related text duplicate text checking technologies that only perform duplicate judgment based on the semantic similarity level, the present invention provides a more accurate duplicate text checking result, thereby resolving the problem of low accuracy in related text duplicate text checking technologies.

[0030] In one implementation, the domain similarity model includes a rerank model. Using the pre-trained domain similarity model, domain similarity calculation is performed based on multiple target keyword clusters, the first text, and the second text to obtain a domain similarity result between the first text and the second text (i.e., S106). This can be specifically performed as follows: Steps A1 to A2: Step A1: Generate a first domain feature vector and a second domain feature vector based on multiple target keyword clusters, a first text, and a second text through a Rerank model.

[0031] The first domain feature vector is a first similarity score between the first text and the target keyword in descending order; the second domain feature vector is a second similarity score between the second text and the target keyword in descending order.

[0032] Rerank model, for example, bge-reranker-v1 model, etc. It should be understood that the rerank model is not limited to the bge-reranker-v1 model.

[0033] Continuing with the above example, the target keyword clusters are: ; Take the first text as x1 and use the following formula to calculate the corresponding first domain feature vector A:

[0034] in, represents the number of eigenvectors of the kth cluster; Represents the first similarity score.

[0035] Take the second text as x2 and use the following formula to calculate the corresponding second domain feature vector B:

[0036] in, represents the number of eigenvectors of the kth cluster; Represents the second similarity score.

[0037] It should be noted that the dimensions of the first domain feature vector and the second domain feature vector are the same as those of the multiple target keyword clusters.

[0038] Step A2: performing similarity calculation based on the first domain feature vector and the second domain feature vector to obtain a domain similarity result.

[0039] In this embodiment, the Rerank model is used to automatically generate domain feature vectors of different texts (first text and second text) based on multiple target keyword clusters, which can improve the accuracy of duplicate checking.

[0040] In one implementation, domain similarity calculation is performed using a domain similarity model, such as Figure 2 Before performing similarity calculation based on the first domain feature vector and the second domain feature vector to obtain the domain similarity result (i.e., step A2), step B1 may also be performed: Step B1: Perform zero-to-one processing on the first domain feature vector and the second domain feature vector to obtain the first domain embedding vector and the second domain embedding vector.

[0041] Continuing with the above example, each first similarity score in the first domain feature vector is and a second similarity score in the second domain feature vector Perform 0-1 processing, and the first similarity score or the second similarity score greater than 0 is treated as 1, and less than 0 is treated as 0. Get the first domain embedding vector and the second domain embedding vector Thus, the situation where the similarity calculation result (domain similarity result) is high due to too many different features between the two texts (the first text and the second text) is avoided.

[0042] The above-mentioned domain similarity model also includes a pre-trained attention model. Accordingly, based on the first domain feature vector and the second domain feature vector, similarity calculation is performed to obtain the domain similarity result (i.e., step A2), which can be specifically performed as follows: step B2: In step B2, the pre-trained attention model is used to perform similarity calculation based on the first domain embedding vector, the second domain embedding vector, and the attention weight matrix to obtain the domain similarity result.

[0043] The attention weight matrix includes the weight for each target keyword cluster.

[0044] The attention model can adopt a deep neural network (DNN) structure.

[0045] Specifically, there are many similarity calculation methods, such as Euclidean distance, cosine similarity, and Pearson correlation coefficient, which are not specifically limited.

[0046] Continuing with the above example, we use the pre-trained attention model to embed the vector according to the first domain and the second domain embedding vector and the attention weight matrix W a ( , dimension is k), the cosine similarity can be calculated using the following formula to obtain the domain similarity result sim2:

[0047] Among them, W a,i Corresponding to each target keyword cluster W i The weight of .

[0048] It should be noted that in the process of training the domain similarity model, the present application updates the attention weight matrix, and evaluates the content repetitiveness in different fields by using the weights in the attention weight matrix obtained through training. The weights in the attention weight matrix not only reflect the importance of the data in each field, but are also directly related to each target keyword cluster. Therefore, the weight of each target keyword cluster can intuitively show the degree of its influence on the entire duplicate checking process. In particular, when business needs change and the duplicate checking requirements for a specific field need to be adjusted, even if there is a lack of new labeled data to retrain the domain similarity model, the present application can also manually adjust the weights (in the attention weight matrix) of the corresponding field, thereby quickly responding to business changes, so that the system can be optimized for specific situations without sacrificing performance, showing good generalization capabilities, and thus being able to meet changing business needs, while ensuring the accuracy of the duplicate checking results while improving adaptability.

[0049] In this embodiment, by introducing an attention mechanism, the system focuses more on parts of the text that are important to specific fields or topics. This allows for customized (personalized) emphasis on specific aspects based on actual business needs (for example, in academia, the focus may be on detecting technological innovations; in legal document review, the focus may be on verifying the consistency of clauses). This makes the duplication detection process more targeted and effectively improves the effectiveness of duplication detection in key areas.

[0050] In one implementation, the following steps C1 to C3 may be performed to obtain multiple target keyword clusters: Step C1, obtaining historical text.

[0051] Step C2: extract keywords from historical texts using a large language model to obtain multiple initial keywords.

[0052] The Large Language Model (LLM) can be, for example, the Tongyi Qianwen Qwen-72b model. Of course, the LLM is not limited to the Qwen-72b model, and can have the same or similar function of extracting keywords from historical texts.

[0053] For example, if the historical text is a description of a historical project, 10 keywords are extracted from the project objectives and content sections of each historical project description text to obtain initial keywords. It should be noted that the above example is only for ease of understanding and does not constitute a specific limitation on the number of initial keywords extracted.

[0054] In step C3, a clustering process and a deduplication process are performed on the multiple initial keywords using a clustering model to obtain multiple target keyword clusters, and the multiple target keyword clusters are stored in a keyword library.

[0055] The specific process of constructing the keyword library can be as follows: Figure 3 shown.

[0056] Acquiring multiple target keyword clusters from the historical text (i.e., S104) can be specifically performed as follows: Step C4: Step C4: Acquire multiple target keyword clusters of the historical text from the keyword library.

[0057] In this embodiment, a plurality of initial keywords of the historical text are first determined through a large language model. Then, considering that the number of keywords will affect the dimension of the vector (first domain feature vector and second domain feature vector) of the content of the text (first text and second text), a dimension that is too high and has low discrimination will reduce the accuracy of the domain similarity calculation. Then, the number of keywords is reduced through clustering and deduplication processing to obtain multiple target keyword clusters.

[0058] In one implementation, a clustering model is used to perform clustering and deduplication processing on multiple initial keywords to obtain multiple target keyword clusters (i.e., step C3). The steps c1 to c4 can be specifically performed as follows: In step c1, the initial keywords are converted into word embeddings through the pre-trained language model.

[0059] The pre-trained language model can be, but is not limited to, a Bidirectional Encoder Representations from Transformers (BERT) model. Any language model that can convert each initial keyword into a word embedding is sufficient. The BERT model is a pre-trained language representation model.

[0060] In step c2, similarity calculation is performed on the word embedding to obtain the similarity results between the initial keywords, and a similarity matrix is ​​constructed based on the similarity results between the initial keywords.

[0061] Among them, the rows and columns of the similarity matrix are the number of initial keywords.

[0062] For example, the cosine similarity calculation can be performed on the word embedding to obtain the similarity result between each pair of initial keywords, thereby obtaining Similarity matrix of , where n is the number of initial keywords.

[0063] Step c3: Using a density-based clustering algorithm, perform a first clustering process on the similarity matrix to obtain multiple keyword clusters.

[0064] Among these, density-based clustering algorithms, such as the Density-Based Spatial Clustering of Application with Noise (DBSCAN) algorithm, define a cluster as the largest set of density-connected points. This algorithm can partition regions with sufficiently high density into clusters and discover clusters of arbitrary shapes in a noisy spatial database. Of course, this density-based clustering algorithm is not limited to the aforementioned DBSCAN algorithm; other density-based clustering algorithms are also acceptable. Numerous density-based clustering algorithms exist in the prior art, and we will not elaborate on them all here.

[0065] Step c4: using a hierarchical clustering algorithm and a preset keyword quantity threshold, performing a second clustering process and deduplication process on the multiple keyword clusters to obtain multiple target keyword clusters.

[0066] Hierarchical clustering is used to identify closely related keyword groups. A preset keyword quantity threshold is then used to select only those keywords that are significantly different. Specifically, the keyword at the cluster center and the keywords closest to the center within the preset keyword quantity threshold are selected. For example, the preset keyword quantity threshold can be set to 3. It should be understood that the keyword quantity threshold in this example does not constitute a specific limitation on the preset keyword quantity threshold.

[0067] In addition, the most representative keywords can be manually selected from each cluster, or if keywords in key areas are missing, keywords can be manually added to obtain the final target keyword list (multiple target keyword clusters), which has a significantly reduced number of keywords but can still effectively represent the content of the original dataset.

[0068] In this embodiment, the most representative keywords can be effectively selected as target keywords through two clustering processes using a density-based clustering algorithm and a hierarchical clustering algorithm and a deduplication process.

[0069] In one implementation, Figure 4 As shown in the figure, the duplicate checking result is generated by the semantic similarity result and the domain similarity result. You can also perform the following step D1 to obtain the semantic similarity result: In step D1, the first text and the second text are semantically embedded using a pre-trained semantic model to obtain a first semantic embedding vector and a second semantic embedding vector, and similarity calculation is performed on the first semantic embedding vector and the second semantic embedding vector to obtain a semantic similarity result.

[0070] The semantic model may be an embedding model, such as the bce-embedding-vase-v1 model. Of course, the semantic model is not limited to the bce-embedding-vase-v1 model, and any model having the same or similar semantic embedding function may be sufficient.

[0071] For example, the following formula can be used to calculate the similarity between the first semantic embedding vector A and the second semantic embedding vector B to obtain a semantic similarity result sim1:

[0072] Based on the domain similarity results, determining the duplicate checking results (i.e., S108) can be specifically performed as follows: Step D2: Generate duplicate checking results based on the domain similarity results and semantic similarity results.

[0073] Continuing with the previous example, the duplicate check result includes the final similarity score of the first text and the second text , In the stage of training the semantic model and the above-mentioned domain similarity model, the dataset used can be: 1000 manually annotated similar text pairs, 1000 randomly selected dissimilar text pairs as negative samples, and the ratio of training set, test set, and validation set can be 8:1:1. By calculating the final similarity score The cross entropy loss with the annotation results is used to update the attention matrix to obtain the trained semantic model and domain similarity model.

[0074] In this embodiment, the duplicate checking result is determined by combining the domain similarity result with the semantic similarity result, thereby improving the accuracy of the duplicate checking.

[0075] Figure 5 This is a flowchart of a text duplication checking method provided by the embodiment of the present application. Figure 5 As shown, the method includes: Step 502: Obtain a first text and a second text.

[0076] The first text is the text to be checked for plagiarism; the second text is the text used to check for plagiarism on the first text.

[0077] Step 504: Obtain historical text.

[0078] Step 506: extract keywords from the historical text using the large language model to obtain multiple initial keywords.

[0079] Step 508 : clustering and de-duplication processing are performed on the multiple initial keywords through a clustering model to obtain multiple target keyword clusters, and the multiple target keyword clusters are stored in a keyword library.

[0080] Step 510: Acquire multiple target keyword clusters of the historical text from the keyword library.

[0081] Step 512 : Generate a first domain feature vector and a second domain feature vector according to the plurality of target keyword clusters, the first text and the second text through a Rerank model of a pre-trained domain similarity model.

[0082] The first domain feature vector is a first similarity score between the first text and the target keyword in descending order; the second domain feature vector is a second similarity score between the second text and the target keyword in descending order.

[0083] Step 514 : performing zero-to-one processing on the first domain feature vector and the second domain feature vector to obtain the first domain embedding vector and the second domain embedding vector.

[0084] In step 516 , the attention model of the domain similarity model is used to perform similarity calculation based on the first domain embedding vector, the second domain embedding vector, and the attention weight matrix to obtain a domain similarity result.

[0085] The attention weight matrix includes the weight for each target keyword cluster.

[0086] In step 518, the first text and the second text are semantically embedded using a pre-trained semantic model to obtain a first semantic embedding vector and a second semantic embedding vector, and similarity calculation is performed on the first semantic embedding vector and the second semantic embedding vector to obtain a semantic similarity result.

[0087] Step 520: Generate duplicate checking results based on the domain similarity results and the semantic similarity results.

[0088] The specific process from step 502 to step 520 has been described in detail in the above embodiment and will not be repeated here.

[0089] In this embodiment, a first text and a second text are obtained; wherein the first text is a text to be checked for duplicates; the second text is a text used to check for duplicates on the first text; multiple target keyword clusters of a historical text are obtained; wherein the target keywords in the target keyword clusters represent domain characteristics of the historical text; domain similarity is calculated based on the multiple target keyword clusters, the first text, and the second text using a pre-trained domain similarity model to obtain domain similarity results for the first text and the second text; and a duplicate checking result is determined based on the domain similarity results. This solution, through the pre-trained domain similarity model, can automatically calculate domain similarity based on the multiple target keyword clusters, the first text, and the second text to obtain domain similarity results for the first text and the second text, thereby determining a duplicate checking result based on the domain similarity results. Furthermore, when the semantic similarity between the first text and the second text is high but the domain similarity is low, a false duplicate checking result is not made. Compared with related text duplicate checking technologies that only judge duplication based on the semantic similarity level, this solution determines a more accurate duplicate checking result, solving the problem of low accuracy of related text duplicate checking technologies.

[0090] Corresponding to the text duplicate checking method provided in the above embodiment, based on the same technical concept, the embodiment of the present invention also provides a text duplicate checking device. Figure 6 Schematic diagram of the structure of a text duplicate checking device according to an embodiment of the present invention, the text duplicate checking device is used to perform Figures 1 to 5 The text duplication detection method described, such as Figure 6 As shown, the text duplicate checking device includes: a first acquisition module 610 , a second acquisition module 620 , a calculation module 630 and a determination module 640 .

[0091] The first acquisition module 610 is used to acquire a first text and a second text; wherein the first text is the text to be checked for duplicate content; and the second text is the text used to check for duplicate content in the first text; The second acquisition module 620 is used to acquire multiple target keyword clusters of the historical text; wherein the target keywords in the target keyword clusters represent the domain characteristics of the historical text; A calculation module 630 is configured to perform domain similarity calculation based on a plurality of target keyword clusters, the first text, and the second text using a pre-trained domain similarity model to obtain a domain similarity result between the first text and the second text; The determination module 640 is used to determine the duplicate checking result based on the domain similarity result.

[0092] In one implementation, the domain similarity model includes a Rerank model. The calculation module 630 includes: A generating unit is configured to generate a first domain feature vector and a second domain feature vector based on the plurality of target keyword clusters, the first text, and the second text using a Rerank model; wherein the first domain feature vector is a first similarity score between the first text and the target keyword in descending order; and the second domain feature vector is a second similarity score between the second text and the target keyword in descending order; The calculation unit is used to perform similarity calculation based on the first domain feature vector and the second domain feature vector to obtain a domain similarity result.

[0093] In one implementation, the text duplication checking device further includes a zero-to-one processing module. The zero-to-one processing module is specifically configured to: The first domain feature vector and the second domain feature vector are zeroed to obtain the first domain embedding vector and the second domain embedding vector.

[0094] The aforementioned domain similarity model also includes a pre-trained attention model. The computing unit is specifically used to: Through the pre-trained attention model, similarity calculation is performed based on the first domain embedding vector, the second domain embedding vector and the attention weight matrix to obtain the domain similarity result.

[0095] The attention weight matrix includes the weight for each target keyword cluster.

[0096] In one implementation, the text duplication checking device further includes a keyword library module. The keyword library module includes: An acquisition unit, used to acquire historical text; An extraction unit, configured to extract keywords from historical texts using a large language model to obtain a plurality of initial keywords; A clustering unit is used to perform clustering and deduplication processing on the multiple initial keywords through a clustering model to obtain multiple target keyword clusters, and store the multiple target keyword clusters in a keyword library; The second acquisition module 620 is specifically configured to: Obtain multiple target keyword clusters of historical texts from a keyword library.

[0097] In one implementation, the clustering unit is specifically configured to: Convert initial keywords into word embeddings through a pre-trained language model; Calculate the similarity of the word embeddings to obtain the similarity results between the initial keywords, and construct a similarity matrix based on the similarity results between the initial keywords; where the rows and columns of the similarity matrix are the number of initial keywords; Using a density-based clustering algorithm, the similarity matrix is ​​subjected to a first clustering process to obtain multiple keyword clusters; Using a hierarchical clustering algorithm and a preset keyword quantity threshold, a second clustering process and a deduplication process are performed on multiple keyword clusters to obtain multiple target keyword clusters.

[0098] In one implementation, the above-mentioned text duplication checking device further includes a semantic computing module. The semantic computing module is specifically used to: The first text and the second text are semantically embedded using a pre-trained semantic model to obtain a first semantic embedding vector and a second semantic embedding vector, and similarity calculation is performed on the first semantic embedding vector and the second semantic embedding vector to obtain a semantic similarity result.

[0099] The determination module 640 is specifically configured to: Generate duplicate checking results based on domain similarity results and semantic similarity results.

[0100] In this embodiment, a first text and a second text are obtained; wherein the first text is a text to be checked for duplicates; the second text is a text used to check for duplicates on the first text; multiple target keyword clusters of a historical text are obtained; wherein the target keywords in the target keyword clusters represent domain characteristics of the historical text; domain similarity is calculated based on the multiple target keyword clusters, the first text, and the second text using a pre-trained domain similarity model to obtain domain similarity results for the first text and the second text; and a duplicate checking result is determined based on the domain similarity results. This solution, through the pre-trained domain similarity model, can automatically calculate domain similarity based on the multiple target keyword clusters, the first text, and the second text to obtain domain similarity results for the first text and the second text, thereby determining a duplicate checking result based on the domain similarity results. Furthermore, when the semantic similarity between the first text and the second text is high but the domain similarity is low, a false duplicate checking result is not made. Compared with related text duplicate checking technologies that only judge duplication based on the semantic similarity level, this solution determines a more accurate duplicate checking result, solving the problem of low accuracy of related text duplicate checking technologies.

[0101] Those skilled in the art should understand that the above-mentioned text duplication checking device can be used to implement the text duplication checking method mentioned above, and the detailed description thereof should be similar to the description of the method mentioned above. To avoid repetition, it will not be repeated here.

[0102] Based on the same technical concept, an embodiment of the present application further provides an electronic device for executing the above-mentioned text duplication checking method. Figure 7The following is a schematic diagram of the structure of an electronic device for implementing various embodiments of the present application. Electronic devices may vary significantly due to different configurations or performances, and may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740. The processor 710, the communications interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 may call a computer program stored in the memory 730 and executable on the processor 710 to perform the following steps: Obtain a first text and a second text; wherein the first text is the text to be checked for duplicate content; and the second text is the text used to check for duplicate content in the first text; Acquire multiple target keyword clusters of the historical text; wherein the target keywords in the target keyword clusters represent domain characteristics of the historical text; Using a pre-trained domain similarity model, domain similarity calculation is performed based on multiple target keyword clusters, the first text, and the second text to obtain a domain similarity result between the first text and the second text; Determine the duplicate checking results based on the domain similarity results.

[0103] In this embodiment, a first text and a second text are obtained; wherein the first text is a text to be checked for duplicates; the second text is a text used to check for duplicates on the first text; multiple target keyword clusters of a historical text are obtained; wherein the target keywords in the target keyword clusters represent domain characteristics of the historical text; domain similarity is calculated based on the multiple target keyword clusters, the first text, and the second text using a pre-trained domain similarity model to obtain domain similarity results for the first text and the second text; and a duplicate checking result is determined based on the domain similarity results. This solution, through the pre-trained domain similarity model, can automatically calculate domain similarity based on the multiple target keyword clusters, the first text, and the second text to obtain domain similarity results for the first text and the second text, thereby determining a duplicate checking result based on the domain similarity results. Furthermore, when the semantic similarity between the first text and the second text is high but the domain similarity is low, a false duplicate checking result is not made. Compared with related text duplicate checking technologies that only judge duplication based on the semantic similarity level, this solution determines a more accurate duplicate checking result, solving the problem of low accuracy of related text duplicate checking technologies.

[0104] The specific execution steps can refer to the various steps of the above-mentioned text duplication checking method embodiment, and can achieve the same technical effect. To avoid repetition, they will not be repeated here.

[0105] It should be noted that the electronic devices in the embodiments of the present application include: servers, terminals, or other devices other than terminals.

[0106] The above electronic device structure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. For example, the input unit may include a graphics processing unit (GPU) and a microphone, and the display unit may be configured as a display panel in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit includes at least one of a touch panel and other input devices. A touch panel is also called a touch screen. Other input devices may include, but are not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be detailed here.

[0107] The memory can be used to store software programs and various data. The memory may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory may include volatile memory or non-volatile memory, or the memory may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct rambus random access memory (DRRAM).

[0108] The processor may include one or more processing units; optionally, the processor may integrate an application processor and a modem processor, wherein the application processor primarily handles operations related to the operating system, user interface, and application programs, and the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into the processor.

[0109] An embodiment of the present application also provides a storage medium on which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the various processes of the above-mentioned text duplication checking method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, they will not be repeated here.

[0110] The processor is the processor in the electronic device in the above embodiment. The storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0111] An embodiment of the present application further provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned text duplication checking method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0112] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0113] An embodiment of the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the various processes of the above-mentioned text duplication checking method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0114] It should be noted that, in this article, the terms "comprises", "includes" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include multitasking and parallel processing according to the functions involved, and may also add, omit, or combine various steps. In addition, the features described with reference to certain examples may be combined in other examples.

[0115] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present application.

[0116] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A text duplication checking method, characterized in that: The method comprises: Obtain a first text and a second text; wherein the first text is the text to be checked for duplicate content; and the second text is the text used to check for duplicate content on the first text; Acquire multiple target keyword clusters of the historical text; wherein the target keywords in the target keyword clusters represent the domain characteristics of the historical text; Performing domain similarity calculation based on the plurality of target keyword clusters, the first text, and the second text using a pre-trained domain similarity model to obtain a domain similarity result between the first text and the second text; Based on the field similarity results, the duplicate checking results are determined.

2. The method according to claim 1, characterized in that The domain similarity model includes a rerank model; the pre-trained domain similarity model performs domain similarity calculation based on the plurality of target keyword clusters, the first text, and the second text to obtain a domain similarity result between the first text and the second text, including: Generate a first domain feature vector and a second domain feature vector based on the plurality of target keyword clusters, the first text, and the second text using a Rerank model; wherein the first domain feature vector is a first similarity score between the first text and the target keyword in descending order; and the second domain feature vector is a second similarity score between the second text and the target keyword in descending order; A similarity calculation is performed based on the first domain feature vector and the second domain feature vector to obtain the domain similarity result.

3. The method according to claim 2, characterized in that Before performing similarity calculation based on the first domain feature vector and the second domain feature vector to obtain the domain similarity result, the method further includes: Performing zero-to-one processing on the first domain feature vector and the second domain feature vector to obtain a first domain embedding vector and a second domain embedding vector; The domain similarity model further includes a pre-trained attention model; performing similarity calculation based on the first domain feature vector and the second domain feature vector to obtain the domain similarity result includes: The pre-trained attention model is used to perform similarity calculation based on the first domain embedding vector, the second domain embedding vector and the attention weight matrix to obtain the domain similarity result; wherein the attention weight matrix includes a weight for each of the target keyword clusters.

4. The method according to claim 1, wherein The method further comprises: Obtaining the historical text; Extracting keywords from the historical text using a large language model to obtain a plurality of initial keywords; Performing clustering and deduplication processing on the multiple initial keywords through a clustering model to obtain multiple target keyword clusters, and storing the multiple target keyword clusters in a keyword library; The step of obtaining multiple target keyword clusters from the historical text includes: A plurality of target keyword clusters of the historical text are obtained from the keyword library.

5. The method according to claim 4, characterized in that The clustering model is used to perform clustering and deduplication processing on the multiple initial keywords to obtain multiple target keyword clusters, including: Converting the initial keywords into word embeddings using a pre-trained language model; Performing similarity calculation on the word embeddings to obtain similarity results between the initial keywords, and constructing a similarity matrix based on the similarity results between the initial keywords; wherein the rows and columns of the similarity matrix are the number of the initial keywords; Using a density-based clustering algorithm, performing a first clustering process on the similarity matrix to obtain a plurality of keyword clusters; A hierarchical clustering algorithm and a preset keyword quantity threshold are used to perform a second clustering process and a deduplication process on the plurality of keyword clusters to obtain a plurality of target keyword clusters.

6. The method according to claim 1, characterized in that The method further comprises: Performing semantic embedding on the first text and the second text using a pre-trained semantic model to obtain a first semantic embedding vector and a second semantic embedding vector, and performing similarity calculation on the first semantic embedding vector and the second semantic embedding vector to obtain a semantic similarity result; Based on the field similarity results, determine the duplicate checking results, including: The duplicate checking result is generated according to the domain similarity result and the semantic similarity result.

7. A text duplication checking device, characterized in that: The device comprises: A first acquisition module is configured to acquire a first text and a second text; wherein the first text is a text to be checked for duplicates; and the second text is a text used to check for duplicates of the first text; A second acquisition module is configured to acquire a plurality of target keyword clusters of the historical text; wherein the target keywords in the target keyword clusters represent the domain characteristics of the historical text; a calculation module, configured to perform domain similarity calculation based on the plurality of target keyword clusters, the first text, and the second text using a pre-trained domain similarity model, and obtain a domain similarity result between the first text and the second text; The determination module is used to determine the duplicate checking result based on the field similarity result.

8. An electronic device, characterized in that: include: processor; as well as A memory arranged to store computer-executable instructions, wherein the executable instructions are configured to be executed by the processor, and the executable instructions include instructions for executing the text duplication checking method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is used to store computer-executable instructions, and the computer-executable instructions enable a computer to execute the text duplication checking method according to any one of claims 1 to 6.

10. A computer program product, characterized in that It includes a computer program, which, when executed by a processor, implements the text duplication checking method as described in any one of claims 1 to 6.