Knowledge data deduplication method and device, equipment and medium

By combining hash clustering and semantic vector models, efficient deduplication of auto insurance claims data is achieved, solving the problems of data redundancy and insufficient deduplication accuracy, and improving information processing efficiency and user experience.

CN120804315APending Publication Date: 2025-10-17CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510681291.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In the digital unmanned scenario of auto insurance claims, existing technologies have problems with data redundancy and insufficient efficiency and accuracy of deduplication technology. Especially when faced with large-scale data, traditional methods find it difficult to effectively identify semantically similar texts, resulting in the continued existence of redundant data.

Method used

The hash clustering algorithm is used to perform text clustering on knowledge documents, and the semantic vector model is combined to perform semantic deduplication within the document cluster. The similarity is calculated by local sensitive hash value and word vector matrix to achieve accurate deduplication of documents.

Benefits of technology

It improves the accuracy of deduplication and the efficiency of processing large-scale data, reduces redundant data storage, effectively solves the problem of data redundancy, and improves information retrieval efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804315A_ABST
    Figure CN120804315A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of big data, and relates to a knowledge data deduplication method, comprising: acquiring target knowledge data from a knowledge document library, the target knowledge data comprising a plurality of knowledge documents; based on a preset Hash clustering algorithm, performing text clustering on the plurality of knowledge documents to obtain a plurality of groups of document clusters; and adopting a preset semantic vector model to perform semantic deduplication on each knowledge document in the plurality of groups of document clusters to obtain a deduplicated document corresponding to the target knowledge data. The invention further provides a device, equipment and a medium. The method can be applied to the business field of financial insurance and the like, and the efficiency and accuracy of knowledge data deduplication can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of human big data, and is applied to online processing business scenarios such as finance and insurance, and particularly relates to a knowledge data deduplication method and device, equipment and a medium. BACKGROUND

[0002] With the rapid progress of artificial intelligence and big data technology, the generation and storage scale of knowledge data are growing explosively. In the digital unmanned scene of car insurance claims, a large model as a digital person needs to accurately identify the customer's intention, and this process will generate a large amount of non-intentional text. The analysis of these data becomes a core task, however, there are many challenges at present.

[0003] On the one hand, the problem of data redundancy is prominent. In actual application, knowledge data often contains a large amount of repeated or similar content, such as news reports, scientific research papers and user-generated content, and multiple documents may express similar information. This not only increases the storage cost, but also causes confusion when users search and obtain information, reduces the information retrieval efficiency, and affects the user experience. On the other hand, the existing deduplication technology has limitations. Most current data deduplication methods have certain effect when dealing with small-scale data or small text differences, but when facing large-scale data, the efficiency and accuracy are difficult to guarantee. Especially when the text surface difference is large but the semantic similarity is high, the traditional method is easy to miss repeated items, resulting in the persistence of redundant data.

[0004] In summary, the existing technology has problems of data redundancy and insufficient efficiency and accuracy of deduplication technology when dealing with non-intentional text data generated in the digital unmanned scene of car insurance claims. SUMMARY

[0005] The purpose of the embodiments of the present application is to provide a knowledge data deduplication method, device, computer equipment and storage medium to solve the problem of data redundancy and insufficient efficiency and accuracy of deduplication technology when the existing technology deals with non-intentional text data generated in the digital unmanned scene of car insurance claims.

[0006] In a first aspect, a knowledge data deduplication method is provided, which adopts the following technical solution:

[0007] Obtaining target knowledge data from a knowledge document library, the target knowledge data including a plurality of knowledge documents; performing text clustering on the plurality of knowledge documents based on a preset hash clustering algorithm to obtain a plurality of document clusters; and performing semantic deduplication on each knowledge document in the plurality of document clusters using a preset semantic vector model to obtain deduplicated documents corresponding to the target knowledge data.

[0008] In a second aspect, a knowledge data deduplication device is provided, which adopts the following technical solution:

[0009] The acquisition module is configured to acquire target knowledge data from the knowledge document library, the target knowledge data including a plurality of knowledge documents;

[0010] The clustering module is configured to perform text clustering on the plurality of knowledge documents based on a preset hash clustering algorithm to obtain a plurality of groups of document clusters;

[0011] The deduplication module is configured to perform semantic deduplication on the knowledge documents in the plurality of groups of document clusters by using a preset semantic vector model to obtain deduplicated documents corresponding to the target knowledge data.

[0012] In a third aspect, a computer device is provided, and the following technical solutions are adopted:

[0013] The target knowledge data is acquired from the knowledge document library, the target knowledge data including a plurality of knowledge documents; the plurality of knowledge documents are subjected to text clustering based on a preset hash clustering algorithm to obtain a plurality of groups of document clusters; and the knowledge documents in the plurality of groups of document clusters are subjected to semantic deduplication by using a preset semantic vector model to obtain deduplicated documents corresponding to the target knowledge data.

[0014] In a fourth aspect, a computer readable storage medium is provided, and the following technical solutions are adopted:

[0015] The target knowledge data is acquired from the knowledge document library, the target knowledge data including a plurality of knowledge documents; the plurality of knowledge documents are subjected to text clustering based on a preset hash clustering algorithm to obtain a plurality of groups of document clusters; and the knowledge documents in the plurality of groups of document clusters are subjected to semantic deduplication by using a preset semantic vector model to obtain deduplicated documents corresponding to the target knowledge data.

[0016] Compared with the prior art, the embodiments of the present application have the following beneficial effects: The target knowledge data is acquired from the knowledge document library, covering a plurality of knowledge documents, to provide a comprehensive data basis for subsequent processing. The hash clustering algorithm is used to perform text clustering on the knowledge documents, which can quickly cluster similar documents into clusters, effectively deal with the data redundancy problem, and reduce the complexity of subsequent processing. The semantic vector model is used to perform semantic deduplication on the knowledge documents in the document clusters. Compared with the traditional method, the model can deeply understand the text semantics, accurately identify and deduplicate even if the text surface difference is large but the semantics are similar. This not only improves the deduplication accuracy, but also improves the efficiency when processing large-scale data, reduces the storage of redundant data, and effectively solves the problems of data redundancy and insufficient efficiency and accuracy of deduplication technology in the prior art. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the solutions in the present application, the drawings needed to be used in the description of the embodiments of the present application will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.

[0018] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;

[0019] Figure 2 a flow chart of one embodiment of the knowledge data deduplication method according to the present application;

[0020] Figure 3 is a structural schematic diagram of one embodiment of the knowledge data deduplication apparatus according to the present application;

[0021] Figure 4 is a structural schematic diagram of one embodiment of the computer device according to the present application. DETAILED DESCRIPTION

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs; the terminology used in the description of the application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application; the description and the claims of the present application and the above-mentioned drawings in the specification illustrate the present application; the terms "comprise" and "have" and any variations thereof in the present application and the claims of the specification and the above-mentioned drawings are intended to cover not exclusive inclusion; the terms "first", "second", etc. in the present application and the claims of the specification or the above-mentioned drawings are used to distinguish different objects, not to describe a particular order.

[0023] In the present application, the phrase "embodiment" means that the specific features, structures or characteristics described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears at various places in the specification does not necessarily refer to the same embodiment, nor is it mutually exclusive or alternative to other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0024] In order to make those skilled in the art better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings.

[0025] As Figure 1As shown, the system architecture 100 can include a terminal device 101, a network 102 and a server 103. The terminal device 101 can be a notebook computer 1011, a tablet computer 1012 or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 can include various connection types, such as wired, wireless communication link or optical fiber cable, etc.

[0026] A user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0027] The terminal device 101 can be various electronic devices with display screens and supporting web browsing, in addition to the notebook computer 1011, the tablet computer 1012 or the mobile phone 1013, the terminal device 101 can also be an electronic book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer and a desktop computer, etc.

[0028] The server 103 can be a server providing various services, such as a background server supporting the pages displayed on the terminal device 101.

[0029] It should be noted that the knowledge data deduplication method provided by the embodiments of the present application is generally executed by a server / terminal device, and accordingly, the knowledge data deduplication apparatus is generally provided in a server / terminal device.

[0030] It should be understood that Figure 1 The number of terminal devices, networks and servers in

[0031] With reference to Figure 2 , a flow chart of one embodiment of the method for service recommendation according to the present application is shown. The knowledge data deduplication method includes the following steps:

[0032] Step S201, obtaining target knowledge data from a knowledge document library, the target knowledge data including a plurality of knowledge documents.

[0033] In the present embodiment, the electronic device (for example, the terminal device 101 or the server 103) on which the knowledge data deduplication method runs can be a terminal device, a server or a terminal device and a server.Figure 1 The server / terminal device shown can obtain target knowledge data through wired or wireless connection. It should be noted that the wireless connection can include, but is not limited to, 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other now known or future developed wireless connection methods.

[0034] The knowledge document library refers to a collection system for storing structured or unstructured knowledge data, usually built by a database, file system or distributed storage architecture, and its sources can include historical claims records, user interaction logs and external public data sources. For example, the collection of accident reports, clause descriptions and user consultation records stored in the car insurance claims scene over the years.

[0035] The target knowledge data refers to a set of text to be processed extracted from the knowledge document library. It represents a set of original data containing semantic duplication risks, which is used for subsequent clustering and deduplication operations. For example, 100,000 user consultation records generated during a car insurance claims peak.

[0036] The plurality of knowledge documents refers to the basic text units that make up the target knowledge data, which can come from different users, different channels or different time nodes, and each document has independent semantic features and is used as the smallest processing object for clustering analysis.

[0037] Step S202, based on the preset hash clustering algorithm, text clustering is performed on the plurality of knowledge documents to obtain a plurality of document clusters.

[0038] The hash clustering algorithm refers to a text grouping method based on the local sensitive hash technology. The text is converted into a fixed number of hash values by simhash algorithm, and the similarity can be measured by Hamming distance.

[0039] Text clustering refers to the technical process of grouping according to the similarity of text content. It is used to reduce the computational complexity of subsequent semantic analysis. For example, 5000 documents related to "vehicle damage assessment process" are clustered into 3 document clusters.

[0040] The plurality of document clusters refers to logical groups formed after text clustering, each group containing a number of documents with similar characteristics. It is used to support fine-grained semantic deduplication processing. For example, the first document cluster contains 200 documents discussing "claims material list".

[0041] Step S203, using a preset semantic vector model, the semantic deduplication is performed on each knowledge document in the plurality of document clusters to obtain the deduplicated document corresponding to the target knowledge data.

[0042] The semantic vector model refers to a text representation model based on deep learning. It is used to break through the semantic limitations of traditional hash algorithms.

[0043] The knowledge documents in each document cluster refer to specific text instances belonging to the same document cluster. After coarse-grained clustering, each cluster contains a number of documents to be analyzed semantically, which are used as input objects for semantic deduplication.

[0044] The semantic deduplication refers to a technical process of comparing the similarity of documents using a semantic vector model based on clustering. It is used to solve the missed detection problem of semantic repetition in traditional methods.

[0045] The deduplicated documents refer to the final text collection formed after hash clustering and semantic deduplication. They represent a high-quality dataset that retains core semantics and eliminates redundancy. For example, 23,000 non-redundant core knowledge documents are selected from 100,000 original records.

[0046] The embodiments of the present application can obtain target knowledge data from a knowledge document library, covering multiple knowledge documents, and providing a comprehensive data basis for subsequent processing. The hash clustering algorithm is used to cluster the knowledge documents, which can quickly cluster similar documents into clusters, effectively addressing the data redundancy problem and reducing the complexity of subsequent processing. The semantic vector model is used for semantic deduplication of knowledge documents in each document cluster. Compared with traditional methods, this model can deeply understand the semantics of the text, accurately identify and deduplicate even if the text has large surface differences but similar semantics. This not only improves the accuracy of deduplication, but also improves the efficiency of processing large-scale data, reduces redundant data storage, and effectively solves the problems of data redundancy and insufficient efficiency and accuracy of deduplication technology in the prior art.

[0047] In some optional implementations of the present embodiment, step 202, based on a preset hash clustering algorithm, text clustering is performed on a plurality of knowledge documents to obtain a plurality of document clusters, specifically including the following steps:

[0048] Feature extraction is performed on the plurality of knowledge documents to obtain a plurality of semantic features of each knowledge document. A weighted feature vector corresponding to each semantic feature is obtained. The weighted feature vector is binarized and reduced in dimension based on a preset hash clustering algorithm to obtain a local sensitive hash value of each knowledge document. The similarity between the local sensitive hash values of any two knowledge documents is calculated. Based on the similarity, the plurality of knowledge documents are grouped to obtain a plurality of document clusters.

[0049] The plurality of semantic features refer to a feature set extracted from the text content to represent different dimensional semantic information. It is used to construct a vector representation that fully reflects the semantics of the text.

[0050] The weighted feature vector is a numerical vector representation formed by assigning weight coefficients to different semantic features. It is used to preserve key semantic information during dimensionality reduction.

[0051] Binarization dimensionality reduction is the process of converting continuous feature vectors into binary codes. This can be achieved using a pre-defined hashing algorithm, such as simhash or a locality-sensitive hashing algorithm. Binarization compresses high-dimensional vectors into fixed-bit binary hash values, improving the efficiency of large-scale text similarity calculations.

[0052] A locally sensitive hash value is a binary code generated by a specific hash function that maintains the similarity of the original data. It is used to quickly filter out potential duplicate documents. For example, the Hamming distance between the locally sensitive hash values ​​[101011] and [101001] of two articles discussing "full responsibility determination" is 2.

[0053] Similarity refers to a quantitative indicator of the similarity between two locally sensitive hash values, which can be obtained by calculating the Hamming distance, etc. It is used to determine whether documents belong to the same cluster.

[0054] In one example, in a digital, unmanned auto insurance claims scenario, an insurance company uses the solution of this embodiment to clean up redundancies in its historical claims knowledge base. First, knowledge documents related to auto insurance claims processes, cases, and clauses can be selected from the auto insurance claims knowledge document library as target knowledge data. For example, 500 documents covering customer consultations and claims cases from different regions can be included. Feature extraction can then be performed on these 500 documents, such as semantic features such as claims steps, types of insurance involved, and accident types. A weighted feature vector corresponding to each semantic feature is obtained. Binarization and dimensionality reduction are performed on the weighted feature vectors using a hash clustering algorithm, converting continuous-valued feature vectors into locally sensitive hash values. The similarity between the locally sensitive hash values ​​of any two knowledge documents, such as the Hamming distance, is calculated. If the Hamming distance between two documents is less than a set threshold, they are considered similar. Based on this similarity, the 500 documents are grouped into 20 document clusters. Subsequently, a preset semantic vector model can be used to perform semantic deduplication on the knowledge documents within each of the 20 document clusters. The model can deeply understand the semantics of documents, identify documents with large surface differences but similar semantics, and remove redundant documents.

[0055] The embodiments of the present application can accurately capture a plurality of semantic features contained in each knowledge document by performing feature extraction on a plurality of knowledge documents. For example, in the financial insurance document scenario, key semantic features such as claim event type, involved insurance type, and claim amount range can be extracted, laying a solid foundation for subsequent processing. The weighted feature vector corresponding to each semantic feature is obtained, different weights are assigned according to the importance of the semantic feature in the document, which can highlight the core semantic information, suppress the interference of secondary information, and improve the accuracy of data processing. The weighted feature vector is binarized and reduced in dimension to obtain a local sensitive hash value, which effectively reduces the data dimension, greatly improves the calculation efficiency, and at the same time preserves the semantic similarity information. The similarity of the local sensitive hash values of any two knowledge documents is calculated, which can quickly quantify the semantic similarity degree between documents. Based on the similarity, the documents are grouped to obtain a plurality of document clusters, and documents with similar semantics are grouped together, realizing the preliminary classification and structured arrangement of massive documents.

[0056] In some optional implementations, the step of "obtaining a weighted feature vector corresponding to each semantic feature" specifically includes the following steps:

[0057] Obtaining the weight of each semantic feature; obtaining a high-dimensional vector of each semantic feature to obtain a plurality of embedding vectors corresponding to each knowledge document, the embedding vectors representing the semantic information of the corresponding semantic features; performing weighted sum operation on the plurality of embedding vectors and the weight of each knowledge document to obtain a weighted feature vector of each knowledge document.

[0058] The weight refers to the relative importance coefficient of a single semantic feature, which is represented by a numerical parameter in the interval [0, 1] and is used to adjust the contribution of different semantic dimensions in the text representation. For example, in the car insurance claim scenario, the "accident type" feature is assigned a weight of 0.5, which is significantly higher than the weight of 0.1 of the "user emotion" feature.

[0059] The plurality of embedding vectors refer to high-dimensional numerical vectors generated by a language model to represent a single semantic feature. The mathematical mapping represents a specific semantic concept and is used to capture the contextual semantic information of a text segment.

[0060] The weighted sum operation refers to a mathematical operation process of linearly combining a plurality of embedding vectors and their corresponding weights, which is realized by vector dot product and weight product accumulation.

[0061] In an example, taking a knowledge document in a plurality of knowledge documents as an example of a car insurance claim document. The semantic features of the car insurance claim document can be extracted, such as "accident type", "vehicle brand", "payout amount range", etc. The weights of each semantic feature can be determined based on historical claim data and business rules. For example, the accident type has a great influence on the claim decision, and is given a higher weight of 0.4. The vehicle brand has a certain influence on the payout amount, and the weight is set to 0.2. The payout amount range directly affects the claim amount, and the weight is set to 0.4. Then, the pre-trained word vector model can be used to obtain a high-dimensional vector for each semantic feature. This operation is performed on all semantic features in each knowledge document to obtain a plurality of embedding vectors corresponding to each document. The plurality of embedding vectors of each knowledge document are weighted and operated with the corresponding weights. For example, the embedding vector corresponding to "accident type" is V1, the weight is 0.4; the embedding vector corresponding to "vehicle brand" is V2, the weight is 0.2; the embedding vector corresponding to "payout amount range" is V3, the weight is 0.4, and the weighted feature vector of the document is

[0062] V = 0.4V1 + 0.2V2 + 0.4V3.

[0063] Embodiments of the present application obtain the weight of each semantic feature, fully considering the importance difference of different semantic features in the knowledge document. For example, in the insurance claim document, the accident liability identification semantic feature has a greater influence on the claim decision than the vehicle color feature, and a higher weight is given to highlight the key information and lay a reasonable foundation for subsequent processing. Obtaining a high-dimensional vector for each semantic feature obtains a plurality of embedding vectors, which accurately represent the semantic feature information in numerical form, convert abstract semantics into computable data, and facilitate subsequent operations. The plurality of embedding vectors are weighted and operated with the weights to obtain a weighted feature vector. This process integrates the importance and semantic information of each semantic feature, retaining the core semantics and suppressing the interference of secondary information. In this way, complex knowledge documents can be converted into simple and representative weighted feature vectors, not only reducing the data dimension and improving the computing efficiency, but also enhancing the feature's expression ability for document semantics.

[0064] In some optional implementations, the step of "performing binary dimension reduction on the weighted feature vector based on a preset hash clustering algorithm to obtain a local sensitive hash value of each knowledge document" specifically includes the following steps:

[0065] Comparing the dimension value of each dimension of the weighted feature vector with a preset threshold value based on a preset hash clustering algorithm. If the dimension value is greater than the threshold value, the corresponding dimension value is changed to an active value. If the dimension value is less than or equal to the threshold value, the corresponding dimension value is changed to an inactive value. Based on the active value and the inactive value, a local sensitive hash value of each knowledge document is generated.

[0066] wherein the dimension value refers to the numerical quantification result of the single coordinate axis direction in the weighted feature vector, used to reflect the relative position of the specific semantic feature in the vector space.

[0067] wherein the activation value refers to the binary state identifier set after threshold comparison, when the dimension value exceeds the preset threshold, it is assigned as "1" state value based on the hash clustering algorithm, derived from the binary dimension reduction processing process, representing the discretization representation of high-dimensional semantic information, used to construct the binary coding unit of the local sensitive hash value.

[0068] wherein the inactivation value refers to the binary state identifier set after threshold comparison, when the dimension value does not exceed the preset threshold, it is assigned as "0" state value, derived from the binary processing process of the feature vector, representing the negative identification of insufficient semantic components, used to construct the sparse hash code together with the activation value.

[0069] In an example, each dimension of the weighted feature vector is judged based on the hash clustering algorithm, such as SimHash algorithm, and the threshold is set to 0. If the dimension value is greater than 0, the corresponding SimHash bit is 1. If the dimension value is less than or equal to 0, the corresponding SimHash bit is 0. Finally, a fixed-length binary string, i.e. SimHash value, is generated, which is the local sensitive hash value of the knowledge document corresponding to the weighted vector. For example, the weighted feature vector A is [+1, -1, +1, -1, +5, +1], and the weighted feature vector B is [+1, -1, -1, +1, +3, +1], then the local sensitive hash value A corresponding to the weighted feature vector A is [101011], and the local sensitive hash value B corresponding to the weighted feature vector B is [100111].

[0070] The embodiment of the present application compares the dimension value of the weighted feature vector with the preset threshold, changes the dimension value to the activation value or the inactivation value, and then generates the local sensitive hash value. On the one hand, it realizes efficient dimension reduction of data. The weighted feature vector often has high dimension, and direct processing has large calculation amount. By converting to activation value and inactivation value through comparison with threshold, the data dimension is greatly reduced, the complexity of subsequent calculation is reduced, and the processing efficiency is improved. On the other hand, the key semantic information is retained. In the dimension reduction process, the threshold is used as the boundary to distinguish the activation value and the inactivation value, which can highlight important semantic features and suppress secondary information, so that the generated local sensitive hash value can still accurately represent the semantic content of the document. In addition, it lays a foundation for document similarity calculation and grouping. The local sensitive hash value generated based on the activation value and the inactivation value facilitates the rapid calculation of the similarity between documents, and then realizes efficient document grouping.

[0071] In some optional implementations, the step of "calculating the similarity between the local sensitive hash values of any two knowledge documents" specifically includes the following steps:

[0072] The number of bits of each local sensitive hash value is obtained; the binary bits of the local sensitive hash values of any two knowledge documents are compared bit by bit to determine the number of different bits; the Hamming distance between the local sensitive hash values of any two knowledge documents is determined based on the number; and the similarity between the local sensitive hash values of any two knowledge documents is determined based on the number of bits and the Hamming distance.

[0073] The number of bits refers to the length of the binary encoding of the local sensitive hash value. It is used to determine the quantization range for the similarity comparison.

[0074] The binary bit refers to the smallest unit of data that makes up the local sensitive hash value, which is generated by binary dimensionality reduction processing. Each bit has a value of 0 or 1, representing the activation state identifier of a single semantic dimension.

[0075] The number of different bits refers to the total number of binary bits that take different values in the same bit sequence between two local sensitive hash values. It is obtained by counting the number of 1s after performing an exclusive OR operation bit by bit. For example, when comparing local sensitive hash value A = [101011] with local sensitive hash value B = [101001], only the 5th bit is different, and the number of different bits is 1.

[0076] The Hamming distance refers to the number of different characters in the corresponding positions of two binary sequences of the same length.

[0077] In an example, from the knowledge document library, such as insurance clause documents, after feature extraction, weighting, and binary dimensionality reduction, the local sensitive hash value of each document is obtained. For example, the local sensitive hash value of a certain insurance clause document is "101101", and the number of bits is 6. The binary bits of the local sensitive hash values of any two knowledge documents are compared bit by bit. For example, the local sensitive hash value of document A is "101101", and the local sensitive hash value of document B is "111001". Comparing bit by bit, it can be seen that the 2nd and 4th bits are different, and the number of different bits is 2, i.e., the Hamming distance is 2. The similarity is determined based on the number of bits and the Hamming distance. Assuming the number of bits is n and the Hamming distance is d, the similarity S = 1 - d / n. In the above example, n = 6 and d = 2, so the similarity S = 1 - 2 / 6 ≈ 0.67.

[0078] The embodiment of the application provides a unified quantification basis for subsequent similarity calculation by acquiring the bit number of each local sensitive hash value. By comparing the binary bits of the local sensitive hash values of any two knowledge documents bit by bit and determining the number of different bits, the embodiment of the application accurately captures the embodiment of the semantic difference between the documents at the hash value level, and reflects the subtle differences in document semantics in the smallest unit of binary bits. The Hamming distance is determined based on the number of different bits, which quantifies the semantic distance between the documents simply and efficiently, and provides a key indicator for similarity calculation. The similarity is determined based on the bit number and the Hamming distance, which comprehensively considers the overall length and difference degree of the hash value, and the calculation result is more scientific and reasonable. Through the process, the similarity between the documents can be quickly and accurately evaluated, and a reliable basis is provided for subsequent document grouping and semantic deduplication.

[0079] In some optional implementations, in step S203, the preset semantic vector model is used to perform semantic deduplication on the knowledge documents in the multiple groups of document clusters to obtain deduplicated documents corresponding to the target knowledge data, and specifically includes the following steps:

[0080] The multiple groups of document clusters are converted into word vector matrices by using the preset semantic vector model, to obtain multiple document word vector matrices in each group of document clusters; the cosine similarity between any two document word vector matrices in each group of document clusters is calculated; and the semantic deduplication is performed on the knowledge documents in the multiple groups of document clusters based on the cosine similarity, to obtain deduplicated documents corresponding to the target knowledge data.

[0081] The word vector matrix refers to a numerical representation set of high-dimensional vectors obtained by mapping text words into the high-dimensional vectors through a pre-trained semantic model.

[0082] The multiple document word vector matrices refer to a set of word vector matrices independently generated for different knowledge documents in the same document cluster, and each matrix corresponds to the semantic encoding result of a single document.

[0083] In an example, in a car insurance claim digital unmanned scene, an insurance company adopts the scheme of the embodiment to perform redundancy cleaning on a historical claim knowledge base. First, knowledge documents related to car insurance claim process, cases, clauses, etc. can be selected from the car insurance claim knowledge document base as target knowledge data, for example, 500 different regional customer consultation and claim case documents. Then, feature extraction can be performed on the 500 documents, such as extracting semantic features such as claim steps, involved insurance types, and accident types. The weighted feature vector corresponding to each semantic feature is obtained. The weighted feature vector is binarized and reduced in dimension, and the continuous value feature vector is converted into a local sensitive hash value. The similarity, such as the Hamming distance, between the local sensitive hash values of any two knowledge documents is calculated. If the Hamming distance of two documents is less than a set threshold, they are considered similar, and the 500 documents are grouped based on the similarity to obtain 20 document clusters. Subsequently, a preset semantic vector model can be used to convert each knowledge document in the 20 document clusters into a word vector matrix to obtain 20 sets of document word vector matrices in the document clusters. Then, the cosine similarity between any two document word vector matrices in the 20 sets of document clusters can be calculated. Finally, based on the cosine similarity, semantic deduplication can be performed on each knowledge document in the 20 sets of document clusters, and the deduplicated documents corresponding to the 500 documents can be obtained.

[0084] The embodiment of the present application converts the knowledge documents in multiple document clusters into word vector matrices through a semantic vector model, converts the text information of the documents into structured numerical representation, accurately captures the semantic information of the words in the documents, and makes documents with different expressions but similar semantics have similar location characteristics in the vector space. By calculating the cosine similarity between any two document word vector matrices in each document cluster, the semantic association degree between the documents is quantified in a mathematical way, which can accurately measure the semantic proximity of the documents and avoid the limitation of relying only on the surface form of the text to judge similarity. Based on the cosine similarity, semantic deduplication can efficiently identify and remove semantically repetitive documents, while retaining the core semantic information of the documents and significantly reducing document redundancy.

[0085] In some optional implementations, the step of "performing semantic deduplication on each knowledge document in the multiple document clusters based on the cosine similarity to obtain deduplicated documents corresponding to the target knowledge data" specifically includes the following steps:

[0086] Based on the cosine similarity, the similarity scores of each document in each document cluster with all documents in the cluster are calculated; the similarity scores of each document in each document cluster are sorted to obtain a sorting result; based on the sorting result, the documents to be retained in each document cluster are determined to obtain the deduplicated documents corresponding to the target knowledge data.

[0087] The similarity score is a quantitative evaluation value formed by aggregating the cosine similarity calculation results of the document and other documents in the cluster.

[0088] In an example, a set of document clusters includes document A, document B, document C, and document D. The similarity score between two documents can be determined by the cosine similarity between the two documents. As shown in the following table, the similarity score between document A and document B is 0.9, the similarity score between document A and document C is 0.8, and the similarity score between document A and document D is 0.7. The similarity score between document B and document A is 0.9, the similarity score between document B and document C is 0.85, and the similarity score between document B and document D is 0.75. The similarity score between document D and document A is 0.7, the similarity score between document D and document B is 0.75, and the similarity score between document D and document C is 0.6.

[0089]

[0090] For each document in the document cluster, the second highest similarity score is found. For example, according to the similarity scores between document A and document B, document C, and document D, the second highest similarity score is 0.8. According to the similarity scores between document B and document A, document C, and document D, the second highest similarity score is 0.85. According to the similarity scores between document C and document A, document B, and document D, the second highest similarity score is 0.8. According to the similarity scores between document D and document A, document B, and document C, the second highest similarity score is 0.7. The second highest similarity scores can then be sorted, and the highest score, 0.85, is found to correspond to document B. Therefore, document B can be selected as the representative sample after deduplication, i.e., the corresponding deduplicated document in the document cluster. Similarly, the deduplicated document of the corresponding document cluster can be determined for other document clusters by the above method. Finally, the deduplicated document corresponding to the target knowledge data is obtained.

[0091] The embodiments of the present application can calculate the similarity score of each document in each set of document clusters with all documents in the cluster by cosine similarity. This process fully utilizes the precise quantitative ability of cosine similarity to measure the semantic similarity of documents in a vector space, comprehensively measures the semantic association degree of documents in the cluster, and makes the similarity score objectively reflect the semantic similarity level of the documents. The similarity scores are sorted to present the similarity differences between the documents in the form of a direct numerical sequence, providing a clear decision basis for document deduplication. Based on the sorting result, the remaining documents in each set of document clusters are determined, which can efficiently remove semantic duplicate documents while retaining core semantic information and significantly reducing document redundancy.

[0092] It should be emphasized that, in order to further ensure the privacy and security of the target knowledge data and the deduplicated document, the target knowledge data and the deduplicated document can also be stored in a node of a block chain.

[0093] The block chain referred to in the present application is a new application mode of distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm and other computer technologies. The block chain is essentially a decentralized database, which is a series of data blocks associated using cryptographic methods, each data block containing information of a batch of network transactions, used to verify the validity (anti-fake) of the information and generate the next block. The block chain can include a block chain underlying platform, a platform product service layer, and an application service layer, etc.

[0094] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by computer readable instructions instructing related hardware, and the computer readable instructions can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments of each method. The storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0095] It should be understood that, although each step in the flowchart of the accompanying drawings is shown in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and they can be executed in other orders. Moreover, at least part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or sub-steps or stages of other steps.

[0096] Further referring to Figure 3 , as an implementation of the method shown in Figure 2 , the present application provides an embodiment of a knowledge data deduplication device, which corresponds to the method embodiment shown in Figure 2 . The device can be specifically applied to various electronic devices.

[0097] As shown in Figure 4 , the knowledge data deduplication device 400 of the present embodiment includes an acquisition module 401, a clustering module 402 and a deduplication module 403. Wherein:

[0098] The acquisition module 401 is configured to acquire target knowledge data from a knowledge document library, the target knowledge data including a plurality of knowledge documents.

[0099] The clustering module 402 is configured to perform text clustering on the plurality of knowledge documents based on a preset hash clustering algorithm to obtain a plurality of document clusters.

[0100] The deduplication module 403 is configured to perform semantic deduplication on the knowledge documents in the plurality of document clusters by using a preset semantic vector model to obtain deduplicated documents corresponding to the target knowledge data.

[0101] In this embodiment, the target knowledge data can be acquired from the knowledge document library, covering a plurality of knowledge documents, to provide a comprehensive data basis for subsequent processing. The hash clustering algorithm can be used to quickly cluster similar documents into clusters, effectively addressing the data redundancy problem and reducing the complexity of subsequent processing. The semantic vector model can be used to perform semantic deduplication on the knowledge documents in the document clusters. Compared with the traditional method, the model can deeply understand the text semantics, accurately identify and deduplicate the knowledge documents even if the texts have large surface differences but similar semantics. This not only improves the deduplication accuracy, but also improves the efficiency when processing large-scale data, reduces the storage of redundant data, and effectively solves the problems of data redundancy and insufficient efficiency and accuracy of the deduplication technology in the prior art.

[0102] In an embodiment, the clustering module 402 includes:

[0103] The extraction sub-module is configured to perform feature extraction on the plurality of knowledge documents to obtain a plurality of semantic features of each knowledge document.

[0104] The vector acquisition sub-module is configured to acquire a weighted feature vector corresponding to each semantic feature.

[0105] The dimension reduction sub-module is configured to perform binary dimension reduction on the weighted feature vector based on a preset hash clustering algorithm to obtain a local sensitive hash value of each knowledge document.

[0106] The first calculation sub-module is configured to calculate the similarity between the local sensitive hash values of any two knowledge documents.

[0107] The grouping sub-module is configured to group the plurality of knowledge documents based on the similarity to obtain a plurality of document clusters.

[0108] The embodiments of the present application can accurately capture a plurality of semantic features contained in each knowledge document by performing feature extraction on a plurality of knowledge documents. For example, in the financial insurance document scenario, key semantic features such as claim event type, involved insurance type, and claim amount range can be extracted, thereby laying a solid foundation for subsequent processing. The weighted feature vector corresponding to each semantic feature is obtained, different weights are assigned according to the importance of the semantic feature in the document, the core semantic information is highlighted, the secondary information interference is suppressed, and the data processing accuracy is improved. The weighted feature vector is binarized and reduced in dimension to obtain a local sensitive hash value, which effectively reduces the data dimension, greatly improves the calculation efficiency, and at the same time preserves the semantic similarity information. The similarity of the local sensitive hash values of any two knowledge documents is calculated, which can quickly quantify the semantic similarity degree between the documents. Based on the similarity, a plurality of document clusters are obtained by grouping the documents, and documents with similar semantics are grouped together, thereby realizing the preliminary classification and structured arrangement of a large number of documents.

[0109] In an embodiment, the vector obtaining submodule is further configured to obtain a weight of each semantic feature; obtain a high-dimensional vector of each semantic feature to obtain a plurality of embedding vectors corresponding to each knowledge document, the embedding vectors representing semantic information of the corresponding semantic features; and perform weighted sum operation on the plurality of embedding vectors and the weight of each knowledge document to obtain a weighted feature vector of each knowledge document.

[0110] The embodiments of the present application obtain the weight of each semantic feature, fully consider the importance difference of different semantic features in the knowledge document, for example, in the insurance claim document, the accident liability identification semantic feature has a greater impact on the claim decision than the vehicle color feature, and a higher weight is assigned to highlight the key information and lay a reasonable foundation for subsequent processing. A plurality of embedding vectors are obtained by obtaining a high-dimensional vector of each semantic feature, which accurately represent the semantic feature information in digital form, convert abstract semantics into calculable data, and facilitate subsequent operations. The weighted sum operation is performed on the plurality of embedding vectors and the weight to obtain a weighted feature vector, which integrates the importance and semantic information of each semantic feature, retains the core semantics, and suppresses the secondary information interference. In this way, the complex knowledge document can be converted into a simple and representative weighted feature vector, which not only reduces the data dimension, improves the calculation efficiency, but also enhances the feature expression ability of the document semantics.

[0111] In an embodiment, the dimension reduction submodule is further configured to compare, based on a preset hash clustering algorithm, a dimension value of each dimension of the weighted feature vector with a preset threshold value; if the dimension value is greater than the threshold value, the corresponding dimension value is changed to an active value; if the dimension value is less than or equal to the threshold value, the corresponding dimension value is changed to an inactive value; and generate a local sensitive hash value of each knowledge document based on the active value and the inactive value.

[0112] The embodiment of the application compares the dimension value of the weighted feature vector with the preset threshold value, changes the dimension value into an activated value or an inactivated value, and then generates a local sensitive hash value. On the one hand, the high-efficiency dimension reduction of data is realized. The weighted feature vector often has a high dimension, and the calculation amount is large if it is directly processed. By comparison with the threshold value, the dimension of data is greatly reduced, the complexity of subsequent calculation is reduced, and the processing efficiency is improved. On the other hand, the key semantic information is retained. In the dimension reduction process, the activated value and the inactivated value are distinguished by the threshold value, the important semantic features are highlighted, and the secondary information is suppressed, so that the generated local sensitive hash value can still accurately represent the semantic content of the document. In addition, the foundation is laid for document similarity calculation and grouping. The local sensitive hash value generated based on the activated value and the inactivated value facilitates the rapid calculation of the similarity between documents, and then the efficient document grouping is realized.

[0113] In an embodiment, the first calculation submodule is further configured to obtain the number of bits of each local sensitive hash value; compare the binary bits of the local sensitive hash values of any two knowledge documents bit by bit, determine the number of different bits; determine the Hamming distance between the local sensitive hash values of any two knowledge documents based on the number; and determine the similarity between the local sensitive hash values of any two knowledge documents based on the number of bits and the Hamming distance.

[0114] The embodiment of the application provides a unified quantitative basis for subsequent similarity calculation by obtaining the number of bits of each local sensitive hash value. By comparing the binary bits of the local sensitive hash values of any two knowledge documents bit by bit and determining the number of different bits, the embodiment accurately captures the embodiment of the semantic difference between documents at the hash value level, and reflects the subtle differences in document semantics in the smallest unit of binary bits. The Hamming distance is determined based on the number of different bits, which quantifies the semantic distance between documents simply and efficiently, and provides a key indicator for similarity calculation. The similarity is determined based on the number of bits and the Hamming distance, which comprehensively considers the overall length and difference degree of the hash value, and the calculation result is more scientific and reasonable. Through the process, the similarity between documents can be quickly and accurately evaluated, and a reliable basis is provided for subsequent document grouping and semantic deduplication.

[0115] In an embodiment, the deduplication module 403 comprises:

[0116] The conversion submodule is configured to convert each knowledge document in each document cluster into a word vector matrix by using a preset semantic vector model, to obtain a plurality of document word vector matrices in each document cluster.

[0117] The second calculation submodule is configured to calculate the cosine similarity between any two document word vector matrices in each document cluster.

[0118] The deduplication submodule is configured to perform semantic deduplication on the knowledge documents in each group of document clusters based on cosine similarity, to obtain deduplicated documents corresponding to the target knowledge data.

[0119] The semantic vector model is used to convert the knowledge documents in each group of document clusters into a word vector matrix, to convert the text information of the documents into a structured numerical representation, to accurately capture the semantic information of the words in the documents, and to make documents with similar semantics but different expressions have similar location characteristics in the vector space. The cosine similarity between the word vector matrices of any two documents in each group of document clusters is calculated, to quantitatively measure the semantic correlation between the documents in a mathematical manner, to accurately measure the semantic proximity of the documents, and to avoid the limitation of relying on the surface form of the text to determine the similarity. The semantic deduplication based on the cosine similarity can efficiently identify and remove the semantically repeated documents, while retaining the core semantic information of the documents and greatly reducing the redundancy of the documents.

[0120] In an embodiment, the deduplication submodule is further configured to calculate the similarity scores of each document in each group of document clusters with all the documents in the cluster based on the cosine similarity, to sort the similarity scores of each document in each group of document clusters, to obtain a sorting result, and to determine the documents to be retained in each group of document clusters based on the sorting result, to obtain the deduplicated documents corresponding to the target knowledge data.

[0121] The similarity scores of each document in each group of document clusters with all the documents in the cluster can be calculated based on the cosine similarity, which fully utilizes the accurate quantitative ability of the cosine similarity for the semantic similarity of the documents in the vector space, comprehensively measures the semantic correlation of the documents in the cluster, and makes the similarity scores objectively reflect the semantic proximity of the documents with other documents. The similarity scores are sorted to present the similarity differences between the documents in the form of a numerical sequence, to provide a clear decision basis for document deduplication. The documents to be retained in each group of document clusters are determined based on the sorting result, to efficiently remove the semantically repeated documents, while retaining the core semantic information and significantly reducing the redundancy of the documents.

[0122] To solve the above technical problems, the embodiment of the present application further provides a computer device. For details, please refer to Figure 4 , Figure 4 The basic structure block diagram of the computer device of the present embodiment is shown in FIG. 1.

[0123] The computer device 4 includes a memory 61, a processor 62, and a network interface 63, which are communicatively connected via a system bus. It should be noted that the computer device 6 is shown as having a memory 61, a processor 62, and a network interface 63, but it should be understood that not all of the illustrated components need be implemented, and that more or fewer components can alternatively be implemented. As will be appreciated by one of ordinary skill in the art, the computer device is a device that can automatically process data and / or information in accordance with a previously programmed or stored set of instructions, and includes, but is not limited to, a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, and the like.

[0124] The computer device can be a desktop computer, a notebook computer, a palm computer, a cloud server, and the like. The computer device can interact with a user through a keyboard, a mouse, a remote controller, a touchpad, a voice control device, and the like.

[0125] The memory 61 includes at least one type of readable storage medium, including a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, and the like), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, and the like. In some embodiments, the memory 61 can be an internal storage unit of the computer device 6, such as a hard disk or a memory of the computer device 6. In other embodiments, the memory 61 can also be an external storage device of the computer device 6, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, and the like. Of course, the memory 61 can include both an internal storage unit and an external storage device of the computer device 6. In the present embodiment, the memory 61 is generally used to store an operating system and various application software installed in the computer device 6, such as computer readable instructions of the knowledge data deduplication method, and the like. In addition, the memory 61 can also be used to temporarily store various data that has been output or will be output.

[0126] The processor 62 may, in some embodiments, be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 62 is generally used to control the overall operation of the computer device 6. In the present embodiment, the processor 62 is used to run computer readable instructions or process data stored in the memory 61, such as computer readable instructions for running the knowledge data deduplication method.

[0127] The network interface 63 may include a wireless network interface or a wired network interface, and is generally used to establish a communication connection between the computer device 6 and other electronic devices.

[0128] The embodiments of the present application can obtain target knowledge data from a knowledge document library, cover multiple knowledge documents, and provide a comprehensive data basis for subsequent processing. The knowledge documents are text clustered based on a hash clustering algorithm, which can quickly cluster similar documents into clusters, effectively deal with data redundancy problems, and reduce the complexity of subsequent processing. The semantic vector model is used for semantic deduplication of each knowledge document in the document cluster. Compared with traditional methods, the model can deeply understand the semantics of the text, accurately identify and deduplicate even if the text has large surface differences but similar semantics. This not only improves the deduplication accuracy, but also improves the efficiency when processing large-scale data, reduces the storage of redundant data, and effectively solves the problems of data redundancy and insufficient efficiency and accuracy of deduplication technology in the prior art.

[0129] The present application also provides another implementation, that is, to provide a computer readable storage medium, the computer readable storage medium stores computer readable instructions, the computer readable instructions can be executed by at least one processor, so that the at least one processor executes the steps of the knowledge data deduplication method as described above.

[0130] The embodiments of the present application can obtain target knowledge data from a knowledge document library, cover multiple knowledge documents, and provide a comprehensive data basis for subsequent processing. The knowledge documents are text clustered based on a hash clustering algorithm, which can quickly cluster similar documents into clusters, effectively deal with data redundancy problems, and reduce the complexity of subsequent processing. The semantic vector model is used for semantic deduplication of each knowledge document in the document cluster. Compared with traditional methods, the model can deeply understand the semantics of the text, accurately identify and deduplicate even if the text has large surface differences but similar semantics. This not only improves the deduplication accuracy, but also improves the efficiency when processing large-scale data, reduces the storage of redundant data, and effectively solves the problems of data redundancy and insufficient efficiency and accuracy of deduplication technology in the prior art.

[0131] Those skilled in the art can clearly understand the above-mentioned embodiment method can be realized by means of software and necessary general hardware platform, of course, can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of software product, and the computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a plurality of instructions to make a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) execute the method of each embodiment of the present application.

[0132] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all the embodiments, and the preferred embodiments of the present application are given in the drawings, but do not limit the patent scope of the present application. The present application can be realized in many different forms, and conversely, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing specific embodiments, or make equivalent replacements to some technical features. Any equivalent structure made by using the contents of the specification and drawings, directly or indirectly applied to other related technical fields, is also within the scope of the patent protection of the present application.

[0133] The non-company software tools or components appearing in the embodiments of the present application are only examples for introduction, not representing actual use.

Claims

1. A knowledge data deduplication method, characterized in that: The steps include: Acquire target knowledge data from a knowledge document library, wherein the target knowledge data includes a plurality of knowledge documents; Based on a preset hash clustering algorithm, text clustering is performed on the plurality of knowledge documents to obtain a plurality of document clusters; A preset semantic vector model is used to perform semantic deduplication on each knowledge document in the plurality of document clusters to obtain deduplicated documents corresponding to the target knowledge data.

2. The method according to claim 1, characterized in that The step of performing text clustering on the plurality of knowledge documents based on a preset hash clustering algorithm to obtain a plurality of document clusters specifically includes: Performing feature extraction on the plurality of knowledge documents to obtain a plurality of semantic features of each knowledge document; Get the weighted feature vector corresponding to each semantic feature; Based on a preset hash clustering algorithm, the weighted feature vector is binarized and dimensionally reduced to obtain a local sensitive hash value of each knowledge document; Calculate the similarity between the local sensitive hash values ​​of any two knowledge documents; Based on the similarity, the multiple knowledge documents are grouped to obtain multiple document clusters.

3. The method according to claim 2, characterized in that The step of obtaining the weighted feature vector corresponding to each semantic feature specifically includes: Obtaining a weight for each semantic feature; Obtaining a high-dimensional vector of each semantic feature to obtain a plurality of embedding vectors corresponding to each knowledge document, wherein the embedding vectors represent semantic information of the corresponding semantic feature; A weighted sum operation is performed on the multiple embedding vectors of each knowledge document and the weight to obtain a weighted feature vector of each knowledge document.

4. The method according to claim 2, characterized in that The step of performing binarization and dimensionality reduction on the weighted feature vector based on a preset hash clustering algorithm to obtain a local sensitive hash value of each knowledge document specifically includes: Based on a preset hash clustering algorithm, the dimension value of each dimension of the weighted feature vector is compared with a preset threshold; If the dimension value is greater than the threshold, the corresponding dimension value is changed to an activation value; If the dimension value is less than or equal to the threshold, the corresponding dimension value is changed to an inactivated value; Based on the activation value and the inactivation value, a local sensitive hash value of each knowledge document is generated.

5. The method according to claim 2, characterized in that The step of calculating the similarity between the local sensitive hash values ​​of any two knowledge documents specifically includes: Get the number of bits of each locality-sensitive hash value; Compare the binary bits of the local sensitive hash values ​​of any two knowledge documents bit by bit to determine the number of different bits; Based on the quantity, determining a Hamming distance between the local sensitive hash values ​​of the arbitrary two knowledge documents; Based on the number of bits and the Hamming distance, a similarity between the local sensitive hash values ​​of the arbitrary two knowledge documents is determined.

6. The method according to claim 1, characterized in that The step of using a preset semantic vector model to perform semantic deduplication on each knowledge document in the plurality of document clusters to obtain deduplication documents corresponding to the target knowledge data specifically includes: Using a preset semantic vector model, each knowledge document in the plurality of document clusters is converted into a word vector matrix to obtain a plurality of document word vector matrices in each document cluster; Calculate the cosine similarity between any two document word vector matrices in each document cluster; Based on the cosine similarity, semantic deduplication is performed on each knowledge document in the multiple document clusters to obtain deduplication documents corresponding to the target knowledge data.

7. The method according to claim 6, characterized in that The step of performing semantic deduplication on each knowledge document in the plurality of document clusters based on the cosine similarity to obtain deduplication documents corresponding to the target knowledge data specifically includes: Calculating a similarity score between each document in each document cluster and all documents in the cluster based on the cosine similarity; Sorting the similarity scores of each document in each group of document clusters to obtain a sorting result; Based on the ranking result, the documents remaining in each group of document clusters are determined to obtain the deduplicated documents corresponding to the target knowledge data.

8. A knowledge data deduplication device, characterized in that: include: An acquisition module, configured to acquire target knowledge data from a knowledge document library, wherein the target knowledge data includes a plurality of knowledge documents; A clustering module, configured to perform text clustering on the plurality of knowledge documents based on a preset hash clustering algorithm to obtain a plurality of document clusters; The deduplication module is used to use a preset semantic vector model to perform semantic deduplication on each knowledge document in the multiple document clusters to obtain deduplication documents corresponding to the target knowledge data.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the knowledge data deduplication method as claimed in any one of claims 1 to 7 when executing the computer-readable instructions.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the knowledge data deduplication method according to any one of claims 1 to 7.