Data processing method and device and electronic equipment

Through the weighted vector and signature vector methods of feature words, the problem of inefficient identification of repetitive problems in the after-sales maintenance knowledge base is solved, and efficient and accurate knowledge base construction and the optimization of intelligent customer service system are achieved.

CN120493915APending Publication Date: 2025-08-15LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510387750.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

When building an after-sales maintenance knowledge base, the existing technology faces inefficiency and information redundancy caused by duplicate problem data. The existing methods have high computational complexity and high resource consumption, so they cannot effectively distinguish between problems such as opposite semantics but similar vocabulary.

Method used

Through the method of weighted vectors, eigenvectors and signature vectors based on feature words, the correlation, hash value and position feature vectors of feature words are used to calculate the similarity of the sample and deduplicate it, reducing the calculation cost and improving processing speed.

Benefits of technology

It effectively eliminates duplicate problems, improves the speed and quality of knowledge base construction, reduces misjudgments caused by location information, and improves the answer accuracy and customer satisfaction of the intelligent customer service system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493915A_ABST
    Figure CN120493915A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device and electronic equipment, and the method comprises the steps: determining a weighted vector of each feature word in a first sample based on the total number of feature words included in the first sample, the occurrence frequency of each feature word, and the correlation degree between any two feature words; wherein if the association degree between any two feature words is greater than a first preset threshold value, the weighted vector of any one feature word in the any two feature words is far smaller than the weighted vector of the other feature word; determining a feature vector of each feature word in the first sample based on the hash value and the position feature vector of each feature word in the first sample; based on the weighted vector of each feature word in the first sample and the feature vector of each feature word, determining a signature vector of each feature word in the first sample; based on the signature vector of each feature word, determining a signature corresponding to the first sample; and based on the signature corresponding to each sample, determining the similarity between the samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a data processing method, device, and electronic device. Background Art

[0002] With the rise of generative language models, more and more organizations are looking to leverage these technologies to build robust knowledge bases, particularly in the after-sales service sector. The challenge in building a knowledge base for after-sales repairs is capturing diverse questions and extracting valuable information from existing cases. However, the large amount of duplicate question data reduces efficiency and increases information redundancy.

[0003] Related technologies use embeddings to determine the repetitiveness of questions. While this method can identify similar questions to a certain extent, calculating embeddings in large datasets consumes a significant amount of computing resources. In practice, this sometimes results in repeated calculations for similar questions, further increasing resource consumption. These factors necessitate a more efficient and cost-effective solution for deduplicating questions, enabling the rapid and accurate organization of knowledge bases. Summary of the Invention

[0004] The present disclosure provides a data processing method, device, and electronic device to at least solve the above technical problems existing in the prior art.

[0005] According to a first aspect of the present disclosure, a data processing method is provided, the method comprising:

[0006] Determining a weighted vector for each feature word in the first sample based on the total number of feature words included in the first sample, the number of occurrences of each feature word, and the degree of association between any two feature words; wherein, if the degree of association between any two feature words is greater than a first preset threshold, the weighted vector of any one of the two feature words is much smaller than the weighted vector of the other feature word;

[0007] Determining a feature vector for each feature word in the first sample based on the hash value and position feature vector of each feature word in the first sample;

[0008] Determine a signature vector for each feature word in the first sample based on the weighted vector of each feature word and the feature vector of each feature word in the first sample; determine a signature corresponding to the first sample based on the signature vector of each feature word;

[0009] Based on the signature corresponding to each sample, the similarity between samples is determined.

[0010] In the above solution, determining the weighted vector for each feature word in the first sample based on the total number of feature words included in the first sample, the number of occurrences of each feature word, and the correlation between any two feature words includes performing the following operations on the first feature word in the first sample:

[0011] Determining a weight of the first feature word in the first sample based on the number of times the first feature word appears in the first sample, the total number of feature words included in the first sample, the total number of samples in the sample set, and the number of samples in the sample set that include the first feature word;

[0012] Determining, based on the sample containing the first feature word, a correlation between the first feature word and other feature words;

[0013] determining a weighted vector of the first feature word in the first sample based on the association between the first feature word and other feature words and the weight of the first feature word in the first sample;

[0014] The weighted vector is used to represent the difference between the feature word and other feature words in the first sample.

[0015] In the above solution, determining the association between the first feature word and other feature words based on the sample containing the first feature word includes performing the following operations on the first feature word and the second feature word:

[0016] The degree of association between the first feature word and the second feature word is determined based on the total number of first-type samples that simultaneously include the first feature word and the second feature word, the number of times the first feature word appears in each first-type sample, the number of times the second feature word appears in each first-type sample, and the total number of second-type samples that include the first feature word or the second feature word; the degree of association between the first feature word and the second feature word refers to the degree to which the first feature word is used to represent the second feature word. The higher the degree of association between the first feature word and the second feature word, the higher the mutual substitutability of the first feature word and the second feature word.

[0017] In the above solution, determining the weighted vector of the first feature word in the first sample based on the association between the first feature word and other feature words and the weight of the first feature word in the first sample includes performing the following operations on the first feature word and the second feature word:

[0018] In response to the weight of the first feature word being less than the product of the second feature word and the association degree, determining that the weighted vector of the first feature word in the first sample is a second preset threshold;

[0019] In response to the weight of the first feature word being greater than the product of the second feature word and the association degree, a weighted vector of the first feature word is determined based on the weight of the first feature word, the weight of the second feature word, and the association degree between the first feature word and the second feature word.

[0020] In the above solution, determining the signature vector of each feature word in the first sample based on the weighted vector of each feature word in the first sample and the feature vector of each feature word includes:

[0021] The signature vector of each feature word is determined based on the product of the weighted vector and the feature vector corresponding to each feature word in the first sample.

[0022] In the above solution, determining the signature corresponding to the first sample based on the signature vector of each feature word includes:

[0023] The sum of the signature vectors of all the feature words is determined as the signature corresponding to the first sample.

[0024] In the above solution, determining the similarity between samples based on the signature corresponding to each sample includes:

[0025] Binarize the signature corresponding to each sample to obtain a binary signature;

[0026] If the degree of overlap of the binary signatures of the two samples is greater than a third preset threshold, the two samples are determined to be similar, and only one of the samples is retained.

[0027] In the above scheme, if the overlap of the binary signatures of the two samples is greater than the third preset threshold, the two samples are determined to be similar, and only one sample is retained, including performing the following operations on the binary signature of the first sample and the binary signature of the second sample:

[0028] Splitting the binarized signature of the first sample into at least two binarized sub-signatures, each of which has the same number of elements;

[0029] Split the binarized signature of the second sample into at least two binarized sub-signatures, each of which has the same number of elements; the number of binarized sub-signatures of the first sample is the same as the number of binarized sub-signatures of the second sample;

[0030] Comparing the binarized sub-signature of the first sample and the binarized sub-signature of the second sample at the same position in parallel;

[0031] If the similarity of any group of binarized sub-signatures meets a preset condition, the first sample is determined to be similar to the second sample, and only one of the first sample and the second sample is retained.

[0032] According to a second aspect of the present disclosure, there is provided a data processing device, the device comprising:

[0033] a weighted vector determining unit, configured to determine a weighted vector for each feature word in the first sample based on the total number of feature words included in the first sample, the number of occurrences of each feature word, and the degree of association between any two feature words; wherein, if the degree of association between any two feature words is greater than a first preset threshold, the weighted vector of any one of the two feature words is much smaller than the weighted vector of the other feature word;

[0034] a feature vector determining unit, configured to determine a feature vector for each feature word in the first sample based on a hash value and a position feature vector of each feature value in the first sample;

[0035] a signature vector determining unit, configured to determine a signature vector for each feature word in the first sample based on a weighted vector of each feature word and a feature vector of each feature word in the first sample; and determine a signature corresponding to the first sample based on the signature vector of each feature word;

[0036] The comparison unit is used to determine the similarity between samples based on the signature corresponding to each sample.

[0037] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0038] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the present disclosure.

[0039] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings, in which several embodiments of the present disclosure are shown by way of example and not limitation, wherein:

[0041] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.

[0042] Figure 1 A first optional flow chart of the data processing method provided by an embodiment of the present disclosure is shown;

[0043] Figure 2 A second optional flow chart of the data processing method provided by the embodiment of the present disclosure is shown;

[0044] Figure 3 A third optional flow chart of the data processing method provided by the embodiment of the present disclosure is shown;

[0045] Figure 4 A data flow chart showing a data processing method provided by an embodiment of the present disclosure is shown;

[0046] Figure 5 An optional structural diagram of a data processing device provided by an embodiment of the present disclosure is shown;

[0047] Figure 6 A schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0048] To make the purposes, features, and advantages of the present disclosure more apparent and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative work shall fall within the scope of protection of the present disclosure.

[0049] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0050] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by those skilled in the art in the art of this disclosure. The terms used in this disclosure are only for the purpose of describing the embodiments of this disclosure and are not intended to limit this disclosure.

[0051] It should be understood that in the various embodiments of the present disclosure, the size of the serial number of each implementation process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure.

[0052] The following solutions are used in related technologies to deduplicate data:

[0053] 1) Based on Embedding and vector database.

[0054] This type of method uses the similarity between text vectors to calculate the degree of text similarity. Specifically, it achieves text deduplication by generating text vectors and searching for similar vectors in a vector database.

[0055] However, this method has high computational complexity. Generating text vectors and calculating the similarity between vectors involves high computational costs, making it inefficient, especially when processing large datasets. Even with clustering optimization measures, high spatiotemporal complexity still exists, and clustering methods are not applicable when there is significant data overlap or large dimensions. In addition, this method is highly model-dependent. The effectiveness of text vectorization is highly dependent on the model used and its preprocessing steps. Different models may show significant differences in performance on different types of text.

[0056] 2) Text similarity method based on minhash.

[0057] This method is computationally expensive and requires high storage costs: multiple hash functions are required, and the signature vectors are long (often exceeding hundreds of dimensions). It only works with aggregate features, cannot process feature weights, and is insensitive to weighted information such as word frequency and position. It is not suitable for deduplication of after-sales repair issues.

[0058] 3) Text similarity algorithm based on Simhash.

[0059] The Term Frequency-Inverse Document Frequency (TF-IDF) algorithm is commonly used to calculate text similarity. This algorithm uses statistics on the frequency of occurrence of feature words in a document and their global distribution to assign weights, reducing the contribution of common words. However, this method fails to represent keyword location information and is significantly affected by the weighting algorithm, making it unsuitable for deduplicating after-sales repair questions. This can result in two questions with completely opposite meanings being identified as duplicates.

[0060] Specifically, in electronic equipment after-sales repair data, TF-IDF assigns low weights to frequent but less discriminative feature words, such as "yes" and "no." These feature words often co-occur with other feature words, leading to ambiguous text signatures. In other words, sentences with opposite meanings may be considered very similar in traditional SimHash. For example, "My graphics card has been replaced" and "My graphics card has not been replaced," although having opposite meanings, are considered similar in traditional SimHash algorithms.

[0061] Furthermore, traditional text deduplication methods typically don't consider location information. However, in electronic device repair data, location information is crucial for correctly understanding and distinguishing different fault descriptions. For example, "The keyboard is locked, so the screen is unresponsive" and "The screen is locked, and the keyboard can't unlock it" describe two different fault conditions, but if location information is ignored, it is difficult to distinguish.

[0062] In response to the problems existing in the related technologies, the embodiments of the present disclosure provide a data processing method to at least solve some or all of the above technical problems.

[0063] It should be noted that the algorithm involved in the embodiment of the present disclosure can be applied to the electronic equipment after-sales maintenance knowledge base, and can also be applied to other scenarios where text deduplication is required.

[0064] Figure 1 A first optional flow chart of the data processing method provided by an embodiment of the present disclosure is shown, and will be explained according to each part.

[0065] Step S101 : determining a weighted vector for each feature word in the first sample based on the total number of feature words included in the first sample, the number of occurrences of each feature word, and the degree of association between any two feature words.

[0066] In some embodiments, a carrier that implements the data processing method (hereinafter referred to as the carrier) performs word segmentation on the first sample to determine the total number of feature words included in the first sample and the number of times each feature word appears; based on the total number of feature words included in the first sample, the number of times each feature word appears and the correlation between any two feature words, a weighted vector for each feature word in the first sample is determined.

[0067] In a specific implementation, when determining the weighted vector of the third feature word in the first sample, the association degree between the two feature words used includes the association degree between the third feature word and other feature words in the first sample.

[0068] Among them, if the correlation between any two feature words is greater than a first preset threshold, the weighted vector of any two feature words is much smaller than the weighted vector of the other feature word; the first preset threshold can be set according to actual needs or experimental results. In the embodiment of the present disclosure, the weighted vector of any feature word is made much smaller than the weighted vector of another feature word in order to reduce the weight of the two feature words with high correlation and reduce the difference between samples; the feature words with low correlation (key feature words) are made to have a higher weight, and the distinguishing feature words between the two samples are also made to have a higher weight, so that more attention is focused on the key feature words that cause sample differences, thereby improving the discrimination between samples with similar vocabulary but opposite semantics.

[0069] The carrier can be a computer program, electronic circuit, database, mobile application, electronic device, cloud computing platform, distributed system, artificial intelligence framework, mathematical model, automation tool and microcontroller, etc., which can implement software or hardware of algorithm and method process.

[0070] Step S102: determining a feature vector for each feature word in the first sample based on the hash value and position feature vector of each feature word in the first sample.

[0071] In some embodiments, the carrier determines a feature vector for each feature word based on the hash value and the position feature vector of the feature word. The position feature vector can be used to distinguish samples with the same feature word but different semantics.

[0072] In some embodiments, the position feature vector is used to represent the position of each feature word in the first sample.

[0073] Step S103 : determining a signature vector of each feature word in the first sample based on the weighted vector of each feature word in the first sample and the feature vector of each feature word.

[0074] In some embodiments, the carrier may determine the signature vector of the feature word directly based on the product of the weight vector of the feature word and the feature word.

[0075] In some embodiments, the carrier determines the signature vector of each feature word in the first sample based on the hash value, position feature vector and weight vector of the feature word, so that the signature vector of the feature word can reflect the position information, occurrence frequency and correlation with other feature words of the feature word, highlighting the attention of key feature words while reducing the attention of non-key words.

[0076] Step S104: Determine the signature corresponding to the first sample based on the signature vector of each feature word.

[0077] In some embodiments, the first sample includes multiple feature words, and a signature corresponding to the first sample is determined based on a signature vector for each feature word in the first sample. It should be noted that the signature corresponding to the first sample is a vector, and each element in the vector is determined based on the signature vector for each feature word. The position of the element in the vector is independent of the position of the feature word in the first sample.

[0078] Step S105: Determine the similarity between samples based on the signature corresponding to each sample.

[0079] In some embodiments, the carrier can determine the similarity between samples by comparing the signature vectors corresponding to each sample. If the overlap of elements at the same position in the signature vectors corresponding to two samples is greater than a third preset threshold, the two samples are similar, and only one sample is retained in the knowledge base. If the overlap of elements at the same position in the signature vectors corresponding to two samples is less than the third preset threshold, the two samples are dissimilar and both are stored in the knowledge base.

[0080] In this way, through the data processing method described in the embodiment of the present disclosure, it is possible to effectively deduplicate questions during the information collection stage. Not only does it reduce the computing cost, but it also significantly improves the processing speed, and can organize accurate knowledge base content more quickly. More importantly, the method provided by the embodiment of the present disclosure can effectively reduce repeated judgments caused by factors such as location information, thereby improving the quality and credibility of knowledge base construction. The high efficiency of early deduplication provides support for subsequent artificial intelligence empowerment, enabling the intelligent customer service system to better answer customer questions and improve customer satisfaction. In addition, by optimizing problem identification in the process of refining the knowledge base, it is easier to discover potential service improvement points, thereby promoting the improvement of overall business value.

[0081] Figure 2 A second optional flow chart of the data processing method provided by the embodiment of the present disclosure is shown, which will be explained according to each part.

[0082] Step S201: Determine the correlation between any two feature words in the sample set.

[0083] In some embodiments, the following description is made by taking determining the correlation between a first feature word and a second feature word as an example.

[0084] In some embodiments, the sample set includes the following types of samples: a first type of sample that includes both the first feature word and the second feature word, the first sample is a first type of sample; a sample that includes the first feature word but does not include the second feature word, and a sample that includes the second feature word but does not include the first feature word, this type of sample is a second type of sample.

[0085] In some embodiments, the carrier determines the total number of first type samples and the total number of second type samples in the sample set; determines the number of times x that the first feature word x appears in each first type sample k , the number of times the second feature word k appears in each first type sample y k ; Wherein, k represents the kth sample in the first type of sample.

[0086] In some embodiments, the carrier determines the ratio of the first type of samples to the second type of samples; the carrier is based on the number of times the first feature word appears x k The number of times the second feature word appears y k Determine the correlation between the first feature word and the second feature word; determine the correlation between the first feature word and the second feature word based on the proportion of the first type of samples in the first type of samples and the second type of samples, and the correlation between the first feature word and the second feature word.

[0087] In some embodiments, the correlation between the first feature word and the second feature word refers to the degree to which the first feature word is used to represent the second feature word. The higher the correlation between the first feature word and the second feature word, the higher the mutual substitutability of the first feature word and the second feature word. If the first feature word and the second feature word appear simultaneously in each sample in the sample set, the correlation between the first feature word and the second feature word is the highest, with a value of 1. The mutual substitutability between the two is the highest, and the first feature word can be used to represent the second feature word.

[0088] Step S202 : determining a weighted vector for each feature word in the first sample based on the total number of feature words included in the first sample, the number of occurrences of each feature word, and the degree of association between any two feature words.

[0089] In some embodiments, the carrier determines the weight of each feature word in the first sample based on the number of times each feature word appears in the first sample, the total number of feature words contained in the first sample, and the total number of samples in the sample set; and determines the weighted vector of each of the any two feature words in the first sample based on the association between the any two feature words and the weight of each of the any two feature words in the first sample.

[0090] In some embodiments, the method of determining the weighted vector of the first feature word in the first sample is taken as an example for description; the method of determining the weighted vectors of other feature words in the first sample is similar.

[0091] In some embodiments, the carrier determines the weight of the first feature word in the first sample based on the number of times the first feature word appears in the first sample, the total number of feature words included in the first sample, the number of samples in the sample set that include the first feature word, and the total number of samples in the sample set.

[0092] Specifically, the carrier determines the frequency of the first feature word appearing in the first sample based on the number of times the first feature word appears in the first sample and the total number of feature words included in the first sample; the carrier determines the probability of the sample including the first feature word based on the number of samples in the sample set including the first feature word and the total number of samples in the sample set; and determines the weight of the first feature word in the first sample based on the frequency of the first feature word appearing in the first sample and the probability of the sample including the first feature word.

[0093] In some embodiments, the carrier may further process the weight of the first feature word in the first sample to determine a weighted vector for the first feature word in the first sample. The weighted vector is used to characterize the difference between the feature word and other feature words in the first sample. If the weighted vector of the feature word is a preset threshold, it indicates that the difference between the feature word and other feature words in the first sample is very small; the larger the weighted vector of the feature word, the greater the difference between the feature word and other feature words in the first sample.

[0094] In some embodiments, in response to the weight of the first feature word being less than the product of the second feature word and the degree of association, the carrier determines that the weight vector of the first feature word in the first sample is a second preset threshold; and in response to the weight of the first feature word being greater than the product of the second feature word and the degree of association, determines the weight vector of the first feature word based on the weight of the first feature word, the weight of the second feature word, and the degree of association between the first feature word and the second feature word. The second preset threshold may be 0.

[0095] During specific implementation, the carrier may determine a weighted vector of the first feature word based on each feature word and other feature words in the first sample.

[0096] Specifically, the carrier determines the weights of all feature words in the first sample, as well as the association between any two feature words; determines the weight of the first feature word based on the weight of the first feature word, the association between the first feature word and each other feature word, and the weights of the other feature words; in the determination process, the carrier determines the weighted vectors of n-1 first feature words, and takes the minimum value as the weighted vector of the first feature word; n is the number of feature words included in the first sample.

[0097] Step S203 : determining a feature vector for each feature word in the first sample based on the hash value and position feature vector of each feature word in the first sample.

[0098] In some embodiments, the carrier determines a position feature vector for each feature word based on the position of each feature word in the first sample. The dimension of the position feature vector is the same as the number of feature words included in the first sample. In the position feature vector corresponding to each feature word, the element corresponding to the position of the feature word in the first sample is 1, and the other elements are 0.

[0099] In some embodiments, the carrier determines a hash value corresponding to each feature word in the first sample; and determines a feature vector of the first feature word based on the hash value corresponding to each feature word and the corresponding position feature vector.

[0100] During specific implementation, the carrier may directly determine the signature vector of the feature word based on the product of the hash value and the position feature vector.

[0101] Alternatively, in a specific implementation, the carrier performs dimensionality reduction processing on the position feature vector, and determines a feature vector based on the position feature vector after dimensionality reduction of each feature word, the position feature vector, and the hash value.

[0102] Step S204: Determine a signature vector for each feature word in the first sample based on the weighted vector of each feature word and the feature vector of each feature word in the first sample; and determine a signature corresponding to the first sample based on the signature vector of each feature word.

[0103] In some embodiments, the carrier may directly determine the signature vector of the feature word based on the product of the weighted vector of the feature word and the feature word. The signature vectors of each feature word in the first sample are accumulated to determine the signature corresponding to the first sample.

[0104] Step S205: Determine the similarity between samples based on the signature corresponding to each sample.

[0105] In some embodiments, the carrier may directly compare the signatures of two samples and determine the similarity between the two samples based on the numerical value.

[0106] In other embodiments, the carrier may perform binarization processing on the signature corresponding to each sample to obtain a binarized signature; compare the binarized signatures of two samples, and determine the similarity between the samples based on the degree of overlap.

[0107] In some further embodiments, the carrier may further split the binary signature of each sample to obtain at least two binary sub-signatures, each of which has the same number of elements. The binary sub-signatures of the two samples are compared to determine the similarity between the samples.

[0108] In a specific implementation, taking the first sample and the second sample as an example, the carrier can split the binary signature of the first sample into at least two binary sub-signatures, each of which has the same number of elements; split the binary signature of the second sample into at least two binary sub-signatures, each of which has the same number of elements; the number of binary sub-signatures of the first sample is the same as the number of binary sub-signatures of the second sample; the binary sub-signatures of the first sample and the binary sub-signatures of the second sample at the same position are compared in parallel; if the similarity of any group of binary sub-signatures meets a preset condition, the first sample is determined to be similar to the second sample, and only one of the first sample or the second sample is retained. The preset condition includes a numerical overlap greater than a fourth threshold; the fourth threshold can be set according to actual needs and experimental results.

[0109] The binary signature of a sample includes many elements. If they are compared one by one, it will take a long time. Therefore, in the embodiment of the present disclosure, the binary signature of the sample is segmented and divided into several binary sub-signatures with the same number of elements. During comparison, multiple binary sub-signatures in two samples can be compared in parallel, thereby improving the efficiency of the comparison process. Moreover, when the similarity of the binary sub-signatures of any two samples meets the preset conditions, the two samples are determined to be similar. This is because the signature vector corresponding to the feature word is dispersed into each element of the signature corresponding to the first sample, and then corresponds to each element of the binary sub-signature. In other words, each element in the binary sub-signature is determined by all the feature words in the first sample and the position information of all the feature words. If the similarity of the binary sub-signatures of any two samples meets the preset conditions, it means that the similarity of other binary sub-signatures also meets the preset conditions, and there is no need to continue comparison.

[0110] Taking the comparison of sample A and sample B as an example, the binary signature of sample A and the binary signature of sample B are both vectors of dimension D; directly comparing the elements in each binary signature takes a long time; sample A is divided into sample a1, sample a2, sample a3 and sample a4; similarly, sample B is divided into sample b1, sample b2, sample b3 and sample b4; the dimension of each sample after division is D / 4; sample a1 and sample b1, sample a2 and sample b2, sample a3 and sample b3, and sample a4 and sample b4 can be compared in parallel. If any set of similarities meets the preset conditions, sample A and sample B are confirmed to be similar samples.

[0111] In this way, through the data processing method described in the embodiment of the present disclosure, it is possible to effectively deduplicate questions during the information collection stage. Not only does it reduce the computing cost, but it also significantly improves the processing speed, and can organize accurate knowledge base content more quickly. More importantly, the method provided by the embodiment of the present disclosure can effectively reduce repeated judgments caused by factors such as location information, thereby improving the quality and credibility of knowledge base construction. The high efficiency of early deduplication provides support for subsequent artificial intelligence empowerment, enabling the intelligent customer service system to better answer customer questions and improve customer satisfaction. In addition, by optimizing problem identification in the process of refining the knowledge base, it is easier to discover potential service improvement points, thereby promoting the improvement of overall business value.

[0112] Figure 3 A third optional flow chart of the data processing method provided in an embodiment of the present disclosure is shown; it will be explained according to each step.

[0113] Step S301: determining the association between the first feature word and the second feature word.

[0114] In some embodiments, the vector performs word segmentation processing on samples in the sample set, and deletes symbols and useless words.

[0115] In some embodiments, the first feature word and the second feature word represent any two feature words included in any sample in the sample set. The sample set includes the following types of samples: the first type of samples including both the first feature word and the second feature word, the number of which is m 11 , the first sample is the first type of sample; the number of samples containing the first feature word but not the second feature word is m 10 , and the number of samples containing the second feature word but not the first feature word is m 01 , this type of sample is the second type of sample.

[0116] In some embodiments, the vector determines a ratio m of the first type of sample to the second type of sample. 11 / (m 11 +m 10 +m 01 ); Based on the number of times the first feature word appears in the k-th sample x k , and the number of occurrences y of the second feature word in the kth sample k Determining the correlation between the first feature word and the second feature word specifically includes:

[0117]

[0118] Furthermore, the carrier determines the correlation J(x, y) between the first feature word x and the second feature word y based on the ratio of the first type of samples in the first type of samples and the second type of samples, and the correlation between the first feature word and the second feature word, including:

[0119]

[0120] Step S302: Determine the weight of the first feature word in the sample set.

[0121] In some embodiments, the carrier may normalize the feature word weights, including the number of times h that the first feature word x appears in the first sample d. dx , the total number of feature words contained in the first sample H d , the total number of samples in the sample set N, the number of samples in the sample set where the first feature word x appears n x , determine the weight of each feature word in the first sample, specifically including:

[0122]

[0123] Step S303 : determining a weighted vector of the first feature word in the first sample based on the association between the first feature word and other feature words and the weight of each feature word in the first sample.

[0124] In some embodiments, the carrier responds that the weight of the first feature word x is less than the product of the second feature word y and the association J(x, y), and then determines the weight vector α of the first feature word in the first sample. ’ dx is the second preset threshold; in response to the weight of the first feature word being greater than the product of the second feature word and the association degree, the weight α of the first feature word is dx , the weight of the second feature word α dy , and the correlation J(x,y) between the first feature word and the second feature word, determine the weighted vector α of the first feature word ’ dx The second preset threshold may be 0. Specifically including:

[0125]

[0126] The significance of the above formula is that when the correlation between feature words x and y is high, if feature words x and y appear simultaneously in the same sample, the weight of one feature word can be reduced. In particular, when the correlation between feature words x and y is 1, either feature word can largely represent the sample information characteristics, thus focusing attention on the feature words that cause sample differences.

[0127] Step S304: Determine the hash value of each feature word in the first sample.

[0128] In some embodiments, the carrier may calculate a hash value of each characteristic word in the first sample based on the murmurHash function to obtain the signature sig0.

[0129] Step S305: Determine the position feature vector of each feature word in the first sample.

[0130] In some embodiments, the carrier determines a position feature vector for each feature word based on the position of each feature word in the first sample. The dimension of the position feature vector is the same as the number of feature words included in the first sample. In the position feature vector corresponding to each feature word, the element corresponding to the position of the feature word in the first sample is 1, and the other elements are 0.

[0131] In some optional embodiments, the carrier may perform dimensionality reduction on the position feature vector corresponding to each feature word to obtain a reduced-dimensional position feature vector sig1; the reduced-dimensional position feature vector sig1 is a binary signature.

[0132] Step S306 : determining a feature vector for each feature word in the first sample based on the hash value and position feature vector of each feature word in the first sample.

[0133] In some embodiments, the carrier performs weighted summation based on the hash value sig0 of each feature word obtained in step S304, the position feature vector u of each feature word obtained in step S305, and the position feature vector after dimensionality reduction to obtain the feature vector of each feature word:

[0134] sig=u*sig0+(1-u)*sig1

[0135] In some embodiments, the vector may set elements with a value of 0 to -1 before performing the operation.

[0136] Step S307 : determining a signature vector for each feature word in the first sample based on the weighted vector of each feature word in the first sample and the feature vector of each feature word.

[0137] In some embodiments, the carrier performs weighted processing on the feature vector of each feature word based on each feature word weighted vector calculated in step S303 to obtain each signature vector.

[0138] In some embodiments, the carrier may determine the signature vector of the feature word directly based on the product of the weight vector of the feature word and the feature word.

[0139] Step S308: Determine the signature corresponding to the first sample based on the signature vector of each feature word.

[0140] In some embodiments, the carrier accumulates the signature vector of each feature word in the first sample to determine the signature corresponding to the first sample.

[0141] Step S309: Determine the similarity between samples based on the signature corresponding to each sample.

[0142] In some embodiments, the carrier performs binarization processing on the signature of each sample, setting elements greater than 0 in the signature of the sample to 1 and elements less than 0 to 0.

[0143] In some embodiments, the carrier performs an XOR operation on the signatures of different samples, comparing their signature values bit by bit. If the value of a bit is different, it is recorded as 1, otherwise it is recorded as 0. The number of 1s obtained is the Hamming distance, and the ratio of the number of 1s to the total number of elements is the overlap. A larger Hamming distance indicates a lower similarity between the two texts, while a smaller Hamming distance indicates a higher similarity.

[0144] In other embodiments, the carrier may divide the signature of each sample into m buckets, and take m as 4 as an example for illustration:

[0145] Each sample's signature is divided into four parts, with each part serving as an index into a bucket. When storing, the signature corresponding to each part is placed into the corresponding bucket. A signature for each sample is generated and added simultaneously to four different buckets to facilitate parallel processing. When inserting a new sample's signature, the signatures in the same bucket are checked, the Hamming distance is calculated, and the presence of duplicate text is determined.

[0146] In this way, through the data processing method provided by the embodiment of the present disclosure, a similarity weighting algorithm is introduced, and Jaccard similarity is used to optimize the weight distribution of feature words. When the correlation between two feature words is high, the weights of the two feature words will be automatically adjusted, and the weight of one of the feature words will be reduced, so that more attention will be focused on the key feature words that cause text differences. In this way, the algorithm can more accurately distinguish texts with opposite meanings or obvious differences, and improve the accuracy of text deduplication; generate a vector of position information, and weightedly fuse it with the hash value of the feature word to ensure that the signature of the final generated sample can effectively reflect the position information of each feature word in the original text. Not only is the efficiency of text deduplication improved, but the accuracy of the algorithm is also enhanced, so that duplicate text can be more accurately identified and excluded when processing after-sales repair data of electronic equipment.

[0147] Figure 4 A data flow chart of the data processing method provided by an embodiment of the present disclosure is shown.

[0148] like Figure 4 As shown, the text is segmented to remove punctuation marks and useless words in the text, and the hash value of each feature word is determined based on the hash algorithm; based on the position of each feature word in the text, the position feature information is determined, and the position feature vector corresponding to each feature word is obtained based on the parity check; the hash value and position feature vector corresponding to each feature word are weighted to obtain the feature vector of each feature word, and the result is multiplied by the weighted vector corresponding to the feature word to obtain the signature vector corresponding to each feature word.

[0149] The signature vectors corresponding to all feature words are merged to obtain the signature corresponding to the text. Before comparison, the signature corresponding to the text is binarized, and elements greater than 0 are set to 1 and elements less than 0 are set to 0. The similarity of the text is then determined by comparing the binarized signatures of the text.

[0150] like Figure 4As shown, the signature vectors corresponding to all feature words are merged to obtain the signature corresponding to the text [14, -321, 10, 5, 124, -523], which is binarized to obtain [101110].

[0151] Figure 5 An optional structural diagram of a data processing device provided by an embodiment of the present disclosure is shown, and will be explained according to each part.

[0152] In some embodiments, the data processing device 500 includes a weight vector determining unit 501 , a feature vector determining unit 502 , a signature vector determining unit 503 and a comparing unit 504 .

[0153] The weighted vector determining unit 501 is configured to determine a weighted vector for each feature word in the first sample based on the total number of feature words included in the first sample, the number of occurrences of each feature word, and the degree of association between any two feature words; wherein, if the degree of association between any two feature words is greater than a first preset threshold, the weighted vector of any one of the two feature words is much smaller than the weighted vector of the other feature word;

[0154] The feature vector determining unit 502 is configured to determine a feature vector for each feature word in the first sample based on the hash value and position feature vector of each feature value in the first sample;

[0155] The signature vector determining unit 503 is configured to determine a signature vector for each feature word in the first sample based on the weighted vector of each feature word and the feature vector of each feature word in the first sample; and determine a signature corresponding to the first sample based on the signature vector of each feature word;

[0156] The comparison unit 504 is configured to determine the similarity between samples based on the signature corresponding to each sample.

[0157] The weighted vector determining unit 501 is specifically configured to perform the following operations on the first feature word in the first sample:

[0158] Determining a weight of the first feature word in the first sample based on the number of times the first feature word appears in the first sample, the total number of feature words included in the first sample, the total number of samples in the sample set, and the number of samples in the sample set that include the first feature word;

[0159] Determining, based on the sample containing the first feature word, a degree of association between the first feature word and other feature words;

[0160] determining a weighted vector of the first feature word in the first sample based on the association between the first feature word and other feature words and the weight of the first feature word in the first sample;

[0161] The weighted vector is used to represent the difference between the feature word and other feature words in the first sample.

[0162] The weighted vector determining unit 501 is specifically configured to perform the following operations on the first feature word and the second feature word:

[0163] The correlation between the first feature word and the second feature word is determined based on the total number of first-type samples that simultaneously include the first feature word and the second feature word, the number of times the first feature word appears in each first-type sample, the number of times the second feature word appears in each first-type sample, and the total number of second-type samples that include the first feature word or the second feature word. The correlation between the first feature word and the second feature word refers to the degree to which the first feature word is used to represent the second feature word. The higher the correlation between the first feature word and the second feature word, the higher the mutual substitutability of the first feature word and the second feature word.

[0164] The weighted vector determining unit 501 is specifically configured to perform the following operations on the first feature word and the second feature word:

[0165] In response to the weight of the first feature word being less than the product of the second feature word and the association degree, determining that the weighted vector of the first feature word in the first sample is a second preset threshold;

[0166] In response to the weight of the first feature word being greater than the product of the second feature word and the association degree, a weighted vector of the first feature word is determined based on the weight of the first feature word, the weight of the second feature word, and the association degree between the first feature word and the second feature word.

[0167] The signature vector determining unit 503 is specifically configured to determine the signature vector of each feature word based on the product of the weighted vector and the feature vector corresponding to each feature word in the first sample.

[0168] The signature vector determining unit 503 is specifically configured to determine the sum of the signature vectors of all feature words as the signature corresponding to the first sample.

[0169] The comparison unit 504 is specifically used to perform binarization processing on the signature corresponding to each sample to obtain a binarized signature;

[0170] If the degree of overlap of the binary signatures of the two samples is greater than a third preset threshold, the two samples are determined to be similar, and only one of the samples is retained.

[0171] The comparison unit 504 is specifically configured to perform the following operations on the binary signature of the first sample and the binary signature of the second sample:

[0172] Splitting the binarized signature of the first sample into at least two binarized sub-signatures, each of which has the same number of elements;

[0173] Split the binarized signature of the second sample into at least two binarized sub-signatures, each of which has the same number of elements; the number of binarized sub-signatures of the first sample is the same as the number of binarized sub-signatures of the second sample;

[0174] Comparing the binarized sub-signature of the first sample and the binarized sub-signature of the second sample at the same position in parallel;

[0175] If the similarity of any group of binarized sub-signatures meets a preset condition, the first sample is determined to be similar to the second sample, and only one of the first sample and the second sample is retained.

[0176] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.

[0177] Figure 6 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0178] like Figure 6 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0179] Multiple components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0180] The computing unit 801 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as the data processing method. For example, in some embodiments, the data processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the data processing method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the data processing method in any other appropriate manner (e.g., by means of firmware).

[0181] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0182] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0183] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0184] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0185] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0186] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0187] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0188] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.

[0189] The above description is merely a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.

Claims

1. A data processing method, comprising: Determining a weighted vector for each feature word in the first sample based on the total number of feature words included in the first sample, the number of occurrences of each feature word, and the degree of association between any two feature words; wherein, if the degree of association between any two feature words is greater than a first preset threshold, the weighted vector of any one of the two feature words is much smaller than the weighted vector of the other feature word; Determining a feature vector for each feature word in the first sample based on the hash value and position feature vector of each feature word in the first sample; Determining a signature vector for each feature word in the first sample based on a weighted vector for each feature word in the first sample and a feature vector for each feature word; Determine the signature corresponding to the first sample based on the signature vector of each feature word; Based on the signature corresponding to each sample, the similarity between samples is determined.

2. The method according to claim 1, wherein determining a weighted vector for each feature word in the first sample based on the total number of feature words included in the first sample, the number of occurrences of each feature word, and the correlation between any two feature words comprises performing the following operations on the first feature word in the first sample: Determining a weight of the first feature word in the first sample based on the number of times the first feature word appears in the first sample, the total number of feature words included in the first sample, the total number of samples in the sample set, and the number of samples in the sample set that include the first feature word; Determining, based on the sample containing the first feature word, a correlation between the first feature word and other feature words; determining a weighted vector of the first feature word in the first sample based on the association between the first feature word and other feature words and the weight of the first feature word in the first sample; The weighted vector is used to represent the difference between the feature word and other feature words in the first sample.

3. The method according to claim 2, wherein determining the association between the first feature word and other feature words based on the sample containing the first feature word comprises performing the following operations on the first feature word and the second feature word: The degree of association between the first feature word and the second feature word is determined based on the total number of first-type samples that simultaneously include the first feature word and the second feature word, the number of times the first feature word appears in each first-type sample, the number of times the second feature word appears in each first-type sample, and the total number of second-type samples that include the first feature word or the second feature word; the degree of association between the first feature word and the second feature word refers to the degree to which the first feature word is used to represent the second feature word. The higher the degree of association between the first feature word and the second feature word, the higher the mutual substitutability of the first feature word and the second feature word.

4. The method according to claim 2, wherein determining the weighted vector of the first feature word in the first sample based on the association between the first feature word and other feature words and the weight of the first feature word in the first sample comprises performing the following operations on the first feature word and the second feature word: In response to the weight of the first feature word being less than the product of the second feature word and the association degree, determining that the weighted vector of the first feature word in the first sample is a second preset threshold; In response to the weight of the first feature word being greater than the product of the second feature word and the association degree, a weighted vector of the first feature word is determined based on the weight of the first feature word, the weight of the second feature word, and the association degree between the first feature word and the second feature word.

5. The method according to claim 1, wherein determining the signature vector of each feature word in the first sample based on the weighted vector of each feature word and the feature vector of each feature word in the first sample comprises: The signature vector of each feature word is determined based on the product of the weighted vector and the feature vector corresponding to each feature word in the first sample.

6. The method according to claim 1, wherein determining the signature corresponding to the first sample based on the signature vector of each feature word comprises: The sum of the signature vectors of all the feature words is determined as the signature corresponding to the first sample.

7. The method according to claim 1, wherein determining the similarity between samples based on the signature corresponding to each sample comprises: Binarize the signature corresponding to each sample to obtain a binary signature; If the degree of overlap of the binary signatures of the two samples is greater than a third preset threshold, the two samples are determined to be similar, and only one of the samples is retained.

8. The method according to claim 7, wherein if the overlap between the binarized signatures of two samples is greater than a third preset threshold, the two samples are determined to be similar, and only one of the samples is retained, comprising performing the following operations on the binarized signature of the first sample and the binarized signature of the second sample: Splitting the binarized signature of the first sample into at least two binarized sub-signatures, each of which has the same number of elements; Split the binarized signature of the second sample into at least two binarized sub-signatures, each of which has the same number of elements; the number of binarized sub-signatures of the first sample is the same as the number of binarized sub-signatures of the second sample; Comparing the binarized sub-signature of the first sample and the binarized sub-signature of the second sample at the same position in parallel; If the similarity of any group of binarized sub-signatures meets a preset condition, the first sample is determined to be similar to the second sample, and only one of the first sample and the second sample is retained.

9. A data processing device, comprising: a weighted vector determining unit, configured to determine a weighted vector for each feature word in the first sample based on the total number of feature words included in the first sample, the number of occurrences of each feature word, and the degree of association between any two feature words; wherein, if the degree of association between any two feature words is greater than a first preset threshold, the weighted vector of any one of the two feature words is much smaller than the weighted vector of the other feature word; a feature vector determining unit, configured to determine a feature vector for each feature word in the first sample based on a hash value and a position feature vector of each feature value in the first sample; a signature vector determining unit, configured to determine a signature vector for each feature word in the first sample based on a weighted vector of each feature word and a feature vector of each feature word in the first sample; and determine a signature corresponding to the first sample based on the signature vector of each feature word; The comparison unit is used to determine the similarity between samples based on the signature corresponding to each sample.

10. An electronic device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 8.