Data processing method and device, electronic equipment and storage medium

By setting similarity thresholds and blocking data, the problem of unclear class boundaries in data clustering is solved, and clear and accurate clustering results are generated to avoid adhesions between classes.

CN120234459APending Publication Date: 2025-07-01SF TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311873949.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-30
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing data clustering methods lead to unclear boundaries between classes and sticky, making it difficult to generate accurate clustering results.

Method used

By setting the first and second thresholds, it is determined that the data in the data set whose similarity to the target data is greater than the first threshold and which is not classified and unblocked data is classified into the same category. The data with similarity between the first and second thresholds is masked data. The masked data does not participate in clustering, and the number of classes is automatically determined using the nature of the data.

Benefits of technology

After clustering, the boundaries between classes are clear and not stuck, and more accurate clustering results are generated without prespecifying the number of classes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234459A_ABST
    Figure CN120234459A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, electronic equipment and a storage medium. Relates to the technical field of data processing. The method comprises the following steps: executing a clustering step: in a data set, determining data of which the similarity with target data is greater than a first threshold value, and if the data meets a preset classification condition, classifying the data and the target data into the same class, the preset classification condition being unclassified and unshielded data, and the target data being each data in the data set; and under the condition that the unclassified data exists in the data set, determining the data, of which the similarity with the target data is smaller than a first threshold value and greater than a second threshold value, in the unclassified data as shielding data, and executing the clustering step on the data except the shielding data. According to the method, data between different thresholds are shielded, so that the problem that data with insufficient similarity values are clustered into the same class, but the data are far away from each other and are easy to overlap with the other class is solved, and the boundaries between the classes obtained after clustering are clear and are not sticky.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data processing, and in particular, to a method, an apparatus, an electronic device, and a storage medium for processing data. Background Art

[0002] With the development of digital technology, there will be a large amount of data in various fields. In order to facilitate data analysis, the problem of clustering data often occurs.

[0003] In the related art, there are many clustering methods for data, but these clustering methods often have some problems, and the unclear boundaries and adhesions between the classes obtained after clustering are problems that need to be solved urgently. Summary of the Invention

[0004] In view of this, the purpose of the present disclosure is to propose a method, an apparatus, an electronic device, and a storage medium for processing data, and the present disclosure can specifically solve existing problems.

[0005] Based on the above purpose, in a first aspect, the present disclosure proposes a method for processing data, where the data is any one of text, voice, and image. The method includes: performing a clustering step: in a data set, determining data whose similarity to the target data is greater than a first threshold, and if the data meets a preset classification condition, classifying the data and the target data into the same class, where the preset classification condition is unclassified and non-blocked data, and the target data is each data in the data set; in the case where there is unclassified data in the data set, determining data whose similarity to the target data is less than the first threshold and greater than a second threshold among the unclassified data as blocked data, and performing the clustering step on data other than the blocked data.

[0006] In a second aspect, there is also provided a data processing apparatus, where the data is any one of text, voice, and image. The apparatus includes: a clustering unit configured to perform a clustering step: in a data set, determining data whose similarity to the target data is greater than a first threshold, and if the data meets a preset classification condition, classifying the data and the target data into the same class, where the preset classification condition is unclassified and non-blocked data, and the target data is each data in the data set; a blocking unit configured to, if there is unclassified data in the data set, determine data whose similarity to the target data is less than the first threshold and greater than a second threshold among the unclassified data as blocked data, and perform the classification step on data other than the blocked data.

[0007] In a third aspect, an electronic device is further provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor runs the computer program to implement the method of the first aspect.

[0008] In a fourth aspect, a computer-readable storage medium is further provided, on which a computer program is stored, and the program is executed by a processor to implement the method of any item of the first aspect.

[0009] Generally speaking, the present disclosure has at least the following beneficial effects: By shielding the data between different thresholds, it avoids the problem that data with insufficient similarity values are clustered into the same class but are likely to overlap with another class due to being far apart, so that the boundaries between the classes obtained after clustering are clear and not sticky. Moreover, the present disclosure does not need to specify the number of classes and can fully utilize the nature of the data for clustering, which helps to generate more accurate clustering results. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In the drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments disclosed in accordance with the present disclosure and should not be regarded as limiting the scope of the present disclosure.

[0011] Figure 1 A flowchart showing a method for processing data according to an embodiment of the present disclosure is shown;

[0012] Figure 2 Another flowchart showing a method for processing data according to an embodiment of the present disclosure is shown;

[0013] Figure 3 A flowchart in the method for processing data is shown;

[0014] Figure 4 A schematic diagram showing a data processing device according to an embodiment of the present disclosure is shown;

[0015] Figure 5 A schematic diagram showing the structure of an electronic device provided by an embodiment of the present disclosure is shown;

[0016] Figure 6 A schematic diagram showing a storage medium provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The present disclosure will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention and are not intended to limit the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.

[0018] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments may be combined with each other. The following will describe the present disclosure in detail with reference to the drawings and in combination with the embodiments.

[0019] Figure 1 A method for processing data of the present disclosure is shown. In an embodiment of the present disclosure, the method includes:

[0020] Step S101, perform a clustering step: in a data set, determine data whose similarity to the target data is greater than a first threshold. If the data meets a preset classification condition, classify the data and the target data into the same class, where the preset classification condition is unclassified and non-blocked data, and the target data is each data in the data set.

[0021] In this embodiment, the above-mentioned execution entity may perform the clustering step. Specifically, the clustering step includes: determining data in the data set, and the similarity between the determined data and the target data is greater than the first threshold. When the determined data meets the preset classification condition, the above-mentioned execution entity classifies the determined data and the target data into the same class. In this way, the above-mentioned execution entity has completed one round of the clustering step. The target data may be any data in the data set. The above-mentioned execution entity may be a terminal or a server, which is not limited herein.

[0022] In this embodiment, the preset classification condition is that the data has not been classified and is non-blocked data. If the data has been classified into any class, or the data is blocked data, it does not meet the preset classification condition. Blocked data refers to data that does not participate in clustering, specifically, data that does not participate in subsequent rounds of clustering.

[0023] Step S102, in the case where there is unclassified data in the data set, determine data whose similarity to the target data is less than the first threshold and greater than a second threshold among the unclassified data as blocked data, and perform the clustering step on the data other than the blocked data.

[0024] In this embodiment, after completing the previous round of the clustering step, if there is unclassified data in the data set, the above-mentioned execution entity may perform the next round of the clustering step. Before performing the next round of the clustering step, the above-mentioned execution entity may determine data whose similarity to the target data is less than the first threshold and greater than the second threshold among the unclassified data as blocked data. Performing the next round of the clustering step on the data other than the blocked data is to perform the clustering step. The first threshold and the second threshold are positive, and the first threshold is greater than the second threshold.

[0025] The data in the data set is any one of text, voice, and image. Specifically, the data in the present disclosure may be sample data or other types of data to be processed. In the case where the data is a sample, the data set is a sample set.

[0026] For example, the above execution entity may perform the following steps on the text data:

[0027] Perform the clustering step of the text data: In the text data set, for the target text data, determine the text data whose similarity to the target text data is greater than the first threshold. If the text data meets the preset classification condition, classify the text data and the target text data into the same class, where the preset classification condition is unclassified and non-blocked data, and the target text data is each data in the data set. In the case where there is unclassified text data in the data set, determine the text data whose similarity to the target text data is less than the first threshold and greater than the second threshold among the unclassified text data as blocked text data, and perform the clustering step of the text data on the text data other than the blocked text data.

[0028] The above execution entity may perform the following steps on the voice data:

[0029] Perform the clustering step of the voice data: In the voice data set, for the target voice data, determine the voice data whose similarity to the target voice data is greater than the first threshold. If the voice data meets the preset classification condition, classify the voice data and the target voice data into the same class, where the preset classification condition is unclassified and non-blocked data, and the target voice data is each data in the data set. In the case where there is unclassified voice data in the data set, determine the voice data whose similarity to the target voice data is less than the first threshold and greater than the second threshold among the unclassified voice data as blocked voice data, and perform the clustering step of the voice data on the voice data other than the blocked voice data.

[0030] The above execution entity may perform the following steps on the image data:

[0031] Perform the clustering step of the image data: In the image data set, for the target image data, determine the image data whose similarity to the target image data is greater than the first threshold. If the image data meets the preset classification condition, classify the image data and the target image data into the same class, where the preset classification condition is unclassified and non-blocked data, and the target image data is each data in the data set. In the case where there is unclassified image data in the data set, determine the image data whose similarity to the target image data is less than the first threshold and greater than the second threshold among the unclassified image data as blocked image data, and perform the clustering step of the voice data on the image data other than the blocked image data.

[0032] In this embodiment, by shielding the data between different thresholds, it is possible to avoid the problem that data with insufficient similarity values are clustered into the same class but are likely to overlap with another class due to being far apart, thereby making the boundaries between the classes obtained after clustering clear and non - sticky. Moreover, the present disclosure does not need to specify the number of classes and can fully utilize the nature of the data for clustering, which helps to generate more accurate clustering results.

[0033] Figure 2 Shows a method for processing data according to an embodiment of the present disclosure. As Figure 2 shown, the method includes:

[0034] Step S201, determine the similarity between each two data in the data set, and generate a similarity set corresponding to the data set.

[0035] In this embodiment, the execution subject of the data processing method can determine the similarity between each two data in the data set, and all the determined similarities form a similarity set. This similarity set corresponds to the data set.

[0036] Step S202, perform a clustering step: in the data set, according to the similarity set, determine the data whose similarity to the target data is greater than a first threshold. If the data meets the preset classification condition, classify the data and the target data into the same class, where the preset classification condition is unclassified and non - shielded data, and the target data is each data in the data set.

[0037] In this embodiment, the above - mentioned execution subject can find the data whose similarity to the target data is greater than the first threshold according to the similarities between the data in the similarity set.

[0038] Step S203, in the case that there is unclassified data in the data set, determine the data whose similarity to the target data is less than the first threshold and greater than a second threshold among the unclassified data as shielded data, and perform the clustering step on the data other than the shielded data.

[0039] In this embodiment, the similarity between each two data can be calculated first to obtain a similarity set, thereby avoiding having to calculate the similarity in real - time every time the similarity between the target data and other data is determined, which helps to improve the clustering efficiency.

[0040] In some optional implementation manners of this embodiment, generating a similarity set corresponding to the data set includes: generating a similarity matrix corresponding to the data set, where each row of the similarity matrix includes the similarities between one data in the data set and all the data.

[0041] In these alternative implementations, the above-mentioned execution entity can generate a similarity matrix for a data set, where each row in the similarity matrix includes the similarities between a data in the data set and all the data.

[0042] Suppose there are three data A, B, and C in the data set. As Figure 3 shown, the figure shows two matrices corresponding to each matrix position. Each matrix position in the upper similarity matrix indicates a similarity value. Each matrix position in the lower matrix indicates which two data the similarity value at the same matrix position in the upper matrix is the similarity value between.

[0043] In the case where the data set is represented as a vector matrix a, the similarity matrix S can be expressed as:

[0044] S = a * a T

[0045] where a T represents the transpose of a, and * represents the dot product, that is, the inner product.

[0046] These implementations can comprehensively and accurately find the data with high similarity to the same data through each row of the similarity matrix.

[0047] In some alternative implementations of this embodiment, determining the data whose similarity to the target data is greater than the first threshold according to the similarity set includes: finding the data whose similarity to the target data is greater than the second threshold according to the similarity set, and in the search results, determining the data whose similarity to the target data is greater than the first threshold.

[0048] In these implementations, the above-mentioned execution entity can first find the data whose similarity to the target data is greater than the second threshold to obtain the search results. Then, the above-mentioned execution entity can continue to narrow the search scope and find the data whose similarity to the target data is greater than the first threshold in the above search results.

[0049] These implementations can achieve finding the data that meets the most accurate search scope by gradually narrowing the search scope of the data.

[0050] In some alternative application scenarios of these implementations, the above-mentioned finding the data whose similarity to the target data is greater than the second threshold according to the similarity set and determining the data whose similarity to the target data is greater than the first threshold in the search results can include: traversing each row of similarities in the similarity matrix to find the data whose similarity to the target data is greater than the second threshold; and in the search results, finding the data whose similarity to the target data is greater than the first threshold.

[0051] In these alternative implementations, the above-mentioned execution entity can traverse the similarity matrix row by row to find data whose similarity to the target data is greater than a second threshold. After the traversal, continue to search in the traversal results to find data whose similarity to the target data is greater than a first threshold.

[0052] For example, the above-mentioned execution entity can find data whose similarity to the target data is greater than the second threshold and record the column index numbers of the found data. The column index number is used to indicate the number of columns in the matrix. After that, the above-mentioned execution entity can traverse the data corresponding to these column index numbers to find data greater than the first threshold among the data corresponding to the column index numbers.

[0053] These implementations can utilize the characteristic that there is the similarity between a certain data in the data set and all data in a matrix row to search for similar data in the similarity matrix, improving the search efficiency.

[0054] In some alternative implementations of any embodiment of the present disclosure, before determining the similarity between each two data in the data set, the method may further include: normalizing the data in the initial data set to obtain a data set.

[0055] In these alternative implementations, the above-mentioned execution entity can normalize each data in the initial data set to obtain the normalized data corresponding to each data, and further obtain the data set corresponding to the initial data set.

[0056] The above-mentioned execution entity can normalize the data in the initial data set in various ways. For example, the above-mentioned execution entity can call a preset normalization model to process the data in the initial data set, thereby obtaining the data set output by the model.

[0057] These implementations can reduce the influence of the data dimension on clustering during the clustering process by normalizing the data, which helps to improve the clustering accuracy.

[0058] In some alternative application scenarios of these implementations, the initial data set is an initial data matrix; normalizing the data in the initial data set to obtain a data set includes: determining the norm of the initial data matrix; dividing the initial data matrix by the norm of the initial data matrix to obtain a quotient, and the quotient is the data set.

[0059] In these application scenarios, the above-mentioned execution entity can determine the norm of the initial data matrix. After that, the above-mentioned execution entity can divide the initial matrix by the above norm to obtain a quotient. The above-mentioned execution entity can use the quotient as the data set. Here, the norm can refer to the square root of the absolute value of the elements in the initial data matrix.

[0060] If the initial data matrix is b, the norm of the initial data matrix is ||b||, and the above quotient x can be expressed as:

[0061]

[0062] These application scenarios can achieve moderate normalization of data through the norm, avoiding the problem of normalizing to too small a numerical range and losing data characteristics.

[0063] The embodiments of the present disclosure provide a data processing device, which is used to execute the data processing method in the above embodiments, and the data is any one of text, voice, and image. As Figure 4 shown, the device includes: a clustering unit 401 configured to execute a clustering step: in a data set, determine data whose similarity to the target data is greater than a first threshold, and if the data meets a preset classification condition, classify the data and the target data into the same class, where the preset classification condition is unclassified and non-screened data, and the target data is each data in the data set; a screening unit 402 configured to, if there is unclassified data in the data set, determine data whose similarity to the target data is less than the first threshold and greater than a second threshold among the unclassified data as screened data, and execute a classification step on data other than the screened data.

[0064] Optionally, the device further includes: a generating unit configured to determine the similarity between every two data in the data set and generate a similarity set corresponding to the data set; the clustering unit 401 is further configured to execute the step of determining data whose similarity to the target data is greater than the first threshold in the following manner: according to the similarity set, determine data whose similarity to the target data is greater than the first threshold.

[0065] Optionally, the generating unit is further configured to execute the step of generating a similarity set corresponding to the data set in the following manner: generate a similarity matrix corresponding to the data set, where each row of the similarity matrix includes the similarity between one data in the data set and all data.

[0066] Optionally, the clustering unit 401 is further configured to execute the step of determining data whose similarity to the target data is greater than the first threshold according to the similarity set in the following manner: according to the similarity set, find data whose similarity to the target data is greater than the second threshold, and among the search results, determine data whose similarity to the target data is greater than the first threshold.

[0067] Optionally, the clustering unit 401 is further configured to perform the following operations: according to the similarity set, find data whose similarity to the target data is greater than a second threshold, and in the search results, determine data whose similarity to the target data is greater than a first threshold: traverse each row of similarities in the similarity matrix to find data whose similarity to the target data is greater than the second threshold; in the search results, find data whose similarity to the target data is greater than the first threshold.

[0068] Optionally, the apparatus is further configured to: before determining the similarity between every two data in the data set, normalize the data in the initial data set to obtain the data set.

[0069] Optionally, the initial data set is an initial data matrix; normalizing the data in the initial data set to obtain the data set includes: determining the norm of the initial data matrix; dividing the initial data matrix by the norm of the initial data matrix to obtain a quotient, and the quotient is the data set.

[0070] The data processing apparatus provided in the above embodiments of the present disclosure and the data processing method provided in the embodiments of the present disclosure are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0071] The embodiments of the present disclosure also provide an electronic device corresponding to the data processing method provided in the foregoing embodiments to execute the above data processing method. The embodiments of the present disclosure do not make limitations.

[0072] Please refer to Figure 5 , which shows a schematic diagram of an electronic device provided in some embodiments of the present disclosure. As Figure 5 shown, the electronic device 50 includes: a processor 500, a memory 501, a bus 502, and a communication interface 503. The processor 500, the communication interface 503, and the memory 501 are connected through the bus 502; a computer program that can run on the processor 500 is stored in the memory 501, and when the processor 500 runs the computer program, it executes the method provided in any of the foregoing embodiments of the present disclosure.

[0073] Among them, the memory 501 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 503 (which can be wired or wireless), a communication connection is realized between the system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0074] The bus 502 can be an ISA bus, a PCI bus, an EISA bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 501 is used to store a program. After receiving an execution instruction, the processor 500 executes the program. The data processing method disclosed in any implementation manner of the foregoing embodiments of the present disclosure can be applied to or implemented by the processor 500.

[0075] The processor 500 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 500 or by instructions in the form of software. The above-mentioned processor 500 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 501, and the processor 500 reads the information in the memory 501 and combines its hardware to complete the steps of the above method.

[0076] The electronic device provided by the embodiments of the present disclosure and the data processing method provided by the embodiments of the present disclosure are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by them.

[0077] The present disclosure embodiment also provides a computer-readable storage medium corresponding to the data processing method provided by the foregoing embodiment. Please refer to Figure 6 which shows that the computer-readable storage medium is an optical disc 60, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the data processing method provided by any of the foregoing embodiments.

[0078] It should be noted that examples of the computer-readable storage medium may further include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical or magnetic storage media, which will not be elaborated one by one here.

[0079] The computer-readable storage medium provided by the above embodiments of the present disclosure and the data processing method provided by the embodiments of the present disclosure are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.

[0080] It should be noted that:

[0081] In the above text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present disclosure is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0082] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present disclosure.

[0083] The embodiments of the present disclosure have been described above in conjunction with the accompanying drawings, which are only specific embodiments of the present disclosure. However, the present disclosure is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present disclosure, those of ordinary skill in the art can also make many forms without departing from the purpose of the present disclosure and the scope protected by the claims, and all of them fall within the protection scope of the present disclosure.

Claims

1. A method for processing data, characterized in that The data is any one of text, speech, and images, and the method includes: Performing a clustering step: In a data set, determining data whose similarity to target data is greater than a first threshold. If the data meets a preset classification condition, classifying the data and the target data into the same class, where the preset classification condition is unclassified and non-blocked data, and the target data is each data in the data set; In the case where there is unclassified data in the data set, determining data whose similarity to the target data is less than the first threshold and greater than a second threshold among the unclassified data as blocked data, and performing the clustering step on the data other than the blocked data.

2. The method according to claim 1, wherein: The method further includes: determining the similarity between each two data in the data set, and generating a similarity set corresponding to the data set; The determining of data whose similarity to the target data is greater than the first threshold includes: determining, according to the similarity set, data whose similarity to the target data is greater than the first threshold.

3. The method according to claim 2, characterized in that, The generating of the similarity set corresponding to the data set includes: Generating a similarity matrix corresponding to the data set, where each row of the similarity matrix includes the similarity between one data in the data set and all data.

4. The method according to claim 3, wherein The determining, according to the similarity set, of data whose similarity to the target data is greater than the first threshold includes: Searching, according to the similarity set, for data whose similarity to the target data is greater than the second threshold, and among the search results, determining data whose similarity to the target data is greater than the first threshold.

5. The method according to claim 4, characterized in that, The searching, according to the similarity set, for data whose similarity to the target data is greater than the second threshold and determining, among the search results, data whose similarity to the target data is greater than the first threshold includes: Traversing each row of similarities in the similarity matrix to search for data whose similarity to the target data is greater than the second threshold; Among the search results, searching for data whose similarity to the target data is greater than the first threshold.

6. The method according to claim 1, characterized in that, Before determining the similarity between each two data in the data set, the method further includes: Normalizing the data in the initial data set to obtain the data set.

7. The method according to claim 6, wherein The initial data set is an initial data matrix; The normalizing of the data in the initial data set to obtain the data set includes: Determining the norm of the initial data matrix; Dividing the initial data matrix by the norm of the initial data matrix to obtain a quotient, and the quotient is the data set.

8. A data processing device, characterized in that, The data is any one of text, speech, and images, and the apparatus includes: A clustering unit configured to perform the clustering step: In a data set, determining data whose similarity to target data is greater than a first threshold. If the data meets a preset classification condition, classifying the data and the target data into the same class, where the preset classification condition is unclassified and non-blocked data, and the target data is each data in the data set; The shielding unit is configured to, if there is unclassified data in the data set, determine the data in the unclassified data whose similarity to the target data is less than the first threshold and greater than the second threshold as shielding data, and perform the classification step on the data other than the shielding data.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor runs the computer program to implement the method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method according to any one of claims 1-7.