Data classification method and device and storage medium

Through the combination of regular matching and text semantic similarity model, the accuracy and adaptability of complex semantic data classification in the existing technology are solved, and efficient classification of financial texts, alarm logs and other data is realized, reducing the need for manual annotation and model training.

CN120372469APending Publication Date: 2025-07-25BEIJING VENUS INFORMATION SECURITY TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510472935.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately classify data involving complex semantics, especially data such as financial texts, alarm logs and health indicators. The existing methods require a large amount of manual annotation and model retraining, which has poor adaptability.

Method used

The method of regular matching combined with text semantic similarity model is adopted to determine data characteristics through regular matching, and the text semantic similarity model is used to calculate the similarity when there is no unique match. The text semantic similarity model is combined with the text semantic similarity model built by the tree data classification specification and the CoSENT framework to realize data classification.

Benefits of technology

The accurate classification of complex semantic data is realized, the need for manual labeling is reduced, the adaptability and classification efficiency is improved, and the confusion problem of using text semantic similarity calculation alone is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372469A_ABST
    Figure CN120372469A_ABST
Patent Text Reader

Abstract

The invention discloses a data classification method and device and a storage medium, and the method comprises the steps: determining the data features of all user data based on a regular matching mode; performing classification operation on each piece of user data; the classification operation comprises the steps of matching data features of the user data with data features of each type of data in stored data classification specifications, and if a unique matching result exists, affiliating the user data to a corresponding type in the matched data classification specifications; and if the unique matching result does not exist, calculating the similarity between the user data and each class of data in the stored data classification specification by utilizing a text semantic similarity model, and if the similarity greater than a preset threshold exists, affiliating the user data to the class corresponding to the similarity, so that accurate classification of the data can be realized, and the user experience is improved. And the method has universality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to data classification technology, particularly to a data classification method, device, and storage medium. Background Art

[0002] Data classification and hierarchical protection is a technology that classifies data according to its sensitivity level and determines corresponding protection levels for each classified data to ensure the security of data during storage, transmission, and processing. This technology not only helps to clearly identify the intrinsic value and potential risks of data but also lays a foundation for subsequent data protection measures, access control, and compliance management.

[0003] Generally speaking, the steps for data classification and grading are as follows: First, according to the existing data classification and grading specifications, user data is divided into specific categories; then, according to the grading corresponding to each category in the data classification and grading specifications, the classified user data is mapped to the corresponding level. Common levels include public level, internal level, confidential level, etc. In data classification and hierarchical protection, accurately classifying user data is the key.

[0004] Common user data classification methods include: The regular matching method has a good matching effect for digital data (such as phone numbers, ID numbers, etc.), but for data involving highly specialized fields that require understanding based on context logic, such as financial texts, alarm logs, health indicators, etc., it cannot match well; for example, financial texts may contain complex descriptions of financial operations, alarm logs involve details of system state changes, and health indicators involve medical terms and descriptions of physiological states. This type of data requires semantic understanding, and regular paradigms (such as cosine similarity) often cannot capture complex semantics; The method based on text classification model matching requires a large amount of labeled text to be provided for the model. For example, XXXXXX text belongs to category A, and YYYYYY text belongs to category B. After training the model with the labeled text, the model can be used to distinguish category A text or category B text. This method requires a large amount of time and effort from relevant technical personnel to label the text and cannot be widely used; moreover, once a new category is introduced, the model needs to be retrained, which is inconvenient to use; and if the user data to be classified is quite different from the training text, the accuracy of the recognition result is relatively low; A method for classification based on TF-IDF vector similarity. The TF-IDF vector is a vector based on word frequency and inverse document frequency index. The prerequisite for this method to be applicable is that the form of user data is similar to the data form recorded in the data classification and grading specification. For example, in user data, "xyz" is uniformly used to represent "minor language". Only when the data form containing "xyz" appears in the data classification and grading specification, the vectors of user data and the data classification and grading specification are similar and can be classified into the same category. In fact, it is impossible to ensure that the data form in the data classification and grading specification is the same as that of user data. Summary of the Invention

[0005] Embodiments of the present application provide a data classification method, apparatus, and storage medium, which can achieve accurate classification of data and have universality.

[0006] The data classification method provided by the embodiments of the present application includes: Determine the data features of all user data based on regular matching; Perform classification operations on each user data respectively; the classification operation includes: matching the data features of the user data with the data features of each type of data in the stored data classification specification. If there is a unique matching result, the user data belongs to the corresponding class in the data classification specification that matches; if there is no unique matching result, calculate the similarity between the user data and each type of data in the stored data classification specification using a text semantic similarity model. If there is a similarity greater than a preset threshold, the user data belongs to the class corresponding to the similarity.

[0007] As an implementation example, the method of determining the data features of all user data based on regular matching includes one or more of the following methods: Determine the data features of all user data based on a regular expression; Determine the data features of all user data based on a regular expression and a preset matching rate; Determine the data features of all user data based on a regular expression and a preset matching priority of data features; Determine the data features of all user data based on a regular expression and case sensitivity; Determine the data features of all user data based on a regular expression and a minimum unique value; Determine the data features of all user data based on a regular expression and data type.

[0008] As an implementation example, the data classification specification has a tree structure, and each type of data in the data classification specification is a leaf node of the tree structure; When the user data is the column data of a table, calculating the similarity between the user data and each type of data in the stored data classification specification by using the text semantic similarity model includes: Obtain the column name vector, column annotation vector, table name vector, and table annotation vector of the column data; Obtain the leaf node vector and parent node vector of each type of data; For each type of data, perform the following operations: Calculate the first similarity between the column name vector and the leaf node vector of this type of data; Calculate the second similarity between the column annotation vector and the leaf node vector of this type of data; Calculate the third similarity between the table name vector and the parent node vector of this type of data; Calculate the fourth similarity between the table annotation vector and the parent node vector of this type of data; Perform a weighted sum of the first similarity, the second similarity, the third similarity, and the fourth similarity to obtain the similarity between the user data and this type of data.

[0009] As an implementation example, calculating the third similarity between the table name vector and the parent node vector of this type of data includes: When the parent node vector of this type of data is not unique, calculate the fifth similarity between the table name vector and each parent node vector; select the maximum third similarity from all the fifth similarities as the third similarity; Calculating the fourth similarity between the table annotation vector and the parent node vector of this type of data includes: When the parent node vector of this type of data is not unique, calculate the sixth similarity between the table annotation vector and each parent node vector; select the maximum sixth similarity from all the sixth similarities as the fourth similarity.

[0010] As an implementation example, if there is a similarity greater than the preset threshold, classifying the user data into the class corresponding to the similarity includes: When the similarities greater than the preset threshold are not unique, classify the user data into the class corresponding to the maximum similarity among the multiple similarities.

[0011] As an implementation example, the establishment method of the text semantic similarity model includes: Build the text semantic similarity model based on the CoSENT framework; Use the multilingual-e5-base model as the vector generation model in the text semantic similarity model.

[0012] As an implementation example, the method further includes: After performing the classification operation on each user data, determine whether there is user data without an assigned class; If there is user data without an assigned class, perform a clustering operation on all user data without an assigned class.

[0013] As an implementation example, the clustering operation includes: Perform dimensionality reduction on the user data without an assigned class through the UMAP dimensionality reduction algorithm; Use the HDBSCAN clustering algorithm to cluster the dimensionality-reduced user data.

[0014] The non-transitory computer-readable storage medium provided by the embodiments of the present application stores one or more program instructions, and the one or more program instructions can be executed by one or more processors to implement the method described in any previous embodiment.

[0015] The data classification device provided by the embodiments of the present application includes: A storage module configured to store computer program instructions that can run on a processor; A processing module configured to execute the computer program instructions to implement the method described in any previous embodiment.

[0016] In the technical solution described in the embodiments of the present application, the data characteristics of all user data are determined based on regular matching; a classification operation is performed on each user data separately; the classification operation includes: matching the data characteristics of the user data with the data characteristics of each class of data in the stored data classification specification. If there is a unique matching result, the user data is assigned to the corresponding class in the data classification specification that is matched; if there is no unique matching result, use a text semantic similarity model to calculate the similarity between the user data and each class of data in the stored data classification specification. If there is a similarity greater than a preset threshold, the user data is assigned to the class corresponding to the similarity. This solution combines regular methods and text semantic similarity calculation methods to determine the classification of data, which can not only classify digital user data well, but also classify data that needs to be understood based on context logic well, and can also avoid the deficiency of confusing similar semantic texts when using the text semantic similarity calculation method alone; in addition, this method does not require manual annotation of texts and does not require the user data to have the same data form as the data in the data classification specification, and has good universality.

[0017] Other features and advantages of the present application will be described in the subsequent specification, and some of them will become obvious from the specification or be understood by implementing the present application. Other advantages of the present application can be realized and obtained through the solutions described in the specification and the drawings. Brief Description of the Drawings

[0018] The drawings are used to provide an understanding of the technical solutions of the present application and form a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solutions of the present application and do not constitute a limitation to the technical solutions of the present application.

[0019] Figure 1 It is an example diagram of the data classification method provided for the embodiments of the present application; Figure 2 It is an example diagram of various data features and their corresponding matching conditions provided for the embodiments of the present application; Figure 3 It is an example diagram of a tree-shaped data classification specification provided for the embodiments of the present application; Figure 4 It is a partial classification example diagram of a data classification specification for the telecommunications industry provided for the embodiments of the present application; Figure 5 It is a module diagram of a data classification device provided for the embodiments of the present application. Detailed Description of the Embodiments

[0020] The present application describes multiple embodiments, but the description is exemplary rather than restrictive, and it is obvious to those of ordinary skill in the art that there can be more embodiments and implementation solutions within the scope covered by the embodiments described in the present application. Although many possible combinations of features are shown in the drawings and discussed in the detailed description, many other combinations of the disclosed features are also possible. Unless specifically restricted, any feature or element of any embodiment can be combined with any other feature or element in any other embodiment, or can replace any other feature or element in any other embodiment.

[0021] The present application includes and contemplates combinations with features and elements known to those of ordinary skill in the art. The embodiments, features, and elements already disclosed in the present application can also be combined with any conventional features or elements to form unique inventive solutions. Any feature or element of any embodiment can also be combined with features or elements from other inventive solutions to form another unique inventive solution. Therefore, it should be understood that any feature shown and / or discussed in the present application can be implemented alone or in any suitable combination. Therefore, except for the limitations made according to the appended claims and their equivalents, the embodiments are not subject to other limitations. In addition, various modifications and changes can be made within the scope of protection of the appended claims.

[0022] In addition, when describing representative embodiments, the specification may have presented the method and / or process as a specific sequence of steps. However, to the extent that the method or process does not depend on the specific order of the steps described herein, the method or process should not be limited to the specific order of steps described. As will be understood by those of ordinary skill in the art, other step sequences are possible. Therefore, the specific order of steps set forth in the specification should not be construed as a limitation on the claims. In addition, the claims directed to the method and / or process should not be limited to performing their steps in the order written, as those skilled in the art can readily understand that these orders can vary and still remain within the spirit and scope of the embodiments of the present application.

[0023] The user data involved in the embodiments of the present application is generally divided into structured data and unstructured data. Structured data mostly exists in the form of tables, such as records in a relational database management system (such as PostgreSQL, MySQL), with a fixed column and row format, facilitating direct querying and analysis; unstructured data has no unified format and structure, such as Office documents, PDF documents, etc.

[0024] The data classification method provided by the embodiments of the present application, as Figure 1 shown, the method includes: Step S101: Determine the data characteristics of all user data based on regular matching; Perform a classification operation on each user data respectively, and the classification operation includes: Step S102: Match the data characteristics of the user data with the data characteristics of each type of data in the stored data classification specification; Step S103: Determine whether there is a unique matching result. If there is, perform Step S106; if not, perform Step S104; Step S104: Calculate the similarity between the user data and each type of data in the stored data classification specification using a text semantic similarity model; The situation where there is no unique matching result may mean that the matching results are not unique. Then, when calculating using the text semantic similarity model, the user data can be compared for similarity only with multiple types of data in the data classification specification that have been matched, which can reduce the amount of data to be matched; the situation where there is no unique matching result may also mean that there is no matching result. At this time, when calculating using the text semantic similarity model, it is necessary to compare the user data with each type of data in the stored data classification specification for similarity; The text semantic similarity (STS) model is a model for evaluating the semantic similarity between different texts. This model uses natural language processing techniques, such as word embedding, deep learning models, etc., to convert the text pairs <A, B> to be compared into vectors, which are considered as sentence embeddings with semantic information. Subsequently, by calculating the similarity between the two vectors, the complex relationships between words are captured, thereby quantifying the semantic similarity degree between texts; Step S105 determines whether there is a similarity greater than a preset threshold. If so, step S106 is executed; Step S106 classifies the user data into the corresponding class in the data classification specification.

[0025] The data classification method described in the embodiments of this application combines the regular method and the text semantic similarity calculation method to determine the classification of data. It can not only classify numerical user data well, but also classify data that needs to be understood based on context logic well. It can also avoid the deficiency of confusing similar semantic texts when using the text semantic similarity calculation method alone; in addition, this method does not require manual annotation of texts, does not require the user data to have the same data form as the data in the data classification specification, and has good universality.

[0026] The data features in the embodiments of this application may include one or more of the following: Phone number, email, name, address, province, decimal, integer, percentage, string, true / false value, English, ID number, bank card number, garbled code; the shown data features are only examples and can be extended based on actual user data.

[0027] In an exemplary embodiment, the method for determining the data features of all user data based on regular matching includes one or more of the following methods: Determine the data features of all user data based on regular expressions; such as determining phone numbers, bank card numbers, etc. based on regular expressions; Determine the data features of all user data based on regular expressions and a preset matching rate; for example, the preset matching rate corresponding to "integer" is 100%. Assuming that the user data is the column data of a table, match the non-empty data in this column. Only when the column data is matched as an integer based on the regular expression and the matching rate reaches 100%, can the data feature of this column data be determined as an integer; Determine the data features of all user data based on regular expressions and the preset matching priority of data features; for example, the preset matching priority of the data feature "string" is lower than that of the data feature "name". When facing user data composed of a string of characters, the regular expression that can determine "name" is preferentially used for matching; Determine the data characteristics of all user data based on regular expressions and whether case is sensitive; for example, when performing gender matching, some user data uses male / female, and some uses Male / Female, so the gender matching regular expression can be set to be case-insensitive. Determine the data characteristics of all user data based on regular expressions and the minimum unique value; for example, when determining the data characteristic "true / false value", not only does the user data need to satisfy the regular expression for determining the "true / false value", but the minimum unique value also needs to be 2 (indicating that there are 2 numerical values in this column of data). Determine the data characteristics of all user data based on regular expressions and data types. For example, for the data characteristic "integer", the corresponding user data may be an integer or a string, but cannot be a decimal, boolean value, etc. Therefore, in the case where the user data is initially determined to be a decimal, there is no need to further determine the data characteristics based on the integer regular expression.

[0028] Figure 2 Examples of various data characteristics and their corresponding matching conditions are given.

[0029] The method for determining data characteristics described in the embodiments of this application can not only obtain data characteristics through conventional regular expressions, such as mobile phone numbers, ID numbers, bank card numbers, etc., but also support a wider range of regular matching methods, enriching the types of data characteristics; moreover, because the embodiments of this application will subsequently make judgments based on text semantic similarity, even if a relatively vague regular item, such as an integer, is matched based on the data characteristics, it will not affect the classification result. In addition, it can also reduce the number of regular matches and improve the matching efficiency.

[0030] In an exemplary embodiment, when matching the data characteristics of the user data with the data characteristics of each category of data in the stored data classification specification, the data characteristics of each category of data in the stored data classification specification can be manually labeled or labeled by a program; even if it is manually labeled, since the labeling can rely on intuition, such as the data characteristic of sales volume is an integer; the object of labeling only depends on the number of classifications in the classification specification, and the number is limited; this determines that the manual labeling method can also be implemented at low cost and the labeling efficiency is not low.

[0031] In an exemplary embodiment, the data classification specification has a tree structure, and each category of data in the data classification specification is a leaf node of the tree structure; Figure 3 A schematic diagram showing a tree-shaped data classification specification is shown.

[0032] Take Figure 4Taking the partial classification intercepted from "YD / T 3813—2020 Classification and Grading Method for Basic Telecommunication Enterprise Data" as an example, 2-1-1-1 product information is a type of data, and 2-2-1-2 channel information is a type of data. The parent node of these two types of data is 2-2-1 service operation service data.

[0033] When the user data is column data in a table, the calculating of the similarity between the user data and each type of data in the stored data classification specification by using the text semantic similarity model includes: Obtaining the column name vector, column annotation vector, table name vector, and table annotation vector of the column data; Obtaining the leaf node vector and parent node vector of each type of data; For each type of data, perform the following operations: Calculating a first similarity between the column name vector and the leaf node vector of this type of data; Calculating a second similarity between the column annotation vector and the leaf node vector of this type of data; Calculating a third similarity between the table name vector and the parent node vector of this type of data; Calculating a fourth similarity between the table annotation vector and the parent node vector of this type of data; Performing a weighted sum of the first similarity, the second similarity, the third similarity, and the fourth similarity to obtain the similarity between the user data and this type of data.

[0034] In an exemplary embodiment, the weights of the first similarity and the second similarity are greater than the weights of the third similarity and the fourth similarity; the weights of the first similarity and the second similarity may be the same or different; the weights of the third similarity and the fourth similarity may be the same or different.

[0035] In an exemplary embodiment, the calculating of the third similarity between the table name vector and the parent node vector of this type of data includes: When the parent node vector of this type of data is not unique, calculating a fifth similarity between the table name vector and each parent node vector; selecting the largest third similarity from all the fifth similarities as the third similarity; The calculating of the fourth similarity between the table annotation vector and the parent node vector of this type of data includes: When the parent node vector of this type of data is not unique, calculating a sixth similarity between the table annotation vector and each parent node vector; selecting the largest sixth similarity from all the sixth similarities as the fourth similarity.

[0036] By implementing the solution recorded in the embodiment, the number of similarities participating in the final weighted sum calculation can be reduced, and the classification efficiency can be improved.

[0037] In an exemplary embodiment, if there is a similarity greater than a preset threshold, attributing the user data to a class corresponding to the similarity includes: When the similarity greater than the preset threshold is not unique, the user data is assigned to a class corresponding to a maximum similarity among multiple similarities.

[0038] In an exemplary embodiment, the text semantic similarity model is established by: Building the text semantic similarity model based on the CoSENT framework; The multilingual-e5-base model is used as a vector generation model in the text semantic similarity model.

[0039] The CoSENT framework generally includes the following modules: The vector generation module converts the input data into a vector of fixed dimension. In this embodiment, the multilingual-e5-base model is used as the vector generation model in the text semantic similarity model. This quantitative comparison can not only reveal the direct vocabulary overlap between texts, but also identify the implicit association based on the context, which is particularly important for understanding and organizing complex database structures. Similarity calculation module, calculates the cosine similarity between vectors. The value of cosine similarity is between -1 and 1, where 1 indicates complete similarity, -1 indicates complete dissimilarity, and 0 indicates orthogonality; The training module uses contrastive learning to optimize parameters during the training process, which usually involves constructing positive sample pairs and negative sample pairs. The training goal is to maximize the cosine similarity between positive sample pairs and minimize the cosine similarity between negative sample pairs. This embodiment can obtain negative sample pairs by disrupting the column data from the same table, that is, randomly staggering the column name and column annotation data. The sorting loss function used in the training can be directly optimized for the similarity calculation, so that the training process is highly consistent with the final reasoning goal, which not only accelerates the model convergence process, but also significantly improves the accuracy and robustness of the model in text similarity judgment. In order to improve the training effect, the data input to the model can be cleaned. The cleaned data can help the model generalize better and improve the performance of the model on unseen data. By reducing the redundancy and irrelevant information in the training row data, the computing resources and time required for model training can be reduced.

[0040] To evaluate the performance of the text semantic similarity model established in the embodiments of this application, the multilingual-e5-base model, the multilingual-e5-small model, and the bge-base-en-v1.5 model are respectively used as the vector generation models in the text semantic similarity model, and the Pearson coefficient and the Spearman coefficient are used to evaluate the model results. As shown in Table 1, it can be seen that when using the multilingual-e5-base model, there is a strong correlation between the prediction results and the actual results of the text semantic similarity model.

[0041] Base model Pearson coefficient Spearman coefficient multilingual-e5-base 0.8169 0.7964 multilingual-e5-small 0.8058 0.7962 bge-base-en-v1.5 0.7759 0.7686 Table 1 To more intuitively display the results of the text semantic similarity model, taking "residential address" as an example of the category in the data classification specification, the column names with vector distances less than the threshold are obtained. As shown in Table 2, user data belonging to the "residential address" category can be found more comprehensively and accurately.

[0042] Column ID Column name Column comment 4142 living place Current place of residence 5433 home address Family address 12963 residence Permanent place of residence 15391 domicile place Place of household registration 28512 cur residence Current place of residence 36800 live address Residential address in Chinese and English 36907 live addr Place of residence 46558 residential address Family residential address 49579 home address Family address Table 2 Using the text semantic similarity calculation model described in the embodiments of this application, columns semantically similar to the categories in the classification specification can be identified from a large number of data columns, and its accuracy far exceeds that of traditional manual rule methods, greatly improving the efficiency and quality of data classification.

[0043] User data is often all-inclusive, and it is difficult for industry classification and grading specifications to cover everything comprehensively. After data classification in the manner of any of the foregoing embodiments, there may still be some user data that has not been classified. Therefore, in another exemplary embodiment, the data classification method may further include: After performing the classification operation on each user data, determine whether there is user data without a belonging category; wherein, the user data without a belonging category refers to that, using the text semantic similarity model, no belonging category can still be found in the data classification specification. For example, when calculating the similarity between user data and category data using the text semantic similarity model, the similarity is not greater than the preset threshold. If there is user data without a belonging category, perform a clustering operation on all user data without a belonging category.

[0044] After the text is vectorized using the text semantic similarity model, high-dimensional dense vectors are usually obtained. However, distance vectors required by clustering algorithms, such as Euclidean distance and Manhattan distance, will lose their meaning when facing high-dimensional data (Curse of Dimensionality). Therefore, in the embodiments of this application, when performing clustering operations, it is necessary to perform dimensionality reduction on the high-dimensional dense vectors. Exemplarily, the UMAP dimensionality reduction algorithm can be selected to perform dimensionality reduction on user data without assigned classes. It has a fast dimensionality reduction speed and can better preserve the global structure of the data.

[0045] In an exemplary embodiment, the HDBSCAN clustering algorithm is used to cluster the dimensionality-reduced user data. The HDBSCAN clustering algorithm reduces the dependence on external parameters by introducing a hierarchical clustering perspective. It only needs to set a basic distance parameter to work, and sometimes it can even automatically estimate the optimal parameters, greatly simplifying the parameter tuning process and improving the usability and intuitiveness of the algorithm. In addition, the unique feature of the HDBSCAN clustering algorithm is that it can identify the natural density grading in the data, which means that even in the case of significant changes in cluster density, the algorithm can robustly identify various clusters. Therefore, HDBSCAN clustering shows higher robustness and accuracy in dealing with complex distributions, variable densities, and noisy data.

[0046] In an exemplary embodiment, the vectors used in the HDBSCAN clustering algorithm can be TF-IDF vectors or semantic text vectors, or both vectors can be used simultaneously. Each of the two vectors has its own advantages and disadvantages. Semantic text vectors contain semantic information, while TF-IDF vectors are good at capturing special words in uncommon scenarios.

[0047] To concretely show the clustering ability of the clustering operation adopted in the embodiments of this application, Table 3 provides an example of clustering data. After obtaining this clustering data, it can be manually labeled as a type of data.

[0048]

[0049] Table 3 Through the clustering method described in this application, it can be ensured that all user data is accurately classified.

[0050] The data classification method described in the foregoing embodiments of this application can also be combined with existing data classification methods. Exemplarily, first use the existing data classification method to classify all user data, and for user data whose assigned class cannot be determined using the existing data classification method, then use the data classification method described in this embodiment for classification.

[0051] The embodiments of the present application further provide a non-transitory computer-readable storage medium storing one or more program instructions, which can be executed by one or more processors to implement the method described in any of the previous embodiments.

[0052] The embodiments of the present application further provide a data classification device, as Figure 5 shown. The device includes: A storage module 501 configured to store computer program instructions executable on a processor; A processing module 502 configured to execute the computer program instructions to implement the method described in any of the previous embodiments.

[0053] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware implementation, the division of the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be executed by several physical components in cooperation. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term "computer storage medium" includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

Claims

1. A data classification method, the method comprising: Determining the data features of all user data based on regular matching; Performing a classification operation on each user data separately; The classification operation includes: matching the data features of the user data with the data features of each type of data in the stored data classification specification. If there is a unique matching result, the user data is classified into the corresponding class in the data classification specification that is matched. If there is no unique matching result, calculating the similarity between the user data and each type of data in the stored data classification specification using a text semantic similarity model. If there is a similarity greater than a preset threshold, the user data is classified into the class corresponding to the similarity.

2. The method according to claim 1, wherein The determining the data features of all user data based on regular matching includes one or more of the following methods: Determining the data features of all user data based on a regular expression; Determining the data features of all user data based on a regular expression and a preset matching rate; Determining the data features of all user data based on a regular expression and a preset matching priority of data features; Determining the data features of all user data based on a regular expression and whether case is sensitive; Determining the data features of all user data based on a regular expression and a minimum unique value; Determining the data features of all user data based on a regular expression and a data type.

3. The method according to claim 1, wherein The data classification specification has a tree structure, and each type of data in the data classification specification is a leaf node of the tree structure; When the user data is column data of a table, the calculating the similarity between the user data and each type of data in the stored data classification specification using a text semantic similarity model includes: Obtaining a column name vector, a column comment vector, a table name vector, and a table comment vector of the column data; Obtaining a leaf node vector and a parent node vector of each type of data; For each type of data, performing the following operations: Calculating a first similarity between the column name vector and the leaf node vector of the type of data; Calculating a second similarity between the column comment vector and the leaf node vector of the type of data; Calculating a third similarity between the table name vector and the parent node vector of the type of data; Calculating a fourth similarity between the table comment vector and the parent node vector of the type of data; Performing a weighted sum of the first similarity, the second similarity, the third similarity, and the fourth similarity to obtain the similarity between the user data and the type of data.

4. The method according to claim 3, wherein The calculating the third similarity between the table name vector and the parent node vector of the type of data includes: When the parent node vector of the type of data is not unique, calculating a fifth similarity between the table name vector and each parent node vector; selecting the maximum third similarity from all the fifth similarities as the third similarity; The calculating the fourth similarity between the table comment vector and the parent node vector of the type of data includes: In the case where the parent node vectors of this type of data are not unique, calculate the sixth similarity between the table annotation vector and each parent node vector; select the maximum sixth similarity from all the sixth similarities as the fourth similarity.

5. The method according to claim 1, wherein the step of, if there is a similarity greater than a preset threshold, attributing the user data to the class corresponding to the similarity, includes: in the case where the similarities greater than the preset threshold are not unique, attributing the user data to the class corresponding to the maximum similarity among the multiple similarities.

6. The method according to claim 1 or 3, wherein the manner of establishing the text semantic similarity model includes: building the text semantic similarity model based on the CoSENT framework; using the multilingual-e5-base model as the vector generation model in the text semantic similarity model.

7. The method according to claim 1, wherein The method further includes: after performing the classification operation on each user data, determining whether there is user data without an attributed class; if there is user data without an attributed class, performing a clustering operation on all the user data without an attributed class.

8. The method according to claim 7, wherein the clustering operation includes: dimensionality reduction of the user data without an attributed class by using the UMAP dimensionality reduction algorithm; clustering the dimensionality-reduced user data by using the HDBSCAN clustering algorithm.

9. A non-transitory computer-readable storage medium storing one or more program instructions, the one or more program instructions being executable by one or more processors to implement the method according to any one of claims 1-8.

10. A data classification device, characterized in that, The apparatus includes: a storage module configured to store computer program instructions executable on a processor; a processing module configured to execute the computer program instructions to implement the method according to any one of claims 1-8.