Data classification method and device, computer equipment and storage medium
By automatically generating data classification rules and using the characteristics of business attribute data, the problem of inefficient manual formulation in the existing technology is solved, and efficient and accurate data classification is achieved.
Patent Information
- Application Number
- CN202510442749.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-08-01
AI Technical Summary
In the prior art, data classification rules rely on manual formulation, are inefficient, have a long implementation cycle and rely on the experience of implementers.
By automatically generating classification rules, using the classification characteristics and hierarchical characteristics of multiple business attribute data, initial classification rules are generated, and through comparison and adjustment, the target classification rules matching the data set to be classified are obtained.
It greatly saves time for manual formulation and improves the efficiency of classification rules generation. The generated rules do not depend on the experience of the implementer, and improves the accuracy and coverage of data classification.
Smart Images

Figure CN120408399A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a data classification method, apparatus, computer device, and storage medium. Background Art
[0002] Data classification is a key step in data security management, and it is usually based on pre-set classification rules to classify data to ensure data security and privacy.
[0003] The pre-set classification rules are usually related to the business attributes of the data, and the corresponding classification rules are manually set through different business attributes. For example, in the payment business, there are relatively high requirements for data privacy and security, and the set classification rules can include specifications for grading data security and encryption, etc.; for another example, in the drug management business, it is necessary to classify and categorize traditional Chinese medicines, Western medicines, prescription drugs, over-the-counter drugs, and the types of drugs for treating corresponding diseases.
[0004] The existing classification rules are usually formulated manually, which has problems of low efficiency, long implementation cycle, and extremely relying on the experience of implementers. Summary of the Invention
[0005] In view of this, this application proposes a data classification method, apparatus, computer device, and storage medium to solve the problems of low efficiency, long implementation cycle, and extremely relying on the experience of implementers in the existing method of manually formulating classification rules.
[0006] The first aspect embodiment of this application proposes a data classification method, including:
[0007] Obtain a data set to be classified;
[0008] Extract multiple first category features of the data set to be classified;
[0009] Adjust the data classes included in the initial classification rules according to the multiple first category features to obtain target classification rules that match the data set to be classified, where the initial classification rules are generated according to multiple business attribute data of the target business to which the data set to be classified belongs;
[0010] Classify the data set to be classified according to the target classification rules.
[0011] The embodiment of this application can greatly save the manual formulation time, improve the generation efficiency of classification rules, and the generated classification rules do not rely on the experience of implementers by automatically generating classification rules that conform to the user's business scenario.
[0012] In the embodiments of the present application, before adjusting the data classes included in the initial classification rule according to the multiple first category features, the following steps are further included:
[0013] For any one of the multiple service attribute data, classification features and grading features are extracted from the service attribute data; any one of the service attribute data includes classification information of the target service.
[0014] Based on the multiple classification features and multiple grading features corresponding to the multiple service attribute data, the initial classification rule is generated.
[0015] By integrating multiple service attribute data of the target service, the embodiments of the present application can improve the coverage of the generated classification rule, thereby greatly improving the accuracy and efficiency of data classification.
[0016] In the embodiments of the present application, generating the initial classification rule based on the multiple classification features and multiple grading features corresponding to the multiple service attribute data includes:
[0017] Select the same classification features from the multiple classification features, and remove the same classification features from the multiple classification features to obtain multiple different classification features.
[0018] Select the same grading features from the multiple grading features, and remove the same grading features from the multiple grading features to obtain multiple different grading features.
[0019] Based on the same classification features, the multiple different classification features, the same grading features, and the multiple different grading features, the initial classification rule is generated.
[0020] In the embodiments of the present application, generating the initial classification rule based on the same classification features, the multiple different classification features, the same grading features, and the multiple different grading features includes:
[0021] According to the same classification features and the same grading features, a first classification rule is generated.
[0022] The first classification rule is adjusted by the multiple different classification features and the multiple different grading features to obtain the initial classification rule.
[0023] In the embodiments of the present application, the initial classification rule includes multiple second category features; adjusting the data classes included in the initial classification rule according to the multiple first category features includes:
[0024] Based on the multiple first-category features and the multiple second-category features, the initial classification rule is initially adjusted to obtain a second classification rule; the initial adjustment includes the addition and deletion of second-category features;
[0025] In response to receiving an adjustment instruction, the second classification rule is secondarily adjusted to obtain the target classification rule; the secondary adjustment includes the modification of second-category features.
[0026] In an embodiment of the present application, based on the multiple first-category features and the multiple second-category features, initially adjusting the initial classification rule to obtain a second classification rule includes:
[0027] Comparing the multiple first-category features and the multiple second-category features;
[0028] Screening out extended category features from the multiple first-category features and adding the extended category features to the multiple second-category features; the extended category features exist in the multiple first-category features but do not exist in the multiple second-category features;
[0029] Screening out trimmed category features from the multiple second-category features and deleting the trimmed category features from the multiple second-category features; the trimmed category features exist in the multiple second-category features but do not exist in the multiple first-category features.
[0030] In an embodiment of the present application, by comparing multiple first-category features and multiple second-category features, screening out extended category features and trimmed category features, and making corresponding adjustments to the initial classification rule according to the extended category features and trimmed category features. By introducing new extended category features, the initial classification rule can better adapt to new business scenarios. At the same time, deleting the no-longer-applicable trimmed category features reduces the redundancy and errors in the initial classification rule, improving the accuracy and reliability of data classification.
[0031] In an embodiment of the present application, in response to receiving an adjustment instruction, secondarily adjusting the second classification rule to obtain the target classification rule includes:
[0032] In response to an adjustment instruction from the client, adjusting the corresponding second classification features in the second classification rule according to the adjustment instruction to obtain a third classification rule; the third classification rule includes the adjusted third classification features and multiple unadjusted fourth classification features;
[0033] For any one of the fourth classification features, calculating the similarity between the fourth classification feature and the third classification feature;
[0034] If the similarity is greater than a preset threshold, the recommendation adjustment information including the fourth classification feature and the third classification feature is sent to the client, so that the client displays the recommendation adjustment information to the user and generates a new adjustment instruction according to the user's adjustment operation;
[0035] If the similarities corresponding to the multiple fourth classification features are all less than or equal to the preset threshold, the third classification rule is used as the target classification rule.
[0036] By manually adjusting the classification rule in the embodiments of the present application, it is possible to ensure that the generated classification rule is more suitable for the business scenario of the dataset to be classified, thereby improving the accuracy of data classification.
[0037] Preferably, in the embodiments of the present application, by calculating the similarity between the unadjusted fourth classification feature and the adjusted third classification feature, features that need to be further optimized can be identified. If the similarity is higher than the preset threshold, the system will recommend adjustment information, and the user can generate new adjustment instructions based on this information to further optimize the classification rule. This process significantly improves the accuracy of the classification rule.
[0038] An embodiment of the second aspect of the present application provides a data classification device, and the device includes:
[0039] A dataset acquisition module, configured to acquire a dataset to be classified;
[0040] A first category feature extraction module, configured to extract multiple first category features of the dataset to be classified;
[0041] A target classification rule generation module, configured to adjust the data classes included in the initial classification rule according to the multiple first category features to obtain a target classification rule that matches the dataset to be classified, where the initial classification rule is generated according to multiple business attribute data of the target business to which the dataset to be classified belongs;
[0042] A data classification module, configured to classify the dataset to be classified according to the target classification rule.
[0043] An embodiment of the third aspect of the present application provides a computer device, which includes a memory and a processor, and the memory and the processor are communicatively connected to each other. Computer instructions are stored in the memory, and the processor executes the computer instructions to execute the data classification method described in the first aspect above.
[0044] An embodiment of the fourth aspect of the present application provides a computer-readable storage medium, and computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause a computer to execute the data classification method described in the first aspect above.
[0045] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, throughout the drawings, the same reference numerals are used to represent the same components.
[0047] In the drawings:
[0048] Figure 1 A flowchart showing a data classification method provided by an embodiment of the present application is shown;
[0049] Figure 2 A schematic structural diagram of a data classification device provided by an embodiment of the present application is shown;
[0050] Figure 3 A schematic structural diagram of a computer device provided by an embodiment of the present application is shown;
[0051] Figure 4 A schematic diagram of a storage medium provided by an embodiment of the present application is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] The exemplary embodiments of the present application will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be fully conveyed to those skilled in the art.
[0053] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in the present application should have the ordinary meaning understood by those skilled in the art to which the present application belongs.
[0054] The technical scenarios related to the embodiments of the present application are described below.
[0055] Data classification and grading is a key step in data security management, and its importance is reflected in many aspects. First of all, it helps to ensure data security and privacy. By classifying and grading data, organizations can identify sensitive data, including personal information and business secrets, and take corresponding protection measures to prevent data leakage and unauthorized access. Secondly, data classification and grading can improve data management efficiency. Clear data classification makes data retrieval, processing, and analysis more efficient, helps optimize resource allocation, and reduces redundancy and data silos. In addition, it also helps with compliance management. Many industries and countries have relevant data protection regulations (Data Security Law, Personal Information Protection Law), and conducting data classification and grading can help organizations ensure compliance with these laws and regulations and reduce compliance risks.
[0056] The method of manually formulating classification rules (including classification and grading) has problems such as low efficiency, long implementation cycle, and the effect being extremely dependent on the experience of the implementers. Based on this, the embodiments of the present application can greatly save the manual formulation time, improve the generation efficiency of classification rules, and the generated classification rules do not depend on the experience of the implementers by automatically generating classification rules that conform to the user's business scenario.
[0057] According to the embodiments of the present application, there is provided an embodiment of a data classification method. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0058] In this embodiment, a data classification method is provided. Figure 1 It is a flowchart of the data classification method according to the embodiments of the present application, as Figure 1 shown, and the process includes the following steps:
[0059] Step S101, obtain the data set to be classified.
[0060] In the embodiments of the present application, the data set to be classified can be understood as the business data of the target user with the target business attribute. For example: in the pharmaceutical management business, the data of all the drugs sold by Pharmacy A and the data of all the drugs sold by Pharmacy B are not completely the same; among them, the pharmaceutical management business is used to represent the target business attribute, Pharmacy A or Pharmacy B is used to represent the target user, and the data of all the drugs sold by Pharmacy A or the data of all the drugs sold by Pharmacy B is used to represent the business data of the target user.
[0061] More specifically, the target business attribute includes but is not limited to government affairs business, financial business, and medical business.
[0062] Before step S102, the method further includes steps S201 - S202:
[0063] Step S201, for any one of the multiple service attribute data, extract classification features and grading features from the service attribute data; any service attribute data includes classification information of the target service.
[0064] In the embodiments of the present application, multiple service attribute data belong to the target service of the dataset to be classified. An example is given to illustrate this: when the target service is the medical service, the multiple service attribute data include but are not limited to: "04 - National Standard - GB_T 39725 - 2020 Information Security Technology Health Medical Data Security Guide", "7. Medical Industry Standard WST306—2023 Classification and Coding Rules for Health Information Datasets". Among them, each service attribute data includes classification information of the target service. Through the classification information in the service attribute data (which can be understood as the initial data classification rules, including classification and grading), the service data of the target service can be classified (including classification and grading). The multiple classification processing results obtained based on multiple different service attribute data are different. For example: service attribute data 1 classifies the service data of the target service to obtain classification processing result 1, and service attribute data 2 classifies the service data of the target service to obtain classification processing result 2, and classification processing result 1 is different from classification processing result 2.
[0065] In some specific embodiments, the classification features and grading features can be extracted from the service attribute data through a first classification model. The first classification model includes but is not limited to GPT - Neo, GPT - J, LLaMA, BERT, T5, OPT, Bloom, and StableLM); the classification features can be understood as the rules for classifying service data, and the rules include multiple categories, the definition, application scenarios, and applicable data ranges of each category. The grading features can be understood as the rules for grading service data, and the rules include multiple levels, and the data features of the service data corresponding to each level (such as data sensitivity, data usage restrictions, etc.).
[0066] Step S202, generate the initial classification rules according to the multiple classification features and multiple grading features corresponding to the multiple service attribute data.
[0067] In the embodiments of the present application, corresponding classification features and grading features can be extracted through each service attribute information, so as to obtain multiple classification features and multiple grading features corresponding to multiple service attribute data.
[0068] In some specific embodiments, the above step S202 includes steps S2021 - S2023:
[0069] Step S2021: Screen out the same classification features from the multiple classification features, and remove the same classification features from the multiple classification features to obtain multiple different classification features.
[0070] In the embodiments of the present application, the same classification features include, but are not limited to: categories with the same quantity, the same category definition, the same category application scenario, and the same applicable data range. An example is given to illustrate "removing the same classification features from the multiple classification features to obtain multiple different classification features": Suppose there are seven classification features n1 - n7, where n1, n3, and n5 belong to the same classification features, then n2, n4, n6, and n7 belong to different classification features.
[0071] More specifically, if the definition of a part of the categories of one classification feature is the same as that of another category feature, the definition of this part of the categories can be used as the same classification feature; similarly, if part or all of the categories are the same, part or all of the category definitions are the same, part or all of the category application scenarios are the same, or part or all of the applicable data ranges of the categories are the same, their similarities can be used as the same classification features.
[0072] Step S2022: Screen out the same grading features from the multiple grading features, and remove the same grading features from the multiple grading features to obtain multiple different grading features.
[0073] In the embodiments of the present application, the same grading features include, but are not limited to: the same level and the same data characteristics of the service data corresponding to the level. If the level of one grading feature is the same as that of another grading feature, or the data characteristics of the service data corresponding to the level are the same, the same level and the data characteristics of the service data corresponding to the level can be used as the same grading features.
[0074] Step S2023: Generate the initial classification rule based on the same classification features, the multiple different classification features, the same grading features, and the multiple different grading features.
[0075] In the embodiments of the present application, the above step S2023 includes steps S301 - S302:
[0076] Step S301: Generate the first classification rule according to the same classification features and the same grading features.
[0077] In the embodiments of the present application, by generating the first classification rule through the same classification features and the same grading features, the common parts of the classification and level in the multiple service attribute data of the target service can be fused to obtain a classification and grading rule that basically conforms to the multiple service attribute data as a whole, that is, the first classification rule.
[0078] Step S302, adjust the first classification rule by using the multiple different classification features and the multiple different grading features to obtain the initial classification rule.
[0079] In the embodiment of the present application, the classification differences and grading differences among multiple service attribute data are incorporated into the first classification rule, so as to obtain a classification and grading rule that fully conforms to the multiple service attribute data, that is, the initial classification rule.
[0080] Step S102, extract multiple first category features of the dataset to be classified.
[0081] In the embodiment of the present application, a second classification model can be used to extract multiple first category features of the dataset to be classified, where the preset classification model refers to an open-source general large model with basic language understanding ability, including but not limited to GPT-Neo, GPT-J, LLaMA, BERT, T5, OPT, Bloom, and StableLM.
[0082] In some specific embodiments, the above first classification model and second classification model can be the same or different, and no specific limitation is made here.
[0083] In the embodiment of the present application, the first category features include the element names in the dataset to be classified and the description information of the elements.
[0084] Step S103, adjust the data classes included in the initial classification rule according to the multiple first category features to obtain a target classification rule that matches the dataset to be classified.
[0085] Among them, the initial classification rule is generated according to multiple service attribute data of the target service to which the dataset to be classified belongs. More specifically, the initial classification rule includes multiple second category features, and each second category feature includes a category name and the description information of the category.
[0086] In some specific embodiments, the above step S103 includes steps S1031 - S1032:
[0087] Step S1031, based on the multiple first category features and the multiple second category features, make a primary adjustment to the initial classification rule to obtain a second classification rule.
[0088] Among them, the primary adjustment includes the addition and deletion of second category features.
[0089] In some specific embodiments, the above step S1031 includes steps S401 - S403:
[0090] Step S401: Compare the multiple first-category features and the multiple second-category features.
[0091] In an embodiment of the present application, each element name and the description information of the element in the first-category features can be compared with each category name and the description information of the category in the second-category features.
[0092] Step S402: Screen out the extended-category features from the multiple first-category features, and add the extended-category features to the multiple second-category features.
[0093] In an embodiment of the present application, the extended-category features exist in the multiple first-category features but do not exist in the multiple second-category features. For example, if the element name and the description information of any element in the dataset to be classified do not exist in the multiple category names and the corresponding category description information in the initial classification rule, then the element name and the description information of this element can be used as the extended-category features and extended to the multiple second-category features of the initial classification rule.
[0094] Step S403: Screen out the trimmed-category features from the multiple second-category features, and delete the trimmed-category features from the multiple second-category features.
[0095] In an embodiment of the present application, the trimmed-category features exist in the multiple second-category features but do not exist in the multiple first-category features. For example, if any category name and the corresponding category description information in the initial classification rule do not exist in the dataset to be classified, then this category name and the description information of this category can be used as the trimmed-category features and deleted from the multiple second-category features of the initial classification rule.
[0096] Step S1032: In response to receiving an adjustment instruction, perform a secondary adjustment on the second classification rule to obtain the target classification rule.
[0097] Among them, the secondary adjustment includes the modification of the second-category features.
[0098] In some specific embodiments, the above step S1032 includes steps S501 - S504:
[0099] Step S501: In response to the adjustment instruction of the client, adjust the corresponding second-classification features in the second classification rule according to the adjustment instruction to obtain a third classification rule.
[0100] In the embodiment of the present application, the user inputs the dataset to be classified through the client. The client will match the second classification rule corresponding to the target business to which the dataset to be classified belongs and present it on the client, so that the user can browse the second classification rule on the client. During the process of browsing the second classification rule, the user can issue an adjustment instruction regarding the second classification rule. Thus, in response to the adjustment instruction of the client, the corresponding second classification feature in the second classification rule is adjusted accordingly to obtain the third classification rule. Herein, the corresponding adjustment can be understood as modifying a certain category name and category description information in the second classification rule.
[0101] In the embodiment of the present application, the adjusted third classification rule includes the adjusted third classification feature (for example: the category whose category name and category description information have been modified) and multiple unadjusted fourth classification features (for example: the category whose category name and category description information have not been modified).
[0102] Step S502, for any fourth classification feature, calculate the similarity between the fourth classification feature and the third classification feature.
[0103] In the embodiment of the present application, the similarity between the fourth classification feature and the third classification feature can be calculated in the following way: convert the category name and category description information of the fourth classification feature into a first embedding vector, and convert the category name and category description information of the third classification feature into a second embedding vector; calculate the cosine similarity between the first embedding vector and the second embedding vector.
[0104] Step S503, if the similarity is greater than the preset threshold, send the recommended adjustment information including the fourth classification feature and the third classification feature to the client, so that the client displays the recommended adjustment information to the user and generates a new adjustment instruction according to the user's adjustment operation.
[0105] Step S504, if the multiple similarities corresponding to the multiple fourth classification features are all less than or equal to the preset threshold, use the third classification rule as the target classification rule.
[0106] In the above steps S503 - S504, the preset threshold can be set according to the actual situation and will not be specifically limited herein. When the similarity is greater than the preset threshold, it is determined that the similarity between the fourth classification feature and the third classification feature is relatively high, and the two may be combined. Therefore, it is necessary to send the recommended adjustment information including the fourth classification feature and the third classification feature to the client, so that the user can readjust the third classification rule according to the recommended adjustment information until the multiple similarities corresponding to the multiple fourth classification features are all less than or equal to the preset threshold, and the final target classification rule is obtained.
[0107] Step S104, classify the dataset to be classified according to the target classification rule.
[0108] In the embodiment of the present application, by generating a target classification rule that matches the dataset to be classified and using the target classification rule to classify the dataset to be classified, the time for manually formulating the classification rule can be greatly saved, the formulation efficiency can be improved, and the generated target classification rule does not depend on the experience of the implementer.
[0109] Corresponding to the implementation manner of the above data classification method, the embodiment of the present application further provides a data classification device for executing the data classification method described in the above embodiment. As Figure 2 shown, the data classification device includes:
[0110] A dataset acquisition module, configured to acquire a dataset to be classified;
[0111] A first category feature extraction module, configured to extract a plurality of first category features of the dataset to be classified;
[0112] A target classification rule generation module, configured to adjust the data classes included in the initial classification rule according to the plurality of first category features to obtain a target classification rule that matches the dataset to be classified, where the initial classification rule is generated according to a plurality of service attribute data of the target service to which the dataset to be classified belongs;
[0113] A data classification module, configured to classify the dataset to be classified according to the target classification rule.
[0114] Optionally, the device further includes:
[0115] A hierarchical and classification feature extraction module, configured to extract classification features and hierarchical features for any one of the plurality of service attribute data before adjusting the data classes included in the initial classification rule according to the plurality of first category features; any one of the service attribute data includes classification information of the target service;
[0116] An initial classification rule generation module, configured to generate the initial classification rule according to the plurality of classification features and the plurality of hierarchical features corresponding to the plurality of service attribute data.
[0117] Optionally, the initial classification rule generation module is further configured to screen out the same classification features from the multiple classification features, and remove the same classification features from the multiple classification features to obtain multiple different classification features; screen out the same grading features from the multiple grading features, and remove the same grading features from the multiple grading features to obtain multiple different grading features; generate the initial classification rule based on the same classification features, the multiple different classification features, the same grading features, and the multiple different grading features.
[0118] Optionally, the initial classification rule generation module is further configured to generate a first classification rule according to the same classification features and the same grading features; adjust the first classification rule through the multiple different classification features and the multiple different grading features to obtain the initial classification rule.
[0119] Optionally, the target classification rule generation module is further configured to perform a primary adjustment on the initial classification rule based on the multiple first category features and the multiple second category features to obtain a second classification rule; the primary adjustment includes the addition and deletion of second category features; in response to receiving an adjustment instruction, perform a secondary adjustment on the second classification rule to obtain the target classification rule; the secondary adjustment includes the modification of second category features.
[0120] Optionally, the target classification rule generation module is further configured to compare the multiple first category features and the multiple second category features; screen out the extended category features from the multiple first category features, and add the extended category features to the multiple second category features; the extended category features exist in the multiple first category features but do not exist in the multiple second category features; screen out the trimmed category features from the multiple second category features, and delete the trimmed category features from the multiple second category features; the trimmed category features exist in the multiple second category features but do not exist in the multiple first category features.
[0121] Optionally, the target classification rule generation module is further configured to, in response to an adjustment instruction from the client, adjust the corresponding second classification features in the second classification rule according to the adjustment instruction to obtain a third classification rule; the third classification rule includes the adjusted third classification features and multiple unadjusted fourth classification features; for any one of the fourth classification features, calculate the similarity between the fourth classification feature and the third classification feature; if the similarity is greater than a preset threshold, send the recommended adjustment information including the fourth classification feature and the third classification feature to the client, so that the client displays the recommended adjustment information to the user and generates a new adjustment instruction according to the user's adjustment operation; if the multiple similarities corresponding to the multiple fourth classification features are all less than or equal to the preset threshold, use the third classification rule as the target classification rule.
[0122] The data classification device provided in the foregoing embodiments of the present application and the data classification method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.
[0123] The embodiments of the present application also provide a computer device to execute the above data classification method. Please refer to Figure 3 , which shows a schematic diagram of a computer device provided in some embodiments of the present application. As Figure 3 shown, the computer device 3 includes: a processor 300, a memory 301, a bus 302, and a communication interface 303. The processor 300, the communication interface 303, and the memory 301 are connected through the bus 302; a computer program that can run on the processor 300 is stored in the memory 301, and when the processor 300 runs the computer program, it executes the data classification method provided in the foregoing embodiments of the present application.
[0124] Among them, the memory 301 may include a high-speed random access memory (Random Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 303 (which can be wired or wireless), a communication connection is established between this system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0125] The bus 302 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 301 is used to store a program, and after receiving an execution instruction, the processor 300 executes the program. The data classification method disclosed in the foregoing embodiments can be applied to the processor 300 or implemented by the processor 300.
[0126] The processor 300 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 300 or the instructions in the form of software. The above-mentioned processor 300 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware decoding processor, or executed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 301, and the processor 300 reads the information in the memory 301 and combines its hardware to complete the steps of the above method.
[0127] The computer device provided in the embodiments of the present application and the data classification method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by them.
[0128] The embodiments of the present application also provide a computer-readable storage medium corresponding to the data classification method provided in the foregoing embodiments. Please refer to Figure 4 , which shows that the computer-readable storage medium is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by the processor, it will execute the data classification method provided in any of the foregoing embodiments.
[0129] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.
[0130] The computer-readable storage medium provided by the above embodiments of the present application and the data classification method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.
[0131] It should be noted that:
[0132] In the specification provided herein, a large number of specific details are set forth. However, it is understood that the embodiments of the present application may be practiced without these specific details. In some instances, well-known structures and techniques have not been shown in detail in order not to obscure the understanding of this specification.
[0133] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting the following schematic: that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present application.
[0134] In addition, those skilled in the art will appreciate that although some of the embodiments described herein include certain features included in other embodiments but not others, the combination of features of different embodiments is within the scope of the present application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0135] The above is only the preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A data classification method, characterized in that, The method comprises: Get the dataset to be classified; Extracting a plurality of first category features of the data set to be classified; Adjusting the data classes included in the initial classification rule according to the plurality of first category features to obtain a target classification rule that matches the data set to be classified, the initial classification rule being generated based on a plurality of business attribute data of a target business to which the data set to be classified belongs; The to-be-classified data set is classified according to the target classification rule.
2. The method according to claim 1, characterized in that, Before adjusting the data classes included in the initial classification rule according to the plurality of first category features, the method further includes: For any one of the plurality of business attribute data, extracting classification features and grading features from the business attribute data; any one of the business attribute data includes classification information of the target business; The initial classification rules are generated according to the multiple classification features and the multiple grading features corresponding to the multiple business attribute data.
3. The method according to claim 2, characterized in that Generating the initial classification rule according to the multiple classification features and the multiple grading features corresponding to the multiple business attribute data includes: Screening out identical classification features from the multiple classification features, and removing the identical classification features from the multiple classification features to obtain multiple different classification features; screening out identical grading features from the plurality of grading features, and removing the identical grading features from the plurality of grading features to obtain a plurality of different grading features; The initial classification rule is generated based on the same classification feature, the multiple different classification features, the same ranking feature, and the multiple different ranking features.
4. The method according to claim 3, wherein Generating the initial classification rule based on the same classification feature, the multiple different classification features, the same grading feature, and the multiple different grading features includes: generating a first classification rule according to the same classification feature and the same grading feature; The first classification rule is adjusted by using the multiple different classification features and the multiple different grading features to obtain the initial classification rule.
5. The method according to claim 1, wherein The initial classification rule includes a plurality of second category features; and adjusting the data classes included in the initial classification rule according to the plurality of first category features comprises: Based on the plurality of first category features and the plurality of second category features, performing a primary adjustment on the initial classification rule to obtain a second classification rule; the primary adjustment includes adding and deleting second category features; In response to receiving the adjustment instruction, the second classification rule is adjusted twice to obtain the target classification rule; the secondary adjustment includes modifying the second category feature.
6. The method according to claim 5, wherein Based on the plurality of first category features and the plurality of second category features, the initial classification rule is initially adjusted to obtain a second classification rule, including: comparing the plurality of first category features with the plurality of second category features; Filtering an expanded category feature from the plurality of first category features, and adding the expanded category feature to the plurality of second category features; the expanded category feature exists in the plurality of first category features but does not exist in the plurality of second category features; Screen out the cropping category features from the multiple second category features, and delete the cropping category features from the multiple second category features; the cropping category features exist in the multiple second category features but do not exist in the multiple first category features.
7. The method according to claim 5 or 6, characterized in that, In response to receiving an adjustment instruction, perform a secondary adjustment on the second classification rule to obtain the target classification rule, including: In response to an adjustment instruction from the client, adjust the corresponding second classification features in the second classification rule according to the adjustment instruction to obtain a third classification rule; the third classification rule includes the adjusted third classification features and multiple unadjusted fourth classification features; For any one of the fourth classification features, calculate the similarity between the fourth classification feature and the third classification feature; If the similarity is greater than a preset threshold, send the recommended adjustment information including the fourth classification feature and the third classification feature to the client, so that the client displays the recommended adjustment information to the user and generates a new adjustment instruction according to the user's adjustment operation; If the multiple similarities corresponding to the multiple fourth classification features are all less than or equal to the preset threshold, use the third classification rule as the target classification rule.
8. A data classification device, characterized in that, The apparatus includes: A data set acquisition module, configured to acquire a data set to be classified; A first category feature extraction module, configured to extract multiple first category features of the data set to be classified; A target classification rule generation module, configured to adjust the data classes included in the initial classification rule according to the multiple first category features to obtain a target classification rule that matches the data set to be classified, where the initial classification rule is generated according to multiple service attribute data of the target service to which the data set to be classified belongs; A data classification module, configured to classify the data set to be classified according to the target classification rule.
9. A computer device, characterized in that, Including: A memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the data classification method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause a computer to execute the data classification method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Business flow classification method and apparatus
CN107729952A
Data processing method and device
CN115563523A
Method and device for classifying and grading data
CN115687725A
Data classification and grading method and device, equipment and storage medium
CN116127372A
Interpretation analysis method, device and equipment based on data classification and grading and medium
CN116204823A