A method, system, medium, and device for data category recognition
Determining data categories through multi-dimensional matching has solved the problems of low efficiency and unstable accuracy of data category identification in the prior art, and achieved efficient and accurate data category identification.
Patent Information
- Application Number
- CN202311236166.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-22
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2043-09-22
AI Technical Summary
In the prior art, data category identification requires users to have professional knowledge, low efficiency and unstable accuracy, especially in the case of large data volumes, which is difficult to meet the real-time classification needs.
By determining the column attributes, associated columns and table information of the data to be identified, multi-dimensional matching is performed with the sample data in the preset sample database, the total matching degree is calculated to determine the data category, reducing user technical requirements and improving identification efficiency and accuracy.
There is no need to have an in-depth understanding of the data content, improve data identification efficiency and accuracy through multi-dimensional matching, reduce technical thresholds, and avoid uncertainty caused by data sampling.
Smart Images

Figure CN117195157B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data identification technology, and in particular to a data category identification method, system, medium and device. Background Art
[0002] In existing technologies, data classification is usually achieved by sampling data and classifying it according to user-defined classification rules. However, this approach has some objective disadvantages:
[0003] First, users need to have an in-depth understanding of the data content before they can propose corresponding rules, such as regular matching and specific algorithms. This requires users to have professional knowledge and experience in related fields, which is a high threshold.
[0004] Secondly, when the amount of data is very large, the data sampling process will take a lot of time, which may lead to low efficiency in processing data and may not meet the needs of real-time or fast data classification;
[0005] In addition, generally, all data are not identified, but sampling is used to represent the overall data. The proportion of different sampled data may lead to differences in recognition accuracy, which will bring certain uncertainties.
[0006] Therefore, when performing data category identification, it is necessary to weigh the above factors and seek more efficient and accurate methods to achieve it. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to overcome the shortcomings of the existing technology and provide a data category identification method, system, medium and equipment to solve the problems in the existing technology of high technical requirements, low efficiency and uneven accuracy of data category identification.
[0008] To achieve the above objectives, the present invention adopts the following technical solutions:
[0009] In a first aspect, the present invention provides a method for identifying a data category, the method comprising:
[0010] Determine the data column to be identified and the data table to be identified where the data to be identified is located;
[0011] According to the data column to be identified and the data table to be identified, obtaining the column attribute information, associated column information and table information of the data to be identified;
[0012] Matching the column attribute information of the data to be identified with the column attribute information of sample data of known data categories in a preset sample database, determining a first matching degree between each type of sample data and the data to be identified, and selecting a preset number of sample data with the highest first matching degree as candidate data;
[0013] Determining associated column information of each candidate data in the sample database, and matching the associated column information of the data to be identified with the associated column information of each candidate data, respectively, to determine a second matching degree between each candidate data and the data to be identified;
[0014] Determining association table information of each candidate data in the sample database, and matching the table information of the data to be identified with the association table information of each candidate data, respectively, to determine a third matching degree between each candidate data and the data to be identified;
[0015] Calculate the total matching degree of each candidate data according to the first matching degree, the second matching degree, and the third matching degree of each candidate data;
[0016] The data category of the candidate data with the highest total matching degree is selected as the data category of the data to be identified.
[0017] Furthermore, determining associated column information of each candidate data in the sample database includes:
[0018] In the sample database, determining other data columns in the same data table as each candidate data as associated columns of each candidate data;
[0019] The data information of the associated column is obtained as the associated column information of each candidate data.
[0020] Furthermore, determining association table information for each candidate data in the sample database includes:
[0021] In the sample database, all data tables where each candidate data type is located are determined as associated tables for each candidate data type;
[0022] The table information of the association table is obtained as the association table information of each candidate data.
[0023] Furthermore, the total matching degree of each candidate data is calculated based on the first matching degree, the second matching degree, and the third matching degree of each candidate data, including:
[0024] The first matching degree, the second matching degree, and the third matching degree of each candidate data are summed to obtain the total matching degree of each candidate data.
[0025] Furthermore, the column attribute information includes at least: column name, column type, column length and column comment.
[0026] Furthermore, the associated column information at least includes: an associated column name.
[0027] Furthermore, the table information at least includes: a table name and a table comment.
[0028] In another aspect, the present invention further provides a data category identification system, comprising:
[0029] A module for determining the data column and data table to be identified, used to determine the data column and data table to be identified where the data to be identified is located;
[0030] The module for acquiring information of data to be identified is used to acquire column attribute information, associated column information and table information of the data to be identified based on the data column to be identified and the data table to be identified;
[0031] a first matching module, configured to match the column attribute information of the data to be identified with the column attribute information of sample data of known categories in a preset sample database, determine a first matching degree between each category of sample data and the data to be identified, and select a preset number of sample data with the highest first matching degree as candidate data;
[0032] a second matching module, configured to determine associated column information of each candidate data in the sample database, and match the associated column information of the data to be identified with the associated column information of each candidate data, respectively, to determine a second matching degree between each candidate data and the data to be identified;
[0033] a third matching module, configured to determine association table information of each candidate data in the sample database, and match the table information of the data to be identified with the association table information of each candidate data, respectively, to determine a third matching degree between each candidate data and the data to be identified;
[0034] A total matching degree calculation module, configured to calculate the total matching degree of each candidate data according to the first matching degree, the second matching degree, and the third matching degree of each candidate data;
[0035] The category identification module is used to select the category of the candidate data with the highest total matching degree as the data category of the data to be identified.
[0036] On the other hand, the present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of any one of the above-mentioned data category identification methods.
[0037] On the other hand, the present invention further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of any one of the above-mentioned data category identification methods.
[0038] The present invention adopts the above technical solution, and the beneficial effects that can be achieved include:
[0039] In the present invention, by determining the data column to be identified and the data table to be identified where the data to be identified is located, the column attribute information, associated column information and table information of the data to be identified are obtained, and the data to be identified is matched with the column attribute information of sample data of known categories, the first matching degree of each type of sample data and the data to be identified is determined, and a preset number of sample data with the highest first matching degree is selected as candidate data, and then the associated column information and associated table information of the data to be identified and the candidate data are matched respectively to determine the second matching degree and the third matching degree of the candidate data. Finally, based on the first matching degree, second matching degree and third matching degree of each candidate data, the corresponding total matching degree is calculated, and the data category of the candidate data with the highest total matching degree is selected as the data category of the data to be identified. This method realizes the identification of data categories by utilizing the relevance and matching degree of metadata, so that users do not need to have an in-depth understanding of the data content or formulate identification rules, thereby reducing technical requirements. At the same time, this method does not require data sampling, which greatly improves the efficiency and accuracy of data identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0041] in:
[0042] Figure 1 is a flow chart of a data category identification method in one embodiment;
[0043] Figure 2 is a structural block diagram of a data category identification system in one embodiment;
[0044] Figure 3 FIG. 1 is a structural block diagram of a computer device in one embodiment.
[0045] Explanation of reference numerals: module for determining data columns and data tables to be identified 100 , module for acquiring information of data to be identified 200 , first matching module 300 , second matching module 400 , third matching module 500 , module for calculating total matching degree 600 , and category identification module 700 . DETAILED DESCRIPTION
[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0047] like Figure 1 As shown, in one embodiment, a data category identification method is provided, which specifically includes the following steps:
[0048] S100, determining the data column to be identified and the data table to be identified where the data to be identified is located;
[0049] S200 : Acquire column attribute information, associated column information, and table information of the data to be identified based on the data column to be identified and the data table to be identified.
[0050] During the specific implementation process, for data that needs to be classified, it is necessary to first determine the data column to be identified and the data table to be identified, and then obtain its column attribute information, associated column information and table information. This can be obtained by parsing data files, querying databases or other methods to collect metadata of the data to be identified. Metadata can provide information about the structure and characteristics of the data, providing a basis for subsequent matching and identification.
[0051] The column attribute information may specifically include: the column name, column type, column length, and column comment of the data column to be identified where the data to be identified is located;
[0052] The associated column information may specifically include: column names of other data columns in the to-be-identified data table where the to-be-identified data is located;
[0053] The table information may specifically include: the table name and table comments of the data table to be identified where the data to be identified is located.
[0054] S300. Match the column attribute information of the data to be identified with the column attribute information of sample data of known data categories in a preset sample database, determine a first matching degree between each type of sample data and the data to be identified, and select a preset number of sample data with the highest first matching degree as candidate data.
[0055] During the specific implementation process, a sample database can be pre-built by collecting and organizing a certain number of sample data of known data categories. The sample data should be representative and able to cover all categories of data to be identified. Specifically, it can come from existing data sets, databases or other data sources.
[0056] Then, for each type of sample data, first extract the column attribute information of the data column, that is, the column name, column type, column length, and column comment. For example, for the sample data of the data category "gender", the column attribute information keywords may include:
[0057] Column names: SEX, GENDER
[0058] Column type and length: CHAR(1), ENUM('M','F'), TINYINT(1)
[0059] Column annotations: "sex", "male", "female", "sex", "gender".
[0060] Then, the column attribute information of the data to be identified is matched with the column attribute information of each type of sample data to determine the first matching degree between each type of sample data and the data to be identified. This process can be implemented using similarity calculation or other matching algorithms, so as to accurately judge the similarity between the data to be identified and the sample data, thereby improving the accuracy of identification.
[0061] The preset number of sample data with the highest first matching degree is selected as candidate data. This is because the higher the first matching degree between the sample data and the data to be identified, the greater the possibility that the two belong to the same data category. In other words, the data category of the sample data is more likely to be the data category of the data to be identified. Therefore, in this embodiment, the sample data and the data to be identified can be sorted by the first matching degree, and the preset number of sample data with high first matching degrees can be selected as candidate data. Subsequently, further screening of the data categories of these candidate data is required to determine the data category of the data to be identified. Specifically, to save computing resources and improve recognition efficiency, the preset number can be set to 3, that is, the sample data of the three categories with the highest first matching degrees are selected as candidate data.
[0062] S400 , determining associated column information of each candidate data in the sample database, and matching the associated column information of the data to be identified with the associated column information of each candidate data respectively, to determine a second matching degree between each candidate data and the data to be identified.
[0063] Furthermore, in some embodiments, determining associated column information for each candidate data type in the sample database includes:
[0064] In the sample database, determining other data columns in the same data table as each candidate data as associated columns of each candidate data;
[0065] The data information of the associated column is obtained as the associated column information of each candidate data.
[0066] During specific implementations, columns of data in the same data table typically have the same data category. Therefore, to further determine which candidate data's data category is closer to the true data category of the data to be identified, this embodiment, after determining the first degree of match between the candidate data and the data to be identified using column attribute information, can also determine the second degree of match between the candidate data and the data to be identified using their respective associated column information. Furthermore, based on the associated column dimension, the matching between the candidate data and the data to be identified can be further determined, providing a basis for subsequent comprehensive determination of the data category of the data to be identified.
[0067] For example: for candidate data of the data category "gender", there may be other data columns such as "name", "age", "birthday", "marital", "education", "occupation" and so on in the same personal information data table as "gender". Then, other data columns such as "name", "age", "birthday", "marital", "education", "occupation" and so on can be used as associated columns of "gender", and the associated column information such as the corresponding column name can be obtained, and then matched with the associated column information of the data to be identified to determine the second matching degree between the candidate data and the data to be identified.
[0068] S500: Determine association table information of each candidate data in the sample database, and match the table information of the data to be identified with the association table information of each candidate data respectively, to determine a third matching degree between each candidate data and the data to be identified.
[0069] Furthermore, in some embodiments, determining association table information for each candidate data type in the sample database includes:
[0070] In the sample database, all data tables where each candidate data type is located are determined as associated tables for each candidate data type;
[0071] The table information of the association table is obtained as the association table information of each candidate data.
[0072] During specific implementations, data of the same data category typically has a high degree of matching across the various data tables it resides in. Therefore, similarly, to further determine which candidate data's data category is closer to the true data category of the data to be identified, this embodiment, after determining the first and second matching degrees between the candidate data itself and the data to be identified, can also determine the third matching degree between the candidate data and the data to be identified using table information from their respective tables. Furthermore, based on the dimensions of the associated tables, the matching between the candidate data and the data to be identified can be further determined, laying the foundation for the subsequent comprehensive determination of the data category to be identified.
[0073] For example, for candidate data of the data category "gender", it may exist in data tables such as "user", "student", "employee", "patient", "order", and "customer". Then, data tables such as "user", "student", "employee", "patient", "order", and "customer" can be used as associated tables of "gender", and the corresponding associated table information such as table name or table comment can be obtained, and then matched with the table information of the data to be identified to determine the third matching degree between the candidate data and the data to be identified.
[0074] S600: Calculate the total matching degree of each candidate data according to the first matching degree, the second matching degree, and the third matching degree of each candidate data.
[0075] Furthermore, in some embodiments, calculating the total matching degree of each candidate data according to the first matching degree, the second matching degree, and the third matching degree of each candidate data includes:
[0076] The first matching degree, the second matching degree, and the third matching degree of each candidate data are summed to obtain the total matching degree of each candidate data.
[0077] S700: Select the data category of the candidate data with the highest total matching degree as the data category of the data to be identified.
[0078] During the specific implementation process, the total matching degree of the data to be identified and the candidate data can be determined based on the matching of multiple dimensional information such as column attribute information, associated column information, and table information between the data to be identified and the sample candidate data. The total matching degree serves as an indicator for evaluating the matching degree between the sample data and the data to be identified. The higher the value, the higher the matching degree, that is, the more likely the data category of the candidate data is to be the true data category of the data to be identified. Therefore, the data category of the candidate data with the highest total matching degree can be selected as the final data category of the data to be identified, so that this embodiment can comprehensively consider the correlation of multiple dimensional information of the data to ensure the accuracy of data category identification.
[0079] In addition, in addition to summing the first, second, and third matching degrees of each candidate data, the method of calculating the total matching degree can also assign different weights to the first, second, and third matching degrees based on the contribution of the matching of column attribute information, associated column information, and table information to the final data category identification according to actual needs, thereby performing a weighted summation of each matching degree to obtain the total matching degree of each candidate data, thereby improving the accuracy of the final data category identification from another perspective.
[0080] The data category identification method described in the above embodiment determines the data column to be identified and the data table to be identified where the data to be identified is located, obtains the column attribute information, associated column information and table information of the data to be identified, and matches the data to be identified with the column attribute information of the sample data of known categories, determines the first matching degree of each type of sample data with the data to be identified, and selects a preset number of sample data with the highest first matching degree as candidate data, then matches the associated column information and associated table information of the data to be identified with the candidate data respectively, determines the second matching degree and the third matching degree of the candidate data, and finally calculates the corresponding total matching degree based on the first matching degree, second matching degree and third matching degree of each candidate data, and selects the data category of the candidate data with the highest total matching degree as the data category of the data to be identified. This method realizes the identification of data categories by utilizing the relevance and matching degree of metadata, so that users do not need to have an in-depth understanding of the data content or formulate identification rules, thereby reducing technical requirements. At the same time, this method does not require data sampling, which greatly improves the efficiency and accuracy of data identification.
[0081] like Figure 2 As shown, in another embodiment, a data category identification system is provided, the system comprising:
[0082] The module 100 for determining the data column and data table to be identified is used to determine the data column and data table to be identified where the data to be identified is located;
[0083] The to-be-identified data information acquisition module 200 is configured to acquire column attribute information, associated column information, and table information of the to-be-identified data based on the to-be-identified data column and the to-be-identified data table;
[0084] A first matching module 300 is configured to match the column attribute information of the data to be identified with the column attribute information of sample data of known categories in a preset sample database, determine a first matching degree between each category of sample data and the data to be identified, and select a preset number of sample data with the highest first matching degree as candidate data;
[0085] The second matching module 400 is configured to determine the associated column information of each candidate data in the sample database, and match the associated column information of the data to be identified with the associated column information of each candidate data, respectively, to determine a second matching degree between each candidate data and the data to be identified;
[0086] The third matching module 500 is configured to determine association table information of each candidate data in the sample database, and match the table information of the data to be identified with the association table information of each candidate data, respectively, to determine a third matching degree between each candidate data and the data to be identified;
[0087] A total matching degree calculation module 600 is used to calculate the total matching degree of each candidate data according to the first matching degree, the second matching degree and the third matching degree of each candidate data;
[0088] The category identification module 700 is configured to select the category of the candidate data with the highest total matching degree as the data category of the data to be identified.
[0089] It should be noted that for other corresponding descriptions of the functional modules involved in the test report generating device provided in this embodiment, reference can be made to the corresponding descriptions of the methods in the above embodiments, which will not be repeated here.
[0090] Figure 3 FIG1 shows an internal structure diagram of a computer device in an embodiment. The computer device can be a terminal or a server. Figure 3 As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor may implement the data category identification method described in any of the above embodiments. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor may implement the data category identification method described in any of the above embodiments. Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0091] In one embodiment, a computer device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the data category identification method described in any of the above embodiments.
[0092] In one embodiment, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the processor performs the steps of the data category identification method described in any of the above embodiments.
[0093] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0094] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0095] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A data category identification method, characterized in that: The method comprises: Determine the data column to be identified and the data table to be identified where the data to be identified is located; According to the data column to be identified and the data table to be identified, obtaining the column attribute information, associated column information and table information of the data to be identified; Matching the column attribute information of the data to be identified with the column attribute information of sample data of known data categories in a preset sample database, determining a first matching degree between each type of sample data and the data to be identified, and selecting a preset number of sample data with the highest first matching degree as candidate data; Determining associated column information of each candidate data in the sample database, and matching the associated column information of the data to be identified with the associated column information of each candidate data, respectively, to determine a second matching degree between each candidate data and the data to be identified; Determining association table information of each candidate data in the sample database, and matching the table information of the data to be identified with the association table information of each candidate data, respectively, to determine a third matching degree between each candidate data and the data to be identified; Calculate the total matching degree of each candidate data according to the first matching degree, the second matching degree, and the third matching degree of each candidate data; The data category of the candidate data with the highest total matching degree is selected as the data category of the data to be identified.
2. The method according to claim 1, characterized in that Determining associated column information for each candidate data in the sample database includes: In the sample database, determining other data columns in the same data table as each candidate data as associated columns of each candidate data; The data information of the associated column is obtained as the associated column information of each candidate data.
3. The method according to claim 1, characterized in that Determining association table information for each candidate data in the sample database includes: In the sample database, all data tables where each candidate data type is located are determined as associated tables for each candidate data type; The table information of the association table is obtained as the association table information of each candidate data.
4. The method according to claim 1, wherein Calculate the total matching degree of each candidate data according to the first matching degree, the second matching degree, and the third matching degree of each candidate data, including: The first matching degree, the second matching degree, and the third matching degree of each candidate data are summed to obtain the total matching degree of each candidate data.
5. The method according to any one of claims 1 to 4, characterized in that The column attribute information includes at least: column name, column type, column length and column comment.
6. The method according to any one of claims 1 to 4, characterized in that The associated column information at least includes: an associated column name.
7. The method according to any one of claims 1 to 4, characterized in that The table information at least includes: a table name and a table comment.
8. A data category identification system, characterized in that: The system comprises: A module for determining the data column and data table to be identified, used to determine the data column and data table to be identified where the data to be identified is located; The module for acquiring information of data to be identified is used to acquire column attribute information, associated column information and table information of the data to be identified based on the data column to be identified and the data table to be identified; a first matching module, configured to match the column attribute information of the data to be identified with the column attribute information of sample data of known categories in a preset sample database, determine a first matching degree between each category of sample data and the data to be identified, and select a preset number of sample data with the highest first matching degree as candidate data; a second matching module, configured to determine associated column information of each candidate data in the sample database, and match the associated column information of the data to be identified with the associated column information of each candidate data, respectively, to determine a second matching degree between each candidate data and the data to be identified; a third matching module, configured to determine association table information of each candidate data in the sample database, and match the table information of the data to be identified with the association table information of each candidate data, respectively, to determine a third matching degree between each candidate data and the data to be identified; A total matching degree calculation module, configured to calculate the total matching degree of each candidate data according to the first matching degree, the second matching degree, and the third matching degree of each candidate data; The category identification module is used to select the category of the candidate data with the highest total matching degree as the data category of the data to be identified.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 7.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the computer program is executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data query method and device, equipment and storage medium
CN113946721A
Field name-free well file field name automatic matching method and system, medium and computer equipment
CN114896352A