Semantic Recognition-Based Data Resource Classification and Organization Method and System
By constructing the classification system and feature vectors of the bank data warehouse and calculating the cosine similarity, the cross-system fragmentation problem of data tables in the bank data warehouse is solved, and the rapid data table classification and query are realized, which improves the data checking rate and accuracy rate.
Patent Information
- Application Number
- CN202111446841.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-11-30
AI Technical Summary
There are problems in bank data warehouses such as cross-system, fragmentation, different standards, scattered business, and lengthy processes, which makes it difficult for business personnel to sort out the themes and semantic relationships between individual tables, and cannot quickly lock in the data table scope corresponding to the concept of demand, resulting in low data check rate and accuracy rate.
Construct a data resource classification organization method based on semantic recognition. By sorting out data topics and business distribution, building a classification system, extracting feature vectors, calculating cosine similarity, dividing the classification number of the data table, providing keyword and similarity search functions, realizing rapid classification and querying of the data table.
It realizes the concept of integrating the content of data tables across systems, quickly locks the semantic closeness and alienation relationship of data tables, improves the data search rate and accuracy rate, lowers the threshold for data use, and improves the data extraction efficiency.
Smart Images

Figure CN114064821B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data table classification, and specifically relates to a method and system for classifying and organizing data resources based on semantic recognition. Background Art
[0002] With the deepening of digital technology and information applications, the data warehouse systems of banks often converge data such as customers, business, and finance of the bank's main systems, providing data support and services for the bank's daily operation analysis, marketing, risk control, financial analysis, internal audit, and regulatory reporting. The data assets in the banking industry are becoming increasingly complex. Problems such as cross-system, fragmented, different standards, scattered business, and long processes are severe problems faced by the integrated application of banking data assets, which hinder the analysis and application of data and the realization of its value. After years of information system construction and the efforts of technical personnel, the data assets have basically achieved unified storage management in terms of form, lineage, and system, and the business domain classification of metadata has been realized. However, there is still a lack of in-depth organizational integration for the content relationship between single tables and the semantic association at the index level.
[0003] When faced with a large number of original data tables, business personnel are difficult to sort out the theme and semantic relationship between single tables, making it impossible to lock the corresponding keywords according to the demand concept, lock the corresponding data table range, and return the main asset content. As a result, business personnel face problems such as not knowing what data there is, not knowing what data should be found, and not knowing where to extract data when analyzing and using. Facing a large amount of data assets, they can only rely on business personnel to temporarily collect relevant data, and it is impossible to control the recall rate and precision rate of data collection, resulting in obvious problems such as omission of relevant data and difficulty in ensuring data accuracy. Summary of the Invention
[0004] The present invention aims to provide a method for classifying and organizing data resources based on semantic recognition, which can calculate the theme semantic relationship of assets according to the metadata content of data tables, quickly lock the semantic space and data asset range corresponding to the demand concept, solve the problem of cross-system, cross-business, and fragmented data asset organization and query, and help business personnel find data, find the right data, and improve the data extraction efficiency.
[0005] The technical solution provided by the present invention is as follows: A method for classifying and organizing data resources based on semantic recognition, including:
[0006] S1: According to the data of the bank's data warehouse system, sort out the theme and business distribution of the data, and construct a classification system, where the classification system includes each level of categories and the corresponding data tables;
[0007] S2: Extract features from each level of categories and data tables and convert them into feature vectors, and construct a category feature vector and a data table feature vector respectively;
[0008] S3: Calculate the cosine similarity between the eigenvectors of each category and the eigenvector of the data table, delimit the similarity threshold, and determine the category to which the data table belongs according to the threshold. Similarly, calculate the cosine similarity between data tables, delimit the similarity threshold, and determine similar data tables according to the threshold.
[0009] S4: Classify the data resources according to the classification system, divide the data classification numbers corresponding to the data tables, and store the classification numbers to which each data table belongs and the similar data tables.
[0010] S5: Organize and utilize the data resources, provide the function of expanding the data tables belonging to the classification number, provide the function of retrieving data tables according to keywords, and provide the function of expanding similar data tables according to the similarity of data tables.
[0011] The working principle and advantages of the present invention are as follows:
[0012] According to the data model of the bank data warehouse system, a classification system is constructed in combination with the logical model of the data warehouse system. It realizes the integration of the content concepts of data tables across systems, sorts out the theme attribution of data tables, and can quickly and automatically identify the category to which a data table belongs.
[0013] Feature extraction is performed on each level of category and data table and converted into eigenvectors, and an eigenvector of the category and an eigenvector of the data table are respectively constructed; the cosine similarity between the eigenvectors of each category and the eigenvector of the data table is calculated, the similarity threshold is delimited, and the category to which the data table belongs is determined according to the threshold. Similarly, the cosine similarity between data tables is calculated, the similarity threshold is delimited, and similar data tables are determined according to the threshold. It realizes sorting out the theme similarity between data tables from the semantic concept level, can sort out the semantic affinity relationship between data tables, and is convenient for quickly locking similar data tables of a certain data table.
[0014] The data resources are classified according to the classification system, the data classification numbers corresponding to the data tables are divided, and the classification numbers to which each data table belongs and the similar data tables are stored. The numerous fragmented data tables are integrated based on their theme classification relationships, and the theme distribution of the entire dispersed data asset system is quickly counted, which is convenient for the management and control of data assets.
[0015] The data resources are organized and utilized, providing the function of expanding the data tables belonging to the classification number, providing the function of retrieving data tables according to keywords, and providing the function of expanding similar data tables according to the similarity of data tables. It realizes an integrated data resource organization system, can effectively improve the query and retrieval efficiency of data tables. Users can search for data tables under the required concept in various ways such as browsing by directory, retrieving by keywords, retrieving by classification number, and expanding specific data tables, which is convenient for users to quickly find the required data assets, improves the recall rate and precision rate of data search, and reduces the data usage threshold in the banking industry.
[0016] The method of the present invention solves the problem of cross-system, cross-business, and fragmented data asset organization and query through the above steps, helping business personnel find data, accurately locate data, and improve data extraction efficiency.
[0017] Further, it is characterized in that: the classification system is a three-level theme classification system.
[0018] Based on the traditional bank data warehouse, the method of the present invention first conducts theme sorting on all existing data, and then constructs a three-level classification system in combination with the logical model of the data warehouse system. The secondary categories and tertiary categories are divided on the basis of the first-level major categories, and this classification system is suitable for the current bank data management mode.
[0019] Further, it is characterized in that: the major themes in the three-level theme classification system include one or more of system names, system business scopes, data table names, data field names, and code values, and the first-level categories in the three-level theme classification system include one or more of customers, agreements, transactions, institutions, products, assets, and general.
[0020] For all existing data, first divide the theme to which the data belongs according to the system name, system business scope, data table name, data field name, and code value, then divide the first-level categories such as customers, agreements, transactions, institutions, products, assets, and general according to different uses, and then refine the classification to the secondary and tertiary categories.
[0021] Further, S2 includes:
[0022] S2-1: Based on the sample wide table data of each level of category, extract the keywords of its data table to construct a string set, perform word processing on the string set to convert the string into word sub-items, then extract semantic features, obtain word vectors, and construct a category feature vector
[0023] S2-2: Based on the keywords of the data table's attribution, extract and construct a string set, perform word processing on the string set, then extract semantic features, obtain word vectors, and construct a data table feature vector.
[0024] The category feature vector is constructed based on the sample wide table data of each level of category, and the data table feature vector is constructed based on the keywords of the data table's attribution. The specific process includes three steps: extracting keywords, word processing, and extracting semantic features.
[0025] Further, the word processing methods in S2-1 and S2-2 include one or more of word segmentation, tokenization, stop word processing, and TF-IDF word frequency calculation.
[0026] Improve the statistical machine understanding language model through the above natural language processing methods.
[0027] Further, in S2-1 and S2-2, semantic feature extraction is performed through the Word2vector model.
[0028] By using the Word2vector model to extract semantic features, each word in natural language is represented as a short vector with a unified meaning and dimension, enabling more generalized analysis of words and sentences.
[0029] Further, it is characterized in that: S3 includes:
[0030] S3-1: Calculate the cosine similarity between the feature vectors of each tertiary category and the feature vector of the data table, calculate the distance between the category feature vector and the data table feature vector based on the cosine similarity, use this distance as the similarity between the data table and the category, define a similarity threshold, determine the candidate categories to which the data table belongs according to the threshold, and perform clustering verification on the candidate categories;
[0031] S3-2: Calculate the cosine similarity between data tables, calculate the distance between the data table feature vectors based on the cosine similarity, use this distance as the similarity between data tables, divide the similarity threshold, and determine similar data tables according to the threshold.
[0032] For the feature vectors of the data table and the feature vectors of each tertiary category, calculate the distance between the data table feature vector and the category feature vector according to the cosine similarity (the cosine value of the angle between vectors), use this distance as the similarity between the data table and the category, and the smaller the matching distance, the higher the similarity. Similarly, the similarity between data tables can be calculated, and the candidate categories to which the data table belongs and similar data tables are determined according to the threshold. To improve the accuracy of categories, clustering verification is performed on the candidate categories.
[0033] Further, the process of performing clustering verification on the candidate categories in S3-1 is as follows: Cluster the data table and the secondary category to divide the range of the secondary category to which it belongs, exclude the candidate tertiary categories that do not meet the requirements, and select the tertiary category to which it belongs according to the similarity, and divide the tertiary category number to which the data table belongs.
[0034] Perform clustering verification on the candidate categories through the kmeans clustering algorithm, classify individuals or objects according to the degree of similarity, so that the similarity between elements in the same class is stronger than that of elements in other classes. Maximize the homogeneity of elements between classes and the heterogeneity of elements between classes.
[0035] Further, S4 is: Classify the data resources according to the classification system, divide the data classification numbers corresponding to the full amount of data tables, organize the corresponding data tables according to the tertiary classification number system, store the tertiary classification numbers to which each data table belongs and similar data tables, and when the metadata structure of the data table changes or a new data table is created, recalculate and update the classification numbers to which the data table belongs.
[0036] The data resources are hierarchically organized according to the subject classification system, providing the ability to expand the subject concepts step by step according to the classification numbers and browse the data table resources belonging to them. On the basis of providing general keyword retrieval, data table retrieval can be provided according to the similarity between data tables, return the similar data tables according to the input data table, and quickly lock the corresponding data table range according to the classification number of the data table. When the metadata structure of the data table changes or a new data table is created, recalculate and update the classification number to which the data table belongs to maintain the accuracy and efficiency of data classification.
[0037] The present invention also provides a data resource classification and organization system based on semantic recognition, characterized in that: the system adopts any one of the above-mentioned data resource classification and organization methods based on semantic recognition. Brief Description of the Drawings
[0038] Figure 1 It is a logical block diagram of an embodiment of the data resource classification and organization method based on semantic recognition of the present invention. Detailed Embodiment
[0039] Embodiment:
[0040] As Figure 1 shown, the present embodiment discloses a data resource classification and organization method based on semantic recognition, specifically including the following steps:
[0041] S1: According to the data of the bank data warehouse system, sort out the theme and business distribution of the data, and construct a classification system, which includes categories at all levels and their corresponding data tables. On the basis of the traditional bank data warehouse, use all the stock data to first sort out the theme. The themes to which the data belongs include system name, system business scope, data table name, data field name, code value, etc. After dividing the major themes, divide the three-level categories according to the use of the data, and cluster on the basis of the first-level major categories and combine expert experience to divide the second-level and third-level categories. The data classification system has a total of 7 first-level categories, 41 second-level categories, and 85 third-level categories, including customers, agreements, transactions, institutions, products, assets, and general.
[0042] S2-1: Based on the sample wide table data of each level of category, extract the keywords of its data table to construct a string set, perform word processing on the string set to convert the string into word items, then extract the semantic features, obtain word vectors, and construct a category feature vector. Extract keywords such as the system name, table name, field name, and code value of the data table to construct a string set, perform word segmentation, word cutting, stop word processing, TF-IDF word frequency calculation and other word processing on the string set to convert the string into word items, and then use Word2vector to extract the semantic features to obtain word vectors, thereby constructing the feature vectors of each level of category.
[0043] S2-2: Extract and construct a string set based on the keywords of the data table's attribution. Perform word processing on the string set, then extract semantic features to obtain word vectors, and construct the data table feature vector. Extract keywords such as the system name, table name, field name, and code value of the data table's attribution, perform word segmentation, tokenization, stop word processing, TF-IDF word frequency calculation, etc. on the words to convert the strings into word items, and then use Word2vector to extract semantic features to obtain the feature vector of the data table.
[0044] S3-1: Calculate the cosine similarity between each three-level category feature vector and the data table feature vector, calculate the distance between the category feature vector and the data table feature vector based on the cosine similarity, use this distance as the similarity between the data table and the category, set a similarity threshold, and determine the candidate categories to which the data table belongs based on the threshold, and perform clustering verification on the candidate categories. Calculate the distance between the data table feature vector and the category feature vector according to the cosine similarity (the cosine value of the angle between the vectors). The smaller the matching distance, the higher the similarity. Find the 5 three-level categories with the highest similarity to the current data table as candidate categories according to the similarity.
[0045] Cluster the current data table with the data tables of each second-level category to find the second-level category to which the data table belongs. First, perform kmeans clustering on the data table samples under each second-level category, with K set to 1 to obtain the initial class center vectors of each second-level category. Then perform second-level category clustering on all data tables, with K set to the number of second-level categories and the initial class center vectors of each second-level category set as the initial class centers. Obtain the category to which each data table belongs according to the clustering results and update the class center vectors of each second-level category. Verify whether the second-level category corresponding to the three-level category in the previous step is consistent with the clustered second-level category, and delete the inconsistent three-level categories. Return the remaining three-level categories. Among the candidate categories of the data table, if the two three-level categories with the highest similarity are selected, if the similarity data of both is above 0.4, accept the candidate categories; if the difference between the two is above 0.1, return the three-level class number with the highest similarity. If the difference between the two is within 0.1 and they are not adjacent three-level categories, then return the highest three-level class number as the belonging class number, and the other as the see also class number.
[0046] S3-2: Calculate the cosine similarity between data tables, calculate the distance between the data table feature vectors based on the cosine similarity, use this distance as the similarity between the data tables, divide the similarity threshold, and determine similar data tables according to the threshold. Set the table names of the top three data tables with a stored similarity above 0.6 as the similar data tables of this data table.
[0047] S4: Classify data resources according to the classification system, divide the data classification numbers corresponding to the full data table, organize the corresponding data tables according to the three-level classification number system, store the three-level classification numbers and similar data tables of each data table, and recalculate and update the classification numbers of the data tables when the metadata structure of the data table changes or a new data table is created
[0048] S5: Organizes and utilizes data resources, providing the ability to expand data tables by classification number, search data tables by keyword, and expand similar data tables based on table similarity. Data is expanded hierarchically according to the category system to find the corresponding data asset range, or the search is performed by entering the data table keyword. When the user selects a data table, the data table range under the corresponding concept can be locked based on the table's classification number or reference category number. Related data tables can also be expanded based on similar tables to quickly locate the data asset range corresponding to the required concept, improving recall and query efficiency.
[0049] This embodiment also discloses a system that is compatible with the above-mentioned data resource classification method based on semantic recognition, and the system uses the above-mentioned method.
[0050] The above are only embodiments of the present invention. Common knowledge such as the known specific structures and characteristics in the scheme are not described in detail here. Ordinary technicians in the field are aware of all common technical knowledge in the technical field of the invention before the application date or priority date, can obtain all existing technologies in the field, and have the ability to apply conventional experimental means before that date. Ordinary technicians in the field can improve and implement this scheme based on their own abilities under the inspiration obtained by this application. Some typical known structures or known methods should not become obstacles for ordinary technicians in the field to implement this application. It should be pointed out that for those skilled in the art, without departing from the structure of the present invention, several variations and improvements can be made, which should also be regarded as the scope of protection of the present invention. These will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.
Claims
1. A method for classifying and organizing data resources based on semantic recognition, characterized in that, Including: S1: Based on the data in the bank data warehouse system, sort out the themes and business distributions of the data, and construct a classification system, where the classification system includes categories at all levels and their corresponding data tables; S2: Extract features from categories and data tables at all levels and transform them into feature vectors, and construct category feature vectors and data table feature vectors respectively; S3: Calculate the cosine similarity between each category feature vector and data table feature vector, delimit a similarity threshold, determine the category to which the data table belongs according to the threshold. Similarly, calculate the cosine similarity between data tables, divide the similarity threshold, and determine similar data tables according to the threshold; S4: Classify the data resources according to the classification system, divide the data classification numbers corresponding to the data tables, and store the classification numbers to which each data table belongs and similar data tables; S5: Organize and utilize the data resources, provide the function of expanding the data tables to which they belong according to the classification number, provide the function of retrieving data tables according to keywords, and provide the function of expanding similar data tables according to the similarity of data tables; The classification system is a three-level theme classification system; The S2 includes: S2-1: Based on the sample wide table data of categories at all levels, extract the keywords of its data tables to construct a string set, perform word processing on the string set to convert the strings into word items, then extract semantic features, obtain word vectors, and construct category feature vectors; S2-2: Based on the keywords to which the data tables belong, extract and construct a string set, perform word processing on the string set, then extract semantic features, obtain word vectors, and construct data table feature vectors; The S3 includes: S3-1: Calculate the cosine similarity between each three-level category feature vector and data table feature vector, calculate the distance between the category feature vector and the data table feature vector according to the cosine similarity, use this distance as the similarity between the data table and the category, delimit a similarity threshold, determine the candidate categories to which the data table belongs according to the threshold, and perform clustering verification on the candidate categories; S3-2: Calculate the cosine similarity between data tables, calculate the distance between data table feature vectors according to the cosine similarity, use this distance as the similarity between data tables, divide the similarity threshold, and determine similar data tables according to the threshold.
2. The method for classifying and organizing data resources based on semantic recognition according to claim 1, characterized in that: In the three-level theme classification system, the major themes include one or more of system name, system business scope, data table name, data field name, and code value. In the three-level theme classification system, the first-level categories include one or more of customer, agreement, transaction, institution, product, asset, and general.
3. The method for classifying and organizing data resources based on semantic recognition according to claim 1, wherein: The word processing methods in S2-1 and S2-2 include one or more of word segmentation, tokenization, stop word processing, and TF-IDF word frequency calculation.
4. The method for classifying and organizing data resources based on semantic recognition according to claim 1, characterized in that: In S2-1 and S2-2, semantic features are extracted through the Word2vector model.
5. The method for classifying and organizing data resources based on semantic recognition according to claim 1, wherein: The process of clustering verification on the candidate categories in S3-1 is: cluster the data tables and the second-level categories to divide the range of the second-level categories to which they belong, exclude the candidate third-level categories that do not conform, and select the third-level categories to which they belong according to the similarity, and divide the third-level category numbers to which the data tables belong.
6. The method for classifying and organizing data resources based on semantic recognition according to claim 1, characterized in that: S4 is as follows: Classify the data resources according to the classification system, assign data classification numbers corresponding to the full-scale data tables, organize the corresponding data tables according to the three-level classification number system, store the three-level classification numbers to which each data table belongs and the similar data tables, and recalculate and update the classification numbers to which the data tables belong when the metadata structure of the data tables changes or new data tables are created.
7. A data resource classification and organization system based on semantic recognition, characterized in that: It is used to run the data resource classification and organization method based on semantic recognition according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method and device for classifying and mapping data tables of HIS systems
CN108595657A
Data table similarity determination method and device
CN112597149A