A method and device for constructing a banking data asset catalog

By constructing a banking data asset catalog and using a keyword extraction model and a named entity recognition model, the mapping relationship between business attributes and technical attributes is clarified. This solves the problems of low cross-departmental communication efficiency and insufficient data sharing in the construction of the banking data asset catalog, and realizes efficient data asset catalog construction and cross-industry data sharing.

CN115964652BActive Publication Date: 2026-04-21THE PEOPLES BANK OF CHINA NAT CLEARING CENT
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
THE PEOPLES BANK OF CHINA NAT CLEARING CENT
Filing Date
2022-10-11
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing technologies, the construction of banking data asset catalogs suffers from problems such as low efficiency in cross-departmental and cross-system communication, inability to meet the requirements of external units when sharing data, and the inability of the data upstream and downstream lineage relationships to encompass the different levels of business logic relationships between all data assets.

Method used

Based on the business requirements of the banking subsystem, a topic word extraction model and a named entity recognition model are used to construct a relationship diagram of topic domain hierarchical classification, business attribute classification and technical attribute classification. The data asset catalog is recorded through a triple structure to clarify the mapping relationship between business attributes and technical attributes.

Benefits of technology

It improves the efficiency of building data asset catalogs, enhances cross-departmental and cross-system communication efficiency, ensures the accuracy and completeness of data sharing, intuitively displays the relationships between data, and facilitates cross-industry data sharing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115964652B_ABST
    Figure CN115964652B_ABST
Patent Text Reader

Abstract

The application provides a bank data asset catalog construction method and device. The bank data asset catalog construction method comprises the following steps: obtaining a subject domain hierarchical classification based on relevant documents of business requirements of a bank subsystem; constructing a business attribute classification based on the business requirements of the bank subsystem, the relevant documents of the business requirements and the subject domain hierarchical classification; constructing a technical attribute classification relationship graph according to the architecture and functions of the bank subsystem, wherein the technical attribute classification relationship graph comprises metadata relationships within the bank subsystem and cross-system metadata relationships; and adding the business attribute classification of the bank subsystem to the technical attribute classification relationship graph. By using the application, the problems in the prior art that the communication efficiency of staffs across departments and systems is low, data cannot meet the requirements when sharing data with external units, and the blood relationship between upstream and downstream data cannot contain different levels of business logic relationships among all data assets can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data asset catalog architecture, specifically a method and apparatus for constructing a data asset catalog using a subject term extraction model and a named entity recognition model, and employing a triplet structure for recording. Background Technology

[0002] As an important technology in the data governance system, the construction of a high-quality data asset catalog is a crucial foundation for enabling data sharing across institutions, regions, and industries.

[0003] The existing data asset catalog is constructed in three main aspects:

[0004] The first aspect is the construction of subject domain hierarchical classification. The purpose of constructing subject domain hierarchical classification is to facilitate managers' understanding of data content. Therefore, when constructing subject domain hierarchical classification, the terminology used must not only conform to the industry standards, but also meet the requirements for terminology when sharing data across industries, and finally, it must also ensure the uniqueness of the term within its own subject domain.

[0005] The second aspect is the construction of business attributes. When constructing business attributes, it is necessary to realize the correspondence between business attributes and subject domain hierarchical classification.

[0006] The third aspect is the construction of technical attribute classification. Technical attribute classification refers to the database and fields where data items are actually stored. When constructing technical attribute classification, it is necessary to ensure that the technical attribute classification corresponds to the business attributes, and that there is no ambiguity between the subject domain hierarchical classification and the data item correspondence.

[0007] There are two main problems in the existing process of building a data asset catalog. First, there is the issue of the correspondence between business attributes and technical attributes. Because business personnel and technical personnel analyze the bank's subsystems from different perspectives, it is difficult to match business attributes with technical attributes. Second, there is the issue of constructing a hierarchical classification system based on subject domains. Since the construction of a hierarchical classification system for subject domains involves all systems within the bank, few staff members are familiar with all business operations and bank subsystems, making it difficult to carry out the work of constructing a hierarchical classification system for subject domains from a global perspective.

[0008] To solve the above problems, existing technologies typically employ the following two methods:

[0009] 1. The data asset catalog is constructed and organized by humans based on their understanding of the classification of business attributes and technical attributes.

[0010] 2. Using the results of metadata management in the field of data governance as a data asset catalog, also known as a "metadata map", metadata displays relevant information such as the storage structure of data from a technical perspective. By analyzing scripts such as data extraction and transformation, the data lineage is obtained, and drill-down queries of the metadata map are completed.

[0011] For the first approach, the number of participants is directly proportional to the number of bank subsystems. The difficulty in constructing the data asset catalog leads to numerous iterative processes, and cross-departmental and cross-system communication is time-consuming and inefficient. For the second approach, the lack of hierarchical classification of subject areas and definition of business attributes, or the absence of data indexes, makes it unsuitable for data sharing with external units. It is limited to internal personnel with some understanding of system construction. Furthermore, the relationships between tables are mostly implemented through program code, but there are currently no security or technical solutions for parsing this code. Therefore, the lineage relationships between upstream and downstream data cannot encompass all relationships between data assets. Summary of the Invention

[0012] One objective of this invention is to provide a method for constructing a data asset catalog in the banking industry, in order to solve the problems in the prior art such as low efficiency of cross-departmental and cross-system communication among staff, inability of data to meet requirements when sharing data with external units, and the inability of the data upstream and downstream lineage relationships to include the different levels of business logic relationships between all data assets.

[0013] To achieve the above objectives, this invention discloses a method for constructing a banking data asset catalog, comprising:

[0014] Based on the relevant documents concerning the business requirements of the banking subsystem, a subject domain hierarchical classification was obtained;

[0015] Based on the business requirements of the banking subsystem, the relevant documents of the business requirements, and the hierarchical classification of the subject domains, a business attribute classification is constructed.

[0016] Based on the architecture and functions of the bank subsystem, a technical attribute classification relationship diagram is constructed, wherein the technical attribute classification relationship diagram includes metadata relationships within the bank subsystem and cross-system metadata relationships;

[0017] Add the business attribute categories of the banking subsystem to the technical attribute category relationship diagram.

[0018] Preferred,

[0019] Collect relevant documents regarding the business requirements of the aforementioned banking subsystem;

[0020] Input the relevant documents of the business requirements of all the bank subsystems into the pre-acquired keyword extraction model to obtain multiple first-level keyword numbers, first-level keywords of each first-level keyword number, the number of documents of each first-level keyword number, and the weight of each first-level keyword in the corresponding first-level keyword number;

[0021] The relevant documents of the business requirements of each of the bank subsystems are input into the keyword extraction model to obtain the secondary keyword number of each bank subsystem, the secondary keywords of the secondary keyword number of each bank subsystem, the number of documents of the secondary keyword of each bank subsystem, and the weight of each secondary keyword in the corresponding secondary keyword number.

[0022] Preferred,

[0023] Based on each primary keyword and its weight, the subject area corresponding to the primary keyword number is determined;

[0024] Based on the parallel relationship between the primary keywords of each primary subject term number and the weight of the primary keywords, a primary keyword is selected as the name of the primary subject term number;

[0025] Based on the parallel relationship between the names of the first-level subject terms corresponding to the same subject domain, the name of the first-level subject term is selected as the name of the first-level category of the subject domain;

[0026] Based on the relevance of each bank subsystem to each subject domain, weights are assigned to each bank subsystem under each subject domain;

[0027] Based on each of the secondary keywords and their weights, the primary category corresponding to the secondary keyword number is determined;

[0028] Based on the weight and parallel relationship of the secondary keywords of each secondary keyword number, select one of the secondary keywords as the name of the secondary keyword number;

[0029] According to the merging rules, the secondary keywords with the same name and corresponding to the same primary category are merged to update the weight of the secondary keywords;

[0030] Based on the weight of the secondary keywords corresponding to the names of the secondary subject terms, the parallel relationship of the names of the secondary subject terms, and the primary category corresponding to the secondary subject terms, the names of the secondary subject terms are selected as the secondary category.

[0031] Preferred,

[0032] The weight of the secondary keyword is multiplied by the weight of its corresponding bank subsystem to obtain the weight of the new secondary keyword.

[0033] The secondary keywords in the secondary subject terms that correspond to the same primary category and have the same name are added together with the corresponding secondary keywords according to the weight of the new secondary keywords.

[0034] Preferred,

[0035] Before constructing the business attribute classification based on the business requirements of the banking subsystem, the relevant documents of the business requirements, and the subject domain hierarchical classification, the following further steps are included:

[0036] Identify named entities, where the named entities are nouns in the second-level category;

[0037] The relevant documents concerning the business requirements of the banking subsystem are annotated according to the annotation requirements of the named entities and named entity recognition models.

[0038] The named body recognition model was trained using the relevant documents of the business requirements of the bank's subsystems that had been annotated.

[0039] Preferred,

[0040] The business requirements of the bank subsystem that has been annotated and the relevant documents required by the business are input into the named entity recognition model that has been trained to obtain a word segmentation list of each named entity.

[0041] The number of bank subsystems involved in each of the segmented words in the statistical word segmentation list;

[0042] Based on the relevance of each bank subsystem to each subject domain, weights are assigned to each bank subsystem under each subject domain;

[0043] The word segmentation weight is obtained by multiplying the number of bank subsystems involved in the word segmentation by the weight of the bank subsystem.

[0044] The business attribute classification of the named entity is determined based on the word segmentation weight of each named entity in the same banking subsystem.

[0045] Preferred,

[0046] Determine the scope of the banking subsystems included in the business attribute classification;

[0047] Among the banking subsystems, select the one with higher weight and a larger total number of word segments;

[0048] Analyze the relationship between the selected bank subsystem and other bank subsystems from the technical attribute classification relationship diagram;

[0049] Analyze the technical attribute classification and related fields from the perspective of the business attribute classification, determine the mapping relationship between the business attributes and the technical attribute classification, and add the business attribute classification of the bank subsystem to the technical attribute classification relationship diagram.

[0050] To achieve the above objectives, this invention discloses an apparatus for constructing a banking data asset catalog, comprising:

[0051] Subject area hierarchical unit, used to obtain subject area hierarchical classification of relevant documents based on the business requirements of the banking subsystem;

[0052] A classification construction unit is used to construct business attribute classifications based on the business requirements of the banking subsystem, related documents of the business requirements, and the subject domain hierarchical classification.

[0053] The relationship diagram construction unit is used to construct a technical attribute classification relationship diagram based on the architecture and functions of the bank subsystem, wherein the technical attribute classification relationship diagram includes metadata relationships within the bank subsystem and cross-system metadata relationships;

[0054] The category addition module is used to add the business attribute categories of the banking subsystem to the technical attribute category relationship diagram.

[0055] Preferred:

[0056] The document collection unit is used to collect relevant documents regarding the business requirements of the banking subsystem.

[0057] The first-level subject term numbering unit is used to input the business requirements of all banking subsystems and the related documents of the business requirements into the pre-acquired subject term extraction model to obtain multiple first-level subject term numbers, the first-level keywords of each first-level subject term number, the number of documents of each first-level subject term number, and the weight of each first-level keyword in the corresponding first-level subject term number.

[0058] The secondary subject term numbering unit is used to input the business requirements and related documents of each bank subsystem into the subject term extraction model to obtain the secondary subject term number of each bank subsystem, the secondary keywords of the secondary subject term number of each bank subsystem, the number of documents of the secondary subject term of each bank subsystem, and the weight of each secondary keyword in the corresponding secondary subject term number.

[0059] Preferred:

[0060] The topic domain determination module is used to determine the topic domain corresponding to the first-level keyword number based on each first-level keyword and the weight of the first-level keyword;

[0061] The primary category selection module is used to select one of the primary keywords as the primary category of the topic domain based on the parallel relationship between the primary keywords and the weight of the primary keywords.

[0062] The weight assignment module is used to assign weights to each of the bank subsystems under each of the subject domains based on the relevance of each of the bank subsystems to each of the subject domains.

[0063] The primary category determination module is used to determine the primary category corresponding to the secondary keyword number based on each secondary keyword and its weight.

[0064] The topic domain acquisition module is used to acquire the topic domain corresponding to the second-level topic word number based on the determined correspondence between the first-level category and the topic domain;

[0065] The weight merging module is used to merge the secondary keywords corresponding to the secondary topic word numbers of the same primary category according to the merging rules, so as to update the weight of the secondary keywords;

[0066] The secondary category acquisition module is used to select standardized words from the updated secondary keywords based on the secondary keywords and their corresponding primary categories to obtain the secondary category.

[0067] Preferred:

[0068] The weight of the secondary keyword is multiplied by the weight of its corresponding bank subsystem to obtain the weight of the new secondary keyword.

[0069] The secondary keywords in the secondary topic terms corresponding to the same primary category are added together with the corresponding secondary keywords according to the weight of the new secondary keywords.

[0070] Preferred:

[0071] A named entity determination module is used to determine named entities, wherein the named entities are nouns in the second-level category;

[0072] The annotation module is used to annotate the relevant documents of the business requirements of the banking subsystem according to the annotation requirements of the named entities and the named entity recognition model.

[0073] Preferred:

[0074] The word segmentation list acquisition module is used to input the business requirements of the bank subsystem that have been annotated and the relevant documents required by the business into the named entity recognition model to obtain the word segmentation list of each named entity;

[0075] The statistics module is used to count the word segmentation list of each named entity in each bank subsystem, the number of words in each word segmentation list, and the frequency of each word segmentation in the business documents of each bank subsystem.

[0076] The named entity correspondence module is used to determine the correspondence between each bank subsystem and the named entity based on the number of words in the word segmentation list of each named entity in the same bank subsystem;

[0077] The word segmentation list merging module is used to merge the word segmentation lists of the same named entity based on the nouns of the second-level category whose named entities are nouns of the second-level category, which are nouns of the second-level category selected from the second-level keywords, and the correspondence between the second-level keywords and the bank subsystem.

[0078] The business attribute classification determination module is used to determine the business attribute classification of the named entity based on the frequency of the occurrence of the words in the merged word segmentation list in the business documents of each of the bank subsystems.

[0079] The classification construction module is used to determine the correspondence between the bank subsystem and the business attribute classification based on the correspondence between the bank subsystem and the named entity, and to construct the business attribute classification.

[0080] Preferred:

[0081] The bank subsystem scope determination module is used to determine the scope of bank subsystems included in the business attribute classification;

[0082] A bank subsystem selection unit is used to select the bank subsystem with a higher weight and a larger total number of word segments from the bank subsystems.

[0083] The bank subsystem analysis module is used to analyze the relationship between the selected bank subsystem and other bank subsystems from the technical attribute relationship diagram;

[0084] The classification addition module is used to analyze the technical attributes and related fields from the perspective of the business attribute classification, determine the mapping relationship between the business attributes and the technical attributes, and add the business attribute classification of the bank subsystem to the technical attribute relationship diagram.

[0085] This application can solve the problems in the prior art, such as low efficiency of cross-departmental and cross-system communication among staff, inability of data to meet requirements when sharing data with external units, and the inability of the data upstream and downstream relationships to include the relationships between all data assets. Attached Figure Description

[0086] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0087] Figure 1 A flowchart illustrating a method for constructing a banking data asset catalog according to an embodiment of the present invention is shown.

[0088] Figure 2 A flowchart illustrating a method for constructing a banking data asset catalog according to an embodiment of the present invention is shown.

[0089] Figure 3 A flowchart illustrating a method for constructing a banking data asset catalog according to an embodiment of the present invention is shown.

[0090] Figure 4 A flowchart illustrating a method for constructing a banking data asset catalog according to an embodiment of the present invention is shown.

[0091] Figure 5 A flowchart illustrating a method for constructing a banking data asset catalog according to an embodiment of the present invention is shown.

[0092] Figure 6 A flowchart illustrating a method for constructing a banking data asset catalog according to an embodiment of the present invention is shown.

[0093] Figure 7 A flowchart illustrating a method for constructing a banking data asset catalog according to an embodiment of the present invention is shown.

[0094] Figure 8 This illustrates a primary classification storage table according to an embodiment of the present invention;

[0095] Figure 9 This illustrates a two-level classification storage table according to an embodiment of the present invention;

[0096] Figure 10 This illustrates a business attribute classification storage table according to an embodiment of the present invention;

[0097] Figure 11 This diagram illustrates the classification relationship of technical attributes of a single-bank subsystem according to an embodiment of the present invention.

[0098] Figure 12 This diagram illustrates the classification relationship of technical attributes of a multi-bank subsystem according to an embodiment of the present invention.

[0099] Figure 13 This illustrates a data attribute relationship diagram containing business attribute relationships according to an embodiment of the present invention;

[0100] Figure 14 This invention illustrates a data asset catalog information storage table according to an embodiment of the present invention;

[0101] Figure 15 A schematic diagram of the structure of a banking data asset catalog construction apparatus according to an embodiment of the present invention is shown. Figure 1 ;

[0102] Figure 16 A schematic diagram of the structure of a banking data asset catalog construction apparatus according to an embodiment of the present invention is shown. Figure 2 ;

[0103] Figure 17 This diagram illustrates the structure of a subject domain hierarchical unit in a banking data asset catalog according to an embodiment of the present invention.

[0104] Figure 18 This diagram illustrates the structure of a classification and construction unit for a banking data asset catalog according to an embodiment of the present invention.

[0105] Figure 19 This diagram illustrates the structure of a classification and addition unit in a banking data asset catalog according to an embodiment of the present invention. Detailed Implementation

[0106] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0107] This embodiment discloses, on the one hand, a method for constructing a banking data asset catalog, such as... Figure 1 As shown, the method for constructing this banking data asset catalog includes:

[0108] S101: Documents related to the business requirements of the banking subsystem, obtaining subject area hierarchical classification;

[0109] Specifically, the subject domain classification system comprises three levels, from top to bottom: subject domain, primary category, and secondary category. When constructing the subject domain classification system, it is required that the terminology be standardized and conform to industry standards.

[0110] In the subject domain classification, subject domains can be selected from the following ten recognized data governance subject domains: parties, products, agreements, events, assets, finance, institutions, regions, marketing, and channels.

[0111] After determining the subject domains, based on the relevant documents of the business requirements of the bank subsystem, the subject term extraction model is used to obtain the primary and secondary categories. Then, the correspondence between the subject domains and the primary categories, and the correspondence between the primary and secondary categories are determined, and finally the construction of the subject domain hierarchical classification is completed.

[0112] Using a topic word extraction model to extract primary and secondary categories not only satisfies the requirement that the sets of primary and secondary categories can encompass the entire topic domain, but also makes the process of constructing primary and secondary categories more scientific.

[0113] The technical principle of the topic word extraction model is to extract topics based on the Chinese financial pre-trained model and the unsupervised method of natural language processing. Taking the mainstream BERTopic algorithm as an example, BERT is the latest word vector technology in the field of natural language processing. BERTopic is a topic modeling technology based on BERT word vectors. It uses BERT embedding and clustering-based TF-IDF to create dense clusters. Moreover, it reduces the embedding dimension to interpret topic information before clustering documents, and at the same time, it can retain representative keywords in topic descriptions.

[0114] The above process standardizes industry terminology and facilitates data sharing. Furthermore, using pre-trained financial models eliminates the need for feature model learning, resulting in more accurate results. The specific implementation of pre-trained financial models involves using documents on standardized terminology in the financial and banking industries, or financial news articles, as training corpora to obtain mature pre-trained models, such as FinBERT, an open-source language model from Entropy Technology. This allows the topic word extraction model based on the pre-trained financial model to identify and output standardized industry terminology.

[0115] S102: Construct a business attribute classification based on the business requirements of the banking subsystem, related documents of the business requirements, and subject domain hierarchical classification;

[0116] Specifically, the process of constructing business attribute categories consists of three steps:

[0117] First, determine the named entity. The named entity should use the standardized name in the secondary category, such as credit business, instant transfer, and cross-border currency payment business, instead of non-standard words such as loan, transfer money, and cross-border institution.

[0118] Then, the correspondence between named entities and bank subsystems is determined. When confirming the correspondence, the number of words in the word segmentation list of each named entity in each bank subsystem is used to determine the corresponding named entity of that bank subsystem.

[0119] Finally, the word segmentation lists of the same named entity are merged across bank subsystems. In the merged word segmentation list, the word segmentation that appears most frequently in the text is selected as the business attribute category of this named entity. Based on the correspondence between the word segmentation and the named entity, the correspondence between the named entity and the business attribute category is determined.

[0120] Named Entities Recognition (NER) is a fundamental task in natural language processing, aiming to identify the named objects to which entities in a corpus belong. This embodiment uses a custom entity type. The implementation process generally involves first labeling the dataset according to the custom entity type, and then selecting a recognition method and tool, such as HMM, MEMM, ME, CRF, SVM, etc. Commonly used named entity recognition tools are HanLP and CRF++.

[0121] S103: Based on the architecture and functions of the bank subsystem, construct a technical attribute classification relationship diagram, which includes metadata relationships within the bank subsystem and metadata relationships across bank subsystems;

[0122] Specifically, the technical attribute classification is defined as the data corresponding to the hierarchical classification of subject domains. This data is stored in the form of tables or fields. When constructing the technical attribute classification relationship graph, it is built using a triple <node, relation, node> structure, where the relation includes:

[0123] Belonging Relationship: The relationship between the table and the banking subsystem is "belonging," meaning the table belongs to the banking subsystem.

[0124] Broadcast relationship: describes the relationship between banking subsystems. For example, if banking subsystem 1 directly transmits information from one or more of its tables to banking subsystem 2, the relationship between banking subsystem 1 and banking subsystem 2 is a broadcast relationship.

[0125] Aggregation relationship: describes the relationship between banking subsystems. For example, banking subsystem 1 summarizes, matches, simplifies and other processes the information from one or more of its tables and then passes it to banking subsystem 2. The relationship between banking subsystem 1 and banking subsystem 2 is an aggregation relationship.

[0126] Containment relationship: describes the relationship between tables. For example, the data in table a comes entirely from table b, and table b contains table a. For example, the archive table contains the original records of the payment business table.

[0127] Change relationship: Describes the relationship between tables. For example, the data in table a is the change information of table b, and table b has a change relationship with table a.

[0128] Summary relationship: describes the relationship between tables. For example, the data in table a is a summary of the information in table b, and table b has a summary relationship with table a.

[0129] Reference relationship: describes the relationship between tables. For example, the data in table a is information about query, cancellation and other operations in table b, or the fields of table a depend on the data range of table b. Table a has a reference relationship with table b.

[0130] Inheritance relationship: describes the relationship between tables. For example, all the information in table a of banking subsystem 1 is directly passed to table b of system 2. Table b has an inheritance relationship with table a.

[0131] Mapping relationship: describes the relationship between fields. For example, field a and field b have a one-to-one mapping relationship, and field a has an association relationship with field b. For example, the transaction identifiers in the archive table and the payment business table have a one-to-one correspondence.

[0132] Association: Describes the relationship between fields. For example, field a in banking subsystem 1 changes as field b in banking subsystem 2 changes. Field a has an association with field b.

[0133] Dependency: Describes the relationship between fields. For example, the data range of field a is field b, and field a has a dependency relationship with field b. For example, the initiating participating institutions in the payment business table are the financial institution code field range in the participating institution table.

[0134] The nodes in the triple <node, relation, node> structure are provided by the metadata information of the banking subsystem.

[0135] S104: Add the business attribute categories of the banking subsystem to the technical attribute category relationship diagram.

[0136] Specifically, based on the relevant fields and database information of the technical attribute classification and the weight of its banking subsystem, the effective information in the relevant fields and database is determined, the correspondence between the effective information and the business attribute classification is determined, the business attribute classification is added to the technical attribute classification relationship graph according to the above correspondence, and finally the metadata and metadata relationship are converted into a table.

[0137] The data asset catalog constructed using the above method includes subject domain hierarchical classification, technical attribute classification, and business attribute classification. The relationship between subject domain hierarchical classification and business attribute classification, as well as technical attribute classification, is displayed in the form of a diagram, making the relationships between data more intuitive and easy to understand, facilitating cross-industry data sharing, and enhancing the value of the data asset catalog.

[0138] In some embodiments, before obtaining the subject domain hierarchical classification from the relevant documents based on the business requirements of the banking subsystem, such as Figure 2 As shown, the method for constructing this banking data asset catalog also includes:

[0139] S201: Collect relevant documents regarding the business requirements of the banking subsystem;

[0140] Specifically, its purpose is to facilitate subsequent operations by inputting the relevant documents of the business requirements of the banking subsystem into the relevant documents separately for each banking subsystem.

[0141] S202: Input the relevant documents of the business requirements of all bank subsystems into the pre-acquired keyword extraction model to obtain multiple first-level keyword numbers, the first-level keywords of each first-level keyword number, the number of documents of each first-level keyword number, and the weight of each first-level keyword in the corresponding first-level keyword number;

[0142] Specifically, in existing technologies, the keyword extraction model cannot output keywords, but can only output them in the form of keyword numbers. At the same time, a primary keyword number can contain multiple primary keywords. When obtaining information such as primary keyword numbers and primary keywords, relevant documents of the business requirements of all bank subsystems are simultaneously input into the keyword extraction model.

[0143] S203: Input the relevant documents of the business requirements of each bank subsystem into the keyword extraction model to obtain the secondary keyword number of each bank subsystem, the secondary keywords of the secondary keyword number of each bank subsystem, the number of documents of the secondary keyword of each bank subsystem, and the weight of each secondary keyword in the corresponding secondary keyword number.

[0144] Specifically, when obtaining secondary subject heading numbers and secondary keywords, the relevant documents of the business requirements of the banking subsystem are input into the subject heading extraction model separately for each banking subsystem.

[0145] The above steps yielded primary subject heading numbers and their keyword information, as well as secondary subject heading numbers and their keyword information, laying the groundwork for determining the primary and secondary classifications.

[0146] In some embodiments, such as Figure 3 As shown, the relevant documents based on the business requirements of the banking subsystem have been classified into subject areas including:

[0147] S301: Determine the subject area corresponding to the first-level keyword number based on each first-level keyword and its weight;

[0148] Specifically, first, the definition of each subject area is determined. Then, by comparing the relationship between the definition of each subject area and the meaning of the keywords in the first-level subject term number, it is determined whether the subject area corresponds to the first-level subject term number. When making the judgment, the first-level keywords with higher weight are given priority.

[0149] Confirming the correspondence between primary subject heading numbers and subject domains prepares for confirming the correspondence between subject domains and primary categories in subsequent steps.

[0150] S302: Based on the parallel relationship between the primary keywords of each primary subject term number and the weight of the primary keywords, select one of the primary keywords as the name of the primary subject term number;

[0151] Specifically, when selecting the name of a first-level subject term number, the weight of the first-level keywords under that first-level subject term number and the parallel relationship between the first-level keywords should be comprehensively considered. The meaning of the selected first-level keywords should cover the meaning of all the first-level keywords in that first-level subject term number as much as possible. In addition, the greater the weight of the first-level keyword, the more likely it is to be selected as the name of the first-level subject term number.

[0152] S303: Based on the parallel relationship between the names of the first-level subject terms corresponding to the same subject domain, select the name of the first-level subject term as the first-level classification of the subject domain;

[0153] Specifically, when selecting the names of the first-level subject headings as the first-level categories, the parallel relationships between the names of the first-level subject headings should be comprehensively considered. A subject domain contains multiple first-level categories. It is necessary to ensure that the selected first-level categories are mutually exclusive in meaning and that the set of selected first-level categories can cover the entire subject domain.

[0154] Please sort the above information as follows Figure 8 Stored in tabular form. Figure 8 The table contains information including subject area, first-level subject term number, first-level keywords, first-level keyword weight, and first-level category.

[0155] S304: Assign weights to each bank subsystem under each subject domain based on the relevance of each bank subsystem to each subject domain;

[0156] Specifically, different bank subsystems have clear correspondences in the subject domain. For example, the working goal of bank subsystem X is bank management, while the working goal of bank subsystem Y is to complete payment transactions. Therefore, bank subsystem X has a deeper connection with the bank subject domain, while bank subsystem Y has a deeper connection with the event subject domain. Under the bank subject domain, the weight assigned to bank subsystem X is greater than the weight assigned to bank subsystem Y.

[0157] S305: Determine the primary category corresponding to the secondary keyword number based on each secondary keyword and its weight;

[0158] Specifically, the process of determining the primary category corresponding to the secondary subject term number is the same as the process of determining the subject area corresponding to the primary subject term number. First, the definition of each primary category must be considered. Then, by comparing the definition of each primary category with the meaning of the keywords in the secondary subject term number, the correspondence between the secondary subject term number and the primary category can be obtained.

[0159] S306: Based on the established correspondence between primary categories and subject areas, obtain the subject area corresponding to the secondary subject term number;

[0160] Specifically, the correspondence between the primary classification and the subject is used as a bridge to obtain the subject domain corresponding to the secondary subject term number.

[0161] S307: Based on the weight and parallel relationship of the secondary keywords of each secondary keyword number, select one of the secondary keywords as the name of the secondary keyword number;

[0162] Specifically, when selecting secondary keywords as names for secondary subject headings, it is necessary to comprehensively consider the weight of the secondary keywords and the parallel relationship between them. The meaning of the selected secondary keywords should cover the meaning of all keywords in the secondary subject heading as much as possible.

[0163] S308: Merge the secondary keywords with the same name and corresponding to the same primary category according to the merging rules, so as to update the weight of the secondary keywords;

[0164] Specifically, the secondary subject term numbers are obtained by inputting the business requirements and related documents of the banking subsystems into the subject term extraction model according to different banking subsystems. Therefore, the secondary subject term numbers and their secondary keywords all correspond to the banking subsystems.

[0165] When merging secondary keywords with secondary subject numbering, the weight of the secondary keyword is first multiplied by the weight of the corresponding bank subsystem, and then the weight of the secondary keyword is added to the weight of the corresponding secondary keyword.

[0166] S309: Based on the weight of the secondary keyword corresponding to the name of the secondary subject term number, the parallel relationship of the names of the secondary subject term numbers, and the primary category corresponding to the secondary subject term number, select the name of the secondary subject term number as the secondary category.

[0167] Specifically, first, the definition of the primary category corresponding to the secondary keywords is determined. Then, the vocabulary that conforms to the banking industry standardization among the merged secondary keywords is selected as the secondary category under the primary category. There can be multiple secondary categories under the same primary category.

[0168] Please sort the above information as follows Figure 9 Stored in tabular form. Figure 9 The information in the table includes subject area, primary category, secondary subject number, system number, system weight, secondary keyword, weight of secondary keyword before merging, weight of secondary keyword after merging, and secondary category.

[0169] In some embodiments, such as Figure 4 As shown, the process for merging second-level keywords based on second-level subject heading numbers is as follows:

[0170] S401: Multiply the weight of the secondary keyword by its corresponding bank subsystem weight to obtain the new weight of the secondary keyword;

[0171] Specifically, the secondary subject term numbers are obtained by inputting the business requirements and related documents of the banking subsystem into the subject term extraction model according to the banking subsystem. Therefore, the correspondence between the secondary subject term numbers and their secondary keywords and the banking subsystem can be easily obtained.

[0172] S402: The secondary keywords in the secondary subject terms that correspond to the same primary category and have the same name are added together with the corresponding secondary keywords according to the weight of the new secondary keywords.

[0173] Specifically, among the secondary subject terms that correspond to the same primary category and have the same name, the weights of the secondary keywords with the same name are added together. If there is no secondary keyword with the same name, the weight of the secondary keyword remains unchanged during the merging process.

[0174] In some embodiments, such as Figure 5 As shown, before constructing the business attribute classification based on the business requirements of the banking subsystem, related documents, and subject domain hierarchical classification, the following further steps are included:

[0175] S501: Determine the named entity, where the named entity is a noun in the second-level category;

[0176] Specifically, named entities are terms involved in the second-level classification. When determining named entity objects, it is important to distinguish between named entities and named entity objects. Taking the event subject domain and the bank subject domain as examples, the second-level classifications of the event subject domain include credit business, instant transfer business, pledge financing business, debit business, and cross-border currency payment business. These five second-level classifications are named entities. The second-level classifications of the bank subject domain include large-scale participating institutions, small-scale participating institutions, and online participating institutions. These three second-level classifications are named entity objects.

[0177] S502: Annotate the relevant documents of the business requirements of the bank subsystem according to the annotation requirements of named entities and named entity recognition models.

[0178] Specifically, the function of a named entity recognition model is to identify named entities defined in a corpus. The techniques used include rule-based and dictionary-based methods, statistical methods, and hybrid methods. Before using a named entity model, the input data needs to be annotated.

[0179] S503: Train the named body recognition model using the relevant documents of the business requirements of the bank subsystem that have been fully annotated.

[0180] Specifically, the first step is to organize the business requirements and related documents for each bank subsystem. Then, the named entity recognition model is trained according to the annotation requirements of the model. The goal of the named entity recognition model is to identify named entities in the corpus. The main methods are divided into three categories: rule-based and dictionary-based methods, statistical methods, and hybrid methods.

[0181] The above steps complete the preparation work for classifying the attributes of the bank subsystem.

[0182] In some embodiments, such as Figure 6 As shown, the business attribute classification is constructed based on the business requirements of the banking subsystem, related documents, and subject domain hierarchical classification. One step includes:

[0183] S601: Input the business requirements of the bank subsystem and the relevant documents required by the business into the trained named entity recognition model to obtain a word segmentation list of each named entity;

[0184] Specifically, the word segmentation list is composed of words that are related to named entities.

[0185] S602: Count the number of bank subsystems involved in each of the segmented words;

[0186] Specifically, the occurrence of each word in the documents of the banking subsystem is counted. As long as the word appears once in the document of the banking subsystem, regardless of how many times it appears in the banking subsystem afterward, the number of banking subsystems involved in the word is increased by 1. For example, if word 1 appears 3 times in banking subsystem A, 10 times in banking subsystem B, and 0 times in banking subsystem C, then the number of banking subsystems involved in the word is 3.

[0187] S603: Assign weights to each of the bank subsystems under each of the subject domains based on the relevance of each of the bank subsystems to each of the subject domains;

[0188] Specifically, the steps are the same as those in S304, and the weights assigned to each bank subsystem under each subject area are the same as those assigned in S304.

[0189] S604: Multiply the number of bank subsystems involved in the word segmentation by the weight of the bank subsystem and sum them up to obtain the word segmentation weight;

[0190] Specifically, the weights of the systems where the word segmentation occurs are multiplied by 1 and accumulated. For example, if word segmentation 1 appears 3 times in bank subsystem A, 10 times in bank subsystem B, and 0 times in bank subsystem C, the weight of bank subsystem A is 3, the weight of bank subsystem B is 1, and the weight of bank subsystem C is 10. Then, the word segmentation weight calculation process for word segmentation 1 is: (1*3)+(1*10)=13. And so on, the word segmentation weight of each word for each named entity is calculated.

[0191] S605: Based on the word segmentation weights of the named entities in the same banking subsystem, determine the business attribute classification of the named entity, and then determine the business attribute classification of the secondary category corresponding to the named entity.

[0192] Specifically, from the word segmentation list of each named entity, the word segmentation with the higher proportion is selected as the business attribute category of the named entity. Since the name of the named entity is taken from the secondary category, the business attribute category of the secondary category is further obtained.

[0193] Store the above information Figure 10 In the table, Figure 10 The table contains information such as subject area, primary category, secondary category, system number, system weight, word segmentation name, word segmentation list, word segmentation frequency, word segmentation weight, and business attribute classification.

[0194] By linking the subject domain classification with the business attribute classification through the above steps, the relationship between data becomes more intuitive, laying the groundwork for adding business attributes to the technical attribute classification relationship diagram in subsequent steps.

[0195] In some embodiments, such as Figure 7 As shown, the process of adding the business attribute categories of the banking subsystem to the technical attribute category relationship diagram includes:

[0196] S701: Determine the scope of the banking subsystems included in the business attribute classification;

[0197] S702: Select the banking subsystem with higher weight and larger total number of word segments from the banking subsystems;

[0198] Specifically, through Figure 10 The table in the document filters out the banking subsystems with higher weights and larger total number of word segments. Systems with higher weights and larger total number of word segments contain more information than other systems, so this system was selected as the object of analysis.

[0199] S703: Analyze the relationship between the selected bank subsystem and other bank subsystems from the technical attribute classification relationship diagram;

[0200] Specifically, in the technical attribute classification relationship diagram, a triplet <node, relationship, node> structure is used to record the relationship between the selected bank subsystem and other bank subsystems.

[0201] From the technical attribute relationship diagram, analyze the relationship between the system and its fields in the previous step and other systems to determine whether it is necessary to add corresponding technical attribute fields for this business attribute.

[0202] S704: Analyze the technical attribute classification and related fields from the perspective of business attribute classification, determine the mapping relationship between business attributes and technical attribute classification, and classify the business attributes of the bank subsystem into the technical attribute classification relationship diagram.

[0203] Specifically, based on the definition of business attribute classification, find the corresponding database tables and fields, refer to the relationships between fields and the weight of the bank subsystem, determine the valid fields, and determine the mapping relationship between the valid fields and the business attribute classification. In addition, it is also necessary to consider, based on experience and expertise, whether the fields in other tables of the relevant bank subsystem meet the definition.

[0204] The final technical attribute classification relationship diagram containing business attribute relationships is converted into a table and stored.

[0205] Through the above steps, the data asset catalog is constructed in the form of a diagram, which fully demonstrates the relationships between tables, libraries, and fields. The logic between data is clear, which is convenient for staff to understand and share data. At the same time, it overcomes the problem that the lack of code parsing leads to the inability to include all the relationships between data assets in the upstream and downstream data lineage.

[0206] In some embodiments, the specific process for constructing a primary classification is as follows: The central bank payment system has five banking subsystems, namely A, B, C, D, and E.

[0207] 1. The business requirements and related documents of these five bank subsystems are used as input data for the keyword extraction model. The output of the keyword extraction model includes the first-level keyword number, the first-level keyword for each first-level keyword number, the number of documents for each first-level keyword, and the weight of the first-level keyword under that keyword number.

[0208] The first-level keywords listed under first-level subject term number **1 include (payment participating institutions, 0.2), (participating institution line number, 0.1), (participating institution full name, 0.06), (line code, 0.05), (participating institution category, 0.05), (bank identification code, 0.04), (direct participating institutions, 0.03), etc. Among them, (payment participating institutions, 0.2) indicates that the weight of the first-level keyword "payment participating institutions" under first-level subject term number **1 is 0.2.

[0209] 2. Based on the definition of the banking subject domain, it can be determined that the first-level subject term number **1 can be classified as the banking subject domain. This is because Teradata defines the banking subject domain as the object served by the central bank's payment system. Among the first-level keywords, payment participating institutions, participating institution bank codes, participating institution full names, participating institution categories, and directly participating institutions are all closely related to the objects served by the central bank's payment system.

[0210] 3. Select the name of the first-level keyword in the first-level subject term number **1. When selecting the name of the first-level subject term number, the weight of the first-level keyword and the parallel relationship between the first-level keywords should be considered. The higher the weight of the first-level keyword, the more likely it is to become a first-level category. In the first-level subject term number **1, the weight of the first-level keyword "payment participating institution" is 0.2, which is higher than the weight of all other first-level keywords. Therefore, "payment participating institution" is selected as the name of the first-level subject term number **1.

[0211] The first-level keywords listed under first-level subject term number **2 include (payment business, 0.1), (information business, 0.1), (funds clearing, 0.05), (credit business, 0.05), (bulk business, 0.03), (real-time business, 0.02), (instant transfer, 0.001), etc. Teradata defines the event subject domain as the financial and non-financial activities between banks and central bank institutions. The method of selecting the first-level subject term number name is the same as that of first-level subject term number **1. After comprehensively considering the weight of the first-level keywords and the parallel relationship between the first-level keywords, payment business is selected as the name of the first-level subject term number.

[0212] Then, the names of the next-level subject terms in each subject domain are selected as the first-level categories of that subject domain. When selecting, the set of all selected first-level categories must cover the entire subject domain, and the first-level categories must be mutually exclusive.

[0213] 4. Arrange the above information according to... Figure 8 The information is stored in the form of a table. Taking the information related to the first-level subject term number **1 as an example, the storage result is as follows:

[0214]

[0215]

[0216] In some embodiments, the process of constructing a secondary classification is as follows:

[0217] 1. Collect the business requirements and related documents of the five bank subsystems A, B, C, D, and E under the central bank's payment system. After collection, input the business requirements and related documents of each bank subsystem into the keyword extraction model. The output will be the secondary keyword number of each bank subsystem, the secondary keywords of each secondary keyword number, the number of documents of each secondary keyword, and the weight of each secondary keyword in the corresponding secondary keyword number.

[0218] The output is as follows:

[0219] Under the secondary subject term ##1 in the bank subsystem A, the secondary keywords include (Credit Business, 0.1), (Instant Transfer, 0.1), (Pledged Financing, 0.1), (Cross-border Payment, 0.05), (End-of-Day Processing, 0.005), (Abnormal Processing, 0.005), and (Payment-related Business, 0.2). Among them, (Credit Business, 0.1) means that the weight of the secondary keyword Credit Business under the secondary subject term ##1 is 0.1.

[0220] Under the secondary subject term number @@1 in the bank subsystem E, the secondary keywords include (credit business, 0.1), (debit business, 0.1), (real-time business, 0.1), (parameter management, 0.05), (batch business, 0.05), (end-of-day processing, 0.0001), (abnormal handling, 0.0001), (payment business, 0.2), etc.

[0221] 2. Based on the degree of association between the banking subsystem and the subject domain, assign a weight to each banking subsystem under each subject domain. The higher the weight assigned, the higher the degree of association between the banking subsystem and the subject domain. The degree of association between the banking subsystem and the subject domain is determined by the work content of the banking subsystem.

[0222] Among the five banking subsystems mentioned above, the job of banking subsystem B is to manage the bank. Therefore, it can be inferred that the correlation between banking subsystem B and the banking subject domain is relatively high, while the correlation with the event subject domain is relatively low. The job of banking subsystem A is to complete payment transactions, so banking subsystem A has a high correlation with the event subject domain and a relatively low correlation with the banking subject domain.

[0223] Therefore, under the banking subject domain and the event subject domain, the weight assignment results for the five banking subsystems are as follows:

[0224] Bank subject areas: (B,5), (C,4), (A,1), (D,1), (E,1);

[0225] Event subject domains: (B,1), (C,2), (A,2), (D,2), (E,2);

[0226] The subject domain (B,5) means that the weight of system B is 5 under the subject domain of banking.

[0227] 3. Determine the correspondence between secondary subject heading numbers and subject areas. The specific process is to first determine the correspondence between secondary subject heading numbers and primary classifications, and then, based on the correspondence between primary classifications and subject areas, determine the correspondence between secondary subject heading numbers and subject areas.

[0228] Under the secondary subject term ##1, the secondary keywords include (credit business, 0.1), (instant transfer, 0.1), (pledged financing, 0.1), (cross-border currency payment, 0.05), (end-of-day processing, 0.005), (abnormal handling, 0.005), and (payment-related business, 0.2). Among them, the secondary keyword "payment-related business" has the same name as the primary category "payment-related business". Therefore, the secondary subject term ##1 and the primary category "payment-related business" are very likely to have a corresponding relationship.

[0229] In the banking subsystem, payment-related transactions are defined as transactions initiated and received by participants through the large-value payment system and settled through the clearing account management system. All other secondary keywords of secondary subject term ##1 conform to this definition. Therefore, the correspondence between secondary subject term ##1 and the primary category of payment-related transactions can be determined. The primary category of payment-related transactions corresponds to the event subject domain, thus the correspondence between secondary subject term ##1 and the event subject domain can also be determined. Similarly, it can be inferred that secondary subject term @@1 in banking subsystem E also corresponds to the primary category of payment-related transactions and the event subject domain.

[0230] Then select the name of each secondary subject term number in the same way as selecting the name of the primary subject term number. The selection results are as follows: the name of secondary subject term number ##1 is payment-related business, and the name of secondary subject term number @@1 is payment-related business.

[0231] 4. Based on the weight of each bank subsystem under the event topic domain, and the correspondence between the names of the secondary topic terms and the bank subsystems, update the weight of the secondary keywords, and merge the secondary keywords that correspond to the same primary category and have the same names of the secondary topic terms.

[0232] Subsystem A's secondary keyword number ##1 and Subsystem E's secondary keyword number @@1 both correspond to the primary category of payment business and the event topic domain, and their keyword numbers have the same name. Subsystem A has a weight of 2 in the event topic domain, and Subsystem E has a weight of 2 in the event topic domain. Multiplying the weights of each subsystem by the weights of the secondary keywords, the weights of the secondary keywords for secondary keyword number ##1 are updated as follows: (Credit business, 0.2), (Instant transfer, 0.2), (Pledge). Financing, 0.2), (Cross-border currency payments, 0.1), (End-of-day processing, 0.01), (Abnormal handling, 0.01), (Payment-related business, 0.4), etc., under the secondary subject term number @@1, the secondary keywords are (Credit business, 0.2), (Debit business, 0.2), (Real-time business, 0.2), (Parameter management, 0.1), (Batch business, 0.1), (End-of-day processing, 0.0002), (Abnormal handling, 0.0002), (Payment-related business, 0.4), etc.

[0233] The keywords listed in systems A and E for the primary category of payment business are merged and their weights are added together. After the update, the weights of the secondary keywords are as follows: (Credit business, 0.4), (Instant transfer, 0.2), (Pledged financing, 0.2), (Debit business, 0.2), (Real-time business, 0.2), (Parameter management, 0.1), (Batch business, 0.1), (Cross-border currency payment, 0.1), (End-of-day processing, 0.0052), (Abnormal handling, 0.0042), (Payment business, 0.8), etc.

[0234] 5. Finally, select the name of the second-level subject term number as the second-level category under the first-level category. When selecting, the parallel relationship between the second-level subject term names should be considered, that is, the set of second-level categories can cover the first-level category and there is mutual exclusion between the second-level categories. In addition, the weight of the updated second-level subject term corresponding to the name of the second-level subject term number should also be considered.

[0235] 6. Save the above information. Figure 9 In the table.

[0236] In some implementations, the process for constructing business attribute categories is as follows:

[0237] 1. First, determine the named entity objects and named entities. In the event subject domain given in the previous embodiment, the first-level category of payment business includes credit business, instant transfer, pledge financing, debit business, and cross-border currency payment business. These five categories are used as named entities. In the banking subject domain, payment participating institutions are used as the first-level category. Its second-level categories include large-amount participating institutions, small-amount participating institutions, and online participating institutions. These three categories are used as named entity objects. Named entities are usually the banking business itself, while named entity objects are the objects served by the banking business.

[0238] 2. Annotate the business requirements and related documents of the five banking subsystems A, B, C, D, and E under the central bank's payment system according to the requirements of the named entity recognition model.

[0239] 3. Using the business requirements, related documents, and requirements specifications that have been annotated as described above, train the named entity model. Then, input the business requirements, related documents, and requirements specifications into the trained named entity recognition model according to the bank subsystems, and output the word segmentation list of each named entity in each bank subsystem.

[0240] 4. Calculate the results of the word segmentation list for each named entity in each bank subsystem, and the number of systems involved in the word segmentation in the word segmentation list. Taking the named entity payment business as an example, the specific process is as follows: through the above steps, it is determined that the word segmentation list of the payment business contains the word "transfer". "Transfer" appears 10 times in bank subsystem A, 8 times in bank subsystem B, 0 times in bank subsystem C, 0 times in bank subsystem D, and 0 times in bank subsystem E. Therefore, the word segmentation "transfer" involves a total of 2 bank subsystems, and the number of bank subsystems involved in the word segmentation "transfer" is 2.

[0241] 5. Based on the relevance of each banking system to each subject domain, weights are assigned to each banking system under each subject domain. The assigned weight values ​​are constructed in the same way as those constructed when building the secondary classification, and the assigned values ​​are also the same. Therefore, the weight assignment results for the five banking subsystems under the banking subject domain and the event subject domain are as follows:

[0242] Bank subject areas: (B,5), (C,4), (A,1), (D,1), (E,1);

[0243] Event subject domains: (B,1), (C,2), (A,2), (D,2), (E,2);

[0244] 6. Add the weights of the involved bank subsystems to obtain the word segmentation weight. Since payment business belongs to the event subject domain and involves bank subsystems A and B, the word segmentation weight can be calculated as 2+1, which is 3, based on the system weights under the event subject domain. Then, compare the word segmentation weights of each word under the same named entity and select the word segmentation weight with the largest weight as the business attribute classification of the named entity. Since the named entity is essentially a secondary classification, the business attribute classification corresponding to the secondary classification is obtained.

[0245] 7. Store the above information to Figure 4 In the table.

[0246] The definition of technical attribute classification is: the book storage database table and fields corresponding to the data items of the subject domain hierarchical classification.

[0247] In some embodiments, the process for constructing the technical attribute classification relationship diagram is as follows:

[0248] 1. First, construct the nodes and relationships according to the bank subsystem.

[0249] A payment transaction is processed by System E. The message information for this transaction is stored in the Archive table BP04, the transaction details are stored in the Payment Transaction table BP01, the message type of this transaction is within the Message Type table BP02, and the participating entities for this transaction are within the Participating Institutions table BP03. This results in... Figure 11 Sino-US relations.

[0250] 2. Then, construct a technical attribute classification relationship diagram across bank subsystems. This mainly involves building a relationship diagram of tables and fields between bank subsystems, such as the relationship between the participating institution table BP03 of bank subsystems A, B, C, and E, for example... Figure 12 The participating institution table BP03 is managed by Bank Subsystem B and broadcast to Bank Subsystem A and Bank Subsystem E through Bank Subsystem C. The participating institution table BP03 of Bank Subsystem B is the same as that of Bank Subsystem C. The participating institution tables BP03 of Bank Subsystem A and Bank Subsystem E are subsets of the participating institution table BP03 of Bank Subsystem C. There are also relationships between the fields of the parameter table BP05 of Bank Subsystem C, Bank Subsystem E and Bank Subsystem D. The participation code field of the parameter table BP05 of Bank Subsystem C is broadcast to the use in Bank Subsystem D and Bank Subsystem E, which are reflected in the BP06 table and BP07 table, respectively.

[0251] In some embodiments, the process of adding a business attribute category is as follows:

[0252] 1. Based on Figure 10The information in the table determines the bank subsystems involved in this business attribute classification.

[0253] 2. Select the bank subsystems with higher weights and more words in the business attribute classification. Identify the database tables related to the subject domain hierarchical classification in this bank subsystem. Analyze the relationship between the database tables of this bank subsystem and other bank subsystems from the technical attribute classification relationship diagram. Bank subsystem B is selected. The participating institution table BP03 of bank subsystem B is related to the basic information of large participating institutions. There is an inheritance relationship between the BP03 table of bank subsystem B and bank subsystems A and C.

[0254] 3. Analyze the technical attribute classification and related fields, define the field range according to the definition and description of the business attribute classification, determine the valid fields by referring to the relationship between fields and the weight of the bank subsystem, and determine the mapping relationship with the business attribute classification. In addition, staff also need to consider, based on experience and expertise, whether the fields in other tables of the relevant bank subsystem meet the requirements.

[0255] The basic information fields for large-scale participating institutions are fields 1, 2, and 3. Based on the relationships between the fields of the three bank subsystems A, B, and C, and the weight of bank subsystem B, it was determined that only three fields from bank subsystem B are needed as attributes of the basic information. In addition, after assessment by multiple staff members, field 4 from bank subsystem A should also be included as basic information. The following is a graphical representation. Figure 13 :

[0256] 4. Finally, Figure 7 Information stored in Figure 14 The table shown contains the following stored results:

[0257]

[0258] Based on the same inventive concept, this application also provides an apparatus for constructing a banking data asset catalog, which can be used to implement the method described in the above embodiments, as described in the following embodiments. Since the principle of the banking data asset catalog construction apparatus in solving the problem is similar to that of the customer review information analysis method, the implementation of the banking data asset catalog construction apparatus can refer to the implementation of the software performance benchmark determination method, and will not be repeated. As used below, the terms "unit" or "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the system described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0259] Figure 15 This is a schematic diagram of the structure of a device for constructing a banking data asset catalog according to the present invention, as shown below. Figure 15As shown, based on the above embodiments, the banking data asset catalog construction device provided in this embodiment of the invention further includes a subject domain hierarchical unit 1501, a classification construction unit 1502, a relationship diagram construction unit 1503, and a classification addition unit 1504.

[0260] Subject area classification unit 1501 is used to obtain subject area classification for relevant documents based on the business requirements of the banking subsystem.

[0261] Specifically, the subject domain hierarchical unit 1501 is used to construct subject domain hierarchical classifications. The subject domain hierarchical classifications include three levels, from top to bottom: subject domain, primary classification, and secondary classification. When constructing subject domain hierarchical classifications, the terminology must be standardized and conform to industry standards.

[0262] In the subject domain classification, the subject domain can be selected from the following ten banking subject domains: parties, products, agreements, events, assets, finance, institutions, regions, marketing, and channels. Since the parties in the banking subsystem are banks, the party subject domain is changed to bank.

[0263] After determining the subject domains, based on the relevant documents of the business requirements of the bank subsystem, the subject term extraction model is used to obtain the primary and secondary categories. Then, the correspondence between the subject domains and the primary categories, and the correspondence between the primary and secondary categories are determined, and finally the construction of the subject domain hierarchical classification is completed.

[0264] Classification building unit 1502 is used to build business attribute classifications based on the business requirements of the banking subsystem, related documents of the business requirements, and subject domain hierarchical classification.

[0265] Specifically, the classification construction unit 1502 is used to construct business attribute classifications. The process of constructing business attribute classifications consists of three steps:

[0266] First, determine the named entity. The named entity should use the standardized name in the secondary category, such as credit business, instant transfer, or cross-border currency payment business, instead of non-standard terms such as loan, money transfer, or cross-border institution.

[0267] Then, the correspondence between named entities and bank subsystems is determined. When confirming the correspondence, the number of words segmented in the word segmentation list of each named entity in each bank subsystem is used to determine the corresponding named entity of that bank subsystem.

[0268] Finally, the word segmentation lists of the same named entity are merged across bank subsystems. In the merged word segmentation list, the word segmentation that appears most frequently in the text is selected as the business attribute category of this named entity. Based on the correspondence between the word segmentation and the named entity, the correspondence between the named entity and the business attribute category is determined.

[0269] The relationship diagram construction unit 1503 is used to construct a technical attribute classification relationship diagram based on the architecture and functions of the bank subsystem. The technical attribute classification relationship diagram includes metadata relationships within the bank subsystem and metadata relationships across systems.

[0270] Specifically, the relationship graph construction unit 1503 is used to construct the technical attribute classification relationship graph. The technical attribute classification is defined as the data corresponding to the hierarchical classification of the subject domain. This data is stored in the form of tables or fields. When constructing the technical attribute classification relationship graph, it is constructed using a triple <node, relation, node> structure.

[0271] The category addition module 1504 is used to add the business attribute categories of the banking subsystem to the technical attribute category relationship diagram.

[0272] Specifically, the category addition module 1504 is used to determine the valid information in the relevant fields and database tables based on the relevant fields and database table information of the technical attribute category and the weight of its banking subsystem, and to determine the correspondence between the valid information and the business attribute category. Based on the above correspondence, the business attribute category is added to the technical attribute category relationship diagram, and finally the metadata and metadata relationship are converted into a table.

[0273] The data asset catalog constructed using the aforementioned device includes subject domain hierarchical classification, technical attribute classification, and business attribute classification. The relationship between subject domain hierarchical classification and business attribute classification, as well as technical attribute classification, is displayed in the form of a diagram, making the relationships between data more intuitive and easy to understand, facilitating cross-industry data sharing, and enhancing the value of the data asset catalog.

[0274] Figure 16 This is a schematic diagram of a device for constructing a banking data asset catalog according to another embodiment of the invention, as shown below. Figure 16 As shown, based on the above embodiments, the banking data asset catalog construction apparatus provided in this embodiment further includes a document collection unit 1505, a primary subject heading numbering unit 1506, and a secondary subject heading numbering unit 1507, wherein:

[0275] Document collection unit 1505 is used to collect relevant documents regarding the business requirements of the banking subsystem;

[0276] The first-level subject term numbering unit 1506 is used to input the business requirements of all bank subsystems and the related documents of the business requirements into the pre-acquired subject term extraction model to obtain multiple first-level subject term numbers, the first-level keywords of each first-level subject term number, the number of documents of each first-level subject term number, and the weight of each first-level keyword in the corresponding first-level subject term number.

[0277] Specifically, in the existing technology, the subject term extraction model cannot output subject terms, so the first-level subject term numbering unit 1602 can only output the results in the form of subject term numbers. At the same time, there can be multiple first-level keywords in one first-level subject term number. When obtaining information such as first-level subject term numbers and first-level keywords, the relevant documents of the business requirements of all bank subsystems are simultaneously input into the subject term extraction model.

[0278] The secondary subject term numbering unit 1507 is used to input the business requirements of each bank subsystem and the related documents of the business requirements into the subject term extraction model to obtain the secondary subject term number of each bank subsystem, the secondary keywords of the secondary subject term number of each bank subsystem, the number of documents of the secondary subject term of each bank subsystem, and the weight of each secondary keyword in the corresponding secondary subject term number.

[0279] Specifically, the secondary subject term numbering unit 1507 is used to obtain the secondary subject term number and secondary keywords. During the acquisition process, the relevant documents of the business requirements of the bank subsystem are input into the subject term extraction model according to the bank subsystem.

[0280] The above units provide the primary subject heading numbers and their keyword information, as well as the secondary subject heading numbers and their keyword information, laying the groundwork for determining the primary and secondary classifications.

[0281] Figure 17 This is a schematic diagram of a device for constructing a banking data asset catalog according to another embodiment of the invention, as shown below. Figure 17 As shown, based on the above embodiments, the subject domain hierarchical unit 1501 provided in this embodiment of the invention further includes:

[0282] The topic domain determination module 1701 is used to determine the topic domain corresponding to the first-level keyword number based on each of the first-level keywords and the weight of the first-level keywords.

[0283] Specifically, the topic domain determination module 1701 determines the topic domain corresponding to the first-level topic keyword number by first defining each topic domain, and then comparing the relationship between the definition of each topic domain and the meaning of the keywords in the first-level topic keyword number to determine whether the topic domain corresponds to the first-level topic keyword number. When making the judgment, the first-level keywords with higher weight are given priority.

[0284] The primary subject term name selection module 1702 is used to select a primary subject term as the name of the primary subject term number based on the parallel relationship between the primary keywords of each primary subject term number and the weight of the primary keywords.

[0285] Specifically, when selecting the first-level subject term name, the first-level category selection module 1702 should consider the weight of the first-level keywords and the parallel relationship between the first-level keywords. The meaning of the selected first-level keywords should cover the meaning of all the first-level keywords in the first-level subject term name as much as possible. In addition, the greater the weight of the first-level keyword, the more likely it is to be selected as the first-level subject term name.

[0286] The primary category selection module 1703 is used to select the name of the primary subject term number as the name of the primary category of the subject domain based on the parallel relationship between the primary subject term number names corresponding to the same subject domain.

[0287] Specifically, when selecting the names of first-level subject headings as first-level categories, the parallel relationships between these names must be comprehensively considered. A subject domain may contain multiple first-level categories; therefore, the selected first-level categories must be semantically mutually exclusive, and the set of selected first-level categories must cover the entire subject domain.

[0288] The weight assignment module 1704 is used to assign weights to each of the bank subsystems under each of the subject domains based on the relevance between each of the bank subsystems and each of the subject domains.

[0289] Specifically, the weighting module 1704 assigns weights to each bank subsystem under each subject domain based on the obvious correspondence between different bank subsystems in the subject domain. For example, the working goal of bank subsystem X is bank management, while the working goal of bank subsystem Y is to complete payment transactions. Therefore, bank subsystem X has a deeper connection with the bank subject domain, while bank subsystem Y has a deeper connection with the event subject domain. Under the bank subject domain, the weight assigned to system X is greater than the weight assigned to system Y.

[0290] The primary category determination module 1705 is used to determine the primary category corresponding to the secondary keyword number based on each of the secondary keywords and the weight of the secondary keywords.

[0291] Specifically, the process for determining the primary category corresponding to the secondary subject term number is the same as the process for determining the subject area corresponding to the primary subject term number. First, the definition of each primary category is considered. Then, by comparing the definition of each primary category with the meaning of the keywords in the secondary subject term number, the correspondence between the secondary subject term number and the primary category is obtained.

[0292] The topic domain acquisition module 1706 is used to acquire the topic domain corresponding to the second-level topic word number based on the determined correspondence between the first-level classification and the topic domain.

[0293] Specifically, the correspondence between the primary classification and the subject is used as a bridge to obtain the subject domain corresponding to the secondary subject term number.

[0294] The secondary keyword name selection module 1707 is used to select one of the secondary keywords as the name of the secondary keyword number based on the weight and parallel relationship of the secondary keywords of each secondary keyword number.

[0295] Specifically, when selecting secondary keywords as names for secondary subject headings, it is necessary to comprehensively consider the weight of the secondary keywords and the parallel relationship between them. The meaning of the selected secondary keywords should cover the meaning of all keywords in the secondary subject heading as much as possible.

[0296] The secondary keyword weight update module 1708 is used to merge secondary keywords that correspond to the same primary category and have the same secondary subject term number according to the merging rules, so as to update the weight of the secondary keywords.

[0297] Specifically, the secondary subject term numbers are obtained by inputting the business requirements and related documents of the banking subsystems into the subject term extraction model according to different banking subsystems. Therefore, the secondary subject term numbers and their secondary keywords all correspond to the banking subsystems.

[0298] The secondary category selection module 1709 is used to select the name of the secondary category based on the weight of the secondary keyword corresponding to the name of the secondary keyword number, the parallel relationship of the names of the secondary keyword numbers, and the primary category corresponding to the secondary keyword number.

[0299] Specifically, first, the definition of the primary category corresponding to the secondary keywords is determined. Then, the vocabulary that conforms to the banking industry standardization among the merged secondary keywords is selected as the secondary category under the primary category. There can be multiple secondary categories under the same primary category.

[0300] Based on the above embodiments, the classification construction unit 1502 provided in this embodiment of the invention further includes:

[0301] The Named Entity Determination module is used to determine named entities, which are nouns in a second-level category.

[0302] Specifically, the named entities determined by the named entity determination module are words involved in the secondary categories. When determining named entity objects, it is important to distinguish between named entities and named entity objects. Taking the event subject domain and the bank subject domain as examples, the secondary categories of the event subject domain include credit business, instant transfer business, pledge financing business, debit business, and cross-border currency payment business. These five secondary categories are named entities. The secondary categories of the bank subject domain include large-scale participating institutions, small-scale participating institutions, and online participating institutions. These three secondary categories are named entity objects.

[0303] The annotation module is used to annotate relevant documents of the business requirements of the banking subsystem according to the annotation requirements of named entities and named entity recognition models.

[0304] Specifically, the annotation module is used to annotate documents related to the business requirements of the banking subsystem, while the named entity recognition model is used to identify named entities defined in the corpus. The technologies used include rule-based and dictionary-based methods, statistical methods, and hybrid methods of the above two. Before using the named entity model, the input data needs to be annotated.

[0305] The model training module is used to train the named body recognition model using relevant documents of the business requirements of the bank's subsystem that have been annotated.

[0306] Specifically, the first step is to organize the business requirements and related documents for each bank subsystem. Then, the named entity recognition model is trained according to the annotation requirements of the model. The goal of the named entity recognition model is to identify named entities in the corpus. The main methods are divided into three categories: rule-based and dictionary-based methods, statistical methods, and hybrid methods.

[0307] Figure 18 This is a schematic diagram of a device for constructing a banking data asset catalog according to another embodiment of the invention, as shown below. Figure 18 As shown, based on the above embodiments, the classification construction unit 1502 provided in this embodiment of the invention further includes:

[0308] The word segmentation list extraction module 1801 is used to input the business requirements of the bank subsystem and the relevant documents required by the business into the trained named entity recognition model to obtain a word segmentation list of each named entity.

[0309] The quantity statistics module 1802 is used to count the number of bank subsystems involved in each of the word segments.

[0310] Specifically, the occurrence of each word in the documents of the banking subsystem is counted. If a word appears once in the documents of the banking subsystem, regardless of how many times it appears in the banking subsystem afterward, the number of banking subsystems involved in the word is increased by 1.

[0311] The system weight module 1803 is used to assign weights to each of the bank subsystems under each of the subject domains based on the relevance between each of the bank subsystems and each of the subject domains.

[0312] The word segmentation weight calculation module 1804 is used to multiply the number of bank subsystems involved in the word segmentation by the weight of the bank subsystem to obtain the word segmentation weight.

[0313] The business attribute classification determination module 1805 is used to determine the business attribute classification of the named entity based on the word segmentation weight of each named entity in the same banking subsystem, and then determine the business attribute classification of the secondary category corresponding to the named entity.

[0314] Specifically, from the word segmentation list of each named entity, the word segmentation with the higher proportion is selected as the business attribute category of the named entity. Since the name of the named entity is taken from the secondary category, the business attribute category of the secondary category is further obtained.

[0315] Figure 19 This is a schematic diagram of a device for constructing a banking data asset catalog according to another embodiment of the invention, as shown below. Figure 19 As shown, based on the above embodiments, the classification and addition unit 1504 provided in this embodiment of the invention further includes:

[0316] The bank subsystem scope determination module 1901 is used to determine the scope of bank subsystems included in the business attribute classification.

[0317] The bank subsystem selection module 1902 is used to select the bank subsystem with higher weight and larger total number of word segments in the bank subsystem.

[0318] Specifically, the bank subsystem selects module 1902 through... Figure 10 The table in the document filters out the banking subsystems with higher weights and larger total number of word segments. Systems with higher weights and larger total number of word segments contain more information than other systems, so this system was selected as the object of analysis.

[0319] The bank subsystem analysis module 1903 is used to analyze the relationship between the selected bank subsystem and other bank subsystems from the technical attribute relationship diagram.

[0320] Specifically, when constructing the technical attribute classification relationship diagram, the banking subsystem analysis module 1903 uses a triplet <node, relationship, node> structure to record the relationship between the selected banking subsystem and other banking subsystems.

[0321] The classification addition module 1904 is used to analyze the technical attributes and related fields from the perspective of the business attribute classification, determine the mapping relationship between the business attributes and the technical attributes, and add the business attribute classification of the bank subsystem to the technical attribute relationship diagram.

[0322] Specifically, the classification addition module 1904 finds the corresponding database tables and fields based on the definition of business attribute classification. It refers to the relationship between fields and the weight of the bank subsystem to determine the valid fields and the mapping relationship between the valid fields and the business attribute classification. In addition, it also needs to consider, based on experience and expertise, whether the fields in other tables of the relevant bank subsystem meet the definition. Finally, the technical attribute classification relationship diagram containing business attribute relationships is converted into a table for storage.

[0323] This embodiment discloses a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer can execute the methods provided in the above-described method embodiments, such as: obtaining a subject domain hierarchical classification based on relevant documents of the business requirements of a bank subsystem; constructing a business attribute classification based on the business requirements of the bank subsystem, the relevant documents of the business requirements, and the subject domain hierarchical classification; constructing a technical attribute classification relationship diagram according to the architecture and functions of the bank subsystem, wherein the technical attribute classification relationship diagram includes metadata relationships within the bank subsystem and cross-system metadata relationships; and adding the business attribute classification of the bank subsystem to the technical attribute classification relationship diagram.

[0324] This embodiment provides a computer-readable storage medium storing a computer program that causes a computer to execute the methods provided in the above-described method embodiments. For example, the methods include: obtaining a subject domain hierarchical classification based on relevant documents concerning the business requirements of a banking subsystem; constructing a business attribute classification based on the business requirements of the banking subsystem, the relevant documents concerning the business requirements, and the subject domain hierarchical classification; constructing a technical attribute classification relationship diagram according to the architecture and functions of the banking subsystem, wherein the technical attribute classification relationship diagram includes metadata relationships within the banking subsystem and cross-system metadata relationships; and adding the business attribute classification of the banking subsystem to the technical attribute classification relationship diagram.

[0325] In the description of this specification, the references to terms such as "an embodiment," "a specific embodiment," "some embodiments," "for example," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0326] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. A method of constructing a banking data asset catalog, characterized by, include: Based on the relevant documents concerning the business requirements of the banking subsystem, a subject domain hierarchical classification was obtained; Based on the business requirements of the banking subsystem, the relevant documents of the business requirements, and the hierarchical classification of the subject domains, a business attribute classification is constructed. Based on the architecture and functions of the bank subsystem, a technical attribute classification relationship diagram is constructed, wherein the technical attribute classification relationship diagram includes metadata relationships within the bank subsystem and cross-system metadata relationships; Add the business attribute categories of the banking subsystem to the technical attribute category relationship diagram; Before obtaining the subject domain hierarchical classification of the relevant documents based on the business requirements of the banking subsystem, the following are also included: Collect relevant documents regarding the business requirements of the aforementioned banking subsystem; Input the relevant documents of the business requirements of all the bank subsystems into the pre-acquired keyword extraction model to obtain multiple first-level keyword numbers, first-level keywords of each first-level keyword number, the number of documents of each first-level keyword number, and the weight of each first-level keyword in the corresponding first-level keyword number; The relevant documents of the business requirements of each of the bank subsystems are input into the keyword extraction model to obtain the secondary keyword number of each bank subsystem, the secondary keywords of the secondary keyword number of each bank subsystem, the number of documents of the secondary keyword of each bank subsystem, and the weight of each secondary keyword in the corresponding secondary keyword number. The relevant documents based on the business requirements of the banking subsystem are classified into subject areas, including: Based on each primary keyword and its weight, the subject area corresponding to the primary keyword number is determined; Based on the parallel relationship between the primary keywords of each primary subject term number and the weight of the primary keywords, a primary keyword is selected as the name of the primary subject term number; Based on the parallel relationship between the names of the first-level subject terms corresponding to the same subject domain, the name of the first-level subject term is selected as the name of the first-level category of the subject domain; Based on the relevance of each bank subsystem to each subject domain, weights are assigned to each bank subsystem under each subject domain; Based on each of the secondary keywords and their weights, the primary category corresponding to the secondary keyword number is determined; Based on the determined correspondence between the primary classification and the topic domain, obtain the topic domain corresponding to the secondary topic term number; Based on the weight and parallel relationship of the secondary keywords of each secondary subject term number, one of the secondary keywords is selected as the name of the secondary subject term number; According to the merging rules, secondary keywords that correspond to the same primary category and have the same secondary subject term number are merged to update the weight of the secondary keywords; Based on the weight of the secondary keywords corresponding to the names of the secondary subject terms, the parallel relationship of the names of the secondary subject terms, and the primary category corresponding to the secondary subject terms, the names of the secondary subject terms are selected as the names of the secondary categories.

2. The method for constructing a bank data asset catalog according to claim 1, characterized in that, The merging rules are as follows: The weight of the secondary keyword is multiplied by the weight of its corresponding bank subsystem to obtain the weight of the new secondary keyword. The secondary keywords in the secondary subject terms that correspond to the same primary category and have the same name are added together with the corresponding secondary keywords according to the weight of the new secondary keywords.

3. The method of claim 2, wherein, Before constructing the business attribute classification based on the business requirements of the banking subsystem, the relevant documents of the business requirements, and the subject domain hierarchical classification, the following further steps are included: Identify named entities, where the named entities are nouns in the second-level category; The relevant documents concerning the business requirements of the banking subsystem are annotated according to the annotation requirements of the named entities and named entity recognition models. The named body recognition model was trained using the relevant documents of the business requirements of the bank's subsystems that had been annotated.

4. The method of claim 3, wherein, Based on the business requirements of the banking subsystem, related documents of the business requirements, and the hierarchical classification of the subject domains, a business attribute classification is constructed, including: Input the business requirements of the bank subsystem that has been annotated and the relevant documents required by the business into the trained named entity recognition model to obtain a word segmentation list of each named entity; The number of bank subsystems involved in each of the segmented words in the word segmentation list; Based on the relevance of each bank subsystem to each subject domain, weights are assigned to each bank subsystem under each subject domain; The word segmentation weight is obtained by multiplying the number of bank subsystems involved in the word segmentation by the weight of the bank subsystem and summing them up by system. The business attribute classification of the named entity is determined based on the word segmentation weight of each named entity in the same banking subsystem.

5. The method of claim 4, wherein, Adding the business attribute categories of the banking subsystem to the technical attribute category relationship diagram includes: Determine the scope of the banking subsystems included in the business attribute classification; Among the banking subsystems, select the one with higher weight and a larger total number of word segments; Analyze the relationship between the selected bank subsystem and other bank subsystems from the technical attribute classification relationship diagram; Analyze the technical attribute classification and related fields from the perspective of the business attribute classification, determine the mapping relationship between the business attributes and the technical attribute classification, and add the business attribute classification of the bank subsystem to the technical attribute classification relationship diagram.

6. A device for constructing a banking data asset catalog, characterized by, include: Subject area hierarchical unit, used to obtain subject area hierarchical classification of relevant documents based on the business requirements of the banking subsystem; A classification construction unit is used to construct business attribute classifications based on the business requirements of the banking subsystem, related documents of the business requirements, and the subject domain hierarchical classification. The relationship diagram construction unit is used to construct a technical attribute classification relationship diagram based on the architecture and functions of the bank subsystem, wherein the technical attribute classification relationship diagram includes metadata relationships within the bank subsystem and cross-system metadata relationships; The category addition module is used to add the business attribute categories of the banking subsystem to the technical attribute category relationship diagram. Also includes: The document collection unit is used to collect relevant documents regarding the business requirements of the banking subsystem. The first-level subject term numbering unit is used to input the business requirements of all banking subsystems and the related documents of the business requirements into the pre-acquired subject term extraction model to obtain multiple first-level subject term numbers, the first-level keywords of each first-level subject term number, the number of documents of each first-level subject term number, and the weight of each first-level keyword in the corresponding first-level subject term number. The secondary subject term numbering unit is used to input the business requirements and related documents of each bank subsystem into the subject term extraction model to obtain the secondary subject term number of each bank subsystem, the secondary keywords of the secondary subject term number of each bank subsystem, the number of documents of the secondary subject term of each bank subsystem, and the weight of each secondary keyword in the corresponding secondary subject term number; The subject area hierarchical unit includes: The topic domain determination module is used to determine the topic domain corresponding to the first-level keyword number based on each first-level keyword and the weight of the first-level keyword; The primary subject keyword name selection module is used to select a primary keyword as the name of the primary subject keyword number based on the parallel relationship between the primary keywords of each primary subject keyword number and the weight of the primary keywords. The primary category selection module is used to select the name of the primary subject term number as the name of the primary category of the subject domain based on the parallel relationship between the names of the primary subject term numbers corresponding to the same subject domain; The weight assignment module is used to assign weights to each of the bank subsystems under each of the subject domains based on the relevance of each of the bank subsystems to each of the subject domains. The primary category determination module is used to determine the primary category corresponding to the secondary keyword number based on each secondary keyword and its weight. The topic domain acquisition module is used to acquire the topic domain corresponding to the second-level topic word number based on the determined correspondence between the first-level category and the topic domain; The secondary keyword name selection module is used to select one of the secondary keywords as the name of the secondary keyword number based on the weight and parallel relationship of the secondary keywords of each secondary keyword number; The secondary keyword weight update module is used to merge secondary keywords that correspond to the same primary category and have the same secondary topic word number according to the merging rules, so as to update the weight of the secondary keywords; The secondary category selection module is used to select the name of the secondary category based on the weight of the secondary keyword corresponding to the name of the secondary keyword number, the parallel relationship of the names of the secondary keyword numbers, and the primary category corresponding to the secondary keyword number.

7. The apparatus of claim 6, wherein, The merging rules are as follows: The weight of the secondary keyword is multiplied by the weight of its corresponding bank subsystem to obtain the weight of the new secondary keyword. The secondary keywords in the secondary topic terms corresponding to the same primary category are added together with the corresponding secondary keywords according to the weight of the new secondary keywords.

8. The apparatus of claim 7, wherein, Also includes: A named entity determination module is used to determine named entities, wherein the named entities are nouns in the second-level category; The annotation module is used to annotate the relevant documents of the business requirements of the bank subsystem according to the annotation requirements of the named entities and the named entity recognition model; The model training module is used to train the named body recognition model using relevant documents of the business requirements of the bank's subsystem that have been annotated.

9. The apparatus of claim 8, wherein, The classification construction unit includes: The word segmentation list extraction module is used to input the business requirements of the bank subsystem and the relevant documents required by the business into the trained named entity recognition model to obtain the word segmentation list of each named entity. The quantity statistics module is used to count the number of bank subsystems involved in each of the word segments; The system weight module is used to assign weights to each of the bank subsystems under each of the subject domains based on the relevance of each of the bank subsystems to each of the subject domains. The word segmentation weight calculation module is used to multiply the number of bank subsystems involved in the word segmentation by the weight of the bank subsystem to obtain the word segmentation weight; The business attribute classification determination module is used to determine the business attribute classification of the named entity based on the word segmentation weight of each named entity in the same banking subsystem, and then determine the business attribute classification of the secondary category corresponding to the named entity.

10. The apparatus of claim 9, wherein, The category addition module includes: The bank subsystem scope determination module is used to determine the scope of bank subsystems included in the business attribute classification; A bank subsystem selection unit is used to select the bank subsystem with higher weight and a larger total number of word segments from the bank subsystem. The bank subsystem analysis module is used to analyze the relationship between the selected bank subsystem and other bank subsystems from the technical attribute relationship diagram; The classification addition module is used to analyze the technical attributes and related fields from the perspective of the business attribute classification, determine the mapping relationship between the business attributes and the technical attributes, and add the business attribute classification of the bank subsystem to the technical attribute relationship diagram.

11. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.

12. A computer readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Theme extraction method and device oriented to science and technology requirements and storage medium

    CN113255340A

  • Automatic data asset checking method and system

    CN113792081A

  • Large watershed hydropower enterprise data governance method

    CN114996247A