A data labeling method and device based on data elements, and a data processing method and device
By using data element-based data labeling and processing methods, the challenges of cross-industry data cleaning and governance have been solved, cross-industry data governance standards have been established, complexity has been reduced, governance progress has been accelerated, and data value has been released.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIAMEN MEIYA PICO INFORMATION CO LTD
- Filing Date
- 2023-05-19
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies cannot effectively achieve cross-industry data cleaning and governance, and data governance systems have limitations in cross-industry applications, failing to meet the needs of mixed industries.
We adopt a data tagging method based on data elements, which involves creating a list of element tag attributes, associating resource IDs with element tags, tagging information groups and setting priorities, and combining data processing methods to recommend data governance and analysis strategies.
It has achieved cross-industry data governance standards, reduced the complexity of data governance, accelerated the governance process, released the value of data, and facilitated the digital expression of data models and the highlighting of key resource characteristics.
Smart Images

Figure CN116701366B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data computer technology, and in particular to a data tagging method, data processing method and apparatus based on data elements. Background Technology
[0002] With the growing demand for data sharing and data mining, big data governance plays a crucial role across various industries. A key step in data governance is data element benchmarking, which is a vital means of transforming diverse and heterogeneous data into standardized data. As the application scope of data governance expands, different industries have different data standards. Therefore, the ability to quickly and effectively perform cross-industry data cleaning is particularly important.
[0003] Currently, the relevant data elements are deeply related to a specific industry or scenario, which cannot meet the needs of cross-industry and mixed-industry big data governance. Furthermore, they are data-centric, and the ability to manage and utilize methods such as models, development, and verification is still insufficient, which limits the application of big data governance systems. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a data tagging method, data processing method, and apparatus based on data elements. Data tagging (element tagging) of data elements can leverage the role of data elements. Simultaneously, by combining with data processing methods, it is easy to characterize data models and facilitate the formation of digital expressions for data resources or rule models. It highlights key resource characteristics, reduces governance complexity, and accelerates governance progress. It is suitable for cross-industry implementation and unlocks data value.
[0005] The present invention adopts the following technical solution:
[0006] Firstly, a data tagging method based on data elements includes:
[0007] Based on the acquired business information, create a list of feature tag attributes, including feature tag fields, from the preset description dimensions;
[0008] Based on the information in the resource table and the list of feature tag attributes, feature tags are created, and resource IDs are associated with feature tags.
[0009] Sequentially label the different information groups in the resource table;
[0010] Based on labeled information groups, identify and label the main entities within the information groups;
[0011] Prioritize the marking of the same subject and the same element.
[0012] Preferably, the preset description dimension includes the following fields: element tag, Chinese name, object type, value range, dictionary code, data type, data meta-encoding, and processing method; the element tag is a unique identifier for a data element; the Chinese name is the designation of one or more Chinese words assigned to the element tag; the object type is used to mark the type of the described subject object; the value range is the set of allowed values for meta-class elements determined according to the data type and representation format specified in the corresponding attributes; the dictionary code is the number of the dictionary code set followed; the data type is the data type that identifies the element; the data meta-encoding is the data meta-encoding corresponding to the industry; and the processing method is the data processing rule corresponding to the element tag.
[0013] Preferably, the information groups represent related markers of the same entity object; the same information groups are marked with the same Arabic numerals, and the marker order is divided according to the importance of the subject.
[0014] Preferably, the subject is a field used to determine the description of the resource subject object, and the subject may not be unique.
[0015] Secondly, a data tagging device based on data elements includes:
[0016] The Element Tag Attribute List Creation Module is used to create a list of element tag attributes, including element tag fields, from a preset description dimension based on the acquired business information.
[0017] The feature tagging module is used to tag features based on the information in the resource table and the feature tag attribute list, and to associate resource IDs with feature tags.
[0018] The information group tagging module is used to sequentially tag different information groups in the resource table;
[0019] The subject tagging module is used to identify and tag subjects within a group of information based on the tags.
[0020] The priority tagging module is used to tag the priority of tags for the same subject and the same element.
[0021] Thirdly, a data processing method based on feature labeling includes:
[0022] Based on the acquired business information, create a list of feature tag attributes, including feature tag fields, from the preset description dimensions;
[0023] Based on the information in the resource table and the list of feature tag attributes, feature tags are created, and resource IDs are associated with feature tags.
[0024] Sequentially label the different information groups in the resource table;
[0025] Based on labeled information groups, identify and label the main entities within the information groups;
[0026] Prioritize marking items with the same subject and the same elements;
[0027] The system associates feature tags in the resource table with at least one preset description dimension in the feature tag attribute list and recommends configuration information for data governance strategies.
[0028] Preferably, the preset description dimension includes the following fields: feature tag, Chinese name, object type, value range, dictionary code, data type, data element encoding, and processing method;
[0029] The data governance strategies include spam filtering, format cleaning, or deduplication.
[0030] Fourthly, a data processing apparatus based on feature labeling includes:
[0031] The Element Tag Attribute List Creation Module is used to create a list of element tag attributes, including element tag fields, from a preset description dimension based on the acquired business information.
[0032] The feature tagging module is used to tag features based on the information in the resource table and the feature tag attribute list, and to associate resource IDs with feature tags.
[0033] The information group tagging module is used to sequentially tag different information groups in the resource table;
[0034] The subject tagging module is used to identify and tag subjects within a group of information based on the tags.
[0035] The priority marking module is used to mark the priority of tags for the same subject and the same feature;
[0036] The data governance module is used to associate feature tags with at least one dimension attribute in the feature tag attribute list through feature tags in the resource table, and recommend configuration information for data governance strategies.
[0037] Fifthly, a data processing method based on feature labeling includes:
[0038] Based on the acquired business information, create a list of feature tag attributes, including feature tag fields, from the preset description dimensions;
[0039] Based on the information in the resource table and the list of feature tag attributes, feature tags are created, and resource IDs are associated with feature tags.
[0040] Sequentially label the different information groups in the resource table;
[0041] Based on labeled information groups, identify and label the main entities within the information groups;
[0042] Prioritize marking items with the same subject and the same elements;
[0043] Based on at least one of the following: element tagging association information group, subject, and priority, configuration information for data analysis is recommended; the data analysis includes extraction and distribution, data quality detection, theme planning, or information presentation.
[0044] Sixthly, a data processing apparatus based on feature labeling includes:
[0045] The Element Tag Attribute List Creation Module is used to create a list of element tag attributes, including element tag fields, from a preset description dimension based on the acquired business information.
[0046] The feature tagging module is used to tag features based on the information in the resource table and the feature tag attribute list, and to associate resource IDs with feature tags.
[0047] The information group tagging module is used to sequentially tag different information groups in the resource table;
[0048] The subject tagging module is used to identify and tag subjects within a group of information based on the tags.
[0049] The priority marking module is used to mark the priority of tags for the same subject and the same feature;
[0050] The data analysis module is used to associate information groups, subjects, and priorities based on feature tags, and to recommend configuration information for data analysis; the data analysis includes extraction and distribution, data quality detection, theme planning, or information presentation.
[0051] In a seventh aspect, a computer-readable storage medium stores computer program code that, when executed by a computer, performs the data element-based data tagging method or the data element-based data processing method.
[0052] The present invention has the following beneficial effects:
[0053] (1) The present invention provides a data tagging method based on data elements, which is a set of attributes for interpreting and analyzing information elements, real-world scenarios, and objective facts described by resources. It is a tagging method for interpreting the objective dimensions of data connotation. The element tagging of the present invention is a tagging method that is detached from the database table structure and focuses only on the object itself. Based on the element tagging, combined with logical operators, global retrieval can be performed to express the objective facts of behavior. It is easy for technical personnel to understand and also convenient for computers to recognize.
[0054] (2) The present invention provides a data processing method based on element tagging, which associates element tags with at least one dimension attribute in the element tag attribute list and recommends configuration information for data governance strategies. This method can realize cross-industry data governance standards, reduce the complexity of data governance, and accelerate the governance process.
[0055] (3) The present invention provides a data processing method based on element tagging, which associates information groups, subjects, and priorities based on element tags and recommends configuration information for data analysis, so as to highlight the key features of resources, help to depict the objective connotation and business meaning of data, and provide an automated and intelligent foundation for subsequent extraction and distribution, data quality detection, theme planning, information display, etc.
[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Attached Figure Description
[0057] Figure 1 This is a flowchart of a data tagging method based on data elements according to Embodiment 1 of the present invention;
[0058] Figure 2 This is a schematic diagram illustrating the relationship between element markers, data elements, and data items in Embodiment 1 of the present invention.
[0059] Figure 3 This is a schematic diagram of multiple sets of standards in the same resource table according to Embodiment 1 of the present invention;
[0060] Figure 4 This is a structural block diagram of a data tagging device based on data elements according to Embodiment 1 of the present invention;
[0061] Figure 5 This is a flowchart of the data processing method according to Embodiment 2 of the present invention;
[0062] Figure 6 This is a structural block diagram of the data processing device according to Embodiment 2 of the present invention;
[0063] Figure 7 This is a flowchart of the data processing method according to Embodiment 3 of the present invention;
[0064] Figure 8 This is a structural block diagram of the data processing device according to Embodiment 3 of the present invention;
[0065] Figure 9This is a schematic diagram of the data governance process using bank user account opening information as an example in Embodiment 4 of the present invention; wherein, (a) represents a diagram of association; and (b) represents a diagram of governance based on (a). Detailed Implementation
[0066] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0067] In the description of this invention, it should be noted that the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0068] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "provided with", "sleeved / connected", "connected", etc., should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be a connection within two components. For those skilled in the art, the specific meaning of the above terms in this invention can be understood according to the specific circumstances.
[0069] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the step identifiers S101, S102, S103, etc. are only used for convenience of description and do not indicate the execution order. The corresponding execution order can be adjusted as needed.
[0070] Example 1
[0071] See Figure 1 As shown in the figure, this embodiment of a data tagging method based on data elements includes:
[0072] S101, Based on the acquired business information, create a list of feature tag attributes, including feature tag fields, from the preset description dimension;
[0073] S102, Based on the information in the resource table and the list of feature tag attributes, perform feature tagging and associate the resource ID with the feature tag;
[0074] S103, sequentially mark different information groups in the resource table;
[0075] S104, Based on the labeled information group, identify the main body in the information group and label it;
[0076] S105, Priority of marking the same subject and the same element.
[0077] In this embodiment, the execution subject of a data tagging method based on data elements is a terminal device, such as a computer terminal or a mobile phone terminal.
[0078] Based on the big data platform, the preset description dimensions include the following fields: element tag, Chinese name, object type, value range, dictionary code, data type, data meta-encoding, and processing method. The element tag is a unique identifier for a data element; the Chinese name is the designation of one or more Chinese characters assigned to the element tag; the object type is used to identify the type of the described subject object; the value range is the set of allowed values for meta-class elements determined by the data type and representation format specified in the corresponding attributes; the dictionary code is the number of the dictionary code set followed; the data type is the data type that identifies the element; the data meta-encoding corresponds to the industry's data meta-encoding; and the processing method is the data processing rule corresponding to the element tag. Specific explanations of each field are shown in Table 1 below.
[0079] Table 1. Description Dimensions of Feature Tags
[0080]
[0081] In actual use, description dimensions can be added as appropriate according to actual needs.
[0082] In this embodiment, the information group, subject, and priority of the feature label are all labeled on the data items of the data resource. The descriptive dimensions of the feature label are combined with the information group, group subject, and priority for annotation. The specific explanations of the information group, subject, and priority are as follows.
[0083] Information group: Represents related tags for the same entity object; the same information group is tagged with the same Arabic letter (1, 2, ...), and the tagging order should be divided according to the importance of the subject. For example, in bank account opening business, the tags such as personnel identity, account opening behavior, and savings behavior should be regarded as information groups for this personnel entity object. Such resources only describe the attributes, behaviors and data level information of the entity object itself. Therefore, the attributes and behaviors related to the entity object should be tagged as a group of information (1), and the data of the bank teller should be tagged as another information group (2).
[0084] Subject: A field that identifies and describes the subject of a resource (the information subject or behavior subject described in the resource). The subject does not have to be unique. For example, in marriage registration, both the man and the woman are subjects; in bank card opening, the account holder is the subject, but the teller can also be a subject.
[0085] Priority: When the same element tags exist for the same subject object, the priority is marked with "1, 2, ...".
[0086] In this embodiment, the relationship between feature tags, data elements, and data items is described in [reference]. Figure 2 As shown.
[0087] The following points can be observed from the diagram:
[0088] (1) Feature tags and data elements have a many-to-one relationship. Depending on the standard data element, one data element can correspond to multiple feature tags. For example, "start time" corresponds to "ys-sj-ks" and "end time" corresponds to "ys-sj-js"; while "date and time" can correspond to "ys-sj, ys-sj-ks, ys-sj-js, ..." depending on the usage in different resource tables.
[0089] (2) The relationship between data elements and data items is one-to-many, meaning one data element can correspond to multiple data items. For example, in the "Resident Household Registration Information Form", fields such as "Owner's Name, Spouse's Name, Eldest Son's Name, and Eldest Daughter's Name" all correspond to the same data element "Name".
[0090] (3) When there are multiple sets of standard information such as feature tags, data elements, and data items in a resource table, and there are conflicts in the relevant configuration information (data length, data type, etc.), the processing priority is: feature tags > data elements > data items.
[0091] See Figure 3 As shown, in a bank's user account opening data table, there are three sets of standards (element tags, general data elements for bank information systems, and the standard for the data items themselves). These three sets of standards have overlapping definitions, such as value ranges, data types, and dictionaries. When a computer performs a corresponding operation, it must execute the appropriate operation according to a predefined priority.
[0092] See Figure 4 As shown, this embodiment also discloses a data tagging device based on data elements, including:
[0093] The feature tag attribute list creation module 401 is used to create a feature tag attribute list, including feature tag fields, from a preset description dimension based on the acquired business information.
[0094] The feature tagging module 402 is used to tag features based on the information in the resource table and the feature tag attribute list, and to associate resource IDs with feature tags.
[0095] Information group marking module 403 is used to mark different information groups in the resource table in sequence;
[0096] The subject tagging module 404 is used to identify and tag subjects in an information group based on the tagged information group.
[0097] Priority marking module 405 is used to mark the priority of the same subject and the same element.
[0098] A data tagging device based on data elements corresponds to a data tagging method based on data elements. The specific implementation of each module can be found in the implementation of the same data tagging method based on data elements, and will not be described again in this embodiment.
[0099] Data elements are essential components of things. Data tagging of data elements not only leverages the role of data units, but also, combined with data processing methods, facilitates the characterization of data models, enabling the formation of digital expressions for data resources or rule models; it helps highlight key resource characteristics, reduces governance complexity, and accelerates governance progress; and it is suitable for cross-industry implementation, further releasing the value of data.
[0100] Example 2
[0101] See Figure 5 As shown in this embodiment, a data processing method based on feature tags includes:
[0102] S501, Based on the acquired business information, create a list of feature tag attributes, including feature tag fields, from a preset description dimension;
[0103] S502, Based on the information in the resource table and the list of feature tag attributes, feature tags are performed, and resource IDs are associated with feature tags;
[0104] S503, sequentially mark different information groups in the resource table;
[0105] S504, Based on the labeled information group, identify the main body in the information group and label it;
[0106] S505, Priority for marking the same subject and the same element;
[0107] S506, associating feature tags in the resource table with at least one preset description dimension in the feature tag attribute list, and recommending configuration information for data governance strategies.
[0108] For the specific implementation of S501 to S505, please refer to Embodiment 1.
[0109] In S506, the preset description dimension includes the following fields: feature tag, Chinese name, object type, value range, dictionary code, data type, data element encoding, and processing method.
[0110] The system associates feature tags in the resource table with at least one preset description dimension in the feature tag attribute list and recommends configuration information for data governance strategies, specifically:
[0111] Data governance strategies can be configured based on at least one of the following descriptions: feature tag, Chinese name, object type, value range, dictionary code, data type, data element encoding, and processing method. Alternatively, information groups, subjects, and priorities can be added to these descriptions to further configure data governance strategies.
[0112] In this embodiment, the data governance strategy includes junk filtering, format cleaning, or deduplication.
[0113] See Figure 6 As shown, this embodiment also discloses a data processing apparatus based on feature labeling, including:
[0114] The feature tag attribute list creation module 601 is used to create a feature tag attribute list, including feature tag fields, from a preset description dimension based on the acquired business information.
[0115] The feature tagging module 602 is used to tag features based on the information in the resource table and the feature tagging attribute list, and to associate resource IDs with feature tags.
[0116] Information group marking module 603 is used to mark different information groups in the resource table in sequence;
[0117] The subject tagging module 604 is used to identify and tag subjects in an information group based on the tagged information group.
[0118] Priority marking module 605 is used to mark the priority of the same subject and the same element;
[0119] The data governance module 606 is used to associate feature tags with at least one dimension attribute in the feature tag attribute list through feature tags in the resource table, and to recommend configuration information for data governance strategies.
[0120] A data processing device based on feature labeling corresponds to a data processing method based on feature labeling. The specific implementation of each module can be found in the implementation of the same data processing method based on feature labeling, and will not be described again in this embodiment.
[0121] This embodiment presents a data processing method based on feature tags. By associating feature tags with at least one dimension attribute in the feature tag attribute list, it recommends configuration information for data governance strategies. This method can realize cross-industry data governance standards, reduce the complexity of data governance, and accelerate the governance process.
[0122] Example 3
[0123] See Figure 7 As shown in this embodiment, a data processing method based on feature tags includes:
[0124] S701, Based on the acquired business information, create a list of feature tag attributes, including feature tag fields, from a preset description dimension;
[0125] S702, Based on the information in the resource table and the list of feature tag attributes, feature tags are performed, and resource IDs are associated with feature tags;
[0126] S703, sequentially mark different information groups in the resource table;
[0127] S704, Based on the labeled information group, identify the main body in the information group and label it;
[0128] S705, Priority for marking the same subject and the same element;
[0129] S706, based on at least one of the element tag association information group, subject, and priority, recommends configuration information for data analysis; the data analysis includes extraction and distribution, data quality detection, theme planning, or information presentation.
[0130] For a detailed implementation of S701 to S505, please refer to Example 1.
[0131] Based on at least one of the following: feature tag association information group, subject, and priority, the system recommends configuration information for data analysis, specifically:
[0132] Data governance strategies can be configured based on at least one of the following descriptions: information group, subject, and priority. Furthermore, data analysis strategies can be configured by adding additional parameters such as feature tags, Chinese names, object types, value ranges, dictionary codes, data types, data element encodings, and processing methods on top of information groups, subjects, and priorities.
[0133] In this embodiment, the data governance strategy includes junk filtering, format cleaning, or deduplication.
[0134] See Figure 8 As shown, this embodiment also discloses a data processing apparatus based on feature labeling, including:
[0135] The feature tag attribute list creation module 801 is used to create a feature tag attribute list, including feature tag fields, from a preset description dimension based on the acquired business information.
[0136] The feature tagging module 802 is used to tag features based on the information in the resource table and the feature tagging attribute list, and to associate resource IDs with feature tags.
[0137] Information group marking module 803 is used to mark different information groups in the resource table in sequence;
[0138] The subject tagging module 804 is used to identify and tag subjects in an information group based on the tagged information group.
[0139] Priority marking module 805 is used to mark the priority of the same subject and the same element marking;
[0140] The data governance module 806 is used to associate information groups, subjects, and priorities based on element tags, and recommend configuration information for data analysis; the data analysis includes extraction and distribution, data quality detection, theme planning, or information presentation.
[0141] This embodiment presents a data processing method based on feature tags. It associates feature tags with at least one of information groups, subjects, and priorities, and recommends configuration information for data analysis. This facilitates highlighting key resource characteristics, helps to depict the objective connotation and business meaning of data, and provides an automated and intelligent foundation for subsequent extraction and distribution, data quality detection, theme planning, and information display.
[0142] Example 4
[0143] This embodiment uses a bank user's account opening information as an example to explain element tagging and data governance. Due to inconsistencies in the individuals filling out the forms, there may be issues such as duplicate submissions and inconsistent descriptions. Based on the element tagging description method, data governance is performed on this resource. First, a list of element tagging attributes is created, as shown in Table 2.
[0144] Table 2. Example list of user account opening information element markings for a certain bank.
[0145]
[0146] Note: Data element coding refers to "JR / T 0015-2004 General Data Elements for Bank Information Technology"
[0147] Based on the example list of feature labels in Table 2, and combining the information group, subject, and priority of the feature labels, preliminary data governance was achieved. The specific steps are as follows:
[0148] S1: Based on the existing list of element tags and the characteristics of the resource table, perform preliminary element tagging. For example, the element tag for "Personnel ID" is "ys-rybh", and the element tag for "Name" is "ys-xm", etc.
[0149] S2: Analyze the information groups in the resource table, using "1, 2, ..." to sequentially label different information groups, with the same number representing the same information group. For example, "Personnel ID, Name, ID Card Number, Document Type 1, Document Number 1, Main Account Type, Main Account, etc." are labeled as the same information group "1";
[0150] S3: Based on the information groups marked in S2, identify the subject within each information group. For example, if the subject of information group 1 is the account holder, then the document number "ID card number" is identified as the subject and marked with "1".
[0151] S4: Use “1, 2, …” to mark the priority of the same subject and the same element; for example, the contact phone number, home phone number, office phone number, etc. of information group 1 “Account Holder” are marked with the priority as contact phone number “1”, office phone number “2”, and home phone number “3”.
[0152] Table 3 shows an example of the annotation of a bank user's account opening information.
[0153] Table 3 Examples of Marking Elements for a Bank's User Account Opening Information
[0154]
[0155]
[0156] Data analysis and data governance can be performed based on Table 3.
[0157] See Figure 9 As shown, after tagging the "Account Opening Information of a Certain Bank Customer," the system associates these tags with attributes such as "Processing Method" and "Data Type" to recommend configuration information for strategies such as spam filtering, format cleaning, and deduplication. After confirmation / modification by the governance personnel, the data governance is completed.
[0158] Feature tagging is a type of tagging that is independent of database tables and focuses solely on the object itself. Based on feature tags, combined with logical operators, global retrieval can be performed, expressing the objective facts of behavior. This makes it easy for technical personnel to understand and also convenient for computers to recognize.
[0159] Example 1: Search condition configuration 1
[0160] The account was opened at Branch A, and the account balance is greater than 1000: (ys-zhmc-kh==”Branch A”)&&(ys-je>=1000)
[0161] Example 2: Search condition configuration 2
[0162] After tagging features in both the source and target tables, a matching strategy is set to perform conditional matching (changing the mapping mode).
[0163] The source table must include "Name", and the "ID Number" must match the "ID Type" and "ID Number": is-xm&&(is-sfzh||(is-zjlx&&is-zjhm)).
[0164] In addition to the data governance described above, other embodiments use information groups, subjects, priorities, and other related information based on element tags to highlight key resource characteristics, help characterize the objective connotation and business meaning of data, and provide an automated and intelligent foundation for subsequent data extraction and distribution, data quality testing, theme planning, and information display.
[0165] The above description is merely a preferred embodiment of the present invention; however, the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and its improved concepts, should be covered within the scope of protection of the present invention.
Claims
1. A data tagging method based on data elements, characterized in that, include: Based on the acquired business information, create a list of feature tag attributes, including feature tag fields, from the preset description dimensions; Based on the information in the resource table and the list of feature tag attributes, feature tags are created, and resource IDs are associated with feature tags. Sequentially label the different information groups in the resource table; Based on labeled information groups, identify and label the main entities within the information groups; Prioritize marking items with the same subject and the same elements; The preset description dimension includes the following fields: feature tag, Chinese name, object type, value range, dictionary code, data type, data element encoding, and processing method; the feature tag is a unique identifier for a data feature. The Chinese name refers to one or more Chinese words assigned to the element label; the object type is used to label the type of the subject object being described; the value range is the set of allowed values for meta-class elements determined according to the data type and representation format specified in the corresponding attributes; the dictionary code is the number of the dictionary code set followed; the data type is the data type that identifies the element; the data element code is the data element code corresponding to the industry; and the processing method is the data processing rule corresponding to the element label. The information group represents related tags for the same entity object; Groups of identical information are marked with the same Arabic numerals, and the order of the markings is determined by the importance of the subject.
2. The data tagging method based on data elements according to claim 1, characterized in that, The subject is a field used to identify the main object describing the resource; the subject may not be unique.
3. A data tagging device based on data elements, characterized in that, The apparatus, based on the method according to any one of claims 1 to 2, comprises: The Element Tag Attribute List Creation Module is used to create a list of element tag attributes, including element tag fields, from a preset description dimension based on the acquired business information. The feature tagging module is used to tag features based on the information in the resource table and the feature tag attribute list, and to associate resource IDs with feature tags. The information group tagging module is used to sequentially tag different information groups in the resource table; The subject tagging module is used to identify and tag subjects within a group of information based on the tags. The priority tagging module is used to tag the priority of tags for the same subject and the same element.
4. A data processing method based on feature labeling, characterized in that, include: Based on the acquired business information, create a list of feature tag attributes, including feature tag fields, from the preset description dimensions; Based on the information in the resource table and the list of feature tag attributes, feature tags are created, and resource IDs are associated with feature tags. Sequentially label the different information groups in the resource table; Based on labeled information groups, identify and label the main entities within the information groups; Prioritize marking items with the same subject and the same elements; Associate the feature tags in the resource table with at least one preset description dimension in the feature tag attribute list, and recommend configuration information for data governance strategies. The preset description dimension includes the following fields: feature tag, Chinese name, object type, value range, dictionary code, data type, data element encoding, and processing method; the feature tag is a unique identifier for a data feature. The Chinese name refers to one or more Chinese words assigned to the element label; the object type is used to label the type of the subject object being described; the value range is the set of allowed values for meta-class elements determined according to the data type and representation format specified in the corresponding attributes; the dictionary code is the number of the dictionary code set followed; the data type is the data type that identifies the element; the data element code is the data element code corresponding to the industry; and the processing method is the data processing rule corresponding to the element label. The information group represents related tags for the same entity object; Groups of identical information are marked with the same Arabic numerals, and the order of the markings is determined by the importance of the subject.
5. The data processing method based on feature labeling according to claim 4, characterized in that, The data governance strategies include spam filtering, format cleaning, or deduplication.
6. A data processing device based on feature labeling, characterized in that, The apparatus, based on the method according to any one of claims 4 to 5, comprises: The Element Tag Attribute List Creation Module is used to create a list of element tag attributes, including element tag fields, from a preset description dimension based on the acquired business information. The feature tagging module is used to tag features based on the information in the resource table and the feature tag attribute list, and to associate resource IDs with feature tags. The information group tagging module is used to sequentially tag different information groups in the resource table; The subject tagging module is used to identify and tag subjects within a group of information based on the tags. The priority marking module is used to mark the priority of tags for the same subject and the same feature; The data governance module is used to associate feature tags with at least one dimension attribute in the feature tag attribute list through feature tags in the resource table, and recommend configuration information for data governance strategies.
7. A data processing method based on feature labeling, characterized in that, include: Based on the acquired business information, create a list of feature tag attributes, including feature tag fields, from the preset description dimensions; Based on the information in the resource table and the list of feature tag attributes, feature tags are created, and resource IDs are associated with feature tags. Sequentially label the different information groups in the resource table; Based on labeled information groups, identify and label the main entities within the information groups; Prioritize marking items with the same subject and the same elements; Based on at least one of the following: feature tagging association information group, subject, and priority, configuration information for data analysis is recommended; the data analysis includes extraction and distribution, data quality detection, theme planning, or information presentation; The preset description dimension includes the following fields: feature tag, Chinese name, object type, value range, dictionary code, data type, data element encoding, and processing method; the feature tag is a unique identifier for a data feature. The Chinese name refers to one or more Chinese words assigned to the element label; the object type is used to label the type of the subject object being described; the value range is the set of allowed values for meta-class elements determined according to the data type and representation format specified in the corresponding attributes; the dictionary code is the number of the dictionary code set followed; the data type is the data type that identifies the element; the data element code is the data element code corresponding to the industry; and the processing method is the data processing rule corresponding to the element label. The information group represents related tags for the same entity object; Groups of identical information are marked with the same Arabic numerals, and the order of the markings is determined by the importance of the subject.
8. A data processing device based on feature labeling, characterized in that, Based on the method of claim 7, the apparatus comprises: The Element Tag Attribute List Creation Module is used to create a list of element tag attributes, including element tag fields, from a preset description dimension based on the acquired business information. The feature tagging module is used to tag features based on the information in the resource table and the feature tag attribute list, and to associate resource IDs with feature tags. The information group tagging module is used to sequentially tag different information groups in the resource table; The subject tagging module is used to identify and tag subjects within a group of information based on the tags. The priority marking module is used to mark the priority of tags for the same subject and the same feature; The data analysis module is used to associate information groups, subjects, and priorities based on feature tags, and to recommend configuration information for data analysis; the data analysis includes extraction and distribution, data quality detection, theme planning, or information presentation.
Citation Information
Patent Citations
Massive electronic data management method
CN107870907A
Data element analysis method and device, electronic device and storage medium
CN112464640A