Data organization and distribution method and device based on element marking and readable medium
By establishing a list of data elements and configuring distribution strategies, the mapping relationship between the data source and the target table is realized, and the problems of high labor costs and low efficiency in the data distribution process in the existing technology are solved, and the automation and intelligence of data distribution are realized.
Patent Information
- Application Number
- CN202510085325.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art requires high business understanding capabilities and a large amount of labor costs in the data distribution process, and the data distribution effect is poor, low efficiency, error-prone, and difficult to achieve automation and intelligence.
By establishing a list of data elements, configuring data feature tags and distribution policy information, establishing a mapping relationship between the source table and the target table, and realizing automatic data distribution.
It improves the efficiency and accuracy of data distribution, reduces labor costs, realizes semi-automation of the data distribution process, and is suitable for data fusion between complex application scenarios and multiple systems.
Smart Images

Figure CN120067104A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology, and particularly relates to a data organization and distribution method, device and readable medium based on element marking. Background Art
[0002] On the one hand, with the in-depth development of the new round of scientific and technological revolution and industrial transformation, the value of data as a key production factor has become increasingly prominent. In recent years, the digital economy has developed rapidly, the scale and energy level of digital infrastructure have been greatly improved, and digital technology and industrial systems have become increasingly mature, laying a solid foundation for better playing the role of data elements. At the same time, there are also problems such as low data supply quality, unsmooth circulation mechanism, and insufficient release of application potential. Implementing the "Data Element ×" action is to give play to the multiple advantages of the existing ultra-large-scale market, massive data resources, rich application scenarios, etc., promote the coordination of data elements with factors such as labor and capital, lead the technology flow, capital flow, talent flow, and material flow with the data flow, break through the constraints of traditional resource elements, and improve the total factor productivity; promote the multi-scenario application and multi-subject reuse of data, cultivate new products and new services based on data elements, realize knowledge diffusion and value multiplication, and open up new space for economic growth; accelerate the integration of diverse data, and promote the innovation and upgrading of production tools with the expansion of data scale and the enrichment of data types, give birth to new industries and new models, and cultivate new drivers of economic development.
[0003] On the other hand, with the development of big data and cloud computing technologies and the accelerating data generation speed, the demand for data has increased sharply, and the amount of data that big data systems need to process has become larger and larger, which puts higher requirements on data processing, distribution, and data organization capabilities. At the same time, it is difficult for data processing personnel to quickly process "multi-source heterogeneous" big data and distribute and store it according to data value density, application usage requirements, etc. Therefore, usually according to the actual business needs, a data distribution strategy is configured, and the required data is extracted from the source table to the target table specifically, so as to achieve the purpose of improving the data value density. Therefore, how to make the data distribution configuration strategy more convenient and achieve automation and intelligence is inevitably the development direction of big data systems.
[0004] In actual engineering practice, to realize the data distribution of Table A to Table B, the engineering implementation steps are as follows:
[0005] a) Determine whether it is necessary to distribute the data of Table A to Table B, and define the data distribution strategy from Table A to Table B;
[0006] b) Establish a mapping from Table A to Table B;
[0007] c) Establish a mapping from the data items of Table A to the data items of Table B;
[0008] d) Execute the data distribution strategy task.
[0009] For the above data distribution implementation plan, there are several disadvantages: First, it requires a relatively high ability to understand business, which means understanding both the source table and the target table. Second, the labor cost is high. Data processing personnel need to participate in the entire data distribution process, analyze the source table and the target table, and manually configure information such as data distribution table mapping, data item mapping, and distribution frequency. This process consumes a large amount of manpower, time, and cost. Third, the data distribution effect is poor. Data processing personnel need to configure mappings one by one, which is inefficient, error-prone, incomplete in mapping, and the data analysis results cannot be reused.
[0010] Therefore, a patent for invention with the publication number CN116991842A applied by the present applicant discloses a data distribution method, device, and readable medium based on data element tags. It establishes a data element list; creates a target table in the target library, configures data element tags with corresponding meanings for the target fields in the target table according to the data element list, and configures the distribution strategy information of the target table. The distribution strategy information includes distribution information and distribution rules, and the distribution rules include data element tags and the logical relationships established based on the data element tags. When the data in the source library is accessed, data element tags with corresponding meanings are configured for the source fields in the source table in the source library according to the data element list; a first mapping relationship is established between the source table and the target table and a second mapping relationship is established between the source field and the target field according to the distribution strategy information, and the data in the source table is distributed to the target table according to the first mapping relationship and the second mapping relationship, which can effectively reduce labor costs and improve distribution efficiency.
[0011] This solution abstracts the previous specific table-to-table and field-to-field mappings into concept matching, realizing semi-automation of the data distribution process, reducing the workload of data distribution and improving the accuracy of distribution. However, this solution realizes semi-automated data distribution for data within the same system. But when it comes to system expansion or data transmission between multiple systems, this solution has great limitations. At the same time, when this solution performs data distribution, subsequent technical personnel need to understand the rule settings and element tag settings of the previous technical personnel before they can use it, which has the defect of inconvenient data analysis and processing. Especially when it comes to data transmission in complex application scenarios, the application of this solution is limited. Summary of the Invention
[0012] A brief overview of the embodiments of the present invention is given below to provide a basic understanding of certain aspects of the present invention. It should be understood that the following overview is not an exhaustive overview of the present invention. It is not intended to identify the key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is merely to present certain concepts in a simplified form as a prelude to the more detailed description to be discussed later.
[0013] To further improve the practicality of the data distribution method based on element tagging and make it widely applicable to various complex application scenarios or data fusion scenarios between different systems, this application proposes a data distribution strategy method based on element tagging. Through unified element tagging and combined with rule strategies, users are recommended to perform tagging to generate a data organization and distribution strategy, realizing the connection configuration between different tables. Based on the unified and standardized data organization and distribution strategy method proposed by this method, it will match the characteristics of the data items in the source table and the target table with the recommended rules in the data organization and distribution strategy knowledge base, and recommend data distribution configurations according to the matching results. When the distribution strategy needs to be adjusted, only the element tagging and adaptation strategy of the corresponding source table need to be adjusted to achieve the distribution mapping configuration with the corresponding target table, thus realizing fast and unified data distribution strategy configuration, improving work efficiency and reducing labor costs. This also helps to cope with changes in data organization and distribution tasks, such as the need to distribute the data of a source table to more than two target tables. This method can effectively improve the scalability and flexibility of the system.
[0014] In a big data system, to meet the requirements of upper-layer applications, data organization design is carried out according to actual software and hardware conditions, data resource characteristics, data processing process requirements, etc., thus forming multiple databases with actual business meanings to store the processed data. Data distribution is an important link in the processing process, which extracts and distributes the source data to the corresponding databases according to the strategy.
[0015] According to the first aspect of this application, a data organization and distribution method based on element tagging is provided, including:
[0016] S1. Establish a data element list, which includes data element tags, element tag parameters, technical element tags, and connection relation expressions; the data element tags are used to describe a certain concept in the real world corresponding to this data item; the element tag parameters are the parameters of the data element tags; the technical element tags are used for technical auxiliary description of this data item; the connection relation expressions express the data element tags as preset rules through connection symbols;
[0017] S2. Establish a target table in the target library, configure data element tags with corresponding meanings for the target fields in the target table according to the data element list, and configure the distribution strategy information of the target table. The distribution strategy information includes distribution information and distribution rules; the distribution rule is: a logical relationship formed based on data element tags, element tag parameters, technical element tags, and connection relation expressions and in accordance with a certain syntax structure;
[0018] Specifically, the distribution rules are based on the data structure and business scenarios of the target table, and generate corresponding policies according to certain grammatical structures based on the information such as the published element tags, element tag parameters, connection symbols, etc. For example, in the personnel theme information, personnel basic information policies, marriage information policies, etc. can be generated. When defining the policies, try to use the core fields such as the key business fields and required fields in the target table to form the recommendation rules. For example, in the personnel basic information policy, the name, ID number, and ID type are used as the matching conditions.
[0019] S3. When the data in the source library is accessed, configure the data element tags with corresponding meanings for the source fields in the source tables in the source library according to the data element list;
[0020] S4. Establish the first mapping relationship between the source table and the target table and the second mapping relationship between the source field and the target field according to the distribution policy information, and distribute the data in the source table to the target table according to the first mapping relationship and the second mapping relationship.
[0021] As a specific implementation solution, the technical element tags are used to assist in the technical description of data items, mainly for expressing the complete technical characteristics of the data items; the technical element tags are mainly used for data fusion and retrieval between different systems. The technical element tags include data classification, data description, data relationship, and multiple description tags associated with the data element tags; the data classification and data description are automatically generated based on the data element list; the multiple description tags associated with the data element tags are obtained by automatically identifying the data element tags based on the data element list and using artificial intelligence (AI) and machine learning technologies. This technology is an existing technology and will not be elaborated here; the data relationship is denoted as Sn, which consists of a central node and several peripheral nodes, and each peripheral node is directly connected to the central node, and the associated peripheral nodes are connected to each other.
[0022] The data relationship Sn has n + 1 nodes, where 1 is the central node and n are the peripheral nodes; two nodes with an association relationship are connected as an edge. The central node is the data description, and the peripheral nodes include the data source set A (the data source can be a table, a database, or an external link, and there can be multiple), and all the data element tag sets B related to the data description; when the data sources in the data source set A are related to each other, the two are connected to form an edge; when the data element tags in the data element tag set B are related to each other, the two are connected to form an edge; when the source of the data element tag in the data element tag set B comes from a certain data source in the data source set A, the two are connected to form an edge.
[0023] Further, to achieve credibility assessment, the data organization and distribution method further includes a process of automatically calculating the credibility of the data source, specifically including:
[0024] For each peripheral node, if more than two data element tags belong to the same data source, a first connection edge is formed between the data element tag and the data source;
[0025] If there is an association relationship between data element tags belonging to the same data source, a second connection edge is formed between the data element tags;
[0026] If there is an association relationship between data element tags that do not belong to the same data source, a third connection edge is formed between the data element tags;
[0027] In practical applications, the more times a node appears repeatedly, the higher its credibility. Therefore, a simple discrimination method is as follows: judge its credibility according to the appearance frequency of each peripheral node in the first connection edge, the second connection edge, and the third connection edge.
[0028] As a further solution, for some special occasions where it is necessary to accurately judge the data source, the discrimination method is as follows: assign an initial credibility to each peripheral node in advance, and comprehensively judge its credibility in combination with the appearance frequency of each peripheral node in the first connection edge, the second connection edge, and the third connection edge (for example, the summation method can be used).
[0029] As an even further solution, the discrimination method is as follows: assign an initial credibility to each peripheral node in advance, and simultaneously set weight values for the first connection edge, the second connection edge, and the third connection edge respectively, and comprehensively calculate the final credibility according to the initial credibility and the weight values.
[0030] Compared with the prior art, the present application creatively evaluates the confidence of data by setting a star-shaped data relationship and according to the appearance frequency of peripheral nodes, which is simple and convenient. At the same time, the data relationship can be conveniently visualized.
[0031] In addition, through the setting of the data relationship, the system can be conveniently extended or data can be transmitted between multiple systems; technicians only need to read the data relationship (visualization or relationship diagram), then they can clearly manage and analyze the data; this solution is especially suitable for data fusion between different systems and has the advantages of being easy to analyze, manage, and expand.
[0032] Among them, data classification is the basis of the data element identification system and is usually carried out according to the following dimensions:
[0033] By data type: structured data (such as databases), unstructured data (such as documents, pictures), semi-structured data (such as log files).
[0034] By business value: core data (directly affecting business decisions), auxiliary data (supporting business operations), and marginal data (potential value to be mined).
[0035] By lifecycle: active data (frequently used), archived data (infrequently used), and discarded data (of no use value).
[0036] Data description annotates the detailed characteristics of data, including but not limited to the following:
[0037] Data attribute code: representing the specific characteristics of data, such as the source, purpose, format, etc. of the data.
[0038] Data quality code: representing the quality of data, such as the accuracy, integrity, consistency, etc. of the data.
[0039] Data label: used to clearly define information such as the characteristics, categories, and attributes of data.
[0040] Specifically, data element marking: It is a marking of a data item, used to explain that this field corresponds to a certain concept in the real world. Denote the element expressing its own meaning as "ys". For example, if a data item is name, then this field can be marked as ys.name, corresponding to the name.
[0041] Specifically, element marking parameters: For some ys element markings with parameters, when forming recommendation rules, it may be necessary to further specify the values of their parameters in order to be used as the basis for matching hits.
[0042] Specifically, connection relation expression: Connect multiple data element markings through special symbols to form a preset rule. For example, to express "both condition A and condition B must be satisfied" as "A&B", and "either condition A or condition B is satisfied" as "A|B", etc.
[0043] According to the second aspect of the present application, there is provided a data organization and distribution device based on element marking, including:
[0044] A data element list establishment module, configured to establish a data element list; the data element list includes data element markings, element marking parameters, technical element markings, and connection relation expressions; the data element markings are used to explain that this data item corresponds to a certain concept in the real world; the element marking parameters are the parameters of the data element markings; the connection relation expression expresses the data element markings as a preset rule through a connection symbol;
[0045] The technical element tags are used for technically assisting in the description of the data item; the technical element tags include data classification, data description, data relationships, and multiple description tags associated with the data element tags; the data classification and data description are automatically generated based on the data element list; the multiple description tags associated with the data element tags are obtained by automatically identifying the data element tags based on the data element list and using artificial intelligence and machine learning technologies; the data relationship is denoted as Sn and consists of a central node and several peripheral nodes, each peripheral node is directly connected to the central node, and the associated peripheral nodes are connected to each other;
[0046] The distribution policy configuration module is configured to create a target table in the target library, configure data element tags with corresponding meanings for the target fields in the target table according to the data element list, and configure the distribution policy information of the target table, where the distribution policy information includes distribution information and distribution rules; the distribution rule is: a logical relationship formed based on the data element tags, element tag parameters, technical element tags, and connection relationship expressions and in accordance with a certain syntax structure;
[0047] The source table configuration module is configured to, when the data in the source library is accessed, configure data element tags with corresponding meanings for the source fields in the source table in the source library according to the data element list;
[0048] The data distribution module is configured to establish a first mapping relationship between the source table and the target table and a second mapping relationship between the source field and the target field according to the distribution policy information, and distribute the data in the source table to the target table according to the first mapping relationship and the second mapping relationship.
[0049] According to a third aspect of the present application, there is provided an electronic device, which includes:
[0050] One or more processors;
[0051] A storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the above data organization and distribution method.
[0052] According to a fourth aspect of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above data organization and distribution method is implemented.
[0053] The solution of the present application improves the prior application mentioned in the background art. Compared with the prior art, it has the following advantages:
[0054] Based on the setting of the technical element tags, by setting multiple description tags for AI recognition, it is convenient to retrieve similar data with different names during data fusion;
[0055] A sub - feature of data relationship is specially designed and is designed as a star structure composed of a central node and several peripheral nodes. Under this structure, weights can be conveniently set for different nodes (the default weight is 1). Its data source can be used for credibility evaluation, and all relevant data element tags can be used for data management and analysis. It can also be extended to realize the data flow between data items corresponding to the data element tags.
[0056] In summary, by adopting the solution of the present invention, multiple advantages such as existing ultra - large - scale markets, massive data resources, and rich application scenarios can be utilized to promote the coordination of data elements with elements such as labor and capital, lead the technology flow, capital flow, talent flow, and material flow with the data flow, break through the constraints of traditional resource elements, and improve the total factor productivity; promote the multi - scenario application and multi - subject reuse of data, cultivate new products and new services based on data elements, realize knowledge diffusion and value multiplication, and open up new economic growth space; accelerate the integration of diverse data, promote the innovation and upgrading of production tools with the expansion of data scale and the enrichment of data types, give birth to new industries and new models, and cultivate new driving forces for economic development. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The present invention can be better understood by referring to the descriptions given below in conjunction with the accompanying drawings, in which the same or similar reference numerals are used throughout all the drawings to denote the same or similar components. The accompanying drawings, together with the following detailed description, are included in this specification and form a part of this specification, and are used to further illustrate the preferred embodiments of the present invention and to explain the principles and advantages of the present invention. In the attached
[0058] In the figures:
[0059] Figure 1 is a schematic flow chart of the data organization and distribution method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] The embodiments of the present invention will be described below with reference to the accompanying drawings. Elements and features described in one drawing or one embodiment of the present invention can be combined with elements and features shown in one or more other drawings or embodiments. It should be noted that, for the sake of clarity, the drawings and the description omit the representation and description of components and processes that are irrelevant to the present invention and are known to those of ordinary skill in the art.
[0061] In a big data system, to meet the requirements of upper-layer applications, data organization design is carried out according to actual software and hardware conditions, data resource characteristics, data processing process requirements, etc., so as to form multiple databases with actual business meanings for storing the processed data. Data distribution is an important link in the processing process, which extracts and distributes the source data to the corresponding databases according to the strategy. The present invention proposes to recommend distribution configurations based on various element tags and in combination with predefined data distribution strategies to drive efficient data governance. See Figure 1 , which specifically includes the following three stages:
[0062] Stage 1: Define element tags
[0063] In this stage, element tags required in data processing, parameters of element tags, technical element tags, connection relation expressions, etc. need to be defined.
[0064] 1. Element tags: They are a kind of tags for data items, used to explain that a certain field corresponds to a certain concept in the real world. Denote the element that expresses its own meaning as "ys". For example, if a certain data item is name, then this field can be tagged as ys.name, corresponding to the name.
[0065] 2. Parameters of element tags: For some ys element tags with parameters, when forming recommendation rules, it may be necessary to further specify the values of their parameters in order to be used as the basis for matching hits.
[0066] 3. Technical element tags (tech element tags): Used for technical auxiliary description of data items, mainly used to express the complete technical characteristics of the data item; this technical element tag is mainly used for data fusion and retrieval between different systems.
[0067] Technical element tags include data classification, data description, data relationship, and multiple description tags associated with data element tags; data classification and data description are automatically generated based on the data element list; multiple description tags associated with data element tags are obtained by automatically identifying data element tags based on the data element list and using artificial intelligence (AI) and machine learning technologies; the data relationship is denoted as Sn, which consists of a central node and several peripheral nodes, and each peripheral node is directly connected to the central node, and the associated peripheral nodes are connected to each other.
[0068] The data relationship Sn has n + 1 nodes, where 1 is the central node and n are the peripheral nodes; two nodes with an association relationship are connected by an edge. The central node is for data description, and the peripheral nodes include the data source set A (the data sources can be tables, databases, or external links, and there can be multiple), and all data element tag sets B related to this data description. When the data sources in the data source set A are related to each other, the two are connected to form an edge; when the data element tags in the data element tag set B are related to each other, the two are connected to form an edge; when the source of a data element tag in the data element tag set B comes from a certain data source in the data source set A, the two are connected to form an edge.
[0069] 4. Connection relationship expression: Connect multiple element tags through special symbols. For example, to express "both condition A and condition B must be satisfied" as "A&B", and "either condition A or condition B is satisfied" as "A|B", etc.
[0070] Phase II: Define strategies
[0071] In this phase, based on the data structure and business scenario of the target table, and based on the information such as the published element tags, element tag parameters, and connection symbols, generate strategies according to a certain grammar structure. For example, in the personnel theme information, strategies for personnel basic information, marriage information, etc. can be generated. When defining strategies, try to use core fields such as key business fields and required fields in the target table to form recommendation rules. For example, in the personnel basic information strategy, use name, ID number, and ID type as matching conditions.
[0072] Phase III: Complete data distribution based on element tags and strategies
[0073] In this phase, analyze the source table, mark the data items with element tags according to the requirements of Phase I, and then match the strategies in Phase II with the element tags of the source table. Based on the matching hit results, the system automatically recommends the data distribution configuration results, including the data organization and distribution relationship from the source table to the target table, and the mapping relationship between the data items in the source table and the target standard data, etc.
[0074] To facilitate the application of the data distribution technology based on element marking to complex application scenarios, this application adds a feature of technical element marking. Different from other existing technologies: 1. By setting multiple description tags for AI recognition, this technical element marking can facilitate the retrieval of similar data with different names during data fusion. 2. A sub-feature of data relationship is specifically designed and is designed as a star structure composed of a central node and several peripheral nodes. Under this structure, weights can be conveniently set for different nodes (the default weight is 1). Its data source can be used for credibility evaluation, and all related data element markings can be used for data management and analysis. It can also be extended to realize the data flow between data items corresponding to the data element markings.
[0075] In addition, during subsequent use, to achieve credibility evaluation, this data organization and distribution method further includes a process of automatically calculating the credibility of the data source, specifically including:
[0076] For each peripheral node, if more than 2 data element markings belong to the same data source, a first connection edge is formed between the data element marking and the data source;
[0077] If there is an association relationship between data element markings belonging to the same data source, a second connection edge is formed between the data element markings;
[0078] If there is an association relationship between data element markings that do not belong to the same data source, a third connection edge is formed between the data element markings;
[0079] In practical applications, the more times a node appears repeatedly, the higher its credibility. Therefore, a simple discrimination method is as follows: Judge its credibility according to the occurrence frequency of each peripheral node in the first connection edge, the second connection edge, and the third connection edge.
[0080] For some special occasions where it is necessary to accurately judge its data source, the discrimination method is designed as: Assign an initial credibility to each peripheral node in advance, and comprehensively judge its credibility in combination with the occurrence frequency of each peripheral node in the first connection edge, the second connection edge, and the third connection edge (for example, the summation method can be used).
[0081] In some more precise occasions, the discrimination method can also be designed as: Assign an initial credibility to each peripheral node in advance, and simultaneously set weight values for the first connection edge, the second connection edge, and the third connection edge respectively, and comprehensively calculate the final credibility according to the initial credibility and the weight values.
[0082] Through the above solution, the present invention establishes a diversified data distribution strategy, enabling data to be interconnected and interoperable among different systems and platforms. This not only helps reduce the overall social cost of data circulation but also improves the availability and security of data. At the same time, most importantly, through the design of technical element tags, the present invention can achieve the discrimination of data credibility, quantify the true value of data elements, and eliminate the uncertainty and heterogeneity of their values.
[0083] Taking the personnel portrait in e-commerce as an example, through research, in this business, the personnel portrait information includes: basic personnel information (name, document type, document number, contact information, gender, date of birth, etc.), marital information (name of spouse, document type, document number, etc.), and educational experience information (name of graduated school, graduation time, major studied, etc.). The source data includes membership registration information, marriage registration information, and educational information, etc. It is necessary to implement a recommendation strategy based on element tags to distribute the source table data to the target table. The implementation steps of this data distribution are as follows:
[0084] According to the requirements of Phase 1, define the element tag information and tech element tag information, as shown in Table 1.
[0085] Table 1 Element Tag List
[0086]
[0087]
[0088] tech.classInfo is used to mark the data items that describe the same entity object in the data resource, group these data items, and form a relationship graph between data element tags. Its expression is tech.classInfo = num, where num is numeric, and the same num value represents the same information group.
[0089] The parent element tag plays a very good role in data distribution and can enhance the matching effect of the data distribution strategy. That is, when the element tag identifier cannot be matched, its parent element tag identifier can be matched.
[0090] Design the distribution strategy according to the requirements of Phase 2. First, analyze the data structure and business scenario of the target table, and configure the element tags for the data items in the target table based on the content of the previous step. The following is the element tag information for basic personnel information.
[0091] Table 2 Basic Personnel Information
[0092] Data Item Identifier Chinese Name of Data Item Element Mark name Name ys.xm certifType Certificate Type ys.zjlx certifNumber Certificate Number ys.zjhm gender Gender Code ys.xbdm nativePlace Native Place Code ys.jgdm mobilePhone Mobile Phone ys.sj nation Nation Code ys.mzdm
[0093] Then, configure the distribution strategy for the target table with element tags already marked. The data distribution strategy consists of multiple element tags, which are connected by connection symbols such as and(&) and or(|). And means "AND", that is, all element tags must be satisfied simultaneously to pass. Or means "OR", that is, only one of the element tags needs to be included to pass. The combination of and and or represents the combination relationship of element tags. Only when this strategy is satisfied can the data be distributed. For example, the data items (such as name, document type, document number, gender, native place code, mobile phone, ethnic group code) in the membership registration information hit the recommendation rule "ys.xm&ys.zjlx&ys.zjhm&ys.xbdm&ys.jgdm&ys.sj&ys.mzdm" of the personnel basic information strategy.
[0094] According to the requirements of Phase 3, mark the data items of the source table with element tags. The content of the source data structure is shown in Tables 4 and 5.
[0095] Table 3 Membership Registration Information
[0096] Data Item Identifier Chinese Name of Data Item Element Mark name Name ys.xm certifType Certificate Type ys.zjlx certifNumber Certificate Number ys.zjhm gender Gender Code ys.xbdm nativePlace Native Place Code ys.jgdm mobilePhone Mobile Phone ys.sj nation Nation Code ys.mzdm title Position rank Rank
[0097] Table 4 Marriage Registration Information
[0098]
[0099]
[0100] When the data items of the source table are marked with element tags, the big data system automatically recommends the data organization and distribution strategy that matches them, and the data distribution task configuration can be constructed.
[0101] In actual big data governance, there may be more than two subject information for some data. To avoid problems such as incorrect configuration and missing data items during data extraction and distribution, grouping can be performed through information groups marked with tech elements, which effectively solves this problem. At the same time, grouping through tech element marking can also achieve "many-to-one" data distribution. Through the analysis of the marriage registration information in Table 4, it is found that there are two subject information, namely the male and female. Using tech element marking, they are respectively marked as information groups 1 and 2. By matching the data items in the information group with the recommendation rules of the personnel basic information strategy, the male subject information (such as male name, male document type, male document number, male mobile phone number, male native place, male ethnicity) can be obtained, which hits the recommendation rule "ys.xm&ys.zjlx&ys.zjhm&ys.xbdm&ys.jgdm&ys.sj&ys.mzdm", and the female subject information (such as female name, female document type, female document number, female mobile phone number, female native place, female ethnicity) hits the recommendation rule "ys.xm&ys.zjlx&ys.zjhm&ys.xbdm&ys.jgdm&ys.sj&ys.mzdm". Here, the parent element marking is also used for auxiliary judgment.
[0102] For a target table, since there can be multiple data organization and distribution strategies, multiple source tables can be matched, and even the same strategy can be matched by the same source table multiple times. The final result is multiple distribution configurations. The big data system makes a comprehensive judgment based on factors such as data freshness and data resource importance, and merges some distribution configurations into one distribution configuration.
[0103] In the present invention, through the abstract concept of data element marking, the previous specific mapping between tables and between data items is abstracted into the matching of concepts. The mapping relationship between the source table and the target table in data organization and distribution is associated with the recommendation rules, realizing the automation and intelligence of the data distribution process. This not only greatly reduces the workload of data distribution, improves the accuracy and flexibility of distribution, and reduces the operation difficulty, but also provides strong support for big data governance and analysis. At the same time, based on the setting of technical element marking, by setting multiple description tags for AI recognition, it is convenient to retrieve similar data with different names during data fusion; a sub-feature of data relationship is specially designed and is designed as a star structure composed of a central node and several peripheral nodes. Under this structure, it is convenient to set weights for different nodes (the default weight is 1). Its data source can be used for credibility evaluation, all relevant data element markings can be used for data management and analysis, and it can also be extended to realize the data flow between data items corresponding to data element markings.
[0104] It should be emphasized that the term "comprising / including", as used herein, refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0105] Although the present invention has been disclosed above by the description of specific embodiments of the present invention, it should be understood that all the above embodiments and examples are exemplary and not restrictive. Those skilled in the art can design various modifications, improvements or equivalents to the present invention within the spirit and scope of the appended claims. These modifications, improvements or equivalents should also be considered to be included within the scope of protection of the present invention.
Claims
1. A method for organizing and distributing data based on element marking, characterized in that: include: S1, establishing a data element list, the data element list includes data element tags, element tag parameters, technical element tags and connection relationship expressions; The data element tag is used to indicate that the data item corresponds to a concept in the real world; the element tag parameter is a parameter of the data element tag; the connection relationship expression expresses the data element tag as a preset rule through a connector; The technical element tag is used to provide technical auxiliary description of the data item; the technical element tag includes data classification, data description, data relationship and multiple description tags associated with the data element tag; the data classification and data description are automatically generated based on the data element list; the multiple description tags associated with the data element tag are obtained by automatically identifying the data element tag based on the data element list and using artificial intelligence and machine learning technology; the data relationship is denoted as Sn, which is composed of a central node and several peripheral nodes, each peripheral node is directly connected to the central node, and the associated peripheral nodes are connected to each other; S2, establish a target table in the target database, configure data element tags with corresponding meanings for target fields in the target table according to the data element list, and configure distribution strategy information of the target table, the distribution strategy information includes distribution information and distribution rules; the distribution rules are: based on data element tags, element tag parameters, technical element tags and connection relationship expressions, and logical relationships formed according to certain grammatical structures; S3, when the data in the source library is accessed, the source field in the source table in the source library is configured with a data element tag with a corresponding meaning according to the data element list; S4, establishing a first mapping relationship between the source table and the target table and a second mapping relationship between the source field and the target field according to the distribution strategy information, and distributing the data in the source table to the target table according to the first mapping relationship and the second mapping relationship.
2. The data organization and distribution method based on element marking according to claim 1 is characterized in that: The data relationship Sn has n+1 nodes, of which 1 is a central node and n are peripheral nodes; two nodes with an association relationship are connected to form an edge; the central node is a data description, and the peripheral nodes include a data source set A and a set B of all data element tags related to the data description; when the data sources in the data source set A are related to each other, the two are connected to form an edge; when the data element tags in the data element tag set B are related to each other, the two are connected to form an edge; when the source of the data element tags in the data element tag set B comes from a data source in the data source set A, the two are connected to form an edge.
3. The data organization and distribution method based on element marking according to claim 2 is characterized in that: In order to achieve credibility evaluation, the data organization and distribution method also has a process of automatically calculating the credibility of the data source, which specifically includes: For each peripheral node, if more than two data element labels belong to the same data source, a first connection edge is formed between the data element label and the data source; If the data element tags belonging to the same data source have an association relationship, a second connection edge is formed between the data element tags; If the data element tags that do not belong to the same data source have an association relationship, a third connection edge is formed between the data element tags; The credibility determination method is as follows: the credibility of each peripheral node is determined according to the frequency of occurrence of the peripheral node on the first connection edge, the second connection edge and the third connection edge.
4. The method for organizing and distributing data based on element marking according to claim 3, characterized in that: The credibility determination method includes: assigning an initial credibility to each peripheral node in advance, and comprehensively determining the credibility of each peripheral node based on the frequency of occurrence of the peripheral node in the first connection edge, the second connection edge, and the third connection edge.
5. The method for organizing and distributing data based on element marking according to claim 3, characterized in that: The credibility determination method includes: assigning an initial credibility to each peripheral node in advance, and setting weight values for the first connection edge, the second connection edge and the third connection edge respectively, and calculating the final credibility based on the initial credibility and the weight values.
6. The data organization and distribution method based on element marking according to claim 1 is characterized in that: The connection relationship expression is to connect multiple data element tags through special symbols to form a preset rule.
7. A data organization and distribution device based on element marking, characterized in that: include: A data element list building module, configured to build a data element list; The data element list includes data element tags, element tag parameters, technical element tags and connection relationship expressions; the data element tags are used to indicate that the data item corresponds to a concept in the real world; the element tag parameters are parameters of the data element tags; the connection relationship expressions express the data element tags as preset rules through connectors; The technical element tag is used to provide technical auxiliary description of the data item; the technical element tag includes data classification, data description, data relationship and multiple description tags associated with the data element tag; the data classification and data description are automatically generated based on the data element list; the multiple description tags associated with the data element tag are obtained by automatically identifying the data element tag based on the data element list and using artificial intelligence and machine learning technology; the data relationship is denoted as Sn, which is composed of a central node and several peripheral nodes, each peripheral node is directly connected to the central node, and the associated peripheral nodes are connected to each other; The distribution strategy configuration module is configured to establish a target table in the target library, configure data element tags with corresponding meanings for target fields in the target table according to the data element list, and configure distribution strategy information of the target table, wherein the distribution strategy information includes distribution information and distribution rules; the distribution rules are: logical relationships formed based on data element tags, element tag parameters, technical element tags and connection relationship expressions and in accordance with certain grammatical structures; The source table configuration module is configured to configure data element tags with corresponding meanings for source fields in the source table in the source library according to the data element list when data in the source library is accessed; The data distribution module is configured to establish a first mapping relationship between the source table and the target table and a second mapping relationship between the source field and the target field according to the distribution strategy information, and distribute the data in the source table to the target table according to the first mapping relationship and the second mapping relationship.
8. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enables the one or more processors to implement the data organization and distribution method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the data organization and distribution method as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Data element tag-based data distribution method and device and readable medium
CN116991842A