A method and device for intelligent processing of data standardization

By creating initial and result message queues, combining data organization specifications and standardized processing strategies, diversified data processing problems are solved, and data standardization and efficient updates are achieved, and data quality and processing efficiency are improved.

CN114547165BActive Publication Date: 2025-07-29INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210060268.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-19
Publication Date
2025-07-29
Estimated Expiration
2042-01-19

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process data in diversified data formats and structures, which leads to difficulty in data utilization and difficulty in forming standardized data, increasing resource investment in data sharing and downstream computing.

Method used

Using the intelligent data standardization processing method, the initial and result message queues are created, combined with data organization specifications, standardized processing strategies and knowledge bases, and standardized processing strategies and knowledge bases are realized, and the message bus is used for cache and data updates.

Benefits of technology

It realizes standardized processing of different data types and structures, reduces manual operation error rate, improves data quality and processing efficiency, and enhances the ease of use and adaptability of the device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114547165B_ABST
    Figure CN114547165B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for intelligent processing of data standardization. The method includes: creating an initial message queue; creating a result message queue; obtaining and organizing the data to be processed from the data warehouse according to the initial data processing strategy, and pushing it to the message bus; parsing and performing standardization processing on the messages in the message bus, and writing the results back to the message bus according to the result data processing strategy; parsing the result messages in the message bus and updating them to the data warehouse. The present invention is instantiated and customized personalized by the user through configuring the data standardization knowledge base, decouples the data source and the data standardization processing function by using the message bus, and has good adaptability and scalability. The present invention realizes the unified standardization processing of data from multiple sources with inconsistent content formats, forms standardized data, improves the intelligence and automation of data standardization processing, reduces the manual operation error rate, and thus improves the data processing efficiency and data quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data governance, and particularly relates to a method and device for intelligent processing of data standardization. Background Art

[0002] Today, with the rapid development of informatization, a vast amount of data is generated every moment. If this data can be utilized and its value fully exploited, it can bring great help to the development of society and economy. However, due to the diverse channels of data generation, information with the same meaning shows diversity in data format, data structure, etc., which brings difficulties to data utilization. Therefore, how to obtain available and easy-to-use data to make the use of data more standardized is an important foundation and prerequisite for realizing the value of data, and also one of the effective ways to improve data utilization efficiency.

[0003] Usability is the basic requirement for data to be utilized. Incorrect or incomplete data cannot effectively assist business and is invalid data for business, lacking usability; ease of use is the basis for data to be widely accepted and used by users. Therefore, defining and processing data in a standardized manner to obtain standardized data will greatly reduce the difficulty of data sharing and the resource investment in downstream data calculation.

[0004] In response to the above application requirements, the present invention proposes a new method to realize the intelligence of the data standardization processing process, simplify user operations, reduce the usage difficulty, and improve the efficiency of data standardization processing, which has practical application value and application prospects. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and device for intelligent processing of data standardization, which is applicable to the data standardization processing of fine-grained (such as database table fields), supports the standardized processing of field information with different data types and different content structures to form standardized data. At the same time, by using technologies such as human-computer interaction visualization, message bus, and plug-and-play component management, the device is capable of quickly adapting to changes in data or business requirements such as the characteristics of field content and processing standardization requirements, improving usability and adaptability.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] A method for intelligent processing of data standardization, the steps of which include:

[0008] 1) Create an initial message queue (IMQ: Initial Message Queue);

[0009] 2) Create a result message queue (RMQ: Result Message Queue);

[0010] 3) According to the Initialize Data Processing Strategy (IDPS), read the field information to be standardized from the data warehouse, organize the field information into a message according to the Data Schema (DS), and write it into the IMQ;

[0011] 4) Read a message from the IMQ, perform standardization processing, and according to the Result Data Processing Strategy (RDPS), write the processing result into the RMQ;

[0012] 5) Sequentially obtain messages from the RMQ, parse and restore the message content to obtain the standardization processing result, and write the result data back to the data warehouse.

[0013] Furthermore, use the message bus system to cache the initial message queue.

[0014] Furthermore, use the message bus system to cache the result message queue.

[0015] Furthermore, the initial data processing strategy includes: defining the fields to be subject to standardization processing, where the fields referred to here can be one or more, and multiple fields can be from one data table or multiple data tables.

[0016] Furthermore, the Data Schema (DS) defines the message content format.

[0017] Furthermore, step 3) can be repeated to write multiple messages into the IMQ in sequence.

[0018] Furthermore, the processing capabilities and processing logic of the standardization processing can be custom-configured by the user through configuring the data standardization knowledge base.

[0019] Furthermore, the data standardization knowledge base includes: the Data Standardization Operation Dictionary (DSOD), the Data Process Unit (DPU), the data standardization rule set, and the data standardization model.

[0020] Furthermore, the Data Standardization Operation Dictionary (DSOD) includes: field identifier, field category identifier, and data standardization processor identifier, and the DSOD is custom-configured by the user during specific implementation.

[0021] Further, the data normalization processor DPU instantiates the normalization processing operation by loading data normalization rules or a data normalization model. Among them, the data normalization rules or the data normalization model are user-defined and configured according to the data normalization requirements during specific implementation.

[0022] Further, the normalization processing includes:

[0023] 4.1) Read a message from the IMQ and parse it, restore the message content in sequence, and organize it in the form of key-value pairs <K i , V i >. K i (Key) is the identifier of field i, and V i (Value) is the attribute value of field i, where i = 1, 2,..., n, and n is the number of fields;

[0024] 4.2) Use <K i , V i > as the input, and implement the normalization processing operation based on the data normalization knowledge base:

[0025] (1) Use K i as the retrieval condition, retrieve based on the data normalization meta-operation dictionary DSOD, and obtain the associated data normalization processor DPU i , where i = 1, 2,..., n. This step is only executed when processing the first message;

[0026] (2) Input V i into DPU i , after executing the normalization processing operation, output the normalized result value SV i (StandardValue), and obtain <K i , SV i >;

[0027] (3) Process <K i , V i > in sequence, where i = 1, 2,..., n, and according to the result data processing strategy RDPS, organize the <K i , SV i > key-value pairs into one or more messages, and write them into the corresponding RMQ j (j = 1, 2,..., m) respectively.

[0028] Further, the result data processing strategy RDPS can be customized by the user. For example: write the result data into a message queue RMQ j (j = 1), or write it into multiple message queues RMQ j (j = 1, 2,..., m).

[0029] Further, messages are obtained sequentially from RMQ j (j = 1, 2, …, m), the message content is parsed and restored to obtain the standardized processing result, and the result data is written back to the data warehouse. Specifically, it includes:

[0030] 5.1) Parse and restore the message content to obtain <K i , SV i > value pairs;

[0031] 5.2) According to K i , update SV i to the data warehouse.

[0032] An intelligent data standardization processing device includes a data synchronization engine, a message bus, a data standardization engine, a data write-back engine, and a data standardization knowledge base;

[0033] The data synchronization engine extracts the data to be standardized from the data warehouse, encapsulates it into a message, and pushes it to the initial message queue of the message bus;

[0034] The data standardization engine reads the message from the initial message queue of the message bus, calls the data standardization knowledge base for data standardization processing, and writes the processing result into the result message queue of the message bus;

[0035] The data write-back engine parses the data in the result message queue of the message bus, writes the standardized result data back to the data warehouse, and realizes data update;

[0036] The data standardization knowledge base is used to implement the logical definition of data standardization processing and the instantiation and assembly of tool components.

[0037] An electronic device includes a memory and a processor, where the memory stores a program for executing the above method.

[0038] A storage medium stores a computer program, where the computer program is set to execute the above method when running.

[0039] Compared with existing methods, through the above-mentioned methods and devices of the present invention, each organization can be specifically instantiated in combination with the characteristics of the data assets it owns, the characteristics of business application requirements, etc. Abstract data standardization requirements according to business application requirements and build a data standardization knowledge base. Build an initial data processing strategy IDPS and a result data processing strategy RDPS according to the characteristics of the data source, help the organization quickly build a set of data standardization intelligent processing devices, realize the standardized processing of data field information of different data types and different content structures, form standardized data, reduce the manual operation error rate, and improve data quality. Moreover, by customizing the initial data processing strategy, the result data processing strategy and the data standardization knowledge base, the flexible adaptation to data and business can be realized, and the usability, adaptability and scalability of the device can be further improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a flowchart of the method according to an embodiment of the present invention.

[0041] Figure 2 It is a schematic diagram of the overall device structure according to an embodiment of the present invention.

[0042] Figure 3 It is a flowchart of constructing the basic operating environment of the device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0044] It should be noted that the terms "including" and "having" in the specification and claims of the present invention and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0045] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0046] A data standardization intelligent processing method according to an embodiment of the present invention is as Figure 1 shown, and includes the following steps:

[0047] S1. Initial data acquisition: According to the initial data processing strategy IDPS, acquire the initial data to be standardized from the data warehouse.

[0048] S2. Initial message production: Organize the acquired initial data into messages according to the requirements of the data organization specification DS, and push them into the corresponding initial message queue IMQ.

[0049] S3. Initial message parsing: Read the messages from the initial message queue IMQ, parse the messages, restore the message content, and organize them in the form of <K i , V i > (i = 1, 2,..., n) key-value pairs.

[0050] S4. Data standardization processing: According to K i , find the corresponding V i , input V i into the data standardization processor DPU i , and output SV i .

[0051] S5. Result message production: According to the result data processing strategy RDPS and the data organization specification DS, organize <K i , SV i > (i = 1, 2,..., n) into result data messages, and write them into the corresponding result message queues RMQ j (j = 1, 2,..., m) respectively.

[0052] S6. Result message parsing: Sequentially obtain the messages from the result message queues RMQ j (j = 1, 2,..., m), parse and restore the message content, and organize them in the form of <K i , SV i > (i = 1, 2,..., n) key-value pairs.

[0053] S7. Result data write-back: According to K i , write the corresponding standardized processing result data SV i back to the data warehouse.

[0054] Further, the initial data processing strategy IDPS can be configured according to domain expert knowledge or user-defined.

[0055] Further, the data organization specification DS adopts the JSON format to encapsulate the <Key, Value> key-value pair data.

[0056] Further, the data normalization processor DPU invokes a preset normalization processing rule or normalization processing model to perform corresponding normalization operations.

[0057] Further, the normalization processing rule is used to clean, transform, filter, etc. the data, and can be configured according to domain expert knowledge or user-defined.

[0058] Further, the normalization processing model mainly realizes entity extraction, relationship extraction, etc. for unstructured data, and can be customized and developed according to data normalization processing requirements.

[0059] Further, the result data processing strategy RDPS can be configured according to domain expert knowledge or user-defined.

[0060] As another aspect of the present invention, there is provided a data normalization intelligent processing device, as Figure 2 shown. For example, the data normalization intelligent processing device in this embodiment can create and apply the following data normalization instances, including:

[0061] 1. Design the external interfaces of the device, including data input and data output interfaces:

[0062] 1) Data input: Support the input of data from multiple data sources and in multiple data formats;

[0063] 2) Data output: After the data is normalized, it is written back to the original data storage area and the original data is updated.

[0064] 2. Design the internal functional components of the device. Include multiple modules such as a data synchronization engine, a message bus, a data normalization engine, a data write-back engine, a data normalization knowledge base, etc.

[0065] 1) The data synchronization engine extracts the data to be normalized from the data warehouse and encapsulates it as a message and pushes it to the message bus.

[0066] Further, the data synchronization engine is instantiated by customizing and configuring the initial data processing strategy IDPS to configure the data input interface of the device.

[0067] Further, the data write-back engine is instantiated by customizing and configuring the result data processing strategy RDPS to configure the data output interface of the device.

[0068] 2) The message bus realizes message caching and shared exchange.

[0069] Furthermore, the message bus supports multiple message queue caches. The message queues are automatically created based on IDPS and RDPS, and the number and message content format of the message queues are jointly determined by IDPS, RDPS, and DS.

[0070] 3) The data normalization engine performs data normalization processing.

[0071] Furthermore, messages are read from the initial message queue of the message bus.

[0072] Furthermore, when the first message in the message queue is read, the initialization of the engine is performed. During the engine initialization process, the data normalization processors corresponding to each field are found respectively and automatically mounted to the data normalization engine. If the mounting is successful, the status of the data normalization processor is set to ready, otherwise it is empty.

[0073] Furthermore, the messages are parsed and the fields contained in the messages are automatically identified.

[0074] Furthermore, if the status of the data normalization processor corresponding to the field is ready, the field value is subjected to data normalization processing.

[0075] Furthermore, the normalization processing results of each field value are organized according to the requirements of RDPS and DS and written into the result message queue in the message bus.

[0076] 4) The data write-back engine is used to parse the data in the result message queue and write the normalized result data back to the data warehouse to achieve data update.

[0077] 5) The data normalization knowledge base is used to implement the logical definition of data normalization processing and the instantiation and assembly of tool components.

[0078] Furthermore, the data normalization knowledge base includes: data normalization meta-operation dictionary, data normalization processor, data normalization processing rules, data normalization processing model, etc.

[0079] Furthermore, the data normalization meta-operation dictionary is configured by the user.

[0080] Furthermore, the data normalization processor is a software program that loads data normalization processing rules or data normalization processing models and performs normalization processing operations.

[0081] Furthermore, the data normalization rules are configured by the user.

[0082] Furthermore, the data normalization model is customized and developed according to the requirements of data processing.

[0083] Furthermore, the data standardization knowledge base can be gradually accumulated during the deployment and application of the device, instantiated according to the characteristics of the data source and the business application, and support expansion.

[0084] 3. Build a basic operating environment for the data standardization intelligent processing device that meets business requirements ( Figure 3 ), including:

[0085] 1) Initialize the data standardization knowledge base: Through the visual operation interface of the data standardization knowledge base, configure the basic information of the data standardization meta-operation dictionary, data standardization processor, data standardization rules, and data standardization model.

[0086] 2) Configure the data standardization processing strategy: When there is a data standardization processing requirement, configure the corresponding IDPS and RDPS through the visual operation interface of the data synchronization engine. For example: Read fields L1, L2, and L3 from Table A, write them into IMQ1, write the result information of the standardized processing of L1 and L2 into RMQ1, and write the result information of the standardized processing of L3 into RMQ2. New data standardization processing requirements can be met by configuring new IDPS and RDPS.

[0087] 3) Expand the data standardization knowledge base: When the existing data standardization knowledge base cannot meet the data standardization processing requirements, the data standardization processor components, rules, models, and operation dictionaries in the data standardization knowledge base can be expanded to adapt to new requirements.

[0088] 4. When the above-built data standardization intelligent processing device specifically processes data, it adopts a data standardization method provided by the present invention, including the following steps:

[0089] First, call the data synchronization engine ① to obtain the initial data to be processed, organize it according to the specification, and write it into the message bus ②.

[0090] Second, the data standardization processing engine ③ reads the message from the message bus ② and performs data standardization processing.

[0091] Furthermore, the data standardization processing engine ③ writes the processing result information into the message bus ②.

[0092] Furthermore, the data write-back engine ④ reads the standardized result information from the message bus ② and updates it to the data warehouse.

[0093] So far, the processes of initial data reading, data standardization processing, and standardized result data update are completed.

[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A method for intelligent processing of data standardization, characterized in that, It includes the following steps: Create an initial message queue; Create a result message queue; According to the initial data processing strategy, read the field information to be standardized from the data warehouse, organize the field information into a message according to the data organization specification, and write it into the initial message queue; Read a message from the initial message queue, perform standardization processing, and write the processing result into the result message queue according to the result data processing strategy; Successively obtain messages from the result message queue, parse and restore the message content, obtain the standardized processing result, and write the result data back to the data warehouse; The said standardization processing is to read a message from the initial message queue and implement the data standardization processing operation based on the data standardization knowledge base; the data standardization knowledge base includes: a data standardization meta-operation dictionary, a data standardization processor, a data standardization rule set, and a data standardization model; the processing ability and processing logic of the data standardization knowledge base are customized and configured by the user to flexibly adapt to the data characteristics and standardization processing requirements; The data standardization meta-operation dictionary includes: field identifier, field category identifier, and data standardization processor identifier; The data standardization processor realizes the instantiation of the standardization processing operation by loading the data standardization rule or the data standardization model; among them, the data standardization rule or the data standardization model is customized and configured by the user according to the data standardization requirements during specific implementation; The said standardization processing includes: Read a message from the initial message queue and parse it. Restore the message content in sequence and organize it in the form of key-value pairs <K i , V i >; K i is the identifier of field i, and V i is the attribute value of field i, where i = 1, 2,..., n, and n is the number of fields; With <K i , V i > as the input, implement the standardization processing operation based on the data standardization knowledge base, including: (1) With K i as the retrieval condition, retrieve based on the data standardization meta-operation dictionary to obtain the associated data standardization processor DPU i , i = 1, 2,..., n; this step is only executed when processing the first message; (2) Input V i into the DPU i . After performing the normalization operation, output the normalized result value SV i , and obtain <K i , SV i >; (3) Process <K i , V i >, where i = 1, 2,..., n, and according to the result data processing strategy, organize <K i , SV i > key-value pairs into one or more messages and write them into the corresponding result message queue RMQ j respectively, where j = 1, 2,..., m.

2. The method according to claim 1, characterized in that, The initial data processing strategy includes: defining the data fields to be subject to standardization processing, and the fields referred to can be one or more, and multiple fields can come from one data table or multiple data tables.

3. The method according to claim 1, characterized in that Define the message content format through the data organization specification.

4. The method according to claim 1, wherein Write the standardized processing result data into one or more result message queues according to the result data processing strategy.

5. An intelligent data standardization processing device using the method according to any one of claims 1-4, characterized in that, It includes a data synchronization engine, a message bus, a data standardization engine, a data write-back engine, and a data standardization knowledge base; The data synchronization engine extracts the data to be subject to standardization processing from the data warehouse, encapsulates it as a message, and pushes it to the initial message queue of the message bus; The data standardization engine reads the message from the initial message queue of the message bus, calls the data standardization knowledge base for data standardization processing, and writes the processing result into the result message queue of the message bus; The data write-back engine parses the data in the result message queue of the message bus, writes the standardized processing result data back to the data warehouse, and realizes data update; The data standardization knowledge base is used to realize the logical definition of data standardization processing and the instantiation and assembly of tool components.

6. A storage medium, characterized in that, The computer program is stored in the storage medium, and the computer program is set to execute the method described in any one of claims 1-4 when running.

7. An electronic device, characterized in that, It includes a memory and a processor, the computer program is stored in the memory, and the processor is set to run the computer program to execute the method described in any one of claims 1-4.

Citation Information

Patent Citations

  • Network data cleaning method and system and related device

    CN112181961A