Data standardization processing method, system, device and medium based on business requirements

CN122840463APending Publication Date: 2026-09-29浪潮智慧科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610736389.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-26
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而,随着企业业务需求的快速迭代和业务系统的频繁变更,这种被动、周期性的标准更新机制逐渐暴露出固有缺陷:数据标准的制定与修订往往滞后于业务变更,导致标准与真实业务数据之间产生脱节,使得数据质量校验规则无法覆盖新增字段或新枚举值,从而引发数据接入失败或质量下降

Benefits of technology

[0014]本发明的有益效果在于,本发明提供的基于业务需求的数据标准化处理方法、系统、设备及介质,通过主动感知业务变更、量化评估影响范围、自动生成标准草案并驱动存量数据自适应清洗,构建了数据标准与业务需求之间的闭环同步机制。相比传统被动更新模式,本方法将标准响应时间从数周级缩短至准实时级,彻底消除标准滞后导致的业务数据脱节问题;同时,通过关联图谱量化分析影响度,确保标准变更的全面性和精确性,避免人工遗漏或误判;自动化生成的存量数据清洗脚本结合灰度发布策略,一次性解决了新老标准并存的历史数据不一致难题,显著降低了人工干预成本与数据治理风险。最终,使数据标准体系具备与业务协同演进的生命力,支撑企业数字化业务的敏捷迭代。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840463A_ABST
    Figure CN122840463A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, and specifically provides a data standardization processing method, system, device and medium based on business requirements, comprising: capturing business system change events in real time and extracting them as structured change information; formalizing the information as a change vector, quantitatively calculating the comprehensive influence degree of each standard item on a pre-constructed business item and standard item association graph to determine the standard item to be changed; using a decision tree to route the change vector attribute to the corresponding draft generator, dynamically generating a standard update draft containing rule parameters; pushing the draft to the review end, and updating to the data standard library after review; and performing real-time standardization processing on incremental data according to the updated standard library, and performing adaptive cleaning and conversion on stock data. The present application realizes quasi-real-time synchronization of data standards and business requirements, and improves the agility and consistency of data governance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data processing technology, specifically relating to a data standardization processing method, system, device, and medium based on business needs. Background Technology

[0002] In the field of data governance, data standardization is a crucial step in ensuring data consistency, availability, and interoperability. Traditional data standardization methods typically employ a static model of "standards first, then application to data," where the data governance team predefines data standards based on business needs and integrates them into the data processing workflow. However, with the rapid iteration of business requirements and frequent changes in business systems, this passive and periodic standard update mechanism has gradually revealed inherent flaws: the formulation and revision of data standards often lag behind business changes, leading to a disconnect between standards and actual business data. This results in data quality verification rules failing to cover newly added fields or enumerated values, causing data access failures or quality degradation. Summary of the Invention

[0003] In view of the above-mentioned shortcomings of the prior art, the present invention provides a data standardization processing method, system, device and medium based on business needs to solve the above-mentioned technical problems.

[0004] In a first aspect, the present invention provides a data standardization processing method based on business needs, comprising: Capture change events from the business system and extract the change events into structured change information; The structured change information is formalized into a change vector. The comprehensive impact of the change vector on each standard item node is quantitatively calculated on the pre-constructed association graph between business item nodes and standard item nodes. Based on the comprehensive impact, the standard item to be changed is determined. The decision tree is used to route the standard item to be changed to the corresponding draft generator according to the attributes of the change vector, so that the draft generator generates a standard update draft containing specific rule parameters; The draft standard update is pushed to the review panel, and the draft standard update that passes the review is updated to the data standard library. Based on the updated data standard library, incremental data is standardized, and existing data is adaptively cleaned and transformed.

[0005] In an optional implementation, change events of the business system are captured, and the change events are extracted into structured change information, including: Collect change signals, including: deploying listeners at the business entity change interface of the business system to capture business model change events in real time; interfacing with the requirements management platform to automatically parse change information in the approved requirements documents; embedding analysis modules in the data flow of the data access layer to detect the addition and change of fields or structures through pattern snapshot comparison, and to detect the addition and drift of enumeration values ​​or value ranges through historical value baseline comparison. The collected change signals are converted into structured change information in a unified format. The structured change information includes the following fields: change subject, change type, change content, change time, change source, and affected data range.

[0006] In an optional implementation, the method for constructing the association map includes: The system constructs nodes, which include business item nodes, standard item nodes, and data field nodes. The business item nodes have a business concept identifier and a change frequency attribute; the standard item nodes have a standard identifier, a standard type, a standard clause, and version attributes; and the data field nodes have a field identifier, a data type, and a sensitivity attribute. Establish relationships between nodes, including: The "belongs to" relationship is used to connect the data field node to the business item node, indicating that the data field belongs to the corresponding business item. A constraint relationship is used to connect the data field node to the standard item node, indicating that the data field is constrained by the corresponding standard item clause. The constraint relationship has a weight attribute that indicates the strength of the constraint. The association relationship is used to connect the business item node to the standard item node, representing the association between the business item and the standard item at the macro level. The association relationship has a relevance attribute that represents semantic similarity or historical co-occurrence frequency.

[0007] In an optional implementation, the comprehensive impact of the change vector on each standard item node is quantified on a pre-built association graph of business item nodes and standard item nodes, and the standard item to be changed is determined based on the comprehensive impact, including: Parse the affected business item identifier, the changed field identifier, the change operator type, the change type, and the change metadata in the change vector; For each standard item node, its comprehensive influence is calculated. The comprehensive influence is a weighted sum of multiple influence components, including: direct association influence, propagation association influence, and semantic similarity. The sum of the weight coefficients of the direct association influence, the propagation association influence, and the semantic similarity is 1. The method for determining the direct correlation influence is as follows: if the change field identifier in the change vector has a corresponding data field node in the correlation graph, then the weight of the constraint relationship between the data field node and the current standard item node is obtained, and the direct correlation influence is calculated in combination with the change operator type and change type of the change vector; otherwise, the direct correlation influence is 0. The method for determining the propagation association impact is as follows: using the affected business item identifier in the change vector to locate the corresponding business item node in the association graph, obtaining the correlation attribute of the association between the business item node and the current standard item node, as well as the shortest path hop count from the business item node to the current standard item node, and attenuating the correlation attribute according to the attenuation function corresponding to the hop count to obtain the propagation association impact. The semantic similarity is determined by calculating the semantic similarity between the change metadata in the change vector and the description text of the current standard item node. Nodes with a comprehensive impact greater than a preset threshold are identified as standard items to be changed.

[0008] In an optional implementation, a decision tree is used to route the criteria item to be changed to the corresponding draft generator based on the attributes of the change vector, including: The standard item to be changed and its corresponding change vector are input into the decision tree, and the judgment conditions are matched layer by layer starting from the root node until the leaf node is reached. Each leaf node is bound to a draft generator, which contains rule templates for specific change scenarios, and the rule templates have dynamically populated parameter bits; Based on the leaf node reached, the corresponding draft generator is invoked; The internal nodes of the decision tree serve as judgment conditions, which are based on the attributes of the change vector. The attributes of the change vector include: change type, change operator type, and standard item type.

[0009] In one optional implementation, the draft generator includes rule templates and dynamically populated parameter bits; the method by which the draft generator generates a standard update draft containing specific rule parameters includes: Extract the change content from the change vector and fill the parameter bits in the rule template. The parameter bits include at least: field name and enumeration value set. For the threshold parameter in the rule template, it is dynamically calculated based on the historical change frequency of the business items involved in the change vector and the sensitivity or importance of the fields involved. The base threshold is subtracted from the product of the historical change frequency and the first coefficient, and then the product of the sensitivity or importance and the second coefficient is added to obtain the final threshold. The filled rule template is assembled into a draft standard update for the standard item to be changed; A confidence score is calculated for the standard update draft, which is the product of the overall impact of the standard item to be changed and the historical accuracy of the draft generator. The confidence score is then appended to the standard update draft and output together.

[0010] In one optional implementation, incremental data is standardized based on the updated data standard library, and existing data is adaptively cleaned and transformed, including: In the data acquisition or processing process, the updated data standard library is loaded in real time, and data cleaning, format conversion, verification and standardized writing operations are performed on the newly generated incremental data based on the updated data standard library. Based on the changes in the draft update of the standard, a data migration script or a data cleaning script is automatically generated to process historical data in batches to make it conform to the updated data standard. The processing adopts a gray release strategy, first processing and verifying a portion of the existing data, and then promoting it to the full amount of existing data after confirming that there are no errors.

[0011] Secondly, the present invention provides a data standardization processing system based on business requirements, comprising: The change capture module is used to capture change events of the business system and extract the change events into structured change information; The impact calculation module is used to formalize the structured change information into change vectors, quantify the comprehensive impact of the change vectors on each standard item node on a pre-constructed association graph of business item nodes and standard item nodes, and determine the standard item to be changed based on the comprehensive impact. The draft generation module is used to route the standard item to be changed to the corresponding draft generator based on the attributes of the change vector using a decision tree, so that the draft generator generates a standard update draft containing specific rule parameters. The draft review module is used to push the standard update draft to the review end, and update the standard update draft that passes the review to the data standard library; The standard processing module is used to standardize incremental data based on the updated data standard library and perform adaptive cleaning and transformation on existing data.

[0012] Thirdly, a device is provided, comprising: Memory, used to store data standardization processing programs based on business needs; A processor is configured to implement the steps of the business-requirement-based data standardization processing method provided in the first aspect when executing the business-requirement-based data standardization processing program.

[0013] Fourthly, a computer-readable medium is provided, on which a data standardization processing program based on business requirements is stored, wherein when the data standardization processing program based on business requirements is executed by a processor, the steps of the data standardization processing method based on business requirements provided in the first aspect are implemented.

[0014] The beneficial effects of this invention lie in the fact that the data standardization processing method, system, equipment, and medium based on business needs provided by this invention construct a closed-loop synchronization mechanism between data standards and business needs by proactively sensing business changes, quantitatively assessing the scope of impact, automatically generating draft standards, and driving adaptive cleaning of existing data. Compared with the traditional passive update mode, this method shortens the standard response time from weeks to near real-time, completely eliminating the problem of business data disconnect caused by standard lag. At the same time, it ensures the comprehensiveness and accuracy of standard changes by quantitatively analyzing the impact through correlation graphs, avoiding human omissions or misjudgments. The automatically generated existing data cleaning script, combined with the canary release strategy, solves the problem of historical data inconsistency where new and old standards coexist in one go, significantly reducing the cost of manual intervention and data governance risks. Ultimately, this enables the data standard system to have the vitality to evolve in synergy with business, supporting the agile iteration of enterprise digital business. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic flowchart of a method according to an embodiment of the present invention.

[0017] Figure 2 This is a schematic topological diagram of the association map of a method according to an embodiment of the present invention.

[0018] Figure 3 This is a schematic flowchart illustrating the method for generating a draft standard update according to an embodiment of the present invention.

[0019] Figure 4 This is a schematic block diagram of a system according to an embodiment of the present invention.

[0020] Figure 5 This is a schematic diagram of the structure of a device provided in an embodiment of the present invention. Detailed Implementation

[0021] To enable those skilled in the art to better understand the technical solutions of this invention, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this invention.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0023] The data standardization processing method based on business requirements provided in this embodiment of the invention is executed by a computer device, and correspondingly, the data standardization processing system based on business requirements runs on the computer device.

[0024] Figure 1 This is a schematic flowchart illustrating a method according to an embodiment of the present invention. Wherein, Figure 1 The executing entity can be a standardized data processing system based on business needs. Depending on different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted.

[0025] like Figure 1 As shown, the method includes: S1. Capture change events from the business system and extract the change events as structured change information; S2. The structured change information is formalized into a change vector. The comprehensive impact of the change vector on each standard item node is quantitatively calculated on the pre-constructed association graph between business item nodes and standard item nodes. Based on the comprehensive impact, the standard item to be changed is determined. S3. Using a decision tree, the standard item to be changed is routed to the corresponding draft generator according to the attributes of the change vector, so that the draft generator generates a standard update draft containing specific rule parameters; S4. Push the draft standard update to the review end, and update the data standard library with the draft standard update that has passed the review; S5. Based on the updated data standard library, standardize the incremental data and perform adaptive cleaning and transformation on the existing data.

[0026] In one embodiment of the present invention, based on step S1, the following will provide a possible embodiment and describe its specific implementation in a non-limiting manner.

[0027] S101: Acquire change signals Business change signals are collected in real time through multiple proactive sensing methods, including at least one of the following: (1) Business System Monitoring Method: Deploy listeners at the change interfaces of key business entities (e.g., "orders", "customers", "products") in business systems (such as CRM systems, ERP systems). The change interfaces include, but are not limited to, the ALTERTABLE log capture component of the database, the request / response interceptor of the API interface, or the entity change event subscription module in the message queue. When the business system performs operations such as adding fields, modifying field types or lengths, or adding enumeration values, the listener captures these business model change events in real time and sends the original event data (such as the structure before and after the change, operation time, and operator) to the change collection center.

[0028] (2) Integration method with the demand management platform: The system is connected to the enterprise's internal demand management platform (such as Jira, ZenTao, TAPD) via REST API or Webhook. The system is pre-configured with a set of keyword rules, including keywords such as "demand", "field", "standard", "add", "modify", and "enumeration". When a demand document or user story that has been reviewed and approved is generated in the demand management platform, the system automatically pulls the document content and uses named entity recognition and dependency parsing techniques in natural language processing to parse out the information related to data item changes, such as "add a 'delivery priority' field to the order table, with a string type and allowed values ​​of high, medium, and low".

[0029] (3) Data Stream Analysis Method: An analysis module is embedded in the data access layer (e.g., a data pipeline built on Kafka, Flink, or Spark Streaming). This module maintains a current schema snapshot for each data source, which records metadata such as field names, field types, and whether they are nullable. Newly incoming structured or semi-structured data (e.g., JSON, Avro format) is parsed in real time, and the parsed schema is compared with the stored schema snapshot: if a new field name is found, it is marked as a "Structure Addition" event; if the type of an existing field changes (e.g., from integer to string), it is marked as a "Structure Change" event. At the same time, for key enumeration fields (e.g., order status), the system records a unique set of its historical values ​​as a baseline (e.g., {"Pending Payment", "Paid", "Shipped"}), and monitors the value of this field in real time in the data stream. When a new value not in the baseline appears (e.g., "Delivered"), an "Enumeration Value Addition" or "Value Domain Shift" warning is immediately triggered.

[0030] The above three acquisition methods can work in parallel and complement each other to ensure comprehensive and real-time capture of business change signals.

[0031] S102: Convert to a unified format for structured change information. The collected raw change signals come from different sources and have varying formats. The system uses an adapter pattern to convert them into unified structured change information. Specifically, a data model for change events is predefined, which includes the following fields: change subject (e.g., the "order" business entity), change type (e.g., "add field", "modify field", "delete field", "add enumeration value"), change content (e.g., "add field 'delivery_priority', type: string, length: 10, enumeration value: [high, medium, low]"), change time (using a unified timestamp format, such as "2026-04-10 10:30:00"), change source (identifying which business system, which requirement task, or which data flow Topic the event comes from), and the scope of affected data (e.g., "all historical data of the order table" or "order data of the last 30 days").

[0032] For example, after capturing the event "A new value 'Delivered' appears in the order status field" from the data stream analysis, the system converts it into the following structured information: Change subject: "Order"; Change type: "Add enumeration value"; Change content: "Add enumeration value 'Delivered' to the order status field"; Change time: "2026-04-10 10:30:00"; Change source: "order_topic data stream"; Affected data scope: "All records containing order status".

[0033] In one embodiment of the present invention, based on step S2, the following will provide a possible embodiment and describe its specific implementation in a non-limiting manner.

[0034] Please refer to Figure 2 The methods for constructing association maps include: I. Node Construction Three types of nodes are extracted from the enterprise's existing metadata management system, data standard library, and business systems: (1) Business Item Nodes: Each business item node corresponds to a business concept, such as "order", "customer", or "product". Node attributes include at least: business concept identifier (unique ID, such as "B_ORDER"), business concept name (such as "order"), and change frequency (volatility_score). Among them, change frequency is a statistic calculated by analyzing historical change records, such as the number or frequency of changes to the business entity in the past year, used to characterize the activity level of the business concept.

[0035] (2) Standard Item Nodes: Each standard item node corresponds to a specific data standard clause. Node attributes must include at least: standard identifier (unique ID, such as "S_DQ_001"), standard type (such as "data classification standard", "encoding standard", "format standard", "data quality standard", "interface standard"), standard clause (specific content description, such as "order amount field must be non-negative and retain two decimal places"), and version number (such as "v2.1"). Standard item nodes can be imported in batches from the enterprise's existing data standard management system and incremental updates are supported.

[0036] (3) Data Field Nodes: Each data field node corresponds to a physical or logical field. Node attributes include at least: field identifier (unique ID, such as "F_ORDER_AMOUNT"), data type (such as "DECIMAL(10,2)", "VARCHAR(50)"), and sensitivity (sensitivity_level, such as "high", "medium", "low", used to identify the importance or security level of the field). Data field nodes can be automatically obtained from the data dictionary or data lineage system.

[0037] II. Relationship Construction Three types of edge relationships are established by parsing existing mapping relationships, semantic matching, or manual configuration: (1) Belongs to Relation (BELONGS_TO): This relation points from a data field node to a business item node, indicating which business concept a data field belongs to. For example, the "Order Amount" field node is connected to the "Order" business item node. This relation is established based on the business affiliation definition in the data dictionary, or is automatically inferred through field naming rules (such as field name prefixes). This relation has no additional attributes.

[0038] (2) Constraint Relationship (CONSTRAINED_BY): This relationship points from a data field node to a standard item node, indicating which standard clause a data field is bound by. For example, the "Order Amount" field node is connected to the "Order Amount is Non-negative and Has Two Decimal Places" standard item node. This relationship has a weight attribute to represent the strength of the constraint. The weight can take values ​​between 0 and 1, for example, "Must Follow" corresponds to a weight of 1.0, "Recommended to Follow" corresponds to a weight of 0.7, and "Reference" corresponds to a weight of 0.3. The weight value can be automatically assigned by parsing mandatory words (such as "should", "must", "recommended") in the standard document, or manually set by data governance experts.

[0039] (3) Relationship (RELATES_TO): This relationship points from the business item node to the standard item node, representing the macro-level association between the business concept and the standard. For example, the "order" business item node can be connected to the "order data quality standard" standard item node. This relationship has a relevance attribute, which is used to represent the semantic similarity or historical co-occurrence frequency between the business item and the standard item. Semantic similarity can be obtained by calculating the textual similarity between the business item name and the standard item clause description (e.g., calculating the cosine similarity after generating vectors using the Sentence-BERT model); historical co-occurrence frequency is the percentage of times the business item and the standard item are mentioned or associated together in the historical standard review records.

[0040] III. Storage and Updating of Atlases The aforementioned nodes and relationships are stored in a graph database (such as Neo4j or JanusGraph), supporting efficient graph traversal and path lookup. The system provides an incremental update mechanism: when a new business item, data field, or standard item is added, the corresponding node creation and relationship calculation are automatically triggered (e.g., automatically recommending associations through semantic matching); when attributes such as the frequency of business changes change, the attribute values ​​of the corresponding nodes are updated periodically.

[0041] The methods for determining the standard items to be changed specifically include: S201: Formal Representation of Changing Vectors The system converts the structured change information output in step S102 into a set of standardized change vectors. Each change vector C is defined as the following quintuple:

[0042] in: Indicates the affected business item identifier, such as "order"; This indicates the data field that has changed. It can be empty if it is a newly added business concept or if the field-level change is unclear. This indicates change operators, including but not limited to: ADD_FIELD (add field), MODIFY_FIELD_TYPE (modify field type), MODIFY_FIELD_LENGTH (modify field length), ADD_ENUM (add enumeration value), DELETE_FIELD (delete field), etc. Indicates the change type, such as SCHEMA change, ENUM change, RULE change; This indicates a change to the metadata vector, which may contain specific information such as new_data_type (new data type), new_length (new length), and new_enum_set (new set of enumeration values).

[0043] For example, if the data stream detects that "the order status field has a new enumeration value 'delivered'", the corresponding change vector would be: , , , , .

[0044] S202: Quantitative Calculation of Overall Impact For each standard term node in the association graph Calculate its overall impact The overall impact is defined as the weighted sum of the three impact components:

[0045] in: Directly related influence; To spread the influence of related information; For semantic similarity; For the weighting coefficients, satisfying .

[0046] The aforementioned weighting coefficients can be obtained through training on historical review data, with the initial recommendation value set to... This means prioritizing direct impact while also considering indirect propagation and semantic relevance.

[0047] (1) Direct correlation influence Calculation The system first determines the field identifier in the change vector. Does a corresponding data field node exist in the association graph? If so, retrieve the data field node and the current standard item node. Weight of the "constraint relationship" between (Value range 0~1), and combined with the change operator and change type Calculate the basic impact value The base impact value is given by a predefined mapping table; for example, for ADD_FIELD, For MODIFY_FIELD_TYPE, For ADD_ENUM, For MODIFY_FIELD_LENGTH, .

[0048] The formula for calculating the direct correlation influence is:

[0049] If field identifier If there is no corresponding data field node in the graph, or if there is no constraint relationship between the current standard item node and the data field node, then... .

[0050] (2) The degree of influence of the spread Calculation The system identifies the business items in the change vector. Locate the corresponding business item node in the association graph. Then, starting from that business item node, traverse along the "association relationship" edge to the current standard item node. Get the relevance attributes of the edges. (Value range 0~1), and the shortest path hop count from the business item node to the standard item node. (generally Indicates a direct association. This indicates an indirect connection via an intermediate node.

[0051] The propagation correlation influence is obtained by attenuating the correlation attribute using an attenuation function:

[0052] Where the attenuation function Exponential decay or linear decay can be used. This embodiment uses exponential decay:

[0053] in This is the attenuation coefficient (recommended value: 0.5). When hour, No decay; when hour, This indicates a moderate reduction in indirect impact. If there is no path between the business item node and the standard item node, then... .

[0054] (3) Semantic similarity Calculation The system calculates metadata in the change vector. With the current standard item node This involves semantic similarity between descriptive texts (such as the standard clause field or standard name). Specifically, a pre-trained Sentence-BERT model is used to... Text fields (such as new field names and enumeration value descriptions) and description texts of standard item nodes are encoded into fixed-length semantic vectors. The cosine similarity between the two vectors is then calculated, with the result ranging from 0 to 1. This similarity can uncover "hidden" influencing standards that are not explicitly linked in the graph but are highly relevant to the changes in content.

[0055] S203: Determination of Standard Items to be Changed The system calculates the overall impact of all standard item nodes. Then, sort them from highest to lowest, and prioritize those with a comprehensive impact greater than a preset threshold. The node is identified as a standard item to be changed. Threshold This value can be set by the system administrator based on business sensitivity; a value of 0.3 is recommended. For standard items below the threshold, the system determines that they are not affected by this change and no update draft needs to be generated.

[0056] In one embodiment of the present invention, based on step S3, the following will provide a possible embodiment and its specific implementation will be described in a non-limiting manner. Please refer to [the relevant documentation]. Figure 3 .

[0057] S301: Decision Tree Route Matching A "Change-Standard Type Decision Tree" is pre-constructed. This decision tree is a multi-branch tree, where internal nodes represent decision conditions and leaf nodes represent specific draft generators. The decision conditions for each internal node are based on one or more of the following attributes of the change vector: change type (e.g., SCHEMA change, ENUM change, RULE change), change operator type (e.g., ADD_FIELD, MODIFY_FIELD_TYPE, ADD_ENUM), and standard item type (e.g., data quality standard, coding standard, format standard). The depth and branching structure of the decision tree can be configured according to the actual business scenario. For example: the root node determines the change type at the first level; the change operator type at the second level; and the standard item type at the third level.

[0058] Once a standard item to be changed and its corresponding change vector are obtained, the change vector is input into the decision tree. Starting from the root node, the attribute values ​​of the change vector are evaluated based on the judgment conditions of the current node. A matching branch is selected to proceed to the next level node, and so on recursively until a leaf node is reached. If no branch can be matched at a certain level (e.g., the change operator type is not within the predefined range), the process is transferred to manual fallback or the default general draft generator is used.

[0059] For example, suppose a change vector has the following attributes: the change type is "ENUM Change", the change operator is "ADD_ENUM", and the type of the standard item to be changed is "Data Quality Standard". The decision tree first matches the "ENUM Change" branch at the root node, then matches the "ADD_ENUM" branch at the second level, and finally matches the "Data Quality Standard" branch at the third level, ultimately reaching the leaf node bound to the "Enumerated Value Compliance Rule Generator".

[0060] S302: Binding and Structure of Draft Generator Each leaf node is pre-bound to a draft generator. The draft generator is an executable rule generation module containing rule templates for specific change scenarios. Rule templates are defined as parameterized strings, with several reserved parameter slots for dynamic population. Parameter slot types include, but are not limited to: field name placeholders (e.g., {field_name}), enumeration value set placeholders (e.g., {enum_set}), threshold placeholders (e.g., {threshold}), and data type placeholders (e.g., {data_type}).

[0061] Different leaf nodes are bound to different draft generators, for example: For the "adding a new field" scenario, the draft generator includes a "field existence check rule" template; For the "modify field type" scenario, the draft generator includes a "type compatibility conversion rule" template; For the "adding new enumeration values" scenario and the standard type is data quality standard, the draft generator includes an "enumeration value compliance rule" template, the content of which is: "The enumeration values ​​of field {field_name} must belong to the set {enum_set}, and the compliance rate threshold = {threshold}%". S303: Call the draft generator and generate a draft. Based on the leaf node reached in step S301, the draft generator bound to that leaf node is invoked. During the invocation, the system passes the specific information from the current change vector as input parameters to the draft generator. The draft generator performs the following operations: (1) Parameter extraction and filling: Extract the relevant information from the change vector and fill the parameter bits in the rule template. For example, for the above "enum value compliance rule" template, the system extracts the newly added enum value set from the change metadata M of the change vector to fill {enum_set}, and extracts the field name from the field identifier F to fill {field_name}.

[0062] (2) Dynamic Threshold Calculation: For threshold parameters (such as {threshold}) in the template, the system does not use a fixed default value, but calculates them dynamically based on the business context. The calculation formula is as follows:

[0063] in: Base is the base threshold, for example, 99.9%; Volatility is the frequency of changes of business item B in the change vector (derived from the volatility_score attribute of the business item node in the association graph). The more frequent the changes, the more relaxed the threshold should be (subtract this item). Criticality represents the sensitivity or importance of a changed field (derived from the sensitivity_level attribute of the data field node in the association graph, which can be quantified numerically). The higher the importance, the stricter the threshold (this item is added). To adjust the coefficients, for example, take 0.1 and 0.05 respectively.

[0064] This dynamic threshold calculation makes the generated rules more closely aligned with actual business risks.

[0065] (3) Draft assembly: Assemble the completed rule template (and possibly multiple rules) into a complete draft update for the current standard item to be changed. The draft content usually includes: standard item identifier, change type, specific rule clauses to be added or revised, suggested effective date, and explanation of the scope of impact.

[0066] (4) Confidence rating: Optionally, the system calculates a confidence score for the generated draft:

[0067] in The comprehensive influence of the standard item calculated in step S103. This represents the accuracy of the draft generator on historical data (continuously updated by the system). The confidence score is output along with the draft for reference in subsequent human-machine collaborative review steps: drafts with high confidence (e.g., >0.8) can be recommended for automatic adoption, while drafts with low confidence require intensive human review.

[0068] In one embodiment of the present invention, based on step S4, the following will provide a possible embodiment and describe its specific implementation in a non-limiting manner.

[0069] S401: Draft Submission and Review Display The draft standard update generated for each standard item to be changed will be assembled into a complete change proposal. This proposal will include at least the following: the source of the change (the original signal that triggered the change, such as a business system monitoring event, a requirement platform task, or data flow analysis result); a change vector summary (affected business items, changed fields, change type, and change content); an impact analysis report (a list of affected standard items and their overall impact and confidence score); and the main text of the generated draft standard update (including specific rules, parameter thresholds, and suggested effective dates for new or revised rules or clauses).

[0070] The above proposals are pushed to the collaborative workbench of the data standards review team via an internal message queue or REST API. This collaborative workbench can be a web application that supports multi-user login, online viewing, annotation, modification, and voting. Upon push, the system automatically assigns tasks based on preset reviewer roles (e.g., business owner, data governance expert, system architect) and notifies relevant personnel via email, instant messaging, or in-system notifications.

[0071] S402: Human-Machine Collaborative Audit and Decision-Making The review panel reviewed the proposals on the collaborative workbench. Specifically: (1) Online viewing and comparison: The system highlights the differences in the draft standard compared to the currently effective standard (e.g., newly added enumeration values, adjusted thresholds). Reviewers can simultaneously view the standard versions before and after the change, the basis for the change (original signal screenshots or requirement links), and the visualization of the impact analysis.

[0072] (2) Modification and Confirmation: Reviewers can directly edit and modify the draft text, such as adjusting threshold parameters, modifying rule descriptions, and adding exception conditions. After modification, click the "Confirm" or "Reject" button. The system records the operation log of each reviewer, including the modification content, confirmation time, and comments.

[0073] (3) Automated approval strategy: For simple, low-risk changes, the system can be configured with automated approval rules, which can be automatically approved without manual intervention. The conditions for automated approval include, but are not limited to: confidence score greater than a preset threshold (e.g., >0.85); the change type is only the addition of enumerated values ​​and does not involve the deletion or modification of existing values; the frequency of changes to the affected business items is lower than a preset level (indicating that the business is stable); and the overall impact is less than a certain low-risk threshold.

[0074] Drafts that meet the criteria for automated approval are directly marked as "automatically approved" by the system and skip the manual review process, achieving effectiveness within seconds. Drafts that do not meet the criteria are then placed in the manual review queue.

[0075] (4) Decision Result Output: After the review is completed, the system summarizes the opinions of all reviewers. For multi-person review scenarios, the decision rules of "one-vote veto" or "majority pass" can be adopted (configurable). The final decision result includes: pass (including pass after modification), rejection, and return for modification (requiring the system to regenerate the draft or supplement information).

[0076] S403: Dynamic updates to the standard library For standard update drafts whose decision result is "passed", the system triggers the standard library update process: (1) Version control: The system reads the currently effective standard item version from the dynamic data standard library, and uses the new draft as the new version of the standard item (the version number is automatically incremented, for example, from v2.1 to v2.2), while retaining historical versions for traceability and rollback.

[0077] (2) Atomic Update: The system writes the new version of standard terms, rule parameters, effective time, approver, approval time, and other information into the corresponding standard item node in the data standard library. The update operation uses database transactions to ensure atomicity and avoid data inconsistency caused by partial writes.

[0078] (3) Notification Release: After the standard library is successfully updated, the system sends a "Standard Update Notification" to all registered subscribers (such as data acquisition systems, ETL tasks, data quality monitoring platforms, and data service gateways) via a message broadcast mechanism. The notification includes the standard item identifier, the new version number, a change summary, and an API endpoint for retrieving the complete standard content. After receiving the notification, subscribers can immediately load the latest standard from the cache or database, achieving "hot loading".

[0079] (4) Audit Log: The system records the complete audit log for this standard update, including: the content of the version before the change, the content of the version after the change, the approver, the approval time, and the automated approval identifier (if any), in order to meet compliance requirements.

[0080] S404: Exception Handling and Rollback If feedback from the business system or an alarm from the monitoring system is received within a certain period (e.g., 24 hours) after the standard update, indicating that the new standard has caused data processing anomalies or business interruptions, the system supports rapid rollback. Authorized administrators can select the "Rollback to Previous Version" operation through the collaborative workbench. The system will automatically restore the corresponding standard items in the data standard library to the version before the update and reissue a notification to all subscribers to reload the old version of the standard.

[0081] In one embodiment of the present invention, based on step S5, a possible embodiment will be given below, and its specific implementation will be described in a non-limiting manner.

[0082] S501: Real-time Standardization Processing of Incremental Data Embed a standard execution engine within the data acquisition or processing workflow. This engine maintains near real-time synchronization with the data standard library. The specific implementation method is as follows: (1) Hot loading of the standard library: When the standard execution engine starts, it loads all currently effective standard items and their rule parameters from the data standard library and caches them in local memory or a distributed cache (such as Redis). At the same time, the engine subscribes to the standard library's "standard update notification" through a message queue (see the release notification mechanism in step S402). Once a notification is received, the engine immediately pulls the latest standard from the cache or database and refreshes the local rule set, achieving millisecond-level hot loading of the standard without restarting the data processing task.

[0083] (2) Incremental data processing: For newly generated incremental data in the data collection or processing process (such as order messages accessed in real time via Kafka, or partition data added in batch ETL tasks), the standard execution engine performs the following operations one by one or in batches according to the latest loaded data standard: Data cleaning: Remove blank characters and special symbols that do not meet the format requirements, and handle missing values ​​(such as filling with default values ​​or marking abnormalities). Format conversion: Convert data to the required type and length, for example, convert string date fields to `YYYY-MM-DD` format, and truncate or report errors for fields that exceed the length according to rules; Validation: Field values ​​are validated according to the rules in the standard (such as enumerated value whitelist, numerical range, regular expression). Data that fails the validation is routed to the exception queue or labeled with a quality tag. Standardized writing: The processed data is written to the target data storage (such as a data lake, data warehouse, or business database) according to the structure defined by the standard, ensuring that the written data fully complies with the currently effective standard.

[0084] For example, if the data standard library adds a new rule: "The enumeration value of the order delivery priority field `delivery_priority` must be {high, medium, low}", then the standard execution engine will mark records in incremental order data with the value of "extremely high" as validation failures, and will follow a preset strategy (such as refusing to write, replacing it with the default value "medium", or writing to the error table and triggering an alarm).

[0085] S502: Adaptive Cleaning and Transformation of Existing Data When data standards change (such as adding fields, modifying enumeration values, or adjusting data types), existing historical data often does not conform to the new standards. The system automatically generates scripts and adopts a canary release strategy to process existing data in batches.

[0086] (1) Automatic generation of migration / cleaning scripts: Analyze the changes in the standard update draft generated in step S303, identify the change types, and call the script generator to automatically generate the corresponding data migration scripts. The script generator supports multiple execution engines, including SQL (suitable for relational databases), Spark jobs (suitable for big data platforms), and Python scripts (suitable for general file processing). Example of specific mapping rules: For "new fields" with default values ​​or calculation rules: generate an `ALTER TABLE ADD COLUMN` statement and attach an `UPDATE` statement to fill in the default values ​​according to the business rules (e.g., "set to 'high' when order amount > 1000, otherwise set to 'low'").

[0087] For "Modify Enumeration Values" or "Expand Value Range": Generate a data cleaning script to map old enumeration values ​​in the existing data to a new set of enumeration values, or correct newly appearing valid values ​​(data that was previously marked due to validation failure).

[0088] For "Modify field type / length": Generate a type conversion script, for example, to remove non-numeric characters from a `VARCHAR` type mobile number field and convert it to `CHAR(11)`.

[0089] For "Delete Field": Generate a script to back up the data in this field and then perform the deletion operation, or perform data anonymization in accordance with compliance requirements.

[0090] (2) Gray-scale release strategy: To control risks, the system adopts a gray-scale release approach to perform existing data cleaning, which is divided into the following stages: Phase 1: Small-scale validation. The system first selects a small representative portion of the existing data (e.g., the time range is the most recent 7 days, the data volume does not exceed 1% of the total, or a specific business area such as "North China orders") as a grayscale sample. An automatically generated cleaning script is executed on this sample, and the processing results are recorded (e.g., the number of processed records, the number of failed records, and the comparison samples before and after conversion).

[0091] Phase Two: Effect Verification and Manual Confirmation. After processing, the system will automatically perform the following verification operations: Verify whether the standard compliance rate of the sample data meets expectations (e.g., ≥99.9%). Check for any unexpected data loss or formatting errors; Compare whether key business metrics (such as total order amount and number of users) remain consistent before and after the processing (for changes that only alter field formats without changing semantics). Simultaneously, the system will push the verification report and sampled data to the review panel, requesting confirmation from data governance experts or business leaders. If issues are found, the changes to the sample data can be rolled back, the script logic adjusted, and the verification re-implemented.

[0092] Phase 3: Phased rollout to full data. After successful verification and confirmation, the system will apply the cleaning script to the remaining existing data in batches. Batch strategies can include time-based sharding (e.g., one batch per month), primary key hash-based sharding, or parallel execution by data partitioning. After each batch is completed, the system checks the execution log and error rate. If the error rate exceeds a preset threshold (e.g., 1%), it will automatically pause and issue an alarm, awaiting manual intervention; otherwise, it will continue to the next batch until all existing data has been processed.

[0093] (3) Exception handling and rollback: The system retains a complete snapshot of the data before and after each data cleaning operation (or records an undo log). If a major defect is found during the full rollout, the authorized administrator can roll back to the state before cleaning with one click, restore the original data, and re-execute the grayscale process after readjusting the script.

[0094] In some embodiments, the data standardization processing system based on business requirements may include multiple functional modules composed of computer program segments. The computer programs for each program segment in the data standardization processing system based on business requirements may be stored in the memory of a computer device and executed by at least one processor to perform (see details). Figure 1 (Description) Functionality for data standardization processing based on business needs.

[0095] In this embodiment, the data standardization processing system based on business needs can be divided into multiple functional modules according to the functions it performs, such as... Figure 4 As shown. The module referred to in this invention is a series of computer program segments that can be executed by at least one processor and perform a fixed function, and is stored in memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0096] The change capture module is used to capture change events of the business system and extract the change events into structured change information; The impact calculation module is used to formalize the structured change information into change vectors, quantify the comprehensive impact of the change vectors on each standard item node on a pre-constructed association graph of business item nodes and standard item nodes, and determine the standard item to be changed based on the comprehensive impact. The draft generation module is used to route the standard item to be changed to the corresponding draft generator based on the attributes of the change vector using a decision tree, so that the draft generator generates a standard update draft containing specific rule parameters. The draft review module is used to push the standard update draft to the review end, and update the standard update draft that passes the review to the data standard library; The standard processing module is used to standardize incremental data based on the updated data standard library and perform adaptive cleaning and transformation on existing data.

[0097] Figure 5 The data standardization processing method based on business needs provided in the embodiments of this application can be applied to devices. Those skilled in the art will understand that the device structure involved in the embodiments of this invention does not constitute a limitation on the device. The device may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. Specifically, the device 500 may include: a processor 510, a memory 520, and a communication unit 530. These components communicate through one or more buses. Those skilled in the art will understand that the server structure shown in the figures does not constitute a limitation on the invention; it can be a bus topology or a star topology, and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.

[0098] The present invention also provides a computer medium, wherein the computer medium may store a program, which, when executed, may include some or all of the steps provided in the embodiments of the present invention. The medium may be a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.

[0099] The same or similar parts between the various embodiments in this specification can be referred to mutually. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple, and the relevant parts can be referred to the description in the method embodiments.

[0100] In the embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or modules may be electrical, mechanical, or other forms.

[0101] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0102] In addition, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0103] Although the present invention has been described in detail with reference to the accompanying drawings and preferred embodiments, the present invention is not limited thereto. Various equivalent modifications or substitutions can be made to the embodiments of the present invention by those skilled in the art without departing from the spirit and essence of the invention, and such modifications or substitutions should all be within the scope of the present invention. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should also be covered within the protection scope of the present invention.

Claims

1. A data standardization processing method based on business needs, characterized in that, include: Capture change events from the business system and extract the change events into structured change information; The structured change information is formalized into a change vector. The comprehensive impact of the change vector on each standard item node is quantitatively calculated on the pre-constructed association graph between business item nodes and standard item nodes. Based on the comprehensive impact, the standard item to be changed is determined. The decision tree is used to route the standard item to be changed to the corresponding draft generator according to the attributes of the change vector, so that the draft generator generates a standard update draft containing specific rule parameters; The draft standard update is pushed to the review panel, and the draft standard update that passes the review is updated to the data standard library. Based on the updated data standard library, incremental data is standardized, and existing data is adaptively cleaned and transformed.

2. The method according to claim 1, characterized in that, Capture change events from the business system and extract the change events into structured change information, including: Collect change signals, including: deploying listeners at the business entity change interface of the business system to capture business model change events in real time; interfacing with the requirements management platform to automatically parse change information in the approved requirements documents; embedding analysis modules in the data flow of the data access layer to detect the addition and change of fields or structures through pattern snapshot comparison, and to detect the addition and drift of enumeration values ​​or value ranges through historical value baseline comparison. The collected change signals are converted into structured change information in a unified format. The structured change information includes the following fields: change subject, change type, change content, change time, change source, and affected data range.

3. The method according to claim 1, characterized in that, The method for constructing the association map includes: The system constructs nodes, which include business item nodes, standard item nodes, and data field nodes. The business item nodes have a business concept identifier and a change frequency attribute; the standard item nodes have a standard identifier, a standard type, a standard clause, and version attributes; and the data field nodes have a field identifier, a data type, and a sensitivity attribute. Establish relationships between nodes, including: The "belongs to" relationship is used to connect the data field node to the business item node, indicating that the data field belongs to the corresponding business item. A constraint relationship is used to connect the data field node to the standard item node, indicating that the data field is constrained by the corresponding standard item clause. The constraint relationship has a weight attribute that indicates the strength of the constraint. The association relationship is used to connect the business item node to the standard item node, representing the association between the business item and the standard item at the macro level. The association relationship has a relevance attribute that represents semantic similarity or historical co-occurrence frequency.

4. The method according to claim 3, characterized in that, On a pre-constructed association graph of business item nodes and standard item nodes, the comprehensive impact of the change vector on each standard item node is quantitatively calculated, and the standard item to be changed is determined based on the comprehensive impact, including: Parse the affected business item identifier, the changed field identifier, the change operator type, the change type, and the change metadata in the change vector; For each standard item node, its comprehensive influence is calculated. The comprehensive influence is a weighted sum of multiple influence components, including: direct association influence, propagation association influence, and semantic similarity. The sum of the weight coefficients of the direct association influence, the propagation association influence, and the semantic similarity is 1. The method for determining the direct correlation influence is as follows: if the change field identifier in the change vector has a corresponding data field node in the correlation graph, then the weight of the constraint relationship between the data field node and the current standard item node is obtained, and the direct correlation influence is calculated in combination with the change operator type and change type of the change vector; otherwise, the direct correlation influence is 0. The method for determining the propagation association impact is as follows: using the affected business item identifier in the change vector to locate the corresponding business item node in the association graph, obtaining the correlation attribute of the association between the business item node and the current standard item node, as well as the shortest path hop count from the business item node to the current standard item node, and attenuating the correlation attribute according to the attenuation function corresponding to the hop count to obtain the propagation association impact. The semantic similarity is determined by calculating the semantic similarity between the change metadata in the change vector and the description text of the current standard item node. Nodes with a comprehensive impact greater than a preset threshold are identified as standard items to be changed.

5. The method according to claim 1, characterized in that, Using a decision tree, the criteria item to be changed is routed to the corresponding draft generator based on the attributes of the change vector, including: The standard item to be changed and its corresponding change vector are input into the decision tree, and the judgment conditions are matched layer by layer starting from the root node until the leaf node is reached. Each leaf node is bound to a draft generator, which contains rule templates for specific change scenarios, and the rule templates have dynamically populated parameter bits; Based on the leaf node reached, the corresponding draft generator is invoked; The internal nodes of the decision tree serve as judgment conditions, which are based on the attributes of the change vector. The attributes of the change vector include: change type, change operator type, and standard item type.

6. The method according to claim 5, characterized in that, The draft generator includes rule templates and dynamically populated parameter bits; the method for the draft generator to generate a standard update draft containing specific rule parameters includes: Extract the change content from the change vector and fill the parameter bits in the rule template. The parameter bits include at least: field name and enumeration value set. For the threshold parameter in the rule template, it is dynamically calculated based on the historical change frequency of the business items involved in the change vector and the sensitivity or importance of the fields involved. The base threshold is subtracted from the product of the historical change frequency and the first coefficient, and then the product of the sensitivity or importance and the second coefficient is added to obtain the final threshold. The filled rule template is assembled into a draft standard update for the standard item to be changed; A confidence score is calculated for the standard update draft, which is the product of the overall impact of the standard item to be changed and the historical accuracy of the draft generator. The confidence score is then appended to the standard update draft and output together.

7. The method according to claim 1, characterized in that, Based on the updated data standard library, incremental data is standardized, and existing data undergoes adaptive cleaning and transformation, including: In the data acquisition or processing process, the updated data standard library is loaded in real time, and data cleaning, format conversion, verification and standardized writing operations are performed on the newly generated incremental data based on the updated data standard library. Based on the changes in the draft update of the standard, a data migration script or a data cleaning script is automatically generated to process historical data in batches to make it conform to the updated data standard. The processing adopts a gray release strategy, first processing and verifying a portion of the existing data, and then promoting it to the full amount of existing data after confirming that there are no errors.

8. A data standardization processing system based on business needs, characterized in that, include: The change capture module is used to capture change events of the business system and extract the change events into structured change information; The impact calculation module is used to formalize the structured change information into change vectors, quantify the comprehensive impact of the change vectors on each standard item node on a pre-constructed association graph of business item nodes and standard item nodes, and determine the standard item to be changed based on the comprehensive impact. The draft generation module is used to route the standard item to be changed to the corresponding draft generator based on the attributes of the change vector using a decision tree, so that the draft generator generates a standard update draft containing specific rule parameters. The draft review module is used to push the standard update draft to the review end, and update the standard update draft that passes the review to the data standard library; The standard processing module is used to standardize incremental data based on the updated data standard library and perform adaptive cleaning and transformation on existing data.

9. A data standardization processing device based on business needs, characterized in that, include: Memory, used to store data standardization processing programs based on business needs; A processor, configured to implement the steps of the data standardization processing method based on business requirements as described in any one of claims 1-7 when executing the data standardization processing program based on business requirements.

10. A computer-readable medium storing a computer program, characterized in that, The readable medium stores a data standardization processing program based on business requirements, which, when executed by a processor, implements the steps of the data standardization processing method based on business requirements as described in any one of claims 1-7.