Digital enterprise data adaptive exchange and intelligent integration method
Through adaptive data agents and multi-agent technology, data redundancy and security issues in enterprise data integration are solved, and cross-department data is achieved quickly, securely and efficiently exchanged and synchronized, adapting to enterprise dynamic changes and ensuring data quality and security.
Patent Information
- Application Number
- CN202510792627.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional static data integration method cannot adapt to the dynamic changes of enterprises and the needs of multiple departments, resulting in repeated data storage and high redundancy, making it difficult to track data relationship updates between departments, and there are hidden dangers of security and permission control.
Through the adaptive data agent, the original data of each department is captured and standardized at the edge, cross-department field alignment and multi-agent collaborative verification are performed, routing decisions are determined based on multi-dimensional feature analysis, data is archived to the data lake and interface is opened to the outside world, forming a closed loop of data cleaning, alignment, permission control and automated routing.
It has achieved rapid adaptation to cross-departmental data changes, precise mapping of synonyms, flexible and efficient data synchronization and transmission path optimization, ensuring data security and compliance, and providing data collaboration that can traceability throughout the life cycle.
Smart Images

Figure CN120455562A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of enterprise informatization and data processing technology, and in particular to a method for adaptive exchange and intelligent integration of digital enterprise data. Background Art
[0002] In the era of digital transformation, data exchange needs within and across enterprises are becoming increasingly complex. Traditional static data integration methods struggle to adapt to the dynamic changes and multi-departmental collaboration demands of enterprises. Adaptive data exchange and intelligent integration within digital enterprises not only enable rapid sharing of unified, high-quality data across departments, but also maintain system flexibility and scalability amidst frequent business structure changes, thereby improving overall operational efficiency and decision-making accuracy.
[0003] Existing technologies often adopt a single data warehouse for centralized management or a relatively fixed cross-departmental data call method (Chinese invention patent, publication number: CN118689882B, title: A method and system for integrating business data based on multiple departments), resulting in the following deficiencies: First, data is stored repeatedly and has high redundancy, making it impossible to efficiently identify the boundaries between cross-departmental shared data and departmental private data; second, it is difficult to track updates to data relationships between departments in a timely manner, which can easily lead to data delays or lags; third, centralized storage may have hidden dangers in security or authority control, and department-specific information can be easily misused or accessed without authorization. Summary of the Invention
[0004] To address the numerous issues with the aforementioned existing technologies, the present invention provides a method for adaptive exchange and intelligent integration of digital enterprise data. This method uses adaptive data agents to capture and standardize raw data from various departments at the edge. After cross-departmental field alignment and multi-agent collaborative verification, compliant corrected data is generated. Multi-dimensional feature analysis is then used to determine routing decisions. The final synchronized data is archived to a data lake and exposed through an external interface. This method forms a closed loop of data cleansing, alignment, permission management, and automated routing, significantly improving the efficiency and security of data exchange between multiple departments within an enterprise and with external partners.
[0005] A method for adaptive exchange and intelligent integration of digital enterprise data, comprising the following steps: Obtaining original business data from each business department and performing field mapping and format unification processing to form first department pre-processed data, and publishing the first department pre-processed data to a message bus; Performing a cross-departmental field alignment operation on the first department pre-processed data, and merging the alignment information with the first department pre-processed data to generate second department pre-processed data; wherein the alignment operation compares and confirms the fields based on a self-organizing multi-agent approach; Performing anomaly detection and data correction on the pre-processed data of the second department to generate compliant corrected data; Perform multi-dimensional feature analysis and adaptive routing decisions on the compliance correction data, distribute the results to various departments to form multi-department synchronized data, archive the multi-department synchronized data in the enterprise data lake and provide an external access interface.
[0006] Preferably, after each business department obtains the original business data, a unique field identifier is assigned to each field, and a data type verification script is used to check the numerical range and text length of the field. Field values that do not meet the inspection standards are corrected according to predefined mapping rules and together with the field values that meet the inspection standards, they constitute the first department preprocessing data.
[0007] Preferably, when performing a cross-departmental field alignment operation on the pre-processed data of the first department, the field definitions of multiple departments are compared in parallel based on the self-organizing multi-agent to generate a matching score list for each field candidate object, and the matching score list is transmitted to the aggregation node; the aggregation node performs statistics and judgment on the matching score list according to a preset matching threshold; if all candidate scores are lower than the matching threshold, the alignment is determined to have failed and manual verification is triggered; if there is a score greater than the matching threshold, it is regarded as a preliminary alignment and the matching score and candidate field are recorded in the alignment information; When the number of field definitions is greater than or equal to 2 and the matching scores all exceed the preset matching threshold, the self-organizing multi-agent compares the usage frequency of the fields within the department in descending order according to the scores, prioritizes the most frequently used fields, and writes the confirmation order and department usage records in the alignment information; After completing the above alignment, the alignment information is stored in the alignment management library in the form of a mapping table. The mapping table contains the field name, corresponding department identifier and matching score. When generating the second department preprocessing data, the mapping table entries are synchronously updated and timestamps are marked, so that subsequent steps can perform anomaly detection or data correction processing based on the latest field correspondence.
[0008] Preferably, when performing anomaly detection and data correction processing on the pre-processed data of the second department, the data range and logical verification conditions are configured for each field according to the preset business rules. For field records whose data range exceeds the preset conditions or there are logical conflicts, the correction script is called to modify the field value and generate an anomaly identification entry, which includes the field name, original value, revised value and verification conditions.
[0009] Preferably, for field records that cannot be corrected through the correction script, a manual review process is initiated after the correction script returns an exception. The reviewer confirms the correction value in the field comparison list and records the field values before and after the review in the correction history table. The correction history table contains the field name, reviewer ID and the final effective value.
[0010] Preferably, when performing multi-dimensional feature analysis and adaptive routing decisions on the compliance correction data, department identification, field sensitivity level and number of data corrections are used as basic features, and a feature vector is constructed in combination with network load information. The feature vector is aggregated through weighted calculation, and the aggregation result is used to determine the distribution order and security policy, and a routing label is added to each data during distribution.
[0011] Preferably, based on the compliance-corrected data whose sensitivity level is higher than the preset value shown in the aggregation result, the data is distributed to the target department after performing a security encryption operation separately; for the compliance-corrected data whose sensitivity level is lower than the preset value, it is packaged with similar data into multi-department synchronization data, and a routing log is generated after the distribution is completed; the routing log contains the routing label, the target department and the encryption status.
[0012] Preferably, after forming multi-department synchronized data, the multi-department synchronized data is archived to the enterprise data lake and the external access interface is configured, and access permission levels and audit modes are set for each external access connection, and the visitor identity and access time are identified through access logs. When a request that does not conform to the registered permission level is detected, an interception strategy is adopted and the interception information is recorded.
[0013] Compared with the prior art, the advantages and beneficial effects of the present invention are: The present invention uses self-organizing multi-agent field alignment and hierarchical data error correction technology to quickly adapt to cross-departmental data changes and accurately map synonymous fields; the present invention uses multi-dimensional feature analysis and adaptive rule engine technology to achieve flexible and efficient data synchronization and transmission path optimization in complex network environments; the present invention uses separate encryption and permission management technology for sensitive data to ensure data security and compliance in multi-department collaboration; the present invention uses data lake archiving and external access auditing technology to achieve traceability of data throughout its life cycle and efficient external collaboration. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 Schematic diagram of the process of the present invention; Figure 2 Schematic diagram of interactive data relationship in the present invention. DETAILED DESCRIPTION
[0015] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure.
[0016] like Figure 1 As shown, a method for adaptive exchange and intelligent integration of digital enterprise data includes the following steps: Obtaining original business data from each business department and performing field mapping and format unification processing to form first department pre-processed data, and publishing the first department pre-processed data to a message bus; In a digital enterprise environment, business departments often have distinct data structures and formats. To ensure a consistent foundation for subsequent alignment, anomaly detection, and intelligent routing, this invention captures raw business data from each department and standardizes it into "first-department pre-processed data," which is then shared across departments using a message bus.
[0017] There are significant differences in the naming and arrangement of data fields in different departments. For subsequent unified management and processing, the original fields need to be converted into a predefined field set. In the present invention, the predefined field set includes a unique identifier and a field type description, which is intended to ensure that the data field names uploaded by each department are consistent and comply with the basic type specifications. If the field names provided by the department deviate significantly, the character similarity between the field name and the predefined set can be calculated. An exemplary similarity can be given by the following formula:
[0018] in Indicates the department field name; Indicates a predefined field name; and are the string lengths, is the indicator function, that is, ), which takes 1 if the characters match and 0 if they do not match.
[0019] When the similarity reaches a threshold, the mapping rule establishes a one-to-one correspondence between the department field name and the predefined field name. This way, field mapping automates most of the name unification process, requiring manual intervention only in the event of duplication or conflict.
[0020] The data formats reported by various departments often differ in types, such as date notation, numerical precision, and currency units. To ensure the consistent data verification process is applied later, format conversion is performed in this step, including standardizing date notation and numerical precision.
[0021] Dates are written in YYYY-MM-DD HH:mm:ss format. Numeric fields retain decimal places according to the preset precision. If there are any amount or other unit fields, they are all based on the specified unit of measurement.
[0022] Once field mapping and format unification are complete, "first department preprocessed data" is generated. This data is published externally via a message bus, allowing subsequent functional modules such as alignment, anomaly detection, and routing decisions to subscribe and perform subsequent processing. The message bus acts as a data exchange hub, connecting all subsequent operations to a single entry point and avoiding the need for excessive custom connections between departments or functional modules.
[0023] For example, within a company, Department A names the "Employee Number" field StaffID, while Department B names the same concept EmpCode. Using the string similarity formula Sim, if a small number of characters in both StaffID and EmpCode match, and the similarity exceeds the threshold of 0.5, the mapping system assigns both StaffID and EmpCode to the predefined "employee_id" field.
[0024] During this process, the string matching function performs cumulative comparisons at the Latin character level. When special characters appear, they can be recorded and manually repaired during alignment or error correction.
[0025] After mapping the data for Department A and Department B simultaneously, the "First Department Preprocessed Data" output by both will contain a field named employee_id, with time notation in the format YYYY-MM-DD HH:mm:ss. Subsequent anomaly detection can uniformly handle invalid employee_id values, thus avoiding conflicts caused by multiple field names.
[0026] Preferably, after each business department obtains the original business data, a unique field identifier is assigned to each field, and a data type verification script is used to check the numerical range and text length of the field. Field values that do not meet the inspection standards are corrected according to predefined mapping rules and together with the field values that meet the inspection standards, they constitute the first department preprocessing data.
[0027] When a digital enterprise is operating, each department will independently generate a large amount of original business data. Because of the different departments and business scenarios, these original data have large differences in field naming, data type, text length, etc. If such scattered and heterogeneous original business data are directly passed to the next step of processing, it will cause field conflicts and abnormal data formats. To solve this problem, after each business department obtains the original business data, the present invention first assigns a unique field identifier and uses a data type verification script to check the numerical range and text length of the field value, so as to ensure the compliance and uniformity of the "first department pre-processing data" at the source stage.
[0028] Raw business data is collected from different departments, and fields may have duplicate or inconsistent names. To provide a basis for subsequent cross-departmental alignment and external release, each field needs to be uniquely identified.
[0029] This invention stores a "unique field identifier" in the field dictionary and associates it with the actual field name. This approach allows the field's corresponding business meaning or definition source to be retrieved later through the field identifier. After assigning a unique field identifier, the field's value range and text length must be checked to ensure there are no violations or abnormal values.
[0030] If a field value for a department exceeds the specified range, it will be considered non-compliant. Similarly, if a text field's length exceeds the upper limit, it will also be considered non-compliant. This mechanism allows for correcting obviously out-of-bounds or illegal data at the source.
[0031] The data type verification script is a relatively mature automated detection method. This invention only emphasizes its specific implementation logic in a multi-department environment, that is, combining the "unique field identifier" to distinguish the importance and range restrictions of the field.
[0032] For field values that do not meet the standards, the present invention defines a "predefined mapping rule" to correct them. The rule can select different mapping methods according to the field attributes (numeric or text type).
[0033] In numerical correction, the minimum and maximum values can be set as and , when the field value Beyond When the interval is set, it is mapped to the allowed interval according to the rules, such as:
[0034] In the above formula, Represents the original value; Indicates the corrected value; Indicates the lower limit of a numeric field; Indicates the upper limit of a numeric field. When modifying text, you can use trimming or padding to truncate overlong text to a predefined length, or fill short fields with placeholders. After the modification process is complete, "first-party preprocessed data" is generated. This data includes all normal field values, as well as replaced or modified field values. Subsequent steps will identify modification markers and perform further verification (such as alignment or error correction).
[0035] By assigning unique field identifiers and performing value range and text length checks after obtaining the original business data from each business department, the spread of non-compliant data can be blocked as early as possible, reducing the burden of subsequent alignment or error correction stages. Invalid fields or out-of-bounds fields will be processed promptly and will not enter the next system release, nor will they conflict with data from other departments. The unique field identifier provides a reliable reference for subsequent processes when identifying field intent, without having to rely on departmental custom naming. The correction mechanism ensures that values and text can be controlled within a specified range, reducing anomalies caused by incompatible field formats during cross-departmental connections. In the subsequent cross-departmental alignment and anomaly detection stages, the data from each department has been quality-assured under the same inspection standards; the field correction records in the "first department pre-processed data" are retained before release and can be associated with the unique field identifier to help subsequent automated processing modules (such as anomaly detection or semantic alignment) achieve accurate analysis.
[0036] Performing a cross-department field alignment operation on the first department pre-processed data, wherein the alignment operation compares and confirms the fields based on a self-organizing multi-agent approach, and merging the alignment information with the first department pre-processed data to generate second department pre-processed data; In the process of digital enterprise data flow, the meaning or naming of fields often differs across departments, resulting in the same business entity displaying different field names or data structures in different departments. If only conventional mapping rules are used, it is difficult to adapt to organizational changes and the dynamic flow of large amounts of cross-departmental data. The present invention uses a self-organizing multi-agent method to align the cross-departmental fields of "first department preprocessed data" to generate unified and adaptable "second department preprocessed data."
[0037] Each department may have an independent field definition library. For example, the "employee number" field name and data requirements may be different in the Finance Department and the Human Resources Department. Traditional methods require manual writing of mapping rules. Once a department adds or deletes a field, the configuration needs to be repeatedly modified, which lacks adaptability. This invention uses a self-organizing multi-agent approach to automatically align fields, reducing maintenance costs.
[0038] In this paper, multi-agent refers to a group of independent agents, each responsible for monitoring and recording the definition and usage information of a specific department or category of fields. When field alignment is required, each agent compares the input field name and type information with the local field definition library and calculates the degree of match. Self-organization means that these agents actively exchange comparison results during the alignment process. If multiple matching conflicts or no suitable match occurs, the agents coordinate decisions or trigger human intervention.
[0039] The final alignment information includes field mapping relationships, such as "Department A Field X" corresponds to "Department B Field Y". This result can be updated over time and queried by other modules of the system.
[0040] In practical applications, the first department's preprocessed data contains fields in a uniform format and their preliminary cleansing and correction values. Through multi-agent comparison and aggregation, an alignment table is generated, listing which fields have the same business meaning or conflict. After merging the alignment table back into the first department's preprocessed data, the "second department's preprocessed data" is formed, which is convenient for use by subsequent modules (such as anomaly detection and intelligent routing).
[0041] Preferably, when performing a cross-departmental field alignment operation on the pre-processed data of the first department, the field definitions of multiple departments are compared in parallel based on self-organizing multi-agents to generate a matching score list for each field candidate object, and the matching score list is passed to the summary node; the summary node counts and judges the matching score list according to a preset matching threshold. If all candidate scores are lower than the matching threshold, the alignment is determined to have failed and manual verification is triggered. If there is a score greater than the matching threshold, it is regarded as a preliminary alignment and the matching score and candidate field are recorded in the alignment information.
[0042] When different departments within an enterprise merge or exchange data, it's common for the same business concept to be represented by different field names or types. For example, "OrderAmount," "SalesMoney," and "Revenue" actually have the same field meaning. This can cause problems for data integration and subsequent analysis.
[0043] To this end, the present invention designs a cross-department field alignment operation so that each department field can be mapped to a common logical name or structure, providing a unified processing basis.
[0044] In this invention, multiple Agents refer to a group of relatively independent software agents, and each Agent has professional knowledge of the definition of a certain department or a certain type of field. When performing alignment, these Agents compare the field information in the external "preprocessed data of the first department" according to the "field definition set" they have mastered. To avoid processing delays caused by the sequential queuing of each Agent, this invention adopts a parallel mechanism, that is, each Agent retrieves its own field definition library and generates a matching score simultaneously, and then passes the result to the summary node for unified judgment.
[0045] The summary node is a central role that aggregates the comparison results of each Agent, and determines whether a certain field can be successfully aligned according to a preset matching threshold. If the matching score ≥ the threshold, it is regarded as successful alignment; if multiple fields reach the threshold simultaneously or no field reaches the threshold, the summary node triggers a manual verification process to prevent incorrect mapping or data loss.
[0046] This invention will record in detail in the alignment information: which fields are matched by which Agents; what the specific matching scores are; if there is a final manual check, the check result will be attached; this record can be used as a reference for subsequent exception correction or routing decision-making to avoid repeated comparison.
[0047] In practical applications, each Agent maintains a department field definition library in the system, including field names, field types, historical usage scenarios, etc.; after receiving the "preprocessed data of the first department", the Agent reads the field names, field types and other possible identification information therein, and generates a matching score using a similarity measure. In this invention, the matching score can synthesize factors such as field name similarity, department preference, historical field usage frequency, etc. to form a comprehensive score. The scoring results of all Agents are submitted concurrently to the summary node without waiting for other Agents to complete.
[0048] The summary node will aggregate the matching scores passed by each Agent for each field, such as: Agent A 、Agent B 、Agent C etc. all give matching scores, and take the highest value or the average value to compare with the "threshold". If the highest matching score ≥ the preset threshold T, it is determined that the corresponding field is successfully aligned; if there are multiple scores ≥ T, it means that there is an alignment conflict and manual confirmation needs to be called. If all matching scores < T, it is determined that the alignment fails and automatically enters the manual verification process (for example, display the recommended candidates and scores for this field on the management platform, and manually confirm the attribution).
[0049] Alignment information typically includes: the original field name (and the pre-processed data record of the first department to which it belongs); the corresponding target field name (the unified name recognized by the agent or the field definition of other departments); a list of matching scores (scored by each agent); the final confirmation method (automatic or manual); and possible timestamps and version numbers for subsequent traceability.
[0050] In parallel comparison mode, there's no need for manual comparison of field definitions against the field definition library. Each agent automatically detects and scores fields, effectively reducing data distortion caused by field conflicts and name inconsistencies. If an enterprise adds a new department or field, the agent system automatically learns and updates the field definitions, eliminating the need to modify core business processes. Through self-organizing agent scoring and threshold judgment at aggregation nodes, potential one-to-many or many-to-one conflicts are promptly detected and manually verified, ensuring the correct mapping of critical business fields. If synonymous fields have inconsistent meanings (e.g., same name but different business semantics), different agents will independently score them, causing the conflict score to fall below the threshold, requiring manual intervention to prevent misalignment. After field alignment is complete, the system generates an "alignment information" that is merged back into the "first department pre-processed data" to become the "second department pre-processed data." Subsequent anomaly detection and adaptive routing only require processing based on the unified field name, significantly simplifying cross-departmental collaboration.
[0051] In actual implementation, enterprises deploy 、 After the aggregation agent, whenever new "first department pre-processed data" enters the platform, and Scan the field names separately and calculate .
[0052] For example , , then submit the results to the summary agent After aggregation, the field is considered to have a certain degree of similarity. If it exceeds 0.80, it is automatically set as a match. Otherwise, it enters the manual verification phase, and the alignment information is updated after the final judgment is made by humans. After a successful alignment, the field name and matching score are loaded into the alignment information and merged into the pre-processed data of the second department. When completing subsequent algorithms and business operations, the aligned name will be treated as a unified field "transport_cost" or other unified name.
[0053] When the number of field definitions is greater than or equal to 2 and the matching scores exceed the preset matching threshold, the self-organizing multi-agent compares the frequency of use of the fields within the department in order from high to low according to the scores, gives priority to confirming the fields with the highest frequency of use, and writes the confirmation order and department usage records in the alignment information.
[0054] When the present invention compares the "first department pre-processed data" with the field definition libraries of multiple departments in a self-organizing multi-agent manner, it often encounters a field that can match more than one candidate object. If the matching scores of all the candidates meet or exceed the preset matching threshold, the system will select a final mapping object from these high-scoring candidates. If the highest-scoring candidate is selected without distinction, the frequency of cross-departmental field usage in actual business operations may be ignored, leading to inconsistencies in subsequent data integration or issues that may not conform to departmental practices.
[0055] This invention introduces "frequency of use within a department" to assist in the matching score. Frequency of use refers to the number of times each field appears or is referenced within the department, the scope of coverage, etc. The higher the frequency, the greater the department's reliance on the field. By combining the score and frequency of use, the system can more reasonably determine matching fields: If two or more candidate fields have the same score or both exceed the threshold, they are compared in descending order based on frequency of use. The most frequently used field is confirmed first, indicating that this field occupies a more important position in the department's business, and aligning with it can reduce subsequent conflicts.
[0056] Since multiple candidate fields can ultimately only determine one optimal mapping target, the solution records the following in the alignment information: "Confirmation Order" - indicating which one was selected for priority confirmation when sorting multiple candidates; "Department Usage Record" - used to retain the number of times the field is cited or how often it is used in each department, to facilitate subsequent audits and modifications.
[0057] In practice, after receiving preprocessed data from the first department, each agent calculates its match score against the known field definition. If multiple candidates for the same field have matching scores above a threshold, a "multi-candidate list" is formed. After discovering multiple candidate lists, the self-organizing multi-agent system notifies the aggregation node to sort them by "frequency of use." Here, each agent can provide the aggregation node with information on the number of departmental references to the candidate field over a recent period or the types of business scenarios it covers.
[0058] Frequency of use can be measured by simple counting, such as the sum of the number of times a field appears in a form, report, or interface call within a period of time. All satisfy the matching score ≥ threshold T, and sort this set from high to low according to the frequency of use: The rest of the fields are arranged in the same way and are confirmed first. .
[0059] At the same time, write "Confirmation Order = 1: C_{1}" in the alignment information. If the scores or frequencies of other candidates are only slightly lower, you can also record "Backup Candidate = 2: C_{2}" for manual reference or error correction when needed.
[0060] After completing this selection, the present invention writes the "confirmation order" and "department usage frequency" into the alignment information. This allows the system to quickly review past records if the same or similar field names appear next time, reducing duplication of work. If a user later discovers that the preferred field was not the most appropriate, the system can review the sorting process and make manual corrections or adjust the threshold.
[0061] If based solely on matching scores, multiple identical scores may appear; frequency of use can be used to balance business weights. This ensures that in complex cross-departmental environments, frequently used fields are given priority for accurate mapping, reducing potential conflicts in key data during subsequent sharing. Once multiple candidate fields have the same or slightly different scores, "common fields" and "rarely used fields" can be distinguished by comparing their frequency of use, avoiding positioning rarely used or temporary fields in the department as primary alignment targets. Enterprise business often evolves, and the frequency of use of a certain field will also change over time. Self-organizing multi-agents can automatically update the alignment order after discovering a significant change in the frequency order, which helps to cope with dynamic scenarios.
[0062] In one embodiment, a company has Department A and Department B, which respectively name the "Customer Payment" field "Payment" and "PayOrder." In the self-organized multi-agent comparison, the field comparison score is 0.88 (greater than the threshold of 0.80), forming multiple candidates.
[0063] The query found that the number of calls to "Payment" in the department A system reached 500 in one month. We also found that "PayOrder" was called 490 times in Department B. Both of these are relatively common; If the priority cannot be determined based solely on the score, the system will make a final comparison based on the specific frequency. If Department A has a higher frequency, "Payment" will be selected as the alignment object, and "Confirmation Order = 1" will be recorded in the alignment information, and "PayOrder" will be marked as "Confirmation Order = 2".
[0064] Step 1: After multiple agents compare "Payment" and "PayOrder" in parallel, the score is 0.88; Step 2: The summary node detection score = 0.88 ≥ threshold 0.80, and the number of candidates ≥ 2; Step 3: and Reporting frequency, assuming Step 4: Sort by frequency (600 > 580), prioritize "Payment," and write the alignment information as "confirm_order = 1 (Payment), 2 (PayOrder)." Step 5: Incorporate this alignment information to generate a new field record. In the pre-processed data for the second department, use "Payment" as the primary field name. If a special case is discovered later requiring a change, manual intervention can be performed to modify the order.
[0065] After completing the above alignment, the alignment information is stored in the alignment management library in the form of a mapping table. The mapping table contains the field name, corresponding department identifier and matching score. When generating the second department preprocessing data, the mapping table entries are synchronously updated and timestamps are marked, so that subsequent steps can perform anomaly detection or data correction processing based on the latest field correspondence.
[0066] After cross-departmental data alignment is complete, the system needs to store and manage the alignment results and ensure that the alignment relationship can be promptly referenced or updated when new "second department pre-processed data" is generated. To this end, the present invention proposes maintaining a "mapping table" and placing it in the alignment management library so that the same field correspondence can be applied in subsequent anomaly detection and data correction stages.
[0067] When cross-departmental fields are compared across multiple agents, the system generates alignment information consisting of the field name, corresponding department ID, and matching score. If this information is only temporarily stored in memory, it cannot be queried or corrected in subsequent steps. Furthermore, if the business changes (such as adding or removing fields by a department) or the system is restarted, the original information will be lost. Therefore, by writing this alignment information into the alignment management library in the form of a "mapping table," it can be persisted and read by other subsystems.
[0068] A mapping table typically contains: field names (such as "OrderID," "EmpCode," etc.); corresponding department identifiers (uniquely identifying departments, such as Finance or Operations); matching scores (generated by multi-agent calculations); and timestamps (reflecting the time the entry was last updated). Combining this mapping table with the "second department preprocessing data" of this invention allows downstream processes (such as the anomaly detection module) to know which department fields each field ultimately corresponds to, and to trace information such as matching scores and last modification times.
[0069] When generating or updating the "second department pre-processing data", the present invention emphasizes marking the timestamp in the mapping table entry; the timestamp can be generated after each field alignment (such as yyyy-MM-dd HH:mm:ss) or marked with a more precise time unit; if there are field changes in other departments subsequently, the system will run the alignment operation again and update the records in the corresponding mapping table, and write the new entry "version" accordingly to assist downstream links in obtaining the latest field correspondence.
[0070] In actual applications, after the cross-departmental field alignment phase is completed, the multi-agent system submits the final alignment results to the alignment management module. The alignment management module compiles mapping table entries based on each result, including: Source field name (the field name in the preprocessed data of the first department); target field name or unified identifier (if already specified); corresponding department identifier (indicating the current ownership or source of the field); matching score (Agent score or manual secondary verification result); timestamp (used to record the latest update time).
[0071] After the entry generation is completed, the mapping table is written into the alignment management library and waits for synchronous call when the "second department preprocessing data" is generated in the next step.
[0072] Whenever new alignment information is written into the mapping table and "second department pre-processed data" is generated, the mapping table must be read synchronously to refresh the fields that need to be updated; if new matches with higher scores are detected for certain fields, they will be written into the second department pre-processed data and the timestamp will be updated; if the mapping of the field is manually confirmed to be changed, the old record will also be overwritten in the mapping table.
[0073] When alignment information needs to be read during anomaly detection or data correction, the most reliable field correspondence can be found by simply querying the latest timestamp record in the mapping table of the alignment management library. If anomaly detection finds that field names conflict in multiple departments, the true field ownership or possible cause of the conflict can be found by comparing the mapping table.
[0074] When the same field appears in subsequent steps, its department affiliation and score can be directly checked from the mapping table, without the need for multiple agents to repeat calculations or frequent human intervention. This reduces the need to "identify from scratch" cross-departmental data information and improves overall processing efficiency and consistency. "Version control" of mapping entries can be achieved through the timestamp mechanism; when the department field definition changes (such as renaming or expanding fields), the entry for the new alignment operation will have a later timestamp, thus replacing the old version. When any problem arises (such as incorrect field mapping is found during anomaly detection), technicians can track the versions of the records in the mapping table to identify the mapping evolution process. Downstream anomaly detection often requires precise knowledge of whether a field corresponds to other department fields and the degree of matching; the data correction link needs to confirm the true meaning behind the field in order to determine the correction standard (such as date format, numerical range). By accessing the mapping table, all key information can be quickly queried.
[0075] like Figure 2 As shown, anomaly detection and data correction processing are performed on the pre-processed data of the second department to generate compliant corrected data; Even after aligning and integrating the fields across departments, data may still contain anomalies, such as out-of-bounds values, incorrect time formats, and logical conflicts (e.g., mismatches between the finance department's account balance and the warehouse department's inventory). If these anomalies aren't detected and corrected promptly, subsequent multi-dimensional analysis and adaptive routing will receive erroneous information, leading to failure of the overall decision-making or exchange process. Therefore, this invention, based on the "second department's pre-processed data," performs anomaly detection and data correction to produce more reliable "compliant and corrected data."
[0076] By pre-setting reasonable ranges for numeric fields or acceptable lengths for text fields, data that exceeds or falls below the range can be flagged or verified. Furthermore, cross-departmental alignment of field meanings can be used to determine whether they conform to common sense within the current enterprise environment.
[0077] In the present invention, the aligned associated information from various departments can be used, such as the consistency rule of "employee number" and "salary amount": if an extreme situation occurs where the salary of the same employee doubles in a short period of time, it can be considered an anomaly.
[0078] If an enterprise already has a rule library or simple statistical model for identifying outliers (such as those based on standard deviation, quantiles, etc.), it can combine the field definitions in the present invention to perform association analysis and automatically mark suspected anomalies.
[0079] For identified common anomalies, established mapping or repair rules can be used. If a value exceeds an upper limit, it can be clipped to the upper limit; if a text field does not match a dictionary, an attempt is made to replace it with the nearest correct value.
[0080] If a serious logical conflict is detected (such as a mismatch between an employee's ID and a department ID) or automatic correction cannot be confirmed, the system will mark the field for manual approval and present the relevant field information and cross-department mapping relationships to the reviewer.
[0081] In the present invention, "compliant corrected data" refers to the final data form that has been automatically or manually corrected and can be used in subsequent intelligent routing decisions and synchronization links, and is accompanied by correction traces and reasons.
[0082] In actual applications, the system first obtains the current values and related attributes (such as data types, logical rules, etc.) of all fields from the aligned dataset.
[0083] Performing anomaly detection includes: Out-of-bounds detection: If the field is numeric, compare the field value and (Preset upper and lower limits) to determine whether or .
[0084] Text detection: If the text length or character set exceeds the limit, it will be recorded as an exception.
[0085] Logical conflict detection: Combine alignment information to check whether multiple fields violate internal rules.
[0086] Call the correction script based on the data type or field location, for example, if the upper limit is exceeded, value:
[0087] in Indicates the original value; Indicates the corrected value; and These are the lower and upper limits of the field respectively; if the correction value cannot be determined, the "pending manual approval" label will be recorded.
[0088] All correction results are consolidated, and any fields that still require manual processing are marked in the output data. Once all remaining fields are corrected, the final "compliant corrected data" is generated, providing reliable input for the next step of multi-dimensional analysis and routing.
[0089] After automatic corrections are made, the system maintains a log indicating the original value, revised value, and the rules or script used. If a human review is involved, the system records the reviewer, time, and final decision. This process can be subsequently audited to verify the appropriateness of the changes.
[0090] After anomaly detection and correction, the "compliant corrected data" that enters subsequent multi-dimensional analysis and adaptive routing decisions can basically eliminate illogical or out-of-bounds error values, laying a solid foundation for corporate decision-making and data exchange. Once this step is normalized in the system, new data entering the platform will first undergo automatic detection and correction, and potential anomalies will be handled at the front end to avoid chain errors when multiple departments are shared or merged later. Since the present invention already has unified field alignment, coupled with the logical conflict detection in this step, if the fields of the finance department and the operations department are consistent but there are abnormal differences, they can be discovered and corrected in time, thereby improving cross-departmental collaboration efficiency and data consistency.
[0091] Preferably, when performing anomaly detection and data correction processing on the pre-processed data of the second department, the data range and logical verification conditions are configured for each field according to the preset business rules. For field records whose data range exceeds the preset conditions or there are logical conflicts, the correction script is called to modify the field value and generate an anomaly identification entry. The anomaly identification entry includes the field name, original value, revised value and verification conditions.
[0092] In the daily operations of digital enterprises, different fields have different business meanings and value limits. To achieve accurate processing during automatic detection and correction, it is necessary to pre-configure business rules for each field in the system. For example: Value range: If a field represents "order amount", it is possible to set a range , when the field value exceeds this range, it is considered abnormal.
[0093] Logical validation conditions: If there is an intrinsic relationship between field A and field B (for example, the inventory number must be greater than or equal to the shipment quantity), you can write validation conditions and mark exceptions if logical conflicts such as negative inventory occur.
[0094] With the above configuration, after receiving the "second department pre-processed data", the system will automatically traverse each field and determine whether it exceeds the numerical range or violates the logical constraints. If an anomaly is found, the system calls the corresponding correction script to perform a repair operation on the field value; the repair method may be to truncate it between the upper and lower limits, or to directly overwrite it with the default value. For logically conflicting fields, corrections can be completed based on pre-defined processing rules (such as taking the average of two fields, or setting it to prioritize the upstream department). If the conflict cannot be resolved automatically, the system retains the mark for manual review.
[0095] When the correction script completes the repair, the system generates a complete exception record for the record and stores it in the exception record table: Field Name: the logical name of the corrected field; Original Value: the abnormal value or status detected before the repair; Corrected Value: the new value after automatic or semi-automatic correction; Verification Condition: the business rule that triggered the exception, such as "Amount < 0" or "Inventory < Shipment Volume." This information can be used for future traceability and auditing, and also provides sufficient basis for manual review.
[0096] In practical applications, the present invention typically stores preset business rules (data ranges and logical validation conditions) in a rule management table. Each rule corresponds to a field or field group, along with upper and lower limits or logical constraints. When processing "Second Department Preprocessed Data," the system matches each record against the rule management table. If the field name or field type meets specific rules, an out-of-bounds or logical conflict determination is performed. Upon detecting an anomaly, the system automatically initiates a correction script based on the current field's correction method. The correction script can perform operations such as simple overwrites and range-limiting operations as designed.
[0097] When the system completes a field value correction, it also writes it to the exception log table, which can be abstracted to the outside world as an "exception identification entry." This entry includes: field name (e.g., OrderAmount); original value (e.g., -50); revised value (e.g., 0); and verification condition (e.g., "Amount ≥ 0"). This system also typically includes a timestamp and department ID to facilitate cross-departmental audits and provide traceability to the next step.
[0098] If two fields depend on the same business logic (such as shipment quantity and inventory quantity) and there is a contradiction, the system can execute a script: and the Shipment Quantity field like , it is marked as abnormal; the script can choose to or give Supplement the difference (the specific rules need to be formulated by the enterprise). Also, when generating abnormal identification entries, record the corresponding verification conditions " or".
[0099] This step prevents data with out-of-bounds or logical conflicts from continuing to the intelligent routing and multidimensional analysis stages, reducing subsequent decision-making errors caused by data anomalies. The anomaly identification entry retains the original value, revised value, and trigger conditions for each correction process. Managers can later review the specific issues and correction reasons, improving system transparency and auditability. Because the rule base can be expanded according to the evolution of the enterprise business, this invention can adapt to changing field definitions and logical requirements, achieving accurate data correction at different stages.
[0100] Preferably, for field records that cannot be corrected through the correction script, a manual review process is initiated after the correction script returns an exception. The reviewer confirms the correction value in the field comparison list and records the field values before and after the review in the correction history table. The correction history table contains the field name, reviewer ID and the final effective value.
[0101] In this invention, the anomaly detection and automatic correction script can handle a large number of out-of-bounds values and simple logical inconsistencies. However, in some complex or unique scenarios, the script may not be able to accurately correct them. For example, a field value may have multiple alternative correction options that cannot be automatically matched; or different departments may have conflicting weights or interpretations of a field, requiring the judgment of business experts. If the correction script returns an anomaly that cannot be corrected, it indicates that the field is in a situation that cannot be handled by automated means, and manual review is required to ensure data quality.
[0102] When the system detects that the correction script returns an exception, it means that the correction operation of the field has failed to give a definite result; then, in the business process specified by the present invention, the system triggers the "manual review" link and displays the field information (including exception type, historical value, context field) to the reviewer for processing.
[0103] The field comparison list is usually an interface or list that lists all field records that cannot be automatically corrected; the reviewer selects the final value to be adopted for the field based on the department's business knowledge or experience (it may also need to be judged in conjunction with other related fields); through this manual intervention method, the system can resolve complex conflicts and ensure that the data results are close to the actual business needs.
[0104] After the manual review is complete, the system writes the original value before the correction and the final effective value designated by the reviewer into a correction history table. This history table contains information such as the field name, reviewer identifier (in the form of a unique ID or work number), and the final effective value, which is used for subsequent audits or rollbacks. This design meets the requirements of digital enterprises for data transparency and traceability, making it easier for management or technical teams to trace the decision-making process when anomalies are discovered.
[0105] In actual applications, a series of rules are usually defined in the correction script. If the field value and the rules cannot match a unique correction path, or multiple conflict scenarios occur, the script will return to the "unable to correct" status. After the system detects this return, it immediately records the field ID and failure reason and calls the manual review process interface.
[0106] The reviewer can view the original value of the field, other matching field information (if there is a cross-departmental association), a summary of the preset rules, etc. on the management platform (or other visual interface) provided by the present invention; the reviewer fills in the final revised value through the drop-down option or text input, and submits it for confirmation.
[0107] The system writes the "original value" and the "value specified by the auditor" into the correction history table at the same time: Field name: indicates the field currently being audited and revised; Auditor ID: a unique identifier such as employee ID = 1002; Final effective value: used to replace the old value of the field in the compliance correction data; Timestamp or version information (optional).
[0108] This mechanism ensures that in the event of any data anomalies or audits in the future, the responsibility and basis for the correction of a field can be accurately located.
[0109] After manual review is completed, the field will receive its final value, which the system will incorporate into the "Compliance Correction Data" for the next step (such as multi-dimensional analysis or adaptive routing). If other logical conflicts related to this field arise later, the history table can quickly provide previous review processes to assist in making further judgments or modifying rules.
[0110] Although automatic scripts can handle most common anomalies, manual review is still required to conduct in-depth intervention in unexpected situations, forming a complete data quality assurance system that combines "automatic + manual". By correcting the historical table to record field names, original values, reviewer identification, final values and other information, each manual intervention can be traced, avoiding the sequelae of blind modifications. When later audits or management need to understand the reasons for data corrections, they can backtrack based on this table to clarify the person in charge and the decision-making process. It is difficult for an enterprise's rule base or scripts to completely cover all new business forms in the short term. Manual review can provide temporary or transitional solutions for new scenarios. When reviewers encounter the same type of problem multiple times, they can subsequently optimize the script or expand the business rules to gradually reduce the frequency of manual intervention.
[0111] Perform multi-dimensional feature analysis and adaptive routing decisions on the compliance correction data, distribute the results to various departments to form multi-department synchronized data, and archive the multi-department synchronized data in the enterprise data lake after formation, and provide an external access interface.
[0112] In the previous steps of the present invention, compliant and corrected data has been generated, which contains cross-departmental aligned field information that has undergone anomaly detection and correction. Due to its high quality and consistency, this data is suitable for use in the multi-dimensional analysis and routing decision-making stages.
[0113] The system extracts features from the fields of compliance revision data, covering the following dimensions: Business priority (e.g., order urgency, financial terms); security level (e.g., fields involving personal privacy or financial confidentiality); data source (different departments or external systems); data temporal distribution or magnitude (size range, number of records). These dimensions can be weighted or customized to construct feature vectors, allowing for comprehensive consideration of downstream processing strategies for each piece of data.
[0114] After obtaining the multi-dimensional feature analysis results, the adaptive routing module of the present invention determines how to distribute the compliance correction data to each department. Specific methods may include: Determine which departments need to synchronize immediately and which departments can wait for batch processing based on priority; choose encrypted transmission or internal visibility only based on security level; allocate routing paths in real time based on data volume and network load conditions; through this adaptive strategy, each department will obtain targeted data at the right time and in an appropriate manner without the need for a unified mandatory model.
[0115] After routing is complete, the system successfully distributes data to each department, which is called multi-department synchronized data. This data is also written to the enterprise data lake for persistent storage, allowing for subsequent analysis, auditing, or review. In this context, the data lake is a centrally managed storage platform that can accommodate a variety of structured and unstructured data.
[0116] To meet the needs of upstream and downstream supply chain or other external partners to query and obtain data, the present invention establishes an external access interface at the enterprise data lake level. This interface can verify the permissions of external identities based on security policies to prevent unauthorized access.
[0117] In practical applications, when processing compliance correction data, a feature vector is generated for each record or field. For example, the feature vector ,in: Indicates the service priority value (0 means the lowest, and larger values indicate more urgent); Indicates the security level identifier (can be different integer levels); This represents the amount of data or record size. Other factors are defined based on the company's specific needs. After feature extraction, the system calculates routing policy recommendations based on preset formulas or rules, such as priority or encryption requirements.
[0118] The system implements routing logic at the message queue or microservice gateway level. High-priority messages are directly pushed to the relevant department's server; high-security messages utilize encrypted channels. The routing module incorporates a "load status" monitoring system. If a department experiences high concurrency, delayed queues or sharded distribution are implemented to ensure overall efficiency.
[0119] After successful routing, each department receives the data it needs, which is called "multi-department synchronized data." To ensure complete records, the system marks each piece of data as successfully delivered after distribution and writes relevant metadata (timestamp, destination department, and delivery status) into the synchronization log table.
[0120] After synchronization among multiple departments is completed, the system will archive the compliance correction data and its distribution records to the enterprise data lake; when the external access interface is opened to partners or other external applications, it can be combined with role permissions and data classification strategies, such as only allowing viewing of some fields or aggregated results; this interface can use API gateways or services to authenticate external users, and perform secondary filtering or encryption when necessary.
[0121] Multi-dimensional feature analysis enables the system to identify high-priority or sensitive data and prioritize its distribution to the appropriate departments. With timely and accurate data, each department can accelerate business processing and decision-making, enhancing overall collaborative efficiency. Adaptive routing decisions, combined with load monitoring and security policies, automatically configure the most appropriate transmission mode for each department, ensuring smooth network operation while balancing data security and the protection of sensitive information.
[0122] Every data release or cross-departmental synchronization is archived in the enterprise data lake, creating a complete historical data version and operation log. If business parties subsequently need to audit or reprocess past data, they can quickly retrieve and access it from the data lake.
[0123] Preferably, when performing multi-dimensional feature analysis and adaptive routing decisions on the compliance correction data, department identification, field sensitivity level and number of data corrections are used as basic features, and a feature vector is constructed in combination with network load information. The feature vector is aggregated through weighted calculation, and the aggregation result is used to determine the distribution order and security policy, and a routing label is added to each data during distribution.
[0124] In the context of multi-departmental collaboration within an enterprise, data distribution must consider not only field sensitivity levels but also departmental privileges, historical correction counts (affecting data credibility and stability), and network load (ensuring overall performance). Simply relying on a single metric to determine transmission priority and security policies can lead to insufficient protection for critical or sensitive data, or excessive encryption that impacts efficiency. To address this, this paper proposes a multi-dimensional feature vector aggregation method to enable refined decision-making for each piece of compliant correction data.
[0125] The present invention combines department identification, field sensitivity level, data correction times, and network load information into a feature vector in a certain order or structure. An exemplary vector can be expressed as:
[0126] in Indicates department identification (can be numerical or categorical code); Indicates the sensitivity level of the field (the higher the value, the higher the security protection requirement); Indicates the number of data corrections (reflecting potential data stability or suspicion); Indicates network load information (such as load factor or congestion index).
[0127] In order to integrate these four characteristics, the present invention adopts a weighted form for aggregation, such as:
[0128] in is the weight, which is used to reflect the proportion of each feature in the routing decision; To aggregate the results and determine important parameters such as the distribution order and encryption strategy, machine learning or scoring formulas known in the art can also perform more complex processing on the vectors, such as classifiers or multidimensional clustering. However, only the core weighted concepts are explained here to avoid irrelevant complex expressions.
[0129] Get the aggregate value The system then determines which data requires encryption and prioritization, which data can be packaged and distributed later, and how to schedule it within the network load. To achieve detailed traceability, each piece of data is assigned a routing tag during distribution, including the current routing batch; the target department or node identifier; and encryption or security policy instructions. These tags play a crucial role in subsequent audits and troubleshooting.
[0130] In practical applications, when extracting feature vectors, the system assigns a department ID to compliance correction data based on the source or target department information in the data; for example, the Finance Department is recorded as 2, the Human Resources Department is recorded as 3, and so on. Each piece of data contains an assessed sensitivity score (e.g., 0 to 10, with higher values indicating higher security requirements), and the system directly reads the value of this field. When the number of automatic corrections or manual reviews accumulates in a record or is saved in a correction history table, this number is incorporated into the feature vector of that data. A higher number indicates a greater history of anomalies upstream. This value is monitored by the system in real time to monitor the current routing link or message queue pressure.
[0131] Different companies can be given sensitivity levels With greater weight, to prioritize confidentiality; if the business scenario attaches great importance to data stability, it can be magnified , to reduce the transmission of data with high error correction frequency; The larger the value, the higher the priority of the distribution and the more likely it is to use a security policy (such as an encrypted channel). If the value is lower, the system can delay sending and use normal channels when the network load is heavy to improve the overall throughput efficiency.
[0132] System basis Automatically decide whether to distribute data encrypted or plain, and write "Security Policy = Encrypted" or "Security Policy = Plain" into the routing tag. The routing tag can also include the distribution batch, target department ID, and timestamp. This facilitates subsequent retrospective queries in the data lake or audit log backend management.
[0133] Preferably, based on the compliance-corrected data whose sensitivity level is higher than the preset value shown in the aggregation result, the data is distributed to the target department after performing a security encryption operation separately; for the compliance-corrected data whose sensitivity level is lower than the preset value, it is packaged with similar data into multi-department synchronization data, and a routing log is generated after the distribution is completed; the routing log contains the routing label, the target department and the encryption status.
[0134] In multi-dimensional feature analysis, the sensitivity level is a quantitative indicator of data's security or privacy-related factors. If the value is higher than the preset value, it indicates that the data has high security protection requirements.
[0135] The present invention extracts the sensitivity level field from the analysis result, and compares its value with a preset threshold to determine whether to enable an encrypted transmission process for the batch of data.
[0136] For compliance-revision data with a sensitivity level above the threshold, the solution implements a separate encryption step. Using either symmetric or asymmetric encryption, data is fully encrypted before being distributed to the target department. This ensures that data cannot be viewed by unauthorized individuals during transmission, meeting corporate confidentiality and compliance requirements.
[0137] If the sensitivity level is below the threshold, the system deems the data batch as not requiring special security protection. It is then packaged with other data of the same level to form multi-department synchronized data, which is distributed to relevant departments in batches or groups. This operation is suitable for common business fields or shared information, reducing repeated transmissions and improving overall data exchange efficiency.
[0138] After the distribution is completed, the solution will generate a routing log, which mainly records: Routing label: distinguishes different types or batches of data flows; target department: indicates the final destination of the data; encryption status: indicates whether the data has been securely encrypted during transmission; this log is written to the synchronization log table or similar storage structure to facilitate tracing the data flow during subsequent audits or troubleshooting.
[0139] In actual application, the system first reads the field from the feature analysis results , indicating the sensitivity level value. Compare: If It is considered highly sensitive data; if It is classified as low-sensitivity data. Defined by the company's internal security policy, it can be fine-tuned according to different business areas.
[0140] For highly sensitive data, the system loads a security module and performs secure encryption on the data blocks. To maintain interoperability in cross-departmental scenarios, each department needs to use the same or compatible decryption key (or corresponding public / private key), which can be maintained uniformly by the key management service outside the solution.
[0141] When low-sensitivity data dominates, the system divides data of the same or similar security levels into batches for multi-department synchronization; in the downstream transmission stage, the relevant modules mark the data with the same routing label (such as "Normal_Sync"), distribute it according to the needs of the other department, and record the distribution time.
[0142] Routing tag: used to identify specific distribution channels (e.g., "HighSecureChannel" for highly sensitive fields, "NormalSync" for other fields); Target department: Finance, Operations, etc.; Encryption status: indicates whether the data is encrypted or plaintext. If highly sensitive data is encrypted, the encryption status is "Encrypted"; if less sensitive data is plaintext, the status is "Plain."
[0143] Encryption strategies are implemented based on the sensitivity level assessed in real time, ensuring the security of high-value sensitive data while avoiding unnecessary encryption overhead for large amounts of less sensitive data. Low-sensitivity data is distributed in batches to reduce network load, while highly sensitive data is transmitted individually to mitigate the risk of leakage, aligning with the digital enterprise's goals of both compliance and efficiency. Routing logs are maintained after each distribution, indicating the routing tag, destination department, and encryption status. This log can be easily queried to verify which confidential fields a department has accessed and whether the correct encryption process was followed.
[0144] Preferably, after forming multi-department synchronized data, the multi-department synchronized data is archived to the enterprise data lake and the external access interface is configured, and access permission levels and audit modes are set for each external access connection, and the visitor identity and access time are identified through access logs. When a request that does not conform to the registered permission level is detected, an interception strategy is adopted and the interception information is recorded.
[0145] Through multi-dimensional analysis and adaptive routing, this solution has already achieved internal data distribution across multiple departments, but there is still a need for external data sharing. By archiving synchronized data from multiple departments in the enterprise data lake, historical versions and access rights can be managed in a unified manner. Building on the data lake structure, the solution also configures external access interfaces, providing channels for compliant partners or authorized entities to query or obtain data externally.
[0146] In a digital enterprise environment, different external access parties have varying levels of data sensitivity and visible content. This solution implements differentiated authorization through a set of access rights grading models. Each external access connection is assigned a grading value (e.g., Lv1, Lv2, Lv3) upon establishment. The corresponding audit mode may include recording information such as access time, access content, and the initiator's identity. Stricter grading values for audit modes often have more audit fields associated with them.
[0147] All access actions are recorded in the access log, including fields such as the visitor's identity (such as an account number or digital certificate), access time, and requested resource, for subsequent review. If the system verifies that the access request's permission level does not match the pre-registered information, an interception policy is implemented and interception information (such as the interception time and requested resource identifier) is recorded. This mechanism prevents unauthorized access or potential intrusion.
[0148] In practice, after multi-department synchronization is complete, the solution archives the corresponding datasets (including metadata) into the data lake. Each record in the data lake is assigned a specific security or sensitivity level, enabling real-time authentication for external access interfaces. During the configuration phase, the external interface specifies the permission levels for external access connections, and a mapping between accessor identities and level values is stored in the permission table.
[0149] When an external system needs to access the data lake, it submits a connection request (including identity credentials and access reason). The system assigns a corresponding level (e.g., Level 2) based on the registered profile. The audit mode settings also specify what information (time, results, content, etc.) should be recorded when accessing the data lake.
[0150] When a visitor initiates a query through an interface, the system queries the permission table to confirm whether the visitor's rating is appropriate for the request. The access action and time are then recorded in the access log. If the request matches the rating, the data is returned. If the request exceeds the permitted range, the system intercepts the request, generates interception information (reason, time, and request target), and can notify the security administrator.
[0151] All requests for bypassing access levels are rejected and violations are logged. Access to sensitive data is protected by dedicated encryption pipelines or secondary authentication strategies, ensuring that highly sensitive fields in the data lake cannot be read by low-level access connections.
[0152] By managing permissions and auditing data across multiple departments, the system ensures that external parties can only access data within their authorized scope, preventing over-disclosure and data abuse. Interception policies take immediate effect in the event of an overreach, mitigating potential risks.
[0153] Access logs detail visitor identities and operation times, facilitating backtracking after security incidents or verifying the legitimacy of access behaviors during periodic audits. The solution supports the continuous addition or modification of external access entities and classification values. Strict auditing can be implemented for scenarios with high security requirements, while lightweight auditing is employed for common external requests, enhancing system flexibility.
[0154] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware.
[0155] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included within the scope of the claims of the present application.
Claims
1. A method for adaptive exchange and intelligent integration of digital enterprise data, characterized in that: The following steps are involved: Obtaining original business data from each business department and performing field mapping and format unification processing to form first department pre-processed data, and publishing the first department pre-processed data to a message bus; Performing a cross-department field alignment operation on the first department pre-processed data, and merging the alignment information with the first department pre-processed data to generate second department pre-processed data; The alignment operation compares and confirms the fields based on the self-organizing multi-agent approach; Performing anomaly detection and data correction on the pre-processed data of the second department to generate compliant corrected data; Perform multi-dimensional feature analysis and adaptive routing decisions on the compliance correction data, distribute the results to various departments to form multi-department synchronized data, archive the multi-department synchronized data in the enterprise data lake and provide an external access interface.
2. The method according to claim 1, characterized in that After each business department obtains the original business data, a unique field identifier is assigned to each field, and the field's numerical range and text length are checked using a data type verification script. Field values that do not meet the inspection standards are corrected according to predefined mapping rules and together with the field values that meet the inspection standards, they constitute the first department's preprocessed data.
3. The method according to claim 1, characterized in that When performing a cross-department field alignment operation on the pre-processed data of the first department, the field definitions of multiple departments are compared in parallel based on the self-organizing multi-agent to generate a matching score list for each field candidate object, and the matching score list is passed to the aggregation node; The aggregation node counts and judges the matching score list based on the preset matching threshold. If all candidate scores are lower than the matching threshold, the alignment is considered a failure and manual verification is triggered. If there is a score greater than the matching threshold, it is considered a preliminary alignment and the matching score and candidate field are recorded in the alignment information. When the number of field definitions is greater than or equal to 2 and the matching scores all exceed the preset matching threshold, the self-organizing multi-agent compares the usage frequency of the fields within the department in descending order according to the scores, prioritizes the most frequently used fields, and writes the confirmation order and department usage records in the alignment information; After completing the above alignment, the alignment information is stored in the alignment management library in the form of a mapping table. The mapping table contains the field name, corresponding department identifier and matching score. When generating the second department preprocessing data, the mapping table entries are synchronously updated and timestamps are marked, so that subsequent steps can perform anomaly detection or data correction processing based on the latest field correspondence.
4. The method according to claim 1, wherein When performing anomaly detection and data correction on the pre-processed data of the second department, the data range and logical verification conditions are configured for each field according to the preset business rules. For field records whose data range exceeds the preset conditions or there are logical conflicts, the correction script is called to modify the field value and generate an anomaly identification entry, which includes the field name, original value, revised value and verification conditions.
5. The method according to claim 4, characterized in that For field records that cannot be corrected through the correction script, the manual review process is initiated after the correction script returns an exception. The reviewer confirms the correction value in the field comparison list and records the field values before and after the review in the correction history table. The correction history table contains the field name, reviewer ID and the final effective value.
6. The method according to claim 1, characterized in that When performing multi-dimensional feature analysis and adaptive routing decisions on the compliance correction data, department identification, field sensitivity level and number of data corrections are used as basic features, and feature vectors are constructed in combination with network load information. The feature vectors are aggregated through weighted calculation, and the aggregation results are used to determine the distribution order and security policy, and routing labels are added to each data during distribution.
7. The method according to claim 6, characterized in that Based on the compliance-corrected data whose sensitivity level is higher than the preset value shown in the aggregation results, a security encryption operation is performed separately and then distributed to the target department; for the compliance-corrected data whose sensitivity level is lower than the preset value, it is packaged with similar data into multi-department synchronization data, and a routing log is generated after the distribution is completed; the routing log contains the routing label, the target department and the encryption status.
8. The method according to claim 1, characterized in that After forming multi-department synchronized data, archive the multi-department synchronized data to the enterprise data lake and configure the external access interface, set access permission levels and audit modes for each external access connection, identify the visitor identity and access time through access logs, and adopt interception strategies and record interception information when requests that do not conform to the registered permission levels are detected.
Citation Information
Patent Citations
A method and system for integrating business data of multiple departments
CN118689882B
Cited By
Data real-time synchronization method and system, terminal equipment and storage medium
CN121117114A