Enterprise multi-source data collaborative management system based on industrial internet platform

The enterprise multi-source data collaborative management system based on the industrial internet platform solves problems such as ambiguous identification of data source types, inaccurate volume assessment, lack of logical collaborative processing, and lag in conflict handling in enterprise multi-source data management. It realizes the orderly and efficient flow and secure management of multi-source data, and enhances the full release of data value.

CN120996752APending Publication Date: 2025-11-21SHENZHEN ZHONGTIAN YUNLIAN TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511116739.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Enterprises face problems such as ambiguous identification of data source types, inaccurate assessment of data volume, lack of logical collaborative processing, delayed conflict handling, and narrow coverage of security audit mechanisms in multi-source data management, leading to disorder and inefficiency in data management.

Method used

The enterprise multi-source data collaborative management system based on the industrial internet platform includes a data access module, a collaborative rule engine module, a conflict resolution module, a quality monitoring module, and a security audit module. By identifying master data and slave data, collecting metadata information, setting collaborative strategies, building conflict prediction models, and combining quality monitoring indicators and security audit mechanisms, it achieves dynamic optimization and overall management.

Benefits of technology

It enables the orderly and efficient flow of multi-source data, accurately identifies the type and volume of data sources, reduces processing deviations, moves conflict resolution forward, enhances security, comprehensively reflects changes in data quality, and improves the release of data value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996752A_ABST
    Figure CN120996752A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial data collaboration, and discloses an enterprise multi-source data collaboration management system based on an industrial internet platform. A data access module of the system obtains multi-source data, determines master and slave data, collects metadata, and identifies data source types and volumes; the collaboration rule engine module sets a collaboration strategy based on the data source type and volume, locates an association relationship and determines a collaboration processing flow; the conflict resolution module identifies core data and a dependency path, constructs a conflict prediction model and sets a conflict resolution rule; the quality monitoring module senses a data processing state, analyzes a circulation trend and sets a quality monitoring index; the security auditing module identifies an access abnormal behavior and sets a security auditing mechanism in combination with a conflict prediction model; and the collaborative management module integrates the indexes, the mechanisms and the rules, carries out collaborative management on the multi-source data, and forms a collaborative management result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial data collaboration technology, specifically to an enterprise multi-source data collaboration management system based on an industrial internet platform. Background Technology

[0002] With the deep penetration of industrial internet technology, the data generated during enterprise operations is experiencing explosive growth. These data sources are diverse, encompassing real-time sensor data from production equipment, supply chain management system data, customer relationship management data, and financial system data, forming a complex and diverse multi-source data system. Different data sources exhibit significant differences in generation frequency, storage format, and update speed. For example, production equipment data is continuously generated at millisecond-level frequencies, while financial statement data is updated monthly or quarterly. This difference presents numerous challenges in the data aggregation and processing process.

[0003] Enterprises generally face numerous pain points in multi-source data management. During the data access phase, the lack of a clear definition of master and slave data, coupled with incomplete metadata collection, leads to ambiguous identification of data source types and inaccurate data volume assessments, thus affecting the targeted nature of subsequent data processing. In the collaborative processing stage, most enterprises rely on traditional experience to set collaborative strategies, failing to dynamically adjust them based on data source type and volume. The relationships between data are poorly defined, resulting in a lack of logic and efficiency in the collaborative processing flow.

[0004] Data conflict resolution is often reactive, with most companies only taking remedial action after a conflict occurs. They fail to identify core data and its dependencies beforehand and lack effective conflict prediction mechanisms, leading to delays in conflict resolution and impacting data flow efficiency. Regarding quality monitoring, existing mechanisms often focus on individual data processing statuses, neglecting the analysis of data flow trends. Monitoring indicators are often detached from metadata information, making it difficult to comprehensively reflect data quality.

[0005] The security audit process also has shortcomings. Identifying abnormal access behavior relies solely on simple permission checks, failing to integrate with data quality monitoring and conflict prediction. This results in a narrow coverage of the security audit mechanism, making it difficult to address data security threats in complex network environments. These intertwined problems lead to a disordered and inefficient management of multi-source data within enterprises, hindering the full realization of data value. Summary of the Invention

[0006] The purpose of this invention is to provide an enterprise multi-source data collaborative management system based on an industrial internet platform to solve the problems mentioned in the background art.

[0007] To achieve the above objectives, the present invention provides an enterprise multi-source data collaborative management system based on an industrial internet platform, the system comprising:

[0008] The data access module is used to acquire multi-source data from the enterprise, determine the master data and slave data of the multi-source data, collect the metadata information of the master data, and identify the data source type and data volume of the multi-source data based on the metadata information.

[0009] The collaborative rules engine module is used to set a collaborative strategy for the multi-source data based on the data source type and the data volume, locate the relationship between the data in the collaborative strategy, and determine the collaborative processing flow of the multi-source data based on the metadata information and the relationship.

[0010] The conflict resolution module is used to identify the core data and the dependency path of the core data in the collaborative processing flow, construct a conflict prediction model of the core data based on the dependency path, and set conflict resolution rules for the multi-source data according to the conflict prediction model and the dependency path.

[0011] The quality monitoring module is used to combine the collaborative processing flow and the conflict resolution rules to perceive the data processing status of the multi-source data in real time, analyze the flow trend of the multi-source data based on the metadata information, and set the quality monitoring indicators of the multi-source data in combination with the data processing status and the flow trend.

[0012] The security audit module is used to identify abnormal access behaviors of the multi-source data based on the quality monitoring indicators; and to set up a security audit mechanism for the multi-source data according to the abnormal access behaviors and the conflict prediction model.

[0013] The collaborative management module is used to perform collaborative management processing on the multi-source data according to the quality monitoring indicators, the security audit mechanism and the conflict resolution rules, and obtain collaborative management results.

[0014] Preferably, determining the collaborative processing flow of the multi-source data based on the metadata information and the association relationship includes:

[0015] Based on the metadata information, identify the business scenarios of the multi-source data;

[0016] Based on the business scenario, schedule the real-time requirements and collaboration strategies for the multi-source data;

[0017] Based on the real-time requirements and the relationships, the collaboration strategy is dynamically optimized to obtain an optimized strategy.

[0018] Based on the real-time requirements and the optimization strategy, the collaborative processing flow for the multi-source data is determined.

[0019] Preferably, the step of constructing a conflict prediction model for the core data based on the dependency path includes:

[0020] Based on the dependency path, calculate the associated data nodes of the core data;

[0021] Identify the non-core data on the associated data nodes;

[0022] Analyze the update frequency of the non-core data relative to the core data;

[0023] Extract the business rules of the core data, and analyze the constraint impact of the business rules on the non-core data based on the update frequency;

[0024] Based on the update frequency and the impact of the constraints, a conflict prediction model for the core data is constructed.

[0025] Preferably, setting conflict resolution rules for the multi-source data based on the conflict prediction model and the dependency path includes:

[0026] Real-time monitoring of the data stream from the multi-source data, and identification of conflicting data within the data stream;

[0027] Based on the conflict prediction model, analyze the synergistic impact of the conflict data on the multi-source data;

[0028] The association logs of the multi-source data and the conflicting data are scheduled, and the cause of the conflict of the conflicting data is identified based on the association logs;

[0029] Based on the synergistic effects and the causes of the conflict, a data correction mechanism for the conflict data is set up;

[0030] Identify the business links of the multi-source data, and set the processing priority of the conflicting data based on the business links and the data correction mechanism;

[0031] Based on the dependency path and the processing priority, set the conflict resolution rules for the multi-source data.

[0032] Preferably, the step of analyzing the flow trend of the multi-source data based on the metadata information includes:

[0033] The metadata information is divided into dimensions to obtain the division results;

[0034] Extract the feature parameters of the partitioning results, and collect the historical flow records of the multi-source data based on the feature parameters;

[0035] Identify the peak and trough periods in the historical data flow records;

[0036] Analyze the traffic fluctuation patterns of the multi-source data based on the peak periods;

[0037] Based on the aforementioned low-period periods, analyze the storage occupancy rate of the multi-source data;

[0038] By combining the traffic fluctuation pattern, the storage occupancy rate, and the characteristic parameters, the flow trend of the multi-source data is analyzed.

[0039] Preferably, the step of setting quality monitoring indicators for the multi-source data by combining the data processing status and the flow trend includes:

[0040] Based on the data processing status, the integrity and consistency of the multi-source data are identified;

[0041] Analyze the correlation between the integrity and consistency of the data and the business value of the multi-source data;

[0042] Based on the flow trend and the correlation with business value, identify the weak links in the quality of the multi-source data, and identify the scope of impact of the weak links;

[0043] Based on the data processing status and the scope of influence, set the quality threshold for the multi-source data;

[0044] Based on the influence range and the quality threshold, quality monitoring indicators for the multi-source data are set.

[0045] Preferably, identifying abnormal access behavior of the multi-source data based on the quality monitoring indicators includes:

[0046] Based on the aforementioned quality monitoring indicators, historical access records of the multi-source data are collected.

[0047] Based on the historical access records, identify the normal access patterns of the multi-source data;

[0048] Identify the current access status of the multi-source data, and combine the current access status with the historical access records to identify access behavior deviations of the multi-source data;

[0049] Based on the normal access mode, set the abnormal access threshold for the multi-source data;

[0050] Based on the abnormal access threshold and the access behavior deviation, abnormal access behavior of the multi-source data is identified.

[0051] Preferably, the step of setting up a security audit mechanism for the multi-source data based on the abnormal access behavior and the conflict prediction model includes:

[0052] Based on the abnormal access behavior, locate the risky data items in the multi-source data;

[0053] Analyze the threat level of the risk data item, and set the protection strategy layer for the risk data item based on the threat level;

[0054] Based on the conflict prediction model, identify the scope of the associated impact of the risk data items;

[0055] Based on the scope of the associated impact, set the isolation path for the risk data item;

[0056] By combining the protection strategy layer and the isolation path, a security audit mechanism for the multi-source data is set up.

[0057] Preferably, the step of performing collaborative management processing on the multi-source data based on the quality monitoring indicators, the security audit mechanism, and the conflict resolution rules to obtain collaborative management results includes:

[0058] Based on the quality monitoring indicators, identify quality problems in the multi-source data;

[0059] Based on the aforementioned quality issues, the triggering conditions for the conflict resolution rules are set, and the data repair method for the multi-source data is also set.

[0060] Based on the triggering conditions and the data repair method, conflict resolution is performed on the multi-source data to obtain the processing result;

[0061] Based on the processing results and the security audit mechanism, the multi-source data is processed collaboratively to obtain collaborative management results.

[0062] Preferably, the step of setting a collaborative strategy for the multi-source data based on the data source type and the data volume includes:

[0063] Analyze the business scenario requirements corresponding to the data source type;

[0064] Based on the business scenario requirements, determine the real-time and accuracy requirements for data collaboration;

[0065] Based on the data volume, assess the computational resource requirements for data processing;

[0066] Based on the aforementioned real-time requirements, accuracy requirements, and computing resource needs, a preliminary collaborative strategy is formulated;

[0067] The feasibility of the preliminary collaboration strategy is verified, the strategy parameters are adjusted and optimized, and the final collaboration strategy is set.

[0068] Compared with the prior art, the beneficial effects of the present invention are:

[0069] This enterprise multi-source data collaborative management system, based on an industrial internet platform, forms a complete multi-source data management system through the organic linkage of its various modules. The data access module clearly defines master and slave data, collects metadata information, and accurately identifies data source types and data volumes, providing a clear basic data framework for subsequent processing and avoiding processing deviations caused by chaotic data sources.

[0070] The collaborative rules engine module sets collaborative strategies based on data source type and volume, accurately locating data relationships and constructing collaborative processing flows by combining metadata information. This ensures data follows a logical order during flow, reducing unnecessary redundancy and making data processing smoother. The conflict resolution module focuses on core data and its dependent paths, building conflict prediction models and setting resolution rules. It moves the data conflict processing nodes forward, reducing processing costs after conflicts occur and minimizing data flow bottlenecks by anticipating potential conflicts.

[0071] The quality monitoring module integrates collaborative processing workflows and conflict resolution rules to monitor data processing status in real time. It also analyzes flow trends based on metadata information, making monitoring indicators more closely aligned with actual data characteristics. This allows for a more comprehensive reflection of data quality changes, facilitating timely detection and adjustment of inappropriate data processing. The security audit module combines quality monitoring indicators with anomaly detection, referencing conflict prediction models to establish a security audit mechanism. This broadens the dimensions of security protection, enabling more accurate capture of abnormal data access behavior and enhancing data management security.

[0072] The collaborative management module integrates quality monitoring indicators, security audit mechanisms, and conflict resolution rules to manage multi-source data in a unified manner. This enables standardized control over the entire data process from access to processing to application. The functions of each module complement each other in this process, avoiding the limitations of single-stage management and allowing multi-source data to flow efficiently within an orderly framework, fully leveraging the collaborative value between different data sources. Attached Figure Description

[0073] Figure 1 This is a schematic diagram illustrating the working principle of the enterprise multi-source data collaborative management system based on the industrial internet platform described in this invention.

[0074] Figure 2 A flowchart of the collaborative processing flow for multi-source data;

[0075] Figure 3 Design diagram of a conflict prediction model for core data;

[0076] Figure 4 Design diagram for quality monitoring indicators of multi-source data. Detailed Implementation

[0077] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0078] Please see Figure 1 This invention provides an enterprise multi-source data collaborative management system based on an industrial internet platform. The system includes: a data access module, a collaborative rule engine module, a conflict resolution module, a quality monitoring module, a security audit module, and a collaborative management module. The specific implementation methods of each module are as follows:

[0079] The data access module is used to acquire multi-source data from the enterprise and determine the master and secondary data. Master data refers to data that has core value and is stable over a long period in the enterprise's business processes, such as basic product information and core customer information. Secondary data is related data generated around the master data, such as real-time production data of products and customer transaction records. The module collects metadata information from the master data, including data structure, data format, data source, and update time. Based on the metadata information, it identifies the data source types of the multi-source data, such as sensor data, ERP system data, and MES system data; it also calculates the data volume, i.e., the storage size and number of records.

[0080] The collaboration rules engine module sets collaboration strategies for multi-source data based on data source type and data volume. For example, a real-time synchronization collaboration strategy is used for sensor data with high real-time requirements; a batch processing collaboration strategy is used for historical production data with large data volumes. It identifies the relationships between data sources within the collaboration strategy, such as the mapping relationship between a specific data source and the master data, and the dependencies between different data sources. Based on metadata information and relationships, it determines the collaborative processing flow for multi-source data, ensuring that data can be efficiently collaborated according to preset rules during the data flow process.

[0081] The conflict resolution module identifies core data and its dependency paths within the collaborative processing flow. Core data refers to data that plays a crucial role in business decisions, and its dependency path represents the stages the data passes through during processing and the order of relationships between these stages. Based on these dependency paths, a conflict prediction model for core data is constructed to anticipate potential data conflicts. According to the conflict prediction model and dependency paths, conflict resolution rules for multi-source data are set, and when data conflicts occur, they are processed quickly according to these rules.

[0082] The quality monitoring module, combining collaborative processing workflows and conflict resolution rules, monitors the real-time data processing status of multi-source data, such as whether data is being transmitted and whether data cleaning has been completed. Based on metadata information, it analyzes the flow trends of multi-source data, such as changes in data transmission volume over different time periods and the time data spends in each business process. Combining the data processing status and flow trends, it sets quality monitoring indicators for multi-source data, such as data integrity indicators and data accuracy indicators, to monitor data quality in real time.

[0083] The security audit module identifies abnormal access behaviors to multi-source data based on quality monitoring indicators, such as unauthorized user access and unusually frequent access. Based on abnormal access behavior and conflict prediction models, a security audit mechanism for multi-source data is established, including access log recording, abnormal behavior alerts, and regular security checks, to ensure the security of data access.

[0084] The collaborative management module performs collaborative management and processing of multi-source data based on quality monitoring indicators, security audit mechanisms, and conflict resolution rules. For example, when quality monitoring indicators show that data quality is substandard, the corresponding processing mechanism is triggered to repair the data; when the security audit mechanism detects abnormal access, measures are taken in a timely manner to prevent unauthorized operations; when data conflicts occur, they are coordinated and processed according to conflict resolution rules. The final collaborative management results provide reliable data support for enterprise business decisions.

[0085] Example 1: Please refer to Figure 2 When determining the collaborative processing flow for multi-source data based on metadata information and relationships, the business scenario of the multi-source data is identified based on the metadata information. Metadata information includes the data's origin, application domain, and associated business modules. Parsing this information clarifies the business scenario to which the data belongs. For example, data labeled "production workshop equipment operating parameters" belongs to the manufacturing process; data labeled "customer order information" belongs to the sales process. The data processing logic and collaboration methods differ across business scenarios; accurately identifying the business scenario is the foundation for subsequent process design.

[0086] The system adapts to the real-time requirements and collaboration strategies of multi-source data based on business scenarios. In manufacturing scenarios, real-time data such as equipment temperature and pressure need to be transmitted and processed within a very short time to support immediate adjustments to the production line. In inventory management scenarios, historical outbound records of goods have lower real-time requirements and can be aggregated and processed at fixed intervals. Simultaneously, corresponding collaboration strategies are matched to different business scenarios. Real-time data synchronization is used in production scenarios to ensure data consistency across all stages; timed batch collaboration is used in inventory scenarios to reduce resource consumption.

[0087] The collaboration strategy is dynamically optimized based on real-time requirements and relationships to obtain an optimized strategy. When the relationship between data and master data changes in a business scenario—for example, material supply data and production planning data that were originally weakly related become strongly related, and the real-time requirement for material supply data increases—the collaboration strategy needs to be adjusted. The original daily data synchronization might be changed to hourly synchronization to adapt to the new relationships and real-time requirements. If the real-time requirement decreases, such as when a certain type of sales report data changes from requiring real-time updates to daily updates, the real-time processing strategy can be optimized to a batch processing strategy to reduce unnecessary resource consumption.

[0088] The collaborative processing flow for multi-source data is determined by combining real-time requirements and optimization strategies. The complete path of data transmission from its entry into the system, through cleaning, transformation, and integration steps, to the target module is clearly defined, specifying the processing time limits and connection methods for each step. For example, in a production scenario, the optimized strategy is real-time processing. The collaborative processing flow must ensure that sensor data, after being received, is cleaned and converted in format within 100 milliseconds and immediately transmitted to the production monitoring module to ensure real-time requirements are met. In a sales scenario, the optimized strategy is timed batch processing. The collaborative processing flow is set to summarize and process the previous day's sales data every morning at midnight and transmit it to the data analysis module before 8:00 AM.

[0089] When setting collaborative strategies for multi-source data based on data source type and data volume, it's essential to analyze the corresponding business scenario requirements for each data source type. IoT device data sources primarily serve device monitoring scenarios, requiring continuous acquisition of device operating status to prevent malfunctions. CRM system data sources serve customer relationship management scenarios, requiring complete recording of customer interaction information to support customer service optimization. The business scenarios corresponding to different data source types differ in their data processing priorities: device monitoring scenarios emphasize real-time performance, while customer relationship management scenarios emphasize data integrity and long-term storage.

[0090] The real-time and accuracy requirements for data collaboration are determined based on the needs of the business scenario. In equipment monitoring scenarios, the real-time requirement for data collaboration is at the second level to ensure that equipment anomalies can be detected in a timely manner, while the accuracy requirement allows for minor errors, such as an acceptable error range of ±0.5℃ for temperature data. In financial accounting scenarios, the accuracy requirement for data collaboration is extremely high, and no calculation errors are allowed. The real-time requirement can be relaxed to the hour level, as long as data collaboration is completed before the end of the day's accounting.

[0091] The computational resource requirements for data processing are assessed based on the data volume. When the data source is a high-frequency sensor network that generates tens of gigabytes of data per hour, a distributed computing cluster needs to be deployed, allocating sufficient CPU, memory, and storage resources to handle the parallel processing of large amounts of data. When the data source is an employee information table with a data volume of only megabytes, the computational resources of a single server can meet the processing requirements, and no additional resources are needed.

[0092] A preliminary collaborative strategy was developed based on real-time requirements, accuracy requirements, and computing resource needs. For high-real-time, large-volume sensor data, the preliminary strategy adopted a stream processing framework for real-time computation, prioritizing the allocation of computing resources while implementing a data compression mechanism to reduce storage consumption. For high-accuracy, small-volume financial data, the preliminary strategy adopted a batch processing framework for fine-grained verification to ensure data accuracy, with computing resources configured as usual.

[0093] The feasibility of the initial collaborative strategy was verified by simulating the data flow process to check whether the strategy could meet the real-time and accuracy requirements under the given computing resources. If the simulation revealed that the processing latency of sensor data exceeded the real-time requirements, the number of computing nodes was increased; if financial data was omitted during the verification process, the verification algorithm was optimized. Based on the verification results, strategy parameters were adjusted, such as modifying the parallelism of the stream processing framework and adjusting the verification frequency of batch processing, to ultimately form a stable and operational collaborative strategy.

[0094] Example 2: Please refer to Figure 3 When constructing a conflict prediction model for core data based on dependency paths, the associated data nodes of the core data are calculated based on the dependency paths. Dependency paths clearly present the various processing nodes, storage nodes, and interacting data nodes involved in the flow of core data. By sorting and analyzing these paths, all data nodes directly or indirectly related to the core data can be located. These associated data nodes may be distributed across different business systems, such as equipment parameter nodes in a production system or transportation status nodes in a logistics system. They are connected to the core data through data transmission, referencing, or processing relationships.

[0095] Identify non-core data on related data nodes. Besides core data, related data nodes contain a large amount of non-core data. This non-core data may include supplementary information, intermediate calculation results, or related business data. For example, on the related node for core data such as production planning, non-core data might include temporary statistics on raw material inventory, production shift scheduling records, etc. While these are not core data, changes to them can affect the accuracy or completeness of the core data.

[0096] Analyze the update frequency of non-core data relative to core data. By recording the update time and number of updates of non-core data within a certain period, calculate the update frequency per unit time. Simultaneously, compare this with the update cycle of core data to determine the temporal relationship between non-core data updates and core data processing. If the update frequency of a certain non-core data item is once per hour, while the processing cycle of core data is once per day, it means that the non-core data may have already been updated multiple times before the core data processing.

[0097] Extract the business rules from the core data and analyze their impact on non-core data based on update frequency. Business rules are the logical specifications that core data must follow during generation, processing, and application. For example, the rule "production planning must be based on raw material inventory data" clearly defines the constraint relationship between raw material inventory data and production planning. Combined with the update frequency of non-core data, analyze whether the business rules can effectively exert their constraint effect when non-core data is updated at this frequency. If the update frequency of non-core data is much higher than the processing frequency of core data, it may lead to the business rules using outdated information from non-core data when processing core data, thus creating a risk of constraint failure.

[0098] A conflict prediction model for core data is constructed based on update frequency and constraint impact. Update frequency and constraint impact are transformed into input variables for the model, where update frequency can be quantified into specific numerical values, and constraint impact can be divided into different levels. An algorithm learns the correlation patterns between these two factors and conflict occurrences in historical data, establishing a model capable of predicting the probability of conflict occurrence based on the input variables. When new non-core data update frequency and constraint impact levels are input, the model can output the likelihood of core data conflict, providing a basis for proactive conflict response.

[0099] When setting conflict resolution rules for multi-source data based on conflict prediction models and dependency paths, the data flow of multi-source data is monitored in real time. Monitoring components deployed on data transmission channels and processing nodes continuously capture data transmission status, format changes, and content information to identify conflicting data. Conflicting data may manifest as the same data item having different values ​​in different sources, data formats not conforming to preset standards, or data loss or duplication during transmission.

[0100] The conflict prediction model analyzes the impact of conflicting data on the synergy of multi-source data. Based on the type of conflicting data, the business processes involved, and its correlation with other data, the model assesses the potential interference with data synergy, such as whether it will lead to errors in subsequent data integration, whether it will affect the normal progress of business processes, and whether it will cause biases in data-driven decision-making. This clarifies the scope and severity of the impact of conflicting data.

[0101] The system tracks the correlation logs between multi-source data and conflicting data. These logs detail the data's source system, transmission path, personnel involved, modification history, and relationships with other data. By retrieving these logs, the source of conflicting data can be traced, such as errors during initial data entry, formatting issues during system interface conversion, or inconsistencies caused by network latency. The system also clarifies the interactions between conflicting data and other data during transmission, providing a complete basis for determining the cause of the conflict.

[0102] A data correction mechanism for conflicting data should be set up based on the collaborative impact and the cause of the conflict. For conflicting data with minor collaborative impact and caused by data entry errors, an automatic system verification and correction mechanism can be used to automatically correct erroneous fields by comparing them with the standard template of the data source. For conflicting data with significant collaborative impact and caused by system interface incompatibility, a manual intervention mechanism needs to be established to work with technical personnel to adjust interface parameters and perform batch verification and correction of the generated conflicting data.

[0103] Identify the business links of multi-source data. A business link is a chain-like structure formed by the flow of data through various business stages, reflecting the correspondence between data and business processes. For example, "order data → inventory data → logistics data → receipt data" constitutes a complete sales business link. Based on the business links and data correction mechanisms, set the processing priority for conflicting data. Conflicting data located at critical nodes in the business link, which would lead to link interruption if not processed in a timely manner, has the highest priority; conflicting data located at non-critical nodes in the link and with a limited scope of impact has a relatively lower priority.

[0104] By combining dependency paths and processing priorities, conflict resolution rules for multi-source data are established. Based on the dependency paths, the nodes that conflicting data must traverse and the related data involved in the processing are clearly defined, ensuring that the processing does not interfere with the flow of other normal data. Simultaneously, conflicting data is prioritized according to processing priority, with high-priority conflicting data processed first. Processing time limits and responsible departments are specified for each priority level, forming a complete rule system covering conflict identification, cause analysis, correction methods, and processing order, ensuring that conflicting data can be processed in an orderly and efficient manner.

[0105] Example 3: Please refer to Figure 4When analyzing the flow trends of multi-source data based on metadata information, the metadata information is divided into dimensions to obtain the division results. Dimensional division can be carried out from aspects such as data generation method, business module, storage format, and update cycle. The data generation method dimension includes automatically collected data and manually entered data; the business module dimension includes data from the procurement module, production module, sales module, etc.; the storage format dimension includes structured table data, unstructured document data, and semi-structured log data; the update cycle dimension includes real-time updated data, hourly updated data, and daily updated data. Through multi-dimensional division, the different attribute characteristics of metadata information can be clearly presented.

[0106] Feature parameters of the segmentation results are extracted, covering the average transmission time of data in each dimension, the dwell time at each processing node, the frequency of data field additions and deletions, and the average daily growth rate of data volume. Based on these feature parameters, historical flow records of multi-source data are collected. These records contain full lifecycle information of data from generation to deletion over the past 12 months, such as the generation time of each data item, the time of its first entry into the system, the processing nodes it passed through, its association operations with other data, changes in storage location, and the time of final archiving or deletion.

[0107] Identify peak and trough periods in historical data transfer records. Through statistical analysis of data transmission and processing volumes in historical records, plot data transfer curves by hour, day, week, and month. Periods significantly above the average level are peak periods, and periods significantly below the average level are trough periods. For example, production module data transfer volume between 8:00 AM and 6:00 PM on weekdays is 3-5 times that of other times, constituting a peak period; while on weekends and public holidays, the transfer volume is only 1 / 10 of the weekday average, constituting a trough period.

[0108] Based on peak periods, analyze the traffic fluctuation patterns of multi-source data. Calculate the start time, peak value time, and end time of data traffic within each peak period, as well as the rate of increase and decrease of traffic during the peak period. Simultaneously observe the interval patterns and traffic variation magnitudes between adjacent peak periods. For example, sales module data experiences a peak period from 9:00 to 11:00 during the last week of each month. Traffic begins at 9:00, increasing by 15% every 10 minutes, reaching its maximum at 10:30, and then decreases by 10% every 10 minutes, returning to normal levels at 11:00. Furthermore, the monthly peak period traffic increases by an average of 8% compared to the previous month.

[0109] Based on off-peak periods, the storage occupancy rate of multi-source data is analyzed. The storage capacity, number of data blocks, and read / write frequency of various data types on different storage media (such as local hard drives, cloud storage, and distributed file systems) are statistically analyzed during off-peak periods to calculate the storage occupancy rate, which is the ratio of actual storage volume to total storage capacity at a given moment. For example, during the off-peak period of 2:00-4:00 AM daily, the storage occupancy rate of production data in the distributed file system remains stable at 45%-50%, while the storage occupancy rate of log data on local hard drives remains stable at 30%-35%.

[0110] By combining traffic fluctuation patterns, storage occupancy rates, and characteristic parameters, the flow trends of multi-source data are analyzed. This allows for a comprehensive assessment of the direction of data traffic changes at different times, the magnitude of increases or decreases in storage occupancy rates, and changes in data flow efficiency at each processing node over the next six months. For example, by combining traffic fluctuation patterns and average daily growth rates during peak sales seasons, it can be predicted that peak traffic for sales data in the next quarter will increase by 25%, storage occupancy rates will reach 70% during peak periods, and data retention time at processing nodes may extend by 10%.

[0111] When setting quality monitoring indicators for multi-source data based on data processing status and flow trends, the integrity and consistency of multi-source data are identified based on the data processing status. Data processing status includes field integrity upon data access, format standardization during processing, and data loss rate during transmission. By checking this status information, data integrity is determined, i.e., whether it contains all fields and records required for the business. Data consistency is determined by comparing the presentation of the same data at different processing nodes, i.e., whether there are content contradictions or format differences.

[0112] The analysis examines the correlation between completeness and consistency of multi-source data and their business value. For product design data, completeness directly impacts the accuracy of the production process; the absence of any design parameter can lead to production errors. Therefore, completeness is highly correlated with business value. For employee attendance data, the lack of completeness has a smaller impact on business value, resulting in a lower correlation. Regarding consistency, the consistency of inventory data is closely related to the formulation of procurement plans, exhibiting a very high correlation. However, the consistency of descriptive data in promotional materials has a relatively low correlation.

[0113] Based on the flow trends and their correlation with business value, weak links in the quality of multi-source data are identified, along with the scope of their impact. For example, by analyzing flow trends, it was found that as data volume increases, the field missing rate of logistics data rises during peak periods. Furthermore, the completeness of logistics data is highly correlated with business value; therefore, the field completeness of logistics data is a weak link in quality, affecting multiple business aspects such as logistics scheduling efficiency, customer satisfaction assessment, and transportation cost accounting.

[0114] Based on the data processing status and scope of impact, set quality thresholds for multi-source data. Quality thresholds are critical values ​​used to determine whether data quality meets standards. For completeness, field missing rate thresholds and record missing rate thresholds can be set; for consistency, data discrepancy rate thresholds can be set. For example, the field missing rate threshold for logistics data can be set to 3%, meaning the proportion of records with missing fields should not exceed 3% of the total number of records; the data discrepancy rate threshold for inventory data can be set to 2%, meaning the proportion of discrepancies between inventory data from different sources should not exceed 2% of the total. The wider the scope of impact and the higher the relevance to business value of the data, the stricter the quality thresholds should be.

[0115] Based on the scope of impact and quality thresholds, establish quality monitoring indicators for multi-source data. These indicators include specific monitoring items, measurement standards, and calculation methods, for example:

[0116] Integrity monitoring metric: Field integrity rate = (Number of complete records / Total number of records) × 100%, with a standard of not less than 97%;

[0117] Consistency monitoring metric: Data consistency rate = (number of records with no differences / total number of records compared) × 100%, with a standard of not less than 98%;

[0118] Timeliness monitoring indicator: Transmission delay rate = (number of delayed transmission records / total number of transmission records) × 100%, with a standard of not exceeding 5%.

[0119] The calculation of transmission latency involves the actual transmission time and the expected transmission time of the data, which can be expressed by the formula:

[0120]

[0121] Perform the calculation, where D r N represents the transmission delay rate. d N represents the number of records with delayed transmission, that is, the number of records whose actual transmission time is later than the expected transmission time. t This represents the total number of records transmitted, which is the total number of all data records that need to be transmitted within this time period.

[0122] These indicators are monitored at different frequencies depending on their scope of influence. Core data with a wide scope of influence is monitored every 10 minutes, while auxiliary data with a smaller scope of influence is monitored every hour, forming a comprehensive and targeted quality monitoring system.

[0123] Example 4: When identifying abnormal access behavior of multi-source data based on quality monitoring indicators, historical access records of multi-source data are collected based on these indicators. The quality monitoring indicators include information such as the data's sensitivity level and the normal access frequency range. Access records requiring special attention are filtered based on this information. For example, for R&D data marked as "highly sensitive," its historical access records must include detailed information such as the accessing user ID, access time, access operation type (e.g., viewing, downloading, modifying), and access duration. The collection period covers access data from the past 12 months to ensure a comprehensive reflection of access characteristics across different time periods.

[0124] Based on historical access records, normal access patterns for multi-source data were identified. Analysis of historical access records for R&D data revealed that users were primarily R&D department employees, access times were mainly between 9:00 AM and 6:00 PM on weekdays, access was mainly for viewing, with downloads occurring no more than 5 times per month, and each access lasting no more than 30 minutes. These characteristics collectively constitute the normal access pattern for R&D data. Normal access patterns differ for different types of data; for example, in the normal access pattern for financial data, access is limited to finance department employees, and modification operations require dual authorization.

[0125] Identify the current access status of multi-source data, including the department of the user currently accessing the R&D data, the access time, the type of operation, and the duration. Compare the current access status with historical access records. If the current user is not an employee of the R&D department, or the access time is between 2:00 AM and 5:00 AM, or a single download operation exceeds 10 times, it is considered an access behavior deviation, deviating from the normal access pattern.

[0126] Based on the normal access mode, abnormal access thresholds are set for multi-source data. For R&D data, abnormal access thresholds include: access by non-R&D department users is considered abnormal; access lasting more than 10 minutes outside of 18:00-9:00 the next weekday is considered abnormal; and more than 8 downloads in a single month is considered abnormal. Abnormal access thresholds for different data types are set separately according to their normal access modes. For financial data, any access operation by non-financial department users is considered abnormal.

[0127] Based on abnormal access thresholds and access behavior deviations, abnormal access behavior for multi-source data is identified. When a non-R&D department user accesses R&D data at 3:00 AM for 45 minutes, and has downloaded R&D data 12 times in a single month, this access behavior deviation exceeds the set abnormal access threshold and is identified as abnormal access behavior.

[0128] When setting up a security audit mechanism for multi-source data based on abnormal access behavior and conflict prediction models, risky data items in the multi-source data are identified based on abnormal access behavior. The aforementioned abnormal access behavior targets core technical parameter data from the R&D department; this data item is the risky data item, containing critical information such as the product's core formula and process parameters. Its leakage would cause serious losses.

[0129] Analyzing the threat level of the risky data items, core technical parameters were repeatedly downloaded by unauthorized users, and this data involves the company's core competitiveness, thus the threat level was determined to be "extremely high." Based on this threat level, a protection strategy layer was set up for the risky data items, including: data transmission using the AES-256 encryption algorithm; requiring two-factor authentication of facial recognition and dynamic password for access; real-time monitoring of the access process, generating an access log every 30 seconds; and requiring individual authorization from the R&D director for download operations.

[0130] Based on the conflict prediction model, the scope of the associated impact of risk data items was identified. Model analysis revealed a direct correlation between core technical parameter data and production process data, as well as material procurement data. If core technical parameter data is tampered with, it will lead to calculation deviations in production process data, resulting in inaccurate material procurement quantities. The scope of the associated impact covers the three core business departments: R&D, production, and procurement.

[0131] Based on the scope of the associated impact, isolation paths are set for risky data items. For core technical parameter data, independent physical servers are deployed for storage, network-isolated from the servers of the production and procurement systems. Data interaction is only conducted through dedicated interfaces, and interface transmissions are filtered by triple firewalls. These isolation paths ensure that the impact of risky data items in the event of abnormal access does not spread to related business systems.

[0132] A multi-source data security audit mechanism is established by combining the protection strategy layer and isolation paths. The security audit includes: hourly checks on the execution status of the protection strategy layer, such as whether the encryption algorithm is running normally and whether two-factor authentication is effectively enabled; daily checks on the network connection logs of the isolation paths to ensure no unauthorized data interactions occur; and real-time recording of all access operations for risky data items to create an immutable audit log. The audit results are processed as follows: when a protection strategy failure is detected, a system alarm is immediately triggered, notifying the security department to handle the issue within 15 minutes; when abnormal connections are detected on the isolation paths, the relevant network ports are automatically disconnected, and a data backup mechanism is initiated; the audit department conducts regular (monthly) compliance checks on the access logs and generates an audit report.

[0133] Example 5: Multi-source data is collaboratively managed based on quality monitoring indicators, security audit mechanisms, and conflict resolution rules. When the collaborative management results are obtained, quality issues in the multi-source data are identified based on the quality monitoring indicators. These indicators track the performance of data in real time during the data flow process. For example, in monitoring production data, if the completeness indicator shows that 30% of the key parameters in the quality inspection data of a batch of products are missing, or if the accuracy indicator shows that the deviation of equipment operating data from the standard value exceeds the allowable range, these are identified as clear quality issues. Different types of data may present different quality problems. Inventory data quality issues may manifest as a large discrepancy between actual inventory and system records, while sales data quality issues may manifest as incomplete customer information fields.

[0134] Based on quality issues, trigger conditions for conflict resolution rules are set, and data repair methods for multi-source data are defined. If the quality issue in production data involves missing critical quality inspection parameters, and this data will be used for subsequent production process adjustments, the conflict resolution rule is set to "immediate trigger." If the quality issue involves missing non-critical remarks, the rule is set to "delayed trigger," and data will be processed uniformly after the day's data is aggregated. For quality issues involving missing quality inspection parameters, the data repair method is to automatically send a retransmission request to the quality inspection system. If no response is received within 15 minutes, a manual reminder mechanism is activated. For issues involving deviations in equipment operation data, the repair method is to call historical data from the same period for comparison and verification, and automatically correct values ​​that exceed the reasonable range.

[0135] Based on the triggering conditions and data repair methods, conflict resolution is performed on multi-source data to obtain the processing result. When a critical quality inspection parameter in the production data is missing, triggering the "immediate trigger" condition, the system automatically sends a retransmission instruction to the quality inspection system. After receiving the retransmission data, the system performs format verification. If the data is correct, it is updated to the production database, forming a preliminary processing result. If the retransmission fails, a manual processing work order is generated, and the quality inspector manually enters the missing parameters, ultimately forming a complete processing result that includes the supplemented data and processing records.

[0136] Based on the processing results and security audit mechanisms, multi-source data undergoes collaborative management to obtain collaborative management results. The processed production data is then linked with the security audit mechanism, which checks access records during the data repair process. Once it confirms that the retransmission operation was performed by authorized personnel and no abnormal access behavior is detected, the data is allowed to proceed to the next processing stage. Simultaneously, in conjunction with conflict resolution rules, consistency verification is performed between the repaired data and associated raw material data to ensure that production data matches raw material consumption data. After the above collaborative management process, the final collaborative management results include complete production data, data repair records, a security audit report, and associated data verification results. These results will be synchronized to the production monitoring platform and data analysis system to support subsequent production decisions and business optimization.

[0137] Throughout the collaborative management process, different modules form a closed-loop linkage. Issues identified by quality monitoring indicators trigger corresponding actions through conflict resolution rules, while a security audit mechanism ensures the security of the processing. The final collaborative management result meets both data quality requirements and security standards, achieving effective end-to-end management of multi-source data. For other types of data, such as sales data, the collaborative management process is similar. Based on the type and severity of quality issues, corresponding conflict resolution rules are triggered, and the security audit mechanism ensures the compliance of data processing, ultimately resulting in a collaborative management outcome adapted to business needs.

[0138] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0139] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An enterprise multi-source data collaborative management system based on an industrial internet platform, characterized in that: include: The data access module is used to acquire multi-source data from an enterprise, determine the master data and slave data of the multi-source data, collect metadata information of the master data, and identify the data source type and data volume of the multi-source data based on the metadata information. The collaborative rules engine module is used to set a collaborative strategy for the multi-source data based on the data source type and the data volume, locate the relationship between the data in the collaborative strategy, and determine the collaborative processing flow of the multi-source data based on the metadata information and the relationship. The conflict resolution module is used to identify the core data and the dependency path of the core data in the collaborative processing flow, construct a conflict prediction model of the core data based on the dependency path, and set conflict resolution rules for the multi-source data according to the conflict prediction model and the dependency path. The quality monitoring module is used to combine the collaborative processing flow and the conflict resolution rules to perceive the data processing status of the multi-source data in real time, analyze the flow trend of the multi-source data based on the metadata information, and set the quality monitoring indicators of the multi-source data in combination with the data processing status and the flow trend. The security audit module is used to identify abnormal access behaviors of the multi-source data based on the quality monitoring indicators. Based on the abnormal access behavior and the conflict prediction model, a security audit mechanism for the multi-source data is set up. The collaborative management module is used to perform collaborative management processing on the multi-source data according to the quality monitoring indicators, the security audit mechanism and the conflict resolution rules, and obtain collaborative management results.

2. The enterprise multi-source data collaborative management system based on an industrial internet platform as described in claim 1, characterized in that, The step of determining the collaborative processing flow of the multi-source data based on the metadata information and the association relationship includes: Based on the metadata information, identify the business scenarios of the multi-source data; Based on the business scenario, schedule the real-time requirements and collaboration strategies for the multi-source data; Based on the real-time requirements and the relationships, the collaboration strategy is dynamically optimized to obtain an optimized strategy. Based on the real-time requirements and the optimization strategy, the collaborative processing flow for the multi-source data is determined.

3. The enterprise multi-source data collaborative management system based on an industrial internet platform as described in claim 1, characterized in that, The construction of the conflict prediction model for the core data based on the dependency path includes: Based on the dependency path, calculate the associated data nodes of the core data; Identify the non-core data on the associated data nodes; Analyze the update frequency of the non-core data relative to the core data; Extract the business rules of the core data, and analyze the constraint impact of the business rules on the non-core data based on the update frequency; Based on the update frequency and the impact of the constraints, a conflict prediction model for the core data is constructed.

4. The enterprise multi-source data collaborative management system based on an industrial internet platform as described in claim 1, characterized in that, The step of setting conflict resolution rules for the multi-source data based on the conflict prediction model and the dependency path includes: Real-time monitoring of the data stream from the multi-source data, and identification of conflicting data within the data stream; Based on the conflict prediction model, analyze the synergistic impact of the conflict data on the multi-source data; The association logs of the multi-source data and the conflicting data are scheduled, and the cause of the conflict of the conflicting data is identified based on the association logs; Based on the synergistic effects and the causes of the conflict, a data correction mechanism for the conflict data is set up; Identify the business links of the multi-source data, and set the processing priority of the conflicting data based on the business links and the data correction mechanism; Based on the dependency path and the processing priority, set the conflict resolution rules for the multi-source data.

5. The enterprise multi-source data collaborative management system based on an industrial internet platform as described in claim 1, characterized in that, The analysis of the flow trend of the multi-source data based on the metadata information includes: The metadata information is divided into dimensions to obtain the division results; Extract the feature parameters of the partitioning results, and collect the historical flow records of the multi-source data based on the feature parameters; Identify the peak and trough periods in the historical data flow records; Analyze the traffic fluctuation patterns of the multi-source data based on the peak periods; Based on the aforementioned low-period periods, analyze the storage occupancy rate of the multi-source data; By combining the traffic fluctuation pattern, the storage occupancy rate, and the characteristic parameters, the flow trend of the multi-source data is analyzed.

6. The enterprise multi-source data collaborative management system based on an industrial internet platform as described in claim 1, characterized in that, The step of setting quality monitoring indicators for the multi-source data by combining the data processing status and the flow trend includes: Based on the data processing status, the integrity and consistency of the multi-source data are identified; Analyze the correlation between the integrity and consistency of the data and the business value of the multi-source data; Based on the flow trend and the correlation with business value, identify the weak links in the quality of the multi-source data, and identify the scope of impact of the weak links; Based on the data processing status and the scope of influence, set the quality threshold for the multi-source data; Based on the influence range and the quality threshold, quality monitoring indicators for the multi-source data are set.

7. The enterprise multi-source data collaborative management system based on an industrial internet platform as described in claim 1, characterized in that, The step of identifying abnormal access behavior of the multi-source data based on the quality monitoring indicators includes: Based on the aforementioned quality monitoring indicators, historical access records of the multi-source data are collected. Based on the historical access records, identify the normal access patterns of the multi-source data; Identify the current access status of the multi-source data, and combine the current access status with the historical access records to identify access behavior deviations of the multi-source data; Based on the normal access mode, set the abnormal access threshold for the multi-source data; Based on the abnormal access threshold and the access behavior deviation, abnormal access behavior of the multi-source data is identified.

8. The enterprise multi-source data collaborative management system based on an industrial internet platform as described in claim 1, characterized in that, The step of setting up a security audit mechanism for the multi-source data based on the abnormal access behavior and the conflict prediction model includes: Based on the abnormal access behavior, locate the risky data items in the multi-source data; Analyze the threat level of the risk data item, and set the protection strategy layer for the risk data item based on the threat level; Based on the conflict prediction model, identify the scope of the associated impact of the risk data items; Based on the scope of the associated impact, set the isolation path for the risk data item; By combining the protection strategy layer and the isolation path, a security audit mechanism for the multi-source data is set up.

9. The enterprise multi-source data collaborative management system based on an industrial internet platform as described in claim 1, characterized in that, The step of performing collaborative management processing on the multi-source data based on the quality monitoring indicators, the security audit mechanism, and the conflict resolution rules to obtain collaborative management results includes: Based on the quality monitoring indicators, identify quality problems in the multi-source data; Based on the aforementioned quality issues, the triggering conditions for the conflict resolution rules are set, and the data repair method for the multi-source data is also set. Based on the triggering conditions and the data repair method, conflict resolution is performed on the multi-source data to obtain the processing result; Based on the processing results and the security audit mechanism, the multi-source data is processed collaboratively to obtain collaborative management results.

10. The enterprise multi-source data collaborative management system based on an industrial internet platform as described in claim 1, characterized in that, The step of setting a collaborative strategy for the multi-source data based on the data source type and the data volume includes: Analyze the business scenario requirements corresponding to the data source type; Based on the business scenario requirements, determine the real-time and accuracy requirements for data collaboration; Based on the data volume, assess the computational resource requirements for data processing; Based on the aforementioned real-time requirements, accuracy requirements, and computing resource needs, a preliminary collaborative strategy is formulated; The feasibility of the preliminary collaboration strategy is verified, the strategy parameters are adjusted and optimized, and the final collaboration strategy is set.