An enterprise project management method and system based on multi-source heterogeneous data integration

Through the multi-source heterogeneous data integration method, the problems of data silos and business process changes in enterprise project management are solved, real-time data processing and resource optimization are realized, and decision-making quality and system performance are improved.

CN120374060BActive Publication Date: 2025-09-02江西展群科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510865176.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-02
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

Traditional enterprise project management systems cannot effectively integrate multi-source heterogeneous data, resulting in serious data silos, unable to achieve real-time analysis and adaptive business process changes, unreasonable resource allocation, affecting decision-making quality and system performance.

Method used

Adopt enterprise project management methods based on multi-source heterogeneous data integration, and by receiving heterogeneous data, classifying and processing, allocating logical timestamps, realizing intelligent conversion and predictive analysis, combining a distributed stream processing framework and an adaptive data acquisition adapter, we monitor business process changes in real time and optimize resource allocation.

Benefits of technology

It realizes the logical consistency of data processing and business process adaptability, supports real-time monitoring and decision-making, eliminates data silos, optimizes resource utilization, and improves decision-making quality and system reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374060B_ABST
    Figure CN120374060B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of enterprise project management, and discloses an enterprise project management method and system based on multi-source heterogeneous data integration, wherein an enterprise project management method based on multi-source heterogeneous data integration includes: receiving heterogeneous data from multiple data sources; classifying the heterogeneous data into fast flow, medium flow and slow flow heterogeneous data flows according to update frequency; allocating logical timestamps to the heterogeneous data flows; processing the heterogeneous data flows based on the logical timestamps to ensure the logical consistency of data processing; converting, aggregating and analyzing the processed data according to a preset data flow processing logic; applying the analysis results to enterprise project management decision support; the present invention realizes the real-time, consistency and integrity of data in enterprise project management, improves the efficiency and decision-making quality of enterprise project management, and solves the problems of data islands, data inconsistency and data processing lag in traditional enterprise project management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of enterprise project management, and more particularly, to an enterprise project management method and system based on multi-source heterogeneous data integration. Background Art

[0002] As enterprises deepen their digital transformation, enterprise project management faces unprecedented data challenges. Traditional enterprise project management methods rely primarily on single data sources or multiple isolated systems. These systems lack effective data integration mechanisms, leading to the following problems:

[0003] Enterprises suffer from a serious "data silo" phenomenon. Project management systems, code management systems, continuous integration / continuous deployment tools, monitoring, and logging systems all operate independently, preventing effective data flow and sharing. This prevents managers from gaining a holistic view of projects and making accurate decisions.

[0004] The heterogeneous data generated by different systems varies significantly in structure, format, and update frequency. For example, code commit records may be updated on a minute-by-minute basis, while project plan documents may be updated on a daily or weekly basis, and monitoring system data may be updated on a second-by-second basis. These varying data streams make it difficult for traditional synchronous processing methods to ensure logical consistency.

[0005] Existing data processing methods mostly use batch processing, which cannot meet the real-time data analysis needs of enterprise project management. When project anomalies or risks arise, managers are often unable to obtain relevant information and take timely measures, which leads to the escalation of the problem.

[0006] Existing project management systems lack the ability to adapt to changes in business processes. When business processes change, data processing logic often needs to be manually adjusted, which not only increases maintenance costs but can also cause data processing to become disconnected from actual business operations.

[0007] When faced with large-scale data processing, traditional project management systems often adopt a unified processing strategy and are unable to perform differentiated processing based on the importance and timeliness of the data, resulting in the inability to reasonably allocate system resources and affecting overall performance.

[0008] Therefore, there is a need for an enterprise project management method and system that can integrate multi-source heterogeneous data, process data streams at different speeds, ensure data logical consistency, support real-time analysis, and adapt to business changes, so as to improve the efficiency of enterprise project management and the quality of decision-making. Summary of the Invention

[0009] The present invention provides an enterprise project management method and system based on multi-source heterogeneous data integration, which solves the technical problems of data time consistency, insufficient adaptability to business process changes, low real-time processing efficiency and high complexity of data source integration in the process of multi-source heterogeneous data integration in related technologies.

[0010] The present invention provides an enterprise project management method based on multi-source heterogeneous data integration, comprising the following steps:

[0011] Receive heterogeneous data from multiple data sources, including data with different structures, formats, and sources;

[0012] Adaptively classify heterogeneous data according to update frequency and data importance to form multi-dimensional heterogeneous data streams with differentiated processing priorities;

[0013] Assign logical timestamps based on business semantics to multi-dimensional heterogeneous data streams and build a temporal consistency framework;

[0014] Intelligently transform, context-aware aggregation, and predictive analysis are performed on multi-dimensional heterogeneous data streams based on logical timestamps, ensuring logical integrity across data sources based on a dynamically evolving data stream processing model.

[0015] The analysis results are applied to the multi-level decision support system of enterprise project management through a visual decision matrix.

[0016] In a preferred embodiment, an enterprise project management method based on multi-source heterogeneous data integration further includes:

[0017] Real-time monitoring of business process changes in enterprise project management;

[0018] Automatically perceive the change patterns and impact scope of business processes through machine learning algorithms;

[0019] According to the changes in business processes, the data flow processing logic is adaptively reconstructed to achieve synchronous evolution of processing logic and business changes.

[0020] In a preferred embodiment, an enterprise project management method based on multi-source heterogeneous data integration further includes:

[0021] Deploy lightweight data acquisition adapters with self-learning capabilities to connect with various data sources and extract data. The lightweight data acquisition adapters can automatically identify changes in data structure and make adaptive adjustments.

[0022] In a preferred embodiment, the heterogeneous data includes at least one of relational database data, non-relational database data, document data, log data, API interface data, Internet of Things device data, and social media data.

[0023] In a preferred embodiment, an enterprise project management method based on multi-source heterogeneous data integration is executed on an elastic and scalable distributed server cluster, which includes an edge data acquisition server, a centralized data processing server and an intelligent application server, and can automatically adjust computing resource allocation according to the data processing load.

[0024] In a preferred embodiment, an enterprise project management method based on multi-source heterogeneous data integration is built on an event-driven distributed stream processing framework, which includes a data flow engine with a fault-tolerant mechanism that can ensure the consistency and integrity of data processing when a node fails.

[0025] In a preferred embodiment, the data source includes at least one of an enterprise's internal project management system, a code management system, a continuous integration and deployment tool, a monitoring and logging system, a communication and collaboration tool, a document management system, an enterprise resource planning system, and an external market data system.

[0026] In a preferred embodiment, the allocation of logical timestamps adopts a multi-level priority algorithm, which comprehensively considers the business criticality, update frequency, data dependency and historical processing mode of the data flow to perform dynamic priority allocation.

[0027] In a preferred embodiment, an enterprise project management method based on multi-source heterogeneous data integration

[0028] In a preferred embodiment, it also includes:

[0029] Implement differentiated processing strategies for multi-dimensional and heterogeneous data streams, achieve millisecond-level response for critical business data, and provide a resource-efficient batch processing mechanism for non-critical data.

[0030] In a preferred embodiment, an enterprise project management system based on multi-source heterogeneous data integration is used to implement an enterprise project management method based on multi-source heterogeneous data integration, including:

[0031] An intelligent data receiving module for receiving heterogeneous data from multiple data sources. Heterogeneous data includes data with different structures, formats, and sources.

[0032] Adaptive data classification module, used to adaptively classify heterogeneous data according to update frequency and data importance, forming a multi-dimensional heterogeneous data stream with differentiated processing priorities;

[0033] The semantic timestamp allocation module is used to assign logical timestamps based on business semantics to multi-dimensional heterogeneous data streams and build a temporal consistency framework;

[0034] The intelligent analysis and processing module performs intelligent conversion, context-aware aggregation, and predictive analysis on multi-dimensional heterogeneous data streams based on logical timestamps, ensuring logical integrity across data sources based on a dynamically evolving data stream processing model.

[0035] The multi-level decision support module is used to apply the analysis results to the multi-level decision support system of enterprise project management through a visual decision matrix.

[0036] The beneficial effects of the present invention are:

[0037] Data processing logic consistency and business process adaptability: This invention solves the problem of data time misalignment and ensures the logical consistency of data processing by classifying heterogeneous data into multi-dimensional heterogeneous data streams and assigning logical timestamps. At the same time, it can monitor business process changes in real time and automatically adjust data processing logic, so that the system can quickly adapt to changes in business needs and maintain consistency between data analysis and actual business.

[0038] Real-time processing and system architecture optimization: The system architecture built on a distributed stream processing framework realizes real-time processing of heterogeneous data streams, supports real-time monitoring and rapid decision-making of enterprise projects, and avoids the delay problem of traditional batch processing. At the same time, the distributed server cluster architecture and priority-based logical timestamp allocation mechanism make the system scalable and fault-tolerant, ensuring the continuity and reliability of data processing.

[0039] Data source integration and elimination of enterprise data silos: Lightweight data acquisition adapters enable plug-and-play data acquisition components, significantly reducing the development and maintenance costs of integrating new data sources. This also integrates data from multiple systems within the enterprise, breaking down data silos in traditional enterprise project management and enabling data sharing and circulation, providing a data foundation for a comprehensive understanding of project status.

[0040] Resource allocation optimization and differentiated processing strategies: Based on the analysis results of multi-source heterogeneous data, it can more accurately predict project resource requirements, optimize resource allocation strategies, and improve enterprise resource utilization efficiency. By implementing differentiated processing strategies for multi-dimensional heterogeneous data streams, it can achieve millisecond-level response for critical business data, while providing a resource-efficient batch processing mechanism for non-critical data, optimizing resource utilization while ensuring system performance.

[0041] Improved decision-making quality and guaranteed system reliability: By converting, aggregating, and analyzing processed data, comprehensive, accurate, and timely decision-making support information is provided for enterprise project management, helping managers make more scientific and reasonable decisions and reduce decision-making risks. At the same time, the event-driven distributed stream processing framework adopted has a complete fault-tolerant mechanism that can ensure the consistency and integrity of data processing in the event of node failure, thereby improving the reliability and stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a flow chart of an enterprise project management method based on multi-source heterogeneous data integration of the present invention;

[0043] Figure 2 is a line graph showing changes in system resource utilization as data flow increases in the present invention;

[0044] Figure 3 It is a scatter plot of the relationship between the business process change adaptation time and the change complexity of the present invention. DETAILED DESCRIPTION

[0045] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.

[0046] At least one embodiment of the present invention discloses an enterprise project management method based on multi-source heterogeneous data integration, such as Figure 1 As shown, the following steps are included:

[0047] Step 1: Receive heterogeneous data from multiple data sources. Heterogeneous data includes data with different structures, formats, and sources.

[0048] The specific steps include:

[0049] Step 1.1, create an adapter registration center to manage various data source adapters;

[0050] The registry contains the adapter metadata repository, adapter instance manager, and adapter monitoring components.

[0051] The adapter metadata database stores configuration information of various adapters, including connection parameters, data structure descriptions, and conversion rules;

[0052] The adapter instance manager is responsible for creating, starting, stopping, and destroying adapters;

[0053] The adapter monitoring component monitors the operating status and performance indicators of each adapter in real time.

[0054] The adapter registry can be implemented using a centralized or distributed architecture.

[0055] In a centralized architecture, a single registry server manages all adapters;

[0056] In a distributed architecture, multiple registration center nodes work together to improve the reliability and scalability of the system.

[0057] In some implementations, the registration center may also support automatic discovery and registration of adapters, simplifying the system configuration and deployment process.

[0058] Step 1.2: Implement the data source adapter standard interface, including the connection interface, data extraction interface, metadata acquisition interface, and health check interface.

[0059] Among them, the connection interface defines the method of establishing and maintaining a connection with a data source;

[0060] The data extraction interface defines the method of obtaining data from the data source and supports full extraction and incremental extraction;

[0061] The metadata acquisition interface defines the method for obtaining data source structure information;

[0062] The health check interface defines methods for checking the running status of the adapter.

[0063] In addition, the design of the standard interface follows the principle of minimization to ensure that the interface is concise and easy to implement.

[0064] For example, a data extraction interface might contain the following core methods:

[0065] Used to obtain full data;

[0066] Used to obtain incremental data after a specified timestamp;

[0067] Used to obtain data based on filter conditions.

[0068] In some embodiments, additional methods may be extended such as To obtain real-time data streams, or Used to query the set of features supported by the adapter.

[0069] Step 1.3: Build adapter plug-ins for different types of data sources based on standard interfaces, including relational database adapters, API interface adapters, message queue adapters, file system adapters, and log system adapters.

[0070] Each adapter optimizes data extraction logic for a specific type of data source. For example, the relational database adapter uses transaction log parsing technology to achieve efficient incremental data extraction; the API interface adapter implements request throttling and caching mechanisms to avoid API throttling caused by frequent requests.

[0071] Therefore, this method can ensure efficient acquisition of various data sources.

[0072] Step 1.4: Build a unified data collection bus to receive the data collected by each adapter and perform preliminary classification.

[0073] The unified data collection bus is implemented using a distributed message queue, supports data partitioning and data pipeline processing, and ensures high throughput and low latency.

[0074] The bus performs time stamping, source information marking and data quality assessment on the received data, providing a basis for subsequent processing.

[0075] Step 2: Adaptively classify heterogeneous data according to update frequency and data importance to form a multi-dimensional heterogeneous data stream with differentiated processing priorities;

[0076] The specific steps include:

[0077] Step 2.1: Classify the collected data streams and divide them into fast streams according to the data update frequency. , medium speed flow and slow flow .

[0078] The classification is based on the specific update frequency range as follows:

[0079] The update cycle of fast streams is in seconds to minutes, such as system monitoring data and real-time logs;

[0080] The update cycle of medium-speed traffic is from hours to days, such as task status updates and code submission records;

[0081] The update cycle of slow streams is weekly to monthly, such as financial data and performance evaluation data.

[0082] In some implementations, data stream classification may employ an adaptive threshold method to dynamically adjust the classification threshold based on historical updated statistical information of the data stream.

[0083] For example, the system can calculate the average update frequency of each data stream over different time periods and automatically classify the data stream into the most appropriate category based on its business importance and access patterns. This approach allows the system to adapt to seasonal changes or long-term trends in data stream update patterns.

[0084] Step 2.2: Create a dependency graph between data flows ,in, Represents a collection of data streams, Indicates the dependencies between data flows.

[0085] For data flows with dependencies , the strength weight of the system maintenance dependency and time sensitivity , used for subsequent logical timestamp calculation.

[0086] The dependency graph can be constructed by combining static definition and dynamic discovery.

[0087] Statically defined dependencies are based on pre-configured business rules, which explicitly specify the dependencies between data flows.

[0088] Dynamic dependency discovery automatically identifies potential dependencies by analyzing access patterns and correlations between data flows.

[0089] For example, the system can apply association rule mining algorithms to analyze the temporal relationships of data flows in business processing, discover frequently co-occurring data flows and establish dependency relationships.

[0090] Step 2.3, apply the logical timestamp algorithm to assign a logical timestamp to each data item.

[0091] The calculation formula for the logical timestamp is:

[0092] ;

[0093] in, The logical timestamp of a data item is a unified time identifier assigned by the system to ensure the logical consistency of data at different rates. Indicates the physical collection time of the data item, that is, the time when the data is actually acquired, the actual timestamp recorded by the system; A logical timestamp set representing other data items that have a dependency relationship with this data item, used to reflect the temporal dependency between data items; Represents the logical timestamp calculation function, which is used to comprehensively consider physical time and dependencies to generate the final logical timestamp value.

[0094] The specific calculation is as follows:

[0095]

[0096] in, is the physical time weight factor, with a value range of [0, 1], and is dynamically adjusted according to the data stream type. When it is close to 1, the logical timestamp is more dependent on the physical acquisition time; when When it is close to 0, the logical timestamp is more dependent on the timestamps of related data items; The physical collection time of the data item, that is, the time when the data is actually obtained; A logical timestamp set of other data items that have a dependency relationship with this data item; The calculated logical timestamp of the data item is used to ensure the logical consistency of data at different rates; It is a timestamp-dependent comprehensive function used to calculate the weighted comprehensive value of the timestamp-dependent data items. It is calculated as:

[0097] ;

[0098] in, A comprehensive function representing dependent timestamps, used to calculate the weighted sum of timestamps of all dependent data items; Indicates the sum of all data flow pairs with dependencies. Representing data flow To the data stream dependencies; Representing data flow arrive The dependency strength weight is used to quantify the importance of the dependency; Representing data flow arrive Time sensitivity, which is used to measure the sensitivity of dependencies to time changes; Indicates dependent data flow The logical timestamp is the logical time value of the upstream data item on which the current data item depends.

[0099] Step 2.4: Build a data stream buffer management mechanism to coordinate the processing of data streams with different rates.

[0100] For fast streaming ,adopting a sliding window processing mechanism to calculate the logical timestamp of the latest data in real time;

[0101] For medium-speed flow ,using time partitioned buffers to aggregate data according to the configured time period;

[0102] For slow flow , adopts a combination of long-term storage and incremental updates to ensure data consistency and availability.

[0103] In the specific application scenario of enterprise project management, the implementation of the adaptive processing algorithm for heterogeneous data streams is as follows:

[0104] Taking software development project management as an example, the system processes data sources from different rates simultaneously:

[0105] Fast streams include system performance monitoring data (updated every second) and development environment status data (updated every minute);

[0106] Medium-speed streams include task status updates (updated every hour or every day) and code submission records (updated irregularly);

[0107] The slow stream includes weekly report data and monthly performance evaluation data.

[0108] The logical timestamp algorithm assigns reasonable physical time weight factors to data of different rates according to the needs of project management. :

[0109] For real-time monitoring scenarios, fast streaming Values ​​close to 1 give priority to physical time;

[0110] For progress tracking scenarios, medium-speed streams The value is about 0.6, which balances physical time and dependencies;

[0111] For performance evaluation scenarios, slow flow A value of about 0.3 takes dependencies into account more.

[0112] During project milestone reviews, the system needs to integrate data from different rates for analysis. Using a logical timestamp algorithm, the system ensures logical consistency between real-time monitoring data (fast stream), daily task completion status (medium stream), and monthly plan completion status (slow stream), avoiding analytical bias caused by differences in data update frequency. For example, when evaluating the completion quality of a development task, the system links the task's status update (medium stream) with related performance monitoring data (fast stream) and quality assessment reports (slow stream) using logical timestamps, providing a comprehensive assessment view.

[0113] Step 3: Assign logical timestamps based on business semantics to multi-dimensional heterogeneous data streams and build a temporal consistency framework;

[0114] The specific steps include:

[0115] Step 3.1: Build a unified data model to standardize heterogeneous data from different sources;

[0116] The unified data model is represented by a graph structure, which includes two basic elements: entity nodes and relationship edges.

[0117] Entity nodes represent business objects (such as projects, tasks, resources, etc.), and relationship edges represent associations between entities (such as subordination, dependency, etc.).

[0118] Each entity node and relationship edge contains a set of attributes for storing specific business data.

[0119] The unified data model can be implemented using different structures. In addition to the graph structure mentioned above, you can also use a hierarchical structure (tree model), a relational structure (tabular model), or a hybrid structure. Choose the most appropriate model structure based on specific business needs and data characteristics.

[0120] For example, a hierarchical structure can be used for organizational structure data with clear hierarchical relationships; a relational structure can be used for transaction-centric financial data.

[0121] In some implementations, the unified data model may also support dynamic extension of the model, allowing new entity types and relationship types to be added during system operation.

[0122] Step 3.2, implement the incremental data stream conversion algorithm, which dynamically selects the optimal conversion strategy according to the nature of data changes.

[0123] The incremental data stream conversion algorithm includes the following computational steps:

[0124] For the input data stream , calculate its difference with the last processed data stream A collection of changes:

[0125] ;

[0126] in, Represents the data stream input at the current moment, including the latest data set; Represents the historical data stream processed last time, serving as a comparison benchmark; Represents the change set, that is, the difference between the current data stream and the historical data stream; Represents a set difference operation, which is used to extract data items in the current data stream that are different from the historical data stream.

[0127] Change Collection Classify and distinguish between new data , update data and delete data ;

[0128] Apply transformation operations to different types of changing data:

[0129] For new data , applying the full conversion function ;

[0130] For updating data , applying the difference transformation function ;

[0131] For deleting data , applying the inverse transformation function ;

[0132] Merge the conversion results with the existing unified data model to obtain the updated data model:

[0133] ;

[0134] in, Represents the updated data model, that is, the latest unified data model after the conversion process is completed; Indicates the data model before the update, that is, the unified data model of the previous version; Represents a model merge operation, which is used to integrate the transformed change data with the existing model; Represents the combined results of various transformation functions, including the comprehensive results of the transformation functions applied to newly added data, updated data, and deleted data respectively; Represents a change set, that is, the difference between the current data stream and the historical data stream.

[0135] In scenarios where data changes frequently but in small amounts, the following optimized variant algorithms can be used:

[0136] To improve the efficiency of processing small-scale frequent changes, the system can implement a change batch processing mechanism. This mechanism collects multiple small-scale changes in a short period of time to form a change batch. , and then merge:

[0137] ;

[0138] in, Represents a merged batch of changes, which is the result of merging multiple small changes; The merge operator represents a change set, which is used to combine multiple change sets into one set; Indicates starting from the first changeset; Indicates the total number of change sets; Indicates the A change set, representing a single data change; Represents a change batch, which is a collection of multiple small-scale changes; 、 、 Respectively represent the first 、 、 independent change sets; Indicates the number of changes included in the batch.

[0139] Merged changeset Then process it according to the above standard algorithm, which significantly reduces the number of processing times and system overhead.

[0140] Step 3.3, build a data transformation rule engine that supports declarative rule definition and automatic rule inference.

[0141] The rule engine includes a rule repository, a rule executor, and a rule learning module;

[0142] The rule warehouse stores predefined transformation rules and rule templates;

[0143] The rule executor is responsible for interpreting and executing the transformation rules;

[0144] The rule learning module automatically derives new conversion rules based on historical conversion data to improve system adaptability.

[0145] Step 3.4: Establish a data quality assessment and repair mechanism to monitor data quality issues during the conversion process in real time and automatically repair them;

[0146] Data quality assessment includes completeness check, consistency check, accuracy check and timeliness check.

[0147] When data quality issues are discovered, the system automatically repairs them according to the preset repair strategy and records the repair log for subsequent analysis and improvement.

[0148] In the actual application scenario of enterprise project management, the implementation and application examples of the incremental data stream conversion algorithm are as follows:

[0149] For example, in a multi-system integrated enterprise project management environment, an enterprise uses multiple systems simultaneously, including project management systems, code management systems, CI / CD systems, and human resources systems. When processing data from these heterogeneous systems, the incremental data stream transformation algorithm adopts different strategies based on the nature of the data changes:

[0150] When the task status in the project management system changes from "In Progress" to "Completed" (which is an update data ), the system only processes the state change part and applies the difference conversion function This incremental processing significantly reduces computation and data transmission compared to processing all task data in full.

[0151] When a new code repository is added (belonging to the new data ), the system applies the complete conversion function , converting the complete metadata of the warehouse and the initial code analysis results into new nodes and relationships in the unified data model. In this case, since it is brand new data, the system needs to perform a complete conversion process.

[0152] When a project member leaves (delete data ), the system applies the inverse conversion function , safely removing the member's association with the project from the unified data model while retaining historical contribution records. This approach ensures data integrity and consistency.

[0153] Step 4: Intelligently transform, context-aware aggregation, and predictive analysis are performed on multi-dimensional heterogeneous data streams based on logical timestamps, ensuring logical integrity across data sources based on a dynamically evolving data stream processing model.

[0154] The specific steps include:

[0155] Step 4.1: Build a business process monitoring system to collect and analyze business activity logs in real time;

[0156] The system includes a log collector, a process identification engine, and a change detector.

[0157] The log collector collects activity logs from various business systems;

[0158] The process identification engine builds the current business process model based on the collected logs;

[0159] The change detector detects changes in business processes by comparing process models at different points in time.

[0160] The detailed composition of the business process monitoring system is as follows:

[0161] The log collector adopts a distributed collection architecture, including log agents, log transmission channels, and log aggregation services.

[0162] The log agent is deployed on each business system and is responsible for the collection and preliminary processing of local logs;

[0163] The log transmission channel is implemented based on the message queue to ensure the reliable transmission of log data;

[0164] The log aggregation service receives all log data and performs unified storage and indexing.

[0165] The process identification engine is based on process mining technology and extracts process models from business activity logs.

[0166] The engine includes a pre-processing module, a process discovery module and a process optimization module.

[0167] The pre-processing module cleans, filters and converts the original logs;

[0168] The process discovery module applies the improved Alpha algorithm to construct a process model from the activity sequence;

[0169] The process optimization module improves the process model by merging similar paths, removing noise and calculating frequency information.

[0170] The change detector identifies changes to business processes by calculating similarities and differences between process models. The detector includes a model comparison algorithm, a change classifier, and a change impact analyzer.

[0171] Model comparison algorithms calculate structural and attribute differences between two process models;

[0172] The change classifier classifies the detected changes into types such as added nodes, deleted nodes, and path changes;

[0173] The Change Impact Analyzer assesses the potential impact of a change on data stream processing.

[0174] Step 4.2: Establish a mapping relationship between business processes and data flow processing;

[0175] The mapping relationship is expressed by the function:

[0176] ;

[0177] in, Represents the data flow processing configuration, which is a complete set of rules for how the system processes, transforms, and routes data; Represents a business process model, including a structured representation of business activity nodes, decision points, flow paths, and their relationships; Represents the characteristics of the data source, including the data format, update frequency, reliability, integrity and other parameter sets that describe the data source attributes; Represents context parameters, including system load, network conditions, user preferences, security policies, and other environmental factors that affect data stream processing.

[0178] The specific mapping calculation method is:

[0179] ;

[0180] in, Represents data flow processing configuration; Represents a set of data routing rules that define how data flows through the system, including rules such as the data source, destination, transmission path, and conditional branches; Represents a set of data conversion rules, which define how to convert raw data into a standard format, including rules for field mapping, format conversion, data cleaning, and validation; Represents a set of data aggregation rules, defining how to merge and aggregate related data, including rules for data grouping, summary calculation, time window aggregation, and associated merging; Represents data processing priority rules, defining the processing order of different data flows, including a priority allocation mechanism based on business importance, timeliness, resource consumption, and dependencies.

[0181] Each rule set is represented by a business process model The nodes and edges in the ,map are generated, for example, the decision points in the business ,process are mapped to data routing rules, and the business activities are ,mapped to data transformation rules.

[0182] Step 4.3: Implement the adaptive data flow reconfiguration mechanism. When a business process change is detected, the system performs data flow reconfiguration according to the following steps:

[0183] Identify changed business process nodes and relationships, and calculate the change set ;

[0184] Calculate the data stream processing configuration that needs to be updated based on the mapping function:

[0185] ;

[0186] in, Indicates the data flow processing configuration that needs to be updated, that is, the set of data processing rules that need to be adjusted due to business process changes; Represents a mapping function, which is used to convert business process changes into corresponding data flow processing configuration changes; Represents a set of changes to a business process, including newly added, modified, or deleted business process nodes and relationships; Represents the characteristics of the data source, including the data format, update frequency, reliability, integrity and other parameter sets that describe the data source attributes; Represents context parameters, including system load, network conditions, user preferences, security policies, and other environmental factors that affect data stream processing.

[0187] Generate a dataflow reconfiguration plan, including configuration update sequences and verification test cases;

[0188] Execute dataflow reconfiguration as planned and monitor the status and impact of the reconfiguration process to ensure system stability.

[0189] Step 4.4: Build a visual process data association monitoring tool to provide managers with a visual view of business processes and data flow processing, and support manual intervention and adjustment.

[0190] The tool displays the current business process model, data flow processing configuration, and the mapping relationship between them. It also provides historical change records and performance indicator analysis to help managers understand the system operation status and optimization direction.

[0191] The following are examples of how business process-aware data stream processing technology can be applied during enterprise organizational transformation:

[0192] A manufacturing company transformed from a functional organization to a matrix organization, which resulted in the project approval process changing from the original step-by-step approval to a parallel approval model.

[0193] By analyzing recent activity logs, the business process monitoring system detected a significant change in the approval path: the original serial path of "project manager → department manager → director → vice president" was transformed into a parallel hybrid path of "project manager → (department manager, functional expert) → comprehensive review committee".

[0194] The change detector identifies this significant process change and triggers the data flow reconfiguration mechanism. Based on the mapping relationship, the system calculates the updated data flow processing configuration: data routing rules are changed from serial forwarding to parallel distribution, data aggregation rules are changed from hierarchical aggregation to a comprehensive scoring mechanism, and processing priority is changed from layer-based to time-based.

[0195] The dataflow reconfiguration plan is executed in two phases:

[0196] In the first phase, a new approval path and data processing logic are established but not activated yet;

[0197] In the second phase, the new process will be activated, and the old process will be retained to handle ongoing approval projects, while all new projects will follow the new process.

[0198] Through this smooth transition, project management data processing during enterprise organizational changes maintains continuity and consistency, avoiding data processing confusion and business interruption during changes in traditional systems.

[0199] Step 5: Apply the analysis results to the multi-level decision support system for enterprise project management through a visual decision matrix;

[0200] The specific steps include:

[0201] Step 5.1, design the process data consistency model and define the consistency constraints between the business process state and the data flow processing state.

[0202] The consistency model is expressed by the following formula:

[0203] ;

[0204] in, Represents the consistency constraint function between the business process state and the data stream processing state; Represents the state of the business process at time point t, that is, a complete description of the configuration, rules, and execution of all business processes in the system at a specific time point t; Represents the data flow processing state at time point t, that is, a complete description of the routing rules, conversion logic, processing priority and execution status of all data flows in the system at a specific time point t; Represents the set of nodes in the system, including all servers, application instances, and processing units involved in data processing and business process execution; Representation node A view of state X, i.e., a node The information perceived about state X, including the node All properties and parameters of state X are accessible; Indicates an equivalence relationship, indicating that the two views are semantically consistent, ensuring that nodes maintain a synchronized understanding of business processes and data flows; Indicates that it is true for all nodes in the system, ensuring global consistency

[0205] This consistency model ensures that each node in the system has an equivalent view of the business process state and data flow processing state, avoiding data processing errors caused by inconsistent states.

[0206] Step 5.2: Implement a distributed consensus algorithm to reach a global consensus when business processes change.

[0207] The distributed consensus algorithm is based on the two-phase commit protocol extension and includes the following steps:

[0208] The process change coordinator sends a pre-commit message to all participating nodes, including a description of the business process change. and dataflow reconfiguration plans ;

[0209] Each node verifies the feasibility of the change and returns a ready or rejected response to the coordinator;

[0210] If all nodes return ready, the coordinator sends a commit message, otherwise it sends an abort message;

[0211] Finally, each node executes or abandons the change according to the coordinator's decision and returns the execution result.

[0212] Step 5.3: Build a business process version management system to support the parallel operation of new and old business processes;

[0213] In addition, the version management system includes a process version repository, a version router, and a version compatibility checker.

[0214] The process version library stores different versions of business process models and corresponding data flow processing configurations;

[0215] The version router directs the data flow to the appropriate version of the processing logic based on data characteristics and context information;

[0216] The version compatibility checker ensures that data exchange between different versions does not result in data loss or errors.

[0217] Step 5.4: Build a process data switching manager to coordinate the smooth transition between the old and new business processes.

[0218] The Switch Manager implements the following functions:

[0219] Data migration planning: Calculate the data mapping relationship between the new and old processes and generate a data migration plan;

[0220] State synchronization control: Maintain state synchronization between the old and new processes during the switching process to ensure data consistency;

[0221] Rollback mechanism: If an exception occurs during the switching process, it can safely roll back to the previous stable state;

[0222] Switching progress monitoring: Real-time monitoring of the progress and status of the switching process, providing visual display and alarm functions.

[0223] The specific implementation and details of the process data synchronization protocol are as follows:

[0224] The distributed consensus algorithm is based on an improved version of the two-phase commit protocol, which adds a pre-verification phase and a dynamic participant management mechanism.

[0225] The pre-verification phase involves comprehensive testing of configuration changes before formal submission to ensure the feasibility of the changes.

[0226] The dynamic participant management mechanism allows for the dynamic addition or reduction of participating nodes during the consensus process, improving the flexibility and fault tolerance of the system.

[0227] The algorithm decides the acceptance or rejection of configuration changes through a voting mechanism, using a weighted majority principle where key nodes have higher voting weights.

[0228] The consistency model is implemented using a hierarchical state synchronization mechanism:

[0229] ;

[0230] in, Represents the consistency constraint function between the business process state and the data stream processing state implemented using the hierarchical state synchronization mechanism; Represents the state of the business process at time point t, including the configuration, rules, and execution status of all business processes; Represents the data flow processing status at time point t, including the routing rules, conversion logic and processing status of all data flows; Represents the set of all nodes in the system, including all servers and application instances involved in processing; Representation node A view or understanding of the state of a business process; Representation node A view or understanding of the state of data flow processing; Represents an equivalence relation, ensuring that the two views are semantically consistent; Indicates that this equivalence relation must be satisfied for all nodes in the system.

[0231] The implementation of this model uses a hierarchical state synchronization mechanism:

[0232] At the bottom layer, a vector clock is used to track the status updates of each node;

[0233] At the middle layer, event tracing technology is used to record all state change events and support state reconstruction and backtracking;

[0234] At the upper level, the visual Figure 1 Consistency checks ensure that all nodes have equivalent views of the business process state and data flow processing state.

[0235] The specific implementation of the business process version management system includes: the process version library uses a graph database to store different versions of process models, supporting version branching, merging, and comparison;

[0236] The version router is implemented based on a context-aware rule engine and dynamically selects the processing version based on factors such as data characteristics, processing stage, and system load;

[0237] The version compatibility checker evaluates the data exchange compatibility between different versions through semantic equivalence analysis and data flow simulation testing.

[0238] In a multi-site collaborative project management scenario, the following are examples of how the process data synchronization protocol can be used:

[0239] A multinational company's R&D project involved collaboration across three R&D centers in Asia, Europe, and North America. Each center had its own project management processes and data processing logic. The company decided to unify its global R&D processes and maintain business continuity during the transition.

[0240] Applications of the process data synchronization protocol in this scenario include:

[0241] Nodes in various locations reach consensus on the new global unified process through a distributed consensus algorithm, and also determine the length of the transition period and the switching strategy;

[0242] The business process version management system configures process variants that adapt to local characteristics for each region while maintaining compatibility with global standard processes;

[0243] The Process Data Switch Manager develops a personalized transition plan for each region and determines the best switching time based on the project cycle and business characteristics of each region.

[0244] During the implementation of new processes at the Asia R&D center, a critical ongoing project needed to continue using the old process until completion. A process data synchronization protocol ensured that data processing for the project continued normally under the old process. Furthermore, a state synchronization control mechanism ensured consistency in project status data between the old and new systems, ensuring data continuity and consistency during the business process change.

[0245] Through the implementation of the above five main steps, the technical solution of this application realizes an enterprise project management method based on multi-source heterogeneous data integration, effectively solving technical problems faced by multi-source heterogeneous data integration in enterprise project management, such as data time consistency, adaptability to business process changes, real-time processing efficiency and data source integration complexity.

[0246] Application examples of this implementation:

[0247] The following example demonstrates the practical application and effectiveness of this method through the transformation of a project management system at a multinational software development company. This company manages multiple software product lines across multiple departments, including R&D, testing, and operations, using multiple information systems and facing the typical challenge of integrating heterogeneous data from multiple sources.

[0248] Application scenario description:

[0249] The company has six R&D centers around the world, running 35 software product projects simultaneously, and employs approximately 2,000 people. The company faces the following specific challenges in project management:

[0250] Data source diversity: The company used 12 different information systems, including the JIRA task management system, the GitLab code management platform, the Jenkins CI / CD system, the SonarQube code quality analysis tool, the Prometheus monitoring system, an employee attendance system, a financial system, and a customer feedback system. These systems employed different data structures and formats, with a variety of database types (MySQL, MongoDB, PostgreSQL, etc.) and data exchange methods (APIs, message queues, file exports, etc.).

[0251] Data update frequencies vary significantly: system monitoring data is updated every 10 seconds, code submission records are generated sporadically (on average dozens of times per hour), task status is updated several times a day, and financial data and personnel assessment data are updated weekly or monthly. Traditional unified batch processing methods either delay the processing of real-time data or waste resources.

[0252] Frequent business process adjustments: Companies adjust their product development strategies quarterly based on market demand and align with project management processes, such as approvals and resource allocation. These changes require manual adjustments to data processing logic, causing project management systems to frequently lag behind actual business processes by one to two weeks, leading to inconsistent data and biased decision-making.

[0253] System performance bottleneck: As projects and data volumes grew, the existing batch data integration solution could no longer meet performance requirements. Data update delays averaged four hours, with peak periods exceeding 12 hours, severely impacting real-time decision-making and anomaly warning capabilities.

[0254] Enterprises hope to build a unified project management platform that can integrate all heterogeneous data sources in real time, automatically adapt to business process changes, and provide a comprehensive and accurate view of project status and decision support.

[0255] Core step implementation example:

[0256] Construction of lightweight multi-source data acquisition adapter system:

[0257] The company first built an adapter registration center, deployed across six R&D centers using a distributed architecture, and set up a master center as a coordination node. The registration center is configured with the following key components:

[0258] Adapter Metadata Repository: Uses a distributed key-value storage system to record the configuration information of each adapter, including:

[0259] Adapter ID: a unique identifier, such as "AD-2023-JIRA-001";

[0260] Connection parameters: data source access credentials, access address, access method (API / database / message queue);

[0261] Data structure description: field mapping relationship, data type conversion rules;

[0262] Collection frequency configuration: collection interval settings that are dynamically adjusted according to data update frequency;

[0263] Standard interface implementation: Based on the abstract factory pattern, a general adapter interface is designed, covering four core functions:

[0264] Connection management: establishing, maintaining, and releasing connections with data sources;

[0265] Data extraction: supports full extraction, incremental extraction, and change capture;

[0266] Metadata acquisition: Automatically discover and map data source structures;

[0267] Health monitoring: self-diagnosis and recovery capabilities;

[0268] Adapter plug-in development: Dedicated adapter plug-ins have been developed for 12 information systems, each optimized for the characteristics of a specific data source:

[0269] For JIRA and GitLab systems, the adapter implements API rate limiting and response caching to effectively address API usage restrictions;

[0270] For relational databases such as MySQL, the adapter uses binlog parsing technology to achieve zero-latency change capture;

[0271] For analysis tools such as SonarQube, the adapter implements an incremental data extraction algorithm to only obtain new or changed analysis results;

[0272] For financial and human resources systems, the adapter integrates file monitoring and automatic parsing capabilities to process regularly generated spreadsheet reports;

[0273] During the deployment, the company also developed an adapter development kit (ADK) to significantly simplify the development of new adapters. Using this toolkit, a developer can complete adapter development for a new system in an average of four hours, compared to an average of three to five days using traditional integration methods.

[0274] All adapters send the collected data to a unified data collection bus, which is implemented based on a distributed message queue and configured as a multi-level topic structure to achieve preliminary classification of the data.

[0275] Automatically add metadata tags when data enters the bus, including:

[0276] Data source identifier: the system that records the source of the data;

[0277] Physical timestamp: records the exact time when the data is collected;

[0278] Collection node information: records the server node that performs the collection;

[0279] Data quality indicators: preliminary assessment results of record completeness and accuracy.

[0280] Application of adaptive processing algorithm for heterogeneous data streams:

[0281] Based on the frequency of data updates, enterprises divide the collected data streams into three levels:

[0282] Fast Stream Configuration:

[0283] Covered data: system monitoring data, user activity logs, and build pipeline status;

[0284] Update frequency: 10 seconds to 5 minutes;

[0285] Processing strategy: real-time processing, with a sliding window size of 30 seconds;

[0286] Physical time weight ( ): 0.85, giving priority to physical time;

[0287] medium speed flow Configuration:

[0288] Covered data: task status changes, code submission records, and quality analysis results;

[0289] Update frequency: 1 hour to 1 day;

[0290] Processing strategy: quasi-real-time processing, with time partitioning of 1 hour;

[0291] Physical time weight ( ): 0.6, balance physical time and dependencies;

[0292] Slow Flow Configuration:

[0293] Data covered: weekly report data, financial data, personnel evaluation data;

[0294] Update frequency: 1 week to 1 month;

[0295] Processing strategy: Process regularly to maintain long-term data consistency;

[0296] Physical time weight ( ): 0.3, prioritize dependencies;

[0297] The company conducted data flow dependency analysis on each product line project and constructed a dependency graph The following is an example of some of the dependencies for core product A:

[0298] Dependency 1: Task Completion Status →Build Status :

[0299] Dependency intensity weight ;

[0300] Time sensitivity ;

[0301] Indicates that the build process should be triggered after the task is completed;

[0302] Dependency 2: Code Submission →Quality Analysis :

[0303] Dependency intensity weight ;

[0304] Time sensitivity ;

[0305] Indicates that quality analysis should be performed after the code is submitted;

[0306] Dependency 3: Build Status →Deployment Status :

[0307] Dependency intensity weight ;

[0308] Time sensitivity ;

[0309] Indicates that the deployment process should be triggered after the build is completed;

[0310] Based on these configurations and dependency graphs, the system calculates a logical timestamp for each data item. Taking a code commit event as an example, the calculation process is as follows:

[0311] Physical collection time:

[0312] ;

[0313] Logical timestamps of dependent data:

[0314] ;

[0315] According to the formula calculate:

[0316] (Medium-speed flow weight);

[0317] (Rely on timestamp synthesis results);

[0318] Final logical timestamp ;

[0319] This logical timestamp mechanism enables the system to coordinate data processing at different rates. For example, when analyzing the release status of a product version in a particular week, the system ensures that all data used (from real-time monitoring data to weekly reports) is consistent in logical time, avoiding inconsistent analysis results caused by varying data update frequencies in traditional systems.

[0320] In actual applications, enterprises configure corresponding processing mechanisms for three types of data flows:

[0321] Fast Stream uses a memory-based stream processing engine with a sliding window mechanism. The window size is dynamically adjusted based on the characteristics of the data stream, ranging from 10 seconds to 5 minutes.

[0322] Medium-speed streams use a near-real-time batch processing engine with a micro-batch processing mechanism and a batch interval of 15 minutes.

[0323] Slow streams use periodic batch processing, combined with an incremental update mechanism, with a processing cycle set to once a day;

[0324] This hierarchical processing mechanism significantly optimizes system resource utilization while ensuring data logical consistency, avoiding resource waste or reduced real-time performance caused by applying a unified processing cycle to all data.

[0325] Application of incremental data stream conversion algorithm:

[0326] The enterprise has built a unified data model based on a graph structure, which includes the following core entities and relationships:

[0327] Core entity nodes:

[0328] Project: represents a product development project;

[0329] Task: represents development tasks and work items;

[0330] Member: represents the project team members;

[0331] Code: represents source code and warehouse;

[0332] Build: represents the CI / CD build task;

[0333] Resource: represents project resources;

[0334] Relationship Edge:

[0335] Contains: The project contains tasks;

[0336] Responsible: Members are responsible for tasks;

[0337] Created: member creation code;

[0338] Triggers: Code triggers the build;

[0339] Uses: Task uses resources;

[0340] The system implements dedicated conversion rules for each data source, such as the JIRA task data conversion rules:

[0341] Field mapping: JIRA task ID → unified model task ID;

[0342] Status transition: JIRA status "InProgress" → unified model status "In Progress";

[0343] Relationship mapping: JIRA task assignment relationship → unified model "responsible" relationship;

[0344] The company implemented an incremental data stream transformation algorithm. The following is an example of handling development task changes in a real case:

[0345] Data change detection: The system detects changes to the task "TASK-1024" in JIRA, including a change in status from "To be developed" to "In progress" and a change in assignee from Team Member A to Team Member B.

[0346] Change set calculation: The system calculates the change set :

[0347] "ChangeSet={

[0348] Updated items: [

[0349] {field: "status", old value: "Under Development", new value: "In Progress"},

[0350] {field: "assignee", old value: "memberA", new value: "memberB"} ]

[0352] }”

[0353] Change classification: The system classifies the change set as updated data (no additions or deletions);

[0354] Differential conversion application: For status changes and assignee changes, the system applies conversion functions only to the changed fields, and does not process unchanged fields such as description and priority.

[0355] "Difference conversion result = {

[0356] Entity Update: {

[0357] Entity Type: "Task",

[0358] Entity ID: "UNIFIED-TASK-1024",

[0359] Update properties: {

[0360] Current Status: In Progress,

[0361] "Status change time": "2023-10-16T09:30:45"

[0362] }

[0363] },

[0364] Relationship Update: {

[0365] Delete a relationship:

[0366] Type: "responsible",

[0367] Source node: "UNIFIED-MEMBER-A",

[0368] Target node: "UNIFIED-TASK-1024"

[0369] },

[0370] New relationship: {

[0371] Type: "responsible",

[0372] Source node: "UNIFIED-MEMBER-B",

[0373] Target node: "UNIFIED-TASK-1024"

[0374] }

[0375] }

[0376] }”

[0377] Data model merging: The system merges the conversion results into a unified data model and updates the task node attributes and related relationships;

[0378] In practice, the system uses a batch processing optimization strategy for frequent changes. For example, during a code review, developers may update the same code file multiple times within a short period of time. The system sets a 5-second change batch processing window, combining multiple updates to the same file within this period into a single process.

[0379] The company has also implemented a declarative transformation rule engine, which enables business personnel to define and modify transformation rules through a visual interface without writing code. The rule engine supports the following functions:

[0380] Rule template library: pre-defined rule templates for common data conversion scenarios, such as task status mapping, personnel role mapping, etc.

[0381] Rule verification: Automatically verify the integrity and consistency of rules and identify potential conflicts;

[0382] Rule version management: Track rule change history and support rollback to previous versions;

[0383] Rule learning: Automatically recommend rule optimization suggestions based on historical data;

[0384] In addition, the company has established a data quality assessment mechanism to monitor data quality issues during the conversion process in real time and trigger automatic remediation. For example, if the system detects that a team member associated with a task does not exist, it automatically marks the data as "pending verification" and generates remediation suggestions. The system also maintains a detailed log of the conversion process, recording the inputs, outputs, and quality indicators of each conversion, supporting problem identification and quality optimization.

[0385] Business process-aware data stream processing technology implementation:

[0386] The company has built a business process monitoring system to track the execution and changes of project management processes in real time. The system includes the following core components:

[0387] Log Collector: A lightweight agent deployed in each business system that captures the following key activity data:

[0388] User operation log: records the user's operation sequence in the project management system;

[0389] System event log: records status changes and process advancements automatically triggered by the system;

[0390] API call log: records the interface calls between systems;

[0391] Database transaction log: records changes in business data;

[0392] Log collection adopts a low-intrusive design and is implemented through interceptors, listeners, and log parsers, with an impact on the original system performance of less than 3%.

[0393] Process Identification Engine: This engine applies a modified Alpha++ algorithm to extract business process models from activity logs. The algorithm consists of the following steps:

[0394] Log preprocessing: clean abnormal data and convert raw events into standard activities;

[0395] Relationship discovery: computing causal and parallel relationships between activities;

[0396] Control flow construction: Build process models that include sequential, selection, parallel and loop structures;

[0397] Decision point identification: key decision points and decision rules in the detection process;

[0398] The process identification engine performs a full analysis every 24 hours and an incremental analysis every 4 hours to maintain the latest process model.

[0399] Change Detector: Detects changes to business processes by comparing process models at different points in time. The following example shows a detected change in the product release approval process:

[0400] Original process path:

[0401] Product manager submits release application → Technical manager approves → Quality manager approves → Department director approves → Release;

[0402] New process path:

[0403] Product manager submits release application → (approved by technical manager and quality manager) [parallel] → approval by release committee → release;

[0404] The change detection algorithm identifies the following change points:

[0405] Newly added node: Release Committee Approval;

[0406] Deletion node: Approval by department director;

[0407] Path change: the approval process of technical managers and quality managers changes from serial to parallel;

[0408] Changes to decision-making rules: The approval requirement has changed from "approval by all approvers" to "approval by a majority of the committee";

[0409] The enterprise has established a mapping relationship model between business processes and data flow processing, and defined functions Specific implementation rules:

[0410] Routing rule mapping: Mapping decision points and conditional branches in the process to data routing rules;

[0411] Example: Map the "approval / rejection" decision point in the release approval process to a data flow routing condition;

[0412] IF approval result == "passed" THEN route to the "release preparation" process;

[0413] ELSE routes to the "publish cancellation" process;

[0414] Transformation rule mapping: Mapping activity nodes in the process to data transformation rules;

[0415] Example: The "Quality Check" activity is mapped to the Quality Data Aggregation transformation rule;

[0416] "aggregate directive = {

[0417] Input: ["unit test results", "integration test results", "performance test results"],

[0418] Operation: "Calculate weighted average score",

[0419] Weights: [0.3, 0.4, 0.3],

[0420] Output: "Overall quality score"

[0421] }”

[0422] Priority rule mapping: mapping activity priorities in the process to data processing priorities;

[0423] Example: Emergency bug fixing process is mapped to high priority data processing mark;

[0424] IF task tag contains "urgent repair" THEN data processing priority = "highest";

[0425] When the system detects a change in the business process, it automatically triggers the data flow reconfiguration mechanism:

[0426] Change impact analysis: Identify the data flow processing logic affected by the process change. In the case of the release approval process change, the system identified the following impact points:

[0427] The approval status data aggregation logic needs to be changed from serial to parallel;

[0428] It is necessary to add data collection and processing related to the "Release Committee";

[0429] The approval decision condition has been changed from "all approved" to "majority approved";

[0430] Configuration update plan: The system generates a configuration update sequence, including:

[0431] Preparation phase: Create a new data flow configuration but do not activate it yet;

[0432] Transition phase: The old and new configurations run in parallel to process in-transit data;

[0433] Switching phase: completely switching to the new configuration;

[0434] Execute configuration updates: The system executes data flow reconfiguration as planned, monitors system status, and automatically rolls back if a node experiences three anomalies within 10 minutes.

[0435] The company has also developed a process data visualization tool to provide managers with an intuitive view of the relationship between business processes and data flow processing. The tool supports the following functions:

[0436] Process visualization: displays the current business process model, including activity nodes, flow paths, and decision points;

[0437] Data flow visualization: Displays the data flow processing flow, including data sources, transformation steps, and target outputs;

[0438] Association mapping visualization: highlighting the mapping relationship between process nodes and data processing steps;

[0439] Historical comparison: supports viewing historical process versions and change history;

[0440] Impact analysis: Preview the potential impact of planned process changes on data processing;

[0441] In practical applications, business-process-aware data stream processing technology enables the system to quickly adapt to business changes. For example, when an enterprise switches from a waterfall development model to an agile development model, the system automatically identifies the process change and adjusts the data processing logic accordingly, including shifting from phased data aggregation to iterative cycle data aggregation and from fixed milestone reporting to continuous integration status reporting. This ensures that data processing is always aligned with actual business needs, avoiding the data processing disconnection issues that often occur in traditional systems during business changes.

[0442] Application of process data synchronization protocol:

[0443] To ensure data consistency during business process changes, the company implemented a distributed consensus-based process data synchronization protocol. This protocol coordinates data processing across multiple R&D centers around the world, ensuring consistent system states while new and existing processes run in parallel.

[0444] Process data consistency model implementation:

[0445] The enterprise defines and implements a consistency model , ensuring that at any time point t, each node n in the system has an equivalent view of the business process state and data flow processing state. The specific implementation includes:

[0446] Hierarchical state synchronization mechanism:

[0447] Bottom layer: Use VectorClock to track the status updates of each node;

[0448] Middle layer: Event sourcing is used to record state change events;

[0449] Upper layer: realize visual Figure 1 Consistency checks to ensure that data processing status is synchronized with business process status;

[0450] Consistency check algorithm: The system performs a consistency check every 30 seconds.

[0451] Application of distributed consensus algorithm:

[0452] The company implemented a distributed consensus algorithm based on an improved two-phase commit protocol to coordinate data flow reconfiguration during business process changes. This algorithm deployed coordination nodes in six R&D centers around the world to ensure consistency of configuration changes. In the case of a product release process change, the algorithm executed as follows:

[0453] Phase 1 (pre-submission):

[0454] The primary coordinating node sends a pre-commit message containing a description of the changes to all participating nodes;

[0455] Each node verifies the feasibility of the change and tests the stability of the new configuration;

[0456] The node returns a ready or rejection response to the coordinator, along with detailed verification results;

[0457] Phase 2 (Submission):

[0458] The main coordinating node collects responses from all participating nodes;

[0459] When more than 80% of the nodes return ready, a commit message is sent;

[0460] Otherwise, send an abort message and record the reason for the failure;

[0461] Each node executes or abandons the change based on the coordinator's decision;

[0462] Pre-verification optimization: Before officially starting the two-phase commit, the system first performs pre-verification of configuration changes, including:

[0463] Configuration syntax verification: Check the syntax correctness of the new configuration;

[0464] Conflict detection: identifying potential conflicts with existing configurations;

[0465] Performance impact assessment: predicting the impact of changes on system performance;

[0466] Business process version management system:

[0467] The company has built a business process version management system to support the parallel operation of new and old business processes:

[0468] Version library implementation: Uses a graph database to store different versions of process models, supporting:

[0469] Version tree structure: records the derivation relationship between process versions;

[0470] Difference comparison: highlight the changes between different versions;

[0471] Tag management: add semantic tags to important versions;

[0472] Version Router Implementation: A context-aware rules engine that dynamically selects the processing version based on the following factors:

[0473] Data source: Select the appropriate version based on the data source system;

[0474] Project phase: Projects at different phases may use different process versions;

[0475] Explicit tagging: supports explicit version tags in data;

[0476] Compatibility Checker: Evaluates version compatibility through static analysis and dynamic testing:

[0477] Static analysis: Check the compatibility of data structures and flow paths;

[0478] Dynamic testing: Use historical data to verify data exchange between different versions;

[0479] Conflict resolution: Provides automatic and manual conflict resolution mechanisms;

[0480] Process Data Switch Manager:

[0481] As the company transitioned from waterfall to agile development, the transition manager performed the following actions:

[0482] Data migration planning: Developed the data mapping relationship between waterfall and agile modes:

[0483] "Mapping rule example = {

[0484] Milestone → Sprint,

[0485] "Phase delivery" → "Iterative increment",

[0486] “Phase Review” → “Iteration Review”

[0487] }”

[0488] Status synchronization control: During the 6-week transition period, the system maintains the project status of both waterfall and agile modes:

[0489] For new projects: directly adopt the agile model;

[0490] For ongoing projects: decide whether to switch based on the degree of completion;

[0491] Completion rate <30%: switch to agile mode;

[0492] Completion > 70%: Maintain waterfall mode;

[0493] Completion rate 30% to 70%: Teams can choose independently;

[0494] Rollback mechanism: To ensure business continuity, the system implements a three-level rollback mechanism:

[0495] Configuration-level rollback: only rolls back the data flow configuration and retains the data status;

[0496] State-level rollback: roll back to the system state at a specified point in time;

[0497] Full rollback: restore to the complete system state before the change;

[0498] Switching monitoring: Real-time monitoring of switching progress and system status:

[0499] Data consistency indicator: monitors the consistency of new and old process data;

[0500] Performance indicators: monitor system response time and resource utilization during the switching process;

[0501] User experience metrics: Track user error rates and satisfaction;

[0502] In agile transformation projects, multiple product lines within an enterprise need to switch development modes at different points in time. The process data synchronization protocol ensures the continuity and consistency of project data during this transformation. For example, during the sprint planning phase of a product, the system seamlessly maps demand planning data from the old process to the sprint planning data of the new process, while maintaining the integrity and relevance of historical data.

[0503] The key value of this protocol is that it enables the system to maintain data processing continuity during business process changes without downtime or data freezes, significantly reducing the impact of business changes on operations. In practice, this protocol ensured that the project management system provided accurate and consistent data support during the company's six-month agile transformation, eliminating data inconsistencies caused by process changes.

[0504] Technical effect verification:

[0505] To verify the technical effectiveness of this application method, the company conducted system performance and effect tests after implementation, focusing on evaluating the two key technical effects of improving real-time performance and enhancing business adaptability.

[0506] Improvement of real-time data processing:

[0507] Based on actual operational data from five major product line projects, the company compared system data processing latency before and after the transformation. The following is a comparison of processing latency for different data sources:

[0508] Fast streaming data processing latency (system monitoring data):

[0509] Before the transformation: average delay was 3 minutes, with peak delays reaching 15 minutes.

[0510] After the transformation: average delay is 0.2 seconds, and peak delay does not exceed 1.5 seconds;

[0511] Improvement effect: Processing delay is reduced by 99.8%, meeting real-time monitoring needs;

[0512] Medium-speed streaming data processing latency (tasks and code submission data):

[0513] Before the transformation: average delay was 2.5 hours, with peak delays reaching 8 hours.

[0514] After the transformation: average delay is 45 seconds, and peak delay does not exceed 3 minutes;

[0515] Improved effect: Processing delay is reduced by 97%, reaching near real-time processing level;

[0516] Processing delays for slow streaming data (financial and personnel assessment data):

[0517] Before the transformation: average delay was 1.5 days, with some data delayed by more than 3 days;

[0518] After the transformation: average delay is 4 hours, and the maximum delay does not exceed 12 hours;

[0519] Improvement effect: Processing delay is reduced by 89%, significantly improving data timeliness;

[0520] In the real-time performance test of a specific scenario, the company selected the key business scenario of "code submission triggering the build process" and compared the end-to-end response time before and after the transformation:

[0521] Before the transformation: It took an average of 28 minutes from code submission to the availability of data analysis results;

[0522] After the transformation: From code submission to data analysis results being available, it takes only 65 seconds on average;

[0523] Improvement effect: End-to-end response time reduced by 96%;

[0524] Improvements in real-time data processing capabilities enable enterprises to achieve the following business value:

[0525] Early problem detection: The system can issue an early warning within an average of 5 minutes after a problem occurs, compared to 3 to 4 hours before the transformation. This reduces the time it takes to repair problems by an average of 63%.

[0526] Improved resource utilization: By gaining real-time insights into resource usage, the company optimized resource allocation, increasing resource utilization from an average of 62% before the transformation to 78%.

[0527] Faster decision-making: Management now receives project status reports at any time, down from once a day. Decision-making cycles are shortened from an average of two days to four hours.

[0528] Business adaptability enhancement effect:

[0529] To verify the improvement in business process change adaptability, the company recorded multiple business process changes that occurred over a six-month period and measured the time and resources required for the system to adapt to these changes. The following is a comparison of the adaptability of three typical business process changes:

[0530] Approval process changes (from serial approval to parallel approval):

[0531] Before the transformation: 3 developers were required to work for 5 days to adjust the system, and data processing was suspended during the change;

[0532] After the transformation: The system automatically identifies process changes and reconfigures them, requiring only one administrator to confirm, without suspending data processing;

[0533] Improvement effect: labor costs reduced by 93% and business continuity improved by 100%;

[0534] Project management methodology change (from waterfall to agile):

[0535] Before the transformation: A dedicated team was required to work for three weeks to restructure the system and manually migrate historical data.

[0536] After the transformation: The system automatically adapted to the process changes, performed data mapping, and smoothly transitioned, which took a total of 5 days;

[0537] Improvement effect: adaptation time is reduced by 76%, and data consistency errors are reduced by 95%;

[0538] Organizational restructuring (department mergers and splits):

[0539] Before the transformation: Manual adjustments to data processing logic and permission settings were required, taking an average of 12 days.

[0540] After the transformation: The system automatically adjusted data flow processing and permission mapping based on organizational changes, which took 2 days;

[0541] Improvement effect: adaptation time is reduced by 83%, and configuration errors are reduced by 89%;

[0542] During periods of high business change (organizational annual adjustments), the company measured the consistency of data processing with actual business processes:

[0543] Before the transformation: After the business change, an average of 42% of the data processing logic was inconsistent with the actual business, and the average correction period was 14 days;

[0544] After the transformation: After the business change, the inconsistency rate dropped to 6%, and the system was able to automatically correct 95% of inconsistencies within 24 hours;

[0545] Improved business adaptability has brought significant business value to enterprises:

[0546] Reduced change costs: The implementation cost of process changes was reduced by an average of 85%, from an average of 42 man-days per change to 6.3 man-days.

[0547] Improved business agility: The cycle from business departments proposing process changes to their full implementation has been shortened from an average of 25 days to 4 days, enabling the company to respond more quickly to market changes.

[0548] Improved user satisfaction: User satisfaction with the project management system increased from 68% before the transformation to 92%, and the number of users reporting that the system did not meet work requirements decreased by 78%.

[0549] Based on the above verification results, the technical solution proposed in this application has been fully verified in the actual enterprise environment, achieving significant improvement in real-time performance and enhanced business adaptability, solving the key technical problems faced by multi-source heterogeneous data integration, and providing strong technical support for enterprise project management.

[0550] like Figure 2 and Figure 3 As shown in the figure, the changes in system resource utilization with the increase in data flow and the relationship between business process change adaptation time and change complexity are respectively shown.

[0551] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.

Claims

1. An enterprise project management method based on multi-source heterogeneous data integration, characterized by: The following steps are involved: Receive heterogeneous data from multiple data sources, including data with different structures, formats, and sources; Adaptively classify heterogeneous data according to update frequency and data importance to form multi-dimensional heterogeneous data streams with differentiated processing priorities; Assign logical timestamps based on business semantics to multi-dimensional heterogeneous data streams and build a temporal consistency framework; Intelligently transform, context-aware aggregation, and predictive analysis are performed on multi-dimensional, heterogeneous data streams based on logical timestamps. The dynamically evolving data stream processing model ensures logical integrity across data sources. Specifically, the following features are included: Build a business process monitoring system to collect and analyze business activity logs in real time; Establish a mapping relationship between business processes and data flow processing; Implement adaptive data flow reconfiguration mechanism; Build a visual process data correlation monitoring tool to provide managers with a visual view of business processes and data flow processing, supporting manual intervention and adjustment; The analysis results are applied to the multi-level decision support system of enterprise project management through a visual decision matrix.

2. The enterprise project management method based on multi-source heterogeneous data integration according to claim 1 is characterized in that: Also includes: Real-time monitoring of business process changes in enterprise project management; Automatically perceive the change patterns and impact scope of business processes through machine learning algorithms; According to the changes in business processes, the data flow processing logic is adaptively reconstructed to achieve synchronous evolution of processing logic and business changes.

3. The enterprise project management method based on multi-source heterogeneous data integration according to claim 1 is characterized in that: Also includes: Deploy lightweight data acquisition adapters with self-learning capabilities to connect with various data sources and extract data. The lightweight data acquisition adapters can automatically identify changes in data structure and make adaptive adjustments.

4. The enterprise project management method based on multi-source heterogeneous data integration according to claim 1 is characterized in that: Heterogeneous data includes at least one of relational database data, non-relational database data, document data, log data, API interface data, IoT device data, and social media data.

5. The enterprise project management method based on multi-source heterogeneous data integration according to claim 1 is characterized in that: It is executed on an elastic and scalable distributed server cluster, which includes edge data collection servers, centralized data processing servers and intelligent application servers, and can automatically adjust computing resource allocation according to data processing load.

6. The enterprise project management method based on multi-source heterogeneous data integration according to claim 1 is characterized in that: It is built on an event-driven distributed stream processing framework, which includes a data flow engine with a fault-tolerant mechanism, which can ensure the consistency and integrity of data processing in the event of node failure.

7. The enterprise project management method based on multi-source heterogeneous data integration according to claim 1 is characterized in that: The data source includes at least one of an enterprise's internal project management system, a code management system, a continuous integration and deployment tool, a monitoring and logging system, a communication and collaboration tool, a document management system, an enterprise resource planning system, and an external market data system.

8. The enterprise project management method based on multi-source heterogeneous data integration according to claim 1 is characterized in that: The allocation of logical timestamps adopts a multi-level priority algorithm, which comprehensively considers the business criticality, update frequency, data dependency and historical processing mode of the data flow to perform dynamic priority allocation.

9. The enterprise project management method based on multi-source heterogeneous data integration according to claim 1 is characterized in that: Also includes: Implement differentiated processing strategies for multi-dimensional and heterogeneous data streams, achieve millisecond-level response for critical business data, and provide a resource-efficient batch processing mechanism for non-critical data.

10. An enterprise project management system based on multi-source heterogeneous data integration, used to implement the enterprise project management method based on multi-source heterogeneous data integration according to any one of claims 1 to 9, characterized in that: include: An intelligent data receiving module for receiving heterogeneous data from multiple data sources. Heterogeneous data includes data with different structures, formats, and sources. Adaptive data classification module, used to adaptively classify heterogeneous data according to update frequency and data importance, forming multi-dimensional heterogeneous data streams with differentiated processing priorities; The semantic timestamp allocation module is used to assign logical timestamps based on business semantics to multi-dimensional heterogeneous data streams and build a temporal consistency framework; The intelligent analysis and processing module performs intelligent conversion, context-aware aggregation, and predictive analysis on multi-dimensional heterogeneous data streams based on logical timestamps, ensuring logical integrity across data sources based on a dynamically evolving data stream processing model. The multi-level decision support module is used to apply the analysis results to the multi-level decision support system of enterprise project management through a visual decision matrix.

Citation Information

Patent Citations

  • Enterprise cloud asset global data integration management method based on data center

    CN119313023A

  • Table-driven and data-driven method, and computer-implemented apparatus and usable program code for data integration system for heterogeneous data sources dependent upon the table-driven and data-driven method

    US20120203790A1