Multi-source heterogeneous data processing flow arrangement engine system

The multi-source heterogeneous data processing workflow orchestration engine system solves the problem that traditional data processing methods are difficult to adapt to changing business needs, and realizes unified data management and efficient and flexible automated processing.

CN121614538APending Publication Date: 2026-03-06BEIJING SUPERMAP SOFTWARE CO LTD +5
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511966087.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Traditional data processing methods rely on manually written scripts or custom-developed interfaces, which are difficult to adapt to changing business needs and different data types, resulting in complex and difficult-to-maintain data integration, transformation and analysis processes.

Method used

This paper provides a multi-source heterogeneous data processing workflow orchestration engine system, including data reading, processing and output modules. It processes data in a unified standard format, uses a rule engine and parameter configuration to build processing sub-modules, supports automatic orchestration and exception handling, and adopts a multi-modal data unified object model and processing interface system to achieve lossless data conversion and efficient processing.

Benefits of technology

It reduces labor costs, improves the efficiency and flexibility of data processing, can adapt to changing business needs and different data types, and realizes unified management and automated processing of multi-source heterogeneous data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121614538A_ABST
    Figure CN121614538A_ABST
Patent Text Reader

Abstract

According to the multi-source heterogeneous data processing flow arrangement engine system, data of different sources and different formats are processed in a unified mode and converted into data of a unified standard format, the complexity of data processing is simplified, and then the processing sub-modules are arranged directly, so that the processing efficiency is improved. According to the method, data processing for multi-source heterogeneous data can be realized without coding, variable business requirements and different data types can be adapted, and compared with a mode of manually compiling a data processing script or customizing a development interface in related technologies, the labor cost can be reduced, the data processing efficiency can be improved, and the data processing efficiency can be improved. And the flexibility of data processing is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing and integration technology, specifically to a multi-source heterogeneous data processing workflow orchestration engine system. Background Technology

[0002] With the rapid development of information systems and sensing devices, massive amounts of multi-source heterogeneous data have been generated in different business domains, including structured, semi-structured, and unstructured data, as well as various data types such as vector, raster, streaming, and document. These data differ significantly in format standards, organizational structure, storage methods, and semantic levels, making data integration, transformation, and analysis processes complex, repetitive, and difficult to maintain.

[0003] Traditional data processing methods typically rely on manually written scripts or custom-developed interfaces to implement the processing flow of multi-source heterogeneous data.

[0004] However, the above methods are difficult to adapt to changing business needs and different data types. Summary of the Invention

[0005] In view of this, this application provides a multi-source heterogeneous data processing workflow orchestration engine system, which can not only reduce labor costs and improve data processing efficiency, but also effectively improve the flexibility of data processing.

[0006] To solve the above problems, the technical solution provided in this application is as follows:

[0007] On the one hand, this application provides a multi-source heterogeneous data processing workflow orchestration engine system, the system including a data reading module, a data processing module, and a data output module:

[0008] The data reading module is used to read multi-source heterogeneous data required by the target business and convert the multi-source heterogeneous data into a unified standard format data;

[0009] The data processing module is used to process the standard format data based on the arrangement of multiple processing sub-modules for the target business, and obtain the data processing result, wherein each processing sub-module corresponds to a data processing step.

[0010] The data output module is used to convert the data processing result into the data format required by the target business, and output the converted data processing result.

[0011] In one possible implementation, the system further includes a first display module, used for:

[0012] If the output of the first processing submodule satisfies the input conditions of the second processing submodule, the second processing submodule is displayed when the first processing submodule is arranged. The first processing submodule is one of the plurality of processing submodules, and the second processing submodule is one of the plurality of processing submodules.

[0013] In one possible implementation, the system further includes a building unit for:

[0014] The multiple processing sub-modules are constructed using a rule engine and parameter configuration.

[0015] In one possible implementation, the system further includes an automatic orchestration module for:

[0016] For the target service, a conditional branch is set for the target processing submodule, and the conditional branch is used to determine the next processing submodule of the target processing submodule;

[0017] If the intermediate data processing result corresponding to the target processing submodule satisfies the first condition in the conditional branch, the processing submodule corresponding to the first condition is taken as the next processing submodule of the target processing submodule.

[0018] In one possible implementation, the processing module includes an exception handling submodule, used for:

[0019] If the third processing submodule malfunctions, the fourth processing submodule will be executed normally. The fourth processing submodule and the third processing submodule are processing submodules that are executed in parallel. The third processing submodule is one of the plurality of processing submodules, and the fourth processing submodule is one of the plurality of processing submodules.

[0020] In one possible implementation, the system further includes a storage module for:

[0021] The arrangement of the multiple processing sub-modules is saved.

[0022] In one possible implementation, the system further includes a second display module for:

[0023] During the data processing of the standard format data, the data processing status corresponding to the multiple processing sub-modules is displayed.

[0024] In one possible implementation, the multi-source heterogeneous data includes data in three dimensions: attribute information, structural description, and spatial objects.

[0025] In one possible implementation, the data reading module is used for:

[0026] Read the multi-source heterogeneous data required by the target business, and transform the data corresponding to different dimensions in the three dimensions of the multi-source heterogeneous data to obtain data in a unified standard format.

[0027] In one possible implementation, the data reading module is used to read multi-source heterogeneous data required by the target business through a standardized data reading interface, and convert the multi-source heterogeneous data into a unified standard format data;

[0028] The data output module is used to convert the data processing result into the data format required by the target business through a standardized data write interface, and output the converted data processing result. The standardized data write interface is an extensible interface, and the standardized data read interface and the standardized data write interface support Java and Python dual language extensions.

[0029] As can be seen from the above technical solution, this solution reads multi-source heterogeneous data required by the target business through a pre-built standardized data input interface, and converts the multi-source heterogeneous data into a unified standard format. This allows for unified processing of data from different sources and in different formats, simplifying the complexity of data processing. Then, the data processing workflow orchestration module processes the standard format data based on the orchestration of multiple processing sub-modules for the target business. Each processing sub-module corresponds to a data processing stage. Finally, the data processing results are converted into the data format required by the target business and output. Thus, data processing of multi-source heterogeneous data can be achieved directly through the orchestration of processing sub-modules without coding. Compared with the methods of manually writing data processing scripts or custom developing interfaces in related technologies, this reduces labor costs, improves data processing efficiency, and performs format conversion when reading or writing data, adapting to changing business needs and different data types, effectively improving the flexibility of data processing. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 A schematic diagram of a multi-source heterogeneous data processing workflow orchestration engine system provided in this application embodiment;

[0032] Figure 2A schematic diagram of a multimodal data unified object model provided in an embodiment of this application;

[0033] Figure 3 This is a schematic diagram of a multi-modal data unified processing interface system provided in an embodiment of this application. Detailed Implementation

[0034] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0035] As described in the background section, traditional data processing methods typically rely on manually written scripts or custom-developed interfaces, lacking a unified abstract model and flexible process management mechanism, making it difficult to adapt to changing business needs and different data types.

[0036] To address the aforementioned issues, this application provides a multi-source heterogeneous data processing workflow orchestration engine system. By directly orchestrating processing sub-modules, it can achieve data processing for multi-source heterogeneous data without coding. Compared to the methods of manually writing data processing scripts or custom developing interfaces in related technologies, it can reduce labor costs, improve data processing efficiency, and perform format conversion when reading or writing data, adapting to changing business needs and different data types, effectively improving the flexibility of data processing.

[0037] The solutions provided in this application relate to the field of data processing and integration technology, and are specifically illustrated through the following embodiments.

[0038] See Figure 1 The diagram shown is a schematic of a multi-source heterogeneous data processing workflow orchestration engine system provided in an embodiment of this application, including a data reading module 101, a data processing module 102, and a data output module 103.

[0039] The data reading module 101 is used to read multi-source heterogeneous data required by the target business and convert the multi-source heterogeneous data into a unified standard format data.

[0040] The target business refers to the task of processing multi-source heterogeneous data. This application does not specifically limit this, but it can be a task of fusing and analyzing multi-source heterogeneous data in application scenarios such as smart cities, geographic information systems or the Internet.

[0041] Multi-source heterogeneous data refers to heterogeneous data obtained from different sources, such as data obtained from the Global Positioning System (GPS), sensor data, third-party interface data, or structured, semi-structured and unstructured data, or vector, raster, streaming, document and other data. This application does not specifically limit this.

[0042] Standard format data refers to the unified object model of multi-modal data, that is, data with a unified standard format. The standard format refers to the pre-set data format. For example, for time and date fields, the numeric format is used as the standard format; for numeric fields, the decimal value with precision is used as the standard format; for Boolean fields, the integer (0,1) is used as the standard format; for GPS data, (precision, dimension) is used as the standard format, and so on.

[0043] refer to Figure 2 The diagram shown is a schematic of a multi-modal data unified object model provided in an embodiment of this application, including data in three dimensions: structural description, attribute information, and spatial objects.

[0044] In one possible implementation, multi-source heterogeneous data includes data in three dimensions: attribute information, structural description, and spatial objects.

[0045] Among them, attribute information is used to describe the non-spatial features of objects, such as place names, classifications, and numerical indicators. Structural descriptions include coordinate systems and attribute structures, which are used to support the spatial consistency and semantic integrity of data. Spatial objects include points, lines, surfaces, rasters, images, etc., which are used to represent entities in geographic space and their geometric relationships.

[0046] Traditional format conversion tools often forcibly map data to a minimalist, uniform format, resulting in information loss. In contrast, the embodiments of this application can summarize and abstract multi-source heterogeneous data through three dimensions of data, masking underlying differences, handling spatial and non-spatial, multi-format mixed data, providing consistent identification and processing capabilities for various data types, and supporting almost all types of geometric shapes (points, lines, surfaces, 3D models, rasters, images, etc.) and various attribute types (text, numerical values, dates, complex structures, etc.), which lays the foundation for reading, converting, and processing multi-source heterogeneous data.

[0047] In one possible implementation, the data reading module 101 is specifically used to: read the multi-source heterogeneous data required by the target business, and convert the data corresponding to different dimensions in the three dimensions of the multi-source heterogeneous data to obtain unified standard format data.

[0048] In this embodiment, attribute information, structural description, and spatial objects are separated: spatial object mapping (such as line object → surface object, surface object → point object), attribute information mapping (field name and field value conversion), and structural description mapping are performed separately.

[0049] This avoids interference between information, enables lossless data conversion, preserves geometric and attribute information in the data as much as possible, and thus improves the richness and accuracy of the data.

[0050] This application does not impose specific restrictions on the reading and conversion methods of multi-source heterogeneous data. For example, data reading interfaces for different data formats can be pre-built, and multi-source heterogeneous data can be read through these interfaces. For example, a field mapping mechanism can be provided, which automatically carries all attributes of the original data source when reading data, and can automatically or manually specify the data type conversion of fields, such as text to numeric, text to date, and modify field names.

[0051] The data processing module 102 is used to process standard format data based on the arrangement of multiple processing sub-modules for the target business, and obtain the data processing results.

[0052] Each processing submodule corresponds to a data processing step. For example, in multi-source heterogeneous spatial data application scenarios, data of different formats and sources often need to go through multiple processes to meet business requirements, such as data reading, coordinate system transformation, spatial clipping, attribute cleaning, format conversion, and output.

[0053] Single data reading and output or isolated data transformation operations cannot meet actual business needs. At the same time, output requirements are also varied, such as map rendering, spatial statistical analysis or cross-system sharing. Without a unified pipeline mechanism, relying solely on manual step-by-step operations will result in highly fragmented and complex coordination between various stages of data processing, leading to insufficient data consistency, lack of automatic task scheduling, and so on.

[0054] Each data processing submodule processes data using a unified data format, eliminating the need for data format conversion. In this embodiment, multiple processing submodules are pre-built. When processing data, these submodules need to be arranged, such as using the output of the first submodule as the input of the second, to achieve seamless connection between them. Automated data processing can then be achieved based on this arrangement.

[0055] This application does not impose specific restrictions on the arrangement of multiple processing sub-modules for the target business. For example, the arrangement can be done manually by the user or automatically based on the input and output content of each processing sub-module to obtain multiple organically connected processing sub-modules. Then, the multi-source heterogeneous data is automatically scheduled and processed according to the order of the multiple processing sub-modules.

[0056] This allows for efficient directed graph-based execution processes, such as parallel execution of tasks that can be executed in parallel; priority execution of tasks that can be executed independently without the output of a preceding task as input; and scheduling execution only after the preceding task has successfully completed, using the output of a preceding task as input. Furthermore, it supports iterative processing of batch data (multiple files in a folder, multiple files from the same data source, etc.) according to a pre-arranged process, iterating through each data point in the data source multiple times until all data is processed, significantly reducing batch data processing time.

[0057] For example, this application embodiment designs a unified multi-source data processing pipeline, which arbitrarily orchestrates and connects input reading, transformation, processing, and output processes in a pipeline manner according to the business needs of the target business, realizing the capability of "data enters once, and the whole process is automatically processed." Furthermore, based on the orchestration method tailored to the target business, it performs automated scheduling, enabling multi-source heterogeneous data to be processed automatically, in batches, and efficiently within the same processing flow, significantly improving data processing and analysis efficiency. This unified multi-source data processing pipeline uniformly transmits data based on standard format data, ensuring consistent information expression between different processing sub-modules and avoiding data loss or semantic inconsistencies.

[0058] In one possible implementation, the system further includes a first display module, used for:

[0059] If the output of the first processing submodule meets the input conditions of the second processing submodule, the second processing submodule will be displayed when the first processing submodule is being arranged.

[0060] The first processing submodule is one of a plurality of processing submodules, and the second processing submodule is one of a plurality of processing submodules.

[0061] This application does not impose specific restrictions on the display method of the second processing submodule. For example, when the user connects the first processing submodule with the previous processing submodule, that is, when the arrangement of the first processing submodule is completed, the next processing submodule is determined to include the second processing submodule based on the expected output information of the first processing submodule. Then, the second processing submodule is displayed at the output position of the first processing submodule. For another example, when the user triggers the first processing submodule, the words "Second Processing Submodule A" are displayed floating on the first processing submodule, so that the user can choose whether to use it as the next data processing step.

[0062] Therefore, during the arrangement process, the first display module shows the next data processing step that can be executed, providing users with choices and helping them improve the arrangement efficiency of data processing steps, thereby improving the efficiency of data processing.

[0063] In one possible implementation, the system also includes building blocks for:

[0064] Multiple processing sub-modules are constructed using a rule engine and parameter configuration.

[0065] By decoupling business logic from code implementation through a rules engine, the data processing strategies of each processing submodule are abstracted into configurable rule sets. Each processing submodule exists as an independent service or plugin, and its input / output format, processing algorithm, triggering conditions, exception thresholds, and other behaviors are entirely driven by parameter configuration. The rules engine is responsible for dynamically loading, parsing, and executing these configurations, enabling "zero-code" adjustments to the data processing pipeline.

[0066] Traditional data processing workflows often require repeated configuration of certain parameters in different scenarios, such as input / output paths and spatial range parameters. This application's embodiments introduce a variable mechanism to break this rigid pattern.

[0067] Variables transform specific file paths, numerical parameters, vector objects, or raster data into unified abstract objects. Users can define a few key variables, eliminating the need to configure specific values ​​in each node. Instead, these variables are passed between nodes, enabling flexible process-driven workflows that automatically adapt to different data sources, spatial ranges, and analytical conditions. Compared to workflows with fixed parameters, pipelines supporting variables significantly reduce repetitive work and greatly improve versatility. Variables support multiple types, including numeric, string, boolean, file paths, vector objects, raster data, and colors. This diverse type design ensures coverage of different business needs. Variables can be dynamically bound to node parameters. When a variable value changes, all nodes that depend on that variable will automatically update their business logic. This mechanism makes the workflow more adaptable.

[0068] Inline variables are a flexible data referencing mechanism. They allow direct referencing of parameter variables within tool input paths, output paths, field names, or expressions. For example, users can dynamically replace part of an input filename with the name of the currently processed vector space object class, or automatically append date and time to the output path. This mechanism avoids hard-coding parameter values, enabling pipelines to adapt to different data sources, naming rules, and operating scenarios, reducing manual modification workload. Through inline variables, users can fine-tune the naming and management of intermediate outputs. For instance, when processing multiple batches of data, users can use inline variables to embed batch numbers, region names, or data timestamps into output filenames, achieving batched and traceable result storage.

[0069] Therefore, through the rule engine and parameter configuration, it can support flexible transformation based on conditions and expressions, which can be applied to data processing scenarios with changing requirements and complex logic. Based on the processing logic or results of the current processing stage, it can provide users with the option of the next processing stage, or automatically orchestrate the next processing stage when it is determined, thereby effectively improving orchestration efficiency.

[0070] In one possible implementation, the system also includes an automatic orchestration module for:

[0071] Set conditional branches for the target processing submodule based on the target business.

[0072] If the intermediate data processing result corresponding to the target processing submodule satisfies the first condition in the conditional branch, the processing submodule corresponding to the first condition will be used as the next processing submodule of the target processing submodule.

[0073] The conditional branch is used to determine the next processing submodule of the target processing submodule. The conditional branch includes multiple conditions, which correspond to different data processing stages.

[0074] Intermediate data processing results refer to the output results corresponding to the target processing submodule.

[0075] In this embodiment of the application, users can set conditional branches according to business scenarios. The conditional branches can use Boolean logic to determine the status of the output result (such as whether the field exists, whether the value exceeds the threshold, whether the spatial object meets the conditions, etc.) and automatically select different subsequent task processes to execute based on the judgment result.

[0076] When the output of the target processing submodule satisfies the first condition in the branch conditions, the processing submodule corresponding to the first condition is taken as the next data processing step of the target processing submodule.

[0077] Therefore, when orchestrating the data processing process, the unified processing pipeline for multi-source data can be made to have "decision-making" capabilities like a program, enabling differentiated processing of different data in business operations, rather than simply executing fixed operations.

[0078] In one possible implementation, the processing module further includes an exception handling submodule, used for:

[0079] If the third processing submodule malfunctions, the fourth processing submodule will be executed normally. The fourth and third processing submodules are executed in parallel.

[0080] The third processing submodule is one of a plurality of processing submodules, and the fourth processing submodule is one of a plurality of processing submodules. The fourth processing submodule and the third processing submodule are processing submodules that are executed in parallel.

[0081] The exception handling submodule in this application adopts an exception isolation mechanism, which is mainly for processing submodules that can be executed in parallel. During the data processing, if the third processing submodule among multiple data processing modules fails, the third processing submodule is isolated and the fourth processing submodule that performs parallel data processing is executed normally.

[0082] Therefore, when there are multiple parallel tasks in the process, the exception isolation mechanism ensures that an error in one task will not cause other parallel tasks to be interrupted, thus avoiding local errors from affecting the overall process and effectively improving the robustness of the system.

[0083] In one possible implementation, the system also includes a storage module for:

[0084] Save the arrangement of multiple processing submodules.

[0085] The orchestrated multi-source data processing pipeline can be saved as a template for easy reuse in different business scenarios, and it also supports nesting in other processes.

[0086] This can further improve the flexibility and reusability of data processing workflow orchestration, thereby enhancing data processing efficiency.

[0087] The data output module 103 is used to convert the data processing results into the data format required by the target business and output the converted data processing results.

[0088] This application does not impose specific restrictions on the data format conversion and output method of the data processing results. For example, data writing interfaces for different business needs can be pre-built, and then the data processing results can be converted into the format required by the target business and output through the writing interface.

[0089] Traditional GIS systems are often limited in their scalability by closed architectures or language-specific bindings, restricting users' ability to quickly build functional modules according to their own business needs, especially in scenarios such as automatic data import and analysis, and embedding industry rules. This system, however, follows the architectural principle of "core stability and plug-in extensions," decoupling the core system from extended functions. The core system provides standard interfaces, and plug-ins interact with the core system through these standard interfaces. All extended functions are injected through a plug-in mechanism, allowing functions to exist as independent modules with low coupling between them.

[0090] To ensure core stability, the software architecture is divided into three layers:

[0091] (1) Core layer: Provides basic capabilities such as plugin loading, tool modeling, general UI controls, software main interface interaction, software startup and restart.

[0092] (2) Extension layer: The extension layer supports functional expansion through the plug-in mechanism. It provides rich capabilities such as data management, data editing, map making and output, data analysis, and data conversion, and supports integration with external systems (databases, cloud services, Web APIs).

[0093] (3) Script engine layer: Embedded Python interpreter, supporting the execution of user-defined script tasks, such as data parsing, transformation, analysis and processing.

[0094] To further lower the barrier to entry for users and improve the efficiency of secondary development, the software natively supports Java and Python dual-language extension development capabilities. Developers can implement plugin modules based on the Java API, or implement lightweight functional extensions through Python API scripts, thereby achieving rapid integration and flexible adaptation for industry applications.

[0095] In one possible implementation, the data reading module reads the multi-source heterogeneous data required by the target business through a standardized data reading interface, and converts the multi-source heterogeneous data into a unified standard format data.

[0096] The data output module is used to convert the data processing results into the data format required by the target business through a standardized data writing interface, and then output the converted data processing results.

[0097] The standardized data input interface and standardized data output interface are built through a multi-modal data unified processing interface system. They support Java and Python dual-language extensions, can convert different source data formats into a unified standard format data, and have scalability to adapt to continuously expanding data types.

[0098] The lack of a unified interface system will lead to high coupling between data parsing, attribute mapping, and processing logic, reducing the engine's scalability and automation. Therefore, in order to achieve unified access, efficient processing, and seamless integration of multi-source heterogeneous data, this application embodiment has constructed a unified multi-modal data processing interface system that matches the "unified multi-modal data object model".

[0099] refer to Figure 3 The diagram shown is a schematic of a multi-modal data unified processing interface system provided in an embodiment of this application, which includes the following four parts: unified object model interface (Feature), data input interface (Reader), data output interface (Writer), and data transformation interface (Transformer).

[0100] The unified object model interface for multi-modal data: The Feature object comprises three core parts: spatial object (Geometry), attribute information (Attributes), and structural description (including coordinate system and attribute structure). Data is represented as Feature objects through the Reader. Whether it's geometric data (points, lines, surfaces, rasters, point clouds) or non-spatial data (attribute tables, business records), it can be converted into a unified Feature. Access is made through a unified interface, then proceeds to the Transformer or Writer, eliminating the need for data format compatibility at each step. Feature attribute fields are not fixed, allowing for dynamic addition, modification, or deletion of attributes during the transformation process. The Feature separates the modeling of spatial objects (Geometry) and attributes (Attributes), enabling the same business record to be flexibly bound to different spatial geometric forms. It features multi-modal data, supporting not only vector spatial objects but also point clouds, rasters, and 3D models. It boasts high scalability: third-party developers can reuse new spatial data types throughout the engine without modifying the kernel, simply by mapping them to the Feature's spatial object (Geometry).

[0101] The Reader function is implemented using the adapter pattern. Internally, it handles the parsing logic for specific data formats, but externally exposes a standardized interface to convert data from different formats and sources into a unified standard format, simplifying subsequent processing. In addition to spatial data, Reader also parses field structures, attribute information, and coordinate system data, ensuring that subsequent processing and conversion steps correctly inherit semantics. Reader is independent of specific business logic, focusing only on "reading and converting to standard format data," decoupling from data processing and writing modules, thus improving engine flexibility and maintainability. Because of its independent design, adding a Reader plugin can support new data formats, avoiding excessive engine coupling. This architecture allows for rapid mapping of any new format to standard format data.

[0102] The Writer function is implemented using the adapter pattern, encapsulating the logic for writing data in various formats while exposing a standardized interface to the outside world. It transforms the data processing results of a unified standard format into data in different target formats. In addition to spatial data, Writer also transforms field structure, attribute information, and coordinate system data, provided the target format is compatible, and outputs them to the target data. Writer is independent of specific business logic and only focuses on "converting the standard format to the target data format". It is decoupled from Reader (data pair module) and Transformer (data transformation module), improving the engine's flexibility and maintainability. Because of its independent design, adding a Writer plugin can support writing new data formats, avoiding excessive coupling within the engine.

[0103] The data transformation interface, Transformer, exists as an independent functional module. For example, each processing sub-module, such as attribute processing, spatial processing, spatial analysis, and projection transformation, corresponds to a Transformer. Users can freely combine multiple Transformers according to their target business to form a customized processing pipeline. This modular design reduces the complexity of data transformation and makes the system scalable and flexible. Some Transformers work internally using a rule engine + parameter configuration approach, supporting flexible transformations based on conditions and expressions. This rule-based mechanism ensures that batch processing and automated cleaning can be achieved in complex scenarios, avoiding manual intervention one by one.

[0104] Therefore, the data interface supports dual-language extension development based on Java and Python, enabling users to quickly customize industry algorithms and business logic, achieve flexible cross-language and cross-platform extensions, improve industry adaptability and secondary development efficiency, and thus effectively enhance the flexibility and reusability of data processing.

[0105] In one possible implementation, the system further includes a second display module for:

[0106] During the data processing of standard format data, the data processing status of multiple processing sub-modules is displayed.

[0107] During the data processing of standard format data, a second display module allows for visual monitoring on the interface, including input and output parameters, execution status, execution time, and exception prompts for each processing step. The multi-modal unified object model permeates the entire data processing flow, from read / write operations to data format conversion nodes. Users can preview the data status at any node, thereby precisely controlling changes in attributes and geometric content, ensuring full traceability.

[0108] Furthermore, the system platform automatically records processing logs, input and output parameters, intermediate data states, and final generated results. This full-process recording not only provides the operation history but also presents it in a flowchart visualization, enabling users to intuitively understand the model's execution logic.

[0109] In traditional manual operations or data processing engines lacking traceability capabilities, if the result does not meet expectations, it often requires a significant amount of time to reproduce the operation step by step and troubleshoot the problem at each stage. This application's embodiments also provide a precise "error location" mechanism through data traceability functionality, allowing users to directly trace back to the data input and output of a specific step, thereby quickly identifying whether the problem stems from an incomplete data source, incorrect parameter configuration, or improper node logic.

[0110] Therefore, this application embodiment, through the second display module, can provide users with a complete data traceability mechanism, ensuring the transparency and auditability of data processing.

[0111] In summary, the multi-source heterogeneous data processing orchestration engine system provided in this application achieves integrated management, automated data transformation and processing, and traceable execution of multi-source, heterogeneous spatial and non-spatial data through a collaborative mechanism of "unified object model for multi-modal data, unified processing interface system for multi-modal data, and unified processing pipeline for multi-modal data." The engine adopts a "core stability and plug-in extension" architecture. The core system provides standard APIs, and functional modules are injected through a plug-in mechanism. It supports dual-language extension development based on Java and Python, enabling users to quickly customize industry algorithms and business logic, achieving flexible cross-language and cross-platform extensions. It possesses both flexible functional extension capabilities and can support the construction and management of complex data processes in a low-cost and high-efficiency manner, thus providing a unified, efficient, and scalable data processing infrastructure for fields such as smart cities, geographic information systems, the Internet of Things, and data governance, with the following beneficial effects:

[0112] (1) Unified heterogeneous data expression and access method: By establishing a unified object model for multi-modal data and a unified processing interface for multi-modal data, the underlying differences are shielded, and the ability to consistently identify and process multiple types of data is provided, realizing unified semantic expression and format-independent processing of data with different spatial and non-spatial structures of multiple types.

[0113] (2) Improve the flexibility and reusability of data processing: By introducing a plug-in operator mechanism, business developers can quickly map data to a multi-modal unified data object model for data formats that are not yet supported, thereby seamlessly integrating them into existing data processing processes.

[0114] (3) Lowering the technical threshold: Through a visual process orchestration method, the data processing process can be orchestrated and reused, allowing users to build complex data processing processes without programming, thereby enhancing the ease of use and maintainability of the engine.

[0115] (4) Achieve efficient automation and scalable processing: Through dynamic scheduling and parallel execution mechanisms, it supports efficient flow and real-time processing of large-scale data, supports conditional branching and variable mechanisms, realizes intelligent and adaptive processes, and meets the performance requirements of multi-scenario applications.

[0116] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0117] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multi-source heterogeneous data processing flow orchestration engine system, characterized in that, The system comprises a data reading module, a data processing module and a data output module. The data reading module is configured to read multi-source heterogeneous data required by a target service and convert the multi-source heterogeneous data into unified standard format data. The data processing module is configured to perform data processing on the standard format data based on an arrangement manner of a plurality of processing sub-modules for the target service, to obtain a data processing result, wherein each processing sub-module corresponds to a data processing link. The data output module is configured to convert the data processing result into a data format required by the target service and output the converted data processing result.

2. The system of claim 1, wherein, The system further comprises a first display module configured to: If an output result of a first processing sub-module meets an input condition of a second processing sub-module, display the second processing sub-module when arranging the first processing sub-module, wherein the first processing sub-module is one of the plurality of processing sub-modules, and the second processing sub-module is one of the plurality of processing sub-modules.

3. The system of claim 1, wherein, The system further comprises a construction unit configured to: Construct the plurality of processing sub-modules by using a rule engine and parameter configuration.

4. The system of claim 3, wherein, The system further comprises an automatic arrangement module configured to: Set a condition branch for a target processing sub-module for the target service, wherein the condition branch is used to determine a next processing sub-module of the target processing sub-module. If an intermediate data processing result corresponding to the target processing sub-module meets a first condition in the condition branch, take a processing sub-module corresponding to the first condition as the next processing sub-module of the target processing sub-module.

5. The system of claim 1, wherein, The processing module comprises an exception processing sub-module configured to: If a third processing sub-module is abnormal, normally execute a fourth processing sub-module, wherein the fourth processing sub-module and the third processing sub-module are processing sub-modules executed in parallel, the third processing sub-module is one of the plurality of processing sub-modules, and the fourth processing sub-module is one of the plurality of processing sub-modules.

6. The system of claim 1, wherein, The system further comprises a saving module configured to: Save the arrangement manner of the plurality of processing sub-modules.

7. The system of claim 1, wherein, The system further comprises a second display module configured to: Display data processing states corresponding to the plurality of processing sub-modules during data processing on the standard format data.

8. The system of claim 1, wherein, The multi-source heterogeneous data comprises attribute information, structure description and spatial object data in three dimensions.

9. The system of claim 8, wherein, The data reading module is configured to: Read multi-source heterogeneous data required by a target service, and convert data corresponding to different dimensions in the three dimensions in the multi-source heterogeneous data, to obtain unified standard format data.

10. The system of claim 1, wherein, The data reading module is configured to read multi-source heterogeneous data required by a target service through a standardized data reading interface, and convert the multi-source heterogeneous data into unified standard format data. The data output module is configured to convert the data processing result into a data format required by the target service through a standardized data write-out interface, and output the converted data processing result, wherein the standardized data write-out interface is an extensible interface, and the standardized data read-in interface and the standardized data write-out interface support Java and Python dual-language extension.