A multi-source data intelligent rule generation and processing method, device and medium

By combining a large language model and a data processing engine, intelligent rule generation and processing of multi-source data is achieved, solving the problems of complex rule engine configuration and difficult cross-system migration, improving rule generation efficiency and data processing capabilities, and ensuring data consistency and reliability.

CN120849450BActive Publication Date: 2025-12-05INSPUR GENERSOFT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511325587.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-12-05
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

Existing rule engines rely on manually written SQL, Groovy, and other scripts, which are complex to configure, costly to maintain, and difficult to migrate across systems. When faced with multi-source heterogeneous data and complex multi-level processing needs, traditional methods have significant shortcomings in terms of rule redundancy and conflict, abnormal data repair, and end-to-end tracing.

Method used

It employs a large language model to automatically convert natural language requirements into structured rules, processes multi-source data through field dictionaries and semantic normalization, generates constraint operator graphs and evaluates execution costs, and combines business data rule engines for classification judgment and assembly processing under accounting rules, providing a full-link traceability mechanism.

Benefits of technology

Significantly reduce manual configuration costs, improve rule generation efficiency, support unified access and standardized processing of multi-source and multi-level data, enhance the system's adaptability to changes in business rules, achieve intelligent anomaly detection and completion, improve data integrity and result reliability, and provide end-to-end traceability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849450B_ABST
    Figure CN120849450B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-source data intelligent rule generation and processing method, equipment and medium, belong to artificial intelligence technical field, for solving in the face of multi-source heterogeneous data and complex multi-level processing demand, there is obvious deficiency in traditional method in rule redundancy conflict, abnormal data repair, full-link traceability aspect technical problem.The method comprises: the unified data representation processing of multi-source data in different business systems is carried out under the relevant field dictionary and semantic normalization, and standard business data is obtained;Standard business data is subjected to hierarchical perception semantic analysis and uniform rule intermediate representation processing, and constraint operator graph is obtained;The execution cost of different operators in constraint operator graph is evaluated, and business data rules stored in data rule table are generated;Standard business data is subjected to classification judgment related to configuration scene, and the screened business data is determined;The screened business data is subjected to re-assembly splicing processing related to accounting rules.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, and in particular to a method and device for intelligent rule generation and processing of multi-source data, and a medium. BACKGROUND

[0002] With the advancement of digital transformation, the amount of data generated and processed by enterprises and institutions is growing exponentially every day. Financial accounting, supply chain management, medical informatization, logistics scheduling, and other fields rely on data-driven intelligent decision-making. However, there has been heterogeneity between different systems for a long time: data formats are not unified, field naming lacks standardization, and semantic definitions differ significantly, which leads to abnormal integration and rule-driven processing of cross-system data.

[0003] The current industry mainstream approach is to rely on traditional rule engines. The core idea is to drive data processing by manually configuring field mapping relationships, setting condition thresholds, and writing script rules. For example, in a financial shared platform, developers usually use SQL or Groovy scripts to implement rules such as journal splitting, loan balance verification, etc. Although this approach meets some needs in the early stage, its shortcomings gradually emerge as business scenarios become more complex:

[0004] First, the rule configuration threshold is too high. Most rule definitions rely on program developers, and business personnel cannot independently complete configuration and modification. This leads to a prolonged response period for requirements, outdated rule library updates, and serious impact on business flexibility.

[0005] Second, rule redundancy and conflicts occur frequently. In complex business scenarios, the same field may be used in multiple rules. Once the business logic is adjusted, a large number of rules must be modified globally, which can easily result in redundancy, conflicts, and even logic vulnerabilities.

[0006] Third, multi-level processing is difficult. In the financial scenario, a batch of data needs to be divided into header, line, and allocation line levels. In the supply chain scenario, an order may need to be split according to region, time window, and weight interval. Traditional rule engines have difficulty in flexible handling of such cross-level processing, often requiring complex script splicing.

[0007] Fourth, the ability to handle exceptions is limited. When data is missing or inconsistent, traditional methods usually rely on manual verification, which is inefficient, has a high error rate, and cannot form an effective automated closed loop.

[0008] With the emergence of large language models (LLM), the natural language understanding and generation capabilities of which can be utilized to automatically extract data processing rules from complex text requirements, providing a more intelligent means of processing multi-source, multi-level data. However, there is a lack of a complete architecture in the existing technology to deeply combine LLM with data processing engines, forming a closed-loop system from rule generation, rule optimization, data processing to exception completion and result tracing. Therefore, there is an urgent need for a multi-source data intelligent rule generation and multi-level processing method and system combined with a large language model to improve the degree of rule generation automation, enhance data processing capabilities, and achieve full-link traceability. SUMMARY

[0009] The embodiments of the present application provide a multi-source data intelligent rule generation and processing method, device and medium, which are used to solve the following technical problems: existing rule engines generally rely on manual scripting of SQL, Groovy, etc., and the rule configuration is complex, the maintenance cost is high, and the cross-system migration is difficult; in the face of multi-source heterogeneous data and complex multi-level processing requirements, traditional methods have obvious deficiencies in rule redundancy conflict, exception data repair, and full-link traceability.

[0010] The embodiments of the present application adopt the following technical solutions:

[0011] On the one hand, the embodiments of the present application provide a multi-source data intelligent rule generation and processing, comprising: performing unified data representation processing on the multi-source data in different business systems with respect to field dictionary and semantic normalization to obtain standard business data; performing hierarchical semantic analysis and unified rule intermediate representation processing on the standard business data through a pre-set large language model to obtain a constraint operator graph; performing execution cost evaluation on different operators in the constraint operator graph, and generating business data rules stored in a data rule table based on structured rules; performing classification judgment on the standard business data with respect to configuration scenarios through a business data rule engine corresponding to the business data rules to determine filtered business data; performing reassembly and splicing processing on the filtered business data with respect to accounting rules to generate processing rule result data; and performing tracing output on the processing rule result data.

[0012] The embodiments of the present application realize the automatic conversion of natural language requirements to structured rules through a large language model, greatly reducing the cost of manual configuration and improving the efficiency of rule generation. It can support unified access and standardized processing of multi-source and multi-level data, ensuring the consistency and integrability of data from different sources. The introduction of rule optimization and hot loading mechanism enhances the adaptability of the system to changes in business rules and enables dynamic expansion. Moreover, intelligent anomaly detection and completion using a large language model can improve data integrity and the reliability of the results. At the same time, the data processing engine supports multi-thread parallel and batch processing, improving performance and throughput in large-scale data scenarios. It also provides a full-link traceability mechanism to ensure that the processing results are explainable, reviewable, and auditable, enhancing the compliance and transparency of the system.

[0013] In a feasible implementation, the multi-source data in different business systems are processed for unified data representation under field dictionary and semantic normalization to obtain standard business data, specifically including: based on multiple access modes of different business systems, obtaining the multi-source data; wherein the multiple access modes at least include: API calling, file uploading and message queue subscription; through a field dictionary library, mapping the original field name and the standard field name of different business environments in the multi-source data in a corresponding relationship to obtain field conversion data; normalizing the data format and data annotation value of different business environments in the multi-source data based on semantic normalization rules to obtain semantic conversion data; according to a batch hierarchical processing mechanism, converting the field conversion data and the semantic conversion data into a general field table structure, and recording the corresponding traceability relationship and system log to obtain the standard business data.

[0014] In an implementable implementation, the standard business data is subjected to hierarchical perception semantic parsing and unified rule intermediate representation processing through a preset large language model, to obtain a constraint operator graph, specifically including: through the natural language requirement information input by the large language model, the standard business data is subjected to task intent recognition processing under hierarchical perception semantic parsing, to obtain intent recognition result data; wherein the intent recognition result includes: screening, splitting, aggregation and scoring; the intent recognition result data is subjected to slot filling processing, to obtain filling recognition result data; wherein the filling recognition result data includes: recognition field, recognition condition, recognition threshold and recognition operator; the filling recognition result data is subjected to hierarchical constraint mapping under correct data hierarchy according to business ontology, and field ambiguity is eliminated, to obtain hierarchical perception semantic parsing data; according to the operator type and constraint information in the executable operator graph, the hierarchical perception semantic parsing data is subjected to element conversion processing based on unified rule intermediate representation under business requirements, to obtain the constraint operator graph; wherein the operator type includes: Filter, GroupSplit, Map, Aggregate, Join and Route; the constraint information includes: field source, data hierarchy and permission tag.

[0015] In an implementable implementation, the different operators in the constraint operator graph are subjected to execution cost evaluation, and business data rules stored in a data rule table are generated based on structured rules, specifically including: through a semantic-cost joint planner, field distribution and statistical information under different execution sequences of the constraint operator graph are subjected to evaluation processing of different operator execution costs, to obtain operator evaluation results; according to the semantic context in the standard business data, the operator evaluation results are subjected to filtering processing of optimal execution, to obtain an optimal execution plan; through a compilation generation field, the constraint operator graph is subjected to generation processing of a domain function, to obtain a structured rule based on the domain function; wherein the compilation generation field includes: rule metadata, operator sequence and field constraint; the domain function at least includes: SQL function and Groovy function; based on the optimal execution plan and the constraint operator graph, the business data rules stored in the data rule table are generated.

[0016] In an implementation, the standard business data is classified and judged in relation to a configuration scenario by a business data rule engine corresponding to the business data rule, and filtered business data is determined, specifically including: importing the standard business data into the business data rule engine; performing a filtering judgment on the standard business data by a filtering definition condition in a filtering group configuration module; the filtering definition condition and the configuration scenario classification are in a header-row relationship; if the standard business data meets the filtering definition condition, the standard business data is input into a corresponding configuration scenario rule execution path; if the standard business data does not meet the filtering definition condition, the standard business data is input into a backup scenario rule execution path or an exception handling mechanism; based on the filtering judgment result, the filtered business data is determined.

[0017] In an implementation, the filtered business data is subjected to reassembly and splicing processing under accounting rules, and processing rule result data is generated, specifically including: performing data row reassembly processing based on target requirements on the filtered business data according to the accounting rules to obtain newly generated requirement data; performing dimension combination processing under a hierarchical structure on the newly generated requirement data to obtain the processing rule result data; the hierarchical structure at least includes: a header table, a row table, and an allocation row; the dimensions at least include: a department, a project, and a subject; the combination includes: aggregation and splitting.

[0018] In an implementation, the processing rule result data is output, specifically including: outputting the processing rule result data to a downstream system corresponding to a business system, and determining output target data; the output target data at least includes: a database, a report system, and an external API; recording rule links, execution batches, and exception completion processes corresponding to the output target data to obtain full-link traceability information about the output target data.

[0019] In an implementation, the processing rule result data is written into a plurality of target business tables based on different business requirement information; the plurality of target business tables are target business tables representing different business subject information under the same configuration scenario; based on the plurality of target business tables, a multi-dimensional tree link under different business requirement information is generated.

[0020] Secondly, embodiments of this application also provide an intelligent rule generation and processing device for multi-source data, the device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, so that the at least one processor can execute the intelligent rule generation and processing method for multi-source data described in any of the above embodiments.

[0021] Thirdly, embodiments of this application also provide a non-volatile computer storage medium, which is a non-volatile computer-readable storage medium storing at least one program. Each program includes instructions, which, when executed by a terminal, cause the terminal to execute a method for generating and processing intelligent rules for multi-source data as described in any of the above embodiments.

[0022] This application provides a method, device, and medium for intelligent rule generation and processing of multi-source data. Compared with the prior art, the embodiments of this application have the following beneficial technical effects:

[0023] 1. By using a large language model, the automatic conversion of natural language requirements into structured rules is achieved, significantly reducing manual configuration costs and improving rule generation efficiency;

[0024] 2. Supports unified access and standardized processing of multi-source and multi-level data, ensuring consistency and fusion of data from different sources;

[0025] 3. Introduce rule optimization and hot-loading mechanisms to enhance the system's adaptability to changes in business rules and achieve dynamic expansion;

[0026] 4. Utilize large language models for intelligent anomaly detection and completion to improve data integrity and result reliability;

[0027] 5. The data processing engine supports multi-threaded parallel and batch processing, improving performance and throughput in large-scale data scenarios;

[0028] 6. Provide a full-chain traceability mechanism to ensure that the processing results are explainable, verifiable, and auditable, thereby enhancing the system's compliance and transparency;

[0029] 7. The intelligent rule generation and processing system forms a closed loop encompassing data access, rule generation, rule optimization, data processing, anomaly completion, and result traceability, comprehensively enhancing its intelligence level and application value. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0031] Figure 1 A flowchart illustrating an intelligent rule generation and processing method for multi-source data provided in this application embodiment;

[0032] Figure 2 This is a schematic diagram of the structure of an intelligent rule generation and processing device for multi-source data provided in an embodiment of this application. Detailed Implementation

[0033] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0034] It should be noted that the intelligent rule generation and processing system of this application consists of six main parts: a data configuration module, a data import module, a data custom function execution module, a data filtering module, a target table data generation module, and a target system data push module. These modules work together to construct a closed-loop system from rule configuration, data import, data transformation, data filtering, data assembly to result push.

[0035] The data configuration module defines the key data that each stage of the entire processing flow depends on. It mainly includes the following:

[0036] Data event definition: Defines documents of different business types, including event code, name, and event message (header structure). The access party pushes data to other systems according to the defined format.

[0037] Scenario Classification: Business documents may involve multiple business scenarios, each with different processing logic. Different scenario classifications have their own data filtering conditions, data processing rules, and data push targets.

[0038] Processing rules: Defines the method for retrieving detailed data from the data push target.

[0039] Filtering criteria: Data grouping and filtering are achieved through filtering criteria. Filtering criteria are applied to processing rules to generate specific target data.

[0040] Custom functions: The system can integrate some data processing functions for data transformation, so that the input and output parameters can achieve various transformations.

[0041] Target table definition: Defines the field format when the final data is pushed to a third-party system. The target table is bound to the scenario category. One processing scenario can push data to N target systems.

[0042] This application provides a method for intelligent rule generation and processing of multi-source data, such as... Figure 1 As shown, the intelligent rule generation and processing method for multi-source data specifically includes steps S101-S106:

[0043] S101. Perform unified data representation processing on multi-source data from different business systems under the relevant field dictionary and semantic normalization to obtain standard business data.

[0044] Specifically, it is necessary to first acquire multi-source data based on various access methods from different business systems. These multiple access methods include at least: API calls, file uploads, and message queue subscriptions.

[0045] Furthermore, by using a field dictionary library, the original field names of different business environments in the multi-source data are mapped to standard field names according to their correspondence, resulting in field-converted data.

[0046] Furthermore, the data formats and data annotation values ​​of different business environments in the multi-source data are normalized based on semantic normalization rules to obtain semantically transformed data.

[0047] Furthermore, based on the batch hierarchical processing mechanism, both the field-converted data and the semantic-converted data are converted into generic field table structures, and the corresponding traceability relationships and system logs are recorded to obtain standard business data.

[0048] In one embodiment, the data import module is responsible for uniformly accessing and standardizing multi-source data from different business systems. It supports various access methods, including API calls, file uploads, and message queue subscriptions. For data from different sources, it performs field dictionary mapping and semantic normalization to form a unified data representation. When different business systems exchange information, there are issues with inconsistent field naming, semantic differences, and data structure differences. For example, System A uses the field VENDOR_ID to represent the supplier number; System B uses the field SUPPLIER_CODE to represent the supplier number; and System C uses PARTNER_ID to represent the same meaning. Direct aggregation would cause the rule engine to fail to correctly identify the fields and would also fail to guarantee data consistency. Therefore, a unified data representation needs to be established through a field dictionary mapping and semantic normalization mechanism.

[0049] Table 1. The field dictionary maintains the correspondence between various systems and standard fields.

[0050]

[0051] As shown in Table 1, during mapping, the original data fields are converted into standard fields.

[0052] Table 2 Semantic Normalization

[0053]

[0054] As shown in Table 2, semantic normalization is required to unify the values ​​of different formats and annotations.

[0055] After data transformation and integration, it is uniformly stored in a generic field table structure for easy access by the rule engine and processing engine. Simultaneously, detailed traceability relationships and system logs are recorded to ensure subsequent auditing and tracking. The format is as follows:

[0056] Table 3. Structure of generic field table

[0057]

[0058] That is, as shown in Table 3, a batch and sub-batch hierarchical processing mechanism can be adopted to ensure the integrity and stability of high-concurrency, large-volume data access.

[0059] Meanwhile, the data import module also supports the mapping and management of multi-level data structures (such as batch, header table, row table, etc.), providing a consistent data foundation for subsequent rule generation and processing.

[0060] S102. Using a pre-defined large language model, standard business data is processed through hierarchical semantic parsing and unified rule intermediate representation to obtain a constraint operator graph.

[0061] Specifically, the natural language requirement information input from the large language model can be used to perform task intent recognition processing on standard business data under hierarchical semantic parsing to obtain intent recognition result data. The intent recognition result includes: filtering, splitting, aggregation, and scoring.

[0062] Furthermore, the intent recognition result data is subjected to slot filling processing to obtain filled recognition result data. This filled recognition result data includes: recognition fields, recognition conditions, recognition thresholds, and recognition operators.

[0063] Furthermore, the data obtained from filling in the recognition results is mapped to the correct data level based on hierarchical constraints related to the business ontology, and field ambiguity is eliminated to obtain hierarchical-aware semantic parsing data.

[0064] Furthermore, based on the operator types and constraint information in the executable operator graph, the hierarchical semantic parsing data undergoes element transformation processing using a unified rule intermediate representation based on business requirements, resulting in a constraint operator graph. The operator types include: Filter, GroupSplit, Map, Aggregate, Join, and Route. The constraint information includes: field source, data level, and permission tags.

[0065] As a feasible implementation method, the data custom function execution module can leverage the natural language understanding capabilities of the large language model to transform user-input business requirements or rule documents into executable custom functions or data rules, which take effect when data processing is performed after matching the scenario.

[0066] In one embodiment, natural language input is first provided, allowing users to directly describe their needs, such as "split orders with sales exceeding 1 million by region." Then, hierarchical semantic parsing is performed, including: Intent Detection: identifying operation types such as "filter," "split," "aggregate," and "rate." Slot Filling: identifying fields, conditions, thresholds, and operators. Next, hierarchical constraint mapping is performed, automatically assigning fields to the correct data hierarchy based on the business ontology defined by data events (e.g., order batches, detail rows, strategy fields), avoiding cross-level ambiguity. Finally, disambiguation and clarification are performed; if a field is ambiguous, such as an amount potentially being considered as either order amount or payment amount, [further clarification is needed].

[0067] In one embodiment, a unified intermediate representation construction is also required. That is, the parsed results do not directly generate SQL, but are transformed into an executable operator graph containing the following elements: Operator type: Filter, GroupSplit, Map, Aggregate, Join, Route; Constraint information: Field source, data level, permission label; User input requirements can be transformed into Filter[field=sales amount, condition=>1000000]; GroupSplit[dimension=region]. Finally, an intermediate representation is constructed using a constrained operator graph (constrained operator graph), facilitating cross-language compilation, conflict detection, and execution optimization.

[0068] S103. Evaluate the execution cost of different operators in the constraint operator graph, and generate business data rules stored in the data rule table based on structured rules.

[0069] Specifically, a predefined semantic-cost joint planner is also needed to evaluate the field distribution and statistical information under various execution orders in the constraint operator graph to obtain the operator evaluation results.

[0070] Furthermore, based on the semantic context in the standard business data, the operator evaluation results are filtered to select the optimal execution plan.

[0071] Furthermore, the constraint operator graph is further processed by generating domain functions through compilation, resulting in structured rules based on these domain functions. The compiled domain includes rule metadata, operator order, and field constraints. The domain functions include at least SQL functions and Groovy functions.

[0072] Furthermore, based on the optimal execution plan and constraint operator graph, business data rules are generated and stored in the data rule table.

[0073] In one embodiment, the data-customized function execution module also includes a semantic-cost joint planner. That is, executable operators may have multiple execution orders. This module introduces a semantic + cost joint optimization mechanism, which can evaluate the cost of executing different operators based on field distribution and statistical information (such as the number of regions and sales distribution); combined with semantic context (e.g., "split" takes precedence over "aggregation"), it determines the optimal execution plan. Then, structured rules are generated, and the executable operator graph is compiled to generate a domain DSL (JSON): containing rule metadata, operator order, and field constraints; then the DSL automatically generates domain functions such as SQL and Groovy functions. Next, the rules are stored, that is, the generated rules are stored in the data rule table.

[0074] As a feasible implementation method, the data-defined function execution module can also adaptively optimize rules. During subsequent rule execution, the system continuously collects runtime data, failure reasons, user feedback, etc., and this information is used to update the domain thesaurus, LLM few-shot hint set, and instruction parameter weights. This results in more accurate subsequent rule generation.

[0075] S104. Using the business data rule engine corresponding to the business data rules, classify and judge the standard business data according to the relevant configuration scenarios to determine the filtered business data.

[0076] Specifically, standard business data is first imported into the business data rule engine.

[0077] Furthermore, the standard business data is filtered and judged based on the filtering definition conditions in the filtering group configuration module. The filtering definition conditions and the configured scenario categories have a header-line relationship.

[0078] Furthermore, if the standard business data meets the filtering definition conditions, the standard business data is input into the corresponding configuration scenario rule execution path. If the standard business data does not meet the filtering definition conditions, the standard business data is input into the backup scenario rule execution path or the exception handling mechanism.

[0079] Furthermore, based on the screening results, the filtered business data is determined.

[0080] As a feasible implementation method, in the data filtering module, after configuring the data judgment rules, the transformed data can participate in the rule judgment and perform multi-condition custom configuration. The filtering condition configuration and the scene classification are in a head-row relationship. The dynamic data of the filtering conditions determines whether to enter the corresponding scene classification.

[0081] In one embodiment, during execution, the transformed data, such as vouchers, orders, and supply chain events represented by unified fields, is imported into the rule engine. First, the filtering group configuration module matches and judges the data. If the filtering conditions defined by the filtering group are met, the data enters the corresponding scenario rule execution path; otherwise, the system can redirect it to a backup rule or an exception handling mechanism. This head-line association ensures the hierarchy and scalability of rule configuration, avoiding the problems of rule redundancy and frequent logical conflicts in traditional rule engines.

[0082] S105. Perform reassembly and splicing processing on the filtered business data under the relevant accounting rules to generate the processing rule result data.

[0083] Specifically, according to accounting rules, the filtered business data must be re-assembled based on the target requirements to obtain newly generated requirement data.

[0084] Furthermore, the newly generated requirement data is processed by combining dimensions within a hierarchical structure to obtain the processing rule result data. The hierarchical structure includes at least a header table, row tables, and allocation rows. Dimensions include at least departments, projects, and subjects; combinations include aggregation and splitting.

[0085] As a feasible implementation method, the processing rule results need to be written into multiple target business tables based on different business requirement information. These multiple target business tables represent different business entity information under the same configuration scenario. Based on these multiple target business tables, a multi-dimensional tree-like link is generated related to different business requirement information.

[0086] In one embodiment, the processing rule transformation step in the target table data generation module needs to take effect after filtering is completed. This module reassembles the filtered data according to accounting rules to generate new data rows and supports multi-level associations. For example, a batch of voucher data can be broken down into a hierarchical structure of "header table - row table - allocation row" according to the configured processing rules. Multiple accounting rules can work together to output combined results of different dimensions, such as aggregation or splitting by department, project, or account. The generated new data not only maintains the logical consistency between fields but also reflects the relationship between business and finance through a tree structure.

[0087] In one embodiment, the target table data generation module also supports cross-table and cross-system data output capabilities. That is, the generated processing rule results can be written to multiple target business tables as needed, achieving "one-time rule configuration, multiple system reuse." For example, the same purchase data can be transformed through rules to simultaneously generate financial entry records, budget occupancy records, and tax return data, thus forming a multi-dimensional financial tree-like business chain. This mechanism not only improves the reusability of rules and data consistency between systems but also significantly shortens the cycle from business requirement changes to rule effectiveness.

[0088] S106. Output the processing rule result data retrospectively.

[0089] Specifically, the processing rule results are first output to the downstream systems corresponding to the business system, and the target output data is determined. The target output data includes at least: databases, reporting systems, and external APIs. Then, the rule chain, execution batches, and exception completion process corresponding to the target output data are recorded to obtain end-to-end traceability information about the target output data.

[0090] As a feasible implementation method, the target system data push module is responsible for outputting the processing results of the target table data generation module to downstream systems and ensuring the traceability of the results. Output targets include databases, reporting systems, external APIs, etc. It can also record the rule chain, execution batch, and exception completion process for each result, forming a full-chain traceability capability. Simultaneously, it supports compliance review and auditing requirements based on logs and traceability chains, ensuring the reliability and authority of the results.

[0091] In one embodiment, a financial scenario: the user inputs "Generate journal entries for all transactions with sales revenue greater than 1 million yuan, categorized by region and tax rate." The system automatically generates the rules, and during execution, breaks down the data, generates headers, rows, and allocation rows, verifies debit and credit balance, and outputs the results to the financial ledger. Abnormal data is flagged by the LLM and the system automatically completes the relevant accounts.

[0092] In one embodiment, a logistics scenario: the user inputs "orders destined for Beijing are divided into light and heavy goods for delivery by weight." The system parses and generates the rules, automatically splits the order into batches, outputs delivery instructions, and detects anomalies (such as missing weight information), which are then completed by the LLM or confirmed manually.

[0093] In one embodiment, a supply chain scenario: the user inputs "mark all orders that have not been completed for more than 30 days as high risk and generate an alert". The system generates conditional rules, executes them, writes high-risk orders into a risk table, and the LLM generates a remedial strategy.

[0094] In addition, embodiments of this application also provide an intelligent rule generation and processing device for multi-source data, such as... Figure 2 As shown, the intelligent rule generation and processing device 200 for multi-source data specifically includes:

[0095] At least one processor 201; and a memory 202 communicatively connected to the at least one processor 201; wherein the memory 202 stores instructions executable by the at least one processor 201 to enable the at least one processor 201 to execute:

[0096] Standard business data is obtained by processing multi-source data from different business systems into a unified data representation under the relevant field dictionary and semantic normalization.

[0097] By using a pre-defined large language model, standard business data is subjected to hierarchical semantic parsing and unified rule intermediate representation processing to obtain a constraint operator graph.

[0098] The execution cost of different operators in the constraint operator graph is evaluated, and business data rules are generated and stored in the data rule table based on structured rules.

[0099] The standard business data is classified and judged according to the relevant configuration scenarios through the business data rule engine corresponding to the business data rules, and the filtered business data is determined.

[0100] The filtered business data is reassembled and spliced ​​according to relevant accounting rules to generate the processing rule result data.

[0101] The processing rule result data will be traced and output.

[0102] This application's embodiments achieve automatic conversion from natural language requirements to structured rules through a large language model, significantly reducing manual configuration costs and improving rule generation efficiency. It supports unified access and standardized processing of multi-source, multi-level data, ensuring consistency and fusion of data from different sources. The introduction of rule optimization and hot-loading mechanisms enhances the system's adaptability to changes in business rules, enabling dynamic expansion. Furthermore, intelligent anomaly detection and completion are performed using a large language model, improving data integrity and result reliability. Simultaneously, the data processing engine supports multi-threaded parallel and batch processing, improving performance and throughput in large-scale data scenarios. A full-link traceability mechanism is also provided to ensure that processing results are interpretable, verifiable, and auditable, enhancing system compliance and transparency.

[0103] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and medium embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the description of the method embodiments.

[0104] The devices and media provided in this application are one-to-one with the methods. Therefore, the devices and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0105] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0106] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0107] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0108] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0109] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0110] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0111] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0112] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0113] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of this specification.

Claims

1. A method for intelligent rule generation and processing of multi-source data, characterized in that, The method includes: Standard business data is obtained by processing multi-source data from different business systems into a unified data representation under the relevant field dictionary and semantic normalization. Using a pre-defined large language model, the standard business data undergoes hierarchical semantic parsing and unified rule intermediate representation processing to obtain a constraint operator graph, specifically including: Using the natural language requirement information input from the large language model, the standard business data is processed for task intent recognition under hierarchical semantic parsing to obtain intent recognition result data; wherein, the intent recognition result includes: filtering, splitting, aggregation, and scoring; The intent recognition result data is subjected to slot filling processing to obtain filled recognition result data; wherein, the filled recognition result data includes: recognition field, recognition condition, recognition threshold and recognition operator; The filled recognition result data is mapped to the hierarchical constraint of the correct data level by the relevant business ontology, and field ambiguity is eliminated to obtain hierarchical-aware semantic parsing data. Based on the operator types and constraint information in the executable operator graph, the hierarchical semantic parsing data is subjected to element transformation processing based on unified rule intermediate representation under business requirements to obtain the constraint operator graph; wherein, the operator types include: Filter, GroupSplit, Map, Aggregate, Join, and Route; the constraint information includes: field source, data level, and permission label; The execution cost of different operators in the constraint operator graph is evaluated, and business data rules stored in the data rule table are generated based on structured rules, specifically including: The semantic-cost joint planner is used to evaluate the field distribution and statistical information under various execution orders in the constraint operator graph, and the execution cost of different operators is evaluated to obtain the operator evaluation results. Based on the semantic context in the standard business data, the optimal execution plan is obtained by filtering the operator evaluation results for the best execution. By compiling and generating a domain, the constraint operator graph is processed to generate domain functions, resulting in structured rules based on the domain functions; wherein, the compiled and generated domain includes: rule metadata, operator order, and field constraints; the domain functions include at least: SQL functions and Groovy functions; Based on the optimal execution plan and the constraint operator graph, the business data rules stored in the data rule table are generated; The standard business data is classified and judged according to relevant configuration scenarios through the business data rule engine corresponding to the business data rules, and the filtered business data is determined. The filtered business data is reassembled and spliced ​​according to relevant accounting rules to generate processing rule result data; The processing rule result data is traced and output.

2. The intelligent rule generation and processing method for multi-source data according to claim 1, characterized in that, Standard business data is obtained by unifying the representation of multi-source data from different business systems using relevant field dictionaries and semantic normalization, specifically including: The multi-source data is obtained based on multiple access methods from different business systems; wherein, the multiple access methods include at least: API calls, file uploads, and message queue subscriptions; By using a field dictionary library, the original field names and standard field names of different business environments in the multi-source data are mapped according to their correspondence to obtain field-converted data; The data formats and data annotation values ​​of different business environments in the multi-source data are normalized based on semantic normalization rules to obtain semantically transformed data. According to the batch hierarchical processing mechanism, both the field-converted data and the semantic-converted data are converted into generic field table structures, and the corresponding traceability relationships and system logs are recorded to obtain the standard business data.

3. The intelligent rule generation and processing method for multi-source data according to claim 1, characterized in that, The standard business data is categorized and judged according to relevant configuration scenarios using the business data rule engine corresponding to the aforementioned business data rules, to determine the filtered business data, specifically including: Import the standard business data into the business data rule engine; The standard business data is filtered and judged based on the filtering definition conditions in the filtering group configuration module; wherein, the filtering definition conditions and the configuration scenario classification have a head-line relationship. If the standard business data meets the filtering definition conditions, then the standard business data is input into the corresponding configuration scenario rule execution path; If the standard business data does not meet the filtering definition conditions, the standard business data will be input into the backup scenario rule execution path or the exception handling mechanism. Based on the screening results, the filtered business data is determined.

4. The intelligent rule generation and processing method for multi-source data according to claim 1, characterized in that, The filtered business data is then reassembled and spliced ​​according to relevant accounting rules to generate processing rule result data, specifically including: According to the accounting rules, the filtered business data is re-assembled based on the target requirements to obtain newly generated requirement data; The newly generated requirement data is processed by combining dimensions under a hierarchical structure to obtain the processing rule result data; wherein, the hierarchical structure includes at least: header table, row table and allocation row; the dimensions include at least: department, project and subject; the combination includes: aggregation and splitting.

5. The intelligent rule generation and processing method for multi-source data according to claim 1, characterized in that, The processing rule result data is traced and output, specifically including: The processing rule result data is output to the downstream system corresponding to the business system, and the output target data is determined; wherein, the output target data includes at least: database, reporting system and external API; Record the rule chain, execution batch, and exception completion process corresponding to the output target data to obtain full-chain traceability information about the output target data.

6. The intelligent rule generation and processing method for multi-source data according to claim 5, characterized in that, Based on different business requirements, the processing rule results are written into multiple target business tables; wherein, the multiple target business tables are target business tables representing different business entity information under the same configuration scenario; Based on the multiple target business tables, a multi-dimensional tree-like link is generated under different business requirement information.

7. An intelligent rule generation and processing device for multi-source data, characterized in that, The device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, enabling the at least one processor to execute a method for intelligent rule generation and processing of multi-source data according to any one of claims 1-6.

8. A non-volatile computer storage medium, characterized in that, The storage medium is a non-volatile computer-readable storage medium, which stores at least one program. Each program includes instructions, which, when executed by a terminal, cause the terminal to perform a method for intelligent rule generation and processing of multi-source data according to any one of claims 1-6.

Citation Information

Patent Citations

  • Calculation method based on rule engine

    CN117950786A

  • Code automatic generation and optimization system based on multiple modes

    CN120276718A