Optimization method and device for rule engine, equipment and medium
By partitioning the database in the Spark environment and creating and caching algorithm rule engine instances in each partition, the problem of low resource utilization of the rule engine is solved, performance is improved and computational overhead is reduced, making it suitable for big data computing scenarios.
Patent Information
- Application Number
- CN202510996414.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-31
AI Technical Summary
In the Spark environment, the algorithm rule engine has low resource utilization, high computational overhead and resource consumption, and cannot meet the needs of large-scale computing, especially with performance degradation under frequent recompilation operations.
The original database is divided into several partitions. An algorithm rule engine instance is created in each partition and stored in the cache. When a call request is received, it is determined whether a pre-compiled algorithm instance exists in the cache. If it exists, it is called; otherwise, it is pre-compiled and stored.
It significantly improves the performance of the rules engine in the big data Spark environment, especially in large-scale computing scenarios with a performance improvement of about 20%, solves the problems of low resource utilization and high computing overhead, and has good scalability.
Smart Images

Figure CN120873024A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of infrastructure operation and maintenance technology, and in particular to an optimization method, apparatus, device and medium for rule engines. Background Technology
[0002] In existing technologies, Oracle, as a storage database, offers fast processing speeds, robust data backup and fault tolerance mechanisms, and high stability. However, with technological advancements, Oracle databases become slower when processing large datasets, impacting the user experience. To improve database computational efficiency, the database is run in a Spark environment.
[0003] In the medical field, Apache Spark, with its high performance, scalability, and rich functionality, has become a key tool for processing and analyzing massive amounts of medical data. Its applications span multiple core scenarios, including disease prediction, drug development, and clinical decision support, significantly improving the efficiency and accuracy of medical data analysis. Spark utilizes the memory caching mechanism of RDDs (Resilient Distributed Datasets) and DataFrames to keep intermediate results in memory, reducing disk I / O. This makes it particularly suitable for machine learning algorithms that require multiple iterations (such as logistic regression and random forests). For example, in genomic data analysis, Spark can increase the computation speed of tasks such as sequence alignment and variant detection by more than 10 times, significantly improving efficiency compared to traditional MapReduce frameworks (such as Hadoop).
[0004] In the financial sector, Spark, with its high performance, high scalability, and rich functionality, has become a key technology support for core business scenarios such as risk assessment, anti-fraud, real-time transaction monitoring, customer credit assessment, and market risk analysis. Its core value lies in its real-time data processing capabilities, multi-source data integration capabilities, flexible machine learning support, and compliance audit assurance.
[0005] To improve the computational efficiency of the database, an algorithm rule engine module was introduced. While the algorithm rule engine reduces code development workload, its performance in a Spark environment cannot meet the demands of large-scale computations. Each call to the algorithm rule engine module requires recompiling the algorithm, resulting in numerous recompiled instances. This doesn't significantly impact performance for small-scale computations, but for large-scale computations, frequent recompilation operations severely degrade algorithm efficiency, leading to performance degradation. Furthermore, the need for recompilation with each call results in frequent allocation and release of resources (such as CPU and memory), leading to low resource utilization and increased system overhead. In a Spark environment, the algorithm rule engine cannot be optimized through broadcasting because it cannot be serialized; this means that in distributed computing, each task node needs to independently create an engine instance, further increasing computational overhead and resource consumption. Summary of the Invention
[0006] In view of the shortcomings of the prior art, the present invention provides an optimization method, apparatus, device and medium for rule engines, aiming to solve the problems of low resource utilization, high computational overhead and resource consumption in the prior art.
[0007] The technical solution of the present invention is as follows:
[0008] The first embodiment of the present invention provides an optimization method for a rules engine, the method comprising:
[0009] Obtain the original database and divide it into several partitions;
[0010] Create an algorithm rule engine instance in each partition and store the algorithm rule engine instance in the cache;
[0011] When an algorithm rule engine call request is detected, it is determined whether a corresponding pre-compiled algorithm instance exists in the cache;
[0012] If a corresponding pre-compiled algorithm instance exists in the cache, then the pre-compiled algorithm instance in the cache is invoked to generate the corresponding compilation result;
[0013] If the corresponding pre-compiled algorithm instance does not exist in the cache, then algorithm pre-compilation is performed to generate a pre-compiled algorithm instance, and the pre-compiled algorithm instance is stored in the algorithm rule engine instance in the cache.
[0014] Another embodiment of the present invention provides an optimization apparatus for a rule engine, the apparatus comprising:
[0015] The partitioning module is used to obtain the original database and divide the original database into several partitions;
[0016] The algorithm instance creation and storage module is used to create an algorithm rule engine instance in each partition and store the algorithm rule engine instance in the cache;
[0017] The monitoring and judgment module is used to determine whether a corresponding pre-compiled algorithm instance exists in the cache when an algorithm rule engine call request is detected.
[0018] The instance invocation module is used to invoke the pre-compiled algorithm instance in the cache if a corresponding pre-compiled algorithm instance exists in the cache, and generate the corresponding compilation result;
[0019] The algorithm compilation and storage module is used to perform algorithm pre-compilation if there is no corresponding pre-compiled algorithm instance in the cache, generate a pre-compiled algorithm instance, and store the pre-compiled algorithm instance in the algorithm rule engine instance in the cache.
[0020] Another embodiment of the present invention provides a computer device, the computer device including at least one processor; and,
[0021] A memory communicatively connected to the at least one processor; wherein,
[0022] The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to perform the steps of the optimization method for the rule engine described above.
[0023] Another embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of the optimization method for a rules engine described above.
[0024] Beneficial Effects: The optimization method, apparatus, device, and medium for a rule engine according to embodiments of the present invention include: acquiring an original database and dividing the original database into several partitions; creating an algorithm rule engine instance in each partition and storing the algorithm rule engine instance in a cache; when an algorithm rule engine call request is detected, determining whether a corresponding pre-compiled algorithm instance exists in the cache; if it exists, calling the pre-compiled algorithm instance in the cache to generate the corresponding compilation result; if it does not exist, performing algorithm pre-compilation to generate a pre-compiled algorithm instance and storing the pre-compiled algorithm instance in the algorithm rule engine instance in the cache. Embodiments of the present invention significantly improve the performance of the rule engine in a big data Spark environment through a pre-compilation caching mechanism and a partitioned instantiation engine. Attached Figure Description
[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a schematic diagram illustrating the application environment of an embodiment of the optimization method for a rule engine according to the present invention;
[0027] Figure 2This is a flowchart of a preferred embodiment of an optimization method for a rules engine according to the present invention;
[0028] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of an optimization device for a rules engine according to the present invention;
[0029] Figure 4 This is a schematic diagram of a preferred embodiment of a computer device according to the present invention;
[0030] Figure 5 This is another structural schematic diagram of a preferred embodiment of a computer device according to the present invention. Detailed Implementation
[0031] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention is further described in detail below. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0032] The embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0033] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Here, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0034] The optimization method for rule engines provided in this invention can be applied to, for example... Figure 1In this application environment, the client communicates with the server via a network. The client accesses the server's network or business platform, and the server can obtain the raw database, divide it into several partitions, create an algorithm rule engine instance in each partition, and store the algorithm rule engine instance in a cache. Upon detecting an algorithm rule engine call request, the server checks if a corresponding pre-compiled algorithm instance exists in the cache. If it exists, the server calls the pre-compiled algorithm instance in the cache to generate the corresponding compilation result; if it does not exist, the server performs algorithm pre-compilation to generate a pre-compiled algorithm instance, which is then stored in the cached algorithm rule engine instance. In this invention, the performance of the rule engine in a big data Spark environment is significantly improved through a pre-compilation caching mechanism and a partitioned instantiation engine. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0035] To address the above problems, embodiments of the present invention provide an optimization method for a rules engine. Please refer to [link to relevant documentation]. Figure 2 , Figure 2 This is a flowchart illustrating a preferred embodiment of an optimization method for a rules engine according to the present invention. Figure 2 As shown, it includes:
[0036] Step S100: Obtain the original database and divide the original database into several partitions.
[0037] This invention primarily addresses the migration from Oracle databases to a Spark big data environment. The Spark environment is a distributed computing ecosystem built on the Apache Spark framework, designed for large-scale data processing, machine learning, and real-time analytics. It supports diverse workloads, from batch processing to stream processing, through in-memory computing, elastic scaling, and a rich API library. The Spark environment includes Spark Core. Spark Core provides fundamental distributed computing capabilities, including task scheduling, memory management, and fault tolerance. Key abstractions include RDD (Resilient Distributed Dataset): an immutable, partitionable, and parallelizable distributed data collection supporting coarse-grained transformations (such as map and filter). DataFrame / Dataset: a structured API that optimizes query execution plans and supports SQL-like operations (such as select and join).
[0038] The Spark environment includes the following extended modules: Spark SQL: supports structured data querying, is compatible with HiveQL, and can directly read formats such as Parquet and JSON. Spark Streaming: micro-batch processing of streaming data (such as Kafka and Flume), with latency in the second range, and is gradually being replaced by Structured Streaming (continuous stream processing based on DataFrames). MLlib: a built-in machine learning algorithm library (classification, regression, clustering) that supports distributed training. GraphX: a graph computing framework used for scenarios such as social network analysis and path planning. The Spark environment also includes cluster managers. Cluster managers include Standalone: a simple cluster manager included with Spark, suitable for testing environments. YARN: a general resource manager for the Hadoop ecosystem, supporting multi-tenant resource allocation. Kubernetes: cloud-native deployment, dynamically scaling worker nodes. Mesos: a general cluster scheduling framework suitable for mixed workloads.
[0039] In Spark, partitioning is the core mechanism of distributed computing. It determines how data is divided and distributed across different nodes in the cluster for execution. A reasonable partitioning strategy can significantly improve parallel efficiency, reduce network overhead, and avoid data skew. Physical partitioning: Logically large datasets (such as RDDs and DataFrames) are split into multiple physical blocks (partitions), each stored on a different node in the cluster. Parallel computing unit: A Spark Task corresponds to a partition, and the number of partitions determines the degree of parallelism. After partitioning, each partition is processed by an independent Task, achieving parallel computing. Data locality: Prioritizes processing local data, avoiding cross-node transfers and reducing network traffic. The number of partitions affects the CPU core utilization of the Executor (usually, the number of partitions ≥ the number of CPU cores), controlling resource allocation.
[0040] Step S100, namely obtaining the original database and dividing the original database into several partitions, includes:
[0041] Step S101: Obtain the original Spark database;
[0042] Step S102: Obtain data attributes from the Spark database;
[0043] Step S103: Divide the Spark database into several partitions according to the data attributes.
[0044] Obtain the raw Spark database and its data attributes. Divide the Spark database into several partitions based on these attributes. Partition types include input partitions, transformed partitions, and explicit partitions. Partitioning methods include, but are not limited to, hash partitioning, range partitioning, and custom partitioning methods. Hash partitioning algorithms take the hash value of the partition key and then modulo it to distribute the data evenly. Suitable for operations requiring even distribution, such as joins and groupByKey. Range partitioning divides the data into partitions based on a range of values (e.g., alphabetical order or numerical range). Suitable for sort and range operations to avoid data skew. Custom partitioning is implemented by inheriting the org.apache.spark.Partitioner class.
[0045] Step S200: Create an algorithm rule engine instance in each partition and store the algorithm rule engine instance in the cache.
[0046] Based on the analysis of the algorithm rule engine and the current state of the big data environment, since the algorithm rule engine cannot be serialized in the Spark environment and cannot be optimized through broadcasting, it is proposed to create an algorithm engine instance for each partition of data. Only one algorithm rule engine instance is maintained in each partition, avoiding the overhead of creating an instance for each algorithm.
[0047] An algorithmic rule engine is a system that separates business rules from application code, allowing non-technical personnel to modify system behavior by configuring rules. It is a core component driving automated decision-making, data processing, or business logic, typically implemented through a predefined rule set or machine learning model. The components of a rule engine include: a rule repository (storing all business rules such as conditions, actions, and priorities); an inference engine (matching and executing rules based on input data); an execution module (triggering actions such as rejecting transactions or sending notifications); and monitoring and feedback (recording rule execution results and supporting dynamic optimization).
[0048] The advantage of a rules engine lies in separating business logic from code, enabling business experts to directly participate in rule formulation and modification without the need for developer intervention, thereby improving the system's flexibility and responsiveness.
[0049] Step S200, which involves creating an algorithm rule engine instance in each partition and storing the algorithm rule engine instance in a cache, includes:
[0050] Step S201: Create pre-compiled algorithm rules;
[0051] Step S202: Load the pre-compiled algorithm rules into the rule engine and create an algorithm rule engine instance;
[0052] Step S203: Store the algorithm rule engine instance in the cache.
[0053] Create pre-compiled algorithm rules, load the pre-compiled algorithm rules into the rule engine, create an algorithm rule engine instance, and store the algorithm rule engine implementation in the cache.
[0054] In the medical field, algorithmic rule engines transform medical knowledge, clinical guidelines, and expert experience into executable logical rules, combining them with machine learning algorithms to automate decision support in scenarios such as disease diagnosis, treatment recommendation, and risk warning. Their core value lies in improving diagnostic and treatment efficiency, reducing human error, and supporting personalized medicine. Medical rule engines typically employ a hybrid architecture, combining hard-coded rules (deterministic logic based on medical guidelines) and soft rules (probabilistic models based on machine learning) to adapt to the complexity and uncertainty of medical knowledge.
[0055] The core components of the rules engine include: a rule base storing medical knowledge rules (such as IF-THEN statements), for example: IF patient age > 50 AND systolic blood pressure > 140 AND diastolic blood pressure > 90 THEN diagnosed as hypertension (grade 2); IF gene testing shows BRCA1 mutation THEN recommended high-risk breast cancer screening; an inference engine matching rules based on input data (such as electronic medical records, test results) to trigger actions (such as diagnostic suggestions, medication reminders); and an interpretation module generating visual reports on the decision-making basis.
[0056] Architectures integrating with machine learning algorithms include: **Rule + Statistical Model:** Rules are used to filter obvious anomalies, and then models such as logistic regression and random forests are used to predict complex risks. Example: In diabetic retinopathy screening, the rule engine first excludes non-diabetic patients, then uses a CNN model to analyze fundus images. **Rule + Deep Learning:** Rules are embedded as prior knowledge into neural networks to improve model interpretability. Example: In cancer treatment recommendations, the rule engine enforces "contraindications to chemotherapy" to prevent the model from outputting dangerous treatment plans.
[0057] In the financial sector, algorithmic rule engines are a key technology. By separating complex business logic and rules from application code, they automate decision-making and logical processing, thereby improving business efficiency, accuracy, and flexibility. The core concept of an algorithmic rule engine is that it's a rule-based system that matches and executes input data using a predefined rule base and inference engine to generate corresponding decisions or actions. In finance, these rules are typically based on business knowledge, regulatory requirements, or risk models. For example: Loan approval: Automatically deciding whether to approve a loan application based on rules such as the customer's credit score, income level, and debt situation. Fraud detection: Identifying abnormal trading behavior by analyzing rules such as transaction amount, frequency, and geographical location. Transaction monitoring: Real-time monitoring of market data to trigger stop-loss or liquidation rules to control risk.
[0058] The technical architecture includes a rule base, which stores business rules and supports CRUD operations on rules. Rules are typically defined in a structured format (such as IF-THEN statements) or in the form of decision tables. The inference engine is the core component, responsible for rule matching and execution. It efficiently traverses the rule base using pattern matching algorithms (such as the Rete algorithm), finding rules that meet the conditions and triggering actions. The Rete algorithm is used to store the factual state during the matching process by building a pattern network and a connection network, avoiding redundant calculations and improving matching efficiency. The interface layer provides interfaces for interaction with external systems, such as API calls, message queues (Kafka), or database connections, supporting real-time data input and result feedback.
[0059] In the financial sector, algorithmic rule engines are used for risk assessment and control, such as credit approval: automatically calculating risk scores and determining loan amounts based on customer characteristics (e.g., age, occupation, income) and credit history. Market risk monitoring: real-time analysis of market data such as stocks and exchange rates triggers VaR (Value at Risk) or volatility threshold rules to adjust investment portfolios. Algorithmic rule engines offer flexibility and maintainability. Business rules are separated from code, allowing non-technical personnel (such as risk control experts) to modify rules through graphical interfaces or low-code tools without redeveloping the system.
[0060] In a further embodiment, step S200, which involves creating an algorithm rule engine instance in each partition and storing the algorithm rule engine instance in a cache, further includes:
[0061] Step S211: Obtain the algorithm source program and the corresponding pre-compiled algorithm instance;
[0062] Step S212: Create an algorithm rule engine based on the algorithm source program and the corresponding pre-compiled algorithm instance;
[0063] Step S213: Store the algorithm rule engine instance in the cache.
[0064] In Apache Spark, caching is a key optimization technique used to persist frequently accessed RDDs or DataFrames / Datasets to memory or disk, avoiding redundant computations and significantly improving job performance. In-Memory Cache stores data entirely in the server process's memory (such as heap memory or off-heap memory), achieving fast lookups through structures like hash tables or skip lists. Common implementations include local caches: Guava Cache: A thread-safe cache provided by Google, supporting expiration policies, size limits, and weak references. Caffeine: A high-performance local cache using the W-TinyLFU eviction algorithm, supporting asynchronous loading. Ehcache: A local cache supporting distributed scaling, configurable for disk overflow.
[0065] Distributed caching: Redis: An in-memory key-value store that supports various data structures (strings, hashes, lists, sets, etc.) and provides persistence options (RDB / AOF). Memcached: A pure in-memory cache with a simple and efficient design, suitable for caching static data (such as images and CSS).
[0066] Data structures include, but are not limited to, hash tables (HashMap): average O(1) lookup time, but hash collisions may occur. Skip lists: ordered structures that support range queries; Redis's ZSET is based on this.
[0067] Memory management methods are as follows: Heap memory: allocated within the JVM heap and affected by GC (e.g., Guava Cache). Off-heap memory (Direct Memory): bypasses JVM GC and reduces pauses (e.g., Netty's ByteBuf, Redis's jemal loc).
[0068] In this embodiment of the invention, the cache is stored in key-value pairs. The key is a unique value for the algorithm, storing the algorithm source code, i.e., a specific algorithm fragment. The value is a pre-compiled algorithm instance. In the caching system, key-value pairs are the data storage format. They allow for quick location of the corresponding value through a unique key, enabling efficient data read and write operations. Keys must be globally unique to avoid conflicts (e.g., user ID user:1001 instead of the simple number 1001). Readability is paramount: meaningful names are used (e.g., order:20231001:status instead of ord123). Simplicity is key: excessively long keys are avoided (RedisKeys are recommended to be <1KB) to reduce memory usage. Structured data is maintained: hierarchical relationships are represented using separators (e.g., colons, |) for easy maintenance and expansion. Examples: User information: user:{userId}:profile; Product inventory: product:{productId}:stock; Session data: session:{sessionId}:data.
[0069] When optimizing value storage based on data type selection, consider the following: For simple types: String, Number (e.g., Redis's INCR counter); For complex types: Hash, List, Set, Sorted Set (ZSET); Serialization: Use JSON, Protobuf, or MessagePack for compressed storage of objects (e.g., Redis's HSET for storing user objects); Use Snappy or Zstandard compression for large text or binary data (e.g., images); Use more compact encoding (e.g., Redis's ziplist for optimizing memory for small lists).
[0070] The key-value implementation of the caching system can use Redis to achieve high-performance key-value storage. For example, the data structures supported include: String: SET key value, supporting GET and SETEX (with expiration time); Hash: HSET user:1 name "Alice" age 30, suitable for storing object fields; List: LPUSH messages "msg1", implementing a message queue; and ZSET: ZADD leaderboard:score user1 100, supporting leaderboards.
[0071] In some other embodiments, Memcached, a pure in-memory cache, can also be used for storage. Pure in-memory caching only supports String types, has no persistence, and is suitable for caching static data. Multi-core CPUs are used to handle concurrent requests with consistent hashing, reducing data migration when nodes are added or removed.
[0072] Step S300: When an algorithm rule engine call request is detected, determine whether there is a corresponding pre-compiled algorithm instance in the cache. If yes, proceed to step S400; otherwise, proceed to step S500.
[0073] The system monitors request data. If an algorithm rule engine call request is detected, it can display the currently invoked rule engine request (method name, parameters, timestamp) on a visual interface. It distinguishes between pending, matching, and completed statuses. Matched rules and their priorities can also be displayed in card format. Manual rule re-evaluation is supported. The system records a complete request processing log (e.g., rule loading, condition matching, action execution) during the request process. It provides viewing of raw request / response data in JSON format. It queries the cache to see if a pre-compiled algorithm instance is already stored and executes the corresponding steps based on the query results.
[0074] Step S300, which involves determining whether a corresponding pre-compiled algorithm instance exists in the cache when an algorithm rule engine call request is detected, further includes:
[0075] Step S301: When an algorithm rule engine call request is detected, obtain the partition corresponding to the call data according to the algorithm rule engine call request;
[0076] Step S302: Obtain the original algorithm source program based on the algorithm rule engine call request;
[0077] Step S303: Generate a cache identifier based on the original algorithm source program;
[0078] Step S304: Determine whether there is a pre-compiled algorithm instance in the cache that corresponds to the cache identifier.
[0079] When an algorithm rule engine call request is detected, the partition corresponding to the call data is obtained based on the request. Within the partition, the original algorithm source code is called, and a cache identifier exists within the original algorithm source code; this cache identifier is the algorithm's unique identifier. When calling the algorithm rule engine, a cache key is first generated based on the algorithm's unique identifier, and it is determined whether a pre-compiled algorithm instance corresponding to the cache key exists in the cache. Subsequent steps are executed based on the determination result.
[0080] Step S400: Call the pre-compiled algorithm instance in the cache to generate the corresponding compilation result.
[0081] If a corresponding pre-compiled algorithm instance exists in the cache, the pre-compiled algorithm instance in the cache is directly called to generate the corresponding compilation result, thereby avoiding recompilation.
[0082] Step S500: Perform algorithm pre-compilation to generate a pre-compiled algorithm instance, and store the pre-compiled algorithm instance in the cached algorithm rule engine instance.
[0083] If the corresponding pre-compiled algorithm instance does not exist in the cache, the compilation process is started to pre-compile the algorithm, generate a pre-compiled algorithm instance, and store the result in the cache for subsequent use.
[0084] Step S500, which involves performing algorithm pre-compilation to generate a pre-compiled algorithm instance if no corresponding pre-compiled algorithm instance exists in the cache, and storing the pre-compiled algorithm instance in the algorithm rule engine instance in the cache, includes:
[0085] Step S501: If there is no corresponding pre-compiled algorithm instance in the cache, the pre-compiled algorithm is called to perform algorithm pre-compilation and generate the pre-compiled calculation result as the pre-compiled algorithm instance.
[0086] Step S502: Store the pre-compiled algorithm instance in the cached algorithm rule engine instance;
[0087] Step S503: Determine whether the current algorithm rule engine call request is the last call request. If the current algorithm rule engine call request is the last call request, then execute step S504. If the current algorithm rule engine call request is not the last call request, then execute step S505.
[0088] Step S504: Release the cache occupied by the algorithm instance;
[0089] Step S505: Obtain the next algorithm rule engine call request and determine whether there is a corresponding pre-compiled algorithm instance in the cache. If there is a corresponding pre-compiled algorithm instance in the cache, call the pre-compiled algorithm instance in the cache to generate the corresponding compilation result; if not, perform algorithm pre-compilation to generate a pre-compiled algorithm instance and store the pre-compiled algorithm instance in the algorithm rule engine instance in the cache.
[0090] If the corresponding pre-compiled algorithm instance does not exist in the cache, the pre-compiled algorithm is invoked to perform algorithm pre-compilation, generating a pre-compiled calculation result, which serves as the pre-compiled algorithm instance. This pre-compiled algorithm instance is then stored in the algorithm rule engine instance within the cache.
[0091] The pre-compiled algorithm instance is stored in the cached algorithm rule engine instance. It is then determined whether the current algorithm rule engine call request is the last call request. If it is, the cache occupied by the algorithm instance is released. If not, the next algorithm rule engine call request is retrieved, and the above steps are repeated. For example, it is further determined whether a corresponding pre-compiled algorithm instance exists in the cache. If it does, the cached pre-compiled algorithm instance is invoked to generate the corresponding compilation result; otherwise, algorithm pre-compilation is performed to generate a pre-compiled algorithm instance, which is then stored in the cached algorithm rule engine instance.
[0092] Step S500, which involves performing algorithm precompilation to generate a precompiled algorithm instance and storing the precompiled algorithm instance in the algorithm rule engine instance in the cache, further includes:
[0093] Step S601: Monitor the cache capacity;
[0094] Step S602: When the cache capacity is greater than or equal to the preset capacity threshold, obtain the usage frequency of the pre-compiled algorithm instance;
[0095] Step S603: Delete pre-compiled algorithm instances whose usage frequency is less than a preset frequency threshold from the cache.
[0096] At the partition level, the same algorithm is pre-compiled and the compilation results are cached. When the cache reaches a certain capacity, the pre-compiled algorithm that is not frequently used is eliminated based on the hit rate, so as to control the cache size and avoid memory overflow.
[0097] Eliminating infrequently used pre-compiled results based on hit rate is a cache management strategy based on access frequency. Its core idea is to prioritize retaining frequently used pre-compiled results (such as SQL execution plans, code compilation caches, etc.) while eliminating entries that have not been accessed for a long time or have a low hit rate.
[0098] Hit rate = (number of times a precompiled entry is used) / (time the entry has existed or total number of visits) or hit rate directly uses the number of visits as the weight (simplified version).
[0099] Periodically scan the pre-compiled cache and calculate the hit rate for each entry. Evict entries with a hit rate below a threshold (e.g., not used in the past N accesses) or the lowest ranking. Variants of LRU (Least Recently Used) or LFU (Least Frequently Used) strategies can be combined.
[0100] Compared to existing technologies, this invention significantly improves the performance of the Oracle rules engine in a big data Spark environment through a pre-compiled caching mechanism and an optimized partition instantiation engine, particularly in high-volume computing scenarios, where performance is improved by approximately 20%. This solution not only addresses the performance bottlenecks of existing technologies but also boasts excellent scalability and applicability, making it widely applicable to other similar big data computing scenarios.
[0101] It should be noted that there is no necessary order between the above steps. Those skilled in the art will understand from the description of the embodiments of the present invention that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.
[0102] Another embodiment of the present invention provides an optimization apparatus for a rule engine, which corresponds one-to-one with the optimization method for a rule engine described in the above embodiments. For example... Figure 3 As shown, device 1 includes:
[0103] The partitioning module 100 is used to obtain the original database and divide the original database into several partitions;
[0104] The algorithm instance creation and storage module 200 is used to create an algorithm rule engine instance in each partition and store the algorithm rule engine instance in the cache;
[0105] The monitoring and judgment module 300 is used to determine whether a corresponding pre-compiled algorithm instance exists in the cache when an algorithm rule engine call request is detected.
[0106] The instance invocation module 400 is used to invoke the pre-compiled algorithm instance in the cache if a corresponding pre-compiled algorithm instance exists in the cache, and generate the corresponding compilation result;
[0107] The algorithm compilation and storage module 500 is used to perform algorithm pre-compilation if there is no corresponding pre-compiled algorithm instance in the cache, generate a pre-compiled algorithm instance, and store the pre-compiled algorithm instance in the algorithm rule engine instance in the cache.
[0108] For specific implementation details, please refer to the method embodiment; they will not be repeated here.
[0109] In one embodiment, the partitioning module 100 is specifically used for:
[0110] Obtain the raw Spark database;
[0111] Retrieve data attributes from a Spark database;
[0112] The Spark database is divided into several partitions based on the data attributes.
[0113] For specific implementation details, please refer to the method embodiment; they will not be repeated here.
[0114] In one embodiment, the algorithm instance creation and storage module 200 is specifically used for:
[0115] Create pre-compiled algorithm rules;
[0116] Load the pre-compiled algorithm rules into the rule engine to create an algorithm rule engine instance;
[0117] The algorithm rule engine instance is stored in the cache.
[0118] For specific implementation details, please refer to the method embodiment; they will not be repeated here.
[0119] In one embodiment, the algorithm instance creation and storage module 200 is further configured to:
[0120] Obtain the algorithm source code and the corresponding pre-compiled algorithm instance;
[0121] An algorithm rule engine is created based on the algorithm source code and the corresponding pre-compiled algorithm instance;
[0122] The algorithm rule engine instance is stored in the cache.
[0123] For specific implementation details, please refer to the method embodiment; they will not be repeated here.
[0124] In one embodiment, the monitoring and judgment module 300 is specifically used for:
[0125] When an algorithm rule engine call request is detected, the partition corresponding to the call data is obtained according to the algorithm rule engine call request;
[0126] Based on the algorithm rule engine, a request is made to obtain the original algorithm source program;
[0127] Generate a cache identifier based on the original algorithm source code;
[0128] Determine whether a pre-compiled algorithm instance corresponding to the cache identifier exists in the cache.
[0129] For specific implementation details, please refer to the method embodiment; they will not be repeated here.
[0130] In one embodiment, the algorithm compilation and storage module 500 is specifically used for:
[0131] If the corresponding pre-compiled algorithm instance does not exist in the cache, the pre-compiled algorithm is called to perform algorithm pre-compilation and generate the pre-compiled calculation result as the pre-compiled algorithm instance.
[0132] The pre-compiled algorithm instance is stored in the cached algorithm rule engine instance;
[0133] Determine if the current algorithm rule engine call request is the last call request;
[0134] If the current algorithm rule engine call request is the last call request, then release the cache occupied by the algorithm instance;
[0135] If the current algorithm rule engine call request is not the last call request, then the next algorithm rule engine call request is obtained, and it is determined whether there is a corresponding pre-compiled algorithm instance in the cache. If there is a corresponding pre-compiled algorithm instance in the cache, then the pre-compiled algorithm instance in the cache is called to generate the corresponding compilation result; if there is no such instance, then algorithm pre-compilation is performed to generate a pre-compiled algorithm instance, and the pre-compiled algorithm instance is stored in the algorithm rule engine instance in the cache.
[0136] For specific implementation details, please refer to the method embodiment; they will not be repeated here.
[0137] In one embodiment, the apparatus further includes a cache capacity monitoring module, which is specifically used for:
[0138] Monitor the cache capacity;
[0139] When the cache capacity is greater than or equal to the preset capacity threshold, the usage frequency of the pre-compiled algorithm instance is obtained.
[0140] Instances of pre-compiled algorithms that are used less frequently than a preset frequency threshold will be removed from the cache.
[0141] For specific implementation details, please refer to the method embodiment; they will not be repeated here.
[0142] This invention provides an optimization device for a rules engine. Through a pre-compiled caching mechanism and a partitioned instantiation engine optimization scheme, it significantly improves the performance of the system's Oracle rules engine in a big data Spark environment, especially in high-volume computing scenarios, where performance is improved by approximately 20%. This solution not only solves the performance bottlenecks in existing technologies but also has good scalability and applicability, and can be widely applied to other similar big data computing scenarios.
[0143] Another embodiment of the present invention provides a computer device, which may be a server, and its internal structure diagram may be as follows. Figure 4As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements server-side functions or steps for an optimization method used in a rules engine.
[0144] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements client-side functions or steps for an optimization method used in a rules engine.
[0145] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0146] Obtain the original database and divide it into several partitions;
[0147] Create an algorithm rule engine instance in each partition and store the algorithm rule engine instance in the cache;
[0148] When an algorithm rule engine call request is detected, it is determined whether a corresponding pre-compiled algorithm instance exists in the cache;
[0149] If a corresponding pre-compiled algorithm instance exists in the cache, then the pre-compiled algorithm instance in the cache is invoked to generate the corresponding compilation result;
[0150] If the corresponding pre-compiled algorithm instance does not exist in the cache, then algorithm pre-compilation is performed to generate a pre-compiled algorithm instance, and the pre-compiled algorithm instance is stored in the algorithm rule engine instance in the cache.
[0151] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0152] Obtain the original database and divide it into several partitions;
[0153] Create an algorithm rule engine instance in each partition and store the algorithm rule engine instance in the cache;
[0154] When an algorithm rule engine call request is detected, it is determined whether a corresponding pre-compiled algorithm instance exists in the cache;
[0155] If a corresponding pre-compiled algorithm instance exists in the cache, then the pre-compiled algorithm instance in the cache is invoked to generate the corresponding compilation result;
[0156] If the corresponding pre-compiled algorithm instance does not exist in the cache, then algorithm pre-compilation is performed to generate a pre-compiled algorithm instance, and the pre-compiled algorithm instance is stored in the algorithm rule engine instance in the cache.
[0157] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0158] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0159] The embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0160] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general-purpose hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can exist in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0161] It should be noted that if any software tools or components not belonging to our company appear in the embodiments of this application, they are merely for illustrative purposes and do not represent actual use.
[0162] Among other things, conditional language such as “can,” “may,” “may,” or “may,” unless otherwise specifically stated or otherwise understood as in the context in which they are used, is generally intended to convey that a particular implementation may include (but not others) certain features, elements, and / or operations. Therefore, such conditional language is also generally intended to imply that features, elements, and / or operations are necessary for one or more implementations in any way, or that one or more implementations must include logic for determining, with or without input or prompting, whether such features, elements, and / or operations are included or will be performed in any particular implementation.
[0163] The contents already described herein in this specification and accompanying drawings include examples of optimization methods and apparatuses capable of providing rules engines. Of course, it is not possible to describe every conceivable combination of elements and / or methods for the purpose of describing the various features of this disclosure, but it will be appreciated that many other combinations and substitutions of the disclosed features are possible. Therefore, it will be apparent that various modifications can be made to this disclosure without departing from the scope or spirit of this disclosure. Furthermore, or in alternatives, other embodiments of this disclosure may become apparent from consideration of this specification and accompanying drawings and from practice of this disclosure as presented herein. It is intended that the examples presented in this specification and accompanying drawings be considered illustrative rather than restrictive in all respects. Although specific terminology is used herein, it is used in a general and descriptive sense and is not intended for limiting purposes.
Claims
1. An optimization method for a rule engine, characterized in that... The method includes: Obtain the original database and divide it into several partitions; Create an algorithm rule engine instance in each partition and store the algorithm rule engine instance in the cache; When an algorithm rule engine call request is detected, it is determined whether a corresponding pre-compiled algorithm instance exists in the cache; If a corresponding pre-compiled algorithm instance exists in the cache, then the pre-compiled algorithm instance in the cache is invoked to generate the corresponding compilation result; If the corresponding pre-compiled algorithm instance does not exist in the cache, then algorithm pre-compilation is performed to generate a pre-compiled algorithm instance, and the pre-compiled algorithm instance is stored in the algorithm rule engine instance in the cache.
2. The optimization method for a rule engine according to claim 1, characterized in that, The process of obtaining the original database involves dividing it into several partitions, including: Obtain the raw Spark database; Retrieve data attributes from a Spark database; The Spark database is divided into several partitions based on the data attributes.
3. The optimization method for a rule engine according to claim 1, characterized in that, The step of creating an algorithm rule engine instance in each partition and storing the algorithm rule engine instance in a cache includes: Create pre-compiled algorithm rules; Load the pre-compiled algorithm rules into the rule engine to create an algorithm rule engine instance; The algorithm rule engine instance is stored in the cache.
4. The optimization method for a rule engine according to claim 1, characterized in that, The step of creating an algorithm rule engine instance in each partition and storing the algorithm rule engine instance in a cache includes: Obtain the algorithm source code and the corresponding pre-compiled algorithm instance; An algorithm rule engine is created based on the algorithm source code and the corresponding pre-compiled algorithm instance; The algorithm rule engine instance is stored in the cache.
5. The optimization method for a rule engine according to claim 1, characterized in that, When an algorithm rule engine call request is detected, the process of determining whether a corresponding pre-compiled algorithm instance exists in the cache includes: When an algorithm rule engine call request is detected, the partition corresponding to the call data is obtained according to the algorithm rule engine call request; Based on the algorithm rule engine, a request is made to obtain the original algorithm source program; Generate a cache identifier based on the original algorithm source code; Determine whether a precompiled algorithm instance corresponding to the cache identifier exists in the cache.
6. The optimization method for a rule engine according to claim 1, characterized in that, If the corresponding pre-compiled algorithm instance does not exist in the cache, then algorithm pre-compilation is performed to generate a pre-compiled algorithm instance, and the pre-compiled algorithm instance is stored in the algorithm rule engine instance in the cache, including: If the corresponding pre-compiled algorithm instance does not exist in the cache, the pre-compiled algorithm is called to perform algorithm pre-compilation and generate the pre-compiled calculation result as the pre-compiled algorithm instance. The pre-compiled algorithm instance is stored in the cached algorithm rule engine instance; Determine if the current algorithm rule engine call request is the last call request; If the current algorithm rule engine call request is the last call request, then release the cache occupied by the algorithm instance; If the current algorithm rule engine call request is not the last call request, then the next algorithm rule engine call request is obtained, and it is determined whether there is a corresponding pre-compiled algorithm instance in the cache. If there is a corresponding pre-compiled algorithm instance in the cache, then the pre-compiled algorithm instance in the cache is called to generate the corresponding compilation result; if there is no such instance, then algorithm pre-compilation is performed to generate a pre-compiled algorithm instance, and the pre-compiled algorithm instance is stored in the algorithm rule engine instance in the cache.
7. The optimization method for a rule engine according to claim 1, characterized in that, If the cache does not contain a corresponding pre-compiled algorithm instance, then algorithm pre-compilation is performed to generate a pre-compiled algorithm instance, and the pre-compiled algorithm instance is stored in the algorithm rule engine instance in the cache. The method further includes: Monitor the cache capacity; When the cache capacity is greater than or equal to the preset capacity threshold, the usage frequency of the pre-compiled algorithm instance is obtained. Instances of pre-compiled algorithms that are used less frequently than a preset frequency threshold will be removed from the cache.
8. An optimization device for a rules engine, characterized in that, The device includes: The partitioning module is used to obtain the original database and divide the original database into several partitions; The algorithm instance creation and storage module is used to create an algorithm rule engine instance in each partition and store the algorithm rule engine instance in the cache; The monitoring and judgment module is used to determine whether a corresponding pre-compiled algorithm instance exists in the cache when an algorithm rule engine call request is detected. The instance invocation module is used to invoke the pre-compiled algorithm instance in the cache if a corresponding pre-compiled algorithm instance exists in the cache, and generate the corresponding compilation result; The algorithm compilation and storage module is used to perform algorithm pre-compilation if there is no corresponding pre-compiled algorithm instance in the cache, generate a pre-compiled algorithm instance, and store the pre-compiled algorithm instance in the algorithm rule engine instance in the cache.
9. A computer device, characterized in that, The computer device includes at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the steps of the optimization method for a rules engine as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of the optimization method for a rules engine as described in any one of claims 1-7.