A general distributed service scheduling and data processing method
Patent Information
- Application Number
- CN202610857996.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-29
AI Technical Summary
针对现有技术的不足,本发明提供了一种通用分布式服务调度与数据处理方法,解决现有技术中存在的协议锁定、适配成本高、存储逻辑复杂、调度能力不足的技术问题
1.通过RPC调度层、通用数据总线层、通用序列化层、通用存储器层的四层解耦架构设计,将通信、序列化、存储、调度四大核心功能通过预定义的统一抽象接口完全隔离,层间无任何业务逻辑耦合,这样开发人员可根据业务需求独立替换任意一层的底层实现,仅需更换对应层级的插件,无需修改上层业务代码与框架核心逻辑,解决了现有框架协议锁定、底层实现与核心逻辑强耦合的技术问题,实现了底层能力的无感知灵活切换,系统的多场景适配灵活性得到质的提升;其次,通过全链路可插拔的插件化扩展机制,将所有底层能力实现封装为独立插件,新增通信协议、序列化格式、存储介质、注册中心适配能力时,开发人员仅需实现对应层级预定义的统一抽象接口,即可将新能力以插件形式接入框架核心,无需对框架核心代码进行任何侵入式修改,同时,新增、替换、卸载插件的全过程无需重启服务节点、不影响线上业务的正常运行,降低系统迭代与扩展的技术成本,显著提升了业务技术迭代的响应效率。
Smart Images

Figure CN122845663A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of distributed systems and microservice architecture technology, and in particular to a general distributed service scheduling and data processing method. Background Technology
[0002] With the large-scale deployment of cloud computing and edge computing technologies, distributed systems and microservice architectures have become the core architectural paradigms for enterprise applications, internet services, and industrial IoT systems. In a distributed architecture, cross-node remote service calls (RPC), cross-service data serialization, and storage management of multi-source heterogeneous data are the three core components supporting the stable and efficient operation of the entire distributed system. Currently, mainstream distributed service scheduling frameworks in the industry are represented by gRPC and Dubbo. These frameworks provide standardized implementation solutions for communication between distributed services, lowering the development threshold of distributed systems to some extent, but they also have the following problems: Firstly, existing distributed service scheduling frameworks, from underlying communication selection, serialization configuration, and storage adaptation to system iteration, all revolve around a strong binding pattern between the core logic and the underlying implementation. However, in the first step of framework implementation—communication layer selection—the core code of existing RPC frameworks is deeply bound to specific communication protocols; for example, gRPC is strongly bound to HTTP / 2, and Dubbo is by default strongly bound to TCP. Developers lock in the underlying transmission method the moment they complete the framework selection. When business scenarios shift from high-throughput cloud-native intranet scenarios to low-latency edge computing scenarios, requiring a change in the transmission protocol, intrusive modifications to the core logic of the framework or even a complete refactoring of the communication module are necessary, resulting in an insurmountable protocol lock. Secondly, in the second step of serialization scheme configuration, existing frameworks do not set a unified abstraction layer for serialization capabilities, leading to different APIs and adaptation methods for different serialization formats. The configuration logic is completely independent, requiring developers to write dedicated adaptation code for different formats such as Protobuf and JSON. Switching serialization schemes necessitates the refactoring of a large amount of business code, resulting in fragmented and locked serialization capabilities. In the third step of multi-source storage adaptation, the existing framework does not provide a unified storage abstraction capability. Developers need to write dedicated driver logic for different storage media such as MySQL, Redis, and OSS. When multiple storage media coexist, multiple sets of independent code need to be maintained, and unified transaction control across storage cannot be achieved, resulting in scenario-locked storage capabilities. This end-to-end coupled binding mode requires intrusive modifications to the core framework code when business development requires the addition of communication protocols, serialization formats, or storage media adaptation capabilities. This not only undermines system stability but also significantly prolongs the technology iteration cycle, ultimately leading to comprehensive limitations on system flexibility, scalability, and iteration efficiency.
[0003] Secondly, the native scheduling capabilities of existing RPC frameworks only cover basic request routing and simple load balancing. In real-world high-concurrency business scenarios, numerous third-party components must be introduced to supplement core scheduling capabilities: message queue components are needed for asynchronous request caching and batch processing; rate limiting and circuit breaking components are needed for service reliability assurance; and tracing components are needed for full-process observability of calls. This fragmented capability patchwork model firstly leads to a significant increase in system architecture redundancy. Adaptation and integration between multiple components introduce additional performance overhead, and state synchronization and data consistency between components are difficult to guarantee. In high-concurrency scenarios, issues such as request blocking, link breakage, and inability to automatically recover from anomalies are highly likely, resulting in severely insufficient support for core business scenarios. Secondly, the deployment, operation, and version compatibility of multiple components require substantial manpower, and the troubleshooting process is lengthy and difficult to pinpoint, significantly increasing the overall system operation and maintenance costs and failure risks. Therefore, we propose a general distributed service scheduling and data processing method. Summary of the Invention
[0004] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a general distributed service scheduling and data processing method, which solves the technical problems of protocol locking, high adaptation costs, complex storage logic, and insufficient scheduling capabilities in existing technologies.
[0005] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: A general distributed service scheduling and data processing method is proposed, the steps of which are as follows: A layered and decoupled distributed processing architecture is constructed, which includes an RPC scheduling layer, a general data bus layer, a general serialization layer, and a general storage layer. The layers interact through a predefined unified abstract interface, and there is no business logic coupling between the layers. The underlying functions of each layer are connected to the architecture in the form of pluggable plug-ins. A unified resource descriptor specification is predefined synchronously, which is used to standardize the metadata format of services and resources. When client nodes and server nodes in a distributed system start up, they encapsulate their own service metadata through the unified resource descriptor specification, and access the adapted registry center through the general data bus layer to register services and initialize node information. The client node receives user service requests, converts them into standard RPC request messages, and calls the general serialization layer to perform serialization processing on the RPC request messages; The client node queries and obtains the communication address and service metadata information of the target server node from the registry center through the general data bus layer, and routes the serialized RPC request message to the target server node based on the pre-configured calling strategy of the RPC scheduling layer. After receiving an RPC request message, the server node performs deserialization processing on the RPC request message through the general serialization layer to obtain the business request parameters. Then, it performs business logic processing based on the business request parameters, generates an RPC response message, and sends it back to the client node after serialization. After receiving the RPC response message, the client node performs deserialization processing on the RPC response message to obtain the business processing result, and then performs subsequent business logic execution based on the business processing result.
[0006] Furthermore, the RPC scheduling layer performs the following scheduling and processing steps for RPC requests: When a client node initiates an RPC request message, it first determines whether the request type of the RPC request message is a synchronous request or an asynchronous request. If it is determined to be a synchronous request, the RPC request message is forwarded to the target server node directly according to the pre-configured calling strategy. If it is determined to be an asynchronous request, the RPC request message is first written to the local message cache queue, and the scheduler performs batch scheduling and distribution processing on the RPC request messages in the message cache queue according to the preset priority rules. During the scheduling process, the RPC scheduling layer records the call chain data, timeout status data, and retry status data of RPC requests in real time through the general memory layer, and realizes request anomaly recovery and full-link tracing based on the recorded data.
[0007] Furthermore, the pre-configured calling strategies of the RPC scheduling layer include four types: load balancing strategy, circuit breaker degradation strategy, rate limiting strategy, and retry strategy; as detailed below: The load balancing strategy is selected from any one of round-robin, weighted, consistent hashing, and least connections. The circuit breaker and degradation strategy triggers the circuit breaker operation by statistically analyzing the timeout rate and error rate of RPC requests, and performs fast failure degradation processing after the circuit breaker is triggered. The rate limiting strategy is implemented based on the token bucket algorithm or the leaky bucket algorithm to control the concurrency and QPS of RPC requests. The retry strategy is pre-configured with the number of retries, the retry interval, and the retry trigger conditions, and automatically retryes failed RPC requests that meet the trigger conditions.
[0008] Furthermore, the transmission protocol adaptation process of the general data bus layer is as follows: The general data bus layer natively includes five transport protocols: AMQP, HTTP, IPC, TCP, and UDP, through pluggable plugins. When adding a custom transmission protocol, the custom transmission protocol is encapsulated as an independent pluggable plug-in access architecture based on the unified transmission abstraction interface predefined in the general data bus layer. The business code does not need to be aware of the type and switching logic of the underlying transmission protocol.
[0009] Furthermore, the service registration and discovery process of the general data bus layer is adapted to six registration centers—Consul, Redis, etcd, ZooKeeper, Kubernetes Service, and DNS—through pluggable plugins. In particular, for edge deployment scenarios, the general data bus layer replaces the centralized registration center with a UDP broadcast mechanism and a P2P node discovery mechanism to perform service discovery and address synchronization among distributed nodes. Throughout the service registration and discovery process, the general data bus layer completes the parsing, matching, and verification of service metadata based on the unified resource descriptor specification.
[0010] Furthermore, the serialization and deserialization processing flow of the general serialization layer is as follows: Business code initiates serialization or deserialization requests through a unified API predefined by the generic serialization layer; The general serialization layer offers four binary serialization formats—Protobuf, MessagePack, FlatBuffers, and Cap'n'Proto—as well as two string serialization formats—JSON and BSON—through pluggable plugins. The framework selects the corresponding serialization format based on pre-configured rules or business scenarios. Specifically, binary serialization format is preferred for intranet communication scenarios, while string serialization format is preferred for public network communication scenarios. During serialization, the general serialization layer performs full parsing and conversion of basic data types and extended data types. The basic data types include STL collections, arrays, and built-in data types, while the extended data types include custom structures, nested structures, and generic data types.
[0011] Furthermore, the serialization process of the RPC request message also includes a large message splitting process based on a preset message data volume threshold; the deserialization process of the RPC request message also includes a corresponding message assembly process, as follows: When the data size of an RPC request message exceeds a preset data size threshold, the framework splits the RPC request message into a message header and a message body. The message header contains the unique index information, data length information, and verification information of the message body. After the split message body is serialized through the general serialization layer, it is stored in the general memory layer, and only the split message header is routed and transmitted. After receiving the message header of the RPC request, the server node retrieves the corresponding message body data by querying the general storage layer based on the unique index information in the message header. The server node performs integrity verification on the message body data based on the verification information in the message header. After the verification passes, the message header and message body are assembled, and then the assembled complete RPC request message is deserialized and the business logic is executed.
[0012] Furthermore, the storage adaptation and data processing of the general-purpose storage layer are adapted to four databases (MySQL, SQLite, PostgreSQL, and MongoDB), three file and object storage systems (local file, object storage OSS, and S3-compatible object storage MinIO), and two memory storage media (Redis and shared memory SHM) through pluggable plug-ins. The general-purpose storage layer exposes a unified CRUD operation API for all compatible storage media, and the business code does not need to be aware of the type, differences and switching logic of the underlying storage media; and it performs distributed transaction processing and data consistency control across different storage media through the two-phase commit (2PC) protocol or eventual consistency protocol.
[0013] Furthermore, the metadata format defined by the Uniform Resource Descriptor specification includes six items: service name, service method, parameter list, version number, communication address, and storage resource information. Its application process is as follows: When a distributed node starts up, it generates standardized service metadata according to the Uniform Resource Descriptor specification and registers the service. When a client initiates an RPC request, it generates matching rules for the target service according to the Uniform Resource Descriptor specification, and performs service discovery and node matching. When performing data operations across storage media, access rules for storage resources are generated according to the Uniform Resource Descriptor specification, and unified data read and write operations are completed through the general-purpose memory layer.
[0014] Furthermore, the full lifecycle management process for the pluggable plug-in is as follows: Based on the predefined unified abstract interface of the plugin's layer, the plugin's functions are developed and encapsulated to generate an independent plugin package. The framework then performs format validation and interface compatibility validation on the generated plugin package. After the validation passes, the plugin is registered and integrated into the architecture. During plugin runtime, the framework calls the plugin's functional logic through a unified abstract interface, and the business code and core architecture code do not directly call the plugin's internal logic; When performing a plugin replacement or uninstallation operation, the framework first stops the traffic scheduling of the corresponding plugin, persists the plugin's running state, and then performs the plugin replacement or uninstallation operation. During this operation, no modifications are made to the core architecture code or business code.
[0015] (III) Beneficial Effects 1. Through a four-layer decoupled architecture design consisting of an RPC scheduling layer, a general data bus layer, a general serialization layer, and a general storage layer, the four core functions of communication, serialization, storage, and scheduling are completely isolated through a predefined unified abstract interface. There is no business logic coupling between the layers. This allows developers to independently replace the underlying implementation of any layer according to business needs, requiring only the replacement of the corresponding layer's plugin. There is no need to modify the upper-layer business code or the core logic of the framework. This solves the technical problems of protocol locking and strong coupling between the underlying implementation and the core logic in existing frameworks, enabling seamless and flexible switching of underlying capabilities and significantly improving the system's multi-scenario adaptability. 2. Through a pluggable extension mechanism across the entire chain, all underlying capabilities are encapsulated as independent plugins. When adding new communication protocols, serialization formats, storage media, or registration center adaptation capabilities, developers only need to implement the predefined unified abstract interface at the corresponding layer to integrate the new capabilities into the framework core as plugins. There is no need to make any intrusive modifications to the framework's core code. At the same time, the entire process of adding, replacing, and uninstalling plugins does not require restarting service nodes or affecting the normal operation of online business, reducing the technical cost of system iteration and expansion and significantly improving the response efficiency of business technology iteration.
[0016] 2. Through the predefined unified resource descriptor specification, the entire chain of service and resource metadata is standardized. Service registration, service discovery, and resource access of all distributed nodes are completed based on this specification, enabling seamless and transparent interoperability between cloud-native clusters, edge embedded devices, and cross-platform heterogeneous services. At the same time, the general data bus layer natively supports all types of network communication protocols, from local inter-process communication such as IPC and shared memory to TCP, HTTP, and UDP. It can be seamlessly deployed in all scenarios, from embedded edge terminals to large-scale cloud-native clusters, and the system's full-scenario compatibility and cross-technology stack interoperability are comprehensively enhanced.
[0017] 3. The RPC scheduling layer natively integrates synchronous and asynchronous dual-mode scheduling, message caching, batch processing, and end-to-end service governance capabilities. Without introducing any third-party components, it can achieve end-to-end scheduling and control including load balancing, circuit breaking and degradation, rate limiting and retrying, end-to-end tracing, and automatic anomaly recovery. Therefore, in actual business operations, asynchronous requests achieve batch scheduling and traffic shaping through a local message cache queue, improving system throughput in high-concurrency scenarios; the end-to-end call state is persisted through a general-purpose storage layer, enabling automatic recovery from request anomalies and end-to-end observability; and the redundancy adaptation capability for multiple registration centers and storage media is addressed from an architectural perspective. This avoids the risk of single points of failure. Moreover, we provide a unified abstract API for the three core capabilities of communication, serialization, and storage. Developers do not need to delve into the technical details of various underlying protocols, serialization libraries, and storage media. They can achieve full-link operations across protocols, formats, and media simply by calling the unified API, which greatly reduces the development threshold and adaptation cost of distributed systems. At the same time, the four-layer decoupled architecture design and plug-in code organization mode make the system's code structure boundaries clear and responsibilities distinct. When problems occur, the corresponding plug-in module can be quickly located without having to investigate the entire code chain, which significantly improves the efficiency of distributed system engineering implementation. Attached Figure Description
[0018] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, the preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0019] Figure 1 This is an overall architecture diagram of an embodiment of the present invention; Figure 2 This is a flowchart illustrating the information distribution process in an embodiment of the present invention. Detailed Implementation
[0020] This application provides a general distributed service scheduling and data processing method to solve the technical problems of protocol locking, high adaptation costs, complex storage logic, and insufficient scheduling capabilities in the prior art. The invention will be further described in detail below with reference to specific embodiments. The general distributed service scheduling and data processing method provided by this invention is implemented based on a four-layer decoupled distributed processing architecture consisting of an RPC scheduling layer, a general data bus layer, a general serialization layer, and a general memory layer, as detailed below: Example 1 This embodiment addresses three core application scenarios: cloud-native enterprise-level microservice clusters, edge computing distributed nodes, and embedded cross-process communication. It implements a general distributed service scheduling and data processing method, primarily consisting of a four-layer core abstract architecture, a pluggable extension mechanism, and a unified resource descriptor specification. The details are as follows: Before implementation, a distributed processing architecture is built and deployed. The architecture consists of four layers: an RPC scheduling layer, a general data bus layer, a general serialization layer, and a general storage layer. Each layer has a predefined unified abstract interface, and interactions between layers are only completed through these abstract interfaces, with no business logic coupling. For the underlying capabilities of each layer, corresponding pluggable plugins are pre-developed. All plugins implement the unified abstract interface of their respective layer and are encapsulated as independent .so / .dll dynamic link library packages. Specifically, these include plugins for the general data bus layer (AMQP, HTTP, IPC, TCP, UDP protocols), Consul, Redis, etcd, ZooKeeper, KubernetesService, and DNS registry center, as well as UDP broadcast node discovery plugins and P2P node discovery plugins adapted for edge scenarios; and Protobuf and Messa plugins for the general serialization layer. The system includes plugins for binary serialization (gePack, FlatBuffers, Cap'n'Proto), string serialization (JSON, BSON); database plugins for the general storage layer (MySQL, SQLite, PostgreSQL, MongoDB), local file system, OSS object storage, MinIO object storage, Redis memory storage, and shared memory SHM); load balancing plugins for the RPC scheduling layer (round-robin, weighted, consistent hashing, minimum connection count); and circuit breaker, rate limiting, and retry strategy plugins. It also predefines a unified resource descriptor specification, whose metadata format includes six items: service name, service method, parameter list, version number, communication address, and storage resource information. This specification serves as the sole standardized metadata format for all services and resources in the architecture, and all distributed nodes adhere to it to encapsulate and parse service metadata.
[0021] When the architecture starts, it first scans the plugin deployment directory and performs format and interface compatibility checks on each plugin package. If the checks pass, the plugin is registered and integrated into the core architecture. Plugins that fail the checks are logged and not integrated. For different application scenarios, the plugins and configurations are adapted and adjusted simultaneously. In cloud-native microservice cluster scenarios, the TCP protocol plugin, HTTP protocol plugin, Consul registry plugin, Protobuf serialization plugin, JSON serialization plugin, MySQL plugin, and Redis plugin are loaded by default to meet the core requirements of high throughput and high availability of the cluster. In edge computing distributed node scenarios, the UDP protocol plugin, P2P node discovery plugin, MessagePack serialization plugin, and SQLite plugin are loaded by default to meet the deployment requirements of low latency and weak network environments of edge nodes. In embedded cross-process communication scenarios, the IPC protocol plugin, shared memory SHM plugin, and FlatBuffers serialization plugin are loaded by default to meet the performance requirements of low power consumption and small size of embedded devices.
[0022] When client and server nodes in a distributed system start up, they sequentially complete the layered initialization of the general data bus layer, general serialization layer, general storage layer, and RPC scheduling layer. After each layer is initialized, it reports its readiness status to the architecture core. Nodes encapsulate their service metadata according to the Uniform Resource Descriptor (URI) specification. For example, the metadata encapsulated by the order server node is: Service Name = order-service, Service Method = createOrder, Parameter List = [orderInfo: OrderStruct, userId: String], Version Number = 1.0.0, Communication Address = 192.168.1.100:8080. Subsequently, the nodes... Through the general data bus layer, based on the pre-configured registry center plugin, the corresponding registry center is accessed. Standardized service metadata is reported to the registry center to complete service registration. The registry center performs format verification, version verification, and parameter verification on the service metadata based on the unified resource descriptor specification. After the verification is passed, the node information is stored and synchronized within the cluster. The node completes the entire process initialization and enters the service-ready state. For edge decentralized deployment scenarios, the node does not need to access the centralized registry center. It broadcasts its own standardized service metadata within the local area network through the UDP broadcast node discovery plugin, while listening to the broadcast information of other nodes, completing service discovery and address synchronization between distributed nodes, and realizing decentralized edge cluster deployment.
[0023] Once the client node is ready, it receives the user's business request and executes the complete RPC request processing flow. First, the client node converts the user's business request into a standard RPC request message and calls the general serialization layer to perform serialization processing on the RPC request message. The business code initiates a serialization request through the predefined unified serialize() API of the general serialization layer, without needing to specify a specific serialization format. The framework automatically selects the optimal serialization format based on pre-configured scenario rules: Protobuf binary serialization format for intranet microservice communication scenarios, JSON string serialization format for public network cross-platform calls, and FlatBuffers zero-copy serialization format for embedded scenarios. The general serialization layer calls the corresponding format serialization plugin to serialize the RPC request message. During processing, it supports full parsing and conversion of basic data types such as STL collections, arrays, and built-in data types, as well as extended data types such as custom structures, nested structures, and generic data types. During serialization, the framework monitors the data size of the RPC request message in real time. Assuming a preset data size threshold of 1MB, if the data size of the RPC request message exceeds 1MB... During MB (Mean-of-the-Moment) processing, the framework automatically performs large message splitting. First, it automatically separates the message header and body of the RPC request message. The message body is serialized and stored in a general-purpose storage. After the client queries the registry for the target server's communication address, it sends only the RPC request header to the server node. Upon receiving the RPC request, the server node queries the general-purpose storage for the message body data and assembles the message. Then, it performs relevant business processing based on the assembled complete message content. After processing, it constructs an RPC response message and sends it to the client node. The client node, upon receiving the RPC response message, continues to execute subsequent processing logic based on the processing results in the message. The split message header contains the message body's unique UUID index information, data length information, and MD5 checksum information. The split message body, after serialization, is stored in a pre-configured Redis storage medium. Only the message header carrying the index information is routed, significantly reducing the amount of data transmitted over the network and avoiding network congestion and request timeouts caused by large message transmission. After receiving the message header, the server retrieves the message body from the general-purpose storage based on the index information, performs integrity verification using the MD5 checksum information, assembles the complete message, and then performs subsequent deserialization and business processing.
[0024] After serialization, the client node queries the registry center through the general data bus layer to obtain the communication address and service metadata information of the target server node. Based on the pre-configured invocation strategy of the RPC scheduling layer, it routes the serialized RPC request message to the target server node. The client node first generates matching rules for the target service according to the Uniform Resource Descriptor specification and sends a query request to the registry center. After the registry center returns a list of server nodes that match the matching rules, the client node selects the target server node based on the pre-configured weighted load balancing strategy and determines the type of RPC request. If it is a synchronous request, the RPC request message is directly sent to the registry center. The RPC request message is sent to the target server node via the TCP protocol plugin of the general data bus layer. If it is an asynchronous request, the RPC request message is first written to the local memory message cache queue. The scheduler performs batch scheduling and distribution of the requests in the message cache queue according to the preset priority rules. During the entire scheduling process, the RPC scheduling layer records the call chain data, timeout status data and retry status data of the RPC request in real time through the general memory layer. Based on this data, automatic recovery of request anomalies and full-link tracing are realized. All link data can be exposed to the outside world through a unified API to connect with monitoring systems such as Prometheus and Grafana to achieve observability.
[0025] After receiving an RPC request message, the server node first completes the message body acquisition and assembly in the case of large messages, then calls the general serialization layer to perform deserialization processing on the complete RPC request message, parses the business request parameters, executes the corresponding business logic processing, generates an RPC response message and completes serialization processing, and sends it back to the client node through the general data bus layer. After receiving the response message, the client node completes deserialization, parses the business processing result, and executes subsequent business logic.
[0026] In addition, the pre-configured call strategy of the RPC scheduling layer is effective throughout the entire request process. The circuit breaker and degradation strategy uses real-time statistics of the timeout rate and error rate of RPC requests. When the request error rate exceeds 50% within 1 minute, the circuit breaker is triggered. After the circuit breaker is triggered, new requests are processed for fast failure degradation. After 5 seconds of circuit breaker, it enters a half-open state, allowing a small number of probe requests to pass. If the request success rate recovers to above 90%, the circuit breaker is closed and normal scheduling is restored. The rate limiting strategy is implemented based on the token bucket algorithm, with a preset QPS limit of 10,000 per node. When the number of requests exceeds the QPS limit, the excess requests are either queued or rejected directly. The retry strategy is pre-configured with 3 retries and a retry interval of 100ms. The retry trigger conditions are request timeout, network error, and server 5xx error. Failed requests that meet the trigger conditions are automatically retried. The above load balancing, circuit breaker and degradation, rate limiting, and retry strategies can all be flexibly adjusted by replacing the corresponding plugins.
[0027] When a business needs to operate on multiple heterogeneous storage media simultaneously, distributed transaction processing across storage media is achieved through a general-purpose storage layer. The business code initiates cross-storage transaction requests through the unified CRUDAPI of the general-purpose storage layer without having to worry about the type, differences and switching logic of the underlying storage media. The general-purpose storage layer completes the transaction commit and data consistency guarantee across storage media through a two-phase commit (2PC) protocol or an eventual consistency protocol, adapting to business scenarios where multiple storage media coexist.
[0028] When adding or replacing underlying capabilities, a non-intrusive hot-swap operation is achieved through a plug-in mechanism. Taking the replacement of the Protobuf serialization plugin with the FlatBuffers serialization plugin as an example, the framework first stops the traffic scheduling of the Protobuf plugin, temporarily switches the serialization requests to the JSON serialization plugin to ensure uninterrupted business operations, and performs the plugin uninstallation operation after completing the state persistence of the Protobuf plugin. No modification to the core framework code or business code is required. Then, the new FlatBuffers plugin package is loaded, and after completing format validation and interface compatibility validation, the plugin registration and access are completed. Finally, the serialization requests are switched to the new FlatBuffers plugin. The entire process does not require restarting distributed nodes and does not affect the normal operation of online business. When adding communication protocols, serialization formats, storage media, or registration center types, it is only necessary to implement the corresponding unified abstract interface and encapsulate it as an independent plugin to access the core architecture without intrusive modification of the core framework code.
[0029] The system incorporates a comprehensive exception handling mechanism covering all aspects of data transmission, plugin operation, and node communication. When any of the following exception scenarios are triggered: network anomaly, plugin loading failure, or node disconnection, the corresponding response strategy is immediately activated. In the event of network timeout or data packet loss, a retry strategy is automatically triggered. If three retries fail, a backup communication protocol and backup service node are switched, and an exception log is recorded. When a plugin loading failure or interface incompatibility is detected, the system automatically rolls back to the previous version of the available plugin and generates a plugin exception warning to the management platform without interrupting the current business process. When a server node is detected to be disconnected or its heartbeat timeout exceeds 3 seconds, the node is automatically removed from the list of available nodes, traffic is switched to other healthy nodes, and node status is synchronized through the registry center to ensure the stable operation of the entire system.
[0030] Example 2 This embodiment mainly addresses the technical problems of strong coupling between the core logic and the underlying implementation of existing RPC frameworks, the need for intrusive modifications to the core code for extensions, and the inability to separate pure business logic from underlying technical details. It focuses on a four-layer decoupled architecture, a plug-in extension mechanism, and a unified resource descriptor specification.
[0031] The core architecture is divided into four layers: RPC scheduling layer, general data bus layer, general serialization layer, and general memory layer. Each layer defines a unified abstract interface of pure virtual functions. The layers interact only through the abstract interface, without any business logic coupling. This achieves complete decoupling of the four core capabilities of communication, serialization, storage, and scheduling. The underlying implementation of any layer can be replaced independently without modifying the business code, solving the problems of existing technology protocol lock-in, high adaptation cost, and poor scalability.
[0032] The RPC scheduling layer predefines the IRpcScheduler unified abstract interface, which contains four core pure virtual functions: SelectNode(), DispatchRequest(), RecordInvokeState(), and ExceptionRecover(). These functions are responsible for service node selection, request scheduling and distribution, call status recording, and exception recovery handling, respectively. Its core implementation is divided into synchronous scheduling unit, asynchronous scheduling unit, state persistence unit, full-link tracing unit, and policy execution unit. The synchronous scheduling unit, for synchronous requests, directly selects nodes and forwards requests based on pre-configured load balancing strategies. During forwarding, it monitors request status in real time and automatically executes retry strategies when timeouts or errors occur. The asynchronous scheduling unit, for asynchronous requests, uses a lock-free circular queue to implement local message caching for asynchronous decoupling. The scheduler employs a multi-threaded pool model, performing batch scheduling and distribution according to request priority, and supports both sequential and concurrent execution modes. The state persistence unit, through a unified API in the general storage layer, persists request call chains, timeouts, retryes, and error states to the storage medium in real time, enabling request breakpoint resumption and automatic recovery in abnormal scenarios. The strategy execution unit has built-in execution entry points for four call strategies: load balancing, circuit breaking and degradation, rate limiting, and retry. All strategies are implemented as plugins, accessible through a unified abstract interface, and can be flexibly replaced with alternative load balancing algorithms such as consistent hashing and least connections, as well as alternative degradation strategies such as rate limiting, circuit breaking, and retry, adapting to the scheduling needs of different business scenarios.
[0033] The general data bus layer predefines two unified abstract interfaces: ITransport and IServiceDiscovery. The ITransport interface defines four core pure virtual functions: Connect(), Send(), Receive(), and Close(), responsible for unified adaptation of the underlying transport protocol. The IServiceDiscovery interface defines four core pure virtual functions: Register(), Discover(), Deregister(), and Heartbeat(), responsible for unified adaptation of service registration and discovery. Its core implementation is divided into a transport protocol adaptation unit, a service registration and discovery unit, and a metadata parsing and verification unit. The transport protocol adaptation unit natively supports five transport protocols—AMQP, HTTP, IPC, TCP, and UDP—through pluggable plugins. When adding a custom transport protocol, only the four core functions of the ITransport interface need to be implemented, encapsulated as an independent plugin for integration. The business code does not need to be aware of the switching of the underlying transport protocol, solving the problem of protocol lock-in in existing technologies. The service registration and discovery unit adapts to six centralized registration centers—Consul, Redis, etcd, ZooKeeper, KubernetesService, and DNS—through pluggable plugins. It also implements UDP broadcast and P2P decentralized node discovery mechanisms for edge scenarios. All registration and discovery implementations follow the IServiceDiscovery interface specification and can be flexibly selected and replaced according to the deployment environment. The metadata parsing and verification unit, based on the unified resource descriptor specification, completes the parsing, matching, and verification of service metadata, ensuring that the format and semantics of service metadata are uniform across different registration centers and transport protocols, thus guaranteeing the accuracy of service calls. This specification defines a unified service and resource metadata format, which is the core foundation for achieving transparent interoperability across protocols and storage, solving the problem of the lack of unified specifications and interoperability difficulties in service and resource metadata in existing technologies.
[0034] The generic serialization layer predefines the ISerializer unified abstract interface, which defines two core pure virtual functions, Serialize() and Deserialize(), responsible for unified adaptation of serialization and deserialization operations. Its core implementation consists of a unified API entry point, a format adaptation unit, a type resolution unit, and a scenario-based automatic selection unit. The unified API entry point provides business code with a single serialize() and deserialize() API. Business code does not need to specify the serialization format; it only needs to pass in the data to be processed and the target type to complete the operation, shielding the differences in underlying serialization libraries and significantly reducing development adaptation costs. The format adaptation unit supports Protobuf, MessagePack, and FlatBull through pluggable plugins. The system supports four binary serialization formats: ffers, Cap'n'Proto, and two string serialization formats: JSON and BSON. Adding a custom serialization format only requires implementing two core functions of the ISerializer interface. The type parsing unit is implemented based on template metaprogramming, natively supporting basic types such as STL collections, arrays, and built-in data types. It also supports full parsing and conversion of extended types such as custom structs, nested structs, and generic data types, without requiring additional adaptation logic in business code. The scenario-based automatic selection unit has a built-in scenario rule engine that automatically selects the optimal serialization format based on intranet / public network scenarios, data volume, and performance requirements. In intranet scenarios, binary formats are prioritized to ensure high performance, while in public network scenarios, string formats are prioritized to ensure compatibility.
[0035] The general-purpose storage layer predefines the IStorage unified abstract interface, which defines four core pure virtual functions: Create(), Read(), Update(), and Delete(), i.e., a unified CRUD operation API. It also defines three transaction-related pure virtual functions: BeginTransaction(), Commit(), and Rollback(), responsible for unified adaptation across storage transactions. Its core implementation is divided into a unified CRUD interface unit, a storage media adaptation unit, and a distributed transaction unit: the unified CRUD interface unit exposes completely consistent read and write APIs for all adapted storage media. Business code only needs to call the unified interface to complete data read and write operations, without needing to know whether the underlying storage media is a relational database. The system addresses the issues of inconsistent storage interfaces and complex adaptation in existing technologies, offering both object storage and in-memory databases. The storage media adaptation unit, through pluggable plugins, supports four databases (MySQL, SQLite, PostgreSQL, and MongoDB), three file and object storage systems (local files, OSS, and MinIO), and two in-memory storage media (Redis and SHM shared memory). Adding a new storage medium only requires implementing the core functions of the IStorage interface. The distributed transaction unit incorporates both two-phase commit (2PC) and eventual consistency protocols, allowing for flexible selection based on business needs. It also enables distributed transaction processing across different storage media, resolving the data consistency guarantee challenge in multi-storage coexistence scenarios.
[0036] The Unified Resource Descriptor (URD) specification uses JSON Schema to define a standardized metadata format. Core fields include service name, service method, parameter list, version number, communication address, and storage resource information. Its application spans the entire system lifecycle: when distributed nodes start, standardized service metadata is generated according to the specification to complete service registration. The registry center only receives metadata that conforms to the specification, ensuring metadata consistency from the source. When a client initiates an RPC request, matching rules for the target service are generated according to the specification. The registry center uses the specification fields to perform accurate service matching and version filtering, ensuring that requests are forwarded to the correct service node. During cross-storage media data operations, access rules for storage resources are generated according to the specification. The general storage layer automatically matches the corresponding storage plugin based on the storage resource information in the specification to complete unified data read and write operations. This specification masks the differences in underlying technology implementations, enabling transparent interaction between different underlying technologies.
[0037] Meanwhile, the plug-in extension mechanism enables full lifecycle management of underlying capabilities. During the plug-in development phase, developers need to implement the predefined unified abstract interface at the corresponding level, compile it into an independent dynamic link library, and supplement the metadata description file. During the plug-in access phase, the framework completes the format, interface compatibility, and dependency verification of the plug-in. After the verification is passed, registration and dynamic loading are completed. During the plug-in runtime phase, the framework only calls the plug-in function through the unified abstract interface. Business code and the core framework do not directly access the internal implementation of the plug-in. During the plug-in replacement / uninstallation phase, the framework first completes traffic segmentation and state persistence, and then performs hot replacement or uninstallation of the plug-in. The entire process does not require modification of the core framework and business code, nor does it require restarting service nodes. It truly realizes a stable core and pluggable extension mode.
[0038] Finally, it should be noted that the above embodiments are merely examples for clearly illustrating the present invention and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A general distributed service scheduling and data processing method, characterized in that, The steps for this processing method are as follows: A layered and decoupled distributed processing architecture is constructed, which includes an RPC scheduling layer, a general data bus layer, a general serialization layer, and a general storage layer. The layers interact through a predefined unified abstract interface, and there is no business logic coupling between the layers. The underlying functions of each layer are connected to the architecture in the form of pluggable plug-ins. A unified resource descriptor specification is predefined synchronously, which is used to standardize the metadata format of services and resources. When client nodes and server nodes in a distributed system start up, they encapsulate their own service metadata through the unified resource descriptor specification, and access the adapted registry center through the general data bus layer to register services and initialize node information. The client node receives user service requests, converts them into standard RPC request messages, and calls the general serialization layer to perform serialization processing on the RPC request messages; The client node queries and obtains the communication address and service metadata information of the target server node from the registry center through the general data bus layer, and routes the serialized RPC request message to the target server node based on the pre-configured calling strategy of the RPC scheduling layer. After receiving an RPC request message, the server node performs deserialization processing on the RPC request message through the general serialization layer to obtain the business request parameters. Then, it performs business logic processing based on the business request parameters, generates an RPC response message, and sends it back to the client node after serialization. After receiving the RPC response message, the client node performs deserialization processing on the RPC response message to obtain the business processing result, and then performs subsequent business logic execution based on the business processing result.
2. The general distributed service scheduling and data processing method according to claim 1, characterized in that, The RPC scheduling layer performs the following scheduling and processing steps for RPC requests: When a client node initiates an RPC request message, it first determines whether the request type of the RPC request message is a synchronous request or an asynchronous request. If it is determined to be a synchronous request, the RPC request message is forwarded to the target server node directly according to the pre-configured calling strategy. If it is determined to be an asynchronous request, the RPC request message is first written to the local message cache queue, and the scheduler performs batch scheduling and distribution processing on the RPC request messages in the message cache queue according to the preset priority rules. During the scheduling process, the RPC scheduling layer records the call chain data, timeout status data, and retry status data of RPC requests in real time through the general memory layer, and realizes request anomaly recovery and full-link tracing based on the recorded data.
3. The general distributed service scheduling and data processing method according to claim 2, characterized in that, The pre-configured call strategies in the RPC scheduling layer include four types: load balancing strategy, circuit breaking and degradation strategy, rate limiting strategy, and retry strategy; as detailed below: The load balancing strategy is selected from any one of round-robin, weighted, consistent hashing, and least connections. The circuit breaker and degradation strategy triggers the circuit breaker operation by statistically analyzing the timeout rate and error rate of RPC requests, and performs fast failure degradation processing after the circuit breaker is triggered. The rate limiting strategy is implemented based on the token bucket algorithm or the leaky bucket algorithm to control the concurrency and QPS of RPC requests. The retry strategy is pre-configured with the number of retries, the retry interval, and the retry trigger conditions, and automatically retryes failed RPC requests that meet the trigger conditions.
4. The general distributed service scheduling and data processing method according to claim 1, characterized in that, The steps for adapting the transmission protocol of the general data bus layer are as follows: The general data bus layer natively includes five transport protocols: AMQP, HTTP, IPC, TCP, and UDP, through pluggable plugins. When adding a custom transmission protocol, the custom transmission protocol is encapsulated as an independent pluggable plug-in access architecture based on the unified transmission abstraction interface predefined in the general data bus layer. The business code does not need to be aware of the type and switching logic of the underlying transmission protocol.
5. The general distributed service scheduling and data processing method according to claim 4, characterized in that, The service registration and discovery process of the general data bus layer is that the general data bus layer adapts to six registration centers, namely Consul, Redis, etcd, ZooKeeper, Kubernetes Service, and DNS, through pluggable plugins. Among them, for edge deployment scenarios, the general data bus layer replaces the centralized registration center with UDP broadcast mechanism and P2P node discovery mechanism to perform service discovery and address synchronization among distributed nodes. Throughout the service registration and discovery process, the general data bus layer completes the parsing, matching, and verification of service metadata based on the unified resource descriptor specification.
6. The general distributed service scheduling and data processing method according to claim 1, characterized in that, The serialization and deserialization process of the general serialization layer is as follows: Business code initiates serialization or deserialization requests through a unified API predefined by the generic serialization layer; The general serialization layer offers four binary serialization formats—Protobuf, MessagePack, FlatBuffers, and Cap'n'Proto—as well as two string serialization formats—JSON and BSON—through pluggable plugins. The framework selects the corresponding serialization format based on pre-configured rules or business scenarios. For intranet communication scenarios, binary serialization format is preferred, while for public network communication scenarios, string serialization format is preferred. During serialization, the general serialization layer performs full parsing and conversion of basic data types and extended data types. The basic data types include STL collections, arrays, and built-in data types, while the extended data types include custom structures, nested structures, and generic data types.
7. The general distributed service scheduling and data processing method according to claim 6, characterized in that, The serialization process of the RPC request message also includes a large message splitting process based on a preset message data volume threshold; the deserialization process of the RPC request message also includes a corresponding message assembly process, as follows: When the data size of an RPC request message exceeds a preset data size threshold, the framework splits the RPC request message into a message header and a message body. The message header contains the unique index information, data length information, and verification information of the message body. After the split message body is serialized through the general serialization layer, it is stored in the general memory layer, and only the split message header is routed and transmitted. After receiving the message header of the RPC request, the server node retrieves the corresponding message body data by querying the general storage layer based on the unique index information in the message header. The server node performs integrity verification on the message body data based on the verification information in the message header. After the verification passes, the message header and message body are assembled, and then the assembled complete RPC request message is deserialized and the business logic is executed.
8. The general distributed service scheduling and data processing method according to claim 1, characterized in that, The storage adaptation and data processing of the general-purpose storage layer are achieved through pluggable plug-ins, which are compatible with four databases: MySQL, SQLite, PostgreSQL, and MongoDB; three file and object storage systems: local file, object storage OSS, and S3-compatible object storage MinIO; and two memory storage media: Redis and shared memory SHM. The general-purpose storage layer exposes a unified CRUD operation API for all compatible storage media, and the business code does not need to be aware of the type, differences and switching logic of the underlying storage media; and it performs distributed transaction processing and data consistency control across different storage media through the two-phase commit (2PC) protocol or eventual consistency protocol.
9. The general distributed service scheduling and data processing method according to claim 1, characterized in that, The metadata format defined by the Uniform Resource Descriptor specification includes six items: service name, service method, parameter list, version number, communication address, and storage resource information. Its application process is as follows: When a distributed node starts up, it generates standardized service metadata according to the Uniform Resource Descriptor specification and registers the service. When a client initiates an RPC request, it generates matching rules for the target service according to the Uniform Resource Descriptor specification, and performs service discovery and node matching. When performing data operations across storage media, access rules for storage resources are generated according to the Uniform Resource Descriptor specification, and unified data read and write operations are completed through the general-purpose memory layer.
10. The general distributed service scheduling and data processing method according to claim 1, characterized in that, The full lifecycle management process for the pluggable plug-in is as follows: Based on the predefined unified abstract interface of the plugin's layer, the plugin's functions are developed and encapsulated to generate an independent plugin package. The framework then performs format validation and interface compatibility validation on the generated plugin package. After the validation passes, the plugin is registered and integrated into the architecture. During plugin runtime, the framework calls the plugin's functional logic through a unified abstract interface, and the business code and core architecture code do not directly call the plugin's internal logic; When performing a plugin replacement or uninstallation operation, the framework first stops the traffic scheduling of the corresponding plugin, persists the plugin's running state, and then performs the plugin replacement or uninstallation operation. During this operation, no modifications are made to the core architecture code or business code.