JDK8-based heterogeneous service large model dynamic adaptation and intelligent scheduling method and system

Through the dynamic adaptation and intelligent scheduling method of heterogeneous service large model based on JDK8, the compatibility, protocol differences and intelligent scheduling problems of multi-platform integration in the JDK8 environment are solved, efficient service scheduling and elastic expansion in low-version environments are achieved, development and operation and maintenance costs are reduced, and system availability and maintainability are improved.

CN120281824APending Publication Date: 2025-07-08浪潮智慧城市科技有限公司 +1
View PDF 0 Cites 14 Cited by

Patent Information

Application Number
CN202510534918.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-08

Smart Images

  • Figure CN120281824A_ABST
    Figure CN120281824A_ABST
Patent Text Reader

Abstract

The invention discloses a JDK8-based heterogeneous service large model dynamic adaptation and intelligent scheduling method and system, and belongs to the technical field of artificial intelligence service integration, and the method comprises a unified service module which provides standardized HTTP / WebSocket access for an application layer and shields the difference of multiple platforms at the bottom layer; the dynamic routing module is used for dynamically selecting service nodes based on service quality indexes and supporting weight calculation, fusing recovery and flow dyeing; the service execution module is used for loading a platform adapter through a dual-mode service factory, integrating a tool calling function and realizing protocol conversion and authentication injection; and the configuration management module supports dual-configuration source hot loading of the static file and the dynamic database and provides versioning rollback and consistency verification. According to the method, the problems of multi-platform protocol difference, service rigid expansion, tool calling coupling and the like can be solved, the complexity of multi-platform management is remarkably reduced, and the flexibility and maintainability of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence service integration, and specifically to a method and system for dynamic adaptation and intelligent scheduling of heterogeneous service large models based on JDK8. Background Art

[0002] With the wide application of artificial intelligence technology, the value of intelligent service integration systems has become increasingly prominent in fields such as finance, healthcare, and industry. However, in the actual implementation process, they still face multiple technical bottlenecks. Existing systems are generally limited by the high-version Java environment dependence. The mainstream frameworks require JDK17 and above versions, resulting in difficulty in upgrading a large number of enterprise legacy systems running in the JDK8 environment. This forces developers to maintain two sets of technology stacks or abandon the iteration of new functions, severely restricting the efficiency of technology evolution.

[0003] At the same time, the significant differences in multi-platform protocols have increased the integration complexity. Different service providers adopt heterogeneous protocols such as HTTP / 1.1, gRPC, and WebSocket, and the data formats cover multiple standards such as JSON, XML, and Protobuf. Developers need to independently develop adaptation code for each platform, resulting in up to 60% of duplicate workload, and it is easy to cause system-level compatibility risks when the protocol changes. The problem of rigid service expansion is particularly prominent. The traditional hard-coded method deeply couples the platform implementation class with the business logic. Adding new services requires modifying the core code and redeploying, and configuration updates rely on downtime operations, making it difficult to meet the dynamic scaling requirements in high-concurrency scenarios.

[0004] In addition, the tool call mechanism lacks a standardized design. The function metadata is scattered in the business code, and it is impossible to dynamically generate interface descriptions that conform to the large model interaction specifications. The load balancing strategy based on simple polling lacks multi-dimensional perception of service quality. It can neither dynamically allocate traffic according to real-time response time and error rate, nor achieve second-level switching in case of node failure, resulting in low resource utilization and potential service stability risks.

[0005] These systematic defects have brought multiple challenges to enterprises in the intelligent transformation, such as high development costs, increased operation and maintenance complexity, and business continuity risks. There is an urgent need to build a new generation of technology system with low-version compatibility, protocol adaptive conversion, dynamic service expansion, and intelligent traffic scheduling. Summary of the Invention

[0006] The technical task of the present invention is to address the above deficiencies and provide a method and system for dynamic adaptation and intelligent scheduling of heterogeneous service large models based on JDK8, which can solve problems such as multi-platform protocol differences, rigid service expansion, and tool call coupling, and provide a highly compatible, highly elastic, and low-maintenance-cost large model service integration solution for fields such as finance, healthcare, and industry, significantly reducing the complexity of multi-platform management and improving the system elasticity and maintainability.

[0007] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0008] A heterogeneous service large model dynamic adaptation and intelligent scheduling method based on JDK8, the implementation of this method includes:

[0009] A unified service module that provides standardized HTTP / WebSocket access for the application layer, shielding the differences of multiple underlying platforms;

[0010] A dynamic routing module that dynamically selects service nodes based on service quality (QoS) indicators, supporting weight calculation, circuit breaker recovery, and traffic coloring;

[0011] A service execution module that loads platform adapters through a dual-mode service factory (ServiceFactory), integrates tool call functions, and implements protocol conversion and authentication injection;

[0012] A configuration management module that supports hot loading of dual configuration sources of static files and dynamic databases, providing versioned rollback and consistency verification.

[0013] This method realizes seamless integration of multiple platforms and efficient service scheduling in a low-version environment through a protocol conversion middleware, a dynamic function call engine, and a dual-mode service factory.

[0014] Furthermore, for the dynamic routing module,

[0015] The weight calculation formula is as follows:

[0016] Weight = α × (1 / response time) + β × (1 - error rate) + γ × (1 / cost); where α + β + γ = 1, by default α = 0.6, β = 0.3, γ = 0.1;

[0017] For the circuit breaker recovery, the circuit breaker mechanism is triggered when the node error rate exceeds 50%, and automatically switches to the standby service;

[0018] For the traffic coloring, it supports the gray release strategy and distributes traffic according to request tags (such as X-Env:canary).

[0019] Furthermore, for the dual-mode service factory,

[0020] In the static mode, service parameters are predefined through YAML / database, supporting environment variable injection;

[0021] In the dynamic mode, third-party adapters are loaded based on the Java SPI mechanism and are registered and take effect immediately at runtime;

[0022] The service instance pool adopts a lazy initialization strategy to create connections on demand and reduce resource consumption.

[0023] Furthermore, the configuration management module,

[0024] A double buffer mechanism is used to implement hot configuration updates, and atomic switching is used to ensure service continuity.

[0025] The version manager records historical snapshots and supports quick rollback by version number;

[0026] Sensitive fields (such as API Key) use memory encryption and regular rotation strategies.

[0027] Furthermore, this method ensures JDK8 compatibility in the following ways:

[0028] Use Retrolambda to convert Lambda expressions to anonymous inner class bytecode;

[0029] Rewrite Java Stream API calls through ASM to replace JDK11+ native implementations;

[0030] OkHttp3 is used to replace Java HttpClient, supporting HTTP / 2 and connection pool reuse.

[0031] Furthermore, when a user initiates a request, the method performs the following steps for intelligent scheduling:

[0032] Step 1: The application uses unified service capabilities to assemble a unified intermediate format UnifiedRequest, calls the system to send a large model request, and performs the following operations:

[0033] Unified entry: receive HTTP / WebSocket and other protocol requests, and complete protocol-independent parsing;

[0034] Traffic scheduling: select the optimal node according to the load strategy (weight, response time) and forward the request to the dynamic routing module;

[0035] Request records: Request records are processed persistently to ensure the subsequent provision of full-link monitoring, log tracking and alarm systems to ensure service stability;

[0036] Step 2: The unified requests enter the dynamic routing module and perform intelligent scheduling:

[0037] Health assessment: Get real-time indicators (response time, error rate, cost weight) of candidate nodes from Redis;

[0038] Weight calculation: perform weight calculation and dynamically sort nodes;

[0039] Optimal node allocation: Select the service node with the highest weight (such as DeepSeek). If this node is in a fusing state, automatically switch to the standby node (such as Zhipu AI).

[0040] Step 3: After the routing decision, the request is passed to the service execution module and the following operations are performed:

[0041] Platform adapter invocation: Obtain the target platform adapter (such as DeepSeekService) through the dual-mode service factory (ServiceFactory), and inject configuration parameters (API key, endpoint).

[0042] Result aggregation: Integrate the large model response with the tool execution result to generate the final answer (such as "It is sunny in Beijing today, 25°C").

[0043] Step 4: After the service execution layer generates the result, the system performs reverse protocol adaptation and the following operations:

[0044] Data deserialization: Convert the platform-specific response (such as the JSON of DeepSeek) into UnifiedResponse.

[0045] Protocol reverse bridging: If the user request protocol is WebSocket, disassemble the HTTP response into long connection shard transmissions.

[0046] Return in unified format: The final data is encapsulated according to the user's initial protocol to ensure seamless parsing by the client.

[0047] The present invention also claims to protect a heterogeneous service large model dynamic adaptation and intelligent scheduling system based on JDK8, including:

[0048] Unified service module, used to provide standardized HTTP / WebSocket access for the application layer, shielding the underlying multi-platform differences;

[0049] Dynamic routing module, used to dynamically select service nodes based on service quality (QoS) metrics, supporting weight calculation, fusing recovery, and traffic coloring;

[0050] Service execution module, used to load the platform adapter through the dual-mode service factory, integrate the tool call function, and implement protocol conversion and authentication injection;

[0051] Configuration management module, supporting hot loading of dual configuration sources of static files and dynamic databases, providing versioned rollback and consistency verification;

[0052] This system realizes the dynamic adaptation and intelligent scheduling of heterogeneous service large models based on JDK8 through the above method.

[0053] Furthermore, the unified service module includes:

[0054] Unified traffic entrance: Provide multi-mode access capabilities, including RESTful API, message queue, event-driven, etc., and support access via HTTP / WebSocket protocols;

[0055] Global traffic management: Implement strategies such as load balancing, gray release, A / B testing, etc., and allocate resources according to business priorities;

[0056] Service governance center: Integrate monitoring, rate limiting, and fault tolerance capabilities, and provide a unified logging, link tracing, and alerting system;

[0057] The dynamic routing module includes the following units to achieve intelligent selection and traffic scheduling of service nodes:

[0058] Health monitoring unit: Real-time collect QoS metrics of service nodes, including response time, error rate, throughput, etc.;

[0059] Weight calculation unit: Dynamically calculate node weights based on multi-dimensional metrics to determine the optimal service node;

[0060] Routing decision unit: Allocate request traffic according to the weight sorting result, and support priority policies and load balancing;

[0061] Fuse management unit: Monitor the node failure status, trigger fusing and switch to standby nodes to ensure service continuity;

[0062] The service execution module interfaces with specific large model platforms to complete request forwarding, protocol adaptation, and result processing, and includes the following units:

[0063] Platform adapter unit: Package the communication protocols and API specifications of different platforms (such as ChatCompletion of OpenAI, Streaming API of DeepSeek);

[0064] Authentication processor unit: Dynamically inject the authentication information required by the platform (API Key, OAuth2 Token, signature);

[0065] Traffic controller unit: Control the request rate based on the token bucket algorithm to avoid triggering platform rate limiting;

[0066] Data serializer unit: Convert the unified request object into a platform-specific format (JSON / XML / Protobuf);

[0067] Error processor unit: Capture platform exceptions and trigger retry or fusing mechanisms;

[0068] The configuration management module, the configuration management layer realizes unified management of system configurations, and includes the following units:

[0069] Configuration Loader Unit: Load configuration data from files (YAML / JSON), databases, or configuration centers (such as Nacos).

[0070] Configuration Parser Unit: Convert the original configuration into Java objects (such as ServiceConfig, AuthConfig).

[0071] Configuration Memory Unit: Store the currently effective configuration using a double-buffering mechanism to avoid read-write conflicts.

[0072] Hot Update Listener Unit: Listen for configuration source change events and trigger dynamic updates.

[0073] Version Manager Unit: Record the historical versions of the configuration and support quick rollback and auditing.

[0074] The present invention also claims a heterogeneous service large model dynamic adaptation and intelligent scheduling device based on JDK8, including: at least one memory and at least one processor;

[0075] The at least one memory is used to store machine-readable programs;

[0076] The at least one processor is used to call the machine-readable program to implement the above method.

[0077] The present invention also claims a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the above method is implemented.

[0078] Compared with the prior art, the heterogeneous service large model dynamic adaptation and intelligent scheduling method and system of the present invention have the following beneficial effects:

[0079] 1. Compatibility with low-version environments reduces upgrade costs: Rewrite Java 8+ features (such as Lambda, Stream API) through bytecode downgrading technology (Retrolambda + ASM), and use OkHttp3 to replace high-version HTTP clients. Users can use modern AI capabilities without upgrading the JDK version, saving hardware upgrade and code refactoring costs, and extending the life cycle of old systems.

[0080] 2. Seamless integration across multiple platforms simplifies the development process: The protocol adaptation layer automatically completes HTTP / gRPC / WebSocket protocol conversion, unifies data model mapping (JSON / XML / Protobuf), shortens the access time for new platforms, and reduces code duplication by 60%; when upgrading platform interfaces, only the adapter configuration needs to be modified, with zero impact on core business logic.

[0081] 3. Intelligent elastic scheduling to ensure high service availability: The dynamic routing algorithm comprehensively considers response time, error rate, and cost coefficient, and combines the circuit breaker mechanism to achieve second-level fault switching. Through intelligent weight allocation, the idle rate of GPU computing resources is reduced, operating costs are saved, and when a node fails, it automatically switches to the standby service. The system availability is increased from 99% to 99.99%, supporting elastic expansion during peak business periods, and significantly improving the throughput of a single node.

[0082] 4. Zero-intrusive function expansion to accelerate business innovation: New platform adapters are loaded through the SPI mechanism. When adding new large model services in the future, the development cycle is significantly reduced; based on the isolation of tool functions from the core system, version iteration conflicts are reduced; third-party developers can contribute new adapters through standardized interfaces, facilitating the rapid expansion of the ecosystem. Brief Description of the Drawings

[0083] Figure 1 It is a flowchart showing the implementation process of the heterogeneous service large model dynamic adaptation and intelligent scheduling method provided by the embodiment of the present invention based on JDK8;

[0084] Figure 2 It is a flowchart showing the implementation process of the unified service module provided by the embodiment of the present invention;

[0085] Figure 3 It is a flowchart showing the implementation process of the dynamic routing module provided by the embodiment of the present invention;

[0086] Figure 4 It is a flowchart showing the implementation process of the service execution module provided by the embodiment of the present invention;

[0087] Figure 5 It is a flowchart showing the implementation process of the configuration management module provided by the embodiment of the present invention. Detailed Embodiment

[0088] The embodiment of the present invention provides a heterogeneous service large model dynamic adaptation and intelligent scheduling method based on JDK8. This method realizes the integration and scheduling of multi-protocol large models based on the following module combinations:

[0089] Unified service module, providing standardized HTTP / WebSocket access for the application layer, shielding the differences of underlying multi-platforms;

[0090] Dynamic routing module, dynamically selecting service nodes based on Quality of Service (QoS) indicators, supporting weight calculation, circuit breaker recovery, and traffic coloring;

[0091] Service execution module, loading platform adapters through a dual-mode service factory (ServiceFactory), integrating tool call functions, and implementing protocol conversion and authentication injection;

[0092] The configuration management module supports hot loading of dual configuration sources, namely static files and dynamic databases, and provides versioned rollback and consistency verification.

[0093] Among them, the dynamic routing module

[0094] The weight calculation formula is as follows:

[0095] Weight = α × (1 / response time) + β × (1 - error rate) + γ × (1 / cost), where α + β + γ = 1;

[0096] For the fuse recovery, the fuse mechanism is triggered when the node error rate exceeds 50%, and it automatically switches to the standby service;

[0097] For the traffic coloring, it supports the gray release strategy and distributes traffic according to request tags (such as X-Env:canary). Thus, intelligent traffic scheduling can be achieved.

[0098] For the dual-mode service factory, in the static mode, service parameters are predefined through YAML / database, and environment variable injection is supported; in the dynamic mode, third-party adapters are loaded based on the Java SPI mechanism, and registration takes effect immediately at runtime; the service instance pool adopts the lazy initialization strategy, creating connections on demand to reduce resource consumption. Service expansion is realized.

[0099] For the configuration management module, configuration hot update is achieved through a double-buffer mechanism, and service continuity is guaranteed through atomic switching; historical snapshots are recorded through a version manager, and quick rollback according to the version number is supported; sensitive fields (such as API Key) adopt in-memory encryption and regular rotation strategies. Dynamic configuration management is realized.

[0100] This method guarantees JDK8 compatibility and achieves low-version compatibility in the following ways:

[0101] Use Retrolambda to convert Lambda expressions into anonymous inner class bytecodes;

[0102] Rewrite Java Stream API calls through ASM to replace the native implementation of JDK11+;

[0103] Adopt OkHttp3 to replace Java HttpClient, supporting HTTP / 2 and connection pool reuse.

[0104] The following further details the specific implementation of this method in conjunction with the accompanying drawings.

[0105] As Figure 1 shown, it includes:

[0106] Unified service module: responsible for providing standardized service interfaces for the application layer, providing multi-mode access capabilities such as RESTful API, message queues, and event-driven, and supporting access protocols such as HTTP / WebSocket.

[0107] Dynamic routing module: Dynamically calculates weights based on real-time quality of service (QoS) data (response time, error rate, resource cost), and distributes requests to the optimal service node through the routing engine.

[0108] Service execution module: calls the target platform interface, handles underlying operations such as data serialization, authentication, and flow control, and returns a response in a unified format.

[0109] Configuration management module: Centrally manage service parameters (API endpoints, authentication keys, protocol types) through a database or file system, support dynamic hot loading, and ensure that configuration changes do not require restarting the service.

[0110] When a user initiates a request (such as "query Beijing weather"), the method performs the following steps:

[0111] Step 1: The application uses unified service capabilities to assemble a unified intermediate format UnifiedRequest, calls the system to send a large model request, and performs the following operations:

[0112] 1. Unified entry: receive HTTP / WebSocket and other protocol requests and complete protocol-independent parsing;

[0113] 2. Traffic scheduling: select the optimal node according to the load strategy (weight, response time) and forward the request to the dynamic routing module;

[0114] 3. Request records: Request records are persisted to ensure the subsequent provision of full-link monitoring, log tracking and alarm systems to ensure service stability.

[0115] Step 2: The unified requests enter the dynamic routing module and perform intelligent scheduling:

[0116] 1. Health assessment: Get real-time indicators (response time, error rate, cost weight) of candidate nodes from Redis;

[0117] 2. Weight calculation: weight calculation is performed according to the formula weight = 0.6 × (1 / response time) + 0.3 × availability + 0.1 × cost, and nodes are dynamically sorted;

[0118] 3. Optimal node allocation: Select the service node with the highest weight (such as DeepSeek). If the node is in a fuse state, it will automatically switch to the backup node (such as Zhipu AI).

[0119] Step 3: After the routing decision, the request is passed to the service execution module and the following operations are performed:

[0120] 1. Platform adapter invocation: Obtain the target platform adapter (such as DeepSeekService) through the dual-mode service factory (ServiceFactory), and inject configuration parameters (API key, endpoint).

[0121] 2. Result aggregation: Integrate the large model response with the tool execution result to generate the final answer (such as "It is sunny in Beijing today, 25°C").

[0122] Step 4: After the service execution layer generates the result, the system performs reverse protocol adaptation and the following operations:

[0123] 1. Data deserialization: Convert the platform-specific response (such as DeepSeek's JSON) to UnifiedResponse.

[0124] 2. Protocol reverse bridging: If the user request protocol is WebSocket, disassemble the HTTP response into long connection shard transmissions.

[0125] 3. Unified format return: The final data is encapsulated according to the user's initial protocol to ensure seamless parsing by the client.

[0126] In the above embodiments, the full-link processing logic of the user request from access to response is fully demonstrated, verifying the significant advantages of the system in terms of compatibility, flexibility, and intelligence. Through the layered architecture and modular design, the following effects are achieved: (1) Low version compatibility: Fully support modern protocols and toolchains in the JDK8 environment; (2) Protocol independence: Users can freely choose the access protocol, and the system automatically adapts to the target platform; (3) Elastic service governance: Intelligent routing and circuit breaking based on real-time data effectively improve resource utilization; (4) Zero-invasion expansion: Dynamic function calls and dual-mode factories support the online deployment of business functions within minutes.

[0127] The embodiment of this method also provides an implementation method of a unified service module. The unified service capability layer serves as the core access point of the system, responsible for providing standardized service interfaces for the application layer, shielding the differences of multiple underlying platforms, and integrating service governance capabilities, including the following units:

[0128] 1. Unified traffic entrance: Provide multi-mode access capabilities such as RESTful API, message queue, and event-driven, and support protocol access such as HTTP / WebSocket.

[0129] 2. Global traffic management: Implement strategies such as load balancing, gray release, and A / B testing, and allocate resources according to business priorities.

[0130] 3. Service Governance Center: Integrates monitoring, traffic limiting, and fault tolerance capabilities, and provides a unified logging, link tracing, and alerting system.

[0131] The unified service module provides a unified interface service externally, including HTTP, WebSocket, etc. The application passes in unified input parameter data according to business needs. At the same time, as the traffic entry, the module monitors and manages all requests, and completely records the link to facilitate subsequent problem tracing.

[0132] Combined with Figure 2 As shown, the following steps are executed in this embodiment:

[0133] Step 1: Unified traffic access and protocol parsing, support protocol access such as HTTP / 1.1, HTTP / 2, WebSocket, etc., provide a unified Endpoint (such as / api / v1 / chat), automatically parse the request header, and extract meta-information such as session ID and request type; convert the original request into an internal unified model UnifiedRequest, which includes business parameters, context information, and metadata.

[0134] Step 2: Global traffic scheduling and governance, use the token bucket algorithm to control the number of requests per second (QPS), isolate by service or user dimension, and perform global traffic limiting; based on elastic allocation of request traffic based on real-time load (CPU, memory), perform dynamic load balancing.

[0135] Step 3: Request recording and service governance, persistently record all requests to facilitate subsequent tracing of the request link.

[0136] In this implementation, the unified service layer effectively shields the protocol differences of multiple platforms through modular design and automated conversion logic, provides high-compatibility and high-elasticity service integration capabilities for enterprise applications in the JDK8 environment, and achieves the following effects:

[0137] (1) Unified service access: Receive requests through multiple protocol entrances (HTTP / WebSocket / gRPC), shielding the differences of the underlying platforms;

[0138] (2) Dynamic service governance: Integrates the registry, load balancing, circuit breaker and traffic limiting, realizing intelligent traffic scheduling and self-healing of faults;

[0139] (3) Global observability: Provides a full-link monitoring, logging tracing, and alerting system to ensure service stability.

[0140] The embodiments of this method also provide an implementation of a dynamic routing module. This module is mainly responsible for intelligently allocating requests to the optimal service nodes according to real-time Quality of Service (QoS) metrics, such as response time, error rate, and cost. At the same time, it also needs to handle the circuit breaker mechanism to ensure the rapid isolation and recovery of faulty nodes. The dynamic routing module consists of the following units, which are responsible for the intelligent selection of service nodes and traffic scheduling:

[0141] 1. Health monitoring unit: Real-time collection of QoS metrics such as response time, error rate, throughput, etc. of service nodes.

[0142] 2. Weight calculation unit: Dynamically calculate node weights based on multi-dimensional metrics to determine the optimal service nodes.

[0143] 3. Routing decision unit: Allocate request traffic according to the weight sorting results, supporting priority policies and load balancing.

[0144] 4. Circuit breaker management unit: Monitor the fault status of nodes, trigger circuit breakers and switch to standby nodes to ensure service continuity.

[0145] As Figure 3 shown, after protocol conversion, the dynamic routing module is responsible for selecting the corresponding routing execution platform for service calls. The specific steps are as follows:

[0146] Step 1, Health data collection. After each interface request call, the health monitoring unit stores node health data using Redis sorted sets (ZSet). The key is service:health:<node_id>, and the value is the timestamp-metric pair. The health data includes:

[0147] (1) Response time: Statistic of the average elapsed time of the last 100 requests (sliding window algorithm);

[0148] (2) Error rate: Calculate the proportion of failed requests within a unit time (configurable, default 5 minutes);

[0149] (3) Resource cost: Define the call cost coefficient of each service node based on the fees required by each platform (e.g., 1.0 for OpenAI and 0.8 for Moonshot).

[0150] Step 2, The weight calculation unit calculates the weight score through the weight formula. The weight formula is as follows:

[0151] Weight = α × (1 / Response time) + β × (1 - Error rate) + γ × (1 / Cost);

[0152] where α + β + γ = 1, default α = 0.6, β = 0.3, γ = 0.1.

[0153] By traversing all candidate nodes, calculating their weight scores, sorting them in descending order of scores, a list of available nodes is generated.

[0154] Step 3: Routing decision and traffic allocation. The default policy of the routing decision unit is to prioritize the highest authority and preferentially select the node with the highest score to process the current request. When the request frequency exceeds the threshold (5 times per second), the routing decision is switched, and it is switched to a random weight distribution, and the traffic is allocated according to the weight ratio (for example, node A has a weight of 60% and node B has a weight of 40%) to adjust the load balancing.

[0155] Step 4: Circuit breaker and recovery. After the error rate exceeds the threshold (50% for 1 minute) or the number of consecutive failures exceeds the limit (10 times), the node circuit breaker mechanism is triggered, and the circuit breaker management unit intervenes and starts the circuit breaker state machine of the current node. Within 5 minutes after the circuit breaker is triggered, the circuit breaker state of the node is Open, at this time all requests are rejected and fail quickly, and this node will not be selected again when selecting subsequent requests; after 5 minutes after the circuit breaker is triggered, the node circuit breaker state enters the Half-Open state. If the probe request is successful, the traffic will be gradually restored; if the probe requests continue to fail within 30 minutes, the node will be completely taken offline and wait for manual intervention to recover.

[0156] In this embodiment, the dynamic routing module realizes the intelligent scheduling and high availability guarantee of service traffic through real-time health monitoring, multi-dimensional weight calculation and circuit breaker recovery mechanism. Combined with the protocol adaptation layer and the dual-mode service factory, the system demonstrates excellent elasticity and stability in the JDK8 environment, providing a reliable technical foundation for enterprise-level AI service integration. Compared with the traditional method, the advantages of this embodiment are:

[0157] (1) Intelligent traffic allocation: Dynamically optimize resource utilization based on real-time QoS metrics, and the response time is reduced;

[0158] (2) Fast fault isolation: The circuit breaker mechanism ensures that the faulty node goes offline in seconds, and the system availability is increased to 99.99%;

[0159] (3) Elastic expansion: Support dynamically adding / removing nodes, and the routing policy automatically adapts to topological changes;

[0160] (4) Cost optimization: Control the call frequency of high-cost services through the cost coefficient to save operating costs.

[0161] This method embodiment also provides an implementation method of a service execution module. As the core execution unit of the system, this module is responsible for docking with specific large model platforms, completing request forwarding, protocol adaptation and result processing, and includes the following units:

[0162] 1. Platform Adapter Unit: Encapsulates the communication protocols and API specifications of different platforms (such as OpenAI's ChatCompletion and DeepSeek's Streaming API).

[0163] 2. Authentication Processor Unit: Dynamically injects the authentication information required by the platform (API Key, OAuth2 Token, signature).

[0164] 3. Traffic Controller Unit: Controls the request rate based on the token bucket algorithm to avoid triggering platform rate limits.

[0165] 4. Data Serializer Unit: Converts the unified request object into a platform-specific format (JSON / XML / Protobuf).

[0166] 5. Error Handler Unit: Catches platform exceptions and triggers the retry or circuit breaker mechanism.

[0167] As Figure 4 shown, after the routing selection is completed, the service execution module is responsible for executing the operation, and the specific steps are as follows:

[0168] Step 1, Adapter Selection and Initialization: The platform adapter unit loads the corresponding adapter through the dual-mode service factory (ServiceFactory) according to the configured platform field (such as passing in platform dictionary encodings like openai, deepseek, etc.). For example, when loading the OpenAIChatAdapter, the API endpoint (https: / / api.openai.com / v1) and authentication key are automatically injected.

[0169] Step 2, Request Serialization and Authentication Injection: The data serialization unit converts the unified request model UnifiedRequest into a platform-specific object (such as OpenAI's ChatCompletionRequest). The unit uses the Jackson library to handle JSON and the Protobuf compiler to handle binary streams; the authentication processor unit adds the authentication header according to the platform type (such as OpenAI's Authorization: Bearer sk-xxx). If the platform requires signature encryption, such as the Zhipu platform, HMAC-SHA256 is used to generate the signature.

[0170] Step 3, Request Execution and Traffic Control: Since each service platform has certain restrictions on QPS, to prevent request failures caused by large application concurrency, an independent token bucket is maintained for each platform. For example, OpenAI is limited to 60 RPM (60 requests per minute). When a request arrives, a token is applied. If it exceeds the limit, the request is queued to limit the frequency.

[0171] Step 4: Response parsing and error handling. Convert the original platform response (such as OpenAI's JSON) into a unified model UnifiedResponse. If there is nested text, such as choices[0].message.content, extract the generated text from the nesting. There is a risk of error in each call. If a network error (such as timeout, connection interruption) occurs, trigger up to 3 retries. If a platform error (such as OpenAI's 429 Too Many Requests) occurs, trigger circuit breaker and notify the routing layer.

[0172] In this embodiment, the service execution module realizes the efficient call and unified response of large models on multiple platforms through modular adapter design, dynamic protocol conversion, and intelligent traffic control. Combining the dynamic routing layer and the protocol adaptation layer, the system shows strong compatibility and stability in the JDK8 environment. Compared with the traditional method, the advantages of this embodiment are as follows:

[0173] (1) Protocol transparency: Decouple user requests from the platform protocol, and the adapter automatically completes protocol conversion;

[0174] (2) Elastic fault tolerance: Through the retry and circuit breaker mechanisms, the service success rate is increased to 99.9%;

[0175] (3) Precise traffic control: Based on the token bucket rate limiting strategy, avoid exceeding the limit of platform API calls;

[0176] (4) Unified error handling: Standardize the exception response format for easy unified parsing by the client.

[0177] This method embodiment also provides an implementation method of a configuration management module. The configuration management layer of this module is responsible for unified management of system configurations, supports flexible management of static and dynamic configurations, and includes the following units:

[0178] 1. Configuration loader unit: Load configuration data from files (YAML / JSON), databases, or configuration centers (such as Nacos).

[0179] 2. Configuration parser unit: Convert the original configuration into Java objects (such as ServiceConfig, AuthConfig).

[0180] 3. Configuration memory unit: Store the currently effective configuration using a double-buffer mechanism to avoid read-write conflicts.

[0181] 4. Hot update listener unit: Listen for configuration source change events and trigger dynamic updates.

[0182] 5. Version manager unit: Record the historical versions of configurations, support quick rollback and auditing.

[0183] As Figure 5 shown, the specific implementation steps are as follows:

[0184] Step 1, Configuration Loading and Initialization. The configuration loader can perform multi-source loading. For static file types, the basic configuration (such as service endpoints, protocol types) is loaded from application.yml, and dynamic configurations (such as API keys, routing policies) are dynamically read through the database + Redis.

[0185] Step 2, Configuration Parsing and Validation. Convert the string values in YAML to Java objects (such as ProtocolType enumeration), and at the same time check the required fields (such as apiKey, endpoint), and verify the compatibility of the protocol type and data format (such as gRPC must use Protobuf).

[0186] Step 3, Hot Update and Listening Mechanism. Use WatchService to listen for configuration file modification events, regularly query the update timestamp of the database configuration table to detect changes, and trigger the refresh logic of modules such as the service factory and routing layer after the configuration is updated. Generate a snapshot each time the configuration is updated and store it in the database or file system, and it can be quickly restored to the specified historical configuration through the version number.

[0187] In this embodiment, the configuration management module realizes centralized and dynamic management of configurations through multi-source loading, hot update listening, and version management. Combined with JDK8 compatibility design (such as using WatchService to replace high-version APIs), the system can still efficiently respond to configuration changes in a low-version environment, providing reliable configuration support for modules such as the protocol adaptation layer and dynamic routing layer. Compared with the traditional method, its effects are as follows:

[0188] (1) Multi-source Integration: Support unified management of multiple configuration sources such as files, databases, and environment variables;

[0189] (2) Zero-downtime Update: The dual-buffer mechanism and hot loading achieve seamless configuration switching without service interruption;

[0190] (3) Secure and Reliable: Configuration verification and version rollback ensure system stability and reduce the risk of human errors;

[0191] (4) Efficient Access: In-memory storage and caching strategy, with a configuration reading latency of <1ms.

[0192] The present invention adopts bytecode downgrading technology to support the operation of features such as Lambda expressions and Stream API in the JDK8 environment, uses OkHttp3 to replace the native HTTP Client in Java 11+ to ensure compatibility with low versions; designs a protocol adaptation middleware to support automatic sniffing and conversion of HTTP / 1.1, HTTP / 2, gRPC, and WebSocket protocols, realizes seamless conversion of JSON / XML / Protobuf data formats and unified mapping of authentication schemes, reduces platform docking code, supports zero-invasive tool extension, and shortens the online time of new functions to the minute level. It supports static configuration and dynamic registration, combines hot update and double-buffering technology to achieve seamless service switching, reduces the development cost of new platform adapters; based on the service quality (QoS) dynamic routing algorithm, it distributes traffic in real time considering response time, error rate, and cost weight; by statistically analyzing the node health, it automatically fuses when the error rate exceeds the threshold, improves resource utilization, and is applicable to the integration of large model services in fields such as smart cities, finance, customer service, and healthcare, significantly reducing the complexity of multi-platform management and enhancing the system's elasticity and maintainability. It solves the problems of compatibility, protocol differences, dynamic expansion, and intelligent scheduling in the integration of large model services across multiple platforms in a low-version Java environment.

[0193] An embodiment of the present invention also provides a heterogeneous service large model dynamic adaptation and intelligent scheduling system based on JDK8, including:

[0194] A unified service module for providing standardized HTTP / WebSocket access to the application layer and shielding the differences of underlying multi-platforms;

[0195] A dynamic routing module for dynamically selecting service nodes based on service quality (QoS) metrics, supporting weight calculation, fuse recovery, and traffic coloring;

[0196] A service execution module for loading platform adapters through a dual-mode service factory, integrating tool call functions, and implementing protocol conversion and authentication injection;

[0197] A configuration management module that supports hot loading of dual configuration sources of static files and dynamic databases, and provides versioned rollback and consistency verification;

[0198] This system realizes the dynamic adaptation and intelligent scheduling of heterogeneous service large models based on JDK8 through the method of dynamic adaptation and intelligent scheduling of heterogeneous service large models based on JDK8 described in the above embodiments.

[0199] The unified service module includes:

[0200] 1. Unified traffic entrance: It provides multi-mode access capabilities such as RESTful API, message queue, and event-driven, and supports protocol access such as HTTP / WebSocket.

[0201] 2. Global traffic management: Implement strategies such as load balancing, gray release, and A / B testing, and allocate resources according to business priorities.

[0202] 3. Service governance center: Integrate monitoring, rate limiting, and fault tolerance capabilities, and provide a unified logging, link tracing, and alerting system.

[0203] The unified service module provides a unified interface service externally, including HTTP, WebSocket, etc. Applications pass in unified input parameter data according to business needs. At the same time, as the traffic entry, the module monitors and governs all requests, and completely records the link for subsequent problem tracing.

[0204] The unified service module performs the following steps:

[0205] Step 1, unified traffic access and protocol parsing, support the access of protocols such as HTTP / 1.1, HTTP / 2, and WebSocket, provide a unified Endpoint (such as / api / v1 / chat), automatically parse the request header, and extract meta-information such as session ID and request type; convert the original request into an internal unified model UnifiedRequest, including business parameters, context information, and metadata.

[0206] Step 2, global traffic scheduling and governance, use the token bucket algorithm to control the number of requests per second (QPS), isolate by service or user dimension, and perform global rate limiting; allocate request traffic elastically based on real-time load (CPU, memory) for dynamic load balancing.

[0207] Step 3, request recording and service governance, persistently record all requests for subsequent request link tracing.

[0208] The unified service layer effectively shields the protocol differences of multiple platforms through modular design and automated conversion logic, and provides high-compatibility and high-elasticity service integration capabilities for enterprise-level applications in the JDK8 environment.

[0209] The dynamic routing module realizes the intelligent selection and traffic scheduling of service nodes, including the following units:

[0210] 1. Health monitoring unit: Real-time collect QoS indicators such as response time, error rate, and throughput of service nodes.

[0211] 2. Weight calculation unit: Dynamically calculate the node weights based on multi-dimensional indicators to determine the optimal service node.

[0212] 3. Routing decision unit: Allocate request traffic according to the weight sorting result, and support priority policies and load balancing.

[0213] 4. Fuse Management Unit: Monitor the node failure status, trigger fusing and switch to the standby node to ensure service continuity.

[0214] After the protocol conversion is completed, the dynamic routing module is responsible for selecting the corresponding routing to execute the platform service call. The specific steps are as follows:

[0215] Step 1, Health data collection. After each interface request call, the health monitoring unit stores the node health data using Redis sorted set (ZSet). The key is service:health:<node_id>, and the value is the timestamp-metric pair. The health data includes:

[0216] (1) Response time: Statistic the average time consumption of the last 100 requests (sliding window algorithm);

[0217] (2) Error rate: Calculate the proportion of failed requests within a unit time (configurable, default 5 minutes);

[0218] (3) Resource cost: Based on the costs required by each platform, predefined the call cost coefficient of each service node (e.g., 1.0 for OpenAI, 0.8 for Moonshot).

[0219] Step 2, The weight calculation unit calculates the weight score through the weight formula. The weight formula is as follows:

[0220] Weight = α × (1 / Response time) + β × (1 - Error rate) + γ × (1 / Cost);

[0221] Among them, α + β + γ = 1, by default α = 0.6, β = 0.3, γ = 0.1.

[0222] By traversing all candidate nodes, calculate their weight scores, sort them in descending order of scores, and generate a list of available nodes.

[0223] Step 3, Routing decision and traffic allocation. The default policy of the routing decision unit is highest privilege first, and preferentially select the node with the highest score to process the current request. When the request frequency exceeds the threshold (5 times per second), then the routing decision is switched, switched to random weight distribution, and allocate traffic according to the weight ratio (e.g., node A weight 60%, node B weight 40%) to adjust the load balancing.

[0224] Step 4: Fusing and Recovery. After the error rate exceeds the threshold (50% for 1 minute continuously) or the number of consecutive failures exceeds the limit (10 times), the node fusing mechanism is triggered. The fusing management unit intervenes and starts the fusing state machine of the current node. Within 5 minutes after the fusing is triggered, the fusing state of the node is Open. At this time, all requests are rejected and quickly failed. When selecting subsequent requests, this node will no longer be selected. After 5 minutes of the fusing being triggered, the node fusing state enters the Half-Open state. If the probing request is successful, the traffic is gradually restored. If the probing requests continue to fail within 30 minutes, the node is completely taken offline and waits for manual intervention to recover.

[0225] Through real-time health monitoring, multi-dimensional weight calculation, and fusing and recovery mechanisms, the dynamic routing module realizes the intelligent scheduling of service traffic and high-availability guarantee. Combined with the protocol adaptation layer and the dual-mode service factory, the system demonstrates excellent elasticity and stability in the JDK8 environment, providing a reliable technical foundation for enterprise-level AI service integration.

[0226] The service execution module, as the core execution unit of the system, is responsible for docking specific large model platforms, completing request forwarding, protocol adaptation, and result processing, and includes the following units:

[0227] 1. Platform Adapter Unit: Encapsulates the communication protocols and API specifications of different platforms (such as ChatCompletion of OpenAI, Streaming API of DeepSeek).

[0228] 2. Authentication Processor Unit: Dynamically injects the authentication information required by the platform (API Key, OAuth2 Token, signature).

[0229] 3. Traffic Controller Unit: Controls the request rate based on the token bucket algorithm to avoid triggering platform traffic limiting.

[0230] 4. Data Serializer Unit: Converts the unified request object into a platform-specific format (JSON / XML / Protobuf).

[0231] 5. Error Handler Unit: Captures platform exceptions and triggers the retry or fusing mechanism.

[0232] After the routing selection is completed, the service execution module is responsible for performing operations. The specific steps are as follows:

[0233] Step 1, Adapter Selection and Initialization: The platform adapter unit loads the corresponding adapter through the Dual-Mode ServiceFactory according to the configured platform field (such as the incoming platform dictionary encodings like openai, deepseek, etc.). For example, when loading the OpenAIChatAdapter, the API endpoint (https: / / api.openai.com / v1) and the authentication key are automatically injected.

[0234] Step 2, Request Serialization and Authentication Injection: The data serialization unit converts the UnifiedRequest into a platform-specific object (such as OpenAI's ChatCompletionRequest). The unit uses the Jackson library to handle JSON and the Protobuf compiler to handle binary streams. The authentication processor unit adds the authentication header according to the platform type (such as OpenAI's Authorization: Bearer sk-xxx). If the platform requires signature encryption, such as the Zhipu platform, HMAC-SHA256 is used to generate the signature.

[0235] Step 3, Request Execution and Traffic Control: Since each service platform has certain limitations on QPS, to prevent request failures caused by a large number of concurrent applications, an independent token bucket is maintained for each platform. For example, OpenAI is limited to 60 RPM (60 requests per minute). When a request arrives, a token is applied. If it exceeds the limit, the request is queued to limit the frequency.

[0236] Step 4, Response Parsing and Error Handling: The original platform response (such as OpenAI's JSON) is converted into the UnifiedResponse. If there is nested text, such as choices[0].message.content, the generated text is extracted from the nested part. There is a risk of error in each call. If a network error (such as timeout, connection interruption) occurs, up to 3 retries are triggered. If a platform error (such as OpenAI's 429 Too Many Requests) occurs, circuit breaking is triggered and the routing layer is notified.

[0237] The service execution module realizes the efficient call and unified response of large models on multiple platforms through modular adapter design, dynamic protocol conversion, and intelligent traffic control. Combined with the dynamic routing layer and the protocol adaptation layer, the system demonstrates strong compatibility and stability in the JDK8 environment.

[0238] The configuration management module, the configuration management layer of this module is responsible for uniformly managing the system configuration, supporting flexible management of static and dynamic configurations, and includes the following units:

[0239] 1. Configuration loader unit: loads configuration data from files (YAML / JSON), databases, or configuration centers (such as Nacos).

[0240] 2. Configuration parser unit: converts raw configuration into Java objects (such as ServiceConfig, AuthConfig).

[0241] 3. Configuration memory unit: Use double buffering mechanism to store the current effective configuration to avoid read and write conflicts.

[0242] 4. Hot update listener unit: listens to configuration source change events and triggers dynamic updates.

[0243] 5. Version manager unit: records configuration history versions and supports fast rollback and auditing.

[0244] The specific steps are as follows:

[0245] Step 1: Configuration loading and initialization. The configuration loader can perform multi-source loading. Static file types load basic configurations (such as service endpoints and protocol types) from application.yml, and dynamic configurations (such as API keys and routing strategies) are dynamically read through the database + Redis.

[0246] Step 2: Configure parsing and validation, convert the string value in YAML into a Java object (such as the ProtocolType enumeration), check the required fields (such as apiKey, endpoint), and verify the compatibility of the protocol type and data format (such as gRPC must use Protobuf).

[0247] Step 3: Hot update and monitoring mechanism. Use WatchService to monitor configuration file modification events, regularly query the update timestamp of the database configuration table, detect changes, and trigger the refresh logic of modules such as the service factory and routing layer after the configuration is updated. A snapshot is generated each time the configuration is updated, stored in the database or file system, and can be quickly restored to the specified historical configuration by version number.

[0248] The configuration management module realizes centralized and dynamic management of configuration through multi-source loading, hot update monitoring and version management. Combined with JDK8 compatibility design (such as WatchService replacing high-version API), the system can still efficiently respond to configuration changes in low-version environments, providing reliable configuration support for modules such as the protocol adaptation layer and dynamic routing layer.

[0249] When a user initiates a request (such as "Query Beijing Weather"), the system performs the following steps:

[0250] Step 1: Apply the unified service capabilities to assemble the unified intermediate format UnifiedRequest, call the system to send the large model request, and perform the following operations:

[0251] 1. Unified entry: Receive requests of protocols such as HTTP / WebSocket, and complete protocol-independent parsing;

[0252] 2. Traffic scheduling: Select the optimal node according to the load strategy (weight, response time), and forward the request to the dynamic routing module;

[0253] 3. Request recording: Persist the request records to ensure subsequent provision of full-link monitoring, log tracking, and alerting systems to guarantee service stability.

[0254] Step 2: The unified request enters the dynamic routing module and performs intelligent scheduling:

[0255] 1. Health assessment: Obtain the real-time metrics (response time, error rate, cost weight) of candidate nodes from Redis;

[0256] 2. Weight calculation: Calculate the weight according to the formula weight = 0.6×(1 / response time)+0.3×availability+0.1×cost, and sort the nodes dynamically;

[0257] 3. Optimal node allocation: Select the service node with the highest weight (such as DeepSeek). If this node is in the fuse state, automatically switch to the standby node (such as Zhipu AI).

[0258] Step 3: After the routing decision, the request is passed to the service execution module and the following operations are performed:

[0259] 1. Platform adapter call: Obtain the target platform adapter (such as DeepSeekService) through the dual-mode service factory (ServiceFactory), and inject configuration parameters (API key, endpoint);

[0260] 2. Result aggregation: Integrate the large model response with the tool execution result to generate the final answer (such as "It is sunny in Beijing today, 25°C").

[0261] Step 4: After the service execution layer generates the result, the system performs reverse protocol adaptation and performs the following operations:

[0262] 1. Data deserialization: Convert the platform-specific response (such as the JSON of DeepSeek) to UnifiedResponse;

[0263] 2. Protocol reverse bridging: If the user request protocol is WebSocket, disassemble the HTTP response into long connection shard transmissions;

[0264] 3. Unified format for return: The final data is encapsulated according to the user's initial protocol to ensure seamless parsing by the client.

[0265] An embodiment of the present invention also provides a heterogeneous service large model dynamic adaptation and intelligent scheduling device based on JDK8, including: at least one memory and at least one processor;

[0266] The at least one memory is used to store machine-readable programs;

[0267] The at least one processor is used to call the machine-readable program to implement the heterogeneous service large model dynamic adaptation and intelligent scheduling method described in the above embodiments based on JDK8.

[0268] An embodiment of the present invention also provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the heterogeneous service large model dynamic adaptation and intelligent scheduling method described in the above embodiments is implemented. Specifically, a system or device equipped with a storage medium can be provided. On this storage medium, software program code for implementing the functions of any one of the above embodiments is stored, and the computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.

[0269] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.

[0270] Embodiments of the storage medium for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROM. Optionally, the program code can be downloaded from a server computer via a communication network.

[0271] In addition, it should be clear that not only can the actual operations be completed in part or in whole by executing the program code read by the computer, but also by the operating system and the like operating on the computer based on the instructions of the program code, so as to implement the functions of any one of the above embodiments.

[0272] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion unit is made to execute part or all of the actual operations, thereby implementing the functions of any one of the above embodiments.

[0273] The present invention has been shown and described in detail above with reference to the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above-mentioned multiple embodiments, those skilled in the art can know that the code review means in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the protection scope of the present invention.

Claims

1. A method for dynamic adaptation and intelligent scheduling of heterogeneous service large models based on JDK8, characterized in that, The implementation of this method includes: Unified service module, providing standardized HTTP / WebSocket access for the application layer, shielding the differences among the underlying platforms; Dynamic routing module, dynamically selects service nodes based on service quality indicators, supports weight calculation, circuit breaker recovery and traffic coloring; The service execution module loads the platform adapter through the dual-mode service factory, integrates the tool call function and implements protocol conversion and authentication injection; The configuration management module supports hot loading of dual configuration sources: static files and dynamic databases, and provides version rollback and consistency verification.

2. The method for dynamic adaptation and intelligent scheduling of heterogeneous service large models based on JDK8 according to claim 1, wherein, The dynamic routing module, The weight calculation formula is as follows: Weight = α × (1 / response time) + β × (1-error rate) + γ × (1 / cost), where α + β + γ = 1; The fuse recovery,the fuse mechanism is triggered when the node error rate exceeds 50%,,automatically switching to the backup service; The traffic coloring supports grayscale release strategy and allocates traffic according to request tags.

3. The method for dynamic adaptation and intelligent scheduling of heterogeneous service large models based on JDK8 according to claim 1, wherein, The dual-mode service factory, Static mode predefines service parameters through YAML / database and supports environment variable injection; Dynamic mode loads third-party adapters based on the Java SPI mechanism, and registration takes effect immediately at runtime; The service instance pool adopts a lazy initialization strategy to create connections on demand to reduce resource consumption.

4. The method for dynamic adaptation and intelligent scheduling of heterogeneous service large models based on JDK8 according to claim 1 or 3, characterized in that The configuration management module, A double buffer mechanism is used to implement hot configuration updates, and atomic switching is used to ensure service continuity. The version manager records historical snapshots and supports quick rollback by version number; Sensitive fields use memory encryption and regular rotation strategies.

5. The method for dynamic adaptation and intelligent scheduling of heterogeneous service large models based on JDK8 according to claim 1, wherein This method ensures JDK8 compatibility in the following ways: Use Retrolambda to convert Lambda expressions to anonymous inner class bytecode; Rewrite Java Stream API calls through ASM to replace JDK11+ native implementations; OkHttp3 is used to replace Java HttpClient, supporting HTTP / 2 and connection pool reuse.

6. The method for dynamically adapting and intelligently scheduling heterogeneous service large models based on JDK8 according to claim 1, wherein When a user initiates a request, the method performs the following steps to perform intelligent scheduling: Step 1: The application uses unified service capabilities to assemble a unified intermediate format UnifiedRequest, calls the system to send a large model request, and performs the following operations: Unified entry: receive HTTP / WebSocket protocol requests and complete protocol-independent parsing; Traffic scheduling: select the optimal node according to the load strategy and forward the request to the dynamic routing module; Request records: Request records are processed persistently to ensure the subsequent provision of full-link monitoring, log tracking and alarm systems to ensure service stability; Step 2: The unified requests enter the dynamic routing module and perform intelligent scheduling: Health assessment: Get real-time indicators of candidate nodes from Redis; Weight calculation: perform weight calculation and dynamically sort nodes; Optimal node allocation: select the service node with the highest weight. If the node is in a fuse state, it will automatically switch to the backup node. Step 3: After the routing decision is made, the request is passed to the service execution module, which performs the following operations: Platform adapter call: Get the target platform adapter through the dual-mode service factory and inject configuration parameters; Result aggregation: Integrate the large model response with the tool execution results to generate the final answer; In step 4, after the service execution layer generates the result, the system reversely executes protocol adaptation and performs the following operations: Data deserialization: Convert the platform-specific response into a UnifiedResponse; Protocol reverse bridging: If the user request protocol is WebSocket, disassemble the HTTP response into long connection shard transmissions; Return in unified format: The final data is encapsulated according to the user's initial protocol to ensure seamless parsing by the client.

7. Heterogeneous service large model dynamic adaptation and intelligent scheduling system based on JDK8, characterized in that Including: Unified service module, used to provide standardized HTTP / WebSocket access for the application layer, shielding the underlying multi-platform differences; Dynamic routing module, used to dynamically select service nodes based on service quality indicators, supporting weight calculation, circuit breaker recovery, and traffic coloring; Service execution module, used to load the platform adapter through a dual-mode service factory, integrate the tool call function, and implement protocol conversion and authentication injection; Configuration management module, supporting hot loading of dual configuration sources of static files and dynamic databases, providing versioned rollback and consistency verification; This system realizes the dynamic adaptation and intelligent scheduling of heterogeneous service large models based on JDK8 through the method described in any one of claims 1-6.

8. The heterogeneous service large model dynamic adaptation and intelligent scheduling system based on JDK8 according to claim 7, characterized in that The unified service module includes the following units: Unified traffic entrance: Provide multi-mode access capabilities, including RESTful API, message queue, event-driven, etc., supporting HTTP / WebSocket protocol access; Global traffic management: Implement load balancing, gray release, A / B test strategies, and allocate resources according to business priorities; Service governance center: Integrate monitoring, rate limiting, and fault tolerance capabilities, providing a unified logging, link tracing, and alerting system; The dynamic routing module realizes the intelligent selection and traffic scheduling of service nodes, including the following units: Health monitoring unit: Real-time collect the QoS indicators of service nodes, including response time, error rate, and throughput; Weight calculation unit: Dynamically calculate the node weights based on multi-dimensional indicators to determine the optimal service node; Routing decision unit: Allocate request traffic according to the weight sorting result, supporting priority strategies and load balancing; Circuit breaker management unit: Monitor the node failure status, trigger the circuit breaker and switch to the standby node to ensure service continuity; The service execution module interfaces with specific large model platforms to complete request forwarding, protocol adaptation, and result processing, including the following units: Platform adapter unit: Encapsulate the communication protocols and API specifications of different platforms; Authentication processor unit: Dynamically inject the authentication information required by the platform; Traffic controller unit: Control the request rate based on the token bucket algorithm to avoid triggering platform rate limiting; Data serializer unit: Convert the unified request object into a platform-specific format; Error processor unit: Capture platform exceptions and trigger the retry or circuit breaker mechanism; The configuration management module, the configuration management layer realizes unified management of system configurations, including the following units: Configuration Loader Unit: Load configuration data from files, databases, or configuration centers; Configuration Parser Unit: Convert the original configuration into Java objects; Configuration Memory Unit: Store the currently effective configuration using a double-buffering mechanism to avoid read-write conflicts; Hot Update Listener Unit: Listen for configuration source change events and trigger dynamic updates; Version Manager Unit: Record the historical versions of the configuration and support quick rollback and auditing.

9. Heterogeneous service large model dynamic adaptation and intelligent scheduling device based on JDK8, characterized in that Including: At least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program to implement the method according to any one of claims 1 to 6.

10. A computer-readable medium, characterized in that, Computer instructions are stored on the computer-readable medium, and when the computer instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Configuration management method, device, system, equipment, storage medium and program product

    CN120547040A

  • Server configuration method and system, server, equipment and medium

    CN120821516A

  • Unified communication system

    CN120935246A

  • Plug-in large model access system for embedded equipment

    CN120980144A

  • Large language model dynamic adaptation method and system based on Java

    CN121012756A