Scheduling method and device of large model in cross-domain computing power center
By unifying external interface modules and adapting parameter data processing, the problems of insufficient computing resources and stability in large model services are solved, flexible scheduling and resource pool expansion across cross-domain computing centers are achieved, and the applicability and stability of large model services are improved.
Patent Information
- Application Number
- CN202510850086.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-17
AI Technical Summary
In existing technologies, the deployment of large model services faces problems such as insufficient supply of computing resources, high cost of high-end computing equipment, and difficulty in ensuring service stability. In addition, the computing network lacks the ability to dynamically allocate heterogeneous computing resources, making hot-swappable resource scheduling across heterogeneous computing centers difficult to achieve, limiting the flexible deployment capabilities of large model services.
Through a unified external interface module, user input is converted into request data based on preset standard rules, the real-time load score of the connected computing center is calculated, the target computing center is dynamically selected, and adaptation parameter data is obtained based on the adaptation tag. The initial dialogue request is converted into the target dialogue request of the target computing center, achieving seamless cross-domain migration and supporting hybrid scheduling of cloud services, bare metal clusters and third-party APIs.
It enhances the deployment flexibility of large models in cross-domain computing centers, breaks through the limitations of a single type of resource pool, realizes an expandable and hot-swappable computing resource pool, significantly expands the available computing power range of large models, and ensures the stability and flexibility of the service.
Smart Images

Figure CN120806130A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of large model deployment, in particular to a large model scheduling method and device in cross-domain computing power centers, equipment and storage medium. BACKGROUND
[0002] With the rapid development of large model technology and the prosperity of open source ecology, more and more users can experience large model services. However, large model service deployment still faces three major challenges: insufficient supply of computing power resources, high cost of high-end computing devices, and difficulty in ensuring service stability.
[0003] The related art adopts a computing power network architecture to build a virtual computing power resource pool by interconnecting multiple computing power centers. The resource pool integrates diversified computing chip resources to provide basic computing power support for large model services, which to some extent alleviates the computing power constraint problem and improves service stability. However, the current computing power network for deploying large models can only provide single-type virtual resources in each computing power center, lacks dynamic allocation capability of heterogeneous computing resources, and cannot realize hot plug resource scheduling of cross-domain heterogeneous computing power centers, which limits the flexible deployment capability of large model services in different application scenarios. SUMMARY
[0004] The main purpose of the embodiments of the present application is to propose a large model scheduling method, device, equipment and storage medium in cross-domain computing power centers, to improve the deployment flexibility of large models in cross-domain computing power centers and to improve the application scope of large models.
[0005] To achieve the above purpose, the first aspect of the embodiments of the present application proposes a large model scheduling method in cross-domain computing power centers, comprising:
[0006] Obtain an initial dialogue request from an asynchronous task queue, wherein the initial dialogue request is request data of a preset standard rule obtained by a unified external interface module in response to user input;
[0007] Calculate the real-time load score of the accessed multiple computing power centers, and select at least one target computing power center from the multiple computing power centers according to the real-time load score, wherein at least two computing power centers are located in a cross-domain resource environment, at least one large model is deployed on each computing power center, and the cross-domain resource environment at least includes at least two of cloud services, bare metal clusters or third-party API gateways;
[0008] Obtain corresponding adaptation parameter data according to the adaptation label of the target computing power center, and convert the initial dialogue request into a target dialogue request corresponding to the target computing power center according to the adaptation parameter data;
[0009] Determine a target large model from the large model deployed by the target computing power center, send the target dialogue request to the target large model for content generation, and obtain an inference result.
[0010] In some embodiments, the obtaining of the corresponding adaptation parameter data according to the adaptation label of the target computing power center comprises:
[0011] If the adaptation label indicates that the adaptation state of the target computing power center is adapted, the adaptation parameter data of the target computing power center is obtained.
[0012] If the adaptation label indicates that the adaptation state of the target computing power center is not adapted, an adaptation operation is performed on the target computing power center to obtain the corresponding adaptation parameter data.
[0013] In some embodiments, the adaptation operation performed on the target computing power center to obtain the corresponding adaptation parameter data comprises:
[0014] Obtaining adaptation reference data of the target computing power center, the adaptation reference data at least comprising: parameter structure data, authentication signature mechanism, platform behavior encapsulation mechanism and response analysis mechanism;
[0015] According to the adaptation reference data, the target computing power center is adapted respectively to obtain corresponding structured parameters, access verification data, request setting rules and standard response output format.
[0016] In some embodiments, the adaptation operation performed on the target computing power center according to the adaptation reference data to obtain the corresponding structured parameters, access verification data, request setting rules and standard response output format comprises:
[0017] Traverse each standardized parameter in the preset standard rule, obtain at least field mapping rule and nested structure from the parameter structure data, convert the standardized parameter into target field parameter according to the field mapping rule, and convert the target field parameter into the structured parameter according to the nested structure;
[0018] Obtain the standard credential field corresponding to the preset standard rule, obtain the authentication rule from the authentication signature mechanism, and generate the access verification data of the target computing power center based on the authentication rule and the standard credential field;
[0019] obtaining the request setting rule according to the platform behavior rule obtained from the platform behavior packaging mechanism, the request setting rule comprising a pre-request rule and / or a post-request rule, the pre-request rule being used to set at least one pre-request processing operation of session initialization, vocabulary uploading, and context preloading before obtaining the initial dialogue request, and the post-request rule being used to set at least one post-request processing operation of polling processing and stream decoding after obtaining the initial dialogue request;
[0020] obtaining a response type from the response analysis mechanism, obtaining corresponding analysis logic according to the response type, and obtaining the standard response output format based on the analysis logic.
[0021] In some embodiments, the obtaining the standard response output format based on the analysis logic comprises:
[0022] If the analysis logic is stream response analysis, obtaining a stream protocol type from the response analysis mechanism, and determining field extraction and semantic reorganization analysis rules according to the stream protocol type.
[0023] In some embodiments, after the sending the target dialogue request to the target large model for content generation to obtain an inference result, the method comprises:
[0024] obtaining thinking process data and final answer data from the inference result;
[0025] displaying the final answer data through the unified external interface module, and determining to perform lazy loading or content folding on the thinking process data according to a preset display rule.
[0026] In some embodiments, the calculating real-time load scores of the accessed multiple computing power centers and selecting at least one target computing power center from the multiple computing power centers according to the real-time load scores comprises:
[0027] calculating the real-time load score according to the obtained task state information, resource load index, and platform capability label of each computing power center;
[0028] taking the computing power center with a real-time load score greater than or equal to a preset score as a candidate computing power center, calculating a candidate score value according to the adaptation level, interface stability parameter, and stream inference capability index of each candidate computing power center, and selecting the candidate computing power center corresponding to the maximum candidate score value as the target computing power center;
[0029] If each real-time load score is less than the preset score, reducing the preset score by a preset step size until the target computing power center is selected.
[0030] To achieve the above object, the second aspect of the embodiment of the present application proposes a large model scheduling device in a cross-domain computing power center, comprising:
[0031] The dialogue request acquisition module is configured to acquire an initial dialogue request from the asynchronous task queue, wherein the initial dialogue request is request data of a preset standard rule obtained by the unified external interface module in response to user input;
[0032] The computing power scheduling module is configured to calculate real-time load scores of a plurality of accessed computing power centers, and select at least one target computing power center from the plurality of computing power centers according to the real-time load scores, wherein at least two of the computing power centers are located in a cross-domain resource environment, at least one large model is deployed on each of the computing power centers, and the cross-domain resource environment at least includes at least two of cloud services, bare-metal clusters, or third-party API gateways;
[0033] The adaptation module is configured to acquire corresponding adaptation parameter data according to the adaptation label of the target computing power center, and convert the initial dialogue request into a target dialogue request corresponding to the target computing power center according to the adaptation parameter data;
[0034] The content generation module is configured to determine a target large model from the large models deployed by the target computing power center, send the target dialogue request to the target large model for content generation, and obtain an inference result.
[0035] To achieve the above object, the third aspect of the embodiment of the present application proposes an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the method of the first aspect when executing the computer program.
[0036] To achieve the above object, the fourth aspect of the embodiment of the present application proposes a storage medium, which is a storage medium, and the storage medium stores a computer program, and the computer program is executed by a processor to implement the method of the first aspect.
[0037] The large model scheduling method across domain computing power centers, device and storage medium are provided in the embodiments of the present application. An initial dialogue request is obtained from an asynchronous task queue, wherein the initial dialogue request is request data of a preset standard rule obtained by a unified external interface module in response to user input; real-time load scores of a plurality of accessed computing power centers are calculated, and at least one target computing power center is selected from the plurality of computing power centers according to the real-time load scores, wherein the at least two computing power centers are located in a cross-domain resource environment, at least one large model is deployed on each computing power center, and the cross-domain resource environment at least includes at least two of cloud services, bare-metal clusters or third-party API gateways; corresponding adaptation parameter data is obtained according to an adaptation label of the target computing power center, the initial dialogue request is converted into a target dialogue request corresponding to the target computing power center according to the adaptation parameter data; a target large model is determined from the large model deployed in the target computing power center, and the target dialogue request is sent to the target large model for content generation to obtain an inference result. In the embodiments of the present application, the user input is converted into the request data of the preset standard rule through the unified external interface module, so that different computing power centers can receive and process the request without difference, and the adaptation problem caused by interface difference is avoided. Next, the request parameters are dynamically adjusted according to the adaptation label of the target computing power center, the seamless migration of the request across domains is realized, the deployment flexibility is enhanced, the mixed scheduling of cloud services, bare-metal clusters and third-party APIs is supported, the limitation of a single type of resource pool is broken through, the scalable and hot-pluggable computing power resource pool is realized, and the available computing power range of the large model is significantly expanded. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 FIG. 1 is a structural schematic diagram of a scheduling system of a cross-domain computing power center corresponding to a large model scheduling method in embodiments of the present application.
[0039] Figure 2 FIG. 2 is a flowchart of a large model scheduling method across domain computing power centers provided by the embodiments of the present application.
[0040] Figure 3 FIG. 3 is a flowchart of calculating real-time load scores of a plurality of accessed computing power centers and selecting at least one target computing power center from the plurality of computing power centers according to the real-time load scores provided by the embodiments of the present application.
[0041] Figure 4 FIG. 4 is a flowchart of obtaining corresponding adaptation parameter data according to an adaptation label of a target computing power center provided by the embodiments of the present application.
[0042] Figure 5 FIG. 5 is a flowchart of performing an adaptation operation on a target computing power center to obtain corresponding adaptation parameter data provided by the embodiments of the present application.
[0043] Figure 6is a flowchart provided by the embodiment of the application, which respectively performs adaptation operation on the target computing power center according to the adaptive reference data, obtains the corresponding structured parameters, access verification data, request setting rules and standard response output format.
[0044] Figure 7 is another schematic diagram of the large model in the scheduling and distribution of the cross-domain computing power center provided by the embodiment of the application.
[0045] Figure 8 is a structural block diagram of the scheduling device of the large model in the cross-domain computing power center provided by another embodiment of the application.
[0046] Figure 9 is a hardware structure schematic diagram of the electronic device provided by the embodiment of the application. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.
[0048] It should be noted that although the functional modules are divided in the device schematic diagram, and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart.
[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0050] First, the several terms involved in the present application are analyzed:
[0051] Artificial Intelligence (AI): is a new technical science to study, develop, simulate, extend and expand human intelligence, and is a branch of computer science. Artificial intelligence aims to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. The research in this field includes robots, language recognition, image recognition, natural language processing and expert systems. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence is also the theory, method, technology and application system of using digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, to perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0052] With the rapid development of large model technology and the flourishing of open source ecology, more and more users can experience large model services. However, there are still three major challenges in deploying large model services: insufficient supply of computing power resources, high cost of high-end computing devices, and difficulty in ensuring service stability. For example, with the emergence of DeepSeek R1 open source large model, high-end computing power resources such as H100 and Ascend910B are in short supply, and a large number of users' access in a short time causes the DeepSeek official service to be temporarily unavailable.
[0053] In terms of using computing power resources, if a single-center large model service can meet all computing power requests during peak periods, it needs to reserve a large amount of resources, which will lead to waste of computing power resources during low request periods. If only limited computing power resources are used, the service stability and reliability may be insufficient to meet the demand during peak periods.
[0054] The related technology adopts a computing power network architecture to build a virtual computing power resource pool by interconnecting multiple computing power centers. The resource pool integrates diversified computing chip resources to provide basic computing power support for large model services. Users do not need to be concerned about the use of the underlying self-owned platform of each computing power center, and can use the computing power resources of multiple computing power centers without awareness, which to some extent alleviates the computing power constraint problem and improves the service stability. However, in the current computing power network for deploying large models, each computing power center can only provide a single type of virtual resource, lacks the ability to dynamically allocate heterogeneous computing resources, and cannot realize hot plug resource scheduling across domain heterogeneous computing power centers, which limits the flexible deployment capability of large model services in different application scenarios.
[0055] Based on this, the embodiments of the present application provide a large model scheduling method, device and equipment in cross-domain computing power centers and a storage medium. Through a unified external interface module, user input is converted into request data of a preset standard rule, ensuring that different computing power centers can receive and process requests without difference, and avoiding adaptation problems caused by interface differences. Next, the request parameters are dynamically adjusted according to the adaptation label of the target computing power center, realizing seamless migration of the request, enhancing deployment flexibility, supporting mixed scheduling of cloud services, bare-metal clusters and third-party APIs, breaking through the limitation of a single type of resource pool, realizing an extensible and hot-pluggable computing power resource pool, and significantly expanding the available computing power range of large models.
[0056] The embodiments of the present application provide a large model scheduling method, device and equipment in cross-domain computing power centers and a storage medium, which are specifically explained by the following embodiments. First, the large model scheduling method in cross-domain computing power centers in the embodiments of the present application is described.
[0057] The embodiments of the present application can acquire and process related data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
[0058] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.
[0059] The large model scheduling method in cross-domain computing power center provided by the embodiments of the present application relates to the technical field of large model deployment. The large model scheduling method in cross-domain computing power center provided by the embodiments of the present application can be applied in a terminal, can also be applied in a server, and can also be a computer program running in a terminal or a server. For example, the computer program can be a native program or a software module in the operating system; it can be a native application program (APP), that is, a program that needs to be installed in the operating system to run, such as a client supporting the scheduling of large models in cross-domain computing power centers, that is, a program that can run only by being downloaded into a browser environment; it can also be a small program that can be embedded into any APP. In short, the above computer program can be any form of application program, module or plug-in. Among them, the terminal communicates with the server through the network. The large model scheduling method in cross-domain computing power center can be executed by the terminal or the server, or cooperatively executed by the terminal and the server.
[0060] In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart watch, or the like. The server can be a standalone server, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms; or a service node in a blockchain system, the service nodes in the blockchain system form a peer-to-peer (P2P) network, and the P2P protocol is an application layer protocol running on the transmission control protocol (TCP) protocol. The terminal and the server can be connected through a communication connection mode such as Bluetooth, universal serial bus (USB), or a network, which is not limited in the embodiment.
[0061] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, in which tasks are performed by remote processing devices connected by a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0062] In an embodiment, referring to Figure 1 , Figure 1 is a structural schematic diagram of a cross-domain computing power center scheduling system corresponding to the scheduling method of the large model in the embodiment of the present application.
[0063] Referring to Figure 1 , the cross-domain computing power center scheduling system includes a unified external interface module, an inference service integration module, and a dynamic large model deployment module. The unified external interface module is used to provide a standardized external service entrance, and uniformly receives user input and returns an initial dialogue request. The inference service integration module is used to be responsible for connecting the initial dialogue request and the backend large model service. The dynamic large model deployment module is used to realize the specific deployment process of the large model service in the computing power center according to the overall service load.
[0064] In one embodiment, the inference service integration module also includes: a request queue and load management module for dynamically monitoring the load of each computing center and reasonably distributing inference requests; an inference interface adaptation module for unifying the interface format of different large model inference services and shielding heterogeneous differences; and an automatic health detection module for periodically detecting the status of large model services, automatically eliminating abnormal nodes, and ensuring stability.
[0065] In one embodiment, the dynamic large model deployment module includes: a large model automatic deployment and recycling layer that automatically starts or releases large model service instances on demand; and a heterogeneous computing power interface adapter module that adapts to the deployment interfaces of different types of computing power centers to ensure a unified calling experience.
[0066] In one embodiment, the cross-domain computing center's scheduling system utilizes a dynamic large model deployment module to connect to the resource management interfaces of each computing center, enabling on-demand deployment, online migration, instance scaling, and dynamic resource release. In actual operation, the cross-domain computing center's scheduling system automatically triggers large model pre-deployment and cooling and recycling operations based on user request trends, the load of each center, and the frequency of large model calls, achieving closed-loop control of the lifecycle management and deployment of large model services across different computing centers.
[0067] Furthermore, the dynamic large model deployment module supports concurrent deployment scheduling for multiple models, versions, and instances, ensuring continuous availability and efficient resource utilization even in the face of high concurrent access or resource fluctuations. Furthermore, deployment behavior is closely integrated with health monitoring mechanisms and task scheduling strategies, enabling fallback, isolation, and automatic redeployment of abnormal instances, providing greater resilience and self-healing capabilities for the system.
[0068] In an embodiment, the large model automatic deployment and recycling layer is mainly responsible for fine management of the life cycle of the large model, and combines task queue pressure, large model activity evaluation and node resource usage state to complete dynamic onboarding and offboarding operations of the large model. Among them, the deployment trigger logic includes but is not limited to the following scenarios: first call of a new large model, first request of an existing large model on a new node, large model request volume reaching expansion threshold, etc. The recycling strategy considers factors such as large model idle time, node resource occupancy rate, and cold resource recycling priority, and uses a hierarchical threshold mechanism to ensure the priority of core large model services and the timely release of edge large models. All deployment operations in this process are issued to the control interface of the target computing center through standardized deployment instructions, supporting deployment orchestration under multi-container platforms (such as Docker, Kubernetes, Slurm, etc.), and pre-warming logic (such as loading large model weights, initializing inference services) can be configured to improve the first response speed. In addition, version control information and node running context of the large model service are maintained during the deployment process, supporting monitoring, rollback and retry of the large model instance state to ensure that the deployment process is observable, controllable and traceable.
[0069] In an embodiment, the heterogeneous computing interface adaptation module is used to shield the interface differences and running environment differences of different computing centers in large model deployment and invocation, and to realize transparent access capability of large model services across platforms, languages and frameworks. Through a unified intermediate adaptation layer, operations such as large model loading, service registration, state monitoring and termination invocation of different platforms are protocol-converted, and heterogeneous service capability interfaces are uniformly abstracted into a standard API template; and the large model format and running dependencies of each domestic computing card are automatically identified and environment-mapped, and a containerized template is used to build the deployment environment to avoid compatibility problems during large model migration. Combined with the inference interface adaptation layer, the input and output data format, error handling mechanism, and invocation context management of the large model service are uniformly abstracted, so that the scheduling system does not need to perceive the differences between the back-end platforms, and can initiate standardized large model service deployment and invocation requests. This layer significantly improves the adaptation capability of the scheduling system to multi-source heterogeneous resources, and provides a foundation for realizing cross-platform large model rapid deployment, dynamic switching and fault rollback.
[0070] In an embodiment, the scheduling system of the large model across domain computing centers adopts a modular and loosely coupled architecture design, ensuring that each module can be independently extended and flexibly combined, and building an efficient, stable and scalable large model dialogue service platform.
[0071] The request input of the user first enters the scheduling system through the unified external interface module, is uniformly delivered to the back-end service after standardized processing, and ensures decoupling of the front-end and the back-end and supports multiple calling scenarios. Subsequently, the reasoning service integration module interfaces the user request and the large model service, the internal request queue and the load management component dynamically schedule according to the real-time state of the computing power center, and the reasoning interface adaptation module is responsible for uniformly calling the API of the large model of various computing power centers, and shielding the interface differences between different suppliers; the automatic health detection module continuously monitors the service availability, automatically removes abnormal nodes, and improves the overall stability and fault tolerance. At the bottom, the large model dynamic deployment module dynamically adjusts the deployment state of the large model according to the current request load, the large model automatic deployment and recovery layer realizes the elastic scaling of the service instance, and the heterogeneous computing power interface adaptation module ensures the compatibility with multiple computing power centers, so that the system can flexibly adapt to different computing resources and large model types. The entire system supports asynchronous processing and high-concurrency scenarios, and can maintain stable response under large-scale access. Through the cooperation of unified entry, intelligent scheduling, automatic deployment and other mechanisms, the rapid response of user requests and the continuous availability of services are effectively guaranteed, and the business needs in various large model service scenarios are met.
[0072] The following describes the scheduling method of the large model in the cross-domain computing power center in the embodiments of the present application. Figure 1 The scheduling method of the large model in the cross-domain computing power center in the embodiments of the present application.
[0073] Figure 2 is an optional flowchart of the scheduling method of the large model in the cross-domain computing power center provided by the embodiments of the present application, Figure 2 The method in the embodiments of the present application can include but is not limited to steps 110-140. It can be understood that the order of steps 110-140 in the embodiments is not limited, and the order of steps can be adjusted or some steps can be reduced or added according to actual needs. Figure 2 The order of steps 110-140 in the embodiments is not limited, and the order of steps can be adjusted or some steps can be reduced or added according to actual needs.
[0074] Step 110: Obtain an initial dialogue request from an asynchronous task queue.
[0075] In an embodiment, referring to Figure 1 , the user generates a content session of the large model through the display interface corresponding to the unified external interface module, inputs dialogue information in the content session, and the unified external interface module converts the dialogue information into request data according to the preset standard rules after responding to the user input, to obtain an initial dialogue request. Considering the possibility that multiple users may initiate a session at the same time, an asynchronous task queue is constructed in the reasoning service integration module, and the initial dialogue request is processed in order.
[0076] For example, if three users (UserA, UserB, and UserC) simultaneously initiate a large-model content session through a unified external interface module, the scheduling system needs to coordinate the resources of the three computing centers to complete the response. The resource environment of the computing center is not limited here and can be a cross-domain resource environment, such as cloud services, bare metal clusters, or third-party API gateways.
[0077] At this point, User A enters the following text into the interface: "Help me generate a popular science article about carbon neutrality, approximately 500 words." User B enters the following text into the interface: "Write a quick sort code in Python, with comments." User C enters the following text into the interface: "Explain the self-attention mechanism of the Transformer large model." At this point, the display interface of the unified external interface module of the scheduling system encapsulates the user input into structured request data as the initial conversation request. For example, User A's initial conversation request could be:
[0078] {"request_id":"REQ-2024-1001",
[0079] "user_input":"Help me generate a popular science article about carbon neutrality, about 500 words",
[0080] "model_type":"llm-text-gen", / / Large model type specified by preset rules
[0081] "priority": "normal" / / Default priority}
[0082] Then the three request data are asynchronously pushed into the task queue (such as Redis / Kafka queue), and a queue prompt is returned to the user on the interface, for example: "Your request has been received and is being queued for processing..."
[0083] Next, the inference service integration module takes out the initial dialogue requests from the asynchronous task queue in sequence, and triggers the scheduling of the cross-domain computing center for each initial dialogue request.
[0084] It can be seen that the unified external interface module serves as the entrance of the scheduling system, is responsible for receiving dialogue requests from users, and forwards them to the back-end reasoning service integration module for processing. It provides a consistent service access method to the outside through a unified API interface, shielding the resource differences of different computing power centers. When the server starts, it automatically loads the information of the currently available computing power centers and initializes whether the services of each computing power center are available, preparing for subsequent intelligent distribution. At the same time, the unified external interface module also regularly checks the service availability status of each computing power center to ensure that only available service nodes are selected during task scheduling. As a key link in the system, the unified external interface module not only bears the responsibility of request access, but also provides stable, efficient, and scalable service access capabilities to the outside by combining databases and scheduling logic, providing a foundation guarantee for subsequent module task processing and resource allocation.
[0085] Step 120: Calculate the real-time load score of the accessed multiple computing power centers, and select at least one target computing power center from the multiple computing power centers according to the real-time load score.
[0086] In an embodiment, at least two computing power centers are located in a cross-domain resource environment, which includes cloud services, bare-metal clusters, third-party API gateways, etc., and is an uncontrollable computing power resource environment. The cross-domain computing power resource environment will face many inconsistencies in API protocols, parameter structures, authentication methods, and response formats. In particular, for those uncontrollable third-party computing power centers, such as OpenAI interface, Ascend computing power service, ModelArts multi-tenant instance, MindIE local cluster, CloudBrain platform, domestic AI inference framework, or self-built API gateway of enterprises and institutions, etc. Due to deployment mechanisms, large model packaging methods or business strategy restrictions, these platforms often cannot use a unified service deployment framework for internal unified packaging. Therefore, the embodiment of the present application mainly faces such uncontrollable computing power network scenarios, realizes highly decoupled, dynamically adaptive, and hot-pluggable large model adaptive deployment.
[0087] It can be understood that each computing power center accesses the heterogeneous algorithm interface adaptation layer of the dynamic large model deployment module through the corresponding computing power service API. And at least one large model is deployed on each computing power center, and the ultimate goal is to select a large model for execution for each initial dialogue request.
[0088] In an embodiment, content distribution needs to be performed according to the real-time load score of each accessed computing power center. Referring to Figure 3 , Figure 3is a flowchart provided by an embodiment of the application for calculating real-time load scores of multiple accessed computing power centers and selecting at least one target computing power center from the multiple computing power centers according to the real-time load scores, and specifically includes the following steps:
[0089] Step 310: Calculate the real-time load score according to the obtained task state information, resource load index and platform capability label of each computing power center.
[0090] In an embodiment, the task state information can be the current number of concurrent tasks of the computing power center, the resource load index can be GPU utilization, memory occupancy, request processing delay, throughput, etc., and the platform capability label can be the operation capacity of the deployed large model. Then the obtained task state information, resource load index and platform capability label are normalized, and weighted calculation is performed according to the preset weighting rule to obtain the real-time load score of each computing power center.
[0091] Step 320: The computing power center with a real-time load score greater than or equal to a preset score is regarded as a candidate computing power center, a candidate score is calculated according to the adaptation level, interface stability parameter and stream inference capability index of each candidate computing power center, and the candidate computing power center corresponding to the maximum candidate score is selected as the target computing power center.
[0092] In an embodiment, if the real-time load score is greater than or equal to the preset score (for example, 50 points), the corresponding computing power center is regarded as a candidate computing power center. At this time, the number of candidate computing power centers can be more than one, so further screening is required.
[0093] Specifically, the adaptation level, interface stability parameter and stream inference capability index of each candidate computing power center are obtained. The adaptation level is used to judge the matching degree of the large model deployed in the candidate computing power center and the current initial dialogue request (for example, a literature optimization large model is preferentially selected for a poem generation task). The interface stability is a stability score based on historical success rate. The stream inference capability index is a Boolean value, full score for supporting stream inference, and zero score otherwise.
[0094] The weights corresponding to the adaptation level, interface stability parameter and stream inference capability index are set according to actual requirements, the candidate scores of each candidate computing power center are calculated in a weighted calculation manner, and the candidate computing power center with the highest candidate score is selected as the target computing power center assigned to the initial dialogue request.
[0095] Step 330: If each real-time load score is less than the preset score, the preset score is reduced by a preset step size until the target computing power center is selected.
[0096] In an embodiment, if all real-time load scores of all computing power centers are less than the preset score or part of the platform responds abnormally during the high incidence period, the preset score is automatically degraded, the preset score is gradually reduced according to the preset step, until the candidate computing power center is selected, and then the target computing power center is selected. The purpose is to ensure the continuity of the service.
[0097] It can be understood that if multiple large models are deployed on the computing power center, the above process of calculating the real-time load score refers to the undistributed large model.
[0098] Step 130: Obtain the corresponding adaptation parameter data according to the adaptation label of the target computing power center, and convert the initial dialogue request into a target dialogue request corresponding to the target computing power center according to the adaptation parameter data.
[0099] In an embodiment, after the target computing power center is obtained, cross-domain computing power center adaptation is needed. Referring to Figure 4 , Figure 4 is a flowchart provided by the embodiment of the present application for obtaining the corresponding adaptation parameter data according to the adaptation label of the target computing power center, and specifically includes the following steps:
[0100] Step 410: If the adaptation label indicates that the adaptation state of the target computing power center is adapted, obtain the adaptation parameter data of the target computing power center.
[0101] In an embodiment, in order to improve the adaptation efficiency, the corresponding adaptation label is set for each computing power center, which is used to indicate whether the computing power center has been adapted. If the adaptation label indicates that the adaptation state of the target computing power center is adapted, the stored adaptation parameter data of the computing power center can be directly obtained.
[0102] Step 420: If the adaptation label indicates that the adaptation state of the target computing power center is not adapted, the target computing power center is adapted to obtain the corresponding adaptation parameter data.
[0103] In an embodiment, if the adaptation label indicates that the adaptation state of the target computing power center is not adapted, that is, the computing power center is accessed to the dispatching system of the cross-domain computing power center for the first time, real-time adaptation is needed. Referring to Figure 5 , Figure 5 is a flowchart provided by the embodiment of the present application for adapting the target computing power center to obtain the corresponding adaptation parameter data, and specifically includes the following steps:
[0104] Step 510: Obtain the adaptation reference data of the target computing power center.
[0105] In one embodiment, the adaptation reference data of the target computing power center includes at least: parameter structure data, authentication signature mechanism, platform behavior encapsulation mechanism and response parsing mechanism. Among them, the parameter structure data includes at least field mapping rules and nested structures. The field mapping rules define the parameter conversion logic within the computing power center (such as field name conversion, enumeration value mapping, priority mapping, etc.) to ensure cross-system parameter compatibility. The nested structure is used to define the hierarchical data design logic of the computing power center. The authentication signature mechanism includes at least authentication rules, which are used to stipulate the signature generation algorithm (such as HMAC-SHA256) of the computing power center, the signature verification process, and the timestamp / random number anti-replay attack security policies. The platform behavior encapsulation mechanism includes at least platform behavior rules, which are used to define the processing flow of the computing power center for requests. The response parsing mechanism includes at least response types, which are used to define the content output method deployed to the large model on the computing power center.
[0106] Step 520: Perform adaptation operations on the target computing power center according to the adaptation reference data to obtain corresponding structured parameters, access verification data, request setting rules and standard response output format.
[0107] In one embodiment, referring to Figure 6 , Figure 6 This is a flowchart of performing adaptation operations on target computing centers according to adaptation reference data provided by an embodiment of the present application to obtain corresponding structural parameters, access verification data, request setting rules, and standard response output formats, specifically including the following steps:
[0108] Step 610: traverse each standardized parameter in the preset standard rules, obtain at least the field mapping rules and nested structure from the parameter structure data, convert the standardized parameters into target field parameters according to the field mapping rules, and convert the target field parameters into structured parameters according to the nested structure.
[0109] In one embodiment, the preset standard rules are rules defined in a unified external interface and are used to convert user input into initial dialog requests. These preset standard rules include multiple standardized parameters, such as prompt, history, and temperature. Field mapping rules and nested structures are obtained from the parameter structure data of the target computing power center, and the parameter structure mapping process is performed accordingly. The standardized parameters are then converted according to the field mapping rules and converted into target field parameters.
[0110] For example, the initial dialog request corresponding to the standardized parameters is expressed as:
[0111] {
[0112] "prompt":"Hello, please write a poem",
[0113] "history":[{"user":"What animal do you like?","bot":"I like pandas."}],
[0114] "temperature":0.7,
[0115] "model_name":"gpt-4"}
[0116] Field mapping rules are used to convert standardized parameters into field names used by the target computing center. For example, the field mapping rules are expressed as:
[0117] prompt:"inputs.text"
[0118] history:"inputs.context"
[0119] temperature:"params.temperature"
[0120] model_name:"url_path"#means inserting URL;
[0121] Therefore, the target field parameter is expressed as:
[0122] {"inputs":{
[0123] "text":"Hello, please write a poem",
[0124] "context":[
[0125] {"user":"What animal do you like?","bot":"I like pandas."}]},
[0126] "params":{"temperature":0.7}};
[0127] Then, the target field parameters are converted into structured parameters according to the nested structure. Assuming the nested structure is: history is converted into role / content format, the structured parameters are expressed as:
[0128] Body:{
[0129] "data":{
[0130] "input":"Hello, please write a poem",
[0131] "past_interactions":[
[0132] {"role":"user","content":"What animal do you like?"},
[0133] {"role":"assistant","content":"I like pandas."}]}
[0134] "config":{"randomness":0.7},
[0135] "model":"gpt-4"};
[0136] It can be understood that the nested structure of the above field mapping rules is only illustrative and does not represent a limitation.
[0137] Step 620: Obtain the standard credential field corresponding to the preset standard rule, obtain the authentication rule from the authentication signature mechanism, and generate the access verification data of the target computing power center based on the authentication rule and the standard credential field.
[0138] In an embodiment, in order to ensure that the initial dialogue request can pass the identity authentication and access verification of different computing power centers, the reasoning interface adaptation module needs to support multiple authentication methods and dynamically adapt to the requirements of the target computing power center. First, obtain the standard credential field corresponding to the preset standard rule, such as "api_key", "secret_key", "auth_type", etc., then obtain the authentication rule from the authentication signature mechanism of the target computing power center, and generate the access verification data of the target computing power center based on the authentication rule and the standard credential field.
[0139] For example, the standard credential field is represented as:
[0140]
[0141] According to the auth_type field, match the rule template in the authentication signature mechanism, and the obtained authentication rule is represented as:
[0142]
[0143] Based on the authentication rule and the standard credential field, the access verification data of the target computing power center is gradually generated. First, generate_timestamp() is called to generate the dynamic parameter timestamp=1710000000, and generate_nonce() is called to generate the dynamic parameter nonce=a1b2c3d4. Next, the string is spliced according to the authentication rule, sign_string=f"{timestamp}{nonce}{standard_credentials['secret_key']}", and then the "secret_key" of the standard credential field is used to calculate the signature value, which is represented as:
[0144] signature=hmac_sha256(secret_key=standard_credentials["secret_key"], data=sign_string)
[0145] Finally, the above data is assembled to obtain the access verification data, which is represented as:
[0146]
[0147] In the above process, the standard credential data configured by the system is decoupled from the specific computing power center, and the generated access verification data (such as signature, timestamp) can be directly injected into the HTTP request through the adaptation process of the authentication rule and the multi-platform adaptation strategy in the authentication signature mechanism, so as to realize "one configuration, multi-platform adaptation", without the need to develop authentication logic for each computing power center. It can be understood that the above authentication rule and standard credential field are only illustrative and do not represent a limitation. The authentication rule can include static APIKey, HMAC-SHA256 dynamic signature, AK / SK structured authentication, BearerToken injection, or dynamic calculation of timestamp and random factor through signature header, etc.
[0148] Step 630: Obtain the request setting rule according to the platform behavior rule obtained from the platform behavior packaging mechanism.
[0149] In an embodiment, the request setting rule of the platform behavior packaging mechanism can be a rule defined by the target computing power center. The request setting rule includes a pre-request rule and / or a post-request rule, and the request setting rule is used to perform corresponding operations on the specific initial dialogue request. It can be understood that the pre-request rule and the post-request rule can exist simultaneously or be selected according to the actual scene. In addition, the request preprocessing operation can be performed to reconstruct the initial dialogue request after execution.
[0150] For example, set before obtaining the initial dialogue request, at least one of the request preprocessing operations of session initialization, preset vocabulary uploading, preloading context is carried out, these request preprocessing operations are taken as the request front rule. It can also be set after obtaining the initial dialogue request, at least one of the request post-processing operations of polling processing, stream decoding is carried out, and these request post-processing operations are taken as the request post rule.
[0151] For example, the request front rule is represented as:
[0152] request_pre_rules:
[0153] "init_session" # session initialization
[0154] "upload_vocab" # vocabulary uploading
[0155] The request post rule is represented as:
[0156] request_post_rules:
[0157] "poll_status" # polling processing
[0158] "decode_stream" # stream decoding
[0159] At this time, before obtaining the initial dialogue request, the request front rule is used to carry out two request preprocessing operations of session initialization and preset vocabulary uploading. Among them, when the session initialization is called, the / session / init interface of the target computing center is called, the inference session is created, and the returned session_id is injected into the subsequent request header, for example, X-Session-ID:sess_123456 is generated. And when the preset vocabulary uploading is uploaded, the user provided domain vocabulary (such as medical terms) is uploaded to the target computing power, and the returned vocab_id of the target computing center is added to the request body, for example, {"prompt":"...","vocab_id":"vocab_789"} is generated.
[0160] After the request preprocessing operation according to the request front rule is carried out, the initial dialogue request is obtained, and the polling processing and stream decoding are carried out. Among them, if the target computing center returns an asynchronous task, the / task / status is continuously polled until it is completed, and finally a unified response is generated, the request is represented as: {"response":"the result obtained by polling...","status":"completed"}. Next, the stream response of the target computing center is blocked according to the punctuation, and each segment is pushed to the display interface of the unified external interface.
[0161] In the above process, the computing center defines platform behavior rules specific to the computing center using a platform behavior encapsulation mechanism, such as pre-dependence (such as session initialization) and post-logic (such as polling). According to the platform behavior rules, the process of setting the request rules can convert the platform behavior rules into executable operation steps, thereby completely decoupling from the platform. Using a Hook mechanism, the request preprocessing operation (Pre-Hook) and the request post-processing operation (Post-Hook) are dynamically inserted to achieve adaptation. Thus, the embodiments of the present application can flexibly extend new computing centers, only need to add the configuration process of the platform behavior rules, without modifying the core code, thereby improving the flexibility of adaptation.
[0162] It can be understood that the above platform behavior rules are only illustrative and do not limit them.
[0163] Step 640: Obtain the response type from the response analysis mechanism, obtain the corresponding analysis logic according to the response type, and obtain the standard response output format based on the analysis logic.
[0164] In an embodiment, the corresponding type can be automatically identified by analyzing the characteristic data of the original response of the target computing center.
[0165] For example, the identification algorithm can be as follows:
[0166]
[0167] Among them, the response type SSE_STREAM is used for the server to push real-time data stream to the client. The response type ASYNC_TASK is used to avoid blocking the main thread when processing time-consuming tasks. The response type JSON_ARRAY is used to return structured batch data. The response type CHUNKED_STREAM is used to send step by step when transmitting data streams of unknown size. The response type BLOCKING_RESPONSE is a general synchronous request-response mode.
[0168] With the response type, the corresponding analysis logic needs to be obtained according to the response type, and the standard response output format is obtained based on the analysis logic. Taking the streaming response SSE_STREAM as an example, the analysis logic is: split the event according to the SSE protocol, extract the text field and mark the completion status.
[0169] Suppose the data to be responded is represented as:
[0170]
[0171] It is understood that the standard response output format is used to indicate how the generated content of the initial dialogue request should be output to the unified external interface. The above response types, parsing logic, standard response output format, etc. are only illustrative and do not represent limitations.
[0172] In one embodiment, if the parsing logic is streaming response parsing, the streaming protocol type is obtained from the response parsing mechanism, and the parsing rules for field extraction and semantic reorganization are determined according to the streaming protocol type. This is because although the response types of some computing power centers are all streaming response parsing, the specific streaming protocol types are different, resulting in different specific streaming output formats, such as SSE event stream format, JSON array multi-segment return format, HTTP long connection package format encoded by chunk, or multi-field nested package format returned by some computing power centers using bare metal server deployment (for example, type="log" / "answer" / "end" structure). Therefore, the embodiment of the present application sets different parsers for different formats, determines the parsing rules for field extraction and semantic reorganization according to the streaming protocol type, automatically identifies each segment identifier, content field and terminator, and performs semantic enrichment. Thereby, a consistent standard response output format is generated for streaming reasoning of different formats.
[0173] After obtaining the adaptation parameter data through the above process, the specific initial dialogue request is filled with data according to the adaptation parameter data and converted into a target dialogue request corresponding to the target computing power center.
[0174] For example, an adaptation parameter data is expressed as:
[0175] {
[0176] "request_format":{
[0177] "base_url":"https: / / api.target-platform.com / v2",
[0178] "model_placement":"path", / / The large model name is in the URL path
[0179] "required_headers":{
[0180] "Authorization":"Bearer{api_key}",
[0181] "X-Session-ID":"{session_id}"},
[0182] "body_structure":{
[0183] "prompt":"inputs.text",
[0184] "temperature":"parameters.temperature",
[0185] "history":"context.messages"}},
[0186] "auth_data":{
[0187] "api_key":"sk_xyz_789",
[0188] "session_id":"sess_123456"},
[0189] "preprocessed_data":{
[0190] "vocab_id":"vocab_789"}}
[0191] At this point, the initial conversation request is expressed as:
[0192] {
[0193] "prompt":"Please explain quantum computing",
[0194] "temperature":0.7,
[0195] "history":[
[0196] {"role":"user","content":"What is a qubit?"},
[0197] {"role":"assistant","content":"A quantum bit is..."}]}
[0198] Fill in the data according to the adaptation parameter data and convert it into the target dialogue request corresponding to the target computing power center, which is expressed as:
[0199] {"url":"https: / / api.target-platform.com / v2 / models / qwen-plus / generate",
[0200] "method":"POST",
[0201] "headers":{
[0202] "Authorization":"Bearersk_xyz_789",
[0203] "X-Session-ID":"sess_123456",
[0204] "X-Vocab-ID":"vocab_789"},
[0205] "body":{
[0206] "inputs":{
[0207] "text":"Please explain quantum computing",
[0208] "context":{
[0209] "messages":[
[0210] {"speaker":"human","message":"What is a qubit?"},
[0211] {"speaker":"bot","message":"Qubits are..."}]}},
[0212] "parameters":{
[0213] "temperature":0.7,
[0214] "stream":true}}}
[0215] Step 140: Determine the target large model from the large models deployed in the target computing center, send the target dialogue request to the target large model for content generation, and obtain the inference result.
[0216] In one embodiment, an available large model is selected from a target computing power center as a target large model, and a target dialogue request is sent to the target large model for content generation to obtain an inference result.
[0217] In one embodiment, the structured decomposition and display logic of the "thinking process" and "final answer" can also be performed in the process of displaying the reasoning results. That is, the thinking process data and the final answer data are obtained from the reasoning results, and the final answer data is displayed through a unified external interface module. The thinking process data is lazy loaded or the content is folded according to the preset display rules. The preset display rules here include lazy loading of the thinking process data, content folding of the thinking process data, or not displaying the thinking process data, thereby improving the consistency of the user experience.
[0218] In addition, for long dialogue or multi-turn interaction scenarios, the inference interface adaptation module can also perform multi-segment stream merging and splicing, content deduplication, and breakpoint resuming, and the like, so as to adapt to abnormal situations such as network instability or inference interruption.
[0219] In an embodiment, in order to improve convenience, an independent interface mapping channel can be set for each type of computing power center in advance, a plug-in architecture is used to form a hot-pluggable adapter component, hard coding dependence is avoided, and the system has real-time configuration and real-time extension access capabilities. The entire inference interface adaptation module can shield the system complexity caused by the heterogeneity of computing power centers, so that the system can efficiently interface with various heterogeneous resources including virtual machines, bare metal nodes, cloud APIs, multi-tenant large model containers, and the like, and realize unified calling and response feedback of intelligent inference across domains, platforms, and organizations.
[0220] In an embodiment, referring to Figure 7 , Figure 7 is another schematic diagram of scheduling and distribution of a large model in a cross-domain computing power center provided by an embodiment of the present application.
[0221] In order to ensure stable access and continuous availability of cross-domain multi-source computing power centers, the scheduling system of the cross-domain computing power center of the present application further sets an automatic health detection module with intelligent judgment and dynamic feedback capability. The module sends a health check request in a non-intrusive manner, continuously monitors the availability status returned by all accessed computing power centers, and returns the availability status to the inference service integration module for scheduling target computing power centers. The availability status can be obtained from multiple dimensions such as network reachability, interface response delay, authentication validity, interface structure change, and inference service availability.
[0222] In addition, the cross-domain computing power center scheduling system periodically initiates concurrent detection tasks for each computing power platform, including: standardized detection requests (such as low-load probe calls of large model inference interfaces), platform-specific heartbeat interface calls, and rapid verification of authentication mechanism effectiveness. The detection mechanism not only judges whether the platform is "accessible", but also focuses on "whether it has the ability to stably process tasks", such as whether the response structure is abnormal, whether the service frequently times out, whether the intermediate state is disordered, and whether the streaming inference is interrupted, and the like. Once a node has a request failure rate increase, interface structure drift, or delay anomaly, and the like, the cross-domain computing power center scheduling system will dynamically adjust its health score according to the pre-defined threshold strategy, and temporarily exclude it from the main scheduling path, avoid allocating new inference tasks, and ensure that the overall service quality is not affected.
[0223] And for the rejected platform, the system enters a fallback monitoring period, periodically initiating a retry probe, and if the platform state returns to above the stable threshold, it will automatically be included in the scheduling candidate set. All health status evaluation results and historical evolution processes are recorded in the internal health status database, supporting subsequent adaptive optimization of scheduling strategies and analysis of platform availability trends. The overall mechanism significantly improves the scheduling system's ability to flexibly schedule and ensure service continuity in the face of uncontrollable computing power environments, especially in scenarios with dynamic platform resource access, high interface instability, or frequent external platform policy changes.
[0224] As can be seen, the scheduling system of the cross-domain computing power center realizes unified management and efficient invocation of different computing power centers through technologies such as dynamic service deployment, API adaptation, intelligent scheduling, and real-time monitoring. In terms of dynamic service deployment, the system adapts the computing power service interfaces of different computing power centers, enabling large models to be deployed in heterogeneous computing power centers; in terms of API adaptation, the system uniformly interfaces with the API interface types provided by different computing power centers such as OpenAI, ModelArts, and MindIE, and standardizes processing in terms of data format, request method, etc., enabling users to access different APIs through a unified interface without additional adaptation. During task scheduling, the system selects the most suitable computing power center to execute tasks based on the real-time load score of each computing power center.
[0225] Through the coordinated work of the above parts, the scheduling system of the cross-domain computing power center can effectively reduce the cost of users deploying model services and providing model service APIs, improve the success rate of service requests, and fully utilize the computing resources of multiple intelligent computing centers to achieve stable and efficient distribution of model service APIs.
[0226] In an embodiment, the scheduling system of the cross-domain computing power center continues to track the status of distributed requests after task scheduling is complete, collects real-time operation feedback from each computing power platform, and dynamically adjusts the real-time load score and health status, thereby realizing adaptive evolution of the scheduling strategy. The module as a whole adopts a decoupled architecture, supports shielding the differences and details of each computing power center at the scheduling layer, encapsulates complex resource access and behavior adaptation logic in the interface adaptation layer and platform configuration, and thereby realizes intelligent, stable, and efficient distribution of requests among multiple source uncontrollable computing power centers.
[0227] Therefore, the large model scheduling method in the cross-domain computing power center has the advantages of low resource demand, flexibility, and service stability. Multiple computing power centers are deployed with corresponding large model services, and the optimal computing power center is automatically selected based on real-time load scores for request distribution, thereby improving the stability and response speed of the overall service. Instead of reserving a large amount of computing power resources in a single computing power center, the scattered computing power of multiple computing power centers is utilized.
[0228] The technical scheme provided by the embodiment of the application obtains an initial dialogue request from an asynchronous task queue, wherein the initial dialogue request is request data of a preset standard rule obtained by a unified external interface module in response to user input; real-time load scores of a plurality of accessed computing power centers are calculated, and at least one target computing power center is selected from the plurality of computing power centers according to the real-time load scores, at least two computing power centers are located in a cross-domain resource environment, at least one large model is deployed on each computing power center, and the cross-domain resource environment at least includes at least two of cloud services, bare-metal clusters or third-party API gateways; corresponding adaptation parameter data is obtained according to an adaptation label of the target computing power center, and the initial dialogue request is converted into a target dialogue request corresponding to the target computing power center according to the adaptation parameter data; a target large model is determined from the large models deployed by the target computing power center, and the target dialogue request is sent to the target large model for content generation to obtain an inference result. In the embodiment of the application, the user input is converted into the request data of the preset standard rule by the unified external interface module, so that different computing power centers can receive and process the request without difference, and the adaptation problem caused by interface difference is avoided. Next, the request parameters are dynamically adjusted according to the adaptation label of the target computing power center, the seamless migration of the request across domains is realized, the deployment flexibility is enhanced, the mixed scheduling of cloud services, bare-metal clusters and third-party APIs is supported, the limitation of a single type of resource pool is broken through, the scalable and hot-pluggable computing power resource pool is realized, and the available computing power range of the large model is significantly expanded.
[0229] The embodiment of the application also provides a large model scheduling device in a cross-domain computing power center, which can realize the large model scheduling method in the cross-domain computing power center. Figure 8 The device comprises:
[0230] The dialogue request obtaining module 810 is configured to obtain an initial dialogue request from an asynchronous task queue, wherein the initial dialogue request is request data of a preset standard rule obtained by a unified external interface module in response to user input.
[0231] The computing power scheduling module 820 is configured to calculate real-time load scores of a plurality of accessed computing power centers, and select at least one target computing power center from the plurality of computing power centers according to the real-time load scores, wherein at least two computing power centers are located in a cross-domain resource environment, at least one large model is deployed on each computing power center, and the cross-domain resource environment at least includes at least two of cloud services, bare-metal clusters or third-party API gateways.
[0232] The adaptation module 830 is configured to obtain corresponding adaptation parameter data according to an adaptation label of the target computing power center, and convert the initial dialogue request into a target dialogue request corresponding to the target computing power center according to the adaptation parameter data.
[0233] The content generation module 840 is configured to determine a target large model from the large models deployed by the target computing power center, send a target dialogue request to the target large model for content generation, and obtain an inference result.
[0234] The large model in the embodiment has a substantially same implementation as the above-mentioned large model in the method for scheduling across computing power centers, and thus will not be described herein.
[0235] The embodiment of the present application further provides an electronic device, which comprises:
[0236] at least one memory;
[0237] at least one processor;
[0238] at least one program;
[0239] The program is stored in the memory, and the processor executes the at least one program to implement the above-mentioned method for scheduling a large model across computing power centers according to the embodiment of the present application. The electronic device can be any intelligent terminal, such as a mobile phone, a tablet computer, a personal digital assistant (PDA), a vehicle-mounted computer, etc.
[0240] Please refer to Figure 9 , Figure 9 The hardware structure of the electronic device of another embodiment is illustrated, which comprises:
[0241] The processor 901 can be implemented in a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute related programs to implement the technical solutions provided by the embodiments of the present application.
[0242] The memory 902 can be implemented in a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 902 can store an operating system and other application programs, and when the technical solutions provided by the embodiments of the present application are implemented by software or firmware, the related program codes are stored in the memory 902 and are called and executed by the processor 901 to implement the method for scheduling a large model across computing power centers according to the embodiments of the present application.
[0243] The input / output interface 903 is configured to realize information input and output.
[0244] The communication interface 904 is configured to realize the communication interaction between the device and other devices. The communication can be realized in a wired manner (for example, a USB, a network cable, and the like) or in a wireless manner (for example, a mobile network, WIFI, Bluetooth, and the like).
[0245] The bus 905 is configured to transmit information between various components (for example, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904) of the device.
[0246] The processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are connected to each other through the bus 905 to realize the communication connection between the device.
[0247] The embodiment of the present application further provides a storage medium. The storage medium is a storage medium, and the storage medium stores a computer program. The computer program is executed by a processor to realize the above-mentioned method for scheduling a large model in a cross-domain computing center.
[0248] The memory is a non-transitory storage medium, and can be used to store a non-transitory software program and a non-transitory computer executable program. In addition, the memory can include a high-speed random access memory, and can further include a non-transitory memory, for example, at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged relative to the processor, and the remote memory can be connected to the processor through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0249] The large model scheduling method, device, equipment and storage medium across the domain computing center are provided in the embodiments of the present application. An initial dialogue request is obtained from an asynchronous task queue, wherein the initial dialogue request is request data of a preset standard rule obtained by a unified external interface module in response to user input; real-time load scores of a plurality of accessed computing centers are calculated, and at least one target computing center is selected from the plurality of computing centers according to the real-time load scores, wherein the at least two computing centers are located in a cross-domain resource environment, at least one large model is deployed on each computing center, and the cross-domain resource environment at least includes at least two of cloud services, bare-metal clusters or third-party API gateways; corresponding adaptation parameter data is obtained according to the adaptation label of the target computing center, the initial dialogue request is converted into a target dialogue request corresponding to the target computing center according to the adaptation parameter data; a target large model is determined from the large model deployed in the target computing center, and the target dialogue request is sent to the target large model for content generation to obtain an inference result. In the embodiments of the present application, the user input is converted into request data of a preset standard rule through a unified external interface module, so that different computing centers can receive and process requests without difference, and adaptation problems caused by interface differences are avoided. Next, the request parameters are dynamically adjusted according to the adaptation label of the target computing center, the seamless migration of the request across the domain is realized, the deployment flexibility is enhanced, the mixed scheduling of cloud services, bare-metal clusters and third-party APIs is supported, the limitation of a single type of resource pool is broken through, the scalable and hot-pluggable computing resource pool is realized, and the available computing power range of the large model is significantly expanded.
[0250] The embodiments described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of technology and the appearance of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0251] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and can include more or fewer steps than the figures shown, or combine certain steps or different steps.
[0252] The device embodiments described above are only schematic, and the units illustrated as separate components can or can not be physically separated, that is, they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments of the present application.
[0253] Those skilled in the art can understand that all or some steps in the above disclosed method, the function modules / units in the system and the device can be implemented as software, firmware, hardware and their appropriate combinations.
[0254] The terms "first", "second", "third", "fourth", and the like in the description and in the claims of this application, if any, are used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of the terms so termed is interchangeable under appropriate circumstances such that the embodiments of the application described herein are, for example, capable of orderly or chronological mundane operation, reverse order operation, based on circuitry availability, based on stated preference or the like, and that "default" or other orderings are thus permissible. Further, the terms "comprise", "comprising", "include", "including", and the like, are specifically intended to be open-ended. That is, references to individual steps and the like do not suhstantially exclude the presence of two or more of a given step or its integral presence in the process, method, system, article, or apparatus having been made with a wider scope. The use of notation such as "first", "second", "third", etc. does not generally limit the areas, but can be used for clarity, and merely establishes the order unless otherwise stated below.
[0255] It should be understood that, in the application, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the relationship between associated objects, which means that there can be three relationships, for example, "A and / or B" can mean that there are only A, only B, and A and B at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or the like means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can mean a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0256] In several embodiments provided in the application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the above units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be omitted or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed objects can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0257] The units described above as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e. they can be located in one place or distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0258] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.
[0259] If the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in the form of a contribution to the prior art, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods in the embodiments of the present application. The foregoing storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0260] The preferred embodiments of the embodiments of the present application are described above with reference to the accompanying drawings, and are not intended to limit the scope of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the embodiments of the present application.
Claims
1. A scheduling method for a large model in a cross-domain computing center, characterized in that: include: Obtaining an initial dialogue request from the asynchronous task queue, wherein the initial dialogue request is request data of a preset standard rule obtained by the unified external interface module in response to user input; Calculate the real-time load scores of multiple connected computing centers, and select at least one target computing center from the multiple computing centers based on the real-time load scores, at least two of which are located in a cross-domain resource environment, and at least one large model is deployed on each computing center. The cross-domain resource environment includes at least two of: cloud services, bare metal clusters, or third-party API gateways. Acquire corresponding adaptation parameter data according to the adaptation tag of the target computing power center, and convert the initial dialogue request into a target dialogue request corresponding to the target computing power center according to the adaptation parameter data; A target large model is determined from the large models deployed in the target computing center, and the target dialogue request is sent to the target large model for content generation to obtain an inference result.
2. The method for scheduling a large model in a cross-domain computing center according to claim 1 is characterized in that: The acquiring corresponding adaptation parameter data according to the adaptation tag of the target computing power center includes: If the adaptation tag indicates that the adaptation status of the target computing power center is adapted, obtaining the adaptation parameter data of the target computing power center; If the adaptation tag indicates that the adaptation status of the target computing power center is not adapted, an adaptation operation is performed on the target computing power center to obtain the corresponding adaptation parameter data.
3. The method for scheduling a large model in a cross-domain computing center according to claim 2 is characterized in that: The adapting operation is performed on the target computing power center to obtain the corresponding adaptation parameter data, including: Obtaining adaptation reference data of the target computing power center, wherein the adaptation reference data includes at least: parameter structure data, authentication signature mechanism, platform behavior encapsulation mechanism, and response parsing mechanism; According to the adaptation reference data, adaptation operations are performed on the target computing power center respectively to obtain corresponding structural parameters, access verification data, request setting rules and standard response output format.
4. The method for scheduling a large model in a cross-domain computing center according to claim 3 is characterized in that: The adaption operation is performed on the target computing power center according to the adaptation reference data to obtain corresponding structural parameters, access verification data, request setting rules and standard response output format, including: Traversing each standardized parameter in the preset standard rule, obtaining at least a field mapping rule and a nested structure from the parameter structure data, converting the standardized parameter into a target field parameter according to the field mapping rule, and converting the target field parameter into the structured parameter according to the nested structure; Obtaining a standard credential field corresponding to the preset standard rule, obtaining an authentication rule from the authentication signature mechanism, and generating the access verification data of the target computing power center based on the authentication rule and the standard credential field; The request setting rules are obtained according to the platform behavior rules obtained from the platform behavior encapsulation mechanism, the request setting rules including pre-request rules and / or post-request rules, the pre-request rules being used to set at least one request pre-processing operation of session initialization, preset vocabulary upload, and context preloading before obtaining the initial dialogue request, and the post-request rules being used to set at least one request post-processing operation of polling processing and streaming decoding after obtaining the initial dialogue request; The response type is obtained from the response parsing mechanism, the corresponding parsing logic is obtained according to the response type, and the standard response output format is obtained based on the parsing logic.
5. The method for scheduling a large model in a cross-domain computing center according to claim 4 is characterized in that: The obtaining of the standard response output format based on the parsing logic includes: If the parsing logic is streaming response parsing, the streaming protocol type is obtained from the response parsing mechanism, and parsing rules for field extraction and semantic reorganization are determined according to the streaming protocol type.
6. The method for scheduling a large model in a cross-domain computing center according to claim 1 is characterized in that: After sending the target dialogue request to the target large model for content generation and obtaining an inference result, the method includes: obtaining thinking process data and final answer data from the reasoning results; The final answer data is displayed through the unified external interface module, and lazy loading or content folding of the thinking process data is determined according to preset display rules.
7. The method for scheduling a large model in a cross-domain computing center according to claim 1 is characterized in that: The calculating real-time load scores of the connected multiple computing power centers and selecting at least one target computing power center from the multiple computing power centers according to the real-time load scores includes: Calculate the real-time load score based on the acquired task status information, resource load indicators, and platform capability labels of each computing power center; The computing power center with the real-time load score greater than or equal to the preset score is selected as a candidate computing power center, and the candidate score is calculated according to the adaptation level, interface stability parameter, and streaming reasoning capability index of each candidate computing power center. The candidate computing power center corresponding to the maximum candidate score is selected as the target computing power center; If each of the real-time load scores is less than the preset score, the preset score is reduced according to a preset step size until the target computing power center is selected.
8. A scheduling device for a large model in a cross-domain computing center, characterized in that: include: A dialog request acquisition module is used to acquire an initial dialog request from an asynchronous task queue. The initial dialog request is request data of a preset standard rule obtained by the unified external interface module in response to user input. Computing power scheduling module: used to calculate the real-time load scores of multiple connected computing power centers and select at least one target computing power center from the multiple computing power centers based on the real-time load scores. At least two of the computing power centers are located in a cross-domain resource environment. At least one large model is deployed on each computing power center. The cross-domain resource environment includes at least two of: cloud services, bare metal clusters, or third-party API gateways. Adaptation module: used to obtain corresponding adaptation parameter data according to the adaptation tag of the target computing power center, and convert the initial dialogue request into a target dialogue request corresponding to the target computing power center according to the adaptation parameter data; Content generation module: used to determine the target large model from the large models deployed in the target computing power center, send the target dialogue request to the target large model for content generation, and obtain the inference result.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the scheduling method of the large model in the cross-domain computing center described in any one of claims 1 to 7.
10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the scheduling method of the large model in a cross-domain computing center according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Method and system for unified scheduling of cross-domain computing power
CN122450688A