Large model service scheduling system and method, electronic equipment and storage medium

The customized prompt vector library is generated by the configured production module, combined with the deep Q learning algorithm and the dynamic scaling system, and efficient scheduling of large-model services is achieved, solving the problems of waste of resources and high operating costs in the existing technology, and improving system performance and operational efficiency.

CN120276858APending Publication Date: 2025-07-08CHINA MOBILE GROUP ZHEJIANG +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510396325.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, large model service scheduling is low efficiency, high cost, and complex maintenance, making it difficult to effectively integrate and optimize monitoring resources, resulting in waste of resources and increased operating costs.

Method used

The configuration production module is used to generate a customized prompt vector library of social governance. Through the unified computing network scheduling system, intelligently dispatch configuration files to idle nodes based on the deep Q learning algorithm, and dynamically adjust resources in combination with the unified dynamic expansion and capacity system to achieve rapid task response and reasonable resource allocation.

Benefits of technology

It improves resource utilization efficiency, reduces operating costs, ensures that the system operates efficiently in complex environments, and adapts to the task needs of different social governance scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276858A_ABST
    Figure CN120276858A_ABST
Patent Text Reader

Abstract

The invention provides a large model service scheduling system and method, electronic equipment and a storage medium, and relates to the technical field of data processing, and the system comprises a configuration production module which determines a configuration file corresponding to a user task vector in a social governance customized prompt vector library according to the user task vector; a unified computing network scheduling system in the configuration scheduling module schedules the configuration file to a computing node for processing; wherein the computing node is a node determined in each idle node through a scheduling model; the idle nodes are determined based on a unified dynamic capacity expansion and contraction system in the configuration scheduling module, and the scheduling model is obtained based on deep Q learning algorithm training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a large model service scheduling system, method, electronic device, and storage medium. Background Art

[0002] In today's society, the pace of digital transformation is accelerating day by day, and the application and innovation of video technology play a crucial role in it. The development of video technology not only greatly enriches the way of information dissemination, but also plays a core role in promoting the modernization transformation of social governance.

[0003] With the rapid increase in video data volume, how to efficiently integrate and optimize monitoring resources has become the key to improving governance capabilities. Therefore, how to more effectively schedule large model services has become an urgent problem in the industry. Summary of the Invention

[0004] The present invention provides a large model service scheduling system, method, electronic device, and storage medium to solve the problem of how to more effectively schedule large model services in the prior art.

[0005] The present invention provides a large model service scheduling system, including: A configuration-based production module, which determines a configuration file corresponding to the user task vector in a social governance customized prompt vector library according to the user task vector; A configuration-based scheduling module, in which a unified computing and network scheduling system in the configuration-based scheduling module schedules the configuration file to a computing node for processing; Wherein, the computing node is a node determined from among each idle node through a scheduling model; the idle node is determined based on a unified dynamic scaling system in the configuration-based scheduling module, and the scheduling model is trained based on a deep Q-learning algorithm.

[0006] According to a large model service scheduling system provided by the present invention, the configuration-based production module is further configured to: Generate a large model according to a preset prompt word, and configure a social governance customized prompt vector library containing a plurality of social governance customized prompt vectors; Wherein, each social governance customized prompt vector includes a task information and a task configuration file corresponding to the task information.

[0007] According to a large model service scheduling system provided by the present invention, the unified dynamic scaling system includes: A monitoring and acquisition module, and the acquisition monitoring system is used to collect real-time computing power network status data of each node, and synchronize it to an index prediction module and a scaling planning module according to the computing power network status data; The index prediction module uses the grey prediction model algorithm to predict the resource load based on the computing power network status data provided by the monitoring and acquisition module, and obtains the resource load prediction information; The scaling planning module determines the expansion or contraction of nodes according to the resource load prediction information, and determines the current idle nodes based on the current nodes after expansion or contraction.

[0008] According to a large model service scheduling system provided by the present invention, the training process of the scheduling model specifically includes: Initialize the Q-network to approximate the state-action value function, select actions using the ε-greedy policy, and store the experience pairs in the experience replay buffer; In the training stage, randomly sample experience pairs from the buffer, calculate the target value using the target Q-network, and construct a loss function to train the agent to optimize the task scheduling strategy; Stop training when the preset iteration condition is met to obtain the scheduling model.

[0009] According to a scheduling method based on any of the above large model service scheduling systems provided by the present invention, it includes: According to the user task vector, determine the configuration file corresponding to the user task vector in the social governance customized prompt vector library; The unified computing network scheduling system in the configuration scheduling module schedules the configuration file to the computing node for processing; Among them, the computing node is the node determined among each idle node through the scheduling model; the idle node is determined based on the unified dynamic scaling system in the configuration scheduling module, and the scheduling model is trained based on the deep Q-learning algorithm.

[0010] According to the scheduling method provided by the present invention, before the step of determining the configuration file corresponding to the user task vector in the social governance customized prompt vector library according to the user task vector, it further includes: Generate a large model according to the preset prompt words, and configure a social governance customized prompt vector library containing multiple social governance customized prompt vectors; Among them, each social governance customized prompt vector contains a task information and the task configuration file corresponding to the task information.

[0011] According to the scheduling method provided by the present invention, before the step of the unified computing network scheduling system in the configuration scheduling module scheduling the configuration file to the computing node for processing, the method further includes: Collect the computing power network status data of each node in real time; Using the grey prediction model algorithm, based on the computing power network state data, perform resource load prediction to obtain resource load prediction information; According to the resource load prediction information, determine node expansion or node contraction, and based on the current nodes after expansion or contraction, determine the current idle nodes.

[0012] According to the scheduling method provided by the present invention, the training process of the scheduling model specifically includes: By initializing the Q-network to approximate the state-action value function, select actions using the ε-greedy strategy, and store the experience pairs in the experience replay buffer; In the training phase, randomly sample experience pairs from the buffer, calculate the target value using the target Q-network, and construct a loss function to train the agent to optimize the task scheduling strategy; When the preset iteration condition is met, stop training to obtain the scheduling model.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the scheduling method described in any one of the above is implemented.

[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the scheduling method described in any one of the above is implemented.

[0015] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the scheduling method described in any one of the above is implemented.

[0016] The large model service scheduling system, method, electronic device, and storage medium provided by the present invention achieve task customization through a configuration-based production module, and match the configuration file according to the user task vector; the unified computing network scheduling system of the configuration-based scheduling module is based on the deep Q-learning algorithm, and intelligently schedules the configuration file to a suitable computing node; the unified dynamic scaling system dynamically adjusts resources to determine idle nodes. This architecture improves resource utilization efficiency, realizes fast response and processing of tasks, adapts to different social governance scenarios, and effectively improves system performance and operation efficiency. Through the collaborative work of the configuration-based production module and the configuration-based scheduling module, fast response of tasks and reasonable allocation of resources are realized. The unified dynamic scaling system and the scheduling model based on deep Q-learning ensure that the system can dynamically adapt to load changes and maintain efficient operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 Schematic diagram of the customized fine-tuning method provided by the related art; Figure 2 Schematic diagram of the main process of the small model mirroring scheduling mode in the related art; Figure 3 Schematic diagram of the structure of the large model service scheduling system provided by the present invention; Figure 4 Schematic diagram of the comparison of the scheduling modes provided by the present invention; Figure 5 Schematic diagram of the process of the scheduling method provided by the present invention; Figure 6 Schematic diagram of the structure of the electronic device provided by the present invention. Detailed implementation manners

[0019] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0020] In the related art, in order to further improve the utilization efficiency of monitoring resources, Figure 1 Schematic diagram of the customized fine-tuning method provided by the related art, as Figure 1 shown, mainly adopts the customized fine-tuning method of small models and combines the scheduling of the computing power network for mirroring deployment.

[0021] Figure 2 Schematic diagram of the main process of the small model mirroring scheduling mode in the related art, as Figure 2 shown, step 1: Customized training and fine-tuning. First, developers use the business data provided by users to customize the training and fine-tuning of existing models to obtain small dedicated models for specific tasks.

[0022] Step 2: Customized pre- and post-processing. To make the model better adapt to the business scenario, developers will carry out customized pre- and post-processing development on the input and output of the model to ensure that the model can receive input data in the required format and produce correct output results.

[0023] Step 3: Inference mirror encapsulation. Integrate the small model obtained through customized training and the relevant pre- and post-processing logics into an inference process, and use mirror technology for encapsulation to ensure the consistency of the software environment required during deployment.

[0024] Step 4: Computing power resource scheduling. According to the specific needs of users, the computing power network intelligently schedules appropriate computing power resources from the idle resource pool to meet the computing power requirements of different tasks.

[0025] Step 5: Service mirror transmission. Send the encapsulated service mirror to a specific computing power cluster through a dedicated data transmission line to ensure the rapid deployment and stability of the service.

[0026] Step 6: Cluster deployment of the service. Deploy the service mirror on the computing power cluster, and utilize the computing power of the cluster to run the inference process to provide users with efficient and reliable services.

[0027] However, the solutions of the traditional small model production scheduling mode have the following disadvantages: 1. High customization cost: In the traditional small model production scheduling mode, due to the diversity and fragmentation of visual business requirements, the user requirements in each scenario may be completely different, resulting in a significant long-tail phenomenon. In this case, due to their limited generalization ability and low reusability, traditional models often need to collect and label a large amount of data for customized training for each specific scenario. This customization process is not only time-consuming and labor-intensive but also costly, because changes in each new task or scenario may mean starting from scratch with model training and adjustment work.

[0028] 2. High computing power cost: In the traditional small model scheduling mode, due to the large variety of models and the lack of a unified scheduling mechanism, algorithms often cannot effectively share and schedule computing power resources. This leads to each model possibly monopolizing and binding specific computing power, causing waste of resources. In addition, small models usually need to be encapsulated and transmitted in the form of mirrors. This transmission method is not only inefficient but also, as the number of models increases, the cost of their transmission and storage rises significantly, especially in scenarios where multiple models need to be updated and deployed frequently.

[0029] 3. High maintenance cost: In the traditional small model production scheduling mode, the deployment of models is usually siloed, that is, each type of task requires the deployment of a corresponding type of model. This means that as business requirements increase, the number of models that need to be deployed and maintained will also increase accordingly, and each model requires separate attention and upgrading. This decentralized deployment method not only increases the complexity of operation and maintenance but also raises the maintenance cost, because the update, repair, and optimization of each model need to be carried out separately and cannot achieve unified management and batch processing.

[0030] Figure 3 It is a schematic structural diagram of the large model service scheduling system provided by the present invention. As Figure 3 shown, it includes: Configuration production module 310. The configuration production module determines the configuration file corresponding to the user task vector in the social governance customization prompt vector library according to the user task vector; In the present invention, in the social governance scenario, the user generates a customized prompt vector according to specific tasks (such as fire detection, car accident detection, etc.). For example, the prompt vector for the fire detection task may contain key information such as "smoke characteristics" and "flame characteristics", which is formed through techniques such as prompt engineering or prompt fine-tuning.

[0031] The social governance customization prompt vector library stores a library of various social governance task vectors and their corresponding configuration files. Each task vector is associated with a specific social governance task, and the vectors in the library are organized by category for quick retrieval and matching.

[0032] The configuration production module finds the most matching task vector by comparing the user task vector with the vectors in the library and obtains its corresponding configuration file. The configuration file contains specific parameters and process definitions required for task execution, such as model input and output configurations, post-processing processes, etc.

[0033] Configuration scheduling module 320. The unified computing and network scheduling system in the configuration scheduling module schedules the configuration file to the computing node for processing; Among them, the computing node is determined among each idle node through a scheduling model; the idle node is determined based on the unified dynamic scaling system in the configuration scheduling module, and the scheduling model is trained based on the deep Q-learning algorithm.

[0034] In the present invention, the unified computing and network scheduling system in the configuration scheduling module 320 is responsible for efficiently scheduling the configuration file generated by the configuration production module to the computing node for processing. The system first receives the configuration file from the configuration production module, which contains information such as task vectors and post-processing processes. Then, the system senses the available nodes in the computing power network, and these nodes are generated after the unified dynamic scaling system performs expansion or contraction operations according to the load prediction results. The unified dynamic scaling system will dynamically adjust the number of microservice instances according to the system load situation to determine the idle nodes. The idle node refers to a node that has not been assigned a task and has available computing resources currently.

[0035] When the unified computing and networking scheduling system schedules the configuration file, it will select the most suitable computing node based on a scheduling model trained by the deep Q-learning algorithm. This model can intelligently select the optimal computing node for task allocation according to the status of idle nodes and task requirements in the current network. In this way, the system can achieve optimal utilization of resources and rapid response to tasks, while ensuring the rationality and efficiency of task allocation. In addition, the scheduling model based on deep Q-learning can also adaptively adjust the scheduling strategy according to changes in the system state, adapt to complex and changeable operating environments, and improve the intelligence and efficiency of task scheduling. This design not only improves the overall performance of the system, but also effectively controls operating costs and avoids waste of resources.

[0036] First, the corresponding prompt vector of the task is obtained through prompt customization, and then the prompt vector is stored in the constructed prompt vector library according to the category. When the user has the need to start the service, the service can be customized by configuring the specific task vector and post-processing process, etc. through the customized request file, and the task vector can be obtained from the prompt vector library.

[0037] First, we deploy the multi-modal large model centrally on all nodes. Then, the unified dynamic scaling system will predict the load metrics of the system by analyzing the monitoring data of the node cluster and the configuration files uploaded by users.

[0038] The metric prediction module uses a grey prediction model to predict the future load change behavior of the system through a small amount of incomplete information. After prediction, we scale the nodes of the service in steps to ensure the reasonable allocation and use of resources. After scaling, idle nodes are obtained, including computing nodes and redundant backup nodes. The computing nodes are scheduled by the node scheduler and serve as the end point of the computing and networking scheduling system. The configuration file passed in by the configuration production module is scheduled to the computing nodes by the computing and networking scheduling system to generate customized services to respond to tasks. Among them, the computing and networking scheduling system adopts a scheduling strategy for unified scheduling of configuration files trained by the deep Q-learning method.

[0039] Optionally, the configuration production module is further configured to: Generate a large model according to a preset prompt word, and configure a social governance customized prompt vector library containing multiple social governance customized prompt vectors; Wherein, each of the social governance customized prompt vectors includes a task information and a task configuration file corresponding to the task information.

[0040] In the present invention, the configuration-based production module first generates a general large model according to a preset prompt. This process involves the processing of the prompt and model training, enabling the generated large model to adapt to the requirements of various social governance tasks. Next, the module constructs a prompt vector library containing multiple social governance customized prompt vectors. Each prompt vector is associated with a specific social governance task and includes a task information and a task configuration file corresponding to the task information. The task information describes the specific content and objectives of the task, while the task configuration file contains all the parameters and process definitions required to execute the task, such as the format of model input and output, post-processing steps, etc.

[0041] In practical applications, when users have a new social governance task, they generate a user task vector according to the task requirements. The configuration-based production module matches the user task vector with the vectors in the social governance customized prompt vector library to find the most similar or matching vector. Once a matching vector is found, the module obtains its corresponding task configuration file and passes it to the next processing link of the system, such as the configuration-based scheduling module, for further task scheduling and execution.

[0042] More specifically, the large model executes different tasks based on the input prompt. Users customize the optimal prompt for the task through means such as prompt engineering and prompt fine-tuning to form a task vector. The mathematical representation form of the task vector is a matrix R∈N∗M, where N is the number of tokens of the task vector required for a specific task, and the dimension of the token is M. The trained prompt vectors are stored in the prompt vector library by category. During inference, users combine the image prompt vectors corresponding to the required categories from the vector database to complete the detection task without a text encoder.

[0043] The user configuration request file includes two parts: a data field and a customized task field. The data field includes a control threshold and a post-processing orchestration process dictionary, which are directly transmitted to the computing power node through the computing network during actual application, act on the large model, and generate a post-processing process. The customized task field is the category prompt vector corresponding to the task, which is input into the large model together with the detected image input by the user during actual application for accurate positioning. Each component of the configuration file is converted into a json serializable structure, and the task vector is transmitted as a matrix file encoded into a base64 string through base64 encoding.

[0044] Based on the property that the vision multi-modal model can output specific task results through a guiding vector, the customized task vector is transmitted through the vector library to control the output of the unified deployed vision multi-modal model. At the same time, a post-processing orchestration system is established, and users configure the post-processing process field in the configuration file to build the task post-processing process.

[0045] In the present invention, the configuration-based production module provides a high degree of customization and flexibility for the system, enabling the system to quickly respond to and process various social governance tasks while ensuring the accuracy and efficiency of task execution.

[0046] Optionally, the unified dynamic scaling system includes: A monitoring and collection module. The collection and monitoring system is used to collect the computing power network status data of each node in real time and synchronize it to the index prediction module and the scaling planning module according to the computing power network status data; The index prediction module uses the grey prediction model algorithm to perform resource load prediction based on the computing power network status data provided by the monitoring and collection module, and obtains resource load prediction information; The scaling planning module determines the expansion or contraction of nodes according to the resource load prediction information, and determines the current idle nodes according to the current nodes after expansion or contraction.

[0047] The dynamic scaling system is one of the core modules in the collaborative scheduling of the computing power network and plays a crucial role in the operation of the computing network application and the processing of user services. This system realizes the dynamic scaling process of microservices, that is, increasing or decreasing microservice instances, through three parts: the monitoring and collection module, the index prediction module, and the scaling planning module, so as to achieve flexible and elastic scheduling of computing power resources.

[0048] The monitoring and collection module is responsible for real-time monitoring of the status of the entire computing power network. It mainly considers two types of resource metrics: system resource level metrics and custom resource metrics. Custom metrics refer to the "business" metric types introduced by third-party monitoring. This module aggregates the metrics to the Aggregator through the metric adapter, and the Aggregator provides the required metrics to the index prediction module and the scaling planning module. The overall load situation of the system is defined as where C, M, D, and B represent the CPU usage, memory usage, disk usage, and network bandwidth usage respectively, and wC, wM, wD, and wB are their weights, satisfying On the one hand, the index prediction module captures the historical resource load metric data of the microservice replica set monitored by the monitoring and collection module, and uses these historical monitoring data and the prediction model to predict the resource load value of the microservice replica set at the next moment; on the other hand, it outputs the resource load prediction value to the scaling planning module to guide the subsequent elastic scaling of microservices. The input of this module is a set of time series data composed of the overall load rates in the previous period of time, and the output is short-term future time series data.

[0049] Prediction sequence = Grey prediction model (historical load data) When using the grey prediction model to predict the system load of computing-network applications' microservices, the result obtained by the metric prediction module is a time series, which represents the possible resource usage in the future for a period of time. The maximum value in the prediction sequence represents the maximum possible resource usage of the microservices within the next n seconds, and is passed as the main metric to the scaling planning module for the next step of planning.

[0050] The grey prediction model adopted is the GM(1,1) model, which is a single-variable first-order differential equation model. The basic form of the model is as follows: where \(x(t)\) is the system state variable (such as resource load). \(a\) and \(b\) are model parameters. \(t\) is time.

[0051] The scaling planning module is the part of the computing power network dynamic scaling system responsible for making scaling decisions. It decides whether to scale out or scale in microservice instances based on the prediction data provided by the metric prediction module to ensure the reasonable allocation and use of resources.

[0052] New replica count = max[(replica count)×(scaling step), maximum replica count] New replica count = min[(replica count)×(scaling step), minimum replica count] Optionally, the training process of the scheduling model specifically includes: Initializing the Q-network to approximate the state-action value function, selecting actions using the ε-greedy policy, and storing the experience pairs in the experience replay buffer; During the training phase, randomly sampling experience pairs from the buffer, calculating the target value using the target Q-network, and constructing a loss function to train the agent to optimize the task scheduling strategy; Stopping the training when the preset iteration condition is met to obtain the scheduling model.

[0053] In the present invention, the configuration file scheduling problem is reformulated as a reinforcement learning (RL) problem, and the deep Q-learning method is applied to train an RL policy for unified configuration file scheduling.

[0054] Step 1: Define the state space: When considering the nth task Tn being ready for scheduling, the state of the RL agent is a vector with m + 1 coordinates. The first m coordinates record the current total workloads of the m worker queues (each coordinate records the workload of one queue). The last coordinate of the state vector records the total workload of the task to be scheduled. This state information is queried from the system whenever a new task is ready for scheduling and logged to reduce the scale of large values. In particular, this information is directly related to the workload balance assigned to the worker threads and the data transfer cost, two key factors that affect the overall system performance.

[0055] Step 2: Define the action space: Since there are m possible actions, the task Tn will be assigned to one of the m worker threads. The policy is specified based on the so-called state-action value table Q(s,a), which evaluates the expected return of taking action a in state s.

[0056] Step 3: Define the reward function: After each task is assigned to a worker thread, the state of the system transitions, and the agent receives a reward signal, which is designed based on two system performance-related characteristics: the balance of the total workload assigned to the worker threads and the data transfer cost caused by scheduling the task Tn.

[0057] Mathematically, the reward signal is defined as follows: Step 4: Select a reinforcement learning algorithm: This paper adopts the deep Q-learning method, aiming to learn the optimal policy to maximize the expected cumulative reward.

[0058] This problem is formulated as follows: where α is a preselected discount factor. In particular, given the state st at time t, the policy generates the action at according to the state-action value function Q as follows: Intuitively, Q(s, a) evaluates the expected return of taking action a in state s, and the policy π simply suggests the action that leads to the highest Q value in a given state. As is well known, the Q function satisfies the following Bellman equation: Step 5: Train the agent: The deep Q-learning algorithm approximates the state-action value function by initializing a Q-network and selects actions using the ε-greedy policy at each time step, while storing the experience pairs in the experience replay buffer. During the training phase, the algorithm randomly samples a batch of experience pairs from the buffer and calculates the target value using the target Q-network And construct the loss function In addition, at regular time steps, the parameters of the target Q-network are copied from the Q-network to ensure the stability of the target network. In this way, the algorithm can learn which actions to take in a given state to obtain the maximum cumulative reward Q(st, at), thereby optimizing the task scheduling strategy. This method is particularly suitable for problems with large state spaces and action spaces, and can effectively allocate tasks to the most suitable computing nodes, improving the overall computing efficiency and resource utilization rate.

[0059] Figure 4 Schematic diagram for comparing scheduling modes provided by the present invention, as Figure 4 shown, this architecture adopts a centralized deployment of large models, and the models deployed on each node are unified. We only need to schedule the configuration file to the under-loaded node through the computing network scheduling system, instead of, like the traditional mode, still having to perform a shaft-like scheduling and deployment of the proprietary model to the corresponding proprietary node through the application demand matching module, historical policy matching module, etc.

[0060] Figure 5 Schematic diagram of the flow of the scheduling method provided by the present invention, as Figure 5 shown, including: Step 510, according to the user task vector, determine the configuration file corresponding to the user task vector in the social governance customized prompt vector library; Step 520, the unified computing network scheduling system in the configuration scheduling module schedules the configuration file to the computing node for processing; wherein, the computing node is determined from each idle node through a scheduling model; the idle node is determined based on the unified dynamic scaling system in the configuration scheduling module, and the scheduling model is trained based on the deep Q-learning algorithm.

[0061] In the present invention, in step 510, the system determines the corresponding configuration file in the social governance customized prompt vector library according to the user task vector. The user task vector is generated by the user according to specific social governance tasks and contains the key information of the tasks. The social governance customized prompt vector library stores various task vectors and their corresponding configuration files. The configuration file contains the specific parameters and process definitions required for task execution.

[0062] In step 520, the unified computing and network scheduling system in the configuration-based scheduling module schedules the configuration file to the computing node for processing. The configuration-based scheduling module is responsible for efficiently scheduling the configuration file to the computing node. The unified computing and network scheduling system is the core component and is responsible for the specific scheduling operation. The computing node is determined from among various idle nodes through the scheduling model, and the idle nodes are determined based on the unified dynamic scaling system. The scheduling model is trained based on the deep Q-learning algorithm and can intelligently select the optimal computing node.

[0063] The entire process realizes the rapid matching and scheduling from the user task vector to the configuration file, ensuring that the task is efficiently executed on the most suitable computing node, while optimizing resource utilization and system performance.

[0064] In the present invention, task customization is achieved through the configuration-based production module, and the configuration file is matched according to the user task vector; the unified computing and network scheduling system of the configuration-based scheduling module, based on the deep Q-learning algorithm, intelligently schedules the configuration file to the appropriate computing node; the unified dynamic scaling system dynamically adjusts resources to determine the idle nodes. This architecture improves resource utilization efficiency, realizes rapid response and processing of tasks, adapts to different social governance scenarios, and effectively improves system performance and operation efficiency. Through the collaborative work of the configuration-based production module and the configuration-based scheduling module, rapid response of tasks and reasonable allocation of resources are achieved. The unified dynamic scaling system and the scheduling model based on deep Q-learning ensure that the system can dynamically adapt to load changes and maintain efficient operation.

[0065] Optionally, before the step in which the unified computing and network scheduling system in the configuration-based scheduling module schedules the configuration file to the computing node for processing, the method further includes: Real-time collection of the computing power network status data of each node; Using the grey prediction model algorithm, based on the computing power network status data, performing resource load prediction to obtain resource load prediction information; According to the resource load prediction information, determining node expansion or node contraction, and based on the current nodes after expansion or contraction, determining the current idle nodes.

[0066] In the present invention, before the unified computing and network scheduling system in the configuration-based scheduling module schedules the configuration file to the computing node for processing, the system will real-time collect the computing power network status data of each node, and use the grey prediction model algorithm to perform resource load prediction based on the collected data to obtain resource load prediction information.

[0067] Then, according to the prediction information, determine whether to perform node expansion or contraction to adjust the current number of nodes, thereby determining the current idle nodes.

[0068] More specifically, first, the monitoring and acquisition module collects the computing power network status data of each node in real time. These data include, but are not limited to, key metrics such as CPU usage, memory occupancy, and network bandwidth. Then, the metric prediction module uses the grey prediction model algorithm to analyze and process the collected data, predict the future resource load situation, and generate resource load prediction information.

[0069] Based on this prediction information, the scaling planning module determines whether the resources of the current system are sufficient to handle future load changes. If resource tension is predicted, the system will expand the nodes, adding new nodes to share the load; conversely, if resource surplus is predicted, the system will scale down the nodes, releasing the redundant node resources to optimize resource utilization.

[0070] In the present invention, the determined current idle node information will provide an important reference basis for the subsequent configuration file scheduling step, ensuring that the configuration file can be efficiently scheduled to the most suitable computing node for processing.

[0071] Figure 6 is a schematic structural diagram of the electronic device provided by the present invention. As Figure 6 shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 complete mutual communication through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the scheduling method, which includes: determining the configuration file corresponding to the user task vector in the social governance customization prompt vector library according to the user task vector; The unified computing network scheduling system in the configuration scheduling module schedules the configuration file to the computing node for processing; wherein, the computing node is determined among each idle node through a scheduling model; the idle node is determined based on the unified dynamic scaling system in the configuration scheduling module, and the scheduling model is trained based on the deep Q-learning algorithm.

[0072] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0073] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the scheduling method provided by the above-mentioned various methods. The method includes: determining a configuration file corresponding to the user task vector in a social governance customized prompt vector library according to the user task vector; The unified computing and networking scheduling system in the configuration scheduling module schedules the configuration file to a computing node for processing; Wherein, the computing node is a node determined from among various idle nodes through a scheduling model; the idle nodes are determined based on the unified dynamic scaling system in the configuration scheduling module, and the scheduling model is trained based on a deep Q-learning algorithm.

[0074] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the scheduling method provided by the above-mentioned various methods. The method includes: determining a configuration file corresponding to the user task vector in a social governance customized prompt vector library according to the user task vector; The unified computing and networking scheduling system in the configuration scheduling module schedules the configuration file to a computing node for processing; Wherein, the computing node is a node determined from among various idle nodes through a scheduling model; the idle nodes are determined based on the unified dynamic scaling system in the configuration scheduling module, and the scheduling model is trained based on a deep Q-learning algorithm.

[0075] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0076] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A large model service scheduling system, characterized in that, It includes: A configuration-based production module that, according to the user task vector, determines the configuration file corresponding to the user task vector in the social governance customization prompt vector library; A configuration-based scheduling module, where the unified computing and network scheduling system in the configuration-based scheduling module schedules the configuration file to a computing node for processing; Among them, the computing node is determined from each idle node through a scheduling model; the idle node is determined based on the unified dynamic scaling system in the configuration-based scheduling module, and the scheduling model is trained based on the deep Q-learning algorithm.

2. The large model service scheduling system according to claim 1, wherein The configuration-based production module is further used for: Generating a large model according to preset prompt words and configuring a social governance customization prompt vector library containing multiple social governance customization prompt vectors; Among them, each social governance customization prompt vector contains a task information and the task configuration file corresponding to the task information.

3. The large model service scheduling system according to claim 1, wherein The unified dynamic scaling system includes: A monitoring and acquisition module, where the acquisition and monitoring system is used to collect the computing power network status data of each node in real time and synchronize it to the index prediction module and the scaling planning module according to the computing power network status data; The index prediction module uses the grey prediction model algorithm to perform resource load prediction based on the computing power network status data provided by the monitoring and acquisition module to obtain resource load prediction information; The scaling planning module determines the expansion or contraction of nodes according to the resource load prediction information, and determines the current idle nodes according to the current nodes after expansion or contraction.

4. The large model service scheduling system according to claim 1, wherein The training process of the scheduling model specifically includes: Initializing the Q-network to approximate the state-action value function, selecting actions using the ε-greedy strategy, and storing the experience pairs in the experience replay buffer; In the training stage, randomly sampling experience pairs from the buffer, calculating the target value using the target Q-network, and constructing a loss function to train the agent to optimize the task scheduling strategy; When the preset iteration condition is met, stop training to obtain the scheduling model.

5. The scheduling method of the large model service scheduling system according to any one of claims 1-4, characterized in that, It includes: According to the user task vector, determining the configuration file corresponding to the user task vector in the social governance customization prompt vector library; The unified computing and network scheduling system in the configuration-based scheduling module schedules the configuration file to a computing node for processing; Among them, the computing node is determined from each idle node through a scheduling model; the idle node is determined based on the unified dynamic scaling system in the configuration-based scheduling module, and the scheduling model is trained based on the deep Q-learning algorithm.

6. The scheduling method according to claim 5, wherein, Before the step of determining the configuration file corresponding to the user task vector according to the user task vector in the social governance customization prompt vector library, it further includes: Generating a large model according to preset prompt words and configuring a social governance customization prompt vector library containing multiple social governance customization prompt vectors; Among them, each social governance customization prompt vector contains a task information and the task configuration file corresponding to the task information.

7. The scheduling method according to claim 5, wherein Before the step in which the unified computing and networking scheduling system in the configured scheduling module schedules the configuration file to a computing node for processing, the method further includes: Collecting in real time the computing power network status data of each node; Using the grey prediction model algorithm to perform resource load prediction based on the computing power network status data to obtain resource load prediction information; Determining node expansion or node contraction according to the resource load prediction information, and determining the current idle nodes based on the current nodes after expansion or contraction.

8. The scheduling method according to claim 5, wherein The training process of the scheduling model specifically includes: Initializing the Q-network to approximate the state-action value function, selecting actions using the ε-greedy policy, and storing the experience pairs in the experience replay buffer; In the training phase, randomly sampling experience pairs from the buffer, calculating the target value using the target Q-network, and constructing a loss function to train the agent to optimize the task scheduling strategy; Stopping the training to obtain the scheduling model when the preset iteration condition is met.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the scheduling method according to any one of claims 5 to 8.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the scheduling method according to any one of claims 5 to 8.