A method for large model cross-domain computing power scheduling

CN121764666BActive Publication Date: 2026-09-25SAIWUZHOU
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511916144.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-09-25
Estimated Expiration
2045-12-18

AI Technical Summary

Technical Problem

当二者进行跨域算力调度时,由于不同域概念的理解不同,调度系统难以准确把握各个域的具体算力需求,无法进行精准的跨域算力调度

Benefits of technology

本发明公开了一种大模型跨域算力调度的方法。该方法首先对不同领域的算力需求描述进行语义分析,提取关键信息形成结构化表示,并采用统一的描述语言实现标准化表达。然后制定跨域调度交流协议,通过可扩展的数据交换机制解决接口差异,实时获取各域资源使用情况。在此基础上,建立算力需求优先级评估机制,从业务重要性、紧急度等维度进行评分,并根据评分结果生成跨域调度方案。本发明还建立了动态调整和监控评估机制,通过与各域交互反馈优化调度效果。该方法实现了跨域算力需求的精准分析、优先级排序和资源匹配,提高了算力资源利用效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121764666B_ABST
    Figure CN121764666B_ABST
Patent Text Reader

Abstract

The application provides a method for large model cross-domain computing power scheduling, comprising the following steps: according to a structured computing power demand representation, using a unified computing power demand description language, defining various attributes and parameters of the computing power demand, forming a standardized computing power demand description method, and realizing consistent expression of the computing power demand between different domains; in a scheduling system, a comparison and priority determination mechanism of the computing power demand is established, the computing power demands of different domains are prioritized according to the standardized computing power demand description, and the urgency and importance of each demand are determined; according to the priority of the computing power demand, combined with the computing power resource situation of each domain, a cross-domain computing power scheduling scheme is generated, the computing power allocation scheme and the scheduling timetable of each domain are determined, so as to realize accurate matching of the computing power demand and the computing power resource.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology, and in particular to a method for cross-domain computing power scheduling of large models. Background Technology

[0002] In large-scale cross-domain computing power scheduling, the different ways of expressing computing power requirements across domains, coupled with the lack of a unified language for describing computing power requirements and standardized communication protocols, pose significant challenges. This is primarily due to the specific applications and understandings of technical terms and concepts in different fields, leading to a lack of consensus on computing power requirements across domains. For example, in computer science and information technology, terms such as "number of CPU cores," "number of GPUs," "memory size," and "storage capacity" are commonly used to describe computing power. However, in particle physics or cosmology simulations, "floating-point operations per second" might be used to measure computing power. When these two domains engage in cross-domain computing power scheduling, the differing understandings of these concepts make it difficult for the scheduling system to accurately grasp the specific computing power requirements of each domain, hindering precise cross-domain scheduling. Furthermore, the lack of standardized communication protocols makes effective communication and negotiation of computing power requirements between different domains difficult. Moreover, the absence of standardized protocols can lead to misunderstandings and errors in communication between different domains, further impacting scheduling effectiveness. In summary, the technical challenge of cross-domain computing power scheduling for large models lies in how to accurately understand and compare computing power requirements when different domains use different languages ​​and protocols to describe these requirements, and then make precise scheduling decisions based on these requirements. Summary of the Invention

[0003] This invention provides a method for cross-domain computing power scheduling of large models, mainly including: Obtain computing power demand description texts from different fields, perform semantic analysis on the computing power demand description texts, and generate structured computing power demand representations; Based on the structured representation of computing power requirements, a unified computing power requirement description language is adopted to form a standardized computing power requirement description. Based on the standardized computing power requirement description, a communication protocol for cross-domain computing power scheduling of large models is formulated, and computing power resource information of each domain is obtained through data exchange and adaptation mechanisms. Based on the standardized computing power requirement description and the computing power resource information, a priority determination mechanism for computing power requirements is established to generate a scheme for cross-domain computing power scheduling of large models. The scheme for cross-domain computing power scheduling of the large model is dynamically adjusted through monitoring and feedback mechanisms to track execution and optimize scheduling effects.

[0004] Furthermore, the step of performing semantic analysis on the computational power demand description text to generate a structured representation of the computational power demand includes: The text describing the computing power requirement was segmented using a word segmentation tool to obtain the segmentation results. High-frequency words are extracted from the word segmentation results based on a preset word frequency threshold. The high-frequency words are then labeled with part-of-speech tags to identify keyword types and obtain a preliminary set of semantic units. The semantic unit set is semantically annotated according to a pre-built multi-domain computing power demand dictionary to identify named entities such as computing power indicators, application scenarios and data scale. Construct a semantic dependency graph, extract the subject-verb-object structure, and determine the core attributes of computing power requirements; The semantic dependency graph is matched with a predefined library of computing power requirement templates, and the template with the smallest edit distance is selected to generate a structured representation of computing power requirements that includes computing power metrics, application scenarios, and data scale.

[0005] Furthermore, the adoption of a unified computing power requirement description language forms a standardized computing power requirement description, including: Key attributes and parameters are extracted from the structured representation of computing power requirements, including computing power, storage capacity, and network bandwidth. Based on the key attributes and parameters, the data type, value range, and unit are determined, resulting in the basic element set of the computing power demand description language. Obtain class hierarchy, attribute features, and constraints using ontology modeling tools; Based on the class hierarchy, attribute features, and constraints, a tag structure and syntax rules are designed. By converting templates, the computing power requirements of different domains are transformed into a standardized format, thereby realizing attribute correspondence and semantic equivalence processing.

[0006] Furthermore, the aforementioned communication protocol for cross-domain computing power scheduling of large models, which obtains computing resource information from each domain through data exchange and adaptation mechanisms, includes: Based on the protocol definition of cross-domain computing power scheduling in the large model, a cross-domain communication network is constructed using a distributed hash table, and the hash value is obtained as the domain node identifier. For cross-domain communication, rate limiting and priority management are set up, and two-way authentication and encrypted data transmission are achieved through encryption protocols; Develop a negotiation process engine that determines the current state based on the state transition matrix and drives the negotiation process based on the state transition results. A lightweight data acquisition agent is generated through a plug-in architecture-based data exchange adapter. A publish-subscribe model is used to transfer data between the agent and the scheduling system, and the acquired data is cleaned, transformed, and standardized.

[0007] Furthermore, the mechanism for prioritizing computing power requirements includes: Key attributes are extracted from the standardized computing power requirement description to construct a multi-dimensional feature vector; The multidimensional feature vector is processed by a pre-established dimensionality reduction model to obtain the dimensionality-reduced feature vector; Based on the reduced feature vectors, rule reasoning is performed through a preset rule engine to determine the comprehensive score for each computing power requirement; A priority queue is used to manage computing power demand. If a new computing power demand is received, the new computing power demand is inserted into the priority queue. For the computing power demand in the priority queue, demand aggregation is performed within a preset time window, and computing power demands of the same type are merged. The priority score of each cluster is calculated based on the clustering results.

[0008] Furthermore, the step of generating a cross-domain computing power scheduling scheme for large models based on the standardized computing power demand description and the computing power resource information includes: A tree structure is used to represent the computing resources of each domain, and the resource type, capacity and utilization information are obtained based on the tree structure; The computing power demand priority and resource constraints are transformed into constraints, and variables are selected through constraint propagation. Different priority time windows are set according to the urgency of the demand, and scheduling is carried out within the time window using a local optimum strategy. If the resource utilization rate exceeds the preset threshold range, task migration in or out is triggered, and task allocation is performed through a weighted round-robin algorithm to generate computing power allocation schemes and scheduling schedules for each domain.

[0009] Furthermore, the scheme for dynamically adjusting the cross-domain computing power scheduling of the large model through monitoring and feedback mechanisms includes: Receive resource status information and task execution information pushed by various domains, and update the monitoring indicator database, which includes resource utilization and task completion time; Calculate the statistical values ​​of the indicators in the monitoring indicator database. If the statistical values ​​of the indicators trigger a preset rule, then execute the scheduling strategy adjustment. Based on past monitoring data, predict future resource needs and update scheduling strategy parameters; Continuously implement monitoring and forecasting steps, dynamically adjust resource allocation and task scheduling schemes, and optimize cross-domain scheduling performance.

[0010] Furthermore, the establishment of a priority determination mechanism for computing power requirements and the generation of a scheme for cross-domain computing power scheduling of large models include: Obtain multi-dimensional evaluation indicators, construct a judgment matrix, and calculate the weights of each dimension using the eigenvector method; Based on the fuzzy evaluation set and the weights, a comprehensive evaluation result is obtained using the weighted average method. Receive demand change events; if the change in demand quantity or priority exceeds a preset threshold, trigger the scheduling system response process. Newly submitted critical requirements are assigned an initial priority. The priority of existing tasks is dynamically adjusted through a recursive adjustment method. If the recursion depth exceeds the preset number of levels, the adjustment is stopped.

[0011] Furthermore, schemes for cross-domain computing power scheduling of large-scale models include: A tree structure is used to represent the computing resources of each domain, where the root node is the total resource pool, the child nodes are the resources of each domain, and the leaf nodes are the specific resource items. Retrieve resource type, capacity, and utilization attribute information based on the tree structure; Transform computing power demand priorities, resource constraints, and task dependencies into constraints; Constraint propagation is performed based on the constraints. If the constraint propagation is complete, then the minimum residual value heuristic is used to select variables; Set three time windows based on the urgency of the need: urgent, normal, and low priority. Within the time window, a shortest job priority strategy is used for locally optimal scheduling. If the time window slides, then local optimal scheduling is performed again; Obtain resource utilization rate; if the resource utilization rate is lower than a preset lower threshold, trigger task migration. If the resource utilization rate is higher than the preset upper limit threshold, the task migration will be triggered. Based on the results of task migration in or out, a weighted round-robin algorithm is used for task allocation.

[0012] Furthermore, in the process of generating a scheme for cross-domain computing power scheduling of large models, resource status information and task execution information carrying domain identifiers are received, and the resource status information and task execution information are pushed by each domain through a message queue. The monitoring indicator database is updated based on resource status information and task execution information. The monitoring indicator database includes multi-dimensional indicators such as resource utilization, task completion time, and scheduling delay. Predict future resource needs and task execution based on multi-dimensional indicators and past monitoring data; The scheduling strategy parameters are updated based on the prediction results. These parameters are used to optimize resource allocation and task scheduling to improve the effectiveness of cross-domain resource scheduling.

[0013] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: This invention discloses a method for cross-domain computing power scheduling of large-scale models. The method first performs semantic analysis on the computing power demand descriptions of different domains, extracting key information to form a structured representation, and uses a unified description language for standardized expression. Then, a cross-domain scheduling communication protocol is established, resolving interface differences through a scalable data exchange mechanism and obtaining real-time resource usage information for each domain. Based on this, a computing power demand priority evaluation mechanism is established, scoring based on dimensions such as business importance and urgency, and generating a cross-domain scheduling plan based on the scoring results. This invention also establishes a dynamic adjustment and monitoring evaluation mechanism, optimizing scheduling effectiveness through interactive feedback with each domain. This method achieves accurate analysis, priority ranking, and resource matching of cross-domain computing power demands, improving the efficiency of computing power resource utilization. Attached Figure Description

[0014] Figure 1 This is a flowchart of a method for cross-domain computing power scheduling of a large model according to the present invention.

[0015] Figure 2 This is a schematic diagram of a method for cross-domain computing power scheduling of a large model according to the present invention.

[0016] Figure 3 This is another schematic diagram of a method for cross-domain computing power scheduling of a large model according to the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0018] like Figures 1-3 This embodiment of a method for cross-domain computing power scheduling of large models may specifically include: Step S101: Based on the computing power demand descriptions provided by different domains, natural language processing methods are used to perform semantic analysis on the computing power demand descriptions, extract key information, and obtain a structured representation of computing power demand to eliminate ambiguity caused by different languages ​​and terms.

[0019] The process involves acquiring descriptions of computing power requirements from different domains, performing word segmentation and part-of-speech tagging on these descriptions, identifying keywords and phrases from the segmentation and tagging results to obtain a preliminary set of semantic units, processing polysemous words, and mapping these polysemous words to a unified concept representation. A semantic dependency graph is constructed using the Stanford Dependency Parser on the unified concept representation, and optimized according to preset semantic rules. The optimized semantic dependency graph is then matched with a predefined standardized computing power requirement template to obtain matching results. Core attributes of the computing power requirements are extracted from the matching results, and a preset computing power resource library is queried based on these core attributes to select the computing power resource combination with the highest matching degree. If multiple computing power resource combinations meet the requirements, they are sorted according to predefined priority rules, and the optimal sorted computing power resource combination is selected.

[0020] The acquired descriptions of computing power requirements from different domains are segmented and part-of-speech tagged to identify keywords and phrases, resulting in a preliminary set of semantic units. Semantic units are then semantically annotated using a pre-built domain ontology knowledge base, such as WordNet or a domain-specific terminology dictionary. The Lesk algorithm is used to process polysemous words, mapping different expressions to a unified conceptual representation, achieving terminology unification and ambiguity elimination. Conditional random fields are employed for semantic role labeling, analyzing the syntactic dependency relationships between semantic units. A preliminary semantic dependency graph is constructed using the Stanford dependency parser, and then optimized using semantic rules to accurately capture the semantic structure of computing power requirements. A graph edit distance algorithm is used to match the constructed semantic dependency graph with predefined standardized computing power requirement templates. These templates contain common computing power requirement patterns, such as compute-intensive and memory-intensive. Through the matching process, a structured representation of computing power requirements conforming to a unified standard is generated, including key parameters such as computing resource type, quantity, and time requirements. Core attributes of computing power requirements, such as computation type, data scale, and time constraints, are extracted from the semantic dependency graph. For the extracted core attributes, a pre-defined computing resource library is queried, and the combination of computing resources with the highest matching degree is selected. If multiple computing resource combinations meet the requirements, they are sorted according to predefined priority rules, such as cost and energy consumption, and the optimal combination is selected. The selected computing resource combination is compared with the original requirement description to generate a structured computing power allocation scheme, including detailed information such as specific hardware configuration, software environment, and scheduling strategy. The obtained computing power requirement descriptions from different domains are segmented and tagged with parts of speech. Chinese word segmentation is performed using the jieba word segmentation tool to obtain word frequency statistics. Nouns, verbs, and adjectives that appear more than 5 times are extracted as keywords to form a preliminary set of semantic units. Based on a pre-built domain ontology knowledge base, such as a WordNet extended version containing 10,000 computing power-related terms, semantic units are semantically annotated. The Lesk algorithm is used to process polysemous words, calculate word semantic similarity, and select the word meaning with the highest similarity for disambiguation.

[0021] For example, the term "core" is mapped to "CPU core" or "core algorithm" in different contexts. Conditional random fields are used for semantic role labeling to identify subject, predicate, and object components. A preliminary semantic dependency graph containing 30 nodes and 50 edges is constructed using the Stanford dependency parser. Twenty predefined semantic rules are applied to optimize the dependency graph, such as merging synonymous nodes and deleting irrelevant modifiers. A graph edit distance algorithm is used to match the optimized semantic dependency graph with 100 predefined standardized computing power requirement templates. The graph edit distance is calculated, and the template with the smallest distance is selected as the best match. A structured representation of computing power requirements is generated, including key parameters such as computing resource type (GPU, CPU, FPGA), quantity (e.g., 8 CPU cores, 32GB memory), and time requirement (e.g., completion within 24 hours). Core attributes of the computing power requirement are extracted from the semantic dependency graph, such as computing type (intensive computing), data size (100TB), and time constraint (48 hours). The system queries a pre-defined computing resource database containing 1000 records, calculates the matching degree using cosine similarity, and selects computing resource combinations with a matching degree greater than 0.8. Based on predefined priority rules such as a cost weight of 0.6 and an energy consumption weight of 0.4, multiple computing resource combinations that meet the requirements are sorted, and the combination with the highest overall score is selected. A structured computing power allocation scheme is generated, including detailed information such as hardware configuration (e.g., 4 NVIDIA V100 GPU servers), software environment (Ubuntu 20.04 LTS, CUDA 11.2), and scheduling policy (dynamic load balancing).

[0022] Semantic analysis is performed on texts describing computing power requirements in different fields, including word segmentation, part-of-speech tagging, and named entity recognition, to obtain a standardized text representation. Key information, including computing power indicators, application scenarios, and data scale, is identified and extracted to form a structured representation of computing power requirements.

[0023] The system receives text describing computing power requirements from different domains and segments the text using the jieba word segmentation tool. High-frequency words are selected from the segmentation results based on a preset word frequency threshold. These high-frequency words are then used for part-of-speech tagging using StanfordPOSTagger. Keyword types such as nouns, verbs, and adjectives are identified within these high-frequency words to obtain a preliminary semantic unit set. Based on a pre-constructed multi-domain computing power requirement dictionary, the semantic unit set is semantically annotated. Named entities such as computing power metrics, application scenarios, and data scale are identified from the semantic unit set. A semantic dependency graph is constructed based on the semantic annotation results, and the subject-verb-object structure is extracted to determine the core attributes of the computing power requirement. The semantic dependency graph is matched against a predefined computing power requirement template library, and the template with the smallest edit distance is selected as the best match, generating a structured representation of the computing power requirement that includes elements such as computing power metrics, application scenarios, and data scale.

[0024] The acquired text describing computing power requirements from different domains was segmented using the jieba word segmentation tool. Word frequencies were statistically analyzed, and words exceeding a preset threshold were selected as high-frequency words. Stanford Postagger was used for part-of-speech tagging to identify keyword types such as nouns, verbs, and adjectives, resulting in a preliminary set of semantic units. Based on a pre-built multi-domain computing power requirement dictionary, semantic units were semantically annotated. This dictionary includes professional terms from computing, storage, and networking fields. Conditional Random Fields (CRF) algorithms were used to identify named entities in the text, including computing power metrics such as FLOPS and bandwidth, application scenarios such as deep learning and big data analysis, and data scale such as PB and TB levels. A Stanford Dependency Parser was used to construct a semantic dependency graph, capturing the syntactic relationships between semantic units. Subject-verb-object structures were extracted from the dependency graph to identify core attributes of computing power requirements, such as computing type (CPU-intensive, GPU-accelerated), storage requirements (high IOPS, large capacity), network bandwidth (low latency, high throughput), etc. Using a graph edit distance algorithm, the constructed semantic dependency graph is matched with a predefined computational power requirement template library, which contains typical computational power requirement patterns for different application scenarios. The template with the smallest edit distance is selected as the best match, generating a structured representation of computational power requirements that includes elements such as computational power indicators, application scenarios, and data scale. The acquired computational power requirement description texts from different domains are segmented using the jieba word segmentation tool, with a word frequency threshold of 10, selecting words with frequencies exceeding this threshold as high-frequency words. Part-of-speech tagging is performed using StanfordPOSTagger, identifying 500 nouns, 300 verbs, and 200 adjectives, forming a preliminary semantic unit set containing 1000 words. Based on a pre-constructed multi-domain computational power requirement dictionary containing 3000 terms from the computing domain, 2000 from the storage domain, and 1500 from the networking domain, semantic annotation is performed on the semantic units. This study utilizes Conditional Random Fields (CRF) algorithms to identify named entities in text, identifying 50 key computing power metrics (e.g., 100 TFLOPS, 100 Gbps bandwidth), 30 application scenarios (e.g., image recognition, natural language processing), and 20 data scale metrics (e.g., 10 PB, 500 TB). A Stanford Dependency Parser is used to construct a semantic dependency graph, capturing syntactic relationships between semantic units and extracting 100 subject-verb-object structures. Core attributes of computing power requirements are identified, including 20 computing types (e.g., 8-core CPU-intensive, 4 V100 GPUs accelerated), 15 storage requirements (e.g., 1 million IOPS, 1 PB capacity), and 10 network bandwidth requirements (e.g., 1 ms latency, 100 Gbps throughput). A graph edit distance algorithm is used to match the constructed semantic dependency graph with a predefined library of computing power requirement templates, containing 100 typical computing power requirement patterns from different application scenarios. The edit distance is calculated, and the template with the smallest distance is selected as the best match.Generate a structured representation of computing power requirements, which includes 15 computing power metrics, 10 application scenarios, and 5 data scales, forming a complete computing power requirement description document.

[0025] Step S102: Based on the structured representation of computing power requirements, a unified computing power requirements description language is adopted to define the various attributes and parameters of computing power requirements, forming a standardized computing power requirements description method, and realizing a consistent expression of computing power requirements across different domains.

[0026] A structured representation of computing power requirements is obtained, from which key attributes and parameters, including computing power, storage capacity, and network bandwidth, are extracted. Based on these key attributes and parameters, the data type, value range, and unit are determined, resulting in a basic set of elements for the computing power requirement description language. Ontology modeling is performed using the Protégé tool. Once the ontology modeling is complete, the class hierarchy, attribute features, and constraints are obtained. An XML Schema is designed based on the class hierarchy, attribute features, and constraints, defining the tag structure, attribute set, and syntax rules. A transformation template is written using XSLT technology. This template is used to convert computing power requirement descriptions from different domains into a standardized XML format. The standardized XML format is generated by XPath queries and template matching to achieve attribute correspondence and semantic equivalence relation processing.

[0027] Key attributes and parameters, such as computing power (FLOPS), storage capacity (GB / TB), and network bandwidth (Gbps), are extracted from the structured computing power demand representation using natural language processing tools. Statistical analysis is used to determine the data type of each attribute, such as integer, floating-point, enumerated type, value range, and unit, constructing the basic element set of the computing power demand description language. Ontology modeling is performed using the Protégé tool, defining class hierarchies such as computing resources, storage resources, and network resources; attribute features such as functional attributes and data attributes; and constraints such as cardinality constraints and value range constraints. An OWL-formatted computing power demand ontology model is constructed to achieve semantic associations and consistent expression between attributes. Based on the computing power demand ontology model, an XML Schema tag structure is designed, such as...<Compute Resource> The system establishes a unified format for describing computing power requirements, including attribute sets such as capacity and bandwidth, and syntax rules such as element nesting relationships and attribute value types. XSLT technology is used to develop transformation templates that convert computing power requirements from different domains into a standardized XML format. XPath queries and template matching are used to handle attribute correspondence and semantic equivalence, achieving consistent expression of cross-domain computing power requirements. NLTK tools are used to extract key attributes and parameters from the structured computing power requirement representation, identifying 20 key attributes such as computing power (e.g., 10 TFLOPS), storage capacity (e.g., 500 TB), and network bandwidth (e.g., 100 Gbps). Statistical analysis determines the data types of each attribute (e.g., 8 integers, 7 floating-point numbers, and 5 enumerations), and sets the value range and unit, constructing a computing power requirement description language set containing 100 basic elements. Ontology modeling was performed using Protégé 5.5.0, defining a three-level class hierarchy: top-level class (resource type), second-level classes (computing resources, storage resources, network resources), and third-level classes (specific resource items). Fifty functional attributes and 30 data attributes were set, along with 20 cardinality constraints and 15 range constraints. An OWL-formatted computing power requirement ontology model containing 200 ontology elements was constructed, achieving semantic associations and consistent expression between 80 attribute pairs. Based on the computing power requirement ontology model, an XML Schema framework was designed, defining 15 main tag structures as follows:<Compute Resource> The system incorporates 30 attribute sets, such as capacity and bandwidth, and 25 syntax rules, including a maximum element nesting depth of 5 levels and attribute value type restrictions, forming a unified format for describing computing power requirements. Using XSLT 3.0 technology, 50 transformation templates were developed, covering the conversion of computing power requirement descriptions from 5 different domains into a standardized XML format. Through 100 XPath query expressions and 80 template matching rules, 200 attribute correspondences and 150 semantic equivalence relations were processed, achieving consistent expression of cross-domain computing power requirements. The final result is a standardized XML document containing 300 computing power requirement description elements.

[0028] Step S103: Based on the standardized computing power demand description, formulate a cross-domain computing power scheduling communication protocol, clarify the process, method and interface for communication and negotiation of computing power demand between various domains, so as to ensure that the scheduling system can accurately obtain computing power demand information of each domain.

[0029] The message format is defined according to the cross-domain computing power scheduling protocol; a cross-domain communication network is constructed using a distributed hash table, and hash values ​​are obtained as domain node identifiers; rate limiting and priority management are implemented for cross-domain communication, and the token bucket capacity and token generation rate are set; the TLS 1.3 protocol is implemented through the Open SSL library, and encryption suites are configured for two-way authentication and encrypted data transmission; a negotiation process engine is developed, which determines the current state based on the state transition matrix if a trigger event is received; if the current state is IDLE and the trigger event is REQUEST SENT, the state is changed to REQUESTING; if the current state is REQUESTING and the trigger event is OFFER RECEIVED, the state is changed to NEGOTIATING; the negotiation process is driven based on the state transition results.

[0030] Based on standardized computing power requirement descriptions, Protocol Buffers are used to define the message format for cross-domain computing power scheduling protocols, including three types: Request Message, Response Message, and Notification Message. Each message type contains common fields such as message ID and timestamp, and specific fields such as request type. The message structure is serialized and encoded. A distributed hash table implemented using the Chord algorithm is used to construct the cross-domain communication network. Each domain node is assigned a 160-bit SHA-1 hash value as an identifier. A finger table and a successor list are implemented to optimize route lookup. A stabilization period of 60 seconds is set to maintain the network topology and establish an efficient point-to-point communication mechanism. A flow control mechanism based on the token bucket algorithm is designed, with a bucket capacity of 100 tokens and a token generation rate of 10 tokens / second. This mechanism performs rate limiting and priority management for cross-domain communication. The OpenSSL library is used to implement the TLS 1.3 protocol, and the ECDHE-ECDSA-AES256-GCM-SHA384 encryption suite is configured for two-way authentication and encrypted data transmission. A negotiation process engine based on a finite state machine was developed, defining five main states: IDLE, REQUESTING, NEGOTIATING, CONFIRMING, and COMPLETED, and ten triggering events such as REQUEST_SENT and OFFER_RECEIVED. A state transition matrix was used to define rules, and event queues and processors were implemented to drive state transitions. Various messages during the negotiation process were handled in an event-driven manner, enabling automated negotiation and decision-making for cross-domain computing power scheduling. Based on standardized computing power requirement descriptions, ProtocolBuffers 3.0 was used to define the message format for the cross-domain computing power scheduling protocol. A Request Message with 20 fields, a Response Message with 15 fields, and a Notification Message with 10 fields were created. Common fields include a 64-bit message ID and a millisecond-level timestamp. Specific fields, such as request type (8 types), response status (6 status codes), and notification type (4 notification types), resulted in an average message structure size of 2KB after serialization. A distributed hash table is implemented using the Chord algorithm. SHA-1 is used to generate 160-bit node identifiers. A cross-domain communication network containing 1000 nodes is constructed. Each node maintains a 160-row finger table and a successor list of 10 successor nodes. The average number of hops for route lookup is approximately log21000 ≈ 10. A stabilization period of 60 seconds is set to perform network maintenance.A token bucket algorithm is implemented for flow control, with a bucket capacity of 100 tokens, a generation rate of 10 tokens / second, and a peak processing capacity of 1000 requests / second. A priority queue is used to manage requests of four priorities. The TLS 1.3 protocol is implemented using OpenSSL version 1.1.1, configured with the ECDHE-ECDSA-AES256-GCM-SHA384 encryption suite, a 256-bit key length, and a symmetric encryption speed of up to 1Gbps. A finite state machine negotiation process engine is developed, defining five main states: IDLE, REQUESTING, NEGOTIATING, CONFIRMING, and COMPLETED. Ten triggering events are designed, including REQUEST_SENT and OFFER_RECEIVED. A 5x10 state transition matrix is ​​used to define 25 transition rules, implementing an event queue with a capacity of 1000 tokens. The processor can process 500 events per second, completing a single negotiation process within an average of three state transitions.

[0031] Based on the cross-domain computing power scheduling communication protocol, an scalable data exchange and adaptation mechanism is adopted to solve the differences in interface specifications and data formats of various domains. The interface parameters of different domains can be quickly added and modified in a configurable manner. Lightweight data acquisition agents are deployed in each domain. Through communication and negotiation with the resource management system within the domain, the resource usage and computing power demand changes of the domain are obtained in real time, and the collected information is reported to the cross-domain scheduling system in a unified format.

[0032] The process involves: acquiring the data exchange adapter information; generating a plug-in architecture-based data exchange adapter using Java's SPI mechanism; receiving the data acquisition agent request; generating a lightweight data acquisition agent based on the request; determining whether the data transmission request has been received; if received, implementing a data transmission mechanism based on Apache Kafka, using a publish-subscribe pattern to transfer data between the agent and the scheduling system; acquiring the raw data information, including data to be cleaned, transformed, and normalized; constructing an Apache Flink-based data processing pipeline based on the raw data information; and performing cleaning, transformation, and normalization on the acquired raw data.

[0033] Design a data exchange adapter based on a plug-in architecture. Utilize Java's SPI mechanism to implement the plug-in architecture. Define the interface specifications and data format mapping relationships for each domain through YAML configuration files. Dynamically load adapter plug-in JAR packages using a Class Loader to achieve runtime plug-in updates and complete data format conversion and protocol adaptation between different domains. Develop a lightweight data acquisition agent, employing Java's Completable Future for asynchronous operations. Create a thread pool of 100 threads to handle concurrent requests. Interact with the domain resource management system using both long polling and WebSocket modes to obtain resource usage and computing power demand changes. Implement a high-performance memory queue cache for acquired data using the Disruptor framework. Implement a data transmission mechanism based on Apache Kafka, using a publish-subscribe pattern to transfer data between the agent and scheduling system. Use Avro for data serialization and compression to achieve batch transmission of 1000 messages per batch. Ensure data reliability through Kafka's at-least-once semantics and manual offset commit by the consumer. A data processing pipeline based on Apache Flink was constructed to clean, deduplicate, fill missing values, transform, unify and standardize units and formats of the collected raw data. The processed time-series data was stored using Influx DB, and a Redis caching layer was used to provide millisecond-level high-performance data query services. A plug-in architecture-based data exchange adapter was designed, using the Java SPI mechanism to implement 10 plugins from different domains. Interface specifications and data format mappings were defined through a 50-line YAML configuration file. Twenty adapter plugin JAR packages were dynamically loaded using a Class Loader, enabling plugin updates within 5 minutes. A lightweight data acquisition agent was developed, using Completable Future to create a thread pool of 100 threads, concurrently processing 1000 requests per second. Resource data was acquired through both long polling at 3-second intervals and real-time push via WebSocket. A 1024-byte Ring Buffer was implemented using the Disruptor framework to cache 100,000 data entries per second. Implement a data transmission mechanism based on Apache Kafka, configure a cluster with 3 brokers and 10 partitions, use a publish-subscribe pattern to transfer data between 50 brokers and 1 scheduling system, use Avro compression to achieve a 30% reduction in original size, achieve batch transmission of 1000 messages per batch with a single batch size of 2MB, and ensure data reliability through at-least-once semantics and manual offset commit at 5-second intervals.An Apache Flink data processing pipeline was built, deploying 5 Task Manager nodes, each with 4 slots, to clean 1 million raw data entries per hour, achieving a 99.9% deduplication rate and 95% missing value imputation. The pipeline transformed data across 5 different unit systems and standardized 20 data formats. Influx DB was used to store the processed time-series data, with a single table capacity of 1TB and a write speed of 500,000 records per second. A caching cluster consisting of 3 Redis nodes provided query response times within 1ms and supported high-concurrency data query services of 100,000 QPS.

[0034] Step S104: In the scheduling system, establish a mechanism for comparing and prioritizing computing power requirements. Based on the standardized description of computing power requirements, prioritize the computing power requirements of different domains and determine the urgency and importance of each requirement.

[0035] Obtain the standardized computing power requirement description, extract key attributes from the description, and construct a multi-dimensional feature vector. Process the multi-dimensional feature vector using a pre-established principal component analysis model to obtain a dimensionality-reduced feature vector. Based on the dimensionality-reduced feature vector, perform rule reasoning using a pre-defined rule engine based on the XGBoost algorithm. The rule engine contains at least one evaluation rule set, the parameters of which are optimized using cross-validation. Determine the comprehensive score for each computing power requirement by judging the result of the rule reasoning. Implement a priority queue using a binary heap data structure. If a new computing power requirement is received, insert it into the priority queue. If the number of computing power requirements in the priority queue exceeds a preset threshold, delete the lowest-priority requirement. Aggregate the computing power requirements in the priority queue within a preset sliding time window. This aggregation includes merging computing power requirements of the same type and similar priority to obtain aggregated computing power requirements. Clustering is performed on the aggregated computing power demand; based on the clustering results, a weighted average method is used to calculate the priority score of each cluster, where the weights of the weighted average method are set based on the original priority.

[0036] Key attributes, including computing resource requirements, storage requirements, network bandwidth requirements, and time constraints, are extracted from standardized computing power requirement descriptions to construct multidimensional feature vectors. Principal component analysis is performed using the scikit-learn library, retaining principal components that explain 95% of the variance. Requirements from different domains are standardized to eliminate the influence of unit dimensions. A rule engine based on the XGBoost algorithm is designed, defining an evaluation rule set of 20 rules covering factors such as urgency, importance, and resource utilization. Model parameters, such as tree depth and learning rate, are optimized through 5-fold cross-validation. Rule inference is performed on the dimensionality-reduced features to calculate the comprehensive score for each computing power requirement. A priority queue is implemented using a binary heap, supporting insertion and deletion operations with O(logn) time complexity. The queue capacity is set to 10,000; when the capacity is exceeded, low-priority requirements are deleted. Computing power requirements are prioritized according to their comprehensive scores, and the ranking position of each requirement in the queue is updated in real time. A demand aggregation mechanism based on a 5-minute sliding time window is implemented, sliding once every 1 minute. Demands of the same type and similar priority are merged within a fixed time window. The K-means algorithm is used to cluster the demands within the window, with the number of clusters automatically determined based on the silhouette coefficient. A weighted average method is used to calculate the priority score after aggregation, with weights based on the original priority settings. Ten key attributes are extracted from standardized computing power demand descriptions, including the number of CPU cores, the number of GPUs, memory capacity, storage capacity, network bandwidth, and task deadline, to construct a 10-dimensional feature vector. Principal component analysis is implemented using the scikit-learn library, with a threshold of 95%, ultimately retaining 6 principal components. The 1000 demands from different domains are standardized to eliminate the influence of unit weight. A rule engine based on the XGBoost algorithm is designed, defining 20 evaluation rules, including 5 urgency rules, 8 importance rules, and 7 resource utilization rules. Model parameters are optimized through 5-fold cross-validation, with a maximum tree depth of 6, a learning rate of 0.1, and 100 iterations. Rule-based reasoning is performed on the dimensionality-reduced features to calculate a comprehensive score for each computing power request, ranging from 0 to 100. A priority queue is implemented using a binary heap, supporting insertion and deletion operations with O(logn) time complexity. The queue capacity is set to 10,000; requests with scores below 50 are deleted when the capacity is exceeded. Computing power requests are prioritized according to their comprehensive scores, and the ranking of each request in the queue is updated every second. A request aggregation mechanism based on a 5-minute sliding time window is implemented, sliding once every 1 minute to process approximately 300 requests within a fixed time window. The K-means algorithm is used to cluster the requests within the window, and the number of clusters is automatically determined by calculating the silhouette coefficient, typically between 3 and 8. A weighted average method is used to calculate the aggregated priority score, with weights based on the original priority settings: high priority has a weight of 0.6, medium priority has 0.3, and low priority has 0.1.

[0037] For standardized computing power requirement descriptions, a priority evaluation index system is used to obtain priority scores for computing power requirements in different domains from multiple dimensions such as business importance, urgency of requirement, and resource utilization efficiency, which serve as the basis for determining the scheduling order. If the computing power requirements of the scheduling system change, the scheduling system's response mechanism is triggered. According to the preset emergency scheduling rules, newly submitted critical requirements are given priority scheduling, and the priority and resource allocation of existing tasks are dynamically adjusted.

[0038] The system acquires multi-dimensional evaluation indicators encompassing business importance, urgency, and resource utilization efficiency. A judgment matrix is ​​constructed based on these indicators, and the weights for each dimension are calculated using the eigenvector method. A fuzzy evaluation set is represented using triangular fuzzy numbers. Based on the fuzzy evaluation set and weights, a weighted average method is used to obtain the fuzzy comprehensive evaluation result. The system receives requirement change events, including additions, deletions, and modifications. If the change in the number of requirements exceeds a preset threshold or the change in priority exceeds a preset threshold, the scheduling system's response process is triggered. Newly submitted requirement information is acquired. If the requirement is a critical requirement, it is assigned a higher initial priority, where the new task inherits part of the priority from the parent task, and the remaining part is calculated based on its own attributes. A depth-first search method is used to recursively adjust the priorities of existing tasks. If the recursion depth exceeds a preset number of levels, the recursive adjustment stops, thereby ensuring the fairness and efficiency of scheduling.

[0039] A multi-dimensional priority evaluation index system is constructed. Saaty's 9-level scaling method is used to build the judgment matrix, and weights are calculated using the eigenvector method. Three main dimensions are set: business importance, demand urgency, and resource utilization efficiency, with 3-5 sub-indicators under each dimension. The fuzzy evaluation set is represented using triangular fuzzy numbers, and a weighted average method is used for fuzzy comprehensive evaluation. Standardized computing power demand descriptions are scored, and a membership function is set to transform qualitative indicators into membership values ​​in the [0,1] interval. The comprehensive priority score for each computing power demand is obtained through weighted summation. An event-driven scheduling response mechanism based on the observer pattern is implemented. Demand change event types are defined, such as addition, deletion, and modification, and trigger conditions and thresholds are set. When the number of demands changes by more than 10% or the priority changes by more than 20%, the scheduling system's response process is triggered, and the priority in the current task queue is reassessed. A dynamic priority adjustment algorithm based on priority inheritance is implemented. According to preset emergency scheduling rules, newly submitted critical requirements are assigned a higher initial priority. New tasks inherit 50% of the priority from their parent tasks, and the remaining 50% is calculated based on their own attributes. Simultaneously, the priorities and resource allocations of existing tasks are recursively adjusted using a depth-first search, limiting the recursion depth to three levels to prevent excessive adjustment and ensure overall scheduling fairness and efficiency. A priority evaluation index system is constructed, using Saaty's 9-level scaling method to build a 3x3 judgment matrix. Weights are calculated using the eigenvector method, resulting in a weight of 0.5 for business importance, 0.3 for requirement urgency, and 0.2 for resource utilization efficiency. Four sub-indicators are set for each dimension, totaling 12 indicators. A triangular fuzzy number (a, b, c) is used to represent the fuzzy evaluation set; for example, business importance is represented by (0.6, 0.8, 1.0) to indicate high importance. 1000 standardized computing power demand descriptions are scored, and a membership function μ(x) = 1 / (1 + ((xc) / (ca))^2) is set to convert qualitative indicators into membership values ​​in the range [0,1]. A weighted average method is used to obtain the comprehensive priority score for each demand, ranging from 0 to 100. An event-driven mechanism based on the observer pattern is implemented, defining three types of demand change events: addition, deletion, and modification. Trigger thresholds are set: a change in the number of demands exceeding 10% or a change in priority exceeding 20%. When 100 new demands or 200 demands experience a priority change exceeding 20 points, a scheduling response is triggered, re-evaluating the priorities of 10,000 tasks. A priority inheritance algorithm is implemented, where a new task inherits 50% of the parent task's priority (e.g., if the parent task's priority is 80, the new task's initial priority is 40). The remaining 50% is calculated based on its own scores across three dimensions (e.g., if the score is 30, the final priority is 70). The recursive adjustment uses a depth-first search, limiting the recursion depth to 3 levels, and the impact of each level is limited to 100 related tasks.

[0040] Step S105: Based on the priority of computing power demand and the computing power resource situation of each domain, generate a cross-domain computing power scheduling scheme, clarify the computing power allocation scheme and scheduling schedule of each domain, so as to achieve accurate matching between computing power demand and computing power resources.

[0041] A tree structure is used to represent the computing resources of each domain, where the root node is the total resource pool, child nodes are the resources of each domain, and leaf nodes are the specific resource items. Resource type, capacity, utilization rate, and other attribute information are obtained based on the tree structure. Factors such as computing power demand priority, resource constraints, and task dependencies are transformed into constraints. Constraint propagation is performed based on these constraints. If constraint propagation is complete, a minimum residual value heuristic is used to select variables. Three time windows—urgent, normal, and low priority—are set based on the urgency of the demand. Within each time window, a shortest job first strategy is used for locally optimal scheduling. If the time window slides, the locally optimal scheduling is re-performed. The resource utilization rate is obtained. If the resource utilization rate is lower than a preset lower threshold, task migration is triggered. If the resource utilization rate is higher than a preset upper threshold, task migration is triggered. Based on the task migration results, a weighted round-robin algorithm is used for task allocation.

[0042] A cross-domain resource pool model is constructed, using a tree structure to represent the computing resources of each domain. The root node represents the total resource pool, child nodes represent the resources of each domain, and leaf nodes represent specific resource items, including types such as computing, storage, and network. Each node contains attributes such as resource type, capacity, and utilization rate. Through resource abstraction, the resource characteristics and capacity of different domains are uniformly described. A scheduling algorithm based on constraint satisfaction is designed, transforming factors such as computing demand priority, resource constraints, and task dependencies into constraints. The AC-3 algorithm is used for constraint propagation, and the minimum residual value heuristic is used to select variables and the minimum constraint value heuristic for assignment. The backtracking depth is limited to 100; if the limit is exceeded, the search is restarted to find the optimal scheduling scheme. A dynamic time window scheduling strategy is implemented, setting three time window sizes based on the urgency of the demand: 1 hour for urgent needs, 4 hours for normal needs, and 12 hours for low priority needs. Within the window, the shortest job first strategy is used for locally optimal scheduling, and the window slides every 15 minutes to achieve global scheduling. An adaptive load balancing mechanism was developed, setting resource utilization thresholds. Task migration is triggered when utilization falls below 30%, and task migration is triggered when utilization exceeds 80%. An exponentially weighted moving average algorithm is used to smooth resource utilization fluctuations, with a weight factor set to 0.3. Task allocation employs a weighted round-robin algorithm, with weights dynamically adjusted based on resource idle ratios to achieve balanced utilization of cross-domain resources. A cross-domain resource pool model is constructed, using a 5-level tree structure to represent the computing resources of 10 domains. The root node represents the total resource pool, containing 1 million CPU cores, 100,000 GPUs, and 1EB of storage capacity. Each domain node contains an average of 100,000 CPU cores, 10,000 GPUs, and 100PB of storage. Leaf nodes are accurate to the individual server level, recording real-time utilization. The AC-3 algorithm is used to process 1000 constraints. The minimum residual value heuristic selects the optimal variable within an average of 10ms, and the minimum constraint value heuristic determines the assignment within 5ms. A backtracking depth limit of 100 is set; exceeding this limit restarts the search. The optimal solution is found after an average of 50 restarts. A dynamic time window strategy is implemented: 1000 jobs are processed within a 1-hour window for urgent tasks, 5000 jobs within a 4-hour window for normal tasks, and 20000 jobs within a 12-hour window for low-priority tasks. The window slides every 15 minutes, using a shortest job first algorithm, reducing the average job completion time by 30%. In the adaptive load balancing mechanism, 10 tasks are migrated in per minute when resource utilization is below 30%, and 5 tasks are migrated out per minute when it is above 80%. The exponentially weighted moving average algorithm uses a weight factor of 0.3, smoothing out resource utilization fluctuations by 50%. The weighted round-robin algorithm dynamically adjusts weights based on the resource idle ratio, reducing load imbalance from 30% to below 5%.

[0043] In step S106, during the generation of the cross-domain computing power scheduling scheme, the scheduling scheme is dynamically adjusted through interaction and feedback with each domain, and a monitoring and evaluation mechanism for cross-domain computing power scheduling is established to track the execution of the scheduling scheme in real time, collect feedback information from each domain, evaluate the scheduling effect, and optimize the computing power scheduling system based on the evaluation results.

[0044] The system receives resource status information and task execution information carrying domain identifiers, which are pushed by each domain through a message queue. It updates the monitoring metric database based on the resource status information and task execution information. The monitoring metric database includes multi-dimensional metrics such as resource utilization, task completion time, and scheduling latency. It calculates the average, maximum, and minimum values ​​of CPU utilization, memory utilization, and network throughput in the monitoring metric database. If a monitoring metric triggers a preset rule, it executes the corresponding scheduling strategy adjustment, where the preset rules include high load rules and task latency rules. It acquires past monitoring data to predict future resource demands and task execution. It updates the scheduling strategy parameters based on the prediction results. The scheduling strategy parameters are used to optimize resource allocation and task scheduling. The above steps are repeated to continuously optimize the cross-domain resource scheduling effect.

[0045] A cross-domain communication network based on a publish-subscribe pattern is constructed. Two topics, "resource-status" and "task-execution," are set up using Apache Kafka. Each domain pushes resource status and task execution information in real time via a message queue every 5 seconds. The central scheduler subscribes to relevant topics and obtains feedback information between domains. A multi-dimensional monitoring indicator system is designed, including resource utilization, task completion time, and scheduling latency. InfluxDB is used to store monitoring data, and three sliding time windows (1 minute, 5 minutes, and 15 minutes) are set to calculate the average, maximum, and minimum values ​​of indicators such as CPU utilization, memory utilization, and network throughput. A dynamic adjustment mechanism based on the Drools rule engine is implemented, defining 10 rules such as "high load rule" and "task latency rule." Rule evaluation is triggered every 30 seconds based on monitoring indicators. If resource utilization exceeds a threshold, load balancing is performed; if task latency exceeds a limit, priority is adjusted. An adaptive learning algorithm based on an LSTM neural network was developed. Inputting monitoring data from the past 24 hours, it predicts resource demand and task execution status for the next hour. The model is retrained and parameters are updated hourly using newly added historical data, and scheduling strategies are optimized based on the prediction results to continuously improve scheduling efficiency. A cross-domain communication network was built using Apache Kafka 2.8.0, configured with three broker nodes and two topics: "resource-status" and "task-execution," each divided into 10 partitions. Each domain pushes a status update every 5 seconds, with an average message size of 2KB. The central scheduler uses a consumer group mode to process messages in parallel, keeping the average latency below 50ms. The monitoring metric system uses InfluxDB 2.0 to store data, with three sliding windows of 1 minute, 5 minutes, and 15 minutes, writing 1000 data points per second. The average, maximum, and minimum values ​​of 10 key metrics, including CPU utilization, memory utilization, and network throughput, were calculated to form 30 derived metrics. The Drools rule engine version 7.59.0 implements dynamic adjustment, defining 10 rules, including "trigger load balancing when CPU utilization exceeds 80%" and "adjust priority when task latency exceeds 5 minutes." Rules are evaluated every 30 seconds, with an average processing time of 100ms. The LSTM neural network uses a 4-layer structure: 24 neurons in the input layer corresponding to 24 hours of historical data, 64 neurons in each of the two hidden layers, and 6 neurons in the output layer, predicting six data points at 10-minute intervals for the next hour. Using the Adam optimizer with a learning rate of 0.001, the model is retrained with 1440 new data points per hour, one data point per minute, achieving a prediction accuracy of 90%. Optimization of the scheduling strategy improves efficiency by 15%.

[0046] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of the embodiments of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.

Claims

1. A method for cross-domain computing power scheduling of large models, characterized in that, include: Obtain computational power demand description texts from different domains, perform semantic analysis on the computational power demand description texts, and generate structured computational power demand representations, including: The text describing the computing power requirement was segmented using a word segmentation tool to obtain the segmentation results. High-frequency words are extracted from the word segmentation results based on a preset word frequency threshold. The high-frequency words are then labeled with part-of-speech tags to identify keyword types and obtain a preliminary set of semantic units. The semantic unit set is semantically annotated according to a pre-built multi-domain computing power demand dictionary to identify computing power indicators, application scenarios and data scale named entities; Construct a semantic dependency graph, extract the subject-verb-object structure, and determine the core attributes of computing power requirements; The semantic dependency graph is matched with a predefined computing power requirement template library, and the template with the smallest edit distance is selected to generate a structured computing power requirement representation that includes computing power indicators, application scenarios and data scale. Based on the structured representation of computing power requirements, a unified computing power requirement description language is adopted to form a standardized computing power requirement description. Based on the standardized computing power requirement description, a communication protocol for cross-domain computing power scheduling of large models is formulated, and computing power resource information of each domain is obtained through data exchange and adaptation mechanisms. Based on the standardized computing power requirement description and the computing power resource information, a priority determination mechanism for computing power requirements is established, and a scheme for cross-domain computing power scheduling of large models is generated, including: A tree structure is used to represent the computing resources of each domain, where the root node is the total resource pool, the child nodes are the resources of each domain, and the leaf nodes are the specific resource items. Retrieve resource type, capacity, and utilization attribute information based on the tree structure; Transform computing power demand priorities, resource constraints, and task dependencies into constraints; Constraint propagation is performed based on the constraints. If the constraint propagation is complete, then the minimum residual value heuristic is used to select variables; Set three time windows based on the urgency of the need: urgent, normal, and low priority. Within the time window, a shortest job priority strategy is used for locally optimal scheduling. If the time window slides, then local optimal scheduling is performed again; Obtain resource utilization rate; if the resource utilization rate is lower than a preset lower threshold, trigger task migration. If the resource utilization rate is higher than the preset upper limit threshold, the task migration will be triggered. Based on the results of task migration in or out, a weighted round-robin algorithm is used for task allocation; The scheme for cross-domain computing power scheduling of the large model is dynamically adjusted through monitoring and feedback mechanisms to track execution and optimize scheduling effects.

2. The method for cross-domain computing power scheduling of large models as described in claim 1, characterized in that, The adoption of a unified computing power requirement description language forms a standardized computing power requirement description, including: Key attributes and parameters are extracted from the structured representation of computing power requirements, including computing power, storage capacity, and network bandwidth. Based on the key attributes and parameters, the data type, value range, and unit are determined, resulting in the basic element set of the computing power demand description language. Obtain class hierarchy, attribute features, and constraints using ontology modeling tools; Based on the class hierarchy, attribute features, and constraints, a tag structure and syntax rules are designed. By converting templates, the computing power requirements of different domains are transformed into a standardized format, thereby realizing attribute correspondence and semantic equivalence processing.

3. The method for cross-domain computing power scheduling of large models as described in claim 1, characterized in that, The aforementioned communication protocol for cross-domain computing power scheduling of large models, which obtains computing resource information from each domain through data exchange and adaptation mechanisms, includes: Based on the protocol definition of cross-domain computing power scheduling in the large model, a cross-domain communication network is constructed using a distributed hash table, and the hash value is obtained as the domain node identifier. For cross-domain communication, rate limiting and priority management are set up, and two-way authentication and encrypted data transmission are achieved through encryption protocols; Develop a negotiation process engine that determines the current state based on the state transition matrix and drives the negotiation process based on the state transition results. A lightweight data acquisition agent is generated through a plug-in architecture-based data exchange adapter. A publish-subscribe model is used to transfer data between the agent and the scheduling system, and the acquired data is cleaned, transformed, and standardized.

4. The method for cross-domain computing power scheduling of large models as described in claim 1, characterized in that, The mechanism for prioritizing computing power requirements includes: Key attributes are extracted from the standardized computing power requirement description to construct a multi-dimensional feature vector; The multidimensional feature vector is processed by a pre-established dimensionality reduction model to obtain the dimensionality-reduced feature vector; Based on the reduced feature vectors, rule reasoning is performed through a preset rule engine to determine the comprehensive score for each computing power requirement; A priority queue is used to manage computing power demand. If a new computing power demand is received, the new computing power demand is inserted into the priority queue. For the computing power demand in the priority queue, demand aggregation is performed within a preset time window, and computing power demands of the same type are merged. The priority score of each cluster is calculated based on the clustering results.

5. The method for cross-domain computing power scheduling of large models as described in claim 1, characterized in that, Based on the standardized computing power requirement description and the computing power resource information, a priority determination mechanism for computing power requirements is established, and a scheme for cross-domain computing power scheduling of large models is generated, including: A tree structure is used to represent the computing resources of each domain, and the resource type, capacity and utilization information are obtained based on the tree structure; The computing power demand priority and resource constraints are transformed into constraints, and variables are selected through constraint propagation. Different priority time windows are set according to the urgency of the demand, and scheduling is carried out within the time window using a local optimum strategy. If the resource utilization rate exceeds the preset threshold range, task migration in or out is triggered, and task allocation is performed through a weighted round-robin algorithm to generate computing power allocation schemes and scheduling schedules for each domain.

6. The method for cross-domain computing power scheduling of large models as described in claim 1, characterized in that, The scheme for dynamically adjusting the cross-domain computing power scheduling of the large model through monitoring and feedback mechanisms includes: Receive resource status information and task execution information pushed by various domains, and update the monitoring indicator database, which includes resource utilization and task completion time; Calculate the statistical values ​​of the indicators in the monitoring indicator database. If the statistical values ​​of the indicators trigger a preset rule, then execute the scheduling strategy adjustment. Based on past monitoring data, predict future resource needs and update scheduling strategy parameters; Continuously implement monitoring and forecasting steps, dynamically adjust resource allocation and task scheduling schemes, and optimize cross-domain scheduling performance.

7. The method for cross-domain computing power scheduling of large models as described in claim 1, characterized in that, The mechanism for prioritizing computing power requirements and generating a scheme for cross-domain computing power scheduling of large models includes: Obtain multi-dimensional evaluation indicators, construct a judgment matrix, and calculate the weights of each dimension using the eigenvector method; Based on the fuzzy evaluation set and the weights, a comprehensive evaluation result is obtained using the weighted average method. Receive demand change events; if the change in demand quantity or priority exceeds a preset threshold, trigger the scheduling system response process. Newly submitted critical requirements are assigned an initial priority. The priority of existing tasks is dynamically adjusted through a recursive adjustment method. If the recursion depth exceeds the preset number of levels, the adjustment is stopped.

8. The method for cross-domain computing power scheduling of a large model according to claim 1, characterized in that, In the process of generating a scheme for cross-domain computing power scheduling of large models, resource status information and task execution information carrying domain identifiers are received. The resource status information and task execution information are pushed by each domain through a message queue. The monitoring indicator database is updated based on resource status information and task execution information. The monitoring indicator database includes multi-dimensional indicators such as resource utilization, task completion time, and scheduling delay. Predict future resource needs and task execution based on multi-dimensional indicators and past monitoring data; The scheduling strategy parameters are updated based on the prediction results. These parameters are used to optimize resource allocation and task scheduling to improve the effectiveness of cross-domain resource scheduling.

Citation Information

Patent Citations

  • Multi-source computing power data integration and intelligent scheduling system and method

    CN118916147A

  • Calculation power scheduling method and system

    CN119127473A