Data, computing power, algorithm and application fusion scheduling method, device and medium

By constructing a trusted data space and a full-element resource map, the security and compliance issues of data circulation have been resolved, and the differentiated needs of multiple industries have been met. This has enabled the full release of data element value and the secure and rational allocation of its resources.

CN122431828APending Publication Date: 2026-07-21JINAN BIG DATA GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JINAN BIG DATA GROUP CO LTD
Filing Date
2026-04-21
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing technologies, the lack of a unified and trustworthy data support environment for data, computing power, algorithms, and application scheduling makes it difficult to guarantee the security and compliance of data circulation, the value of data elements cannot be fully released, and the supply and demand matching is inefficient, making it difficult to adapt to the differentiated needs of multiple industries.

Method used

Construct a trusted data space, a full-element resource map, and a trusted constraint rule set; receive user demand requests; generate matching schemes through a supply-demand matching objective function; combine a scheduling optimization objective function and multiple scheduling strategies to generate optimal scheduling instructions; monitor task execution in real time; and dynamically adjust the trusted constraint rule set.

Benefits of technology

It improves the security and compliance of data circulation, meets the differentiated needs of multiple industries, and realizes the full release of the value of data elements and the security and rationality of their allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122431828A_ABST
    Figure CN122431828A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of computers and provides a fusion scheduling method for data, computing power, algorithms and applications, which comprises the following steps: constructing a trusted data space, a full-element resource graph and a trusted constraint rule set; receiving a user demand request and analyzing and mapping the user demand request to the full-element resource graph; performing trusted constraint screening on the full-element resource graph based on the trusted constraint rule set; generating a supply-demand matching scheme through a supply-demand matching target function in a trusted candidate resource set; optimizing the matching scheme by using a scheduling optimization target function and combining multiple scheduling strategies to generate optimal scheduling instructions; recording a task execution process in real time; and after the task is completed, calculating a reinforcement learning reward value, iteratively optimizing supply-demand matching target function parameters and scheduling optimization target function parameters, and the trusted constraint rule set. The application realizes real-time participation in scheduling decision-making by the trusted constraint, improves the safety and compliance of data circulation, and fully releases the value of data elements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a method, device, and medium for the integrated scheduling of data, computing power, algorithms, and applications. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] With the development of the digital economy, computing power and algorithms, as core production requirements, have become key indicators of digital competitiveness in terms of their circulation efficiency and integration level. Currently, the digital economy faces prominent problems such as inefficient factor circulation, inefficient supply-demand matching, and a lack of customized services. On the one hand, computing resources are unevenly distributed, with a large amount of idle computing power failing to effectively connect with actual needs. Algorithm transformation paths are limited, and the implementation of application scenarios faces the "last mile" obstacle. On the other hand, enterprises and institutions are increasingly demanding personalized and secure computing power and algorithm services during their digital transformation, but the market lacks integrated solutions that can integrate factor transactions, customized development, and integrated scheduling, resulting in a gap between factor supply and demand.

[0004] In existing technologies, the scheduling of data, computing power, algorithms, and applications largely relies on traditional centralized management models. These models lack a unified and trustworthy data support environment, making it difficult to guarantee the security and compliance of data circulation. Data security strategies and scheduling decisions are disconnected, making it impossible to achieve real-time driving of the scheduling process by security constraints. This makes it difficult to adapt to the differentiated needs of various industries such as urban management, emergency management, and healthcare, thus restricting the full release of the value of data elements. Summary of the Invention

[0005] To overcome the shortcomings of the prior art, this invention provides a method, device, and medium for the integrated scheduling of data, computing power, algorithms, and applications, aiming to solve the technical problem that data security strategies and scheduling decisions are separated in the prior art, and the value of data elements cannot be fully released.

[0006] To achieve the above objectives, the technical solution provided by the present invention is as follows: In a first aspect, the present invention provides a method for the integrated scheduling of data, computing power, algorithms, and applications, comprising: Construct a trusted data space, a full-element resource map, and a trusted set of constraint rules; The set of trusted constraint rules includes data out-of-domain constraints, computing power security domain constraints, access permission constraints, and compliance policy constraints. The system receives user demand requests, parses the requests and maps them to a full-element resource map. Based on a set of trustworthy constraint rules, it performs trustworthy constraint screening on candidate resources in the full-element resource map to generate a set of trustworthy candidate resources. Within the set of trustworthy candidate resources, it generates a matching scheme that includes data, computing power, algorithms and applications through a supply and demand matching objective function. The matching scheme is optimized using the scheduling optimization objective function, and the optimal scheduling instruction is generated by combining scheduling strategies such as elastic scaling scheduling, data affinity scheduling, load balancing scheduling and algorithm-based selection scheduling. During task execution, real-time monitoring is performed, and the entire process operation log and resource usage record are recorded. After the task is completed, the reinforcement learning reward value is calculated, and when the triggering conditions are met, the parameters of the supply and demand matching objective function and the scheduling optimization objective function are iteratively optimized. The operation log and resource usage record are fed back to the trusted data space, and the trusted constraint rule set is dynamically adjusted.

[0007] Optionally, the trusted data space includes secure data storage, fine-grained access control, and compliance audit traceability.

[0008] Optionally, the full-element resource map includes nodes and resource correlation. The nodes include data, computing power, algorithms, and applications. The resource correlation includes data format compatibility, computing power architecture adaptability, security level consistency, and application scenario matching between resources.

[0009] Optionally, the user requirement request includes business scenario type, computing power specification requirements, algorithm function requirements, data access permissions, security level requirements, and service quality indicators.

[0010] Optionally, the supply and demand matching objective function is a weighted sum of supply and demand distance, correlation, scheduling cost, and trust constraint penalty term.

[0011] Optionally, the scheduling optimization objective function is a weighted sum of resource utilization, execution efficiency, latency, and reliable scheduling score.

[0012] Optionally, the trusted scheduling score is calculated using the computing power security certification level, the highest security level, and the resource compliance indication function; The resource compliance indicator function takes the value 1 when the resource satisfies all the rules in the trusted constraint rule set; otherwise, it takes the value 0.

[0013] Optionally, the reinforcement learning reward value is obtained by weighted summation of user satisfaction, execution performance, and error rate.

[0014] In a second aspect, the present invention provides an electronic device including a processor and a memory, wherein the memory is used to store a computer program, the computer program being loaded and executed by the processor to implement the full-element fusion scheduling method of data, computing power, algorithms and applications as described in the first aspect above.

[0015] Thirdly, the present invention provides a computer-readable storage medium for storing a computer program; wherein the computer program, when executed by a processor, implements the full-element fusion scheduling method of data, computing power, algorithms and applications as described in the first aspect above.

[0016] As can be seen from the above technical solutions, the present invention has the following beneficial effects: The technical solution of this invention provides a method, device, and medium for the integrated scheduling of data, computing power, algorithms, and applications. First, it constructs a trusted data space, a full-element resource map, and a trusted constraint rule set. Next, it receives user demand requests, parses and maps these requests to the full-element resource map, and filters the map based on the trusted constraint rule set. A supply-demand matching objective function generates a matching scheme on the filtered map. Then, it optimizes the matching scheme using a scheduling optimization model and generates scheduling instructions by combining various scheduling strategies. Finally, it records the entire process operation log and resource usage records during task execution. After the task is completed, it calculates the reinforcement learning reward value and optimizes the parameters of the supply-demand matching objective function, or adjusts the trusted constraint rule set based on the reward value. In this invention, trusted constraints participate in the scheduling decision-making process in real time, improving the security and compliance of data flow, meeting the differentiated needs of multiple industries, fully releasing the value of data elements, and improving the security and rationality of scheduling. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the data, computing power, algorithm and application fusion scheduling method provided in Embodiment 1 of the present invention. Detailed Implementation

[0019] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0020] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0021] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0022] Example 1 For the scheduling of all elements including data, computing power, algorithms, and applications, existing scheduling methods lack a unified and reliable data support environment. This compromises the security and compliance of data circulation, hindering the full realization of the value of data elements. Figure 1 As shown, this embodiment provides a method for the integrated scheduling of all elements, including data, computing power, algorithms, and applications, comprising: Step S11: Construct a trusted data space, a full-element resource map, and a trusted constraint rule set; The set of trusted constraint rules includes data out-of-domain constraints, computing power security domain constraints, access permission constraints, and compliance policy constraints. Step S12: Receive user demand requests and parse the demand requests to map them to a full-element resource map. Based on the set of trustworthy constraint rules, perform trustworthy constraint screening on candidate resources in the full-element resource map to generate a set of trustworthy candidate resources. Within the set of trustworthy candidate resources, generate a matching scheme containing data, computing power, algorithms, and applications through a supply and demand matching objective function. Step S13: Optimize the matching scheme using the scheduling optimization objective function, and generate the optimal scheduling instruction by combining scheduling strategies such as elastic scaling scheduling, data affinity scheduling, load balancing scheduling, and algorithm-based selection scheduling. Step S14: Real-time monitoring is performed during task execution, and the entire process operation log and resource usage record are recorded; after the task is completed, the reinforcement learning reward value is calculated, and when the triggering condition is met, the parameters of the supply and demand matching objective function and the scheduling optimization objective function are iteratively optimized; and the operation log and resource usage record are fed back to the trusted data space to dynamically adjust the trusted constraint rule set.

[0023] The following content provides a more detailed description of the all-element fusion scheduling method for data, computing power, algorithms, and applications provided in this embodiment.

[0024] In step S11, a trusted data space, a full-element resource map, and a trusted constraint rule set are constructed. In this embodiment of the application, the trusted constraint rule set includes data out-of-domain constraints, computing power security domain constraints, access permission constraints, and compliance policy constraints; Understandably, the trusted constraint rule set is generated based on the data security classification results, access control policies, and compliance audit requirements in the trusted data space.

[0025] Specifically, data outbound constraints, based on data security classification results, limit the physical cluster range within which data of different security levels is allowed to be computed and stored, preventing highly sensitive data from flowing out of the specified security boundary; the specific rule is: when the data security level ≥ L threshold In this case, the computation, storage, and transmission of the data are only permitted on computing nodes with a security certification level no lower than the data's security level. For example, video surveillance data with a security level of L4 (the highest level) is only allowed to be processed within the L4 private network GPU server cluster A and is prohibited from being transmitted to general computing clusters with lower security levels; traffic checkpoint data with a security level of L3 is allowed to be computed within clusters A and B, but is prohibited from being processed on cluster C, which only has L1 certification.

[0026] Computational power security domain constraints set access admission conditions for computing power clusters that carry high-security-level data, allowing only authenticated and authorized user tasks to execute on the cluster, preventing unauthorized tasks from occupying security domain resources or accessing sensitive data; the specific rule is: security domain level S k The computing cluster only accepts user security credentials of level C. user ≥S k Task submission. For example: Private network GPU cluster A is designated as the highest security domain, and only accepts tasks submitted by authorized users (such as users from designated departments such as the Emergency Management Bureau and the Public Security Bureau) who hold private network security tokens and have passed two-factor authentication; unauthorized ordinary user tasks may not be scheduled to this cluster even if computing resources are idle.

[0027] Access restrictions are used to approve and control data resource access across organizations and departments, ensuring that data sharing occurs within the scope authorized by the data owner. Specifically, the rules are: when the organization requesting the task belongs to O... req Organizational affiliation with data owner O data In case of inconsistencies, access is only permitted after approval and authorization from the data owner. For example, when an emergency management department requests traffic checkpoint data from the traffic management bureau, the system automatically triggers a cross-departmental data access approval process. The data access token can only be obtained after approval by the traffic management bureau's data administrator. The approval record is included in the trusted audit log, and if the approval timeout (default 30 minutes), it will be automatically rejected and an alternative data source will be recommended.

[0028] Compliance policy constraints impose compliance protection requirements on data processing behaviors involving personal privacy and sensitive information, ensuring that the data processing process complies with national regulations and industry standards. Specifically, when data tags contain sensitive categories such as "personal privacy," "biometrics," or "location trajectory," the processing process must meet the corresponding compliance protection algorithm requirements. For example, processing tasks involving personal privacy data (such as facial recognition results, license plate numbers, etc.) must pass through a differential privacy processing engine before the data enters the computing node, following the ε-differential privacy mechanism (ε≤1.0, δ≤10). 5 Add Laplace noise to ensure that individual personal information cannot be reverse-identified; only output the calculation results.

[0029] The aforementioned trusted constraint rule set technical solution serves as the core constraint input for subsequent supply and demand matching and scheduling optimization. It ensures that scheduling decisions throughout the entire process are executed within the trusted boundary, thereby enabling the trusted data space to drive the scheduling of data, computing power, algorithms, and applications, and ultimately realizing the full release of the value of data elements.

[0030] In one specific implementation, the trusted data space includes secure data storage, fine-grained access control, and compliance audit traceability.

[0031] In one specific implementation, the full-element resource map includes nodes and resource correlation. The nodes include data, computing power, algorithms, and applications. The resource correlation includes data format compatibility, computing power architecture adaptability, security level consistency, and application scenario matching between resources.

[0032] Understandably, the data is obtained by integrating structured data, unstructured data, and semi-structured data, and the data undergoes de-identification processing, integrity verification, and trusted authentication.

[0033] Specifically, the desensitization process employs a desensitization algorithm based on differential privacy.

[0034] Understandably, the computing power includes GPU servers, CPUs, memory, etc.; the algorithms include computer vision, natural language processing, large models, etc.; and the applications include application scenario templates, service orchestration patterns, industry adaptation solutions, etc.

[0035] Specifically, the resource correlation degree is expressed as: ; in, Representing resources and The degree of correlation, , It is a resource in data, computing power, algorithms, and applications; Indicates the first Class attribute weight, Indicates the first The function for calculating the similarity of class attributes, where n is the total number of attributes, and n=4.

[0036] Understandably, The class attributes are data format compatibility, computing power architecture adaptability, security level consistency, and application scenario matching.

[0037] It should be noted that data format compatibility is denoted as... =1, when calculating the resource correlation between data and algorithms. The Jaccard similarity coefficient is used to calculate the degree of overlap between the data and the set of data formats supported by the algorithm. ,in , Resources , Supported data formats.

[0038] The computing architecture adaptability is denoted as =2, when calculating the resource correlation between computing power and algorithm. The structural fit between the algorithm and the computing power is calculated using a binary matching function. If the algorithm and the computing power are matched, then... Partial compatibility If incompatible .

[0039] Security level consistency is recorded as =3, calculate the matching degree between data security level and computing power security certification level. =1- ,in, Indicates the data security level. Indicates the computing power security certification level. This indicates the highest level of security.

[0040] Application scenario matching degree is denoted as =4, using cosine similarity to calculate the similarity between application scenario vectors and algorithm or data vectors. or .

[0041] in, Represents an application scenario vector. Represents the algorithm vector. Represents a data vector.

[0042] The aforementioned resource correlation technology solution breaks down information silos among the four elements of data, computing power, algorithms, and applications, establishes explicit and quantitative correlations among these four elements, and forms a reliable data space. This provides accurate and reliable resource recommendations for subsequent resource supply and demand matching and scheduling optimization, avoiding scheduling failures caused by unclear compatibility between elements.

[0043] In step S12, user demand requests are received and parsed to map to a full-element resource map. Based on the set of trustworthy constraint rules, candidate resources in the full-element resource map are screened using trustworthy constraints to generate a set of trustworthy candidate resources. Within the set of trustworthy candidate resources, a matching scheme containing data, computing power, algorithms, and applications is generated through a supply and demand matching objective function.

[0044] To address the issue of poor accuracy in traditional supply and demand matching schemes, this invention designs a supply and demand matching objective function that includes supply and demand distance, correlation, scheduling cost, and credibility constraint penalty terms, thereby improving the accuracy of supply and demand matching scheme generation.

[0045] In one specific implementation, the user requirement request includes business scenario type, computing power specification requirements, algorithm function requirements, data access permissions, security level requirements, and service quality indicators.

[0046] Specifically, based on the set of credible constraint rules, candidate resources in the full-element resource map are screened using credible constraints to generate a set of credible candidate resources, as follows: Trusted constraint rule set built based on trusted data space Candidate resources in the full-element resource map are screened, and resource combinations that do not meet the constraints of data out-of-domain, computing power security domain, access permission, and compliance policy are eliminated to generate a set of trustworthy candidate resources. .

[0047] In one specific implementation, the supply and demand matching objective function is a weighted sum of supply and demand distance, correlation, scheduling cost, and trust constraint penalty term.

[0048] Specifically, the supply and demand matching objective function is expressed as: ; in, This represents the user's request vector. Represents a vector of credible candidate resource combinations. Indicates the distance between supply and demand. This indicates the degree of correlation between a user's request and a combination of trusted candidate resources. Indicates scheduling cost, This represents a credible constraint penalty term. , , , This represents the weighting coefficient.

[0049] Furthermore, the supply-demand distance is a weighted Euclidean distance in a multi-dimensional feature space, which comprehensively measures the degree of deviation between user demand requests and candidate resource combinations in four dimensions: computing power specifications, algorithm capabilities, data coverage, and service quality. It is expressed as: ; in, This indicates the computing power specifications requested by the user. This indicates the specifications of computing power in a portfolio of trusted candidate resources. This indicates the algorithmic capabilities required by the user's request. This indicates the capability of algorithm instances within a credible candidate resource combination. This indicates the data coverage of the user's request. This indicates the coverage of available datasets in a combination of trustworthy candidate resources. This represents a service quality indicator that reflects the user's request. The estimated service quality index is calculated based on the service quality statistics of similar historical tasks. δ1, δ2, δ3, and δ4 represent the weighting coefficients of each distance, and δ1+δ2+δ3+δ4=1.

[0050] Furthermore, the correlation between user demand requests and trusted candidate resource combinations is calculated by aggregating resource correlation in the full-element resource map, and is expressed as: ; in, This represents the i-th resource feature node parsed from user request X. A resource in data, computing power, algorithms, or applications. This represents the total number of nodes in the full-element resource map corresponding to the resources requested in user request X. This represents the j-th candidate resource node in the trusted candidate resource combination Y. This represents the total number of candidate resource nodes in the trusted candidate resource combination Y. Represents the demand resource characteristic nodes in the total factor resource map. With candidate resource nodes The degree of resource correlation between them This represents the resource characteristics required for the i-th resource. Find the resource node with the highest resource correlation among all candidate resources of Y.

[0051] Furthermore, the scheduling cost, which comprises the computing power usage cost, data transmission cost, algorithm scheduling cost, and security compliance cost incurred by candidate resource combination Y in executing user demand requests, is expressed as: ; in, Indicates the cost of using computing power. , Traverse candidate resource combinations All computing nodes, For computing power nodes The unit time price (yuan / core·hour). For computing power nodes Execute user request Estimated duration of use (in hours). This represents the estimated resource utilization rate of the node. Indicates data transmission cost. , For user requests The data set that needs to be called For dataset To computing nodes The amount of data (GB) that needs to be transferred between them. The available bandwidth of the transmission link (GB / s). The data transfer rate is per unit (RMB / GB), when the dataset is already cached on the target node. This encourages data affinity scheduling; Indicates the cost of algorithm invocation. , Traverse candidate resource combinations All algorithm instances included. Algorithm Instance The cost per call license (RMB / call). To request based on user needs Prediction algorithm Number of calls; for open-source or self-developed algorithms, ; Indicates the cost of security compliance. , In candidate resource combinations computing nodes in The above is a request for user needs. The datasets involved The computational cost of performing security and compliance processing (such as data desensitization, differential privacy noise addition, encrypted transmission, etc.) (converted to yuan based on equivalent computing power time). For dataset Security level weighting (the higher the security level, the more complex and costly the compliance process). , , , The weighting coefficients for each cost item satisfy the following conditions: .

[0052] Furthermore, the credible constraint penalty term is expressed as: ; in, This indicates the highest security level of the data involved in user request X. For example, if the security level of video surveillance data is L4, then... =L4, This represents the security authentication level of the i-th computing node in the trusted candidate resource combination Y, i.e., the security level authenticated when the computing node registers in the trusted data space. For example, the security authentication level of the private network GPU cluster A is L4. This represents the penalty coefficient.

[0053] The above technical solution, based on trusted candidate resources, considers supply-demand distance, correlation, scheduling cost, and trusted constraint penalty terms, and performs solution matching under trusted constraints, thereby improving the accuracy and rationality of supply-demand matching solution generation.

[0054] In step S13, the matching scheme is optimized using the scheduling optimization objective function, and the optimal scheduling instruction is generated by combining scheduling strategies such as elastic scaling scheduling, data affinity scheduling, load balancing scheduling, and algorithm-based selection scheduling.

[0055] Existing scheduling optimization schemes only consider resource utilization and execution efficiency, lacking optimization in terms of trustworthiness. This invention optimizes scheduling from the aspects of resource utilization, execution efficiency, latency, and trustworthiness scheduling score. It also designs a variety of scheduling strategies such as elastic scaling scheduling, data affinity scheduling, load balancing scheduling, and algorithm-based optimization scheduling, which improves the effect of scheduling optimization, makes the task execution process highly secure, and maintains high resource utilization and execution performance even in complex application scenarios.

[0056] In one specific implementation, the scheduling optimization objective function is a weighted sum of resource utilization, execution efficiency, latency, and reliable scheduling score.

[0057] Specifically, the scheduling optimization objective function is expressed as: ; in, Indicates resource utilization rate, Indicates execution efficiency. Indicates a delay. Indicates the trusted scheduling score. , , , This indicates the weights to be optimized.

[0058] Furthermore, resource utilization rate is expressed as: ; Where M represents the total number of computing nodes participating in the scheduling. This represents the amount of resources used by computing node j. This represents the total resource amount of computing node j.

[0059] Execution efficiency is expressed as: ; in, Indicates the baseline execution time (based on statistics of similar historical tasks). Indicates the actual execution time. This represents the number of subtasks that were successfully completed. This represents the total number of subtasks.

[0060] The trusted scheduling score is represented as: ; in, This represents the security authentication level of the i-th computing power node in the trusted candidate resource combination Y, that is, the security level authenticated when the computing power node registers in the trusted data space. For example, the security authentication level of private network GPU cluster A is L4, and the security authentication level of private network GPU cluster B is L3. This indicates the highest security level defined in the system. This indicates a compliance instruction function.

[0061] In one specific implementation, the trusted scheduling score is calculated using the computing power security certification level, the highest security level, and the resource compliance indication function; The resource compliance indicator function takes the value 1 when the resource satisfies all the rules in the trusted constraint rule set; otherwise, it takes the value 0.

[0062] In this embodiment of the invention, elastic scaling scheduling is used to monitor the load rate of computing nodes in real time. ,when (Upper limit threshold, default 80%), and consistently exceeding When the alarm time window (default 30 seconds) is reached, a capacity expansion operation is triggered. New computing power nodes are selected from the trusted resource pool and added to the scheduler according to the security level matching principle. The new nodes must meet the trusted constraint rule set. The computational power security domain constraint in the context; when (Lower threshold, default 20%), and consistently exceeding When the release time window (default 120 seconds) is reached, a scaling-down operation is triggered, prioritizing the release of the node with the highest trust level redundancy.

[0063] Data affinity scheduling is used to compute each candidate computing node. Tasks to be scheduled Data affinity between : ; in, For task T i The required data set The data set already cached on computing node j; prioritize scheduling tasks to The node with the highest value is selected to reduce cross-node data transmission overhead; based on this, a data out-of-domain constraint in the trusted space is superimposed: when When the data contains high-security-level data that is restricted from leaving the domain, scheduling is only allowed to nodes within the security domain that meet the out-of-domain constraints.

[0064] Load balancing scheduling is based on the real-time load rate of each computing node. and remaining resources The calculation yields the following result, which is expressed as: ; in, This represents the load rate of computing node j. This represents the current remaining available resources of computing node j (including the number of idle GPU / CPU cores, available memory capacity, available storage bandwidth, etc., calculated in equivalent computing units). This represents the total resource quantity of computing node j (i.e., the total amount of all hardware resources of computing node j after being converted into equivalent computing power units); tasks are prioritized for scheduling to The highest-ranking node ensures that the overall load of the cluster is evenly distributed.

[0065] Algorithm-based optimal scheduling, when multiple algorithm instances can satisfy the same task requirements, calculates the algorithm's optimal score based on the algorithm's historical performance metrics (accuracy (Acc), throughput (Thr), and resource consumption (Rsc)). : ; Where 'a' represents a candidate algorithm instance (i.e., one of multiple candidate algorithms that meet the same task requirements, such as the Tongyi Thousand Questions instance, the local fine-tuning model instance, etc.). This represents the accuracy of candidate algorithm instance a. This represents the throughput of candidate algorithm instance a. This represents the resource consumption of candidate algorithm instance a. , , For optimal weighting; select The highest-level algorithm instance performs the task.

[0066] It should be noted that, among the scheduling strategies mentioned above, elastic scaling scheduling, data affinity scheduling, load balancing scheduling, and algorithm-based optimal scheduling are not necessarily all used during task execution. One or more of these scheduling strategies can be used depending on the requirements.

[0067] The optimal scheduling instruction is generated based on the above scheduling strategy. Before being issued, the scheduling instruction needs to be verified for compliance through the secure approval interface of the trusted data space. After the verification is passed, the scheduling instruction is issued to the corresponding resource node through the trusted interface to start task execution. During task execution, resource usage status and task execution progress data are collected at random time intervals of t∈[1,20] seconds to dynamically adjust the scheduling strategy.

[0068] The above technical solution optimizes the matching scheme using the scheduling optimization objective function, and combines scheduling strategies such as elastic scaling scheduling, data affinity scheduling, load balancing scheduling, and algorithm-based optimization scheduling to generate the optimal scheduling instruction. This improves the scheduling optimization effect, makes the task execution process more secure, and maintains high resource utilization and execution efficiency even in complex application scenarios.

[0069] In step S14, real-time monitoring is performed during task execution, and the entire process operation log and resource usage record are recorded. After the task is completed, the reinforcement learning reward value is calculated, and when the triggering condition is met, the parameters of the supply and demand matching objective function and the scheduling optimization objective function are iteratively optimized. The operation log and resource usage record are fed back to the trusted data space, and the trusted constraint rule set is dynamically adjusted.

[0070] To address the issue that existing methods only optimize supply and demand matching schemes, resulting in low accuracy in scheduling across all elements including data, computing power, algorithms, and applications, this invention designs a reinforcement learning reward value. This not only optimizes the parameters of the supply and demand matching objective function but also the parameters of the scheduling optimization objective function and the set of trustworthy constraint rules. This meets the differentiated scheduling needs of multiple industries, improves scheduling accuracy, and fully releases the value of data elements.

[0071] In one specific implementation, the reinforcement learning reward value is obtained by weighted summation of user satisfaction, execution performance, and error rate.

[0072] Specifically, the reinforcement learning reward value is represented as: ; in, Indicates user satisfaction. Indicates execution performance. Indicates the error rate. , , This represents the return weighting coefficient.

[0073] Furthermore, user satisfaction is expressed as: ; in, Indicates the actual service quality. This indicates the target service quality requested by the user. Indicates user ratings, Indicates the user's waiting time. This indicates the user's maximum tolerable waiting time. , , Indicates the weighting coefficient. + + =1.

[0074] Execution performance is expressed as: ; in, Indicates the base time. Indicates the actual execution time. Indicates resource utilization rate, Indicates actual throughput. Indicates the expected throughput. , , This represents the weighting coefficient.

[0075] Error rate is expressed as: ; in, This indicates the number of subtasks that failed to execute. This indicates the number of subtasks that timed out. This represents the number of subtasks that violate trust constraints. This represents the total number of subtasks.

[0076] Understandably, when the return value continuous Number of sampling periods (default) (Below the preset threshold) (Right now continued (per cycle), or single return value The decline exceeded (Right now This triggers parameter updates for the matching model and scheduling strategy.

[0077] When the return value In continuous One optimization cycle (default) Fluctuation range within ) Less than the convergence threshold (default ),and The mean is higher than the target threshold When the model is deemed converged, the current iteration of optimization is stopped.

[0078] The following example, using a multi-source data fusion and analysis scenario in urban emergency management, illustrates the specific implementation of this invention.

[0079] I. Construction of Trusted Data Space Establish a trusted data space architecture that includes a basic resource layer, a security layer, and an element catalog layer.

[0080] Basic resource layer: integrates distributed storage clusters, heterogeneous computing power clusters (including dedicated network GPU server cluster A with 4 units equipped with PPUs, general computing power cluster B with 8 units equipped with 910B GPUs, and CPU cluster C with 16 general servers) and a diverse algorithm repository (including 12 types of algorithms such as video AI analysis algorithms, weather prediction models, traffic simulation algorithms, and large model judgment algorithms).

[0081] Security layer: Deploy a data anonymization engine (according to the ε-differential privacy mechanism (ε≤1.0, δ≤10)). 5 The system includes: adding Laplace noise, access control module, and operation auditing system. Specifically, video surveillance data is marked as security level L4 (highest level), with an out-of-domain constraint rule set to "only allowed to be calculated within private network cluster A"; traffic checkpoint data is marked as security level L3, allowing calculation within clusters A and B; weather forecast data is marked as security level L2, allowing cross-domain scheduling after anonymization; and geographic information basic data is marked as security level L1, with no out-of-domain restrictions.

[0082] The element catalog layer establishes four major categories of element catalogs: data, computing power, algorithms, and applications. The application element catalog includes templates such as "Urban Flood Early Warning Application Template," "Traffic Management and Analysis Application Template," and "Emergency Command Large Screen Application Template." Regarding the correlation of computing resources, for example, when calculating the correlation between the "Rainstorm Flood Prediction Algorithm" and the "Weather Forecast Dataset," data format compatibility is considered. (The algorithm supports NetCDF and CSV formats, and the dataset is provided in NetCDF format.) Computing power architecture adaptability. (The algorithm supports GPU acceleration, and cluster B has a GPU), security level consistency. (Data security level L2, cluster B security authentication L3), application scenario matching degree (The algorithm is highly correlated with the "Urban Flood Early Warning Application Template") After weighted calculation , marked as strong association.

[0083] Construction of a comprehensive resource map: Based on the four types of resource elements registered in the aforementioned trusted data space—data, computing power, algorithms, and applications—each resource is treated as a node. The resource correlation degree between any two resource nodes is calculated. When the correlation degree exceeds a preset threshold, a correlation edge is established and labeled with the correlation strength level (strong correlation ≥ 0.8, medium correlation ≥ 0.6), forming a comprehensive resource map. For example, the resource correlation degree between the "Rainstorm and Urban Flooding Prediction Algorithm" node and the "Meteorological Forecast Dataset" node is 0.85, marked as strong correlation; the correlation degree with the "Urban GIS Data" node is 0.72, marked as medium correlation. This map provides a structured basis for resource relationship calculations and demand mapping in subsequent supply and demand matching.

[0084] Constructing a set of trusted constraint rules: Based on data security classification results, access control policies, and compliance audit requirements, a set of trusted constraint rules is generated. The constraints include: (1) Data outbound constraint, L4 data is limited to cluster A; (2) Computing power security domain constraint, cluster A only accepts tasks from authorized users; (3) Access permission constraint, cross-department data access requires approval; (4) Compliance policy constraint, all processing involving personal privacy data must meet differential privacy requirements.

[0085] II. Intelligent matching of supply and demand A city issued a rainstorm warning, and the emergency management department submitted an emergency assessment request through the platform: "Based on multi-source data (video surveillance, traffic checkpoint data, real-time meteorological data, and urban geographic information), conduct urban flooding risk assessment, deploy video AI water level detection algorithms, meteorological prediction models, and traffic simulation models, and use large models for comprehensive assessment to assist decision-making, output risk levels and evacuation plans, and require a response delay of no more than 60 seconds and an accuracy rate of no less than 85%."

[0086] Demand parsing and mapping to the full-element resource map: After receiving the above demand requests, the system extracts structured demand features through scene recognition algorithms, including data requirements (video surveillance, traffic checkpoint data, meteorological data, GIS data), computing power requirements (GPU acceleration), algorithm requirements (video AI, weather forecasting, traffic simulation, large model analysis), service quality indicators (latency ≤ 60 seconds, accuracy ≥ 85%), etc., and maps each demand feature to the corresponding node in the full-element resource map, and retrieves matching candidate resources along the associated edges to form a candidate resource set.

[0087] Trusted constraint pre-filtering: First, load the trusted constraint rule set. Upon detecting that the requirement involved video surveillance data (L4 security level), clusters B and C, which did not meet the out-of-domain constraints, were immediately removed from the candidate computing power resources for the "Video AI Water Level Detection Algorithm," retaining only the private network cluster A as a trusted candidate computing power node for the algorithm. The processing candidate scope for traffic checkpoint data (L3) was limited to clusters A and B, while meteorological data (L2), after being anonymized, could be processed by all three clusters.

[0088] Demand characteristics are analyzed using scene recognition algorithms, and a matching scheme incorporating data, computing power, algorithms, and applications is generated through a supply-demand matching objective function. This includes a trust constraint penalty term. Ensure that all resource combinations that do not meet the security constraints receive extremely high penalty values ​​and are excluded. The final matching result is:

[0089] III. Dynamic Scheduling Execution Based on the matching scheme and combined with the real-time resource status (cluster A GPU load rate 28%, cluster B GPU load rate 45%, cluster C CPU load rate 15%, all algorithms available, dataset ready), the optimal scheduling strategy is solved by optimizing the scheduling objective function.

[0090] Step 1: Trusted Security Verification: After the scheduling instruction is generated, it is first verified through the trusted data space's security approval interface to confirm that the video AI water level detection task has indeed been scheduled to cluster A (meeting L4 data out-of-domain constraints), that the meteorological data has been anonymized (meeting compliance policies), and that cross-departmental data access permissions have been approved (meeting permission constraints). The scheduling instruction is issued after successful verification.

[0091] Step 2: Data Affinity Scheduling: Calculate the data affinity of each task. The traffic simulation model needs to use both traffic checkpoint data and GIS data simultaneously. Cluster B has already cached a portion of the GIS data fragments; the affinity... Cluster A has not yet cached GIS data, resulting in low affinity. However, the security level of traffic checkpoint data is L3, allowing computation on clusters A and B. Considering both data affinity and trust constraints, the traffic simulation model is scheduled to cluster B.

[0092] Step 3: Algorithm Optimization and Scheduling: For the large-scale model comprehensive evaluation task, there are two candidate algorithm instances (the General 1000 Questions instance and the Local Fine-tuning Model instance). Based on historical performance (General 1000 Questions accuracy 88%, throughput 12 qps, high resource consumption; Local Model accuracy 91%, throughput 8 qps, medium resource consumption), the algorithm optimization score is calculated, and the Local Fine-tuning Model instance is selected for execution.

[0093] Step 4: Task Execution and Elastic Scaling Scheduling: Collect computing power utilization and algorithm execution progress data at random 5-second intervals. When it is detected that the GPU load rate of cluster A rises to 85% due to a sudden increase in video streaming (10 streams → 50 streams) during a rainstorm, exceeding... After 30 seconds, elastic scaling scheduling is automatically triggered. At this time, a backup GPU node that meets the L4 security domain constraint is searched from the trusted resource pool, and two backup servers in cluster A are added to the schedule, reducing the load rate to 52% to ensure uninterrupted video AI analysis.

[0094] IV. Full-process monitoring and closed-loop optimization The task execution process was monitored in real time, and various indicators were recorded: Cluster A computing power usage time was 6.2 hours, video AI water level detection accuracy was 93%, meteorological prediction model accuracy was 89%, traffic simulation deviation rate was 6%, data transmission was encrypted throughout the process with no security incidents, and the overall accuracy of large model analysis was 91%.

[0095] After the task is completed, collect user feedback: Emergency Management Department Rating (Normalized), actual response delay 42 seconds (target 60 seconds), overall judgment accuracy 91% (target 85%).

[0096] Calculate user satisfaction: .

[0097] Computational execution performance: baseline execution time The average execution time for similar emergency assessment tasks in the past was 5.0 hours, and the actual execution time was [missing data]. It takes 6.2 hours, which is efficient in terms of time. / =5.0 / 6.2≈0.806; Resource utilization rate The weighted average utilization rate of each participating cluster is calculated as follows: Cluster A has an average utilization rate of 65%, Cluster B has an average utilization rate of 60%, and Cluster C has an average utilization rate of 35%. ≈0.82 (weighted by the number of subtasks undertaken by each cluster); actual throughput =14 qps, expected throughput =15 qps, throughput ratio / =14 / 15≈0.933, execution performance =0.35×0.806+0.30×0.82+0.35×0.933=0.282+0.246+0.327=0.855.

[0098] Calculate the error rate There were a total of 28 sub-tasks, with 2 timed out, 0 failed, and 0 security violations. .

[0099] Calculate the return value: .

[0100] Higher than the preset threshold The current cycle does not trigger parameter updates. However, the execution data for this cycle is recorded, and when accumulated... Convergence is determined after one cycle.

[0101] Trusted space feedback enhancement: The task execution log, including scheduling path, data flow record, and security event statistics, is sent back to the trusted data space to update the following: (1) Elastic scaling performance evaluation data of cluster A in the rainstorm scenario, enriching the trusted behavior profile; (2) Accuracy distribution of video AI algorithm under rainstorm and low visibility conditions, updating the algorithm capability profile; (3) Confirming that there are no security events in this cross-departmental data collaboration process, enhancing the credibility rating of relevant data sharing strategies. Thus, a complete closed loop of "trusted space → scheduling decision → execution feedback → trusted space enhancement" is completed.

[0102] Example 2 This embodiment provides an electronic device, including a processor and a memory, wherein the memory is used to store a computer program, which is loaded and executed by the processor to implement the data, computing power, algorithm and application fusion scheduling method as described in Embodiment 1.

[0103] Example 3 This embodiment provides a computer-readable storage medium for storing a computer program; wherein the computer program, when executed by a processor, implements the data, computing power, algorithm, and application fusion scheduling method as described in Embodiment 1.

[0104] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A method for the integrated scheduling of data, computing power, algorithms, and applications, characterized in that, include: Construct a trusted data space, a full-element resource map, and a trusted set of constraint rules; The set of trusted constraint rules includes data out-of-domain constraints, computing power security domain constraints, access permission constraints, and compliance policy constraints. Receive user request requests, parse the request requests and map them to a full-element resource map, and perform trust constraint screening on candidate resources in the full-element resource map based on a trust constraint rule set to generate a trust candidate resource set; Within a set of trusted candidate resources, a matching scheme containing data, computing power, algorithms, and applications is generated through a supply and demand matching objective function. The matching scheme is optimized using the scheduling optimization objective function, and the optimal scheduling instruction is generated by combining scheduling strategies such as elastic scaling scheduling, data affinity scheduling, load balancing scheduling and algorithm-based selection scheduling. During task execution, real-time monitoring is performed, and the entire process operation log and resource usage record are recorded. After the task is completed, the reinforcement learning reward value is calculated, and when the triggering conditions are met, the parameters of the supply and demand matching objective function and the scheduling optimization objective function are iteratively optimized. The operation log and resource usage record are fed back to the trusted data space, and the trusted constraint rule set is dynamically adjusted.

2. The data, computing power, algorithm, and application fusion scheduling method as described in claim 1, characterized in that, The trusted data space includes secure data storage, fine-grained access control, and compliance audit traceability.

3. The fusion scheduling method for data, computing power, algorithms, and applications as described in claim 1, characterized in that, The full-element resource map includes nodes and resource correlation. The nodes include data, computing power, algorithms, and applications. The resource correlation includes data format compatibility, computing power architecture adaptability, security level consistency, and application scenario matching between resources.

4. The fusion scheduling method for data, computing power, algorithms, and applications as described in claim 1, characterized in that, The user request includes business scenario type, computing power specifications, algorithm function requirements, data access permissions, security level requirements, and service quality indicators.

5. The fusion scheduling method for data, computing power, algorithms, and applications as described in claim 1, characterized in that, The objective function for supply and demand matching is a weighted sum of supply and demand distance, correlation, scheduling cost, and credibility constraint penalty term.

6. The fusion scheduling method for data, computing power, algorithms, and applications as described in claim 1, characterized in that, The scheduling optimization objective function is a weighted sum of resource utilization, execution efficiency, latency, and reliable scheduling score.

7. The data, computing power, algorithm, and application fusion scheduling method as described in claim 5, characterized in that, The trusted scheduling score is calculated using the computing power security certification level, the highest security level, and the resource compliance indicator function. The resource compliance indicator function takes the value 1 when the resource satisfies all the rules in the trusted constraint rule set; otherwise, it takes the value 0.

8. The fusion scheduling method for data, computing power, algorithms, and applications as described in claim 1, characterized in that, The reinforcement learning reward value is obtained by weighted summation of user satisfaction, execution performance, and error rate.

9. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory is used to store a computer program, which is loaded and executed by the processor to implement the full-element fusion scheduling method of data, computing power, algorithms and applications as described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, Used to store computer programs; wherein the computer programs, when executed by a processor, implement the full-element fusion scheduling method of data, computing power, algorithms and applications as described in any one of claims 1-8.