Resource allocation method and system based on edge cloud, electronic equipment and storage medium

By collecting, analyzing, and monitoring the dynamic adjustment of computing resources in the edge cloud environment, the problem of lack of adaptive allocation of edge cloud computing resources is solved, and efficient resource utilization and precise adaptation are achieved.

CN121008904APending Publication Date: 2025-11-25BEIJING CHINA POWER INFORMATION TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510880153.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

In existing technologies, the allocation of computing resources in edge cloud environments lacks adaptive capabilities and cannot respond to changes in user behavior, fluctuations in device performance, and dynamic adjustment needs of application load, resulting in low resource utilization efficiency or performance degradation.

Method used

The data acquisition module collects environmental information, the data analysis module analyzes and processes the data to determine the initial resource allocation strategy and optimize it into the target strategy, the resource allocation module allocates computing resources according to the target strategy, and the monitoring and feedback module monitors and adjusts the allocation strategy in real time, forming a dynamic optimization mechanism.

Benefits of technology

It enables adaptive resource allocation to the dynamic environment of edge cloud, improving the global utilization of computing resources and the ability to accurately adapt to needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121008904A_ABST
    Figure CN121008904A_ABST
Patent Text Reader

Abstract

The invention provides a resource allocation method and system based on edge cloud, electronic equipment and a storage medium. The method comprises the steps that a data acquisition module acquires environment information in an edge cloud environment; wherein the environment information comprises application load information, equipment performance information and user behavior information; the data analysis module analyzes and processes the environment information to obtain an application load prediction result, an equipment capability evaluation result and a behavior analysis result, and determines an initial resource allocation strategy according to the application load prediction result, the equipment capability evaluation result and the behavior analysis result; optimizing the initial resource allocation strategy to obtain a target resource allocation strategy; the resource allocation module allocates the computing resources according to the target resource allocation strategy; and the monitoring feedback module monitors the use state of the computing resources in real time and feeds back the use state to the data analysis module, so that the data analysis module adjusts the target resource allocation strategy according to the use state.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of data processing, and particularly relates to a resource allocation method and system based on edge cloud, an electronic device and a storage medium. BACKGROUND

[0002] With the development of Internet of Things technology, a large number of devices access the network to develop an edge cloud distributed computing mode. In the edge cloud environment, the management of computing resources faces many challenges. However, the current static resource allocation strategy is generally used to allocate computing resources. The core mechanism of the static resource allocation strategy is to perform offline configuration of resources based on a fixed proportion set in advance, which leads to a lack of adaptive ability of the system to the dynamic environment of the edge cloud.

[0003] Therefore, how to realize the adaptation to the dynamic environment of the edge cloud when allocating computing resources has become a technical problem to be solved. SUMMARY

[0004] Therefore, the purpose of the present disclosure is to provide a resource allocation method and system based on edge cloud, an electronic device and a storage medium to solve or partially solve the above technical problems.

[0005] To achieve the above purpose, a first aspect of the present disclosure provides a resource allocation method based on edge cloud, applied to a resource allocation system based on edge cloud, the system comprising a data collection module, a data analysis module, a resource allocation module and a monitoring feedback module. The method comprises:

[0006] The data collection module collects environment information in the edge cloud environment, wherein the environment information comprises application load information, device performance information and user behavior information.

[0007] The data analysis module analyzes and processes the environment information to obtain application load prediction results, device capability evaluation results and behavior analysis results, determines an initial resource allocation strategy according to the application load prediction results, the device capability evaluation results and the behavior analysis results, and optimizes the initial resource allocation strategy to obtain a target resource allocation strategy.

[0008] The resource allocation module allocates computing resources according to the target resource allocation strategy.

[0009] The monitoring feedback module monitors the use state of the computing resources in real time, and feeds back the use state to the data analysis module, so that the data analysis module adjusts the target resource allocation strategy according to the use state.

[0010] Based on the same inventive concept, a second aspect of the present disclosure provides an edge cloud-based resource allocation system, the system comprising:

[0011] The data collection module is configured to collect environment information in the edge cloud environment; wherein the environment information comprises application load information, device performance information and user behavior information;

[0012] The data analysis module is configured to analyze and process the environment information to obtain application load prediction results, device capability evaluation results and behavior analysis results, and determine an initial resource allocation strategy according to the application load prediction results, the device capability evaluation results and the behavior analysis results, and optimize the initial resource allocation strategy to obtain a target resource allocation strategy;

[0013] The resource allocation module is configured to allocate computing resources according to the target resource allocation strategy;

[0014] The monitoring feedback module is configured to monitor the use state of the computing resources in real time, and feed back the use state to the data analysis module, so that the data analysis module adjusts the target resource allocation strategy according to the use state.

[0015] Based on the same inventive concept, a third aspect of the present disclosure provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable by the processor, wherein the processor implements the method described above when executing the computer program.

[0016] Based on the same inventive concept, a fourth aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described above.

[0017] It can be seen from the above that the edge cloud-based resource allocation method, system, electronic device and storage medium provided by the present disclosure. The data acquisition module acquires environmental information in the edge cloud environment; wherein the environmental information includes: application load information, device performance information and user behavior information. In this way, the data acquisition module can construct a three-dimensional perception system of applications, devices and users based on multi-modal data fusion, and improve the matching accuracy of resource supply and demand. The data analysis module analyzes and processes the environmental information to obtain application load prediction results, device capability evaluation results and behavior analysis results, and determines an initial resource allocation strategy according to the application load prediction results, the device capability evaluation results and the behavior analysis results, and optimizes the initial resource allocation strategy to obtain a target resource allocation strategy. In this way, the data analysis module can automatically optimize the target resource allocation strategy of the computing resources according to the environmental information, maximize the global utilization rate, and realize the precise adaptation of the computing resources on demand. The resource allocation module allocates computing resources according to the target resource allocation strategy. The monitoring feedback module monitors the use state of the computing resources in real time, and feeds back the use state to the data analysis module, so that the data analysis module adjusts the target resource allocation strategy according to the use state. In this way, the monitoring feedback module can feed back the use state of the computing resources monitored in real time to the data analysis module, and the data analysis module can adjust the target resource allocation strategy in real time and accurately according to the use state of the computing resources in real time. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present disclosure or the related art, the following will briefly introduce the drawings needed to be used in the embodiments or related art descriptions. Obviously, the drawings in the following description are only embodiments of the present disclosure, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0019] Figure 1 Flowchart of the edge cloud-based resource allocation method of the embodiment of the present disclosure;

[0020] Figure 2 Structure diagram of the edge cloud-based resource allocation system of the embodiment of the present disclosure;

[0021] Figure 3 Structure diagram of the edge cloud-based resource allocation system of another embodiment of the present disclosure;

[0022] Figure 4 Structure diagram of the data acquisition module of the embodiment of the present disclosure;

[0023] Figure 5 Structure diagram of the data analysis module of the embodiment of the present disclosure;

[0024] Figure 6Structure diagram of a resource allocation module of an embodiment of the present disclosure;

[0025] Figure 7 Structure diagram of a monitoring feedback module of an embodiment of the present disclosure;

[0026] Figure 8 Flow chart of a pre-allocation process of computing resources of an embodiment of the present disclosure;

[0027] Figure 9 Flow chart of a device overload triggered dynamic scheduling process of an embodiment of the present disclosure;

[0028] Figure 10 Flow chart of a user behavior optimized cache strategy process of an embodiment of the present disclosure;

[0029] Figure 11 Flow chart of a multi-modal data collaborative sampling process of an embodiment of the present disclosure;

[0030] Figure 12 Flow chart of a demand-device collaborative evaluation process of an embodiment of the present disclosure;

[0031] Figure 13 Flow chart of an anomaly detection triggered self-healing closed loop process of an embodiment of the present disclosure;

[0032] Figure 14 Flow chart of a simulation verification process of an embodiment of the present disclosure;

[0033] Figure 15 Flow chart of an energy efficiency optimized driven load migration process of an embodiment of the present disclosure;

[0034] Figure 16 Structure diagram of an electronic device of an embodiment of the present disclosure. DETAILED DESCRIPTION

[0035] In order to make the objectives, technical solutions and advantages of the present disclosure clearer, the present disclosure is further described in detail below with reference to specific embodiments and with reference to the accompanying drawings.

[0036] It should be noted that, unless otherwise defined, technical terms or scientific terms used in the embodiments of the disclosure shall be understood as the common meaning understood by one of ordinary skill in the art to which the embodiments of the disclosure belong. The terms "first", "second" and similar words used in the embodiments of the disclosure do not represent any order, number or importance, but are only used to distinguish different components. The terms "include" or "contain" and similar words mean that the elements or objects before the words cover the elements or objects listed after the words and their equivalents, and do not exclude other elements or objects. The terms "connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to represent relative positional relationships, and when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0037] Based on the description of the background, under the driving of the Internet of Things technology, the massive device access to the network has given rise to the edge cloud, a distributed computing mode, which reduces the latency by sinking computing resources to the vicinity of the data source. However, intelligent home, industrial Internet of Things and other different Internet of Things applications have significantly different demands for computing resources, and edge devices are diverse and heterogeneous in performance, so traditional cloud computing resource management methods are difficult to adapt to their distributed and resource-constrained characteristics. Current approximate solutions mostly use static resource allocation strategies, that is, a fixed resource proportion is set for different application types in advance. Although this is simple to implement, it cannot respond to the dynamic adjustment needs of user behavior changes, device performance fluctuations and application loads in actual operation, such as being unable to cope with traffic peaks in video monitoring scenarios or device performance degradation, resulting in low resource utilization efficiency or application performance degradation, and there is an urgent need for an intelligent management scheme that can adaptively perceive environmental changes and dynamically optimize resource allocation.

[0038] The related art generally uses a static resource allocation strategy, the core mechanism of which is to perform offline resource configuration based on a pre-set fixed proportion, resulting in a lack of adaptive ability of the system to the dynamic environment of the edge cloud. Specifically, this strategy does not establish a real-time response mechanism for the dynamic nature of user behavior, the time-varying nature of device performance and the scenario-based application demand, and when the user access mode fluctuates, the edge node performance degrades or the application load characteristics change, the resources allocated by the fixed proportion cannot adapt to the actual demand as needed, thereby causing performance degradation due to resource competition in high-load scenarios, or energy efficiency loss due to resource idling in low-load scenarios. This static allocation paradigm has become a key bottleneck restricting the resource utilization rate and service reliability of the edge cloud system.

[0039] In the edge cloud environment, the management of computing resources faces many challenges. With the continuous expansion of application scenarios of edge cloud services, such as Internet of Things device access and real-time data processing, different types of applications have different demands for computing resources. At the same time, the performance and computing resource capacity of edge cloud devices vary greatly. How to efficiently manage computing resources in such a complex environment to meet the needs of various applications and improve the overall utilization of computing resources is a technical problem to be solved by the present invention.

[0040] The terms involved in the present disclosure are explained as follows:

[0041] Multimodal Technology (MMT): an important branch of artificial intelligence, focusing on integrating and processing multiple modalities of data, supporting the processing of text, images, audio, video, sensor data, etc., to improve the understanding, reasoning and generation capabilities of models.

[0042] Adaptive Sampling (AS): a technology that dynamically adjusts the sampling strategy, aiming to optimize sampling efficiency and resource utilization according to current data distribution, task requirements or system state. Adaptive sampling is widely used in machine learning, signal processing, sensor networks, reinforcement learning, etc., with the core goal of reducing redundant sampling, improving information acquisition quality or reducing computing / storage overhead.

[0043] Quality of Service (QoS): a measure of network differentiated guarantee capability for data flow, ensuring real-time audio / video, industrial control and other critical businesses to obtain deterministic guarantee of delay, jitter and packet loss rate in congested environment through priority scheduling, bandwidth reservation and congestion control mechanisms. Core implementation includes: traffic classification and marking, queue scheduling algorithm, congestion management, etc.

[0044] Knowledge Graph (KG) and Requirement Knowledge Graph (RKG): Knowledge Graph is the core technology carrier of "knowledge representation and reasoning", while Requirement Knowledge Graph is its deep application in demand analysis and resource management scenarios. Both of them solve semantic understanding and decision support problems in complex systems through structured knowledge modeling and intelligent reasoning. In technical documents, Requirement Knowledge Graph is often combined with machine learning (demand extraction) and reinforcement learning (decision optimization) to form a complete link of "data collection → knowledge modeling → intelligent decision-making".

[0045] Device Fingerprint Database (DFDB): A database system used to store and manage device fingerprint information. Device fingerprint is a unique identifier or feature set generated by collecting and analyzing multi-dimensional characteristics of device hardware, software, and network environment, used to identify and distinguish different devices. DFDB centrally stores and manages these device fingerprint information and their related attributes (such as device type, operating system, browser version, IP address, etc.) for subsequent comparison, query and analysis.

[0046] Spatiotemporal Graph Neural Network (ST-GNN): A method that combines spatiotemporal information and graph neural networks, capturing the dependencies of nodes in time and space, modeling and analyzing data with spatiotemporal characteristics (such as traffic flow, video frame sequences), suitable for dynamic system modeling, spatiotemporal prediction and other tasks.

[0047] Lightweight AI Agent (LAIA): A low-resource consumption and high-efficiency intelligent agent system, designed to implement intelligent task processing in devices with limited computing power (such as mobile devices, embedded devices, edge computing nodes) or scenarios with high real-time requirements. This type of agent usually simplifies model structure, optimizes algorithm efficiency or uses lightweight framework to achieve a balance between high performance and low energy consumption.

[0048] Computer Vision Analysis (CVA): Using computer technology to process, analyze and understand images or videos, extracting visual features, recognizing objects, detecting scenes or behaviors. Widely used in security monitoring, autonomous driving, medical image diagnosis and other fields, it is an important branch of artificial intelligence.

[0049] Operational Temporal Pattern Graph (OTPG): Records the pattern of operation sequence over time through a graph structure, revealing the temporal dependency and rules between operations. Can be used for user behavior analysis, system operation log mining, and auxiliary anomaly detection, process optimization and predictive maintenance.

[0050] Conditional Generative Adversarial Network (CGAN): Based on Generative Adversarial Network (GAN), introduces conditional constraints, makes the generator generate data according to the input conditions. Suitable for image generation, data augmentation, etc. Can generate images of specific styles or simulate data distribution under specific conditions.

[0051] Differential Privacy Mechanism (DPM): By adding controllable noise to data, individual information is hidden while overall statistical characteristics are preserved, preventing data leakage and privacy infringement. Suitable for sensitive data publishing, machine learning training, etc. To ensure the balance between data availability and privacy protection.

[0052] Automated Feature Engineering (AFE): Use algorithms to automatically extract, select and transform features from raw data, reduce manual intervention and improve feature construction efficiency. Suitable for high-dimensional data and complex data types, can optimize model performance and speed up machine learning process.

[0053] Multi-Task Learning Model (MTLM): By sharing underlying feature representation, simultaneously learning multiple related tasks, using the correlation between tasks to improve model generalization ability. Suitable for natural language processing, computer vision, etc. Can reduce data requirements and improve learning efficiency.

[0054] Requirement Causality Graph (RCG): Describes the causal relationship between requirements in a graph structure, clearly defines the propagation path and impact range of requirement changes. Can be used for requirement management, change impact analysis, and auxiliary requirement priority sorting and risk assessment.

[0055] Device Performance Benchmark Library (DPBL): A library that stores device performance benchmark data, provides standardized testing methods and indicators, and is used to evaluate and compare the performance of different devices. Suitable for hardware selection, performance optimization, etc. Provides data support for decision-making.

[0056] User-Application Interaction Graph (UAIG): Records the interaction behavior between users and applications in a graph structure, reveals user usage patterns and application function associations. Can be used for user behavior analysis, application recommendation system, and auxiliary product optimization and user experience improvement.

[0057] Non-Cooperative Game Theory Framework (NCGTF): A framework that studies how multiple participants make decisions without cooperation, analyzing strategic interactions and equilibrium outcomes. It is applicable in economics, political science, and other fields to predict market behavior, policy impacts, etc.

[0058] Multi-Objective Optimization Engine (MOOE): An algorithm or tool that optimizes multiple objective functions simultaneously, seeking Pareto optimal solutions that satisfy multiple constraints. It is applicable in engineering design, resource allocation, etc., balancing conflicts among multiple objectives.

[0059] Decision Tree Visualization (DTV): A graphical representation of decision tree models, visually presenting decision rules and branching logic. It is applicable in model interpretation, result analysis, helping users understand model decision-making processes and improve model interpretability.

[0060] Case-Based Reasoning (CBR): A reasoning method that compares new problems with historical cases for similarity, reusing or modifying historical solutions to solve problems. It is applicable in medical diagnosis, legal consultation, etc., reducing repetitive work and improving problem-solving efficiency.

[0061] Semantic Network Construction (SNC): Constructing a semantic network representing concepts, entities, and their relationships for knowledge representation, reasoning, and querying. It is applicable in natural language processing, knowledge graph construction, etc., supporting intelligent question answering, information retrieval, etc.

[0062] Heterogeneous Resource Generalization Framework (HRGF): A framework that processes and integrates different types of heterogeneous resources, achieving unified management and utilization through abstraction and conversion. It is applicable in cloud computing, big data, etc., improving resource utilization and system flexibility.

[0063] Topology-Aware Scheduling Algorithm (TASA): A scheduling algorithm that considers network topology structure, optimizing task allocation and resource scheduling based on node location and connection relationships. It is applicable in distributed systems, data centers, etc., improving task execution efficiency and system performance.

[0064] Deep Reinforcement Learning (DRL) combines deep learning and reinforcement learning techniques to learn optimal policies through interaction between an agent and its environment. It is applicable to fields such as games and robot control, and can handle high-dimensional state spaces and complex decision-making problems.

[0065] Spatiotemporal Demand Forecasting (STDF) is a method for predicting future demand based on spatiotemporal data, taking into account the impact of time and space factors on demand. It is applicable to fields such as transportation and energy, providing decision support for resource allocation and scheduling.

[0066] A Markov Decision Process (MDP) is a mathematical model that describes decision-making problems in stochastic processes. It possesses the Markov property, meaning that future states depend only on the current state and the decision made. It is applicable to fields such as dynamic programming and reinforcement learning, and is used to solve for optimal policies.

[0067] Dynamic Voltage and Frequency Scaling (DVFS): A technology that dynamically adjusts processor voltage and frequency based on system load, reducing power consumption and improving performance. It is suitable for mobile devices, data centers, and other scenarios to achieve energy efficiency optimization.

[0068] Geographic Load Balancing (GLB) is a technology that distributes load based on geographical location, optimizing system response speed and resource utilization through intelligent routing and resource allocation. It is suitable for scenarios such as content delivery networks and cloud computing, improving user experience.

[0069] Checkpoint Optimization Technique (COT): A technique that optimizes checkpoint setup and recovery processes, reducing checkpoint overhead and improving system fault tolerance. It is suitable for distributed systems, high-performance computing, and other scenarios, ensuring rapid system recovery after a failure.

[0070] Progressive Cache Warming (PCW) is a technique that loads cached data gradually to avoid performance issues caused by loading large amounts of data at once. It is suitable for high-concurrency, high-data-volume scenarios, improving system response speed and stability.

[0071] Dynamic Policy Optimization Algorithm (DPOA): An optimization algorithm that dynamically adjusts policies based on environmental changes, improving policy adaptability through online learning and feedback mechanisms. Suitable for autonomous driving, robot control, and other fields that deal with complex and dynamic environments.

[0072] Multimodal Monitoring Architecture (MMA): An architecture that combines various monitoring modalities such as logs, metrics, and traces, providing comprehensive and multi-dimensional system monitoring capabilities. Suitable for microservices, distributed systems, and other scenarios, enabling rapid fault localization and performance optimization.

[0073] Digital Twin Technology (DTT): A technology that creates a digital copy of a physical entity or system, optimizing entity behavior through real-time data synchronization and simulation analysis. Suitable for intelligent manufacturing, smart cities, and other fields, enabling predictive maintenance and resource optimization.

[0074] Temporal Convolutional Network (TCN): A convolutional neural network specifically designed for processing time-series data, capturing long-term dependencies through causal and dilated convolutions. Suitable for speech recognition, time series prediction, and other scenarios, improving model performance and efficiency.

[0075] Model Predictive Control (MPC): A model-based control method that predicts the future behavior of a system and optimizes control sequences to achieve goals. Suitable for industrial process control, autonomous driving, and other fields, handling multivariate and constraint optimization problems.

[0076] Federated Learning Framework (FLF): A distributed machine learning framework that allows multiple participants to collaboratively train models without sharing data. Suitable for privacy-sensitive fields such as healthcare and finance, protecting data privacy while improving model performance.

[0077] What-If Analysis Engine (WIAE): An analysis tool that simulates the results of different hypothetical conditions by adjusting parameters or scenarios to predict potential impacts. Suitable for decision support, risk management, and other scenarios, helping users make more informed decisions.

[0078] 3D Topology Visualization Component (3DTVC): A visualization component for displaying three-dimensional topology structures, providing intuitive presentation of complex networks or geographic information through interactions such as rotation and scaling. Suitable for network management, geographic information systems, and other scenarios, improving data understanding and analysis capabilities.

[0079] Complex Event Querying (CEQ): A technology for querying complex event patterns, extracting meaningful information or patterns from a large number of events by defining event patterns and conditions. Suitable for real-time monitoring, log analysis, and other scenarios, enabling event correlation analysis and anomaly detection.

[0080] Residual Network 50 (ResNet50): ResNet50 is a classic deep convolutional neural network architecture, belonging to the Residual Network (ResNet) series. It solves the gradient problem in deep network training through residual blocks and skip connections, supports ultra-deep network construction, significantly improves model performance, and is widely used in computer vision tasks such as image classification and detection.

[0081] As mentioned above, how to achieve self-adaptation to the dynamic environment of edge cloud when allocating computing resources has become an important research problem.

[0082] Based on the above description, as shown in the method for allocating resources based on edge cloud proposed by the embodiment, applied to a resource allocation system based on edge cloud, as shown in the system includes a data acquisition module 201, a data analysis module 202, a resource allocation module 203, and a monitoring feedback module 204; the method includes: Figure 1 Figure 2

[0083] Step 101, the data acquisition module acquires environment information in the edge cloud environment; wherein the environment information includes application load information, device performance information, and user behavior information.

[0084] In specific implementation, the data acquisition module 201 is used to collect environment information in the edge cloud environment. Through multi-modal technology, full-factor environment monitoring is realized, device, network, and application data are dynamically collected, and physical parameters are combined with visual perception; adaptive mechanism is used to optimize sampling frequency and data enhancement to ensure data quality, and finally, semantic network construction based on knowledge representation and reasoning is used to support intelligent query, forming a complete perception system.

[0085] The data acquisition module 201 uses multi-modal technology to dynamically monitor the running state of devices, networks, and applications, and combines adaptive sampling strategies and semantic network construction methods to form a full-factor perception system.​​

[0086] In step 102, the data analysis module analyzes the environment information to obtain application load prediction results, device capability evaluation results and behavior analysis results, determines an initial resource allocation strategy according to the application load prediction results, the device capability evaluation results and the behavior analysis results, and optimizes the initial resource allocation strategy to obtain a target resource allocation strategy.

[0087] In implementation, the data analysis module 202 analyzes the collected environment information to formulate a computing resource allocation strategy. Key data features are extracted through machine learning-based automatic feature engineering, and application demand, device performance and user behavior prediction are cooperatively optimized through a machine learning-based multi-task learning model. Computing resource allocation rules are revealed through knowledge representation and reasoning-based causal reasoning, and decision-making basis is generated through machine learning and knowledge representation and reasoning-based explainable artificial intelligence. The computing resource competition problem in a multi-tenant scenario is solved through reinforcement learning and decision optimization-based game theory modeling, and intelligent decision-making is achieved.

[0088] The data analysis module 202 cooperatively predicts application demand, device performance and user behavior through machine learning-based automatic feature engineering and multi-task learning model, and generates explainable intelligent decisions through knowledge representation and reasoning-based causal reasoning and reinforcement learning and decision optimization-based game theory modeling.

[0089] In step 103, the resource allocation module allocates computing resources according to the target resource allocation strategy.

[0090] In implementation, the resource allocation module 203 allocates and schedules computing resources according to the decision result. Multiple types of computing resources are unified as schedulable units through a heterogeneous resource generalization framework, dynamic scheduling is achieved through a reinforcement learning algorithm, prediction models based on machine learning are integrated to predict computing resource demand and allocate in advance, transfer learning based on machine learning is integrated to quickly recover from faults, and dynamic frequency adjustment technology based on edge and distributed technology is adopted to optimize energy efficiency, forming an intelligent, efficient and reliable computing resource management mechanism.

[0091] The resource allocation module 203 realizes dynamic scheduling based on a heterogeneous resource generalization framework, reinforcement learning and prediction models based on machine learning, and combines transfer learning based on machine learning and dynamic frequency adjustment technology based on edge and distributed technology to optimize energy efficiency.

[0092] In step 104, the monitoring and feedback module monitors the use state of the computing resources in real time and feeds back the use state to the data analysis module, so that the data analysis module adjusts the target resource allocation strategy according to the use state.

[0093] In implementation, the monitoring feedback module 204 monitors the usage state of the computing resources in real time and feeds back the usage state to the data analysis module 202 to adjust the target resource allocation strategy. The system state is transparentized through the visual monitoring platform, the computing resource anomaly is quickly identified by the artificial intelligence anomaly detection based on machine learning, the performance bottleneck is accurately located by the knowledge graph based on knowledge representation and reasoning, the management strategy is dynamically adjusted by the PID / MPC control loop based on reinforcement learning and decision optimization, the global strategy is optimized by the federated learning based on edge and distributed technology, the simulation deduction is supported by the digital twin based on digital twin and simulation, and the closed-loop optimization mechanism is formed through continuous iteration.

[0094] The monitoring feedback module 204 monitors the resource state in real time through the visual platform combined with the intelligent anomaly detection based on machine learning, locates the performance bottleneck by the knowledge graph based on knowledge representation and reasoning, and constructs the dynamic closed-loop optimization mechanism by the PID / MPC control based on reinforcement learning and decision optimization and the federated learning based on edge and distributed technology.

[0095] Figure 3 The structure schematic diagram of the resource allocation system based on edge cloud of another embodiment of the present disclosure is shown. As shown in Figure 3 The artificial intelligence technologies such as machine learning, deep learning, reinforcement learning and decision optimization, knowledge representation and reasoning, edge and distributed technology, digital twin and simulation, meta-learning and automatic learning empower the four aspects of data collection, data analysis, resource allocation and monitoring feedback of the resource allocation system based on edge cloud. Among them, the data collection includes the functions of application collection, device collection, behavior collection, collection coordination, etc.; the data analysis includes the functions of preprocessing, demand analysis, capability evaluation, behavior analysis, decision engine, decision coordination, etc.; the resource allocation includes the functions of heterogeneous resource generalization, reinforcement learning scheduling, pre-allocation, fault-tolerant migration engine, energy efficiency optimization, etc.; the monitoring feedback includes the functions of monitoring instrument, anomaly detection, feedback control, learning optimization, simulation, etc.

[0096] Through the above embodiment, the data collection module collects environment information in the edge cloud environment; wherein, the environment information includes: application load information, device performance information and user behavior information. In this way, the data collection module can construct a three-dimensional perception system of applications, devices and users based on multi-modal data fusion, and improve the resource supply-demand matching accuracy. The data analysis module analyzes and processes the environment information to obtain application load prediction results, device capability evaluation results and behavior analysis results, and determines an initial resource allocation strategy according to the application load prediction results, the device capability evaluation results and the behavior analysis results, and optimizes the initial resource allocation strategy to obtain a target resource allocation strategy. In this way, the data analysis module can automatically optimize the target resource allocation strategy of the computing resources according to the environment information, maximize the global utilization rate, and realize the precise adaptation of the computing resources on demand. The resource allocation module allocates the computing resources according to the target resource allocation strategy. The monitoring feedback module monitors the use state of the computing resources in real time, and feeds back the use state to the data analysis module, so that the data analysis module adjusts the target resource allocation strategy according to the use state. In this way, the monitoring feedback module can feed back the use state of the computing resources monitored in real time to the data analysis module, and the data analysis module can adjust the target resource allocation strategy in real time and accurately according to the real-time use state of the computing resources.

[0097] In some embodiments, as shown in FIG. 1, the data collection module includes an application collection unit, a device collection unit, a behavior collection unit, and a collection coordination unit; step 101 includes: Figure 4

[0098] Step 1011, the application collection unit collects application document data from the application software through a multi-modal analysis algorithm and a natural language processing algorithm, collects application running features from the application software through an application programming interface, and takes the application document data and the application running features as application load information.

[0099] In specific implementation, the application collection unit 2011 realizes dynamic perception of cross-domain resource demand through a lightweight intelligent agent. A multi-modal analysis framework is adopted to fuse QoS constraints such as “video analysis ≤50ms latency” in the application document based on machine learning NLP, and to capture runtime features such as TensorFlow job GPU memory occupation through API hook. Combined with DQN dynamic adjustment of sampling granularity based on reinforcement learning and decision optimization, a demand knowledge graph based on knowledge representation and reasoning is constructed to support semantic query, which significantly improves the resource supply-demand matching accuracy in heterogeneous edge scenarios.

[0100] ​At step 1012, the device collection unit collects device storage parameters of the edge device through the probe, collects the device physical state of the edge device through visual analysis, and takes the device storage parameters and the device physical state as the device performance information.

[0101] In implementation, the device collection unit 2012 collects multi-dimensional performance indicators of servers, gateways and sensors through lightweight probes, and tracks key parameters such as CPU / GPU utilization, memory bandwidth and storage IOPS in real time. Based on the federated learning construction of edge and distributed technology, a device fingerprint library is constructed to identify differences in heterogeneous hardware computing power, and the fan speed and LED indicators of the physical state of the computer vision analysis device are combined to form a software and hardware collaborative portrait. An LSTM model based on deep learning is used to predict the end-of-life risk of SSD and provide early warning, and a dynamic power consumption model is established through power metering to reduce the idle state power consumption of GPU, supporting predictive maintenance and energy efficiency optimization of edge systems.

[0102] At step 1013, the behavior collection unit collects the interaction log of the user operation record, constructs an operation time sequence mode atlas according to the interaction log, determines the predicted user behavior according to the time sequence mode atlas through a time sequence attention algorithm, detects abnormal user behavior in real time according to the time sequence mode atlas through an isolation forest algorithm, and takes the predicted user behavior and abnormal user behavior as the user behavior information.

[0103] In implementation, the behavior collection unit 2013 focuses on user behavior modeling and anomaly detection in edge scenarios, collects the interaction log of user click, input, page jump and other operation records through non-inductive collection, constructs an operation time sequence mode atlas, and realizes behavior prediction using a deep learning-based time sequence attention Transformer architecture. To address data collection discontinuity or log missing problems, a deep learning-based conditional generative adversarial network (CGAN) is deployed to generate synthetic data to fill data gaps; at the same time, an isolation forest algorithm based on machine learning is used to detect abnormal click streams such as high-frequency abnormal operations and unexpected operation sequences in real time. To protect privacy and security, a differential privacy mechanism based on knowledge representation and reasoning is introduced to add noise disturbance to the original behavior sequence, ensuring data availability while achieving privacy budget control.

[0104] At step 1014, the collection coordination unit performs spatio-temporal correlation modeling on the application load information, the device performance information and the user behavior information to obtain the environment information, and adjusts the sampling mode of the application collection unit, the device collection unit and the behavior collection unit according to the multi-modal fusion result of the environment information.

[0105] In specific implementation, the collection coordination unit 2014 dynamically optimizes the sensor sampling strategy based on the dynamic strategy optimization algorithm of reinforcement learning and decision optimization, realizes energy consumption reduction and network bandwidth elastic allocation, constructs a spatio-temporal graph neural network based on deep learning to fuse device state, application demand and user behavior ternary data, and forms global context perception. Through complex event query, multi-source heterogeneous data collection and cross-domain knowledge fusion are effectively coordinated.

[0106] Through the above scheme, the application collection unit 2011 provides resource demand profiling, the device collection unit 2012 monitors hardware state, and the behavior collection unit 2013 characterizes operation mode. The data of the three are time and space correlated modeling through the collection coordination unit 2014. The collection coordination unit 2014 dynamically optimizes the data sampling strategy based on the multi-modal fusion result, such as increasing the sampling granularity of the application collection unit 2011 during burst traffic; adjusting the resource allocation priority, such as increasing the GPU monitoring frequency of the device collection unit 2012 during video analysis, and driving the adaptive adjustment of the abnormal detection threshold of the behavior collection unit 2013, finally realizing the joint optimization and global intelligent decision of "demand-resource-behavior".

[0107] In some embodiments, as shown in Figure 5 The data analysis module includes a preprocessing unit, a demand analysis unit, a capability evaluation unit, a behavior analysis unit, a decision engine unit, and a decision coordination unit. Step 102 includes:

[0108] Step 1021, the preprocessing unit pre-processes the environment information to obtain standard data, and analyzes the standard data to obtain application load information, device performance information and user behavior information.

[0109] In specific implementation, the preprocessing unit 2021 realizes data quality improvement through a multi-modal data processing pipeline. An AutoEncoder based on deep learning constructs an anomaly detection model, which can automatically correct outliers of indicators such as CPU utilization, and generates synthetic data conforming to the original distribution through a conditional GAN based on deep learning to fill in missing fields. A BERT model based on deep learning is used to realize semantic standardization of natural language description, and a device performance ontology library based on knowledge representation and reasoning is used to complete automatic mapping of heterogeneous hardware indicators. A dynamic time warping algorithm based on machine learning is used to align asynchronous time series data, providing a high-quality data foundation for edge intelligent analysis.

[0110] At step 1022, the demand analysis unit analyzes the application load information to obtain predicted storage demand and application document data, determines predicted load demand from the application load information by using a pre-constructed load prediction model, and detects abnormal demand from the application load information by using a pre-constructed anomaly detection model. The predicted storage demand, the predicted load demand, and the abnormal demand are taken as application load prediction results.

[0111] In specific implementation, the demand analysis unit 2022 accurately describes application resource demand by multi-dimensional modeling. A multi-task learning framework is constructed by using a deep learning-based Transformer to synchronously predict the demand for CPU, GPU, memory, and the like, and an NLP engine based on machine learning is integrated to automatically analyze QoS constraints in configuration documents. A Prophet model based on machine learning is used to implement load prediction, and an LSTM-AE based on deep learning is used to detect burst demand anomalies. A demand causality graph is constructed by using a PC algorithm based on knowledge representation and reasoning, key influencing factors are quantified, and cross-domain resource optimization decisions are supported.

[0112] At step 1023, the capability evaluation unit determines a hardware computing power index of the edge device based on a pre-constructed device performance benchmark library according to the device performance information, and determines a solid state disk (SSD) predicted life of the edge device based on a long short-term memory (LSTM) network according to the device performance information. The hardware computing power index and the SSD predicted life are taken as device capability evaluation results.

[0113] In specific implementation, the capability evaluation unit 2023 dynamically evaluates edge device capability by multi-dimensional modeling. A device performance benchmark library is constructed by using federated learning based on edge and distributed technologies, heterogeneous hardware computing power indexes are standardized, and device physical states are analyzed by using computer vision based on machine learning. An LSTM model based on deep learning is used to predict the life of the SSD and provide early warning, a digital twin technology based on digital twin and simulation is used to simulate load pressure, an energy efficiency model is integrated to quantify the energy consumption of unit computing power, and dynamic capability boundaries are provided for resource scheduling.

[0114] At step 1024, the behavior analysis unit determines a predicted click rate from the user behavior information by using a pre-constructed user prediction model, determines an association mode between a user and an application from the user behavior information by using a graph attention network, and determines a predicted access space-time from the user behavior information by using a spatio-temporal graph convolution network. A user portrait is updated based on the predicted click rate, the association mode, and the predicted access space-time to obtain a behavior analysis result.

[0115] In specific implementation, the behavior analysis unit 2024 realizes accurate characterization of user intent through multi-dimensional behavior modeling. A hybrid prediction model is constructed based on DeepFM of deep learning, which significantly improves the prediction accuracy of click rate. The implicit association patterns in the "user-application" interaction graph are mined by combining the graph attention network (GAT) based on deep learning. The spatiotemporal prediction of future access hot zones is realized by developing the spatiotemporal graph convolution network (STGCN) based on deep learning, and the user portrait is dynamically updated through the online learning mechanism based on machine learning. The isolation forest algorithm based on machine learning is integrated to detect abnormal behavior in real time, and to provide behavior-driven decision support for resource elasticity.

[0116] In step 1025, the decision engine unit performs resource allocation according to the application load prediction result, the device capability evaluation result and the behavior analysis result through multi-level optimization to obtain an initial resource allocation strategy.

[0117] In specific implementation, the decision engine unit 2025 realizes intelligent decision of resource allocation through a multi-level optimization mechanism. A multi-tenant competition model is constructed based on reinforcement learning and decision optimization non-cooperative game theory, and Nash equilibrium solution is used to realize Pareto optimal allocation. Distributed constraint optimization (DCOP) based on reinforcement learning and decision optimization is used to solve cross-domain conflicts. SHAP values based on machine learning are used to quantify decision causality, and decision tree visualization based on machine learning is used to present resource scheduling logic chain. A PID controller based on reinforcement learning and decision optimization and a deep reinforcement learning component based on reinforcement learning and decision optimization are integrated to form a closed-loop feedback mechanism, which realizes long-term effective optimization while ensuring QoS.

[0118] In step 1026, the decision coordination unit optimizes the initial resource allocation strategy based on a pre-constructed multi-objective optimization engine to obtain a target resource allocation strategy.

[0119] In specific implementation, the decision coordination unit 2026 realizes global decision optimization through a multi-level collaborative mechanism. A multi-objective optimization engine is constructed based on NSGA-III algorithm of reinforcement learning and decision optimization, which synchronously optimizes resource utilization, energy consumption and QoS indicators. The dynamic configuration of optimization weights is realized by combining the analytic hierarchy process (AHP) based on knowledge representation and reasoning. A decision knowledge base is established, historical strategies are reused through case-based reasoning (CBR) based on knowledge representation and reasoning, and a meta-learning component based on meta-learning and automatic learning is integrated to accelerate adaptation to new scenarios, forming a strategy iteration closed loop to enhance the self-adaptation ability of the system.

[0120] Through the above scheme, the preprocessing unit 2021, the demand analysis unit 2022, the capability evaluation unit 2023, the behavior analysis unit 2024, the decision engine unit 2025, and the decision coordination unit 2026 constitute a data-driven intelligent decision closed-loop system. The preprocessing unit 2021 provides a standardized data base for the demand analysis unit 2022, the capability evaluation unit 2023, and the behavior analysis unit 2024; the output results of the three are used to generate an initial resource allocation strategy by the decision engine unit 2025, and then the multi-objective optimization and knowledge reuse are performed by the decision coordination unit 2026.

[0121] In some embodiments, as shown in Figure 6 The resource allocation module includes a heterogeneous resource generalization unit, a reinforcement learning scheduling unit, a pre-allocation unit, a fault-tolerant migration engine unit, and an energy efficiency optimization unit. Step 103 includes:

[0122] In step 1031, the heterogeneous resource generalization unit determines the resource dependency relationship based on a graph neural network and manages the computing resources according to the resource dependency relationship.

[0123] In specific implementation, the heterogeneous resource generalization unit 2031 realizes efficient management of heterogeneous resources through a multi-level abstraction mechanism. A multi-dimensional resource tag system including computing power types, energy efficiency levels, and topological positions is constructed, a graph neural network (GNN) based on deep learning is used to model resource dependency relationships to optimize slicing combination strategies, and a topology-aware scheduling algorithm is developed to reduce cross-node communication overhead. A natural language interface based on machine learning is integrated to support semantic resource queries, and a resource description framework (RDF) based on knowledge representation and reasoning is used to realize standardized encapsulation and dynamic arrangement of heterogeneous resources, providing a unified scheduling interface for upper-layer applications.

[0124] In step 1032, the reinforcement learning scheduling unit constructs a scheduling strategy for the target resource allocation strategy based on a proximal policy optimization algorithm, determines the execution time of each task in the target resource allocation strategy through a long short-term memory network, optimizes the scheduling strategy according to the execution time to obtain an optimized scheduling sequence, and schedules the computing resources according to the scheduling sequence.

[0125] In specific implementation, the reinforcement learning scheduling unit 2032 realizes dynamic resource optimization through deep reinforcement learning. A PPO algorithm based on reinforcement learning and decision optimization is used to construct a scheduling strategy. The state space integrates queue length, resource utilization and other dimensional characteristics. The reward function balances the three objectives of utilization rate, QoS compliance rate and energy consumption. A LSTM network based on deep learning is deployed to predict task execution time to optimize the scheduling sequence, and a deep learning based attention mechanism is used to dynamically allocate scheduling weights. Through the experience replay buffer based on reinforcement learning and decision optimization, the continuous iteration strategy can adapt to changes in workloads, effectively cope with sudden loads and resource competition, and ensure QoS compliance rate of critical tasks.

[0126] In step 1033, the pre-allocation unit determines the predicted spatio-temporal demand according to the target resource allocation strategy through a time series prediction model, calibrates the current allocation proportion of the target resource allocation strategy based on a proportional-integral-derivative controller to obtain a target allocation proportion, and adjusts the scheduling of the computing resources according to the target allocation proportion.

[0127] In specific implementation, the pre-allocation unit 2033 drives resource forward-looking scheduling through multi-scale prediction. A Transformer based on deep learning and a Prophet-LSTM model based on machine learning are integrated to realize spatio-temporal demand prediction. A Monte Carlo Tree Search (MCTS) based on reinforcement learning and decision optimization is used to optimize the multi-resource collaborative reservation strategy, and a progressive cache warm-up mechanism is combined. A PID controller based on reinforcement learning and decision optimization is deployed to dynamically calibrate the pre-allocation proportion, effectively reducing the risk of resource contention under sudden loads, and improving the timeliness of critical task response.

[0128] In step 1034, the fault-tolerant migration engine unit detects anomalies in the edge device through an Isolation Forest algorithm and a Long Short-Term Memory network, and repairs the anomalies existing in the edge device.

[0129] In specific implementation, the fault-tolerant migration engine unit 2034 ensures task continuity through a multi-level fault-tolerant mechanism. Isolation Forest based on machine learning and LSTM model based on deep learning are deployed to realize real-time detection and life prediction of device anomalies. A TrAdaBoost based on machine learning is used to accelerate the adaptation of the migration learning model, and checkpoint optimization technology is combined to reduce the amount of state migration data. A Markov Decision Process (MDP) based on reinforcement learning and decision optimization is constructed to optimize migration target selection. Multi-objective optimization algorithms are integrated to balance recovery time, migration cost and QoS impact, realize fast fault recovery, and significantly improve edge service availability.

[0130] At step 1035, the energy efficiency optimization unit determines the predicted device power consumption based on the pre-constructed device power consumption model, adjusts the voltage frequency of the power grid according to the predicted device power consumption, and adjusts the scheduling of the computing resources according to the voltage frequency of the power grid and the carbon intensity data.

[0131] In specific implementation, the energy efficiency optimization unit 2035 realizes green computing through multi-dimensional energy efficiency management. A fine-grained device-level power consumption model is constructed, and a support vector regression (SVR) based on machine learning is used to realize the whole cabinet power consumption prediction. A deep deterministic policy gradient (DDPG) algorithm based on reinforcement learning and decision optimization is used to realize the continuous adjustment of the voltage frequency, and a load-aware hibernation strategy is used to reduce the idle energy consumption. The power grid carbon intensity data and the geographic load balancing mechanism are integrated, and the computing tasks are migrated to the clean energy rich areas, which significantly improves the data center energy efficiency ratio (performance / watt).

[0132] Through the above scheme, the five units of the resource allocation module 203 form a closed-loop intelligent resource management system. The heterogeneous resource generalization unit 2031 provides a standardized resource view and topology-aware scheduling basis, supporting the reinforcement learning scheduling unit 2032 to make dynamic decisions based on a multi-dimensional state space; the pre-allocation unit 2033 drives forward-looking scheduling through spatiotemporal prediction, and the results are input into the scheduler to optimize real-time allocation; the fault-tolerant migration engine unit 2034 intervenes to ensure continuity in the event of a fault, and the energy efficiency optimization unit 2035 adjusts the resource working state through dynamic voltage frequency adjustment control and carbon-aware scheduling, with the five units working together to realize high-availability, low-energy-consumption edge intelligent resource management.

[0133] In some embodiments, as shown in Figure 7 The monitoring feedback module includes a monitoring instrument unit, an anomaly detection unit, a feedback control unit, a learning optimization unit, and a simulation unit. Step 104 includes:

[0134] At step 1041, the monitoring instrument unit monitors the usage state of the computing resources in real time and constructs the association relationship between the edge devices, application software, and computing resources.

[0135] In specific implementation, the monitoring instrument unit 2041 realizes global perception of the edge environment through a multi-modal monitoring architecture. A Prometheus+Grafana architecture is used to construct a million-level time series data pipeline, combined with a 3D topology visualization component to dynamically map the "device-application-resource" association relationship. A machine learning-based isolation forest and a deep learning-based LSTM-AE dual-engine are deployed to detect anomalies such as memory leaks and latency surges in real time, a machine learning-based Prophet model is integrated to realize fast resource bottleneck prediction, and a digital twin technology based on digital twin and simulation is used to continuously calibrate the "physical-virtual" state deviation, providing visual data support for operation and maintenance decisions.

[0136] At step 1042, the anomaly detection unit performs anomaly detection on the computing resource according to the usage state by using the pre-trained anomaly detection model to obtain time sequence detection results and graph detection results, and takes the time sequence detection results and the graph detection results as the anomaly detection results.

[0137] In specific implementation, the anomaly detection unit 2042 implements accurate fault diagnosis through a multi-modal detection mechanism. A deep learning-based time sequence convolution network (TCN) and a deep learning-based graph anomaly detection double engine are constructed to capture time sequence anomalies such as cache hit rate mutation and cross-node associated faults such as storage I / O competition, respectively. A knowledge graph based on knowledge representation and reasoning and a Bayesian network based on knowledge representation and reasoning are integrated to automatically infer the root cause of performance bottlenecks and quantify the contribution degree of influencing factors. Reinforcement learning is used to dynamically adjust the monitoring threshold to adapt to the peak and valley characteristics of the business, thereby improving the accuracy of anomaly detection and reducing false positives.

[0138] At step 1043, the feedback control unit performs repair on the key indicators of the edge device based on a proportional-integral-derivative controller according to the anomaly detection results to obtain device parameters, and sends the device parameters to the data analysis module, so that the data analysis module adjusts the target resource allocation strategy according to the usage state.

[0139] In specific implementation, the feedback control unit 2043 implements automatic closed loop of resource management through a multi-level control mechanism. A model predictive control (MPC) based on reinforcement learning and decision optimization is used to optimize multivariable parameters, and a PID controller based on reinforcement learning and decision optimization is used to realize rapid regulation and control of key indicators. A self-healing library with pre-defined repair strategies is built, and a case-based reasoning (CBR) based on knowledge representation and reasoning is used to generate customized solutions. A Monte Carlo simulation evaluation strategy based on reinforcement learning and decision optimization is deployed to ensure the safety of adjustment and form a complete closed loop from monitoring anomalies to executing repair.

[0140] At step 1044, the learning optimization unit trains the anomaly detection model in the anomaly detection unit based on federated learning to obtain a trained anomaly detection model.

[0141] In specific implementation, the learning optimization unit 2044 implements distributed strategy co-evolution through a federated learning framework. A secure aggregation protocol and a differential privacy mechanism are constructed to jointly train the anomaly detection model across nodes. A meta-learning component based on meta-learning and automatic learning is integrated to accelerate the convergence of new node models and support AB testing driven strategy iteration. A genetic algorithm based on machine learning and an AutoML engine based on machine learning are deployed to automatically optimize monitoring thresholds and control parameters, thereby forming a global resource management optimization closed loop under data privacy protection.

[0142] Step 1045: The simulation unit verifies the target resource allocation strategy using a pre-built digital twin simulation model.

[0143] In practical implementation, simulation unit 2045 supports edge environment strategy verification through high-fidelity modeling and hypothesis analysis. A 3D model of the device is constructed based on deep learning's Neural Radiation Field (NeRF), and a resource contention simulator based on digital twins and simulation is developed to reproduce concurrent multi-application scenarios. A What-If analysis engine based on digital twins and simulation is integrated to support strategy pre-evaluation, and a reinforcement learning sandbox is used to simulate long-term strategy effects. Fault injection via ChaosMesh verifies system resilience, achieving a zero-cost strategy trial-and-error and optimization loop.

[0144] Through the above scheme, the five units of the monitoring feedback module 204 constitute a data-driven intelligent operation and maintenance closed loop. The monitoring instrument unit 2041 provides the foundation for global status perception and digital twin calibration, supporting the anomaly detection unit 2042 to perform "time-series-graph" dual-modal diagnosis; the feedback control unit 2043 executes automated repair based on the detection results, and after the strategy effect is verified by the simulation unit 2045, the global strategy evolution is achieved through the federated learning framework of the learning optimization unit 2044, forming a complete self-iterative link of "monitoring-diagnosis-repair-verification-optimization", which significantly improves the level of intelligence in edge cloud resource management.

[0145] In some embodiments, the pre-allocation process of computing resources includes: the pre-allocation unit receiving application load information and sending a computing resource query request to the heterogeneous resource generalization unit; the pre-allocation unit receiving candidate resources from the heterogeneous resource generalization unit and sending a reservation request to the reinforcement learning scheduling unit, so that the reinforcement learning scheduling unit can generate a reservation contract based on the reservation request.

[0146] In practice, Figure 8 This is a flowchart illustrating the pre-allocation process of computing resources according to an embodiment of this disclosure. Figure 8 As shown, the application demand drives the pre-allocation process of computing resources. The application acquisition unit 2011 detects an increase in service load, triggering QoS constraints such as "latency < 50ms" and sending them to the pre-allocation unit 2033. The pre-allocation unit 2033 then queries the heterogeneous resource generalization unit 2031 to obtain a list of computing resource pools that meet energy efficiency standards. After the heterogeneous resource generalization unit 2031 returns candidate resources, the pre-allocation unit 2033, combined with the decision model of the reinforcement learning scheduling unit 2032, submits a reservation request including Service Level Agreement (SLA) parameters. The reinforcement learning scheduling unit 2032 optimizes through dynamic game theory, ultimately identifying the target resource node and generating a reservation contract, synchronously updating the global resource view, thus completing the closed loop of demand-driven predictive resource pre-allocation.

[0147] In some embodiments, a device overload triggers a dynamic scheduling process, including: the reinforcement learning scheduling unit receiving device performance information and sending a feasible domain query request to the energy efficiency optimization unit; the reinforcement learning scheduling unit receiving the voltage and frequency adjustment range from the energy efficiency optimization unit, determining that hot migration is required based on the performance loss and heat dissipation requirements of the edge device, and then sending a hot migration start request to the fault-tolerant migration engine unit so that the fault-tolerant migration engine unit can perform hot migration according to the hot migration start request.

[0148] In practice, Figure 9 This is a flowchart illustrating the device overload-triggered dynamic scheduling process according to an embodiment of this disclosure. Figure 9 As shown, when the device acquisition unit 2012 detects that the GPU temperature exceeds the limit (e.g., "GPU temperature > 85℃"), it immediately sends an overload alarm to the reinforcement learning scheduling unit 2032. The reinforcement learning scheduling unit 2032 then queries the energy efficiency optimization unit 2035 to obtain a dynamic frequency adjustment strategy for dynamic voltage and frequency regulation (e.g., reducing the frequency from 1.2GHz to 0.8GHz). After obtaining the feasible domain, the reinforcement learning scheduling unit 2032 comprehensively evaluates the performance loss and heat dissipation requirements, and decides to trigger the fault-tolerant migration engine unit 2034 to execute a hot migration of the TensorFlow job. After the migration process starts, the device load is gradually transferred to the backup node, the GPU temperature begins to drop, the system synchronously updates the resource topology and marks the overheated node as a maintenance state, completing the closed-loop handling.

[0149] In some embodiments, the user behavior optimization caching strategy process includes: the heterogeneous resource generalization unit receiving an information coordination instruction and sending a preloading instruction to the pre-allocation unit, so that the pre-allocation unit can perform pre-caching of user behavior information according to the preloading instruction and send the real-time monitored cache hit rate to the data acquisition module.

[0150] In practice, Figure 10 A flowchart illustrating the process of optimizing caching strategies based on user behavior in embodiments of this disclosure. Figure 10 As shown, the behavior analysis unit 2013 updates the access hotspot prediction based on the spatiotemporal map, triggering the acquisition coordination unit 2014 to dynamically optimize the SSD cache allocation. After the coordination command is issued, the heterogeneous resource generalization unit 2031 injects the video stream preloading strategy into the pre-allocation unit 2033. The pre-allocation unit 2033 executes pre-caching and monitors in real time, ultimately reporting back to the behavior analysis unit 2013 the optimization effect of a 27% increase in hit rate, forming a closed-loop optimization chain of "behavior analysis → resource optimization → preloading execution → effect feedback", significantly improving the distribution efficiency of hot content.

[0151] In some embodiments, the multimodal data collaborative sampling process includes: the acquisition coordination unit receiving a traffic alarm from the application acquisition unit and a high load alert from the device acquisition unit, controlling the sampling coordination mechanism to start, sending the start command of the sampling coordination mechanism to the device acquisition unit, and sending a command to the behavior acquisition unit to reduce the log sampling rate.

[0152] In practice, Figure 11 This is a flowchart of the multimodal data collaborative sampling process according to an embodiment of this disclosure. Figure 11 As shown, the application acquisition unit 2011 triggered a sudden traffic alarm due to microservice link overload, interface call overload, and cache breakdown caused by data sampling. The device acquisition unit 2012 simultaneously reported high CPU load conditions such as "CPU utilization > 90% for 10 seconds." The acquisition coordination unit 2014 initiated dynamic collaborative sampling, sending an instruction to the device acquisition unit 2012 to increase the monitoring frequency from 1Hz to 10Hz to enhance performance tracing. Simultaneously, it sent an instruction to the behavior acquisition unit 2013 to reduce the log sampling rate from 100% to 10% to release resources. Through precise allocation of multimodal data acquisition, the system optimized the overall monitoring load while ensuring the accuracy of key indicator monitoring, constructing a resource-efficient adaptive sampling closed loop.

[0153] In some embodiments, the demand-device collaborative evaluation process includes: the demand analysis unit sending a query instruction for the device computing power boundary to the capability evaluation unit and receiving a latency benchmark value under the federated learning scenario from the capability evaluation unit; the demand analysis unit dynamically adjusting the resource demand matrix according to the latency benchmark value; the demand analysis unit sending the updated resource demand matrix to the decision engine unit; and the decision engine unit sending the updated resource demand matrix to the decision coordination unit so that the decision coordination unit can trigger multi-objective optimization based on the updated resource demand matrix.

[0154] Figure 12 This is a flowchart of the requirements-device collaboration evaluation process according to an embodiment of this disclosure. Figure 12 As shown, the demand analysis unit 2022 initiates a computing power boundary query to the capability assessment unit 2023 to obtain the baseline value of ResNet50 inference latency in the federated learning scenario. Based on device performance feedback, the demand analysis unit 2022 dynamically adjusts the resource demand matrix, increasing the GPU core count requirement by 20%. The updated resource demand matrix is ​​injected into the decision engine unit 2025, and after multi-objective optimization game by the decision coordination unit 2026, a balanced solution that takes into account a 5% improvement in QoS and a 2% increase in cost is generated, completing the closed-loop collaborative assessment of "demand-device" capabilities.

[0155] In some embodiments, after step 1042, further comprising:

[0156] Step 1042A, the anomaly detection unit detects that the storage input / output interface has an anomaly, and sends an emergency response instruction to the feedback control unit.

[0157] Step 1042B, the feedback control unit sends a simulation start instruction to the simulation unit after receiving the emergency response instruction.

[0158] Step 1042C, the simulation unit simulates the isolation of the fault node based on the simulation start instruction, verifies the migration feasibility, and sends the migration feasibility to the feedback control unit.

[0159] Step 1042D, the feedback control unit sends a hot migration execution instruction to the resource allocation module according to the migration feasibility.

[0160] In specific implementation, Figure 13 is a flowchart of the self-healing closed-loop process triggered by the anomaly detection of the embodiments of the present disclosure. As Figure 13 shown, the anomaly detection unit 2042 detects that the storage I / O delay is abnormal, triggering the feedback control unit 2043 to start the emergency response. The feedback control unit 2043 calls the simulation unit 2045 to build a sandbox environment, simulates the isolation of the fault node and verifies the migration feasibility. The simulation result shows that "the application instance can be migrated to the standby node without feeling within 60 seconds", and the feedback control unit 2043 immediately sends an instruction to the fault-tolerant migration engine unit 2034 to execute the hot migration operation. After the migration is completed, the storage I / O competition alarm is automatically released, the system returns to the steady state, and the full-link self-healing closed loop of "detection-simulation-repair" is completed.

[0161] Through the above scheme, after the anomaly detection unit 2042 detects that the storage input / output interface has an anomaly, and the simulation unit 2045 verifies the migration feasibility, the feedback control unit sends a hot migration execution instruction to the fault-tolerant migration engine unit 2034. In this way, after the fault-tolerant migration engine unit 2034 executes the hot migration operation, the alarm of the storage input / output interface anomaly will be automatically released, the system returns to the steady state, and the full-link self-healing closed loop of "detection-simulation-repair" is completed.

[0162] In some embodiments, step 1044 comprises:

[0163] Step 1044A, the learning optimization unit trains the anomaly detection model in the anomaly detection unit based on federated learning to obtain an updated anomaly detection model, and sends the updated anomaly detection model to the simulation unit.

[0164] Step 1044B, the simulation unit detects the synthetic fault by using the updated anomaly detection model to detect the fault in the simulation environment, and sends the synthetic fault to the monitoring instrument unit.

[0165] Step 1044C, the monitoring instrument unit acquires the fault event in real time, and determines the detection time delay according to the synthetic fault and the fault event, and sends the detection time delay to the simulation unit.

[0166] Step 1044D, the simulation unit determines that the updated anomaly detection model converges according to the detection time delay, and then takes the updated anomaly detection model as the trained anomaly detection model.

[0167] In specific implementation, Figure 14 is a flowchart of the simulation verification process of the embodiment of the present disclosure. As Figure 14 shown, the learning optimization unit 2044 deploys a new generation of anomaly detection model (F1-score is improved by 15%) to the simulation unit 2045, triggering the simulation environment to inject synthetic faults such as memory leakage. The monitoring instrument unit 2041 captures the simulated fault event in real time, and optimizes the detection delay from 12 seconds to 8 seconds. The verification data is fed back to the simulation unit 2045. After the simulation unit 2045 confirms the model convergence (loss value <0.02) based on the historical data, it feeds back the verification report to the learning optimization unit 2044, completes the closed-loop strategy evolution of “model iteration → simulation test → performance evaluation”, and promotes the continuous optimization of anomaly detection capability.

[0168] Through the above scheme, the learning optimization unit 2044 sends the updated anomaly detection model to the simulation unit 2045, the simulation unit 2045 verifies the updated anomaly detection model through simulation, and after confirming that the updated anomaly detection model converges, takes the updated anomaly detection model as the trained anomaly detection model, and feeds back the verification report. In this way, the closed-loop strategy evolution of “model iteration → simulation test → performance evaluation” can be completed, and the continuous optimization of anomaly detection capability can be promoted.

[0169] In some embodiments, the energy efficiency optimization driven load migration process includes: the energy efficiency optimization unit sends a query request for power grid carbon intensity data to the learning optimization unit, and receives a clean energy node list sent by the learning optimization unit; the energy efficiency optimization unit sends a load migration instruction to the fault-tolerant migration engine unit, so that the fault-tolerant migration engine unit confirms SLA compliance according to the load migration instruction.

[0170] Figure 15 is a flowchart of the energy efficiency optimization driven load migration process of the embodiment of the present disclosure. As Figure 15As shown, the energy efficiency optimization unit 2035 triggers the fault-tolerant migration engine unit 2034 to perform cross-region load migration based on the grid carbon emission data request, locks the western clean energy node (carbon intensity <50gCO2 / kWh) through the learning optimization unit 2044, and triggers the fault-tolerant migration engine unit 2034 to perform cross-region load migration, seamlessly transferring 50% of the batch processing tasks to the low-carbon data center. After the migration is completed, the system automatically checks the task completion timing and confirms that the time deviation <2% meets the SLA compliance, achieving the dual-target optimization closed loop of "green computing power scheduling-business continuity guarantee".

[0171] The core inventive points of the embodiments of the present disclosure are:

[0172] 1) Based on multi-modal data fusion, a "device-application-user" three-dimensional perception system is constructed, cross-domain data association modeling is realized combined with a knowledge graph, and the matching accuracy of resource supply and demand is improved.

[0173] 2) Based on a dynamic sampling strategy and a priority adjustment mechanism, an adaptive closed loop is formed, the behavior threshold is adaptively optimized through execution result feedback driving, and a "perception-decision-execution-optimization" cycle is constructed.

[0174] 3) Based on a multi-task learning framework, the demand for heterogeneous resources is predicted synchronously, the key influencing factors are quantified combined with causal reasoning, and the cross-domain resource conflict rate is reduced.

[0175] 4) Based on a graph neural network, a heterogeneous resource topology perception scheduling is realized, and the cross-node communication overhead is reduced through physical resource logical slicing management.

[0176] 5) Based on dual-mode diagnosis and digital twinning, a "monitoring-diagnosis-repair-verification-optimization" self-iterative closed loop is constructed.

[0177] 6) Based on a demand causal graph, a device fingerprint library, and a user behavior graph, a knowledge system is constructed, and the accuracy of abnormal diagnosis is improved.

[0178] 7) Based on the whole chain of "perception-decision-execution-optimization", an intelligent resource management system is constructed, and the task response time delay and energy consumption overhead are reduced through dynamic frequency adjustment technology.

[0179] The embodiments of the present disclosure can achieve the following technical effects:

[0180] 1) Full-factor perception and accurate decision-making: Based on multi-modal data acquisition technology, the multi-dimensional state of devices, networks, and applications is integrated, a full-factor perception system is formed combined with adaptive sampling (AS) and semantic network construction technology, and a demand evolution model is constructed through a causal reasoning engine to realize intelligent pre-allocation of resources. Through the causal reasoning engine to construct a demand evolution model, intelligent pre-allocation of resources to high-priority tasks is realized, effectively avoiding resource contention and inefficient configuration.

[0181] 2) Dynamic scheduling and resource utilization improvement: Real-time tracking of load dynamics using reinforcement learning algorithms, combined with heterogeneous resource generalization framework (HRGF) and dynamic voltage and frequency scaling technology (DVFS), optimized resource quota allocation strategies through machine learning prediction models. Automatically optimize the allocation of computing resources, maximize global utilization, reduce inefficient energy consumption, and maintain low resource fragmentation.

[0182] 3) Real-time monitoring and anomaly detection: Construct a holographic mirror image of the system based on digital twin technology, combine knowledge graph (KG) and machine learning anomaly detection algorithms to realize real-time index analysis and bottleneck tracing. Real-time analysis of key indicators, intelligent tracing of bottleneck nodes through knowledge graph, and shortening of abnormal response time.

[0183] 4) Rapid adjustment and fault recovery: Activate PID / MPC collaborative control mechanism when equipment fails, complete task migration and resource dynamic reconstruction in real time, reduce business interruption time, and ensure multi-tenant SLA continuity. Real-time task migration through PID / MPC collaborative control mechanism and federated learning framework (FLF), and multi-tenant SLA continuity through game theory modeling.

[0184] 5) User behavior prediction and personalized allocation: Deeply analyze user behavior characteristics based on multi-task learning framework, build QoS priority protection channel to realize resource on-demand adaptation, and automatically build QoS priority protection channel for high-frequency interaction applications to realize resource on-demand precise adaptation.

[0185] 6) Priority scheduling and low latency guarantee: Apply Nash equilibrium game scheduling strategy (non-cooperative game theory framework NCGTF), combine Markov decision process (MDP) to optimize delay-sensitive business response efficiency, and Nash equilibrium game scheduling strategy for key task applications to significantly improve QoS compliance rate and user interaction real-time response efficiency of delay-sensitive business.

[0186] 7) Multi-factor collaborative decision optimization: Integrate application demand, device performance and energy efficiency target to build a joint optimization model, generate a globally optimal solution through dynamic strategy optimization algorithm (DPOA), dynamically generate a globally optimal resource allocation scheme, and significantly increase task processing throughput.

[0187] 8) Heterogeneous resource adaptation and energy efficiency improvement: Rely on unified resource templates to realize intelligent matching of CPU / GPU / FPGA and other heterogeneous units, combine time convolution network (TCN) prediction model and dynamic voltage and frequency scaling (DVFS) to realize energy efficiency step-by-step optimization, and combine dynamic voltage and frequency scaling dynamic frequency adjustment strategy to promote the step-by-step increase of artificial intelligence inference scenario energy efficiency ratio.

[0188] Through the above embodiment, the data collection module collects environment information in the edge cloud environment; wherein, the environment information includes: application load information, device performance information and user behavior information. In this way, the data collection module can construct a three-dimensional perception system of application, device and user based on multi-modal data fusion, and improve the resource supply-demand matching accuracy. The data analysis module analyzes and processes the environment information to obtain application load prediction results, device capability evaluation results and behavior analysis results, and determines an initial resource allocation strategy according to the application load prediction results, the device capability evaluation results and the behavior analysis results, and optimizes the initial resource allocation strategy to obtain a target resource allocation strategy. In this way, the data analysis module can automatically optimize the target resource allocation strategy of the computing resource according to the environment information, realize the maximization of global utilization rate, and realize the on-demand precise adaptation of the computing resource. The resource allocation module allocates the computing resource according to the target resource allocation strategy. The monitoring feedback module monitors the use state of the computing resource in real time, and feeds back the use state to the data analysis module, so that the data analysis module adjusts the target resource allocation strategy according to the use state. In this way, the monitoring feedback module can feed back the use state of the computing resource monitored in real time to the data analysis module, and the data analysis module can adjust the target resource allocation strategy in real time and accurately according to the use state of the computing resource in real time.

[0189] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of the present embodiment can also be applied to a distributed scenario, and be completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only execute one or more steps in the method of the embodiments of the present disclosure, and the multiple devices can interact with each other to complete the method.

[0190] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than those described above and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.

[0191] Based on the same inventive concept, the present disclosure also provides an edge cloud-based resource allocation system corresponding to the method of any of the above embodiments.

[0192] Reference Figure 2 , the edge cloud-based resource allocation system, the system comprising: a data collection module, a data analysis module, a resource allocation module and a monitoring feedback module;

[0193] The data collection module is configured to collect environment information in the edge cloud environment; wherein the environment information includes application load information, device performance information and user behavior information;

[0194] The data analysis module is configured to analyze and process the environment information to obtain application load prediction results, device capability evaluation results and behavior analysis results, determine an initial resource allocation strategy according to the application load prediction results, the device capability evaluation results and the behavior analysis results, and optimize the initial resource allocation strategy to obtain a target resource allocation strategy;

[0195] The resource allocation module is configured to allocate computing resources according to the target resource allocation strategy;

[0196] The monitoring feedback module is configured to monitor the use state of the computing resources in real time, and feed back the use state to the data analysis module, so that the data analysis module adjusts the target resource allocation strategy according to the use state.

[0197] In some embodiments, the data collection module includes an application collection unit, a device collection unit, a behavior collection unit and a collection coordination unit;

[0198] The application collection unit is configured to collect application document data from application software through a multi-modal analysis algorithm and a natural language processing algorithm, collect application running characteristics from application software through an application programming interface, and use the application document data and the application running characteristics as application load information;

[0199] The device collection unit is configured to collect device storage parameters of edge devices through a probe, collect device physical states of edge devices through visual analysis devices, and use the device storage parameters and the device physical states as device performance information;

[0200] The behavior collection unit is configured to collect interaction logs of user operation records, construct an operation timing mode graph according to the interaction logs, determine predicted user behavior according to the timing mode graph through a timing attention algorithm, detect abnormal user behavior in real time according to the timing mode graph through an isolation forest algorithm, and use the predicted user behavior and abnormal user behavior as user behavior information;

[0201] The collection coordination unit is configured to perform spatio-temporal correlation modeling on the application load information, the device performance information and the user behavior information to obtain the environment information, and adjust the sampling mode of the application collection unit, the device collection unit and the behavior collection unit according to the multi-modal fusion results of the environment information.

[0202] In some embodiments, the data analysis module comprises: a preprocessing unit, a demand analysis unit, a capability evaluation unit, a behavior analysis unit, a decision engine unit and a decision coordination unit;

[0203] The preprocessing unit is configured to preprocess the environment information to obtain standard data, and parse the standard data to obtain application load information, device performance information and user behavior information;

[0204] The demand analysis unit is configured to parse the application load information to obtain predicted storage demand and application document data, determine predicted load demand from the application load information through a pre-constructed load prediction model, and detect abnormal demand from the application load information through a pre-constructed anomaly detection model, and take the predicted storage demand, the predicted load demand and the abnormal demand as an application load prediction result;

[0205] The capability evaluation unit is configured to determine a hardware computing power indicator of an edge device based on a pre-constructed device performance benchmark library according to the device performance information, and determine a solid state disk predicted life of the edge device based on a long short-term memory network according to the device performance information, and take the hardware computing power indicator and the solid state disk predicted life as a device capability evaluation result;

[0206] The behavior analysis unit is configured to determine a predicted click rate from the user behavior information through a pre-constructed user prediction model, determine an association pattern between a user and an application from the user behavior information through a graph attention network, determine a predicted access space-time from the user behavior information through a spatio-temporal graph convolution network, and update a user portrait to obtain a behavior analysis result according to the predicted click rate, the association pattern and the predicted access space-time;

[0207] The decision engine unit is configured to perform resource allocation to obtain an initial resource allocation strategy according to the application load prediction result, the device capability evaluation result and the behavior analysis result through multi-level optimization;

[0208] The decision coordination unit is configured to optimize the initial resource allocation strategy to obtain a target resource allocation strategy based on a pre-constructed multi-objective optimization engine.

[0209] In some embodiments, the resource allocation module comprises: a heterogeneous resource generalization unit, a reinforcement learning scheduling unit, a pre-allocation unit, a fault-tolerant migration engine unit and an energy efficiency optimization unit;

[0210] The heterogeneous resource generalization unit is configured to determine a resource dependency relationship based on a graph neural network, and manage the computing resources according to the resource dependency relationship;

[0211] The reinforcement learning scheduling unit is configured to construct a scheduling strategy of the target resource allocation strategy based on a proximal policy optimization algorithm, determine an execution time of each task in the target resource allocation strategy through a long short-term memory network, optimize the scheduling strategy according to the execution time to obtain an optimized scheduling sequence, and schedule the computing resources according to the scheduling sequence.

[0212] The pre-allocation unit is configured to determine a predicted space-time demand according to the target resource allocation strategy through a time series prediction model, calibrate a current allocation ratio of the target resource allocation strategy based on a proportional-integral-derivative controller to obtain a target allocation ratio, and adjust the scheduling of the computing resources according to the target allocation ratio.

[0213] The fault-tolerant migration engine unit is configured to detect an anomaly of the edge device through an isolation forest algorithm and a long short-term memory network, and repair the anomaly of the edge device.

[0214] The energy efficiency optimization unit is configured to determine a predicted device power consumption based on a pre-constructed device power consumption model, adjust a voltage frequency of a power grid according to the predicted device power consumption, and adjust the scheduling of the computing resources according to the voltage frequency of the power grid and carbon intensity data.

[0215] In some embodiments, the monitoring feedback module includes a monitoring instrument unit, an anomaly detection unit, a feedback control unit, a learning optimization unit, and a simulation unit.

[0216] The monitoring instrument unit is configured to monitor a usage state of the computing resources in real time, and construct an association relationship among the edge device, the application software, and the computing resources.

[0217] The anomaly detection unit is configured to detect an anomaly of the computing resources according to the usage state through a pre-trained anomaly detection model to obtain a time series detection result and a graph detection result, and take the time series detection result and the graph detection result as an anomaly detection result.

[0218] The feedback control unit is configured to repair a key indicator of the edge device according to the anomaly detection result based on a proportional-integral-derivative controller to obtain a device parameter, and send the device parameter to the data analysis module, so that the data analysis module adjusts the target resource allocation strategy according to the usage state.

[0219] The learning optimization unit is configured to train an anomaly detection model in the anomaly detection unit based on federated learning to obtain a trained anomaly detection model.

[0220] The simulation unit is configured to verify the target resource allocation strategy by using a pre-constructed digital twin simulation model.

[0221] In some embodiments,

[0222] The anomaly detection unit is further configured to detect that there is an anomaly in the storage input / output interface, and send an emergency response instruction to the feedback control unit.

[0223] The feedback control unit is further configured to send a simulation start instruction to the simulation unit after receiving the emergency response instruction.

[0224] The simulation unit is further configured to simulate isolation of a fault node based on the simulation start instruction, verify migration feasibility, and send the migration feasibility to the feedback control unit.

[0225] The feedback control unit is further configured to send a hot migration execution instruction to the resource allocation module according to the migration feasibility.

[0226] In some embodiments,

[0227] The learning optimization unit is further configured to train an anomaly detection model in the anomaly detection unit based on federated learning to obtain an updated anomaly detection model, and send the updated anomaly detection model to the simulation unit.

[0228] The simulation unit is further configured to detect a synthetic fault in a simulation environment by using the updated anomaly detection model, and send the synthetic fault to the monitoring instrument unit.

[0229] The monitoring instrument unit is further configured to acquire a fault event in real time, determine a detection time delay according to the synthetic fault and the fault event, and send the detection time delay to the simulation unit.

[0230] The simulation unit is further configured to determine that the updated anomaly detection model converges according to the detection time delay, and use the updated anomaly detection model as the trained anomaly detection model.

[0231] For the sake of convenience, the above-described apparatus is described in various modules in terms of functions. Of course, the functions of each module can be implemented in one or more software and / or hardware when implementing the present disclosure.

[0232] The apparatuses of the above-described embodiments are used to implement the corresponding resource allocation method based on edge cloud in any of the preceding embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be described here.

[0233] Based on the same inventive concept, the disclosure also provides an electronic device corresponding to any of the above-mentioned embodiment methods, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the edge cloud-based resource allocation method of any of the above-mentioned embodiments.

[0234] Figure 16 A more specific hardware structure of an electronic device provided by the embodiment is shown, which can include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 for communication within the device.

[0235] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, etc., for executing related programs to implement the technical solutions provided by the embodiments of the present disclosure.

[0236] The memory 1020 can be implemented by a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs, and when the technical solutions provided by the embodiments of the present disclosure are implemented by software or firmware, the related program codes are stored in the memory 1020 and executed by the processor 1010.

[0237] The input / output interface 1030 is used to connect input / output modules to realize information input and output. The input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0238] The communication interface 1040 is configured to connect a communication module (not shown in the figure) to realize the communication interaction between the device and other devices. The communication module can realize communication through wired mode (for example, USB (Universal Serial Bus), network cable, etc.), or can realize communication through wireless mode (for example, mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).

[0239] The bus 1050 includes a path for transmitting information between various components (for example, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040) of the device.

[0240] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only contain the components necessary for the implementation of the embodiments of the present specification, and does not have to contain all the components shown in the figure.

[0241] The electronic device of the above embodiment is used to realize the corresponding edge cloud-based resource allocation method in any of the preceding embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here.

[0242] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the edge cloud-based resource allocation method according to any of the above embodiments.

[0243] The computer-readable medium of the present embodiment includes permanent and non-permanent, removable and non-removable media, which can be realized by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0244] The computer instructions stored in the storage medium of the above embodiments are used to make the computer execute the edge cloud-based resource allocation method according to any one of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which are not described here again.

[0245] Based on the same inventive concept, the present application also provides a computer program product corresponding to the method of any of the above embodiments, comprising computer program instructions, which, when executed on a computer, make the computer execute the edge cloud-based resource allocation method according to any one of the above embodiments, have the beneficial effects of the corresponding method embodiments, which are not described here again.

[0246] It can be understood that, before using the technical solutions of various embodiments in the present disclosure, the type, use range, use scenario, etc. of the personal information involved will be informed to the user in a proper manner, and the authorization of the user will be obtained.

[0247] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be executed will require the acquisition and use of personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware such as electronic devices, application programs, servers or storage media that execute the technical solutions of the present disclosure according to the prompt information.

[0248] As an optional but not limited implementation manner, in response to accepting the active request of the user, the manner of sending prompt information to the user may, for example, be a pop-up window manner, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0249] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0250] Those skilled in the art will understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure is limited to these examples; under the idea of the present disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes of different aspects of the embodiments of the present disclosure as described above. In order to be brief, they are not provided in details.

[0251] Additionally, to simplify the description and discussion, and so as not to obscure the understanding of the embodiments of the disclosure, the well-known power / ground connections of integrated circuits (ICs) and other components can or can not be shown in the provided figures. Furthermore, devices can be shown in block diagram form in order to avoid obscuring the understanding of the embodiments of the disclosure, and in keeping with the fact that the specific details of the implementation of such block devices are highly dependent on the platform within which the embodiments of the disclosure are to be implemented (i.e., these details should be well within the understanding of one of ordinary skill in the art). Where specific details are set forth in order to describe an example embodiment of the disclosure, it should be apparent to one skilled in the art that the embodiments of the disclosure can be practiced without, or with variation of, these specific details. The description is thus to be considered as illustrative and not restrictive, and the only purpose is to exemplify the principles of the disclosure.

[0252] While the disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications and variations will be apparent to those skilled in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) can use the embodiments discussed.

[0253] The embodiments of the disclosure are intended to cover all such alternatives, modifications and variations as falling within the scope of the disclosure. Accordingly, any omission, modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the disclosure should be included in the protection scope of the disclosure.

Claims

1. A method for resource allocation based on edge cloud, characterized in that, The application is applied to a resource allocation system based on an edge cloud, and the system comprises a data collection module, a data analysis module, a resource allocation module and a monitoring feedback module; the method comprises: The data collection module collects environmental information in an edge cloud environment; wherein the environmental information comprises application load information, device performance information and user behavior information; The data analysis module analyzes and processes the environmental information to obtain application load prediction results, device capability evaluation results and behavior analysis results, determines an initial resource allocation strategy according to the application load prediction results, the device capability evaluation results and the behavior analysis results, and optimizes the initial resource allocation strategy to obtain a target resource allocation strategy; The resource allocation module allocates computing resources according to the target resource allocation strategy; The monitoring feedback module monitors the use state of the computing resources in real time and feeds back the use state to the data analysis module, so that the data analysis module adjusts the target resource allocation strategy according to the use state.

2. The method of claim 1, wherein, The data collection module comprises an application collection unit, a device collection unit, a behavior collection unit and a collection coordination unit; The data collection module collects environmental information in an edge cloud environment, comprising: The application collection unit collects application document data from application software through a multi-modal analysis algorithm and a natural language processing algorithm, collects application running characteristics from application software through an application programming interface, and takes the application document data and the application running characteristics as application load information; The device collection unit collects device storage parameters of edge devices through a probe, collects device physical states of edge devices through a visual analysis device, and takes the device storage parameters and the device physical states as device performance information; The behavior collection unit collects interactive logs of user operation records, constructs an operation timing mode graph according to the interactive logs, determines predicted user behaviors according to the timing mode graph through a timing attention algorithm, detects abnormal user behaviors in real time according to the timing mode graph through an isolation forest algorithm, and takes the predicted user behaviors and abnormal user behaviors as user behavior information; The collection coordination unit performs spatio-temporal correlation modeling on the application load information, the device performance information and the user behavior information to obtain the environmental information, and adjusts the sampling mode of the application collection unit, the device collection unit and the behavior collection unit according to the multi-modal fusion result of the environmental information.

3. The method of claim 1, wherein, The data analysis module comprises a preprocessing unit, a demand analysis unit, a capability evaluation unit, a behavior analysis unit, a decision engine unit and a decision coordination unit; The data analysis module analyzes and processes the environmental information to obtain application load prediction results, device capability evaluation results and behavior analysis results, determines an initial resource allocation strategy according to the application load prediction results, the device capability evaluation results and the behavior analysis results, and optimizes the initial resource allocation strategy to obtain a target resource allocation strategy, comprising: The preprocessing unit pre-processes the environmental information to obtain standard data, and analyzes the standard data to obtain application load information, device performance information, and user behavior information; The demand analysis unit analyzes the application load information to obtain predicted storage demand and application document data, determines predicted load demand from the application load information through a pre-constructed load prediction model, detects abnormal demand from the application load information through a pre-constructed anomaly detection model, and takes the predicted storage demand, the predicted load demand, and the abnormal demand as an application load prediction result; The capability evaluation unit determines a hardware computing power index of an edge device from the device performance information based on a pre-constructed device performance benchmark library, and determines a solid state disk predicted life of the edge device from the device performance information based on a long short-term memory network, and takes the hardware computing power index and the solid state disk predicted life as a device capability evaluation result; The behavior analysis unit determines a predicted click rate from the user behavior information through a pre-constructed user prediction model, determines an association mode between a user and an application from the user behavior information through a graph attention network, determines a predicted access space-time from the user behavior information through a spatio-temporal graph convolution network, and updates a user portrait based on the predicted click rate, the association mode, and the predicted access space-time to obtain a behavior analysis result; The decision engine unit performs resource allocation based on the application load prediction result, the device capability evaluation result, and the behavior analysis result through multi-level optimization to obtain an initial resource allocation strategy; The decision coordination unit optimizes the initial resource allocation strategy based on a pre-constructed multi-objective optimization engine to obtain a target resource allocation strategy.

4. The method of claim 1, wherein, The resource allocation module includes a heterogeneous resource generalization unit, a reinforcement learning scheduling unit, a pre-allocation unit, a fault-tolerant migration engine unit, and an energy efficiency optimization unit; The resource allocation module allocates computing resources based on the target resource allocation strategy, including: The heterogeneous resource generalization unit determines a resource dependency relationship based on a graph neural network, and manages the computing resources based on the resource dependency relationship; The reinforcement learning scheduling unit constructs a scheduling strategy of the target resource allocation strategy based on a proximal policy optimization algorithm, determines an execution time of each task in the target resource allocation strategy through a long short-term memory network, optimizes the scheduling strategy based on the execution time to obtain an optimized scheduling sequence, and schedules the computing resources based on the scheduling sequence; The pre-allocation unit determines a predicted space-time demand based on the target resource allocation strategy through a time series prediction model, calibrates a current allocation ratio of the target resource allocation strategy based on a proportional-integral-derivative controller to obtain a target allocation ratio, and adjusts the scheduling of the computing resources based on the target allocation ratio; The fault-tolerant migration engine unit detects anomalies of an edge device through an isolation forest algorithm and a long short-term memory network, and repairs the anomalies of the edge device. The energy efficiency optimization unit determines predicted device power consumption based on a pre-constructed device power consumption model, adjusts the voltage frequency of the power grid according to the predicted device power consumption, and adjusts the scheduling of the computing resources according to the voltage frequency of the power grid and the carbon intensity data.

5. The method of claim 1, wherein, The monitoring feedback module comprises a monitoring instrument unit, an anomaly detection unit, a feedback control unit, a learning optimization unit, and a simulation unit. The monitoring feedback module monitors the usage state of the computing resources in real time and feeds back the usage state to the data analysis module, so that the data analysis module adjusts the target resource allocation strategy according to the usage state, comprising: The monitoring instrument unit monitors the usage state of the computing resources in real time and constructs the association relationship between the edge device, the application software and the computing resources; The anomaly detection unit detects the computing resources according to the usage state through the pre-trained anomaly detection model to obtain time series detection results and graph detection results, and takes the time series detection results and the graph detection results as anomaly detection results; The feedback control unit repairs the device parameters based on the key indicators of the edge device according to the anomaly detection results through the proportional-integral-derivative controller, and sends the device parameters to the data analysis module, so that the data analysis module adjusts the target resource allocation strategy according to the usage state; The learning optimization unit trains the anomaly detection model in the anomaly detection unit based on federated learning to obtain a trained anomaly detection model; The simulation unit verifies the target resource allocation strategy through a pre-constructed digital twin simulation model.

6. The method of claim 5, wherein, After the anomaly detection unit detects the computing resources according to the usage state through the pre-trained anomaly detection model to obtain time series detection results and graph detection results, and takes the time series detection results and the graph detection results as anomaly detection results, it further comprises: The anomaly detection unit detects that the storage input / output interface is abnormal, and sends an emergency response instruction to the feedback control unit; The feedback control unit sends a simulation start instruction to the simulation unit after receiving the emergency response instruction; The simulation unit simulates the isolation of the fault node based on the simulation start instruction, verifies the migration feasibility, and sends the migration feasibility to the feedback control unit; The feedback control unit sends a hot migration execution instruction to the resource allocation module according to the migration feasibility.

7. The method of claim 5, wherein, The learning optimization unit trains the anomaly detection model in the anomaly detection unit based on federated learning to obtain a trained anomaly detection model, comprising: The learning optimization unit trains the anomaly detection model in the anomaly detection unit based on federated learning to obtain an updated anomaly detection model, and sends the updated anomaly detection model to the simulation unit; The simulation unit detects the fault in the simulation environment using the updated anomaly detection model to obtain a synthetic fault, and sends the synthetic fault to the monitoring instrument unit; The monitoring instrument unit acquires a fault event in real time, determines a detection time delay according to the synthetic fault and the fault event, and sends the detection time delay to the simulation unit; The simulation unit determines that the updated anomaly detection model converges according to the detection time delay, and regards the updated anomaly detection model as the trained anomaly detection model.

8. An edge cloud based resource allocation system, characterized by, The system comprises a data acquisition module, a data analysis module, a resource allocation module and a monitoring feedback module; The data acquisition module is configured to acquire environment information in an edge cloud environment; wherein the environment information comprises application load information, device performance information and user behavior information; The data analysis module is configured to analyze and process the environment information to obtain application load prediction results, device capability evaluation results and behavior analysis results, determine an initial resource allocation strategy according to the application load prediction results, the device capability evaluation results and the behavior analysis results, and optimize the initial resource allocation strategy to obtain a target resource allocation strategy; The resource allocation module is configured to allocate computing resources according to the target resource allocation strategy; The monitoring feedback module is configured to monitor a use state of the computing resources in real time, and feed back the use state to the data analysis module, so that the data analysis module adjusts the target resource allocation strategy according to the use state.

9. An electronic device, comprising: The non-transitory computer readable storage medium stores computer instructions for causing a computer to execute the method of any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium, comprising: The non-transitory computer readable storage medium stores computer instructions for causing a computer to execute the method of any one of claims 1 to 7.

Citation Information

Cited By

  • Dynamic adjustment method and device for traction control of heavy haul train and electronic equipment

    CN121291160A

  • Cross-domain deterministic data center network-oriented service flow multi-target routing method

    CN121396874A

  • Method for multi-objective routing of traffic flows in a cross-domain deterministic data center network

    CN121396874B

  • Wireless communication resource allocation method and system based on artificial intelligence

    CN121397753A

  • Cloud platform container scheduling method and device, electronic equipment and storage medium

    CN121636175A