Cloud machine resource scheduling method and device, electronic equipment and storage medium

By collecting and analyzing user operation data to generate target intent tags, and combining them with spatiotemporal graph neural networks for cloud machine resource scheduling, the problems of resource idleness and shortage in existing technologies are solved, and dynamic differentiated allocation of resources is realized, improving utilization and allocation efficiency.

CN121979671APending Publication Date: 2026-05-05CHINA MOBILE INTERNET CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE INTERNET CO LTD
Filing Date
2025-12-30
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing cloud server resource scheduling solutions fail to effectively differentiate the resource requirements of different applications, resulting in idle or insufficient resources, an inability to cope with load fluctuations caused by user operations, and low resource utilization and allocation efficiency.

Method used

By collecting and analyzing user operation data, accurate target intent labels are generated. Combined with spatiotemporal graph neural networks, resource prediction and scheduling are performed to achieve dynamic and differentiated resource allocation.

Benefits of technology

It improved the utilization and allocation efficiency of cloud server resources, reduced operating costs, and enhanced the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121979671A_ABST
    Figure CN121979671A_ABST
Patent Text Reader

Abstract

The invention provides a cloud machine resource scheduling method and apparatus, an electronic device and a storage medium, relates to the technical field of cloud computing, and provides a basis for cloud resource scheduling by generating a target intention label accurately representing an application scene and a user intention through collection and intelligent analysis of user operation data by an end side. The target intention label can effectively distinguish resource demand differences of different application types and dynamically reflect load changes brought by user operation, so that the cloud can implement dynamic resource allocation matched with real demands according to real-time and refined intention information, and the user experience is improved. Therefore, the problem of resource supply and demand mismatching caused by a static allocation scheme in related technologies is solved. Dynamic alignment of cloud machine resource supply and actual complex requirements is achieved, the overall utilization efficiency and scheduling accuracy of resources are remarkably improved, and therefore resource redundancy and operation cost are effectively reduced while the service quality is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud computing technology, and in particular to a cloud machine resource scheduling method and apparatus, electronic device and storage medium. Background Technology

[0002] Cloud phone technology provides users with a remote mobile experience by virtualizing a mobile operating system in the cloud. Cloud-based applications, as a lightweight form of cloud-based technology, focus on providing single-application services. Current resource scheduling schemes typically employ a static partitioning strategy, which pre-divides physical servers into several equal logical resource units. When a cloud-based application starts, the scheduling system searches the resource pool for idle units with the corresponding application installed and allocates them as a whole.

[0003] However, current resource scheduling schemes ignore the differentiated needs of applications for resources such as the Central Processing Unit (CPU), Graphics Processing Unit (GPU), and memory, uniformly allocating fixed-specification resources. This results in idle resources for low-load applications, while potentially insufficient resources for high-load applications. Simultaneously, it neglects load fluctuations caused by user operations, failing to achieve precise resource allocation. This leads to low resource utilization and allocation efficiency, increasing operating costs and potentially impacting user experience.

[0004] Therefore, how to achieve dynamic and accurate scheduling of cloud server resources to improve resource utilization and allocation efficiency is an urgent problem to be solved. Summary of the Invention

[0005] This application provides a cloud server resource scheduling method, apparatus, electronic device, and storage medium. Its main objective is to address how to achieve dynamic and accurate scheduling of cloud server resources to improve resource utilization and allocation efficiency.

[0006] According to a first aspect of this application, a cloud server resource scheduling method is provided, wherein the method is applied to the edge application layer and includes: Collect operation data generated during the operation of cloud-based applications. The operation data includes time-series behavioral data reflecting the usage habits of cloud-based applications and static data identifying the application type of cloud-based applications. Both behavioral data and static data contain at least one type of operation content. The operation data is classified and labeled to obtain the intent labels corresponding to each type of operation content in the operation data. The intent label and operation data are input into the locally trained intent recognition model for analysis and processing to obtain the target intent label representing the current usage information of the cloud application; The target intent tag is transmitted to the cloud-side service layer so that the cloud-side service layer can perform differentiated resource prediction and resource scheduling based on the target intent tag.

[0007] According to a first aspect of this application, a cloud server resource scheduling method is provided, wherein the method is applied to a cloud-side service layer and includes: It receives target intent tags uploaded by multiple end-side application layers, and obtains local hardware status information and network status information. Spatiotemporal graph data is constructed by using multiple edge application layers as nodes and target intent tags as the temporal change features of the nodes. Spatiotemporal graph data, hardware status information, and network status information are input into a locally trained spatiotemporal network model for spatiotemporal joint modeling to obtain the predicted resource distribution for multiple cloud applications within a preset future time period. Based on the predicted resource distribution, a dynamic resource allocation strategy is executed to allocate differentiated cloud machine resources to multiple cloud applications.

[0008] According to a third aspect of this application, a cloud server resource scheduling device is provided, wherein the device is configured in the edge application layer, comprising: The data collection unit is used to collect operation data generated during the operation of cloud-based applications. The operation data includes time-series behavioral data reflecting the usage habits of cloud-based applications and static data identifying the application type of cloud-based applications. Both behavioral data and static data contain at least one type of operation content. The annotation unit is used to classify and annotate the operation data to obtain the intent label annotations corresponding to each type of operation content in the operation data. The analysis unit is used to input the intent label annotation and operation data into the locally trained intent recognition model for analysis and processing, so as to obtain the target intent label representing the current usage information of the cloud application; The transmission unit is used to transmit the target intent tag to the cloud-side service layer, so that the cloud-side service layer can provide a basis for differentiated resource prediction and resource scheduling based on the target intent tag.

[0009] According to a fourth aspect of this application, a cloud server resource scheduling device is provided, wherein the device is configured in the cloud-side service layer, comprising: The receiving unit is used to receive target intent tags uploaded by multiple end-side application layers. The acquisition unit is used to acquire local hardware status information and network status information; The building unit is used to construct spatiotemporal graph data by using multiple edge application layers as nodes and target intent tags as the temporal change features of the nodes. The modeling unit is used to input spatiotemporal graph data, hardware status information and network status information into the locally trained spatiotemporal network model to perform spatiotemporal joint modeling and obtain the predicted resource distribution for multiple cloud applications within a preset future time period. The allocation unit is used to execute dynamic resource allocation strategies based on the predicted resource distribution, and to allocate differentiated cloud machine resources to multiple cloud applications.

[0010] According to a fourth aspect of this application, a cloud server resource scheduling system is provided, comprising: at least one end-side application layer and a cloud-side service layer. At least one end-side application layer is configured with a cloud machine resource scheduling device as described in the third aspect above. The cloud-side service layer is configured with the cloud machine resource scheduling device as described in the fourth aspect above; In this case, at least one application layer on the client side is connected to the service layer on the cloud side.

[0011] According to a fifth aspect of this application, an electronic device is provided, comprising: At least one processor; and A memory that is communicatively connected to at least one processor; wherein, The memory stores instructions that can be executed by at least one processor, such that the at least one processor is able to perform the method of the first aspect or the method of the second aspect described above.

[0012] According to a sixth aspect of this application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to perform the method of the first aspect or the method of the second aspect.

[0013] According to a seventh aspect of this application, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method as described in the first aspect or the method as described in the second aspect.

[0014] The cloud server resource scheduling method, apparatus, electronic device, and storage medium provided in this application generate target intent tags that accurately represent application scenarios and user intentions through the collection and intelligent analysis of user operation data on the edge, providing a basis for cloud resource scheduling. These target intent tags can effectively distinguish the differences in resource requirements among different application types and dynamically reflect the load changes brought about by user operations. This allows the cloud to implement dynamic resource allocation that matches actual needs based on real-time and refined intent information, thereby overcoming the resource supply-demand mismatch problem caused by static allocation schemes in related technologies. It achieves dynamic alignment between cloud server resource supply and actual complex needs, significantly improving overall resource utilization efficiency and scheduling accuracy, thus effectively reducing resource redundancy and operating costs while ensuring service quality.

[0015] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0016] The accompanying drawings are provided for a better understanding of this solution and do not constitute a limitation of this application. Wherein: Figure 1 A flowchart illustrating a cloud server resource scheduling method provided in an embodiment of this application; Figure 2 This is a schematic diagram illustrating an intent labeling process provided in an embodiment of this application. Figure 3 A schematic diagram of a model federated learning process provided in an embodiment of this application; Figure 4 A flowchart illustrating another cloud server resource scheduling method provided in this application embodiment; Figure 5 This is a schematic diagram of a resource distribution prediction process provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a dynamic resource manipulator provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a cloud server resource scheduling system provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of a cloud server resource scheduling device provided in an embodiment of this application; Figure 9 This is a schematic diagram of another cloud server resource scheduling device provided in an embodiment of this application; Figure 10 This is a schematic diagram of another cloud server resource scheduling device provided in an embodiment of this application; Figure 11 This is a schematic diagram of another cloud server resource scheduling device provided in an embodiment of this application; Figure 12 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0017] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0018] This application relates to at least the interdisciplinary field of cloud computing and artificial intelligence, specifically providing a cloud server resource scheduling method. This method is primarily deployed at the client-side application layer and the cloud-side service layer. Therefore, the method is designed as a two-layer system: the client-side application layer, whose main responsibility is to collect user operation records and predict user intent; and the cloud-side service layer, whose main responsibility is to process the user intent tags collected by the client-side application layer, and combine this with a spatiotemporal graph neural network to predict cloud server resource demands, thereby achieving intelligent resource allocation and improving cloud server resource utilization.

[0019] The client-side application layer refers to the software logic layer that runs on user terminal devices (such as the browser environment of smartphones, tablets, or personal computers). Its core responsibility is to collect user interaction information at close range and perform preliminary intelligent analysis, working in collaboration with the remote cloud-side service layer.

[0020] The cloud-side service layer refers to the software service logic layer running on cloud data centers or edge computing node clusters. Located at the back end of the entire system architecture, it is responsible for aggregating information from massive numbers of terminals and performing global analysis, decision-making, and resource management. Unlike the edge application layer, which focuses on single-user intent recognition, the core mission of the cloud-side service layer is to collaboratively optimize the aggregated multi-user needs and the underlying infrastructure status from a macro perspective, achieving efficient and intelligent global resource scheduling.

[0021] The cloud server resource scheduling method, apparatus, electronic device, and storage medium of this application are described below with reference to the accompanying drawings.

[0022] Figure 1 This is a flowchart illustrating a cloud server resource scheduling method provided in an embodiment of this application.

[0023] like Figure 1 As shown, this method is applied to the terminal application layer and includes the following steps: Step 101: Collect operation data generated during the operation of cloud-based applications. The operation data includes time-series behavioral data reflecting the usage habits of cloud-based applications and static data identifying the application type of cloud-based applications. Both behavioral data and static data contain at least one type of operation content.

[0024] In the embodiments of this application, the aim is to accurately identify user intent by finely analyzing multi-dimensional data generated during the operation of cloud-based applications, thereby providing a key basis for subsequent intelligent scheduling of cloud resources. A cloud-based application refers to an application form based on a cloud computing architecture. Essentially, it deploys the execution environment (including the operating system and the application itself) of traditional mobile or desktop applications on a remote cloud server. Users receive the application's image and sound output through a client (such as a browser or lightweight client) in a streaming manner, and upload local input operations (such as touch, keyboard, and mouse commands) to the cloud, thus obtaining a near-native application experience. A single cloud-based application typically focuses on providing a single application service function, such as cloud gaming, cloud video editing, or cloud office software.

[0025] Operational data is a composite dataset that contains information reflecting users' usage status of cloud applications from different dimensions. Operational data is systematically divided into at least two main categories: time-series behavioral data and static data.

[0026] Temporal behavioral data refers to data sequences that are strongly correlated with time order and can dynamically reflect user operating habits and interaction patterns. The characteristics of temporal behavioral data are that it has a sequential order, each data point carries a timestamp or implicit temporal context, and its change patterns imply user habits, preferences, and real-time intentions. Typical temporal behavioral data may include, but is not limited to, sequences of real-time gesture trajectories of users on cloud application interfaces, sequences of user gaze focus movement on the screen, etc.

[0027] Static data refers to background information that remains relatively fixed or does not change frequently with continuous user interaction during a specific cloud application usage session. Static data is primarily used to identify the application type of a cloud application. Application types are categorized based on the application's core functions and service content; for example, they can be classified as graphics-intensive applications (e.g., 3D games, video rendering software), compute-intensive applications (e.g., scientific computing software), interaction-intensive applications (e.g., office software), or streaming media applications. Static data is typically determined when the user launches or accesses the cloud application.

[0028] Both temporal behavioral data and static data contain at least one type of operational content. Operational content refers to the specific information category represented by the data. For example, in temporal behavioral data, gesture trajectory is one type of operational content, and gaze focus is another. In static data, application identifiers or preset functional category labels can be considered as operational content. By simultaneously collecting behavioral data containing temporal dynamics and static data providing contextual background, a complete data foundation is built for subsequent intent analysis, encompassing both micro-level operational details and macro-level scene cognition.

[0029] Step 102: Classify and label the operation data to obtain the intent labels corresponding to each type of operation content in the operation data.

[0030] In the embodiments of this application, classification labeling is a data preprocessing and information extraction process. Its purpose is to parse, clean and standardize the raw, unstructured operational data according to the category of the operational content to which it belongs, and to assign a symbolic identifier that can initially represent its semantics or potential intent to each type of data, namely intent labeling.

[0031] Intent labeling refers to one or more discrete or continuous labels generated for a specific segment or type of operation. It is a preliminary, intermediate-level abstract representation of the user's behavioral tendencies or purposes implied in the data segment. For example, a continuous gesture trajectory data, after classification and labeling, may be labeled as intent labels such as rapid swiping, fine-touch operation, or continuous pressing; a gaze focus sequence may be labeled as attention concentration, rapid gaze movement, or focusing on a specific area; while static application type data, either on its own or after transformation, can serve as a type of intent label, such as game applications, document processing applications, etc. Then, the temporal behavioral data and static data are combined to determine the final intent label, such as labeling it as launching a game or playing a video.

[0032] Specifically, regarding the annotation of intent tags on operational data, this application provides a flowchart illustrating the process of intent tagging, such as... Figure 2 As shown, the messy raw data is transformed into structured data with clear category attributes and preliminary semantic information through intent labeling.

[0033] Step 103: Input the intent label annotation and operation data into the locally trained intent recognition model for analysis and processing to obtain the target intent label representing the current usage information of the cloud application.

[0034] In the embodiments of this application, the locally trained intent recognition model refers to a machine learning model that has been pre-downloaded and deployed on the user-side device. This model has been trained to understand the complex mapping relationship from operational data to high-level user intent.

[0035] A locally trained intent recognition model can be a temporal model based on a long short-term memory network architecture, which excels at handling sequence data with long-term dependencies. The model receives input from different data sources (corresponding to different operations), performs feature fusion and deep analysis through its internal network structure, and finally outputs a comprehensive, higher-level judgment result, namely the target intent label. The target intent label is the model's final intent judgment for the current user's operation session. For example, the model might combine tags for rapid swiping gestures, distracted gaze, and static tags from video streaming applications, ultimately generating a target intent label indicating that the user is in a fast video browsing mode, where image quality requirements may be lower, but smoothness requirements are high. This label more directly points to the user's potential needs for underlying cloud resources (such as GPU rendering capabilities and network bandwidth).

[0036] Step 104: Transmit the target intent tag to the cloud-side service layer so that the cloud-side service layer can perform differentiated resource prediction and resource scheduling based on the target intent tag.

[0037] In the embodiments of this application, the cloud application layer refers to the service logic layer deployed on a cloud server cluster, which is responsible for receiving massive amounts of information from the client side and performing global analysis and decision-making. The purpose of the client side uploading the target intent tag is to provide the cloud application layer with refined input information guided by the user's real-time intent.

[0038] The cloud application layer, based on received target intent tags from numerous users and combined with global information such as server hardware resource status and network topology, performs differentiated resource prediction and scheduling. Differentiated resource prediction and scheduling means that the cloud no longer adopts a simple, fixed resource allocation strategy, but rather performs forward-looking resource demand prediction and dynamic, personalized resource allocation and scheduling based on the resource requirements implied by user intent (e.g., gaming applications require more GPU resources, while video conferencing applications require stable bandwidth and CPU resources), and the differences in needs that different users may have even when using the same application due to different habits (e.g., some users prefer high image quality, while others are more sensitive to latency). For example, for users identified as having high-intensity graphical interactive gaming intent, the cloud may reserve or allocate more graphics processing unit resources on the server node they will connect to in advance; for users with document editing and saving intent, it may focus more on allocating stable computing resources and input / output throughput capabilities.

[0039] This application integrates temporal behavioral data and static application type data, giving intent recognition a dual dimension of dynamic habits and static scenarios. This significantly improves the accuracy and comprehensiveness of intent judgment, avoiding potential misjudgments from relying on a single data source. Preliminary intent recognition and tag generation are completed on the device side, leaving the original data rich in user privacy locally and only uploading anonymized abstract intent tags. This protects user data security while greatly reducing the data processing pressure and network transmission load on the cloud. It provides direct, clear, and semantically rich demand guidance for cloud resource scheduling, enabling cloud resource allocation to shift from guesswork to on-demand intelligent allocation. This lays a crucial data foundation for improving the overall utilization of cloud machine resources, reducing operating costs, and enhancing the personalized experience for end users. This constitutes the front-end perception link of an efficient, secure, and intelligent cloud resource scheduling system.

[0040] In one possible implementation of this application embodiment, when collecting operation data generated during the operation of a cloud application, the following methods may be used, but are not limited to: in response to opening the cloud application, directly extracting the static data of the cloud application, the static data including at least the application context information of the cloud application; collecting gesture coordinates during the operation of the cloud application to obtain gesture trajectory information, and encoding and blurring the gesture trajectory information to obtain a gesture trajectory sequence, the temporal behavior data including at least the gesture trajectory sequence and the gaze focus area sequence; collecting eye gaze point data and eye movement data during the operation of the cloud application through an image acquisition device with preset permissions, and generating a gaze focus area sequence based on the eye gaze point data and eye movement data.

[0041] In the embodiments of this application, the data collection process is triggered immediately when a user launches a cloud application by clicking a link, icon, or other interactive method. In this initial stage, the direct extraction of static data is performed first. The launch action not only refers to the start of the application process but also marks the beginning of a user-cloud service interaction session. Responding to this launch event, background information that can be determined without continuous user interaction can be directly obtained from the cloud application's configuration information, loaded metadata, or preset identifiers. The extracted static data includes at least application context information. Application context information defines the basic environment and functional scope of the current cloud application operation. Application context information may include, but is not limited to: a unique identifier for the cloud application, preset functional category tags for the cloud application (e.g., 3D real-time rendering games, high-definition video streaming players, document collaborative editing software), descriptions of the underlying resource characteristics required by the cloud application (e.g., whether GPU acceleration, high input / output (I / O) throughput, or low-latency network are required), and preset configuration parameters for the cloud application to adapt to different end-user environments. Application context information is determined when a cloud application starts up, providing an indispensable scenario framework for understanding subsequent temporal behavioral data. This allows the analysis of dynamic behaviors such as gestures or gaze to be conducted within the correct application functional context, avoiding ambiguity.

[0042] Simultaneously, it will initiate fine-grained capture of user interaction behavior. For manual operations, it will continuously collect the raw gesture coordinates generated by the user during the operation of the cloud application. Gesture coordinates refer to the real-time position data of the user's operation points reported by the input device on the touch screen or the interaction plane simulated by the mouse. They are usually in the form of (x, y) coordinate pairs and corresponding timestamps. By continuously recording gesture coordinates, gesture trajectory information describing the user's operation path is obtained. Gesture trajectory information is a temporal set of raw coordinate points, intuitively reflecting the dynamic characteristics of the user's finger or cursor movement path, speed, acceleration, and pause. However, directly transmitting or using raw coordinates may pose a privacy risk (potentially leading to the deduction of interface content) and result in inconsistent data granularity. Therefore, the gesture trajectory information is encoded and blurred. Encoding and blurring is a data transformation and desensitization technique that aims to remove the absolutely precise position information of the data while retaining its relative spatial relationships and motion pattern characteristics.

[0043] In one specific implementation, the method for collecting gesture trajectory information can divide the continuous screen coordinate space into uniform blocks (e.g., a virtual grid), mapping all coordinate points falling within the same block to a unique identifier for that block. Furthermore, space-filling curves (such as Hilbert curves) can be used to encode the block identifiers, transforming the two-dimensional spatial relationship into a one-dimensional sequence encoding. This allows for the representation of the approximate direction of gesture movement and area transfer patterns while concealing precise locations. After this processing, the original coordinate sequence is transformed into a standardized, desensitized symbol sequence, namely, the gesture trajectory sequence. The gesture trajectory sequence is a sequence of encoded symbols representing different screen areas or movement patterns arranged chronologically; it is a secure and effective abstract representation of the original gesture information. The temporal behavioral data includes at least this gesture trajectory sequence and the gaze focus area sequence described below, together forming a timeline characterizing user behavior from tactile and visual attention dimensions.

[0044] To achieve a deeper understanding of intent, visual attention capture is further introduced by invoking an image acquisition device on the terminal with pre-defined permissions to collect richer biometric behavioral data. The image acquisition device typically refers to the terminal's front-facing camera. Obtaining pre-defined permissions means that before executing this step, explicit user consent to the use of the camera for specific data collection must be obtained through a standardized user authorization process (e.g., a system-level permission pop-up request). This ensures the compliance of the technology implementation and protects user privacy. Under authorization, the image acquisition device is invoked to collect eye gaze point data and eye movement data during the user's operation of cloud applications. Eye gaze point data refers to the location of the user's visual attention focus on the screen, estimated using computer vision algorithms (such as the pupil-corneal reflex method), and is usually represented in coordinate form. Eye movement data includes microscopic patterns of eye movement, such as the type, speed, and amplitude of fixation (the eye remaining relatively still in one place), saccades (the eye rapidly jumping between two points), and tracking movements (the eye smoothly following a moving target). By integrating time-synchronized eye fixation point data and eye movement data, a higher-level sequence of gaze focus regions is generated through analytical algorithms (e.g., clustering consecutive fixations and combining them with saccade path analysis).

[0045] The sequence of visual focus areas is not a series of coordinate points, but rather a temporal logic representing the shift of a user's visual attention between different functional areas or content blocks on the screen. For example, the sequence might be: [Area A (menu bar) gaze for 500ms] - [quick scan] - [Area B (video playback window) gaze for 2000ms] - [Area C (comments section) brief gaze for 300ms].

[0046] This application extracts application context information directly at startup, providing accurate and immediate scene anchors for all subsequent behavioral data analysis. This allows the same gesture or gaze pattern to be correctly interpreted as different intentions in different application types (e.g., rapid clicking in a game might be an attack, while in office software it might be a selection), greatly improving the scene relevance and accuracy of intent recognition. Secondly, gesture coordinates are encoded and blurred, embedding a privacy protection mechanism at the source of data utilization. This transforms raw coordinates that may contain sensitive interface information into symbol sequences that cannot reconstruct specific content but retain behavioral patterns, achieving data usability without visibility and providing a technical path for processing sensitive data on the client side. By integrating gaze focus analysis into a standardized authorization process, a user attention dimension that is difficult to capture with traditional interaction data is introduced. This allows for differentiation between active, detailed operations and passive browsing, thus more accurately determining the intensity and type of the user's actual resource needs. This multimodal (application context, gesture, gaze) and multi-level (raw signal, processed sequence) data collection constitutes a rich perceptual foundation.

[0047] In one possible implementation of this application embodiment, when classifying and labeling operation data, the following methods can be used, but are not limited to: performing behavior pattern recognition on each type of operation content in the operation data to obtain the behavior pattern corresponding to each type of operation content; and labeling each type of operation content based on the behavior pattern to obtain the intent label label corresponding to each type of operation content.

[0048] In the embodiments of this application, behavior pattern recognition is the process of extracting typical action or state features that are statistically regular, repeatable, and rich in semantic information from raw or pre-processed data streams. For time-series behavior data, such as gesture trajectory sequences, behavior pattern recognition may identify specific behavior patterns such as rapid one-way swiping, precise multi-touch, long press and drag, and irregular shaking by analyzing the speed, acceleration, trajectory shape (such as straight lines, curves, and loops) of coordinate changes and pause intervals. For static data, its behavior patterns may be directly related to the typical resource usage patterns implied by its application type, such as graphics rendering intensive, real-time interactive response, or background computing intensive. Behavior patterns are a higher level of abstraction than raw data; they strip away the specific numerical details of the data while retaining the essential features that can distinguish different user operating habits or application scenarios.

[0049] Behavioral pattern-based labeling is a relatively technical process of mapping pattern features to more semantic and business-guided symbolic labels. This process assigns appropriate intent labels to each identified behavioral pattern based on predefined or learned mapping rules. For example, when a user's gesture trajectory reveals rapid one-way swipes and brief, scattered clicks, and their gaze sequence reveals rapid browsing, the intent label for this data segment might be "rapid content browsing" based on this combination of patterns. For static application context information, the corresponding behavioral patterns can be directly or after simple transformation to be labeled with intent labels such as those for applications with high GPU demands. The labeling process imbues the data with explicit semantic information that can be directly understood and processed by machine learning models. Intent labeling is an intermediate interpretation of a user's possible intent within the corresponding time period of a data segment, derived after behavioral pattern analysis.

[0050] This application decomposes the complex annotation task into feature extraction (pattern recognition) and semantic mapping (annotation), allowing for independent optimization of each step. Annotation based on behavioral patterns ensures the consistency, objectivity, and repeatability of the annotation results. Pattern recognition eliminates noise and irrelevant details from the original data, enabling the same user intent to be identified as the same behavioral pattern across different specific operational data, thus obtaining consistent labels and significantly improving annotation quality and reliability. The resulting intent labels are rich in refined structured information, directly reflecting the essential characteristics of user interaction. This enhances the efficiency and accuracy of the entire edge intent recognition process.

[0051] In one possible implementation of this application embodiment, when inputting intent label annotation and operation data into a locally trained intent recognition model for analysis and processing, the following methods can be used, but are not limited to: preprocessing and extracting features from temporal behavior data to obtain temporal feature vectors, and performing feature encoding on static data to obtain type feature vectors; inputting the temporal feature vectors into a Long Short-Term Memory (LSTM) network, updating the hidden state of the temporal feature vectors through the gating mechanism of the LSTM network to obtain a target hidden state vector, wherein the intent recognition model includes an LSTM network; concatenating the target hidden state vector and type feature vector to obtain a target feature vector, and performing classification calculation on the target feature vector to obtain a target intent label.

[0052] In the embodiments of this application, preprocessing aims to standardize the original or pre-annotated time-series data to eliminate dimensional differences, handle missing values, and make it suitable for model input. This may include, but is not limited to, normalization, sequence alignment, and imputation. Feature extraction, based on preprocessing, extracts a subset of information or transformed representations that effectively characterize the inherent patterns of the time-series data. For example, it extracts movement speed, acceleration statistics, and direction change frequency from gesture trajectory sequences, and gaze duration, saccade speed, and region shift probability from gaze sequences. After this operation, the original time-series behavioral data is transformed into a numerical, information-rich time-series feature vector. The time-series feature vector is a multi-dimensional vector, where each dimension represents a key feature extracted from the original time-series data. This vector encapsulates the dynamic characteristics of user behavior over time. In parallel, static data is feature-encoded. Feature encoding is the process of converting non-numerical or categorical static information (such as application type labels) into a numerical form that can be directly processed by machine learning models. Common methods include one-hot encoding and embedding encoding. Through feature encoding, static data is transformed into a type feature vector. The type feature vector is a fixed-length numerical vector that represents the inherent, time-invariant attributes and contextual information of the cloud application itself.

[0053] The gating mechanism mainly consists of an input gate, a forget gate, and an output gate: the input gate determines how much new information needs to be stored in the cell state; the forget gate determines how much old information needs to be discarded from the previous cell state; and the output gate determines the hidden state content based on the current cell state. At each time step, the cell state and hidden state are updated internally through the gating mechanism calculation based on the current input temporal feature vector and the hidden state of the previous time step. The hidden state is the vector output by the Long Short-Term Memory (LSTM) network at each time step, containing the LSM network's memory and understanding of the processed sequence information up to the current time step. As the sequence is processed step by step, the hidden state is continuously updated. When the entire temporal feature vector sequence has been processed, the hidden state output at the final time step is adopted as the target hidden state vector. The target hidden state vector is a contextual representation that integrates all temporal behavioral information throughout the entire observation period; it encodes the temporal patterns and evolution of user operating habits.

[0054] Concatenation is a feature fusion method that involves linking two or more vectors end-to-end in a dimension to create a new vector with a higher dimension. Through concatenation, the target hidden state vector representing dynamic behavioral temporal patterns and the type feature vector representing static application background are integrated to form a comprehensive target feature vector. This target feature vector simultaneously contains both the temporal dynamic information of user operations and the static context information of the application scenario in which the user operates.

[0055] Classification calculations are typically performed by a fully connected layer (possibly combined with a softmax or sigmoid activation function). Its function is to map high-dimensional feature vectors to a predefined intent category space and calculate the probability distribution of each potential intent category. Based on this probability distribution (e.g., selecting the category with the highest probability), the final judgment result, i.e., the target intent label, is output.

[0056] Specifically, when determining the target intent label, temporal behavioral data and static data are distinguished and classified. Then, the intent recognition model preprocesses and extracts features from the input data. For gesture data, normalization, sequence alignment, and padding are performed. For gaze data, masking, sequence alignment, and padding are performed. Static data is feature-encoded and concatenated into a long static feature vector. The model architecture uses a temporal model to process gesture and gaze sequences separately, extract high-level temporal features, and fuse the extracted temporal features with static application context features. Intent classification is then performed based on the fused features.

[0057] This application employs targeted feature extraction and encoding for both temporal behavior and static data, respecting the heterogeneity of the data and ensuring that the most relevant and discriminative information can be extracted from various data types. By introducing a Long Short-Term Memory (LSTM) network to process temporal features, the model can effectively capture complex temporal dependencies and long-term patterns in user operations, enhancing the depth of intent recognition's understanding of the coherence of user behavior. Finally, by concatenating and fusing the output deep temporal representation with static application background features, the final classification decision simultaneously considers dynamic behavioral details and static scene constraints, resulting in more accurate, comprehensive, and consistent target intent labels that align with actual business logic. This forms a hierarchical, fine-grained to macroscopic inference chain, enhancing the ability of edge devices to independently complete complex intent recognition.

[0058] In one possible implementation of this application embodiment, it is necessary to effectively learn and optimize the intent recognition model to obtain a locally trained intent recognition model that can accurately infer user intent. Specifically, the following methods may also be used, but are not limited to: obtaining training operation data including real intent labels, training the intent recognition model based on the training operation data, and obtaining a locally trained intent recognition model; wherein, training the intent recognition model includes: using classification cross-entropy as a loss function, calculating the difference between the predicted intent label and the real intent label predicted by the intent recognition model, and updating the model parameters of the intent recognition model using the backpropagation algorithm.

[0059] In the embodiments of this application, the training operation data is a collection of fully annotated historical operation data samples, with the same format as the real-time collected operation data, including time-series behavioral data and static data. Crucially, each training sample is associated with a true intent label. The true intent label is pre-annotated by experts, verified through other reliable methods, or deduced from subsequent explicit user behavior, and is considered to be the correct intent category.

[0060] Training is an iterative optimization process whose core objective is to adjust the model's internal parameters (including but not limited to the weight matrix and bias vector) so that the model's predicted output on the training data is as close as possible to the corresponding true intention label. The classification cross-entropy loss function is a standard metric in machine learning used to measure the difference between the probability distribution predicted by the model and the true label distribution. The loss value is zero when the model's predictions are completely correct; the more inaccurate the predictions, the larger the loss value.

[0061] Backpropagation is an efficient method for calculating the gradient of the loss function with respect to each parameter in the model. Its working principle is as follows: First, forward propagation is performed, where training data is input and the predicted intent label and corresponding loss value are calculated. Then, starting from the output layer, the algorithm calculates the partial derivatives (i.e., gradients) of the loss value with respect to each layer's parameters. These gradients indicate the direction and magnitude of parameter adjustments to reduce the overall loss. Using the calculated gradients, the optimizer (such as stochastic gradient descent or its variants) updates the model parameters according to preset rules such as the learning rate. Through repeated iterations of forward propagation, loss calculation, backpropagation, and parameter updates on a large amount of training data, the model's internal parameters are continuously adjusted, eventually converging to a better state, thus obtaining a locally trained intent recognition model.

[0062] In one specific implementation, the training of the model can be carried out in, but is not limited to, the following manner: During the training process, because there are multiple types of data input, categorical cross-entropy loss is used to process the data. The formula for the categorical cross-entropy loss function is: . Batch size The total number of intent categories, For the sample True intent label (if it is a category) If the value is 1, then it is 1; otherwise, it is 0, which is one-hot encoding. That is, predicting the category label, predicting the sample Category The probability of.

[0063] For a Long Short-Term Memory (LSTM) cell (taking the gaze branch as an example), at each time step t, it receives the current input and the hidden state and cell state of the previous time step, and calculates the new state.

[0064] During training, backpropagation, using a time-based algorithm, updates the weights and biases of all models based on the gradients calculated from the loss function. Finally, accuracy is used as the evaluation metric to derive the user intent recognition label.

[0065] In one possible implementation of this application embodiment, after the model is trained locally on the device side, it needs to perform secure and efficient information interaction and model synchronization with the cloud-side service layer. This allows the intent recognition model deployed on the device side to continuously absorb collective wisdom and constantly improve its recognition accuracy and generalization ability. Specifically, the following methods can also be used, but are not limited to: obtaining parameter update information of the intent recognition model during model training and encrypting the parameter update information to obtain encrypted update information; uploading the encrypted update information to the cloud-side service layer; receiving aggregated model parameters from the cloud-side service layer and updating the locally trained intent recognition model based on the aggregated model parameters, wherein the aggregated model parameters are obtained by the cloud-side service layer through aggregation and updating based on the encrypted update information.

[0066] In the embodiments of this application, parameter update information refers to the change in the model's current parameters compared to the parameters before training (i.e., gradient information) after a round or stage of local training, or directly the new model parameters obtained after training. To ensure user privacy and communication security when transmitting this information to the cloud, this parameter update information is encrypted. Encryption involves using cryptographic algorithms (such as symmetric encryption, asymmetric encryption, or homomorphic encryption) to transform the parameter update information, thereby preventing potential information leakage or model reverse engineering attacks. The encrypted data becomes the encrypted update information.

[0067] In the cloud, the cloud-side service layer acts as a coordination center, collecting encrypted update information from numerous end-user application layers. The cloud then performs aggregation on this parameter update information from different users and scenarios, either by decryption (if asymmetric encryption is used, the cloud holds the private key) or directly in encrypted form (if advanced technologies such as homomorphic encryption are used). Aggregation combines scattered model updates, which may contain local data biases, into a global, more universally representative direction for model parameter improvement using algorithms (such as the FedAvg federated averaging algorithm). After aggregation, the cloud generates new, improved model parameters—the aggregated model parameters.

[0068] Finally, the cloud-side service layer distributes the aggregated model parameters to each edge application layer. Upon receiving the parameters, the edge application layer does not directly replace its local model. Instead, it updates its locally trained intent recognition model based on the aggregated model parameters. This update operation typically means replacing or partially adjusting the existing parameters of the local model with the received global parameters, allowing the local model to instantly acquire the experience learned from the global data and achieve synchronous model evolution.

[0069] Specifically, regarding the updating of the locally trained intent recognition model, this application provides a flowchart illustrating a model federated learning process, such as... Figure 3 As shown, after model training is completed, the model training information is uploaded to the server via encrypted gradients. The server aggregates the gradients from each client to update the model parameters, and then sends the updated model parameters to the client. The client receives and updates its local model. This stage is executed periodically, for example, automatically once every 24 hours.

[0070] In one specific implementation, the implementation details of the client-side application layer in this application can be illustrated through, but are not limited to, the following example: In the client-side application layer, the cloud-based mobile application is used by users through configured H5 links, with the business scenario being either a Node.js environment or a browser environment. At this time, it is necessary to obtain the user's operation sequence, which includes various types of operations: real-time user gesture trajectory, user gaze focus area sequence, and application context. For the real-time user gesture trajectory, operation data can be collected using industry-standard methods such as event tracking. After obtaining the event tracking information, the user's original gesture coordinates are processed by block encoding (e.g., mapping to a grid and generating position codes using Hilbert curves) to avoid transmitting the original coordinates. The user's gaze focus can be collected and recorded using the phone's front-facing camera, with user authorization, by analyzing the user's eye movements and gaze points to generate an attention heatmap. As for the application context information, the application type is already distinguished when the user clicks on the configured link, such as video, game, or office software. After obtaining the operation data, the front-end model can be trained using the TensorFlowJS framework. Specifically, after the browser initializes the link, it preloads and deploys a lightweight LSTM model locally on the device. Information collected during the user's use of the cloud application (which can be expanded to support more types in the future) is used as input to the LSTM model. It's important to note that the LSTM model is downloaded in its entirety the first time the user visits the cloud application; subsequent downloads are only updates, avoiding wasting bandwidth resources on every complete download. After the raw data is collected on the device, it needs to be categorized and labeled with intent tags before being fed into the LSTM model.

[0071] After the data is labeled with intent tags, the local data is anonymized before model training begins, followed by the federated learning process. The initial purpose of federated learning on local user data is to obtain more representative data and more diverse user usage analysis, so as to better cover usage data in various user scenarios. This results in more comprehensive and representative training data, reducing the risk of data homogeneity caused by single user or regional limitations. Specifically, each edge-side LSTM model receives the anonymized local data and then adds differential noise during training.

[0072] After training is completed, user intent recognition labels are obtained. At the same time, encrypted gradients are uploaded to the server. The server aggregates the gradients of each user to update the model parameters, and then sends the updated model to the user. The user receives and updates the local LSTM model.

[0073] Figure 4 This is a flowchart illustrating a cloud server resource scheduling method provided in an embodiment of this application.

[0074] like Figure 4 As shown, this method is applied to the cloud-side service layer and includes the following steps: Step 401: Receive target intent tags uploaded by multiple end-side application layers, and obtain local hardware status information and network status information.

[0075] In the embodiments of this application, each target intent tag originates from an independent user terminal (device-side application layer), representing the final determination result of the device-side application layer regarding the user's intent in the current cloud application usage session. It is high-level semantic information extracted after in-depth analysis by the device-side local model. These target intent tags are asynchronously uploaded from multiple devices, constituting the raw signals for the cloud to understand the global user demand situation.

[0076] Meanwhile, the cloud-side service layer also needs to obtain local hardware and network status information. Hardware status information refers to the real-time operational status data of the physical or virtual servers that constitute the cloud server resource pool, which includes at least the CPU utilization, GPU utilization, memory usage, storage I / O load, and the topological location of each server. Network status information describes the quality of the network channels connecting the user end (client-side application layer) and the cloud-side service layer, as well as connecting different cloud servers, which includes at least bandwidth availability, network latency, packet loss rate, and network topology.

[0077] Step 402: Using multiple edge application layers as nodes and target intent tags as the temporal change features of the nodes, construct spatiotemporal graph data.

[0078] In the embodiments of this application, in the spatiotemporal graph data structure, each user terminal using a cloud-based application, or its represented session, is considered a node in the graph. These nodes are not isolated; they may have implicit connections due to geographical proximity, use of the same application, or being in the same network exchange domain. The cloud-side service layer further uses target intent tags as the temporal change characteristics of the nodes. This means that each node is accompanied by a continuously updated sequence of characteristics over time, i.e., a continuously uploaded stream of target intent tags.

[0079] The temporal variation characteristics depict the dynamic evolution of the behavior intentions of a node (end-side). This data structure, which jointly represents nodes with dynamic characteristics and their interrelationships (edges), is called spatiotemporal graph data. Spatiotemporal graph data is a composite data structure that can simultaneously depict spatial relationships between entities (graph structure) and temporal changes in entity attributes (node ​​feature sequences). It is suitable for describing complex systems such as massive user bases and cloud server clusters that are both spatially interconnected and temporally evolving.

[0080] Step 403: Input the spatiotemporal graph data, hardware status information and network status information into the locally trained spatiotemporal network model to perform spatiotemporal joint modeling, and obtain the predicted resource distribution for multiple cloud applications within a preset future time period.

[0081] In the embodiments of this application, the locally trained spatiotemporal network model is a machine learning model pre-deployed on the cloud side, and its core architecture is usually a spatiotemporal graph neural network. This spatiotemporal network model is designed specifically for processing spatiotemporal graph data and can simultaneously capture the dynamic patterns of node features in the time dimension (e.g., certain user groups experience peak gaming intent at fixed times each day) and the diffusion or correlation effects in the spatial dimension (e.g., an overloaded server may affect the experience of users on its neighboring servers, or a popular application concentrated in a specific area leads to local resource contention).

[0082] The model takes the aforementioned information as input and, through its complex internal graph convolution (capturing spatial dependencies) and temporal modeling modules (such as recurrent neural networks, capturing temporal dependencies), outputs a forward-looking analysis result: a predicted resource distribution for multiple cloud applications within a pre-defined future timeframe. This predicted resource distribution is a multi-dimensional demand heatmap or quantitative prediction matrix that indicates the expected demand for various computing resources (such as GPU power, CPU cores, and memory capacity) for different types of cloud applications (such as games, video, and office applications) across different geographical regions and server clusters within a future timeframe (e.g., the next 5 or 15 minutes). This quantitative prediction provides precise guidance for pre-allocation of resources.

[0083] Step 404: Execute a dynamic resource allocation strategy based on the predicted resource distribution to allocate differentiated cloud machine resources to multiple cloud applications.

[0084] In the embodiments of this application, the dynamic resource allocation strategy is a set of decision rules for resource scheduling based on predicted resource distribution and real-time constraints. Its core objective is to proactively and in advance re-plan and allocate computing, storage, and network resources in the cloud machine resource pool based on predicted resource demand hotspots and intensities. This dynamic resource allocation strategy is not a simple equal distribution, but rather allocates differentiated cloud machine resources to multiple cloud applications. This differentiated allocation is reflected in the following ways: for user groups predicted to be entering high-load graphics rendering scenarios, their corresponding cloud application instances will be scheduled or have more high-performance GPU resources reserved; for real-time interactive office applications, priority may be given to ensuring their CPU response speed and low network latency paths; simultaneously, the strategy also considers load balancing to avoid over-concentrating high-demand applications on a few servers. This allocation is dynamic and continuously adjusted with the periodic updates of the prediction results.

[0085] In one specific implementation, the prediction of resource distribution in the cloud-side service layer of this application can be illustrated by, but is not limited to, the following example: After receiving the user intent tag uploaded from the client side, the server invokes the capabilities of the Spatio-Temporal Graph Neural Network (STGNN) framework to predict cloud server resource allocation. This framework integrates neural networks and various temporal learning methods to accurately predict user cloud server resource needs in different scenarios, comprehensively analyzing user intent tags, hardware resource status, and network topology. User behavior intent tags serve as the temporal dimension input to the STGNN engine, covering its periodic operation patterns and analyzing user behavior habits from a temporal perspective. Cloud server service hardware resource status serves as the spatial dimension input to the STGNN engine, analyzing the load balancing between servers. Real-time network status serves as a constraint condition for the STGNN engine, providing resource scheduling paths. Joint modeling and analysis are performed from multiple dimensions including time, space, and network constraints, ultimately outputting a key resource demand heatmap (including predicted resource distribution). The general process is as follows: Figure 5 As shown, Figure 5 This is a flowchart illustrating a resource distribution prediction process provided in this application.

[0086] This application constructs spatiotemporal graph data and uses a spatiotemporal network model for joint modeling, enabling the cloud to understand, with high precision and globality, the periodic and trend changes of user demand over time, as well as the propagation and correlation patterns in space (server topology, network regions), thereby making predictions. By integrating hardware and network status information as inputs and constraints to the model, it ensures that resource prediction and scheduling decisions are not only based on user demand but also closely integrated with the actual supply capacity of infrastructure and network accessibility, making the scheduling scheme both demand-satisfying and physically feasible. Based on accurate predicted resource distribution, a dynamic resource allocation strategy is executed, realizing on-demand allocation and off-peak scheduling of resources, improving the overall utilization and elasticity of the cloud server resource pool, effectively alleviating the contradiction between resource fragmentation and local overload, and improving the smoothness and stability of the end-user experience through forward-looking resource guarantees.

[0087] In one possible implementation of this application embodiment, spatiotemporal joint modeling can be performed in the following ways, but is not limited to: performing graph convolution operation on the target temporal change features of the target node at each time step, aggregating the temporal change features of the neighboring nodes associated with the target node, and generating a target feature sequence with spatial context enhancement for the target node within a continuous time window, wherein the target node is any node in the spatiotemporal graph data; inputting the target feature sequence into the gated recurrent unit of the locally trained spatiotemporal network model to learn the dynamic evolution law in the time dimension and obtain spatiotemporal fusion features; mapping the spatiotemporal fusion features to resource demand prediction values, and generating a resource heat map representing the spatial and temporal distribution of resource demand based on the resource demand prediction values, wherein the resource heat map includes the predicted resource distribution.

[0088] In the embodiments of this application, any node in the spatiotemporal graph data is selected as the target node for the current analysis. Each node is associated with a sequence of target intent labels that serve as its temporal change features. To understand the situation of the target node in the spatial network, a graph convolution operation is performed on the target temporal change features of the target node at each time step. Graph convolution is a computational process defined on spatiotemporal graph structure data for aggregating information about nodes and their neighbors. Its core idea is to make the feature representation of a node not only depend on itself, but also be influenced by the features of its directly connected or neighboring nodes within a certain number of hops. Through this operation, the temporal change features of the neighboring nodes associated with the target node are aggregated. For example, if a node with the intent to start a game has multiple neighboring nodes with the same high-intensity interaction intent, the feature representation of that node will be enhanced, reflecting a localized group behavior trend or resource competition pressure. After performing this spatial aggregation on each time step within a continuous time window, the temporal feature sequence of the original single node is transformed into a series of new feature representations rich in spatial context, thereby generating a target feature sequence of the target node with enhanced spatial context within a continuous time window. This target feature sequence not only preserves the change of the target node's own intent over time, but also encodes the spatial impact of the evolution of the intent of other nodes in its local network environment.

[0089] Gated Recurrent Units (GRUs) are a variant of recurrent neural networks similar to Long Short-Term Memory (LSTM) networks. They also control the flow of information through gating mechanisms (update and reset gates), but with a simpler structure. They excel at capturing dependencies in time series and learning their dynamic evolution. A GRU processes the spatially enhanced feature sequence sequentially, updating its hidden state at each step based on the current input and the previous state, thereby learning the short-term fluctuations, long-term trends, and cyclical patterns inherent in the sequence. After processing the entire time window sequence, the hidden state vector output by the GRU is considered the spatiotemporal fusion feature. This feature simultaneously integrates the intentional evolution patterns of the target node itself and its spatial neighbors within the observation time window, representing a unified representation deeply interwoven with spatial correlation and temporal dynamics.

[0090] Finally, the abstract spatiotemporal fusion features need to be transformed into specific, actionable quantitative indicators of resource demand. This is accomplished through a mapping function (typically implemented by one or more fully connected layers) that maps the spatiotemporal fusion features to predicted resource demand values. These predicted resource demand values ​​can be multidimensional vectors, with each dimension corresponding to the expected demand for a specific type of cloud server resource (such as GPU computing power units, CPU core count, or memory GB) at a future time. To more intuitively and globally display the prediction results, a resource heatmap representing the spatial and temporal distribution of resource demand is generated based on the predicted resource demand values. A resource heatmap is a visual and computable data structure that graphically identifies the intensity of resource demand in different regions at different future points in time on a geographical or logical topology map of a cloud server cluster, using varying color intensity or numerical values. Essentially, a resource heatmap summarizes, statistically analyzes, and visualizes the dispersed predicted values ​​obtained from the model calculations of massive numbers of nodes, according to their corresponding server locations and time points. Therefore, the complete information contained in the resource heatmap constitutes the predicted resource distribution.

[0091] In one possible implementation of this application embodiment, when allocating differentiated cloud machine resources to multiple cloud applications, the following methods can be used, but are not limited to: allocating a basic resource guarantee zone to multiple cloud applications, wherein the basic resource guarantee zone is used to provide fixed quota resources that meet the minimum operating requirements of each of the multiple cloud applications; configuring a thermal pre-allocation resource pool for multiple cloud applications, and dynamically allocating cloud machine resources in the thermal pre-allocation resource pool to multiple cloud applications according to the predicted resource distribution; configuring a backup elastic resource pool for multiple cloud applications, wherein the backup elastic resource pool is used to supplement resources when the predicted resource distribution indicates that the cloud machine resources in the thermal pre-allocation resource pool are insufficient or the resource demand of the cloud applications increases.

[0092] In the embodiments of this application, the basic resource guarantee area is a dedicated resource portion of the cloud machine resource pool that is pre-divided and fixedly allocated to each cloud application instance. Its core function is to provide a fixed quota of resources to meet the minimum operating requirements of multiple cloud applications. The minimum operating requirements refer to the amount of computing, memory, and network resources necessary to ensure that the cloud application can start, maintain basic interface responsiveness, and execute its core functions. For example, a cloud gaming application requires at least GPU computing power to render a basic scene and CPU and memory to maintain process operation. Fixed quota resources mean that once these resources are allocated to an application instance, they are generally not reclaimed or allocated to other instances during their lifecycle, thus providing the application with deterministic resource guarantees unaffected by other applications. The establishment of this area ensures that regardless of changes in the overall system load, each cloud application can obtain the most basic resource supply, fundamentally avoiding situations where applications cannot start or become completely unavailable due to resource contention, providing a baseline guarantee for service quality.

[0093] The thermal pre-allocated resource pool occupies the majority of resources in the cloud machine resource pool. It is a globally shared, non-fixed resource collection. The allocation of this resource pool is entirely driven by intelligent prediction. Specifically, based on predicted resource distribution, the cloud machine resources in the thermal pre-allocated resource pool are dynamically allocated to multiple cloud applications. Real-time analysis of the predicted resource distribution generated by the spatiotemporal network model (e.g., resource heatmaps) identifies cloud applications and user sessions that will enter a high-resource-demand state in the near future (e.g., complex scene rendering in games, encoding processing in video editing). Subsequently, the appropriate type and quantity of resources (e.g., additional GPU cores, higher CPU frequency quotas, more memory space) are proactively allocated from the thermal pre-allocated resource pool, pre-allocated or immediately distributed to application instances predicted to have high demand. This allocation is dynamic, meaning that resources are not permanently occupied but allocated and reclaimed according to the rhythm of the prediction cycle. When the predicted high-demand period for an application ends, the excess resources it obtained from the pool are released back into the pool for use by other applications with demand.

[0094] To address prediction biases and unforeseen surges, the strategy also incorporates a safety buffer mechanism: configuring backup elastic resource pools for multiple cloud applications. These backup elastic resource pools are a portion of the resource pool that are in a standby state and do not participate in regular dynamic allocation. Their core function is to replenish resources when the predicted resource distribution indicates insufficient cloud machine resources in the thermal pre-allocated resource pool, or when the resource demands of cloud applications increase. This primarily occurs in two scenarios: first, when the rate or magnitude of actual user demand growth exceeds the expected predicted resource distribution, leading to the depletion of idle resources in the thermal pre-allocated resource pool and an inability to meet the needs of all high-demand applications; second, when a cloud application experiences a new, sudden, high-intensity task outside the predicted period due to user actions. When the monitoring system detects such resource gaps, it triggers an emergency mechanism to urgently allocate resources from the backup elastic resource pool to ensure a consistent user experience and complete task completion.

[0095] The establishment of the basic resource guarantee zone in this application separates deterministic resource guarantees from flexible resource sharing. This satisfies the most basic functional requirements of each application while avoiding the overall low resource utilization caused by excessively reserving static resources to meet minimum guarantees. The thermal pre-allocation resource pool uses dynamic scheduling based on accurate predictions, improving the efficiency of matching resources and demands in time and space. The backup elastic resource pool effectively absorbs the uncertainty of the prediction model and sudden fluctuations in business, enhancing overall resilience and adaptability to complex and ever-changing environments, and ensuring the stability of service quality levels.

[0096] In one possible implementation of this application embodiment, when dynamically allocating cloud machine resources in the thermal pre-allocation resource pool to cloud applications, it can be achieved in the following ways, but not limited to: determining the resource allocation priority of each of the multiple cloud applications according to the predicted resource demand values ​​of each of the multiple cloud applications in the predicted resource distribution, wherein the larger the predicted resource demand value, the higher the resource allocation priority; and dynamically allocating cloud machine resources in the thermal pre-allocation resource pool to cloud applications based on the resource allocation priority.

[0097] In the embodiments of this application, the resource demand forecast is a quantitative estimate by a spatiotemporal network model of the various cloud machine resources (such as GPU computing power, CPU cores, and memory capacity) required by each cloud application instance (or the user session behind it) within a preset future time period. This value directly reflects the model's judgment on the load pressure or performance requirements that the application will face. Resource allocation priority is a relative ordinal indicator used to determine the order in which multiple applications acquire resources and the degree to which they are satisfied when multiple applications are simultaneously competing for resources in the thermal pre-allocated resource pool.

[0098] A higher predicted resource demand corresponds to a higher priority in resource allocation. This means that a cloud gaming application predicted by the model to be entering a high-intensity graphics rendering phase will have a significantly higher predicted resource demand than a cloud document application predicted to be in a text browsing state. Therefore, the gaming application will be given a higher priority. The magnitude of the predicted resource demand is directly related to the sensitivity of user experience and business criticality. High predicted demand usually means that the application is processing or about to process computationally intensive, real-time-critical tasks. The timeliness and sufficiency of resource supply are crucial to preventing lag, latency, and other experience degradation. By mapping the magnitude of the predicted value to the level of priority, the automatic identification and ranking of potential service risks is achieved.

[0099] Priority-based allocation is a systematic scheduling decision-making process. At the beginning of each resource scheduling cycle, the scheduler processes resource requests from cloud applications in descending order of priority. For the highest priority application, it first attempts to fully satisfy the predicted resource demand from the currently available resources in the hot pre-allocated resource pool. If resources are sufficient, the application receives all the pre-allocated resources it needs. Then, the scheduler continues to process the next priority application, allocating resources from the remaining resources, and so on. When the available resources in the pool cannot fully satisfy the predicted demand of a certain priority application, they may be allocated proportionally or partially satisfied according to a preset strategy to ensure that higher priority applications always receive more adequate resource guarantees. This process is dynamic, meaning that the priority ranking and resource allocation results are not static but are continuously updated as the predicted resource distribution changes within each scheduling cycle. An application that receives high priority and a large amount of resources in the current cycle due to high predicted demand may have its priority reduced in the next cycle if its predicted demand decreases, and some of the previously allocated excess resources may be reclaimed from the pool for use by other currently higher priority applications.

[0100] In one specific implementation, the resource allocation of the cloud-side service layer in this application can be illustrated by, but is not limited to, the following example: The STGNN engine can uniformly process the temporal periodicity of user behavior (such as daily game peaks) and the spatial distribution of hardware resources (such as edge node topology), outputting the required resource demand heatmap. After obtaining the accurate resource heatmap, the cloud resource platform can use the resource heatmap to analyze the focus of resources, and at the same time combine the dynamic resource adjuster to optimize the allocation of cloud machine resources, maximizing the use of cloud machine cloudification applications and efficient resource utilization. The allocation strategy of the dynamic resource adjuster is roughly as follows: basic area resources are provided to each cloudification application, and each cloud machine cloudification application will be pre-allocated a basic specification of cloud machine resources to meet the basic use of cloudification applications; the heatmap pre-allocation area accounts for the largest proportion, mainly using the accurate resource heatmap output by the STCNN engine for resource coordination, and using the heatmap to predict resource allocation for different types of cloudification applications and key tags such as user habits and user intent. For example, cloud-based gaming applications receive more GPU and other resources than regular cloud applications. Different users using the same cloud resources may also have different usage patterns depending on their habits and intentions. Additionally, cloud servers have a certain percentage of extra allocation space. This space is primarily used as a safety net, ensuring that resources are allocated only when cloud applications require more resources under full load. In normal scenarios or when the hot allocation area is not fully utilized, the resources in the extra allocation area will not be used. The dynamic resource allocater structure is as follows: Figure 6 As shown, Figure 6 This is a schematic diagram of the structure of a dynamic resource manipulator provided in this application.

[0101] In one possible implementation of this application embodiment, the spatiotemporal network model deployed on the cloud needs to be trained using historical data to obtain predictive capabilities, so that the cloud can achieve accurate spatiotemporal joint modeling and resource demand prediction. Specifically, but not limited to the following methods can also be used: acquiring training resource data, which includes at least historical intent tags, historical hardware status information, and historical network status information for historical time periods; inputting the training resource data into the spatiotemporal network model for spatiotemporal joint modeling to obtain a preset predicted resource distribution for subsequent time periods; using the actual resource distribution for the preset subsequent time periods as the training objective, optimizing the parameters of the spatiotemporal network model by minimizing the error between the predicted resource distribution and the training objective, and obtaining a locally trained spatiotemporal network model.

[0102] In the embodiments of this application, a historical time period refers to a past time period that has completed a full operational cycle. Historical intent tags are sequences of user intent tags actually uploaded and recorded by each endpoint within that historical time period; they accurately reflect the behavioral patterns and changing needs of the user group at that time. Historical hardware status information and historical network status information are time-series data corresponding to that historical time period, collected from the cloud platform monitoring system, representing server resource utilization, load metrics, and network quality metrics, respectively. These three types of data are strictly aligned on the timeline.

[0103] Training resource data is input into the spatiotemporal network model for joint spatiotemporal modeling, yielding the predicted resource distribution for training in subsequent time periods. In this step, the model is placed in a simulation environment similar to actual operation. The training program extracts a continuous segment from historical data as input, containing historical intent labels, hardware, and network state information within a time window T. Based on this past context, the model performs its internal forward propagation computation, attempting to output its inference about resource demand within the immediately following future time window T—the predicted resource distribution for training.

[0104] The training objective is to use the actual resource distribution for subsequent time periods as the training target. The actual resource distribution is data extracted from historical records showing actual resource usage observed within a time window T (e.g., the actual CPU / GPU utilization matrix for each server). The model's learning objective is to make its output, the predicted resource distribution for training, as close as possible to this actual resource distribution. To achieve this, the training process optimizes the parameters of the spatiotemporal network model by minimizing the error between the predicted resource distribution and the training objective. This is specifically achieved by defining a loss function (e.g., mean squared error loss). Then, the gradient of the loss function with respect to all adjustable parameters of the model (i.e., weights and biases) is calculated using the backpropagation algorithm, and these parameters are optimized along the direction of reducing error using optimization algorithms such as gradient descent. Through iteration, the model's parameters are continuously fine-tuned, and its ability to capture spatiotemporal dependencies in historical data and extrapolate predictions based on these dependencies gradually improves.

[0105] After sufficient training iterations, when the model's prediction error on the validation dataset converges to a satisfactory low level, it is considered that the model has learned the inherent operating rules of the system, and at this point, a locally trained spatiotemporal network model is obtained.

[0106] In one possible implementation of this application embodiment, the cloud acts as a coordination hub, which needs to securely integrate local learning results from multiple terminals and feed back the fused global knowledge, thereby achieving a joint improvement in the distributed model capabilities. Specifically, it can also adopt, but is not limited to, the following methods: receiving encrypted update information uploaded by multiple terminal application layers, performing information aggregation and update processing on the multiple encrypted update information to obtain aggregated model parameters; and sending the aggregated model parameters to multiple terminal application layers respectively.

[0107] In the embodiments of this application, the goal of information aggregation and update processing is to merge numerous scattered local model updates, which may contain noise or bias, into a more robust and generalizable global model update direction. Specifically, the cloud side first decrypts the received ciphertext information in a secure environment (if using an asymmetric encryption scheme), or performs specific calculations directly in the ciphertext state (if using advanced cryptographic techniques such as homomorphic encryption). Subsequently, the cloud application uses a predetermined aggregation algorithm (e.g., federated averaging) to mathematically integrate all the decrypted update information. This algorithm typically calculates the average or weighted average of all local update parameters, where the weights can be determined based on the amount of data on each end, the magnitude of model updates, or other reliable indicators. Through this aggregation process, data bias or random noise on individual devices is weakened in the averaging operation, while common patterns prevalent across multiple devices are strengthened and highlighted. The result of this series of calculations is the aggregated model parameters. The aggregated model parameters represent the consensus-based knowledge improvement formed after a round of collective learning, integrating the behavioral patterns of a broad user group, and are theoretically more comprehensive and stable than any single local model parameter.

[0108] Corresponding to the cloud server resource scheduling method described above, this application also proposes a cloud server resource scheduling system.

[0109] The cloud server resource scheduling system includes at least one terminal application layer and a cloud service layer. At least one application layer on the client side is configured to perform the above operations. Figure 1 Cloud machine resource scheduling device in the middle step; The cloud-side service layer configuration can execute the above. Figure 4 Cloud machine resource scheduling device in the middle step; In this case, at least one application layer on the client side is connected to the service layer on the cloud side.

[0110] In the embodiments of this application, to facilitate understanding of the system, a schematic diagram of a cloud server resource scheduling system is provided, as shown below. Figure 7As shown, the system starts by classifying user intent and applications at the client-side application layer, and then differentiates cloud applications based on usage habits and different purposes. On the cloud service side, the system uses models to infer and optimize the allocation strategy, thereby solving the problems of low utilization rate and low allocation efficiency of cloud machine resources occupied by current cloud applications.

[0111] Since the embodiments of the cloud server resource scheduling system in this application correspond to the above-described method embodiments, details not disclosed in the embodiments of the cloud server resource scheduling system can be referred to the above-described method embodiments, and will not be repeated in this application.

[0112] Corresponding to the cloud server resource scheduling method described above, this application also proposes a cloud server resource scheduling device. Since the device embodiment of this application corresponds to the method embodiment described above, details not disclosed in the device embodiment can be referred to the method embodiment described above, and will not be repeated here.

[0113] Figure 8 This is a schematic diagram of a cloud server resource scheduling device provided in an embodiment of this application. The device is configured in the application layer on the endpoint, such as... Figure 8 As shown, it includes: The acquisition unit 81 is used to collect operation data generated during the operation of cloud-based applications. The operation data includes time-series behavioral data reflecting the usage habits of cloud-based applications and static data identifying the application type of cloud-based applications. Both the behavioral data and the static data contain at least one type of operation content. Labeling unit 82 is used to classify and label the operation data to obtain the intent label labels corresponding to each type of operation content in the operation data. Analysis unit 83 is used to input intent label annotation and operation data into the locally trained intent recognition model for analysis and processing, so as to obtain the target intent label representing the current usage information of the cloud application; The transmission unit 84 is used to transmit the target intent tag to the cloud-side service layer so that the cloud-side service layer can provide a basis for differentiated resource prediction and resource scheduling based on the target intent tag.

[0114] Furthermore, in one possible implementation of this application embodiment, the acquisition unit 81 is specifically used for: In response to enabling cloud-based applications, the static data of the cloud-based applications is directly extracted. The static data includes at least the application context information of the cloud-based applications. The gesture coordinates during the operation of cloud-based applications are collected to obtain gesture trajectory information. The gesture trajectory information is then encoded and blurred to obtain a gesture trajectory sequence. The temporal behavior data includes at least the gesture trajectory sequence and the gaze focus area sequence. By acquiring image acquisition devices with preset permissions, eye gaze point data and eye movement data are collected during the operation of cloud-based applications, and a sequence of gaze focus areas is generated based on the eye gaze point data and eye movement data.

[0115] Furthermore, in one possible implementation of this application embodiment, the annotation unit 82 is specifically used for: Behavioral pattern recognition is performed on various types of operational content in the operational data to obtain the corresponding behavioral patterns for each type of operational content. Based on behavioral patterns, various types of operational content are labeled to obtain the corresponding intent labels for each type of operational content.

[0116] Furthermore, in one possible implementation of this application embodiment, the analysis unit 83 is specifically used for: Preprocessing and feature extraction are performed on time-series behavioral data to obtain time-series feature vectors, and feature encoding is performed on static data to obtain type feature vectors; The temporal feature vector is input into the Long Short-Term Memory (LSTM) network, and the hidden state of the temporal feature vector is updated through the gating mechanism of the LTM network to obtain the target hidden state vector. The intent recognition model includes the LTM network. The target hidden state vector type feature vector is concatenated to obtain the target feature vector. Classification calculation is performed on the target feature vector to obtain the target intent label.

[0117] Furthermore, in one possible implementation of the embodiments of this application, such as Figure 9 As shown, the cloud server resource scheduling device also includes: Training unit 85 is used to acquire training operation data including real intent labels, and to train the intent recognition model based on the training operation data to obtain a locally trained intent recognition model. The training process for the intent recognition model includes: using classification cross-entropy as the loss function, calculating the difference between the predicted intent label and the real intent label by the intent recognition model, and updating the model parameters of the intent recognition model using the backpropagation algorithm.

[0118] Furthermore, in one possible implementation of the embodiments of this application, such as Figure 9 As shown, the cloud server resource scheduling device further includes: an update unit 86, which is used for: Obtain the parameter update information of the intent recognition model during the model training process, and encrypt the parameter update information to obtain the encrypted update information; The encrypted update information is uploaded to the cloud-side service layer; It receives aggregated model parameters from the cloud-side service layer and updates the locally trained intent recognition model based on the aggregated model parameters. The aggregated model parameters are obtained by the cloud-side service layer through aggregation and updating based on encrypted update information.

[0119] Figure 10 This is a schematic diagram of a cloud server resource scheduling device provided in an embodiment of this application. The device is configured in the cloud-side service layer, such as... Figure 10 As shown, it includes: The receiving unit 1001 is used to receive target intent tags uploaded by multiple end-side application layers; The acquisition unit 1002 is used to acquire local hardware status information and network status information; Construction unit 1003 is used to construct spatiotemporal graph data by using multiple end-side application layers as nodes and target intent tags as the temporal change features of the nodes. Modeling unit 1004 is used to input spatiotemporal graph data, hardware status information and network status information into the locally trained spatiotemporal network model to perform spatiotemporal joint modeling and obtain the predicted resource distribution for multiple cloud applications within a preset future time period. The allocation unit 1005 is used to execute a dynamic resource allocation strategy based on the predicted resource distribution to allocate differentiated cloud machine resources to multiple cloud applications.

[0120] Furthermore, in one possible implementation of this application embodiment, the construction unit 1003 is specifically used for: At each time step, a graph convolution operation is performed on the target temporal change features of the target node, and the temporal change features of the neighboring nodes associated with the target node are aggregated to generate a spatial context-enhanced target feature sequence of the target node within a continuous time window, where the target node is any node in the spatiotemporal graph data; The target feature sequence is input into the gated recurrent unit of the locally trained spatiotemporal network model to learn the dynamic evolution law in the time dimension and obtain spatiotemporal fusion features. The spatiotemporal fusion features are mapped to predicted resource demand values, and a resource heat map representing the spatial and temporal distribution of resource demand is generated based on the predicted resource demand values. The resource heat map includes the predicted resource distribution.

[0121] Furthermore, in one possible implementation of this application embodiment, the allocation unit 1005 is specifically used for: A basic resource guarantee zone is allocated to multiple cloud applications. The basic resource guarantee zone is used to provide a fixed quota of resources that meet the minimum operating requirements of each of the multiple cloud applications. Configure a thermal pre-allocation resource pool for multiple cloud applications, and dynamically allocate cloud machine resources in the thermal pre-allocation resource pool to multiple cloud applications based on the predicted resource distribution. Configure backup elastic resource pools for multiple cloud applications. These backup elastic resource pools are used to supplement resources when cloud machine resources in the predicted resource distribution indicator thermal pre-allocation resource pool are insufficient or when the resource demand of cloud applications increases.

[0122] Furthermore, in one possible implementation of this application embodiment, the allocation unit 1005 is specifically used for: Based on the predicted resource demand values ​​of multiple cloud applications in the predicted resource distribution, the resource allocation priority of each cloud application is determined. The larger the predicted resource demand value, the higher the resource allocation priority. Based on resource allocation priority, cloud machine resources in the thermal pre-allocation resource pool are dynamically allocated to cloud applications.

[0123] Furthermore, in one possible implementation of the embodiments of this application, such as Figure 11 As shown, the cloud server resource scheduling device further includes: a training unit 1006, which is used for: Acquire training resource data, which includes at least historical intent labels, historical hardware status information, and historical network status information for historical time periods. The training resource data is input into the spatiotemporal network model for spatiotemporal joint modeling to obtain the predicted distribution of training resources for the preset subsequent time periods; Using the actual resource distribution in the subsequent time period as the training objective, the parameters of the spatiotemporal network model are optimized by minimizing the error between the predicted resource distribution used for training and the training objective, thus obtaining a locally trained spatiotemporal network model.

[0124] Furthermore, in one possible implementation of the embodiments of this application, such as Figure 11 As shown, the cloud server resource scheduling device further includes: an aggregation unit 1007, which is used for: Receive encrypted update information uploaded by multiple end-side application layers, and perform information aggregation and update processing on multiple encrypted update information to obtain aggregated model parameters; The aggregated model parameters are sent to multiple edge application layers respectively.

[0125] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of the embodiments of this application, and the principle is the same. Therefore, the embodiments of this application are not limited thereto.

[0126] According to embodiments of this application, this application also provides an electronic device, a readable storage medium, and a computer program product.

[0127] Figure 12 A schematic block diagram of an example electronic device 1200 that can be used to implement embodiments of this application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.

[0128] like Figure 12 As shown, device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes based on a computer program stored in ROM (Read-Only Memory) 1202 or loaded from storage unit 1208 into RAM (Random Access Memory) 1203. RAM 1203 may also store various programs and data required for the operation of device 1200. The computing unit 1201, ROM 1202, and RAM 1203 are interconnected via bus 1204. I / O (Input / Output) interface 1205 is also connected to bus 1204.

[0129] Multiple components in device 1200 are connected to I / O interface 1205, including: input unit 1206, such as keyboard, mouse, etc.; output unit 1207, such as various types of monitors, speakers, etc.; storage unit 1208, such as disk, optical disk, etc.; and communication unit 1209, such as network card, modem, wireless transceiver, etc. Communication unit 1209 allows device 1200 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0130] The computing unit 1201 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 performs the various methods and processes described above, such as cloud server resource scheduling methods. For example, in some embodiments, the cloud server resource scheduling method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1200 via ROM 1202 and / or communication unit 1209. When the computer program is loaded into RAM 1203 and executed by the computing unit 1201, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured to perform the aforementioned cloud machine resource scheduling method by any other suitable means (e.g., by means of firmware).

[0131] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0132] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0133] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0134] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0135] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.

[0136] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0137] It's important to note that artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0138] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application can be achieved, and this is not limited herein.

[0139] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A cloud server resource scheduling method, characterized in that, Applied to the edge application layer, including: The operation data generated during the operation of cloud-based applications is collected. The operation data includes time-series behavioral data reflecting the usage habits of the cloud-based applications and static data identifying the application type of the cloud-based applications. Both the behavioral data and the static data contain at least one type of operation content. The operation data is classified and labeled to obtain the intent labels corresponding to each type of operation content in the operation data; The intent label and the operation data are input into a locally trained intent recognition model for analysis and processing to obtain a target intent label that represents the current usage information of the cloud application. The target intent tag is transmitted to the cloud-side service layer so that the cloud-side service layer can perform differentiated resource prediction and resource scheduling based on the target intent tag.

2. The cloud server resource scheduling method according to claim 1, characterized in that, The operational data generated during the cloud-based application of the data acquisition process includes: In response to opening the cloud application, the static data of the cloud application is directly extracted, and the static data includes at least the application context information of the cloud application. The gesture coordinates during the operation of the cloud application are collected to obtain gesture trajectory information, and the gesture trajectory information is encoded and blurred to obtain a gesture trajectory sequence. The temporal behavior data includes at least the gesture trajectory sequence and the gaze focus area sequence. The eye gaze point data and eye movement data during the operation of the cloud application are collected by an image acquisition device with preset permissions, and the gaze focus area sequence is generated based on the eye gaze point data and eye movement data.

3. The cloud server resource scheduling method according to claim 1, characterized in that, The process of classifying and labeling the operation data to obtain the intent label annotations corresponding to each type of operation content in the operation data includes: Behavioral pattern recognition is performed on each type of operation content in the operation data to obtain the corresponding behavioral pattern for each type of operation content. Based on the behavioral patterns, the various types of operation content are labeled to obtain the intent tags corresponding to each type of operation content.

4. The cloud server resource scheduling method according to claim 1, characterized in that, The step of inputting the intent label and the operation data into a locally trained intent recognition model for analysis and processing to obtain the target intent label representing the current usage information of the cloud application includes: The time-series behavioral data is preprocessed and features are extracted to obtain a time-series feature vector, and the static data is feature-encoded to obtain a type feature vector; The temporal feature vector is input into a long short-term memory network, and the hidden state of the temporal feature vector is updated through the gating mechanism of the long short-term memory network to obtain the target hidden state vector. The intent recognition model includes a long short-term memory network. The target hidden state vector is concatenated with the type feature vector to obtain the target feature vector. Classification calculation is performed on the target feature vector to obtain the target intent label.

5. The cloud server resource scheduling method according to claim 1, characterized in that, The method further includes: Obtain training operation data including real intent labels, and train the intent recognition model based on the training operation data to obtain the locally trained intent recognition model. The training process for the intent recognition model includes: using classification cross-entropy as a loss function, calculating the difference between the predicted intent label and the real intent label predicted by the intent recognition model, and updating the model parameters of the intent recognition model using the backpropagation algorithm.

6. The cloud server resource scheduling method according to claim 5, characterized in that, After training the intent recognition model based on the training operation data to obtain the locally trained intent recognition model, the method further includes: Obtain the parameter update information of the intent recognition model during the model training process, and encrypt the parameter update information to obtain encrypted update information; The encrypted update information is uploaded to the cloud-side service layer; The system receives aggregated model parameters from the cloud-side service layer and updates the locally trained intent recognition model based on the aggregated model parameters. The aggregated model parameters are obtained by the cloud-side service layer through aggregation and updating based on the encrypted update information.

7. A cloud server resource scheduling method, characterized in that, Applied to the cloud-side service layer, including: It receives target intent tags uploaded by multiple end-side application layers, and obtains local hardware status information and network status information. Using the multiple edge application layers as nodes and the target intent tag as the temporal change feature of the node, a spatiotemporal graph data is constructed. The spatiotemporal graph data, the hardware status information, and the network status information are input into a locally trained spatiotemporal network model to perform spatiotemporal joint modeling, thereby obtaining the predicted resource distribution for multiple cloud applications within a preset future time period. Based on the predicted resource distribution, a dynamic resource allocation strategy is executed to allocate differentiated cloud machine resources to the multiple cloud applications.

8. The cloud server resource scheduling method according to claim 7, characterized in that, The step of inputting the spatiotemporal graph data, the hardware status information, and the network status information into a locally trained spatiotemporal network model for spatiotemporal joint modeling to obtain the predicted resource distribution for multiple cloud applications within a preset future time period includes: At each time step, a graph convolution operation is performed on the target temporal change features of the target node, and the temporal change features of the neighboring nodes associated with the target node are aggregated to generate a spatial context-enhanced target feature sequence of the target node within a continuous time window, wherein the target node is any node in the spatiotemporal graph data; The target feature sequence is input into the gated recurrent unit of the locally trained spatiotemporal network model to learn the dynamic evolution law in the time dimension and obtain spatiotemporal fusion features; The spatiotemporal fusion features are mapped to predicted resource demand values, and a resource heat map representing the spatial and temporal distribution of resource demand is generated based on the predicted resource demand values, wherein the resource heat map includes the predicted resource distribution.

9. The cloud server resource scheduling method according to claim 8, characterized in that, The step of executing a dynamic resource allocation strategy based on the predicted resource distribution to allocate differentiated cloud machine resources to the multiple cloud applications includes: A basic resource guarantee zone is allocated to the multiple cloud applications, and the basic resource guarantee zone is used to provide a fixed quota of resources that meet the minimum operating requirements of each of the multiple cloud applications. Configure a thermal pre-allocation resource pool for the multiple cloud applications, and dynamically allocate cloud machine resources in the thermal pre-allocation resource pool to the multiple cloud applications according to the predicted resource distribution; Configure a backup elastic resource pool for the plurality of cloud applications. The backup elastic resource pool is used to supplement resources when the predicted resource distribution indicates that the cloud machine resources in the thermal pre-allocation resource pool are insufficient or the resource demand of the cloud applications increases.

10. The cloud server resource scheduling method according to claim 9, characterized in that, The step of dynamically allocating cloud server resources in the thermal pre-allocation resource pool to the cloud application based on the predicted resource distribution includes: Based on the predicted resource demand values ​​of each of the multiple cloud applications in the predicted resource distribution, the resource allocation priority of each of the multiple cloud applications is determined, wherein the larger the predicted resource demand value, the higher the resource allocation priority. Based on the resource allocation priority, the cloud machine resources in the thermal pre-allocation resource pool are dynamically allocated to the cloud application.

11. The cloud server resource scheduling method according to claim 7, characterized in that, The method further includes: Acquire training resource data, which includes at least historical intent tags, historical hardware status information, and historical network status information for historical time periods; The training resource data is input into the spatiotemporal network model for spatiotemporal joint modeling to obtain the predicted distribution of training resources for a preset subsequent time period. Using the actual resource distribution of the preset subsequent time period as the training objective, the parameters of the spatiotemporal network model are optimized by minimizing the error between the predicted resource distribution used for training and the training objective, thereby obtaining the locally trained spatiotemporal network model.

12. The cloud server resource scheduling method according to claim 7, characterized in that, The method further includes: Receive encrypted update information uploaded by each of the multiple end-side application layers, and perform information aggregation and update processing on the multiple encrypted update information to obtain aggregated model parameters; The aggregated model parameters are sent to the multiple end-side application layers respectively.

13. A cloud server resource scheduling device, characterized in that, include: The data acquisition unit is used to collect operation data generated during the operation of cloud-based applications. The operation data includes time-series behavioral data reflecting the usage habits of the cloud-based applications and static data identifying the application type of the cloud-based applications. Both the behavioral data and the static data contain at least one type of operation content. The annotation unit is used to classify and annotate the operation data to obtain the intent label annotations corresponding to each type of operation content in the operation data. The analysis unit is used to input the intent label and the operation data into a locally trained intent recognition model for analysis and processing, so as to obtain the target intent label representing the current usage information of the cloud application; The transmission unit is used to transmit the target intent tag to the cloud-side service layer so that the cloud-side service layer can provide a basis for differentiated resource prediction and resource scheduling based on the target intent tag.

14. A cloud server resource scheduling device, characterized in that, include: The receiving unit is used to receive target intent tags uploaded by multiple end-side application layers. The acquisition unit is used to acquire local hardware status information and network status information; The construction unit is used to construct spatiotemporal graph data using the multiple end-side application layers as nodes and the target intent tag as the temporal change feature of the node; The modeling unit is used to input the spatiotemporal graph data, the hardware status information and the network status information into the locally trained spatiotemporal network model to perform spatiotemporal joint modeling and obtain the predicted resource distribution for multiple cloud applications within a preset future time period. The allocation unit is used to execute a dynamic resource allocation strategy based on the predicted resource distribution to allocate differentiated cloud machine resources to the multiple cloud applications.

15. A cloud server resource scheduling system, characterized in that, include: At least one terminal application layer and a cloud service layer, Each of the at least one end-side application layer is configured with a cloud machine resource scheduling device as described in claim 13; The cloud-side service layer is configured with the cloud machine resource scheduling device as described in claim 14; In this configuration, at least one terminal application layer is communicatively connected to the cloud-side service layer.

16. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-6 or the method of any one of claims 7-12.

17. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6 or any one of claims 7-12.

18. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-6 or any one of claims 7-12.