Cloud data processing system based on cloud server
Through the cloud data processing system based on cloud servers, the problem of efficient, intelligent and secure processing in a multi-source heterogeneous data environment is solved, the unified analysis of heterogeneous data and the integration of security features are realized, the intelligent scheduling capability of the system is improved, the unified analysis of heterogeneous data and the integration of security features are realized, an efficient data processing system is realized, the improvement of the map construction capability in the existing technology is solved, and the efficient processing and secure aggregation of heterogeneous data are realized.
Patent Information
- Application Number
- CN202510770436.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Existing technologies make it difficult to achieve efficient, intelligent, and secure data fusion and processing in a multi-source heterogeneous data environment, especially in task scenarios with strong real-time requirements, complex structures, and dynamic changes. Traditional systems lack encoding strategies for source information perception, have limited graph structure construction capabilities, and are subject to the risk of information leakage and privacy reconstruction.
A cloud data processing system based on a cloud server is provided, which includes a data access and encoding module, a graph structure generation module, a scheduling optimization module and a security feature aggregation module. Through source-aware embedding processing, semantic association relationship construction, dynamic scheduling and privacy protection feature fusion, efficient processing and secure aggregation of heterogeneous data are achieved.
It improves the fusion processing capability of heterogeneous data, optimizes task scheduling efficiency and matching accuracy, realizes real-time subgraph recognition and priority judgment, and enhances the system's streaming processing response capability and data security.
Smart Images

Figure CN120687434A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a cloud data processing system based on a cloud server. Background Art
[0002] With the rapid development of cloud computing and big data technologies, the demand for data processing in distributed environments is growing. Especially in the context of the continuous convergence of heterogeneous data from multiple sources, achieving efficient, intelligent, and secure data fusion and processing has become a key technical challenge. Traditional data processing systems rely on static scheduling models and rule-driven feature extraction methods, making them difficult to adapt to real-time, complex, and dynamically changing task scenarios. Furthermore, with the increasing demand for privacy protection of multi-source data, centralized processing models pose risks of information leakage and privacy reconstruction, and lack effective mechanisms for secure feature fusion.
[0003] In existing technologies, although some systems support heterogeneous data access and preliminary feature expression, they lack encoding strategies for source information perception, making it difficult to establish stable cross-source semantic mapping relationships. Furthermore, traditional graph construction methods fail to effectively combine high-dimensional embedding representations with multi-granularity semantic associations, resulting in limited graph structure expressiveness and an inability to support resource scheduling optimization for complex tasks. At the scheduling mechanism level, static weights or heuristic rule matching methods are commonly used, which cannot fully express the high-dimensional feature adaptation relationship between tasks and resources and lack the ability to dynamically evolve.
[0004] Therefore, there is an urgent need for a comprehensive data processing system for cloud server environments that can support heterogeneous access, graph structure construction, task scheduling optimization, real-time subgraph recognition and security feature fusion, so as to improve the system's intelligent scheduling capabilities, real-time analysis capabilities and data security level. Summary of the Invention
[0005] Based on the above-mentioned shortcomings of the prior art, the purpose of the present invention is to provide a cloud data processing system based on a cloud server to solve the above-mentioned technical problems.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a cloud data processing system based on a cloud server, comprising:
[0007] The data access and encoding module receives raw data from multiple heterogeneous data sources and performs structured parsing and source-aware embedding on the data;
[0008] The graph structure generation module constructs node representations based on the encoded data, generates semantic associations between nodes, and forms a graph structure representation that can be deployed at different storage levels;
[0009] The scheduling optimization module establishes a scheduling mapping relationship based on the attributes of the processing tasks and the status of cloud computing resources, realizing the dynamic allocation of task resources and the construction of a scheduling graph;
[0010] The streaming subgraph analysis module is used to construct a local graph structure for real-time data access, identify subgraph regions with structural aggregation or behavioral compactness, and mark them as high-priority tasks;
[0011] The secure feature aggregation module is used to perform privacy-protected feature fusion on the local feature representations of multiple data nodes and complete secure aggregation processing of distributed data through the cloud.
[0012] The present invention is further configured such that the data access and encoding module includes:
[0013] Perform protocol identification and structural analysis on raw data from multiple heterogeneous data sources to extract data fields that conform to the preset semantic structure;
[0014] Identify the data source identifier and structural feature labels of the original data, and establish a representation mapping corresponding to the source attributes;
[0015] Generate data source-aware embedded representations based on data fields and data source identifiers;
[0016] The validity of the embedding representation is evaluated and data instances that do not meet the set rules are eliminated.
[0017] The present invention is further configured such that the source-aware embedding representation forms a distinguishable source semantic subspace in a high-dimensional embedding space by jointly modeling the context structure between data fields and the relationship vectors between data source labels, so as to improve the normalized parsing capability of heterogeneous data.
[0018] The present invention is further configured such that the graph structure generating module includes:
[0019] Generate graph node representation based on the encoded data embedding representation and establish a node set with source-aware attributes;
[0020] Perform multi-dimensional matching analysis on the semantic features between nodes, construct composite semantic relationships and generate edge sets annotated with semantic weights;
[0021] Construct a semantic graph structure based on node sets and edge sets, and support layer representations of different granularities based on conditions such as semantic density and association strength;
[0022] Map multi-layer graph structures to different cloud storage and computing environments to form a graph data organization that can be flexibly scheduled and accessed on demand.
[0023] The present invention is further configured such that the scheduling optimization module includes:
[0024] Generate a unified structure of task representation vector based on the attribute information of the processing task to characterize the resource demand characteristics and execution constraints of the task;
[0025] Collect and encode the current operating status of each cloud computing resource node and construct a resource representation vector for scheduling matching;
[0026] A heterogeneous matching mechanism is constructed based on the high-dimensional feature relationship between task representation vectors and resource representation vectors to generate a set of adaptation scores for tasks and resources.
[0027] Construct a scheduling graph structure based on the adaptation score results and generate a globally optimal task-resource mapping relationship;
[0028] The scheduling graph structure is mapped to the task control engine to trigger task scheduling execution, and a resource status feedback channel is established to support dynamic scheduling iterative updates.
[0029] The present invention is further configured such that the scheduling graph structure has the ability to trace node states and dynamically update edge weights, and supports triggering local incremental reconstruction of the graph structure when resource states change, so as to achieve adaptive adjustment of task scheduling paths.
[0030] The present invention is further configured such that the streaming subgraph analysis module includes:
[0031] Based on the structured data output by the encoding module, a local graph structure with time-aware properties is generated within a preset time window;
[0032] Extract structural embedding features from local graph structures and identify graph regions with structural aggregation based on high-dimensional correlations between nodes;
[0033] Analyze the triggering behavior of nodes in the time series dimension, construct the behavior compactness feature vector, and evaluate the dynamic activity of the local node set;
[0034] By integrating the structural aggregation and behavioral compactness features, a priority discrimination mechanism for subgraph sorting is constructed, and high-priority subgraph regions are selected based on the discrimination results;
[0035] The subgraph areas identified as high priority are marked as urgent processing tasks, and the relevant structural information is passed to the scheduling optimization module to trigger rapid resource scheduling.
[0036] The present invention is further configured such that the security feature aggregation module includes:
[0037] The local feature vector of each node is embedded, perturbed, and mapped to a unified feature representation space.
[0038] Construct nonlinear cross features based on node interactions and form a global aggregate representation and apply high-order perturbations to prevent inverse reconstruction for encryption;
[0039] The encrypted global aggregation representation is transmitted to the graph structure generation module and the scheduling optimization module.
[0040] The present invention is further configured such that the construction of the nonlinear cross-features is based on tensor interaction mapping across node feature dimensions, and is combined with random rotation perturbation and weight mask injection in a high-order perturbation mechanism to enhance the privacy security and reconstruction irreversibility of the global aggregated features.
[0041] The present invention provides a cloud data processing system based on a cloud server. The method receives raw data from multiple heterogeneous data sources through a data access and encoding module, and performs structured parsing and source-aware embedding processing on the data. A graph structure generation module constructs node representations based on the encoded data, generates semantic associations between nodes, and forms a graph structure representation that can be deployed at different storage levels. A scheduling optimization module establishes a scheduling mapping relationship based on the attributes of processing tasks and the status of cloud computing resources to achieve dynamic allocation and scheduling graph construction between task resources. A streaming subgraph analysis module is used to construct a local graph structure for real-time accessed data, identify subgraph areas with structural aggregation or behavioral compactness, and mark them as high-priority tasks. A security feature aggregation module is used to perform feature fusion under privacy protection on the local feature representations of multiple data nodes, and complete secure aggregation processing of distributed data through the cloud. The beneficial effects produced include:
[0042] 1. Improved heterogeneous data access capabilities: By introducing a source-aware embedded coding mechanism, unified parsing and semantic alignment of data sources from various protocols, structures, and semantic backgrounds can be achieved. This significantly enhances the system's ability to integrate and process heterogeneous data, providing a highly consistent feature foundation for subsequent graph modeling and task scheduling.
[0043] 2. Optimizing task scheduling efficiency and matching accuracy: A heterogeneous matching mechanism is constructed based on high-dimensional representations of tasks and resources. The adaptive scoring graph is combined to generate task-resource mapping relationships. This approach offers global optimality and dynamic evolution capabilities, significantly improving the rationality of resource allocation and computational efficiency, and adapting to scheduling requirements in multi-task concurrent scenarios.
[0044] 3. Real-time subgraph identification and priority determination: By jointly analyzing the structural aggregation and behavioral compactness characteristics in the local graph, a multi-dimensional subgraph priority determination mechanism is constructed to achieve rapid identification and marking of high-value task areas. The scheduling linkage mechanism supports the immediate dispatch of resources and enhances the system's streaming processing response capabilities.
[0045] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts. In the drawings:
[0047] Figure 1 The flowchart of a cloud data processing system based on a cloud server is shown as an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0048] The following describes the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art will readily appreciate the other advantages and benefits of the present invention from the disclosure herein. The present invention may also be implemented or applied through various other specific embodiments, and the various details in this specification may be modified or altered based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are intended only to illustrate the present invention and are not intended to limit the scope of protection of the present invention.
[0049] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. Therefore, the illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.
[0050] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the embodiments of the present invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.
[0051] Example 1
[0052] A cloud data processing system based on a cloud server, such as Figure 1 As shown, including:
[0053] The data access and encoding module receives raw data from multiple heterogeneous data sources and performs structured parsing and source-aware embedding on the data;
[0054] The graph structure generation module constructs node representations based on the encoded data, generates semantic associations between nodes, and forms a graph structure representation that can be deployed at different storage levels;
[0055] The scheduling optimization module establishes a scheduling mapping relationship based on the attributes of the processing tasks and the status of cloud computing resources, realizing the dynamic allocation of task resources and the construction of a scheduling graph;
[0056] The streaming subgraph analysis module is used to construct a local graph structure for real-time data access, identify subgraph regions with structural aggregation or behavioral compactness, and mark them as high-priority tasks;
[0057] The secure feature aggregation module is used to perform privacy-protected feature fusion on the local feature representations of multiple data nodes and complete secure aggregation processing of distributed data through the cloud.
[0058] The present invention is further configured such that the data access and encoding module includes:
[0059] Perform protocol recognition and structured parsing on raw data from multiple heterogeneous data sources to extract data fields that conform to the preset semantic structure. Specifically, protocol recognition automatically identifies the format standards for data transmission or storage, such as HTTP, MQTT, CSV, etc., to correctly parse the raw data structure. Structured parsing converts unstructured or semi-structured raw data into a data structure with clearly defined fields and types. The input data field set is Where M is the number of fields and the field feature dimension is d f , each field is represented as a vector The data source identifier set is Where N is the number of data sources, and each data source corresponds to a label vector The inter-field context structure is represented by the adjacency tensor A∈{0,1} M×M×K Indicates that K is the number of context relationship types, A i,j,k =1 indicates field f i With f j In the kth relationship, there is a dependency, and the relationship between data source labels is expressed as a tensor. Indicates that L is the dimension of label relationship features; map field features to a unified embedding space and define the embedding matrix Among them, e i is the field embedding, W kis the linear transformation weight matrix of the k-th contextual relationship, σ(·) is a nonlinear activation function, which combines the features of adjacent fields through weighted contextual relationships. After linear transformation, the context-aware embedding of the fields is extracted to capture the diverse dependencies between fields.
[0060] Identify the data source identifier and structural feature label of the original data, and establish a representation mapping corresponding to the source attributes using relational tensor mapping: in, is the data source label embedding, α j,m,l is the attention weight, satisfying ∑ m,l α j,m,l =1, used to adjust the interaction between labels, Q l is the projection matrix of the lth attention channel;
[0061] Generate data source-aware embedding representation based on data fields and data source identifiers i ;
[0062] Evaluate the effectiveness of the embedding representation and eliminate data instances that do not meet the set rules, and design the mapping function Calculate the embedding validity score: Among them, w is the weight vector, b is the bias term, and v i is the embedding validity score, threshold screening: if v i ≥θ, θ∈(0,1), keep z i .
[0063] The present invention is further configured such that the source-aware embedding representation forms a distinguishable source semantic subspace in a high-dimensional embedding space by jointly modeling the context structure between data fields and the relationship vectors between data source labels, thereby improving the normalized parsing capability of heterogeneous data. Specifically, the field and all data source label embeddings are bilinearly fused to construct a data source-aware field embedding, ensuring that the embedding forms a semantic subspace structure in the high-dimensional space, and the field embedding and the data source label embedding are fused through tensor product: To fuse the embeddings, the final source-aware embedding is defined as N is the number of data sources.
[0064] The present invention is further configured such that the graph structure generating module includes:
[0065] Based on the encoded data embedding representation, a graph node representation is generated and a node set with source-aware attributes is established. Specifically, the graph node representation is transformed into the encoding module output embedding by the mapping function φ(·): i =φ(z i )=LayerNorm(Pzi +b v ), where v i is the graph node representation, LayerNorm is layer normalization to improve numerical stability, P is the weight matrix, b v is the bias term;
[0066] Perform multi-dimensional matching analysis on the semantic features between nodes, construct complex semantic relationships and generate edge sets annotated with semantic weights. Define the multi-dimensional semantic similarity matrix between nodes as tensor S. Based on the cross-mapping of node feature tensors, combine bilinear mapping with element cross-action, capture the multi-dimensional complex semantic associations between nodes, generate multi-dimensional semantic weights S′, and determine the existence of edges based on the threshold matrix T: k is the number of semantic relations;
[0067] Construct a semantic graph structure based on node sets and edge sets, and support the division of layer representations of different granularities based on conditions such as semantic density and association strength, and introduce semantic density function Based on the sum of adjacent edge weights: p is the cubic norm, is the graph node set, M is the number of nodes, w i,j is a multidimensional semantic weight vector, emphasizing the weight influence of stronger associated edges, and defining the layer partitioning function When τ is satisfied l -1≤ρ(v i )<τ l , Among them, {τ0,τ1,…,τ L} is the threshold sequence for semantic density division;
[0068] Map multi-layer graph structures to different cloud storage and computing environments to form a graph data organization that can be flexibly scheduled and accessed on demand.
[0069] The present invention is further configured such that the scheduling optimization module includes:
[0070] Based on the attribute information of the processing task, a task representation vector with a unified structure is generated to characterize the resource requirement characteristics and execution constraints of the task. Specifically, let each task T i The attribute set is Each attribute is embedded by the function φ k Converted into a vector, and finally formed the task representation vector: Among them, ⊕ represents the vector splicing operation, is the embedding function of attribute k;
[0071] Collect and encode the current operating status of each cloud computing resource node, construct the resource representation vector for scheduling matching, and each cloud resource node R jThe state feature set is Using the embedding function ψ m Convert to resource representation vector:
[0072] According to the high-dimensional feature relationship between the task representation vector and the resource representation vector, a heterogeneous matching mechanism is constructed to generate a set of adaptation scores for tasks and resources, and calculate the task T i With Resource R j The adaptation score ρ ij , introducing three-dimensional tensor kernel mapping definition: Among them, t i [p] is the pth dimension of the task vector, r j [q] is the qth dimension of the resource vector, is the coefficient of the tensor kernel in the three-dimensional coordinates (p, q, z), γ z is the task scheduling preference coefficient, which supports fine-tuning of system strategies, ρ ij Score the final adaptation for scheduling decisions;
[0073] A scheduling graph structure is constructed based on the adaptation score results, and a globally optimal task-resource mapping relationship is generated. The globally optimal design of the scheduling graph ensures optimal resource utilization in multi-task and high-concurrency scenarios.
[0074] The scheduling graph structure is mapped to the task control engine to trigger task scheduling execution, and a resource status feedback channel is established to support dynamic scheduling iterative updates.
[0075] The present invention is further configured such that the scheduling graph structure has the ability to trace node states and dynamically update edge weights, and supports triggering local incremental reconstruction of the graph structure when resource states change, so as to achieve adaptive adjustment of task scheduling paths. Specifically, all task-resource scores are constructed as a weighted bipartite graph G = (U, V, E, ρ), where: U = {t i} is the set of task nodes; V = {r j} is the resource node set, E={(i,j)|ρ ij >τ} is the set of candidate matching edges, ρ ij The final adaptation score is used as the edge weight, τ is the preset threshold, and the scheduling function is defined to achieve the optimal scheduling path. sti,j are not allocated repeatedly, st is used to introduce constraints, and the global optimality design of the scheduling graph ensures the optimization of resource utilization in multi-task high concurrency scenarios.
[0076] The present invention is further configured such that the streaming subgraph analysis module includes:
[0077] Based on the structured data output by the encoding module, a local graph structure with time-aware attributes is generated within a preset time window. Within the preset time window Δt, the structured data stream output by the encoding module is collected, and a time-aware local graph G is generated based on the timestamp events between entities. t =(V t ,E t ), where each node v i ∈V t With edge e ij ∈E t All are accompanied by a time tag τ i , τ ij , so that the graph has the ability to evolve in time series;
[0078] Extract structural embedding features from the local graph structure and identify graph regions with structural aggregation based on the high-dimensional correlation between nodes. Specifically, the structural aggregation calculation defines the structural aggregation feature μ for node i in the local graph. i for: in: is the neighbor set of node i, is the structural similarity weight between nodes i and j, z i , is the embedding vector of the node, Φ(x)=x 2 +λ1x is the structural nonlinear amplification function, is anisotropic tensor interaction mapping, is the structural nonlinear amplification factor, ⊙ is the element-wise multiplication;
[0079] Analyze the triggering behavior of nodes in the time series dimension, construct the behavior compactness feature vector, evaluate the dynamic activity of the local node set, and calculate the behavior compactness feature ψ for node i. i Defined as: in, is the trigger frequency of node i at time t; is the behavior state vector at that moment, is a nonlinear cross function, is the behavioral nonlinear amplification factor; t0 is the window start time, ΔT is the window length;
[0080] Based on the characteristics of structural aggregation and behavioral compactness, a priority discrimination mechanism for subgraph sorting is constructed, and high-priority subgraph areas are screened based on the discrimination results. The subgraph priority scores are calculated for local subgraphs. The priority ratings are: in: is the structure-behavior fusion feature of node i, is the subgraph importance scoring function, To stack the fusion vectors into a matrix, det is the determinant function, which represents the degree of feature space span, and rank is the matrix rank function, which measures independence. is the rank enhancement weight factor;
[0081] The subgraph area identified as high priority is marked as an urgent processing task, and the relevant structural information is passed to the scheduling optimization module to trigger rapid resource scheduling. If the priority score is greater than the preset screening threshold, the subgraph is marked as a high-priority task, and the downstream scheduling module is triggered to quickly dispatch resources.
[0082] The present invention is further configured such that the security feature aggregation module includes:
[0083] The local feature vector of each node is embedded, perturbed, mapped, encoded, encrypted, and mapped to a unified feature representation space. Specifically, the original embedded feature of each node i is Introducing perturbation terms and encoding mappings, Among them, x i is the original embedding vector of the node, is the local perturbation vector, which is dynamically generated by the encryption control module, ⊕ is the site XOR perturbation operation, Π i (·)=M i ·(·)+b i Is a private mapping function, containing a weight matrix and the bias term b i , is the security feature representation after perturbation mapping;
[0084] Construct nonlinear cross features based on node interactions and form a global aggregate representation and apply high-order perturbations to prevent inverse reconstruction for encryption;
[0085] The encrypted global aggregation representation is transmitted to the graph structure generation module and the scheduling optimization module.
[0086] The present invention is further configured such that the construction of the nonlinear cross-feature is based on tensor interaction mapping across node feature dimensions, and is combined with random rotation perturbation and weight mask injection in the high-order perturbation mechanism to enhance the privacy security and reconstruction irreversibility of the global aggregated feature. Specifically, for any node pair (i, j), its high-order cross-feature is constructed as follows: in, For the outer product operation, generate ⊙ is the element-wise multiplication operation, is the feature block after tensor interaction mapping, To define the nonlinear interaction function across nodes.
[0087] It should be noted that the specific manner in which each module and unit performs operations in the cloud data processing system based on a cloud server provided in the above embodiment has been described in detail in the method embodiments and will not be repeated here. In actual applications, the cloud data processing system based on a cloud server provided in the above embodiment can, as needed, allocate the above functions to different functional modules, that is, divide the internal structure of the system into different functional modules to complete all or part of the functions described above, and this is not limited here.
[0088] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0089] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.
[0090] In this application, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.
[0091] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0092] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0093] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0094] In the several embodiments provided in this application, it should be understood that the disclosed system can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0095] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0096] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0097] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0098] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A cloud data processing system based on a cloud server, characterized in that: include: The data access and encoding module receives raw data from multiple heterogeneous data sources and performs structured parsing and source-aware embedding on the data; The graph structure generation module constructs node representations based on the encoded data, generates semantic associations between nodes, and forms a graph structure representation that can be deployed at different storage levels; The scheduling optimization module establishes a scheduling mapping relationship based on the attributes of the processing tasks and the status of cloud computing resources, realizing the dynamic allocation of task resources and the construction of a scheduling graph; The streaming subgraph analysis module is used to construct a local graph structure for real-time data access, identify subgraph regions with structural aggregation or behavioral compactness, and mark them as high-priority tasks; The secure feature aggregation module is used to perform privacy-protected feature fusion on the local feature representations of multiple data nodes and complete secure aggregation processing of distributed data through the cloud.
2. A cloud data processing system based on a cloud server according to claim 1, characterized in that: The data access and encoding module includes: Perform protocol identification and structural analysis on raw data from multiple heterogeneous data sources to extract data fields that conform to the preset semantic structure; Identify the data source identifier and structural feature labels of the original data, and establish a representation mapping corresponding to the source attributes; Generate data source-aware embedded representations based on data fields and data source identifiers; The validity of the embedding representation is evaluated and data instances that do not meet the set rules are eliminated.
3. A cloud data processing system based on a cloud server according to claim 2, characterized in that: Source-aware embedding representation forms a distinguishable source semantic subspace in a high-dimensional embedding space by jointly modeling the contextual structure between data fields and the relationship vectors between data source labels, thereby improving the normalized parsing capability of heterogeneous data.
4. A cloud data processing system based on a cloud server according to claim 2, characterized in that: The graph structure generation module includes: Generate graph node representation based on the encoded data embedding representation and establish a node set with source-aware attributes; Perform multi-dimensional matching analysis on the semantic features between nodes, construct composite semantic relationships and generate edge sets annotated with semantic weights; Construct a semantic graph structure based on node sets and edge sets, and support layer representations of different granularities based on conditions such as semantic density and association strength; Map multi-layer graph structures to different cloud storage and computing environments to form a graph data organization that can be flexibly scheduled and accessed on demand.
5. A cloud data processing system based on a cloud server according to claim 1, characterized in that: The scheduling optimization module includes: Generate a unified structure of task representation vector based on the attribute information of the processing task to characterize the resource demand characteristics and execution constraints of the task; Collect and encode the current operating status of each cloud computing resource node and construct a resource representation vector for scheduling matching; A heterogeneous matching mechanism is constructed based on the high-dimensional feature relationship between task representation vectors and resource representation vectors to generate a set of adaptation scores for tasks and resources. Construct a scheduling graph structure based on the adaptation score results and generate a globally optimal task-resource mapping relationship; The scheduling graph structure is mapped to the task control engine to trigger task scheduling execution, and a resource status feedback channel is established to support dynamic scheduling iterative updates.
6. A cloud data processing system based on a cloud server according to claim 5, characterized in that: The scheduling graph structure has the ability to trace node status and dynamically update edge weights, and supports triggering local incremental reconstruction of the graph structure when resource status changes to achieve adaptive adjustment of task scheduling paths.
7. A cloud data processing system based on a cloud server according to claim 2, characterized in that: The streaming subgraph analysis module includes: Based on the structured data output by the encoding module, a local graph structure with time-aware properties is generated within a preset time window; Extract structural embedding features from local graph structures and identify graph regions with structural aggregation based on high-dimensional correlation relationships between nodes; Analyze the triggering behavior of nodes in the time series dimension, construct the behavior compactness feature vector, and evaluate the dynamic activity of the local node set; By integrating the structural aggregation and behavioral compactness features, a priority discrimination mechanism for subgraph sorting is constructed, and high-priority subgraph regions are selected based on the discrimination results. The subgraph areas identified as high priority are marked as urgent processing tasks, and the relevant structural information is passed to the scheduling optimization module to trigger rapid resource scheduling.
8. A cloud data processing system based on a cloud server according to claim 1, characterized in that: The security feature aggregation module includes: The local feature vector of each node is embedded, perturbed, and mapped to a unified feature representation space. Construct nonlinear cross features based on node interactions and form a global aggregate representation and apply high-order perturbations to prevent inverse reconstruction for encryption; The encrypted global aggregation representation is transmitted to the graph structure generation module and the scheduling optimization module.
9. A cloud data processing system based on a cloud server according to claim 8, characterized in that: The construction of nonlinear cross-features is based on tensor interaction mapping across node feature dimensions, and is combined with random rotation perturbations and weight mask injection in high-order perturbation mechanisms to enhance the privacy security and reconstruction irreversibility of global aggregated features.
Citation Information
Patent Citations
Task scheduling method and device based on distributed cloud platform and related equipment
CN117707797A
AR (Augmented Reality) application calculation unloading strategy for mobile edge calculation
CN117896781A
Cloud computing task storage resource intelligent scheduling method based on adaptive graph neural network
CN119440832A
Cloud data processing system based on artificial intelligence algorithm
CN119473645A
Business data analysis method and device based on big data, equipment and storage medium
CN119669309A