A Method for Constructing an Edge Computing Knowledge Graph Based on Representation Learning
By constructing an edge computing knowledge graph based on representation learning, the problems of inefficiency and low quality in traditional methods are solved, and efficient task-resource matching is achieved, especially intelligent matching support is provided in dynamic environments.
Patent Information
- Application Number
- CN202411772791.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-12-04
AI Technical Summary
In an edge computing environment, traditional knowledge graph construction methods are difficult to effectively process large amounts of distributed data, resulting in inefficient construction and low quality, unable to extract valuable information, and unable to effectively match tasks and resources.
Using a representation learning method, entities, entity attributes and relational data are extracted abstractly, preprocessed and data annotated, and using representation learning technology to map entities and relationships into low-dimensional vectors, structured knowledge graphs are constructed, and unknown or missing entity relationships are automatically inferred and completed through efficient link prediction technology.
It improves the efficiency and quality of the knowledge graph construction, can extract deep information from complex data, optimize the matching of tasks and resources, and especially provide efficient task allocation and resource matching support in dynamic environments.
Smart Images

Figure CN119690669B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of edge computing, and particularly relates to a method for constructing an edge computing knowledge graph based on representation learning. Background Art
[0002] In edge computing scenarios, there are a large number of edge devices, which are usually distributed in various locations and have their own computing capabilities and storage resources. At the same time, edge devices need to process various types of tasks, such as data collection, data processing, data analysis, etc. These tasks have their own requirements, such as computing requirements, storage requirements, network requirements, etc. One of the main challenges of edge computing is how to effectively match tasks and resources. Due to the large number of edge devices, diverse types of tasks, and the dynamic and uncertain nature of the edge computing environment, the problem of task and resource matching is very complex. To solve this problem, many methods have been proposed, such as optimization-based methods, rule-based methods, learning-based methods, etc. Among them, learning-based methods have received increasing attention because they can utilize historical data and domain knowledge, as well as handle complex and uncertain environments. In edge computing, nodes are interconnected with complex relationships, and using a knowledge graph, a graph structure, can effectively model its semantic relationships and structural characteristics. However, due to the huge amount of data and obvious distributed characteristics, traditional knowledge graph construction methods are often difficult to effectively process, resulting in low construction efficiency and poor quality of the knowledge graph. In addition, due to the complexity of edge computing data, traditional methods often cannot extract valuable information from it, making the constructed knowledge graph lack depth and breadth.
[0003] Based on the above problems, a method for constructing an edge computing knowledge graph based on representation learning is proposed. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for constructing an edge computing knowledge graph based on representation learning to solve the problems in the background art.
[0005] To achieve the above purpose, the present invention provides a method for constructing an edge computing knowledge graph based on representation learning, including the following steps:
[0006] S1. Abstractly extract entities, entity attributes, and relationships between entities in the edge computing system, and collect key entity, entity attribute, and relationship data between entities;
[0007] S2. Preprocess and data annotate the data collected in S1;
[0008] S3. Automatically extract valuable feature representations from the data obtained by using the representation learning technique S2, map entities and relationships into low-dimensional vectors, realize the numerical representation of entities and relationships in the knowledge graph, and construct a structured knowledge graph;
[0009] S4. Through an efficient link prediction technique, automatically infer and complete unknown or missing entity relationships in the edge computing environment, and quickly match the optimal execution device when a new task appears.
[0010] Preferably, in the S2, the preprocessing process includes the following steps:
[0011] 1) Determine the data sources for the data collected in S1. The data sources include edge device logs, sensor data, and system monitoring data;
[0012] 2) Design different data collection methods for different data sources; specifically:
[0013] For edge device logs, use a log data collection engine to efficiently collect data from multiple sources; the collected logs include the device's network traffic logs and task execution records;
[0014] For sensor data, use a sensor interface to collect data from each sensor;
[0015] For system monitoring data, use a system monitoring and data collection engine.
[0016] 3) Clean and format the raw data;
[0017] 4) Label the data processed in step 3).
[0018] Preferably, in the S2, data cleaning includes removing noise data, removing duplicate data, and filling in missing data. Among them, removing noise data means deleting irrelevant or incorrect data, such as debug information in logs, abnormal readings of sensors, etc.; removing duplicate data means deleting duplicate records to ensure the uniqueness of the data; filling in missing data means using methods such as interpolation and mean filling to complete the missing data;
[0019] Formatting includes unifying the data format and unifying the time format; among them, unifying the data format means converting data from different sources into a unified format, such as JSON format, and unifying the time format is to ensure that the timestamps of all data are consistent;
[0020] The specific process of data annotation is to mark the entities and relationships involved in the data according to the collected data content.
[0021] Preferably, the S3 specifically includes the following content:
[0022] S31. Automatically learn valuable feature representations from the preprocessed data, map entities and relationships to a low-dimensional vector space, numerically represent the entities and relationships in the knowledge graph, and perform pre-training to obtain the embedding vector representations of the basic nodes and relationships;
[0023] S32. Use the graph attention mechanism to aggregate information according to neighbor nodes and relationships, and update the semantic vectors of the nodes;
[0024] S33. Form triples based on the aggregated semantic vectors and predict the rationality of the triples, where the composition of the triples is entity-relationship-entity;
[0025] S34. Compare the prediction results with the true results and use deep learning methods to update the network information.
[0026] Preferably, the S31 specifically includes the following steps:
[0027] 1) For each entity x e calculate the d e -dimensional vector e e through a randomly initialized embedding layer, and for each relationship x r calculate the d r -dimensional vector e r , and assign a corresponding randomly sized d r ×d e random mapping matrix M r ; where d e and d r can be unequal, allowing entities and relationships to be embedded in spaces of different dimensions;
[0028] The embedding layer calculation process is: define a randomly initialized embedding matrix M e , the size of the matrix is d e ×d x , where d x is the feature dimension size of the input data, which varies according to whether the input is an entity vector or a relationship vector, and d e is the dimension size of the output embedding vector; for the data x to be embedded, the embedding process is expressed as:
[0029] e = M e x;
[0030] where e is the embedding vector; the random initialization process of the embedding layer is to randomly initialize the embedding matrix M e ;
[0031] 2) Adopt the mapping matrix M rMap the entity from the original space to the relational space. For a given relation r, the head entity h and the corresponding embedding vector e h , the tail entity t and the corresponding embedding vector e t are transformed through the mapping matrix M r into:
[0032] h r = M r e h , t r = M r e t ;
[0033] 3) To evaluate the rationality of the positive sample triple (h, r, t), define the scoring function as f r (h, t), then the distance between the mapped entity vectors is expressed as:
[0034] f r (h, t) = ||h r + e r - t r ||;
[0035] 4) To optimize the embedding representations of entities and relations, use a margin-based ranking loss function. For each positive sample triple (h, r, t), generate a set of negative sample triples (h′, r, t′). The negative samples are not real triples but are randomly sampled. Define the loss function L kg as:
[0036] L kg = ∑ (h,r,t)∈S ∑ (h′,r,t′)∈S ′ [γ + f r (h, t) - f r (h′, t′)];
[0037] where S represents the set of positive samples, S′ represents the set of negative samples, and γ is the margin parameter;
[0038] 5) Use the loss function L kg to cooperate with the Adam optimizer to complete the pre-training of the embedding layer;
[0039] Learn the embedding representations of entities and relations by minimizing the above loss function, and prepare for subsequent graph attention calculation. Use the Adam algorithm to complete this optimization process.
[0040] Preferably, the S32 specifically includes the following steps:
[0041] 1) Calculate the attention scores between two adjacent nodes. The calculation formula is as follows:
[0042]
[0043] f n (x) = σ(W n x + b n );
[0044] Among them, is the attention score, a hrt is the normalized attention score, N h is the set of all neighbor nodes of the head node, f n is a learnable single-layer perceptron, and n is 1, 2, 3... k;
[0045] In order to consider the relationship between entities when calculating the attention score, subtract the embedded tail entity vector from the relationship entity vector to keep the data distribution the same as when the node is embedded; at the same time, because the proximity between h r +e r and t r has been constrained during node embedding, so a single-layer perceptron f(x) is used to further process the node information during attention score calculation to model deeper semantic information;
[0046] 2) Aggregate the information of all current neighbor nodes weighted according to the obtained attention score, and the calculation process is expressed as:
[0047]
[0048] Among them, is the neighbor node aggregation vector;
[0049] In order to consider the relationship between nodes during aggregation, and at the same time consider the aggregated tail entity and the relationship, add them together with the semantics learned by the embedding layer, and use a single-layer perceptron f(x) to deepen the semantics;;
[0050] 3) Update the original node embedding vector according to the aggregated node vector, and the update method is expressed as:
[0051]
[0052] Among them, f update is the node vector update function used to update the node representation. is the node vector representation after one update;
[0053] 4) Repeat the operation steps 1) - 3), and after t times, obtain the node vector representation updated t times
[0054] Preferably, in S33, the formula for predicting node relationships for the given sample triple (h, r, t) in the training set is:
[0055]
[0056] This calculation formula calculates the correlation between the tail node vector plus the relation vector and the head node vector in the way of a class attention mechanism. After that, a learnable single-layer perceptron is passed, and finally the node relation prediction score is output. The higher the node relation prediction score of the triple in the sample, the better. The function represented by this prediction formula is used as the objective function and trained and optimized using the Adam optimizer to obtain the final model.
[0057] Preferably, step S4 specifically includes the following steps:
[0058] 1) Using the entity embedding vector and the relation mapping matrix obtained in S3, construct the potential link between the task and the device;
[0059] 2) Use the scoring function to evaluate the rationality of the triple, traverse all possible device matching schemes, and select the top three devices with the highest scores as the candidate set of the optimal matching results;
[0060] 3) Select the task execution device from the candidate set of the optimal matching results as the final execution device.
[0061] Therefore, a method for constructing an edge computing knowledge graph based on representation learning of the present invention has the following beneficial effects:
[0062] (1) On the one hand, the present invention proposes a method for constructing an edge computing knowledge graph based on representation learning, solves the problem of constructing a knowledge graph in edge computing, improves the construction efficiency and quality, and extracts valuable information from complex edge computing data, thereby optimizing the effective matching problem between tasks and resources in a large-scale, dynamic and uncertain edge computing environment; on the other hand, it also provides a method for predicting and complementing edge computing knowledge graph relationships. Through efficient link prediction technology, it automatically infers and complements unknown or missing entity relationships in the edge computing environment, mainly focusing on when a new task appears in the scenario, and automatically matches the optimal execution device of the task by giving the resources and data required by the task.
[0063] (2) The present invention significantly improves the efficiency of knowledge graph construction by automatically learning feature representations from a large amount of unlabeled data. At the same time, the introduction of representation learning technology makes it possible to extract deep-level information from complex data, thus ensuring the high quality of the knowledge graph. In addition, the present invention can optimize the calculation task allocation and achieve the efficient matching of tasks to be completed and system resources.
[0064] (3) The present invention applies the graph attention mechanism to the construction process of the edge computing knowledge graph. By calculating the attention weights of node neighbors, the node embedding representation is dynamically adjusted, which can better capture the complex relationships between nodes. Compared with traditional static embedding methods, the graph attention mechanism not only improves the adaptability of the model to complex edge environments but also significantly enhances the modeling ability of the knowledge graph for dynamic relationships between edge devices.
[0065] (4) Through triple prediction and link prediction techniques, the present invention complements unknown or missing entity relationships in the edge computing environment, especially providing effective support for the automatic matching of new devices or new tasks. Compared with existing simple relationship reasoning methods, the link prediction technique of the present invention can efficiently infer missing relationships and improve the accuracy of task allocation and resource matching.
[0066] (5) Through the representation learning method based on the knowledge graph, the present invention constructs a more accurate task and resource matching model. Compared with traditional methods based on rules or simple optimization, the present invention can combine historical data and domain knowledge to achieve efficient task allocation in a dynamic and uncertain edge computing environment. Especially in the case of multi-device concurrency and dynamic task changes, the overall efficiency of the system is significantly improved.
[0067] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0068] Figure 1 It is a schematic diagram of entities, entity attributes, and relationships between entities in an embodiment of the present invention;
[0069] Figure 2 It is a process diagram of the knowledge graph construction in an embodiment of the present invention;
[0070] Figure 3 It is a schematic diagram of data preprocessing in an embodiment of the present invention;
[0071] Figure 4 It is a schematic diagram of the node embedding module in an embodiment of the present invention. Detailed Embodiments
[0072] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0073] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0074] Embodiment
[0075] The present invention provides a method for constructing an edge computing knowledge graph based on representation learning, comprising the following steps:
[0076] S1. Abstract and extract entities, entity attributes, and relationships between entities in the edge computing system, and collect key entity, entity attribute, and relationship data between entities;
[0077] As Figure 1 shown, examples of entities, entity attributes, and relationships between entities in the edge computing system are given:
[0078] (1) Entities:
[0079] a) Devices: Include various edge devices, such as sensors, drones, in-vehicle devices, smartphones, etc.
[0080] i. Attribute: Device type
[0081] b) Data: Data generated or processed on edge devices.
[0082] i. Attribute: Data type
[0083] c) Resources: Computing resources of edge devices, such as CPU, memory, storage, etc.
[0084] i. Attributes: Resource type, resource quantity
[0085] d) Tasks: Tasks that need to be executed on edge devices.
[0086] i. Attribute: None
[0087] (2) Relationships:
[0088] a) Device - owns - Resource
[0089] b) Device - produces - Data
[0090] c) Task - requires - Resource
[0091] d) Device - executes - Task
[0092] e) Task - produces - Data
[0093] f) Data - inputs - Task
[0094] By abstracting the relationships between devices, tasks, resources, and data in edge computing, complex associations between nodes can be effectively captured, and knowledge graphs with deeper semantic depth can be constructed using these relationships to help achieve key functions such as resource scheduling optimization, task allocation, and device management, and improve the intelligent decision-making ability in edge computing scenarios. The construction process of the knowledge graph is as Figure 2 shown.
[0095] S2. As shown in Figure 3 , preprocess the data collected in S1; specifically:
[0096] 1) Determine the data sources. To comprehensively cover various types of information in the edge computing system, multiple data sources need to be determined:
[0097] Edge device logs:
[0098] Content: including device network traffic logs, task execution logs, etc.
[0099] Sensor data:
[0100] Content: data from various sensors, such as temperature, humidity, pressure, GPS, etc.
[0101] System monitoring data:
[0102] Content: including system performance metrics such as network traffic, CPU usage, memory usage, etc.
[0103] 2) Data collection methods. Design different data collection methods for different data sources, specifically as follows:
[0104] For edge device logs, use a log data collection engine to efficiently collect data from multiple sources, clean and preprocess the data, and send the processed data to a specified destination for subsequent analysis and knowledge graph construction. The collected logs include:
[0105] The network traffic logs of the device
[0106] The log fields include:
[0107] A. Timestamp: records the time of the log entry.
[0108] B. Device ID: the ID that uniquely identifies the device.
[0109] C. Network traffic protocol type: the network protocol used by the data packet.
[0110] D. Source IP address: the source IP address of the data packet.
[0111] E. Source port: the source port number of the data packet.
[0112] F. Destination IP: the destination IP address of the data packet.
[0113] G. Destination port: the destination port number of the data packet.
[0114] H. Data packet size: the size of the data packet (in bytes).
[0115] I. Total number of data packets: The number of data packets transmitted within a specific time period.
[0116] J. Total traffic: The total amount of data transmitted within a specific time period (in bytes).
[0117] Task execution record
[0118] Log fields include:
[0119] A. Timestamp: The time when the log entry is recorded.
[0120] B. Task ID: The ID that uniquely identifies the task.
[0121] C. Task name: The name or description of the task.
[0122] D. Task status: The current status of the task.
[0123] E. Execution device ID: The unique identifier of the device that executes the task.
[0124] F. Resource usage: The resource information used during the task execution.
[0125] G. Input data: The input data required for the task execution.
[0126] H. Output data: The output data generated by the task execution.
[0127] I. Error information: The error or exception information that occurred during the task execution.
[0128] J. Execution duration: The execution time of the task from start to end.
[0129] For sensor data, use the sensor interface to collect data from each sensor, clean and preprocess the data, and send the processed data to the specified destination for subsequent analysis and knowledge graph construction.
[0130] The collected data fields include:
[0131] A. Timestamp: The time when the data is collected.
[0132] B. Task ID: The ID that uniquely identifies the sensor.
[0133] C. Sensor type: The type or category of the sensor.
[0134] D. Measured value: The specific value measured by the sensor.
[0135] E. Unit: The unit of the measured value.
[0136] F. Location: The location where the sensor is located or the location where the data is collected.
[0137] For system monitoring data, the system monitoring and data collection engine is used to efficiently collect data from multiple sources, clean and preprocess the data, and send the processed data to the specified destination for subsequent analysis and knowledge graph construction.
[0138] The collected data fields include:
[0139] A. Timestamp: Records the time of data collection.
[0140] B. Task ID: The ID that uniquely identifies the sensor.
[0141] C. Data type: System monitoring data type.
[0142] D. Data measurement value: The specific value of the monitoring data.
[0143] Data preprocessing: Cleans and formats the raw data of edge computing to lay the foundation for representation learning and knowledge graph construction. It consists of the following steps:
[0144] Data cleaning:
[0145] Remove noisy data: Delete irrelevant or incorrect data, abnormal readings of sensors. Set thresholds according to the specifications and expected measurement ranges of sensors, consider readings beyond the thresholds as abnormal, and eliminate them accordingly. At the same time, adopt filtering algorithms to smooth the data and reduce the impact of random noise.
[0146] Remove duplicate data: Delete duplicate records. Use the TF-IDF algorithm and Jaccard similarity calculation method to find duplicate entries in the text set to ensure data uniqueness.
[0147] Fill in missing data: For missing data, use forward filling, that is, fill with the previous observation value to retain the time series characteristics in the data.
[0148] Data formatting:
[0149] Unify data format: Convert data from different sources into a unified JSON format.
[0150] Timestamp processing: Unify the time format to ensure that the timestamps of all data are consistent.
[0151] Data annotation:
[0152] Annotate relevant entities and relationships: According to the data content, indicate the entities and relationships involved in the data. The relationships refer to the entities, attributes, and relationships defined in the above text.
[0153] The annotation rules are as follows
[0154] A. Deduplicate the device IDs for all data, identify all devices existing in the edge computing system, mark them as device entities, and mark the device type attributes according to their data types (host logs or sensor data).
[0155] B. Mark the data entities according to the input data and output data fields in the task running logs, and assign them unique IDs and data types.
[0156] C. Mark the task entities according to the task running logs and mark the task attributes.
[0157] D. Mark the resource entities according to the system monitoring data and mark the resource attributes.
[0158] E. Mark the relationship between device output data according to the sensor data.
[0159] F. Mark the resource relationships owned by the system according to the system resource data.
[0160] G. Mark the connection relationships between device entities according to the network traffic data.
[0161] H. Mark the tasks running on the devices according to the task running logs.
[0162] I. Mark the input-output relationships between tasks and data according to the task running logs.
[0163] J. Mark the resource relationships required by the tasks according to the task running logs.
[0164] S3. Automatically extract valuable feature representations from the data obtained by using the representation learning technology S2, map the entities and relationships into low-dimensional vectors, realize the numerical representation of the entities and relationships in the knowledge graph, and construct a structured knowledge graph; specifically:
[0165] S31. As Figure 4 shown, automatically learn valuable feature representations from the preprocessed data, map the entities and relationships into a low-dimensional vector space, and realize the numerical representation of the entities and relationships in the knowledge graph; perform pre-training before formal training to obtain the basic vector embeddings;
[0166] 1) For each entity x e compute the d e -dimensional vector e e through a randomly initialized embedding layer, and for each relationship x r compute the d r -dimensional vector e r through a randomly initialized embedding layer, and assign a corresponding randomly sized mapping matrix M r ×d e to each relationship; where d r ; among them, de and d r may not be equal, allowing entities and relationships to be embedded into spaces of different dimensions;
[0167] The calculation process of the embedding layer is as follows: Define a randomly initialized embedding matrix M e , the size of the matrix is d e ×d x , where d x is the size of the feature dimension of the input data, which varies according to whether the input is an entity vector or a relationship vector, and d e is the dimension size of the output embedding vector; for the data x to be embedded, the embedding process is expressed as:
[0168] e = M e x;
[0169] where e is the embedding vector; the random initialization process of the embedding layer is to randomly initialize the embedding matrix M e .
[0170] 2) Use the mapping matrix M r to map entities from the original space to the relationship space. For a given relationship r, the head entity h and the corresponding embedding vector e h , the tail entity t and the corresponding embedding vector e t are converted through the mapping matrix M r to:
[0171] h r = M r e h , t r = M r e t ;
[0172] 3) To evaluate the rationality of the positive sample triple (h, r, t), define the scoring function as f r (h, t), then the distance between the mapped entity vectors is expressed as:
[0173] f r (h, t) = ||h r + e r ― t r ||;
[0174] 4) To optimize the embedding representation of entities and relationships, use a margin-based ranking loss function. For each positive sample triple (h, r, t), generate a set of negative sample triples (h′, r, t′). The negative samples are not real triples but are randomly sampled. Define the loss function L kg as:
[0175] L kg = ∑(h,r,t)∈S ∑ (h′,R,T′)∈s′ [γ + f r (h, t) ― f r (h′, t′)];
[0176] Among them, S represents the set of positive samples, s′ represents the set of negative samples, and γ is the margin parameter;
[0177] 5) Use the loss function L kg Cooperate with the Adam optimizer to complete the pre-training of the embedding layer. Learn the embedding representations of entities and relationships by minimizing the above loss function, and prepare for subsequent graph attention calculation. Use the Adam algorithm to complete this optimization process.
[0178] S32. Use the graph attention mechanism to aggregate information according to neighbor nodes and relationships, and update the semantic vectors of nodes;
[0179] 1) Calculate the attention scores between two adjacent nodes. The calculation formula is as follows:
[0180]
[0181] f n (x) = σ(W n x + b n );
[0182] Among them, is the attention score, a hrt is the normalized attention score, N h is the set of all neighbor nodes of the head node, f n is a learnable single-layer perceptron, and n is 1, 2, 3... k; In order to consider the relationship between entities when calculating the attention score, subtract the embedded tail entity vector from the relationship entity vector to keep the data distribution the same as when the nodes are embedded; At the same time, because the proximity between h r + e r and t r has been constrained during node embedding, so use the single-layer perceptron f(x) to further process the node information during attention score calculation to model deeper semantic information.
[0183] 2) Aggregate the information of all current neighbor nodes weighted according to the obtained attention scores. The calculation process is expressed as:
[0184]
[0185] Among them, is the neighbor node aggregation vector;
[0186] To consider the relationships between nodes during aggregation, and also consider the aggregated tail entity and the relationship, add them by means of the semantics learned by the embedding layer, and use a single-layer perceptron f(x) to deepen the semantics; 3) Update the original node embedding vector according to the aggregated node vector, and the update method is expressed as:
[0187]
[0188] where f update is the node vector update function used to update the node representation. is the node vector representation after one update;
[0189] 4) Repeat the operation steps 1) - 3), and after performing t times, obtain the node vector representation updated t times
[0190] S33. According to the aggregated semantic vectors, form triples and predict the rationality of the triples. Among them, the composition of the triples is entity - relationship - entity;
[0191] The formula for predicting node relationships for the given sample triples (h, r, t) in the training set is:
[0192]
[0193] This calculation formula calculates the correlation between the tail node vector plus the relationship vector and the head node vector in the way of a class attention mechanism. After that, through a learnable single-layer perceptron, finally output the node relationship prediction score. The higher the node relationship prediction score of the triples in the sample, the better. Take the function represented by this prediction formula as the objective function and use the Adam optimizer for training optimization to obtain the final model.
[0194] S34. Compare the prediction results with the true results and use deep learning methods to update the network information.
[0195] S4. Through efficient link prediction technology, automatically infer and complete unknown or missing entity relationships in the edge computing environment, and when a new task appears, quickly match the optimal execution device; specifically:
[0196] 1) Use the entity embedding vectors and relationship mapping matrices obtained in S3 to construct potential links between tasks and devices;
[0197] 2) Use a scoring function to evaluate the rationality of the triples, traverse all possible device matching schemes, and select the top three devices with the highest scores as the candidate set of the optimal matching results;
[0198] 3) Select the task execution device from the candidate set of the optimal matching results as the final execution device.
[0199] In the above method, noise and null data are processed in the data preprocessing part to make the system more stable and robust, enabling it to handle various extreme situations; the model training method of using pre-training plus post-training allows the information in the embedding layer to obtain basic semantic information in advance in a simple training process, enabling the model to converge faster in subsequent formal training and improving training efficiency; by calculating the attention weights of node neighbors and dynamically adjusting the node embedding representation, complex relationships between nodes can be better captured. Compared with traditional static embedding methods, the graph attention mechanism not only improves the adaptability of the model to complex edge environments, but also significantly enhances the ability of the knowledge graph to model dynamic relationships between edge devices.
[0200] Therefore, the present invention relates to a method for constructing an edge computing knowledge graph based on representation learning, which solves the problem of constructing a knowledge graph in edge computing and optimizes the effective matching problem of tasks and resources in a large-scale, dynamic and uncertain edge computing environment; and it can quickly complete the missing relationships in the knowledge graph when facing new devices and tasks, realizing intelligent matching between devices and tasks.
[0201] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that: they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for constructing an edge computing knowledge graph based on representation learning, characterized in that, It includes the following steps: S1. Abstract and extract entities, entity attributes, and relationships between entities in the edge computing system, and collect key entity, entity attribute, and relationship data between entities. Specifically as follows: (1) Entities: a) Devices: Include various edge devices, sensors, drones, in-vehicle devices, and smartphones; i. Attributes: Device type; b) Data: Data generated or processed on edge devices; i. Attributes: Data type; c) Resources: Computing resources of edge devices, including CPUs, memory, and storage; i. Attributes: Resource type, resource quantity; d) Tasks: Tasks that need to be executed on edge devices; i. Attributes: None; (2) Relationships: a) Device - owns - Resource; b) Device - produces - Data; c) Task - requires - Resource; d) Device - executes - Task; e) Task - produces - Data; f) Data - inputs - Task; By abstracting the relationships between devices, tasks, resources, and data in edge computing, complex inter-node associations can be effectively captured, and knowledge graphs with deeper semantic depth can be constructed using these relationships to help achieve key functions such as resource scheduling optimization, task allocation, and device management, and improve the intelligent decision-making ability in edge computing scenarios; S2. Preprocess the data collected in S1, including the following steps: 1) Determine the data sources for the data collected in S1. To comprehensively cover various types of information in the edge computing system, multiple data sources need to be determined. The data sources include edge device logs, sensor data, and system monitoring data. Specifically: Edge device logs: Content: Include device network traffic logs and task execution logs; Sensor data: Content: Data from various sensors, such as temperature, humidity, pressure, GPS; System monitoring data: Content: Include system performance metrics such as network traffic, CPU usage, and memory usage; 2) Design different data collection methods for different data sources, specifically as follows: For edge device logs, use a log data collection engine to collect data from multiple sources, clean and preprocess the data, and send the processed data to a specified destination for subsequent analysis and knowledge graph construction; For sensor data, use a sensor interface to collect data from each sensor, clean and preprocess the data, and send the processed data to a specified destination for subsequent analysis and knowledge graph construction; For system monitoring data, use a system monitoring and data collection engine to efficiently collect data from multiple sources, clean and preprocess the data, and send the processed data to a specified destination for subsequent analysis and knowledge graph construction; 3) Clean and format the raw data. Data cleaning includes removing noise data, removing duplicate data, and filling in missing data. Among them, removing noise data means deleting irrelevant or incorrect data, abnormal readings of sensors. Set thresholds according to the specifications and expected measurement ranges of sensors, and regard readings exceeding the thresholds as abnormal and eliminate them accordingly. At the same time, use filtering algorithms to smooth the data and reduce the impact of random noise; Removing duplicate data means deleting duplicate records. By using the TF-IDF algorithm and the Jaccard similarity calculation method, duplicate entries in the text collection are found to ensure the uniqueness of the data; Filling in missing data means that for missing data, forward filling is used, i.e., filling with the previous observation value, to retain the time series characteristics in the data; Formatting includes unifying the data format and the time format; 4) Data annotation is performed on the data processed in step 3). The specific process of data annotation is to mark the entities and relationships involved in the data according to the collected data content. The annotation rules are as follows: A. De-duplicate the device ID for all data, find all devices existing under the edge computing system, mark them as device entities, and mark the device type attribute according to their data types; B. Mark data entities according to the input data and output data fields in the task running log, and assign a unique ID to them, along with the data type; C. Mark task entities according to the task running log and mark task attributes; D. Mark resource entities according to the system monitoring data and mark resource attributes; E. Mark the relationship between device output data according to the sensor data; F. Mark the relationship between the resources owned by the system according to the system resource data; G. Mark the connection relationship between device entities according to the network traffic data; H. Mark the tasks running on the device according to the task running log; I. Mark the input-output relationship between tasks and data according to the task running log; J. Mark the resource relationship required by the task according to the task running log; S3. Use representation learning techniques to automatically extract valuable feature representations from the data obtained in S2, map entities and relationships into low-dimensional vectors, realize the numerical representation of entities and relationships in the knowledge graph, and construct a structured knowledge graph, which specifically includes the following contents: S31. Automatically learn valuable feature representations from the preprocessed data, map entities and relationships into a low-dimensional vector space, perform numerical representation of entities and relationships in the knowledge graph, complete pre-training, and obtain the embedding vector representations of basic nodes and relationships; S32. Use the graph attention mechanism to aggregate information according to neighbor nodes and relationships, and update the semantic vectors of nodes; S33. Form triples according to the aggregated semantic vectors and predict the rationality of the triples. Among them, the composition of the triples is entity-relationship-entity; S34. Compare the prediction results with the real results and use deep learning methods to update the network information; S4. Through efficient link prediction techniques, automatically infer and complete unknown or missing entity relationships in the edge computing environment, and quickly match the optimal execution device when a new task appears. Specifically: 1) Use the entity embedding vectors and relationship mapping matrices obtained in S3 to construct potential links between tasks and devices; 2) Use a scoring function to evaluate the rationality of triples, traverse all possible device matching schemes, and select the top three devices with the highest scores as the candidate set of the optimal matching results; 3) Select the task execution device from the candidate set of the optimal matching results as the final execution device.
2. The method for constructing an edge computing knowledge graph based on representation learning according to claim 1, wherein: The specific steps of S31 are as follows: 1) Define the randomly initialized embedding matrix as M e , with size d e ×d x , where d x is the size of the feature dimension of the input data, and d e is the dimension size of the output embedding vector; for the data x to be embedded, the embedding process is expressed as: e = M e x; where e is the embedding vector; 2) Use the mapping matrix M r Map the entity from the original space to the relational space. For a given relation r, the head entity h and the corresponding embedding vector e h , the tail entity t and the corresponding embedding vector e t Through the mapping matrix M r The conversion results in: h r = M r e h 、t r = M r e t ; 3) To evaluate the rationality of the positive sample triple (h, r, t), the scoring function is defined as f r (h, t), and the distance between the mapped entity vectors is expressed as: f r (h,t) = ||h r + e r − t r ||; 4) For each positive sample triple (h, r, t), generate a set of negative sample triples (h′, r, t′), and define the loss function L kg as: L kg = ∑ (h,r,t)∈S ∑ (h′,r,t′)∈S′ [γ + f r (h, t) ― f r (h′, t′)]; Among them, $S$ represents the set of positive samples, $S'$ represents the set of negative samples, and $\gamma$ is the margin parameter; 5) Use the loss function L kg Complete the pre-training of the embedding layer in conjunction with the Adam optimizer.
3. The method for constructing an edge computing knowledge graph based on representation learning according to claim 2, characterized in that: The specific steps of $S32$ are as follows: 1) Calculate the attention scores between two adjacent nodes, and the calculation formula is as follows: f n (x) = σ(W n x + b n ); Among them, is the attention score, a hrt is the normalized attention score, N h is the set of all neighbor nodes of the head node, f n is a learnable single-layer perceptron, where n = 1, 2, 3... k; 2) Aggregate the information of all current neighbor nodes weighted according to the obtained attention scores, and the calculation process is expressed as: Among them, is the aggregated vector of neighbor nodes; 3) Update the original node embedding vector according to the aggregated node vector, and the update method is expressed as: Among them, f update is the node vector update function for updating the node representation, is the node vector representation after one update; 4) Repeat the operation steps 1) - 3), and after t times, the node vector representation updated t times is obtained.
4. A method for constructing an edge computing knowledge graph based on representation learning according to claim 3, characterized in that: In $S33$, the formula for predicting node relationships for the given sample triple $(h, r, t)$ in the training set is: Use this prediction formula as the objective function and use the Adam optimizer for training optimization to obtain the final model.
Citation Information
Patent Citations
Metalearning-based medical common sense knowledge graph automatic construction method
CN116861001A
Multi-modal knowledge graph establishment method and application
CN117131933A