Data processing method and device

By building a knowledge graph to determine the popularity of data entries and partitioning the storage, the rationality of hot and cold data storage in big data scenarios is solved, and data access performance and storage efficiency are improved.

CN120764646APending Publication Date: 2025-10-10LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510873182.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

In big data scenarios, it is difficult to store data reasonably while taking into account data access performance, especially the distinction and storage of hot data and cold data.

Method used

By constructing a knowledge graph, the dynamic attributes and association relationships of nodes are used to determine the popularity of data entries, and then the data entries are stored in the storage area corresponding to the corresponding popularity attribute category to achieve partitioned storage of hot data and cold data.

Benefits of technology

It achieves reasonable data storage while taking into account data access performance, improves the efficiency and accuracy of data storage, and ensures fast access to hot data and stable storage of cold data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120764646A_ABST
    Figure CN120764646A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, and the method comprises the steps: obtaining a knowledge graph corresponding to a plurality of data entries, the knowledge graph comprises a plurality of nodes, each node represents an entity determined in one data entry, and the attributes of the nodes at least comprise dynamic attributes, the dynamic attributes of the nodes are used for representing accessed information of data entries corresponding to the nodes; for a target node in the knowledge graph, determining the popularity of a target data entry corresponding to the target node based on the dynamic attribute of the target node and the association relationship between the target node and other nodes in the knowledge graph; based on the popularity of the target data entry, determining a popularity attribute category of the target data entry, the popularity attribute category being used for indicating that the target data entry is hot data or cold data; and storing the target data entry into a target storage area corresponding to the popularity attribute category of the target data entry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data processing method and device. Background Art

[0002] With the continuous advancement of digitalization, the data generated by enterprises and institutions is rapidly increasing. This increasing amount of data has placed higher demands on data storage. However, in big data scenarios, it is difficult to store data reasonably while taking into account data access performance. Summary of the Invention

[0003] In one aspect, the present application provides a data processing method, comprising:

[0004] Obtaining a knowledge graph corresponding to the plurality of data entries, the knowledge graph comprising a plurality of nodes, each node representing an entity determined in a data entry, the attributes of the node comprising at least a dynamic attribute, the dynamic attribute of the node being used to characterize accessed information of the data entry corresponding to the node;

[0005] For a target node in the knowledge graph, determining the popularity of a target data entry corresponding to the target node based on the dynamic attributes of the target node and the association relationship between the target node and other nodes in the knowledge graph;

[0006] determining a heat attribute category of the target data entry based on the heat of the target data entry, the heat attribute category being used to indicate whether the target data entry is hot data or cold data;

[0007] The target data entry is stored in a target storage area corresponding to the heat attribute category of the target data entry.

[0008] In one possible implementation, obtaining the knowledge graphs corresponding to the multiple data entries includes at least one of the following:

[0009] Based on the multiple data entries to be stored, construct a knowledge graph corresponding to the multiple data entries;

[0010] In response to the currently obtained new data entry, adding a node corresponding to the new data entry to the constructed knowledge graph to obtain an updated knowledge graph corresponding to the multiple data entries;

[0011] In response to an update in accessed information of a data entry corresponding to a node in a knowledge graph, the dynamic attributes of the node corresponding to the data entry in the knowledge graph are updated to obtain an updated knowledge graph.

[0012] In another possible implementation, determining, for a target node in the knowledge graph, the popularity of a target data entry corresponding to the target node based on dynamic attributes of the target node and associations between the target node and other nodes in the knowledge graph, includes:

[0013] For a target node in the knowledge graph whose popularity has not yet been determined or whose dynamic attributes have been updated, if the time between the generation time of the target data entry corresponding to the target node and the current moment is greater than the set time, the popularity of the target data entry corresponding to the target node is determined based on the dynamic attributes of the target node and the association between the target node and other nodes in the knowledge graph, where the association includes at least one of the degree and centrality of the target node;

[0014] If the generation time of the target data entry corresponding to the target node is not longer than the set time from the current moment, the heat of the target data entry corresponding to the target node is determined based on the heat corresponding to at least one neighbor node of the target node in the knowledge graph and the common access time between the target node and the neighbor node. The heat corresponding to the neighbor node is the heat of the data entry corresponding to the neighbor node, and the common access time is the time when the target data entry corresponding to the target node and the data entry corresponding to the neighbor node were last jointly accessed.

[0015] In another possible implementation, determining the heat of the target data entry corresponding to the target node based on the heat corresponding to at least one neighbor node of the target node in the knowledge graph and the common access time between the target node and the neighbor node includes:

[0016] Determine at least one neighbor node of the target node from the knowledge graph;

[0017] For each neighbor node, determine a weight corresponding to the neighbor node based on the time between the common access time of the target node and the neighbor node and the current time;

[0018] Based on the heat and weighted weight corresponding to each of the neighbor nodes, the heat of the target data entry corresponding to the target node is determined.

[0019] In another possible implementation, determining at least one neighbor node of the target node from the knowledge graph includes:

[0020] Sampling at least two neighbor nodes from a set of neighbor nodes belonging to the target node in the knowledge graph, wherein the at least two neighbor nodes include at least one neighbor node whose corresponding data entry belongs to hot data and at least one neighbor node whose corresponding data entry belongs to cold data;

[0021] The determining the heat of the target data entry corresponding to the target node based on the heat and weighted weight corresponding to each of the neighboring nodes includes:

[0022] Determine, from the knowledge graph, a target number of layers of sub-knowledge graphs extending from the target node through the at least two neighboring nodes;

[0023] Based on the heat and weighted weight corresponding to each neighbor node in the sub-knowledge graph, as well as the heat corresponding to each layer of child nodes of the neighbor node, the heat corresponding to each neighbor node in the sub-knowledge graph and the heat corresponding to each layer of child nodes of the neighbor node are weightedly aggregated in combination with the graph neural network algorithm to obtain the heat of the target data entry corresponding to the target node.

[0024] In yet another possible implementation, the method further includes:

[0025] If there is at least one to-be-archived node in the knowledge graph whose corresponding data entry meets the archiving condition, determining the edge weights of the edges between the nodes in the knowledge graph based on the association relationships between the nodes in the knowledge graph;

[0026] Divide the knowledge graph into at least one graph community based on the popularity of data entries corresponding to each node in the knowledge graph and the edge weights of edges between different nodes, each graph community including at least one node;

[0027] Determine a target graph community where the node to be archived is located, determine at least one data entry to be archived from data entries corresponding to each node in the target graph community, and archive the at least one data entry to be archived.

[0028] In another possible implementation, dividing the knowledge graph into at least one graph community based on the popularity of data entries corresponding to each node in the knowledge graph and the edge weights of edges between different nodes includes:

[0029] Determine the initial community to which each node in the knowledge graph belongs, and take the initial community to which the node belongs as the first community to which the node currently belongs;

[0030] For each node, determining the node connection closeness corresponding to the first community where the node is located based on the popularity of data entries corresponding to each node in the first community where the node is located and the edge weights of the edges between the nodes;

[0031] Optimizing the first community where each node is located to obtain the second community where each node is located, wherein the node connection closeness corresponding to the second community where the node is located is higher than the node connection closeness corresponding to the first community where the node is located;

[0032] The second community where the node is located is determined as the graph community to which the node belongs, to obtain at least one graph community.

[0033] In another possible implementation, determining the node connection closeness corresponding to the first community where the node is located based on the popularity of data entries corresponding to each node in the first community where the node is located and the edge weights of edges between the nodes includes:

[0034] For a node pair consisting of any two nodes in the first community where the node is located, determining a weight of the node pair based on an edge weight of an edge between the two nodes in the node pair and the popularity of data entries corresponding to the two nodes in the node pair;

[0035] The node connection closeness of the first community where the node is located is determined based on the weight of each node pair in the first community where the node is located and the popularity of the data entry corresponding to each node in each node pair.

[0036] In yet another possible implementation, storing the target data entry in a target storage area corresponding to the heat attribute category of the target data entry includes:

[0037] If the target data entry is a data entry that has not been stored, storing the target data entry in a target storage area corresponding to the heat attribute category of the target data entry;

[0038] If the target data entry belongs to a stored data entry and the current storage area where the target data entry is currently located does not match the heat attribute category of the target data entry, the target data entry is migrated from the current storage area to a target storage area that matches the heat attribute category of the target data entry.

[0039] In another aspect, the present application further provides a data processing device, comprising:

[0040] a graph obtaining unit, configured to obtain a knowledge graph corresponding to a plurality of data entries, wherein the knowledge graph includes a plurality of nodes, each node representing an entity determined in a data entry, and the attributes of the node include at least a dynamic attribute, wherein the dynamic attribute of the node is used to represent accessed information of the data entry corresponding to the node;

[0041] a popularity determination unit, configured to determine, for a target node in the knowledge graph, the popularity of a target data entry corresponding to the target node based on the dynamic attributes of the target node and the association relationship between the target node and other nodes in the knowledge graph;

[0042] a category determining unit, configured to determine a heat attribute category of the target data entry based on the heat of the target data entry, wherein the heat attribute category is used to indicate whether the target data entry is hot data or cold data;

[0043] The data storage unit is configured to store the target data entry into a target storage area corresponding to the heat attribute category of the target data entry. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0045] Figure 1 A flowchart of the data processing method provided in this application;

[0046] Figure 2 A flowchart of another data processing method provided by this application;

[0047] Figure 3 A schematic diagram of a process for determining the heat of a target data entry corresponding to a target node in this application;

[0048] Figure 4 A flowchart of another data processing method provided by this application;

[0049] Figure 5 A schematic diagram of an implementation process for dividing a knowledge graph into at least one graph community in this application;

[0050] Figure 6 A schematic diagram of the structure of a data processing device provided in this application;

[0051] Figure 7 A schematic diagram of another structure of the data processing device provided by this application;

[0052] Figure 8 A schematic diagram of the composition architecture of the electronic device provided in this application. DETAILED DESCRIPTION

[0053] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation methods of the present application are only used to explain the specific embodiments of the present application and are not intended to limit the present application. It is known to those skilled in the art that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0054] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0055] like Figure 1 , shows a flow chart of the data processing method provided by the present application. The method of this embodiment can be applied to electronic devices, which can be personal computers, servers, or nodes in cloud platforms, etc., without limitation. This embodiment includes:

[0056] S101, obtaining a knowledge graph corresponding to multiple data entries.

[0057] The knowledge graph consists of multiple nodes, each representing an entity identified in a data entry. The entity represented by a node is extracted from a data entry. Therefore, each node corresponds to a data entry, and the entity represented by the node is an object such as an event, person, product, order, or material in the data entry.

[0058] In this application, a data entry may include entity-related attributes. The entity attributes in a data entry may include, but are not limited to, one or more of the entity's name, model, status, generation time, origin, and amount. The type and number of entity-related attributes included in a data entry may vary depending on the domain from which the data entry originates.

[0059] For example, taking the data entry as data related to a product, the data entry may include attribute data such as the product name, model, material number, and production time.

[0060] For another example, if a data entry is data related to an order, then the data entry may include attribute data such as order ID, order time, and order amount.

[0061] In this application, each data entry is also associated with access information for that data entry. This access information can reflect how the data entry has been accessed before the current moment. Therefore, this access information can also be referred to as historical access information for the data entry. It is understood that accessing a data entry not only includes reading (or querying) the data entry, but also includes adding data content to the data entry, deleting or modifying part of the data entry's information, etc. Therefore, the access information for the data entry can include, but is not limited to, one or more of: the access time of the data entry (such as the access time of the most recent access within a set period or the access time of the most recent access), the number of accesses, the access frequency, the update frequency (data entry updates due to additions, modifications, or deletions of part of the data entry's content), and the number of updates. The access information for a data entry can be stored as the data content of the data entry, or in association with the data entry, without limitation.

[0062] It can be seen that the subject described by each data entry can be abstracted as a node in the knowledge graph, so that one node corresponds to one data entry. There is no restriction on the specific implementation of extracting entities from data entries.

[0063] It is understood that each node in the knowledge graph has its own attributes. In this application, the attributes of a node include at least its dynamic attributes. The dynamic attributes of a node are used to represent the access information of the data entry corresponding to the node. For example, the dynamic attributes of a node may include, but are not limited to, one or more of the access time, number of accesses, or access frequency of the data entry corresponding to the node within the most recent set period of time.

[0064] Furthermore, to clarify the relationships between nodes in the knowledge graph (such as whether there are edges between nodes), node attributes also include static attributes of the node. These static attributes include the attributes of the entity determined from the data entry corresponding to the node. Because the attributes of each entity in the data entry do not change with changes such as the number of times the data entry is accessed, the attributes of the entity in the data entry can be called the static attributes of the node.

[0065] Based on this, the knowledge graph also includes edges between different nodes, where there are edges between nodes with associated static attributes, and there are no edges between nodes with no associated static attributes.

[0066] For example, suppose the knowledge graph includes nodes 1, 2, and 3, where node 1 represents product aa extracted from data entry 11, node 2 represents order bb extracted from data entry 22, and node 3 represents order cc extracted from data entry 33. Data entry 22 of order bb describes static attributes related to the order, such as the quantity and amount of product aa purchased, while order cc is an order to purchase a certain device and has no association with product aa. Therefore, in the knowledge graph, there will be an edge between nodes 1 and 2, but there will not be an edge between nodes 1 and 3. Of course, if the supplier or buyer of order cc is the same as that of order bb, or there is some other association, then an edge can also exist between nodes 2 and 3.

[0067] In this application, the knowledge graph obtained can be a pre-constructed knowledge graph, a knowledge graph constructed in real time, or a currently updated knowledge graph (such as the presence of new nodes or updates to static attributes of nodes or dynamic data, etc.), without specific restrictions.

[0068] S102, for a target node in the knowledge graph, determine the popularity of a target data entry corresponding to the target node based on the dynamic attributes of the target node and the association relationship between the target node and other nodes in the knowledge graph.

[0069] In this application, the popularity of the target data entry corresponding to the target node may also be referred to as the target node's popularity, which is used to characterize the access characteristics of the target data entry corresponding to the target node. The popularity of the target data entry corresponding to the target node can reflect part or all of the information such as the frequency of access to the target data entry and whether the time from the last access to the current time is too long. For example, the popularity of the target data entry corresponding to the target node can characterize the frequency of access to the target data entry.

[0070] The target node is a node in the knowledge graph. Depending on the application scenario, the storage status of each data entry in the knowledge graph, and the update status of the knowledge graph, the target node can have multiple possibilities.

[0071] For example, in the first possible scenario, the target node can be any node in the knowledge graph. For example, if the popularity of the data entries corresponding to each node in the knowledge graph is not determined, it means that the data entries corresponding to each node in the knowledge graph have not been properly stored. Therefore, each node in the knowledge graph can be used as a target node to determine the popularity of the data entries corresponding to each node.

[0072] In the second possible scenario, the target node can be a newly added node in the knowledge graph. For ease of description, the newly added node in the knowledge graph is referred to as a newly added node. It is understandable that if there is a newly added new data entry, then it is necessary to add a node corresponding to the new data entry to the knowledge graph. The popularity of the new data entry corresponding to the newly added node in the knowledge graph has not yet been determined. In order to determine the storage area suitable for the new data entry, it is necessary to determine the popularity of the new data entry corresponding to the newly added node. Based on this, the target node can be a newly added node in the knowledge graph.

[0073] It is understandable that in the above two possible situations, in essence, nodes in the knowledge graph whose corresponding popularity has not yet been determined (i.e., the popularity of the corresponding data entry has not yet been determined) are used as target nodes. For example, in the first possible situation, each node in the knowledge graph needs to be used as a target node separately, and step S102 is executed. If a node in the knowledge graph has already been used as a target node and the popularity of the corresponding data entry has been determined, there is no need to repeatedly use the node as a target node. Therefore, only the nodes in the knowledge graph whose corresponding popularity has not yet been determined need to be used as target nodes.

[0074] In the third possible scenario, considering that the access information of the data entry corresponding to the node in the knowledge graph changes, the popularity of the data entry will change, and the change in the access information of the data entry will inevitably synchronize the update of the dynamic attributes of the node corresponding to the data entry in the knowledge graph. Based on this, the present application can calculate the popularity of the data entry corresponding to the node when the dynamic attributes corresponding to the node change. Therefore, the target node can be the node whose dynamic attributes are updated in the knowledge graph.

[0075] In this application, the association relationship between the target node and other nodes in the knowledge graph can be the association information between the target node and other nodes other than the target node analyzed from the knowledge graph, and the importance of the target node can be reflected through this association relationship. For example, the association relationship between the target node and other nodes can include: the number of nodes directly connected to the target node, the popularity of other nodes connected to the target node, and the impact of the target node on accessing data entries corresponding to other nodes, etc., without specific limitation.

[0076] It can be understood that, unlike determining the popularity of a data entry based solely on its historical access frequency, the present application extracts entities from the data entry as nodes in the knowledge graph, thereby enabling a more comprehensive and efficient analysis of the association between the data entry and the data entries corresponding to other nodes based on the knowledge graph. On this basis, the access information of the data entry represented by the dynamic attributes of the node and the association between the node and other nodes can more comprehensively reflect the possible future access situations of the data entry, thereby enabling a more accurate determination of the popularity of the data entry.

[0077] S103 : Determine a popularity attribute category of the target data entry based on the popularity of the target data entry.

[0078] The heat attribute category is used to indicate whether the target data entry is hot data or cold data. Hot data refers to data with a relatively high access frequency or a relatively high probability of being accessed in the future, while cold data refers to data with a relatively low access frequency or a relatively low probability of being accessed in the future.

[0079] For example, if the heat of a target data entry exceeds a heat threshold, the target data entry is determined to be hot data; if the heat of a target data entry does not exceed the heat threshold, the target data entry is determined to be cold data. The heat threshold can be set as needed and is adjustable. For example, the heat threshold can be dynamically adjusted based on the usage of the storage area corresponding to the heat attribute category. For example, if the usage of the storage area corresponding to the hot data exceeds 80%, the specific value of the heat threshold can be increased.

[0080] S104: Store the target data entry into a target storage area corresponding to the heat attribute category of the target data entry.

[0081] In this application, different heat attribute categories correspond to different storage areas, so that data can be reasonably stored according to the heat attribute category of the data.

[0082] For example, hot data has a relatively high access frequency or a high probability of being accessed recently. Therefore, the storage area corresponding to hot data can be a storage area of ​​a memory with a high access speed, such as a cache. Cold data has a relatively low access frequency or a relatively low probability of being accessed recently. Therefore, the storage area corresponding to cold data can be a storage area of ​​a storage medium with a low access speed.

[0083] It can be understood that in the case that the knowledge graph is newly constructed and the dynamic attributes corresponding to the nodes in the knowledge graph are updated, the heat of the target data entry corresponding to the target node needs to be determined, and therefore the target data entry corresponding to the target node can be a data entry that has not been stored or can have been stored in a target storage area corresponding to a heat attribute category. Based on this, for any target data entry corresponding to a target node, if the target data entry is a data entry that has not been stored, the target data entry is stored in a target storage area corresponding to a heat attribute category of the target data entry; if the target data entry is a stored data entry, and a current storage area where the target data entry is currently located does not match a heat attribute category of the target data entry that is currently determined, the target data entry is migrated from the current storage area to a target storage area that matches the heat attribute category of the target data entry that is currently determined.

[0084] From the above, in the present application, the knowledge graph is used to represent the relationship between multiple data entries. For a target node in the knowledge graph, the dynamic attribute of the target node and the association relationship between the target node and other nodes can not only represent the access situation of the data entry corresponding to the target node, but also represent the association relationship between the data entry corresponding to the target node and other data entries. Therefore, based on the dynamic attribute of the target node and the association relationship between the target node and other nodes in the knowledge graph, the access feature of the data entry can be more comprehensively analyzed, and the heat of the data entry can be more accurately determined. On this basis, based on the heat attribute category to which the heat of the data entry belongs, the data entries of different heat attribute categories are stored in different storage areas, so that the data entries can be more accurately and reasonably stored while the access performance of the data entries is taken into account.

[0085] It can be understood that the data read and write speed of the storage area corresponding to the hot data is relatively fast, but the problem of data power loss cannot be effectively avoided. In order to ensure the storage reliability of the data entries, for the data entries belonging to the hot data, after the data entries are stored in the storage area corresponding to the hot data, the data entries belonging to the hot data are also synchronized to the storage area corresponding to the cold data.

[0086] In particular, when synchronizing data entries belonging to hot data to the storage area corresponding to cold data, the data synchronization priority can also be determined based on the update frequency of the data entries belonging to the hot data and the popularity of other nodes associated with the node corresponding to the data entry. Specifically, data entries with high update and access frequencies, or with relatively high popularity of other associated nodes, will have a relatively low synchronization priority. Therefore, data entries with low update and access frequencies, or with relatively low popularity of other associated nodes, can be prioritized for synchronization to the storage area corresponding to cold data to reduce data fluctuations.

[0087] It is understandable that the solution of this application can achieve the purpose of storing multiple data entries in different storage areas according to cold data and hot data with the help of knowledge graph. On this basis, this application can be executed when storing multiple data entries that need to be partitioned, or when there are updates to the partitioned stored data entries, or when there are new data entries. Based on this, there are many possible situations in which the knowledge graph obtained in this application can be:

[0088] In the first possible scenario, a knowledge graph corresponding to the multiple data entries to be stored can be constructed based on the multiple data entries to be stored. The multiple data entries to be stored are multiple data entries that have not been stored using the solution of this application, that is, multiple data entries that have not yet been stored according to the heat attribute category. In this case, since a knowledge graph does not yet exist, it is necessary to construct a knowledge graph.

[0089] Among them, the specific implementation of constructing a knowledge graph based on multiple data entries can be unrestricted. For example, for each data entry, the entity in the data entry can be extracted to generate a node for representing the entity; the attribute corresponding to the entity in the data entry is determined, and the attribute corresponding to the entity in the data entry is determined as the static attribute of the node corresponding to the data entry. Based on the accessed information of the data entry, the dynamic attribute of the node corresponding to the data entry is determined. On this basis, based on the static attributes of each node, the nodes that are associated with each other can be determined. For any two nodes that are associated with each other, an edge between the two nodes can be constructed, thereby obtaining a knowledge graph containing multiple nodes and edges between nodes.

[0090] In the second possible scenario, in response to the new data entry currently obtained, a node corresponding to the new data entry is added to the constructed knowledge graph to obtain an updated knowledge graph corresponding to multiple data entries.

[0091] The new data entry is a newly generated data entry, and the constructed knowledge graph is a knowledge graph constructed using multiple existing data entries before obtaining the new data entry. The constructed knowledge graph may be the knowledge graph constructed in the first possible scenario.

[0092] For example, when a company adds a new order data, in order to determine the popularity attribute category corresponding to the order data and reasonably store the order data, it is necessary to add a node corresponding to the order data to the knowledge graph constructed based on historical order data to update the knowledge graph and determine the popularity of the newly added order data based on the updated knowledge graph.

[0093] The implementation of building a node for the new data entry can be found in the previous related introduction. It can be understood that adding the node of the new data entry to the knowledge graph not only includes adding the node of the new data entry to the knowledge graph, but also includes establishing edges between the node corresponding to the new data entry and other nodes already in the knowledge graph. In particular, the static data of the node corresponding to the new data entry can be combined to determine the associated nodes associated with the node of the new data entry, and establish edges between the node of the new data entry and its associated nodes.

[0094] In the third possible scenario, in response to an update in the accessed information of a data entry corresponding to a node in the knowledge graph, the dynamic attributes of the node corresponding to the data entry in the knowledge graph are updated to obtain an updated knowledge graph. It is understandable that after the dynamic attributes of a node change, the popularity of the data entry corresponding to the node will also change. Therefore, it is necessary to redetermine the popularity of the data entry corresponding to the node and redefine the target storage area where the data entry should be stored.

[0095] In the present application, based on the different possible situations of obtaining the knowledge graph, there may be multiple situations for the target nodes that need to calculate the corresponding heat. In different situations of the target nodes, the generation time of the target data entry corresponding to the target node and the amount of information of the accessed information associated with the target data entry will be different. Accordingly, in order to more accurately determine the heat of the target data entry corresponding to the target node, the specific implementation of determining the heat of the target data entry corresponding to the target node will also be different in different situations of the target node.

[0096] The following combination Figure 2 For explanation. Figure 2 Another flow chart of the data processing method provided by the present application is shown. The method of this embodiment may include:

[0097] S201, obtaining a knowledge graph corresponding to multiple data entries.

[0098] The knowledge graph includes multiple nodes, each node represents an entity determined in a data entry, and the attributes of the node include at least dynamic attributes. The dynamic attributes of the node are used to represent the accessed information of the data entry corresponding to the node.

[0099] This step can be referred to the relevant introduction of the previous embodiment and will not be repeated here.

[0100] S202: For any target node in the knowledge graph whose popularity has not yet been determined or whose dynamic attributes have been updated, if the generation time of the target data entry corresponding to the target node is longer than the set time from the current moment, the popularity of the target data entry corresponding to the target node is determined based on the dynamic attributes of the target node and the correlation between the target node and other nodes in the knowledge graph.

[0101] In this embodiment, a node in the knowledge graph whose popularity has not yet been determined or whose dynamic attributes have been updated is called a target node.

[0102] For example, if the knowledge graph obtained in step S201 is a newly constructed knowledge graph, then the heat corresponding to each node in the knowledge graph (that is, the heat of the data entry corresponding to the node) is not determined. Therefore, it is necessary to calculate the heat of the data entry corresponding to each node. Therefore, each node in the knowledge graph will be used as a target node.

[0103] For example, after the popularity of the data entries corresponding to each node in the knowledge graph has been determined and stored in the corresponding storage area, if there are new data entries, then new nodes will inevitably appear in the knowledge graph. At this time, the node corresponding to the new data entry is the target node.

[0104] For example, after the popularity of the data entries corresponding to each node in the knowledge graph has been determined and stored in the corresponding storage area, if the access information of the data entries in the multiple data entries corresponding to the knowledge graph is updated, for example, the data entry is accessed or the time of the last access changes, etc., then the dynamic attributes of the node corresponding to the data entry in the knowledge graph will also be updated, and the popularity of the data entry may also change, so it is necessary to determine the popularity corresponding to the data entry as the target node to re-determine the popularity corresponding to the data entry.

[0105] Among them, the correlation degree can reflect the importance of the target node in the knowledge graph. The higher the importance of the target node, the higher the possibility that the target data entry will be accessed separately or simultaneously when accessing the data entries corresponding to other nodes.

[0106] For example, the relevance includes at least one of the degree and centrality of the target node.

[0107] The degree of the target node is the number of edges connected to the target node in the knowledge graph.

[0108] The centrality of a target node is used to characterize the relative importance or influence of the target node in the knowledge graph. It is not only related to the edges connecting the target nodes, but also needs to consider information such as the position of the target node in the knowledge graph and the path information flow. For example, the centrality of a node can include, but is not limited to, at least one of the node's closeness centrality, betweenness centrality, and eigenvector centrality, without specific restrictions. For example, the closeness centrality of a node can be the inverse of the average shortest path length from the node to other nodes in the knowledge graph.

[0109] Among them, based on the dynamic attributes of the target node and the correlation between the target node and other nodes in the knowledge graph, there are many possibilities for determining the heat of the target data entry corresponding to the target node, which are not specifically limited.

[0110] For example, the weighted values ​​of the access frequency, update frequency, and correlation of the target data entry represented by the dynamic attributes of the target node can be calculated, and the calculated weighted values ​​can be used as the heat of the target data entry. For example, the heat H of the target data entry can be expressed as: H = s1 * access frequency + s2 * update frequency + s3 * correlation, where s1, s2, and s3 are different weighted coefficients set respectively. Specifically, the weighted coefficients corresponding to the three parameters of access frequency, update frequency, and correlation can be set according to the application scenario or actual needs.

[0111] It can be understood that if the generation time of the target data entry corresponding to the target node is longer than the set time from the current moment, it means that the target data entry has been generated for a period of time, then the access information of the target data entry can reflect the access characteristics such as the access frequency and update frequency of the target data entry, that is, the dynamic attributes of the target data entry can accurately reflect the access situation of the target data entry. Therefore, the popularity of the target data entry corresponding to the target node can be determined based on the dynamic attributes of the target node and the correlation between the target data node and other nodes.

[0112] On the contrary, if the generation time of the target data entry corresponding to the target node is not more than the set time from the current moment, since the generation time of the target data entry is short, the possibility of the target data entry being accessed is low, and the access information of the target data entry may be empty or the amount of data is small, which cannot truly reflect the situation in which the target data entry may actually be accessed. Therefore, the subsequent step S203 can be executed to determine the heat of the target data entry in combination with the heat of the neighboring nodes associated with the node corresponding to the target data entry.

[0113] For example, when constructing a knowledge graph based on multiple data entries, given the low probability that each data entry is newly generated and the undetermined popularity of the nodes corresponding to each data entry, to simplify the calculation, it may be assumed that the time from the generation of each data entry to the current moment is greater than a set time. Of course, it is also possible to determine whether the time from the generation of the target entry to the current moment is greater than a set time based on the actual generation time of each data entry.

[0114] For another example, for a newly added data entry, the generation time of the new data entry is relatively short, and the access information of the new data entry may be empty. Therefore, it is not appropriate to determine its popularity based on the dynamic attributes and correlation of the node corresponding to the new data entry.

[0115] S203, for any target node in the knowledge graph whose popularity has not yet been determined or whose dynamic attributes have been updated, if the generation time of the target data entry corresponding to the target node is not greater than the set time from the current moment, the popularity of the target data entry corresponding to the target node is determined based on the popularity corresponding to at least one neighbor node of the target node in the knowledge graph and the common access time between the target node and the neighbor node.

[0116] The target node's neighbor nodes refer to nodes in the knowledge graph that are directly connected to the target node via edges. The heat corresponding to the neighbor node is the heat of the data entry corresponding to the neighbor node.

[0117] The common access time is the time when the target data entry corresponding to the target node and the data entry corresponding to the neighboring node are most recently commonly accessed.

[0118] It is understandable that due to the relationships between different data items, in many cases, accessing a data item may require the use of information from other data items, or querying a certain information may require querying two or more data items simultaneously. This may result in situations where two or more data items need to be accessed simultaneously, causing two or more data items to be accessed together. For example, if a user wants to determine the sales volume of a product over a period of time, they need to query the data item containing the product's attributes. Based on the product model and product number, they need to query all sales orders for the product during that period to ultimately determine the product's sales volume.

[0119] It can be understood that in the case that the generation time of the target data entry corresponding to the target node is more, the accessed information of the target data entry can not truly reflect the access frequency and other access characteristics of the target data entry being accessed, but through the heat of the neighbor nodes of the target node and the common access time of the target node and the neighbor nodes being commonly accessed, the frequency and other conditions of the target data entry corresponding to the target node being accessed in the future can be reflected, and naturally the heat of the target data entry can also be accurately reflected.

[0120] S204, determining a heat attribute category of the target data entry based on the heat of the target data entry.

[0121] The heat attribute category is used to indicate that the target data entry is hot data or cold data.

[0122] S205, storing the target data entry into a target storage area corresponding to the heat attribute category of the target data entry.

[0123] The above steps S204 and S205 can refer to the related description of the foregoing embodiments, and will not be repeated here.

[0124] In the embodiments, Figure 2 The specific implementation of determining the heat of the target data entry corresponding to the target node based on the heat of at least one neighbor node of the target node and the common access time between the target node and the neighbor node can also have many possibilities, and is not limited. The following will be described taking an implementation manner as an example. As shown in Figure 3 Fig. 1 shows an implementation flowchart of the present application for determining the heat of the target data entry corresponding to the target node based on the heat of at least one neighbor node of the target node and the common access time between the target node and the neighbor node. The implementation flowchart can include:

[0125] S301, if the time length of the generation time of the target data entry corresponding to the target node from the current time is not greater than the set time length, determining at least one neighbor node of the target node from the knowledge graph.

[0126] Among them, it can be determined from the knowledge graph that all neighbor nodes of the target node.

[0127] In an optional manner, considering that the number of neighbor nodes of the target node can be more, resulting in a large amount of calculation, based on this, in order to reduce the amount of calculation, the present application can also sample at least two neighbor nodes from the neighbor node set belonging to the target node in the knowledge graph.

[0128] The neighbor node set includes all neighbor nodes of the target node in the knowledge graph. In order to accurately and reasonably determine the popularity of the target data entry corresponding to the target node while reducing the amount of computation, the at least two sampled neighbor nodes may include at least one neighbor node whose corresponding data entry belongs to hot data and at least one neighbor node whose corresponding data entry belongs to cold data.

[0129] Of course, in practical applications, the proportion or number of the neighboring nodes whose corresponding data entries belong to hot data and the neighboring nodes whose corresponding data entries belong to cold data among the at least two sampled neighboring nodes can also be set, without specific restrictions.

[0130] S302 : For each neighbor node, determine a weight corresponding to the neighbor node based on the time length between the common access time between the target node and the neighbor node and the current time.

[0131] For example, considering that the closer the common access time of the neighbor node is to the current time, the more important its influence on determining the popularity of the target node is, therefore, the longer the time between the common access time between the neighbor node and the target node and the current time, the smaller the weighted weight corresponding to the neighbor node, that is, there is an inversely proportional relationship between the weighted weight corresponding to the neighbor node and the time between its corresponding common access time and the current time.

[0132] For example, in one possible implementation, the weighted weight w of the neighbor node v corresponding to the target node u is uv It can be calculated by the following formula:

[0133]

[0134] Among them, t current is the current moment, t uv is the common access time between the target node u and its corresponding neighbor node v. λ is the set decay coefficient, and λ = 0.1 means a 10% decay every day.

[0135] S303 : Determine the heat of the target data entry corresponding to the target node based on the heat and weighted weight corresponding to each neighboring node.

[0136] For example, the heat of each neighboring node is weightedly summed based on the weighted weight corresponding to each neighboring node to obtain the heat of the target data entry.

[0137] For example, in one possible implementation, a target number of layers of sub-knowledge graphs extending from the at least two sampled neighbor nodes of the target node can be first determined from the knowledge graph. Then, based on the heat and weighted weight corresponding to each neighbor node in the sub-knowledge graph, as well as the heat corresponding to each layer of sub-nodes of each neighbor node, the heat corresponding to each neighbor node in the sub-knowledge graph and the heat corresponding to each layer of sub-nodes of each neighbor node are weightedly aggregated in combination with a graph neural network algorithm to obtain the heat of the target data entry corresponding to the target node.

[0138] The sub-knowledge graph belongs to a knowledge graph with a target number of layers. The sub-knowledge graph can be considered as a partial knowledge graph with the target number of layers, which is obtained by extending the target node as the root node through the at least two sampled neighbor nodes to the sub-nodes of each layer of the sampled neighbor node in the knowledge graph. Therefore, in addition to the target node and the sampled neighbor nodes, the nodes in the sub-knowledge graph can also include other nodes directly or indirectly connected to the neighbor nodes.

[0139] The target number can be set as needed. For example, assuming the target number of layers is 2, then the sub-knowledge graph only includes the target node and the sampled neighbor nodes; if the target number of layers is 3, then the sub-knowledge graph includes not only the target node and the sampled neighbor nodes, but also the sub-nodes directly connected to each sampled neighbor node; if the target number of layers is 3, then the sub-knowledge graph includes not only the target node and the sampled neighbor nodes, but also the first-layer sub-nodes directly connected to each sampled neighbor node and the sub-nodes directly connected to the first-layer sub-nodes (i.e., the second-layer sub-nodes of the neighbor node).

[0140] In this possible implementation method, the present application uses a graph neural network algorithm to perform weighted aggregation on the heat corresponding to each neighbor node and each layer of child nodes of the neighbor node in the sub-knowledge graph. Since the weighted weight of each neighbor node in the weighted aggregation process is related to the common visit time of each neighbor node and the target node, the influence of the heat of the neighbor nodes that have been recently visited together with the target node on the corresponding heat of the target node can be strengthened in the weighted aggregation process, and the heat of the target data entry corresponding to the target node can be determined more reasonably.

[0141] In this application, there are many possibilities for graph neural network algorithms, without any specific restrictions.

[0142] As in a possible implementation, the graph neural network algorithm employed by the present application can be a GraphSAGE (Graph Sample and Aggregated) algorithm. GraphSAGE is a graph neural network model for graph node embedding learning, which aggregates the information of neighboring nodes to the target node through sampling and aggregation, thereby learning the representation vector of the node. However, unlike the traditional GraphSAGE algorithm which treats the vectors of all neighboring nodes equally in weighted aggregation, the present application assigns corresponding weighted weights to different neighboring nodes.

[0143] As an example, taking the target number of layers as K layers, the process of weighted aggregation of each neighboring node and each layer of sub-node of the neighboring node in the sub-knowledge graph using the GraphSAGE algorithm can be represented by the following Formula Two and Formula Three:

[0144]

[0145] wherein, is the heat aggregation result of the target node u at the kth layer of the sub-knowledge graph, N(u) represents the set of neighboring nodes sampled by the target node u, w uv is the weighted weight of the neighboring node v. It can be seen that the weighted weight of each neighboring node is considered in the heat aggregation process of the present application. Wherein, k is an integer from 2 to K. Wherein, in the sub-knowledge graph, the layer where the target node is located is the Kth layer, and the layer farthest from the target node is the 1st layer.

[0146] is the heat of the neighboring node v at the k-1th layer of the sub-knowledge graph. The heat of the neighboring node at the k-1th layer of the sub-knowledge graph is aggregated from the heat of each sub-node of the neighboring node at the 1st layer to the k-1th layer of the sub-knowledge graph. In particular, since the neighboring node is located at the 2nd layer of the sub-knowledge graph, the heat of the neighboring node at the 2nd layer is aggregated from the heat of the neighboring node and the heat of the neighboring node at the 3rd layer.

[0147] is the final heat of the target node u at the kth layer of the sub-knowledge graph. Correspondingly, the final heat of the target node u at the Kth layer of the sub-knowledge graph is the heat of the target data entry corresponding to the target node.

[0148] wherein, W (K) is the learned weight matrix. σ represents the activation function, and CONCAT() represents the concatenation function.

[0149] Of course, the above is an example of an implementation method of the GraphSAGE algorithm to perform weighted aggregation of the heat of each neighbor node and its child nodes. In actual applications, there may be other implementation possibilities, which are not specifically limited.

[0150] It is understandable that after storing each data entry in partitions according to the popularity attribute category corresponding to the data entry, some data entries that have not been accessed for a long time can be archived. Archiving data entries means transferring infrequently accessed data entries to independent storage devices for long-term storage.

[0151] In this application, in order to reduce the number of data archiving times and improve the efficiency of data archiving, this application can also determine other data entries that are related to the data entry to be archived and are also suitable for archiving based on the association relationship of the data entries, and archive them together. Figure 4 For explanation. Figure 4 Another flow chart of the data processing method provided by the present application is shown. The method of this embodiment may include:

[0152] S401, obtaining a knowledge graph corresponding to multiple data entries.

[0153] The knowledge graph includes multiple nodes, each of which represents an entity identified in a data entry. The attributes of each node include at least dynamic attributes, and the dynamic attributes of the node are used to represent the accessed information of the data entry corresponding to the node.

[0154] S402, for a target node in the knowledge graph, determine the popularity of a target data entry corresponding to the target node based on the dynamic attributes of the target node and the association relationship between the target node and other nodes in the knowledge graph.

[0155] S403 : Determine the popularity attribute category of the target data entry based on the popularity of the target data entry.

[0156] The heat attribute category is used to indicate whether the target data entry is hot data or cold data.

[0157] S404: Store the target data entry into a target storage area corresponding to the heat attribute category of the target data entry.

[0158] For the above steps S401 to S404 , reference can be made to the relevant introduction of the previous embodiment and will not be repeated here.

[0159] S405: If there is at least one to-be-archived node in the knowledge graph whose corresponding data entry meets the archiving conditions, determine the edge weights of the edges between the nodes in the knowledge graph based on the association relationship between the nodes in the knowledge graph.

[0160] The archiving conditions can be set according to actual needs. The fact that a data entry meets the archiving conditions indicates that the data entry has not been accessed for a set period of time or has a low probability of being accessed. The determination of the presence of a node to be archived that meets the archiving conditions in the knowledge graph can be determined by active detection by the electronic device; or it can be determined based on an archiving instruction initiated by the user for at least one node to be archived that meets the archiving conditions, where the archiving instruction indicates that there is at least one node to be archived that meets the archiving conditions.

[0161] For the sake of easy distinction, the nodes corresponding to the data entries that meet the archiving conditions are called to-be-archived nodes.

[0162] Among them, there are many possible ways to determine the edge weight of an edge in a knowledge graph. The current method for determining the edge weight of an edge between two nodes in a network structure can be adopted, without specific restrictions. For example, the edge weight of an edge between different nodes can be determined by combining the access link of the data entry corresponding to the node in the knowledge graph and the association of the node in the knowledge graph. For example, to access node C, it is necessary to visit node A and node B in sequence before accessing node C. Then node B is more important, and the edge between node A and node B has a more important influence on accessing node C. Therefore, the weight of the edge between node A and node B will be greater than the weight between node B and node C.

[0163] S406: Divide the knowledge graph into at least one graph community based on the popularity of the data entries corresponding to each node in the knowledge graph and the edge weights of the edges between different nodes.

[0164] Each graph community includes at least one node. A graph community is part of a knowledge graph and is a node group consisting of at least one node and at least one edge between the nodes.

[0165] In this application, the heat of the data entries corresponding to each node in the knowledge graph and the edge weight of each edge can classify nodes with the same heat attribute category to which the corresponding heat belongs and are closely connected into the same graph community.

[0166] For example, this application can use an unsupervised community discovery algorithm to divide the knowledge graph into at least one graph community. However, unlike the traditional unsupervised community discovery algorithm that only considers the edge weights of the edges between nodes, this application has made some improvements to the unsupervised community discovery algorithm. In the process of using the unsupervised community discovery algorithm to discover graph communities in the knowledge graph, the heat of the data entries corresponding to each node in the knowledge graph and the edge weights of the edges between nodes are combined to determine the closeness between each node in each graph community, and finally find at least one graph community with the highest closeness.

[0167] S407: For each node to be archived, determine the target graph community where the node to be archived is located, determine at least one data entry to be archived from the data entries corresponding to each node in the target graph community, and archive the at least one data entry to be archived.

[0168] Among them, the data entries to be archived in the target graph community are also data entries that meet the archiving conditions. There is no restriction on the specific method of determining the data entries to be archived. For example, if the time from the last access time of the data entry corresponding to the node in the target graph community to the current moment exceeds the set time, the data entry is determined to be a data entry to be archived.

[0169] It is understandable that, since the present application will determine the graph community suitable for each node in combination with the heat of the data entry corresponding to the node during the graph community discovery process, the heat corresponding to each node in the same graph community should be the same, and the possibility of nodes corresponding to cold data and hot data appearing in the same graph community is relatively low. On this basis, the possibility of other nodes in the same target graph community as the node to be archived being suitable for archiving is also relatively high. Therefore, the present application can determine at least one data entry to be archived from the target graph community where the node to be archived is located, and archive at least one data entry together, thereby realizing batch archiving of multiple data entries, reducing the number of frequent data archiving executions, and improving data archiving efficiency.

[0170] For example, when a user initiates data archiving, not only the data item to be archived designated by the user can be archived, but also other data items related to the item to be archived and that meet the archiving criteria. Alternatively, when an electronic device initiates an archiving operation for a portion of data items, it can simultaneously detect all related data items that meet the archiving criteria and archive them together.

[0171] Archiving the data entry may include storing the data entry in a fixed storage device outside the storage area corresponding to the hot data and the cold data, and there is no specific limitation.

[0172] exist Figure 4 In the embodiment of the present invention, there are many possibilities for the specific implementation of dividing the knowledge graph into at least one graph community. The following is an example of an implementation method. Figure 5 , shows a schematic diagram of an implementation process of dividing a knowledge graph into at least one graph community based on the popularity of data entries corresponding to each node in the knowledge graph and the edge weights between different nodes in this application. This process may include:

[0173] S501, determining the initial community to which each node in the knowledge graph belongs, and taking the initial community to which the node belongs as the first community in which the node currently resides.

[0174] Each node can only belong to one initial community, and the initial community to which the node belongs can be set randomly. For example, each node can be considered as an initial community.

[0175] The purpose of using the initial community to which the node belongs as the current first community of the node is to subsequently obtain the second community to which the node is most suitable by continuously updating the first community to which the node belongs.

[0176] S502 : For each node, determine the node connection closeness corresponding to the first community where the node is located based on the popularity of data entries corresponding to each node in the first community where the node is located and the edge weights of the edges between the nodes.

[0177] The node connection closeness corresponding to the first community can reflect the similarity of the popularity between the nodes in the first community and the closeness of the connection between the nodes.

[0178] In the first possible case, the node connection closeness corresponding to the first community can be calculated according to the closeness calculation formula between heat and edge weight.

[0179] In the second possible scenario, the calculation formula for calculating the connection density of a community in the unsupervised community discovery algorithm can be improved by replacing the edge weights of the two nodes involved in the calculation formula for calculating the connection density of a community in the community discovery algorithm with the weighted sum of the edge weight of the edge between the two nodes and the heat corresponding to the two nodes, and then calculating the node connection density corresponding to the first community based on the improved formula.

[0180] For example, in the second possible scenario, specifically, for a node pair consisting of any two nodes in the first community where the node is located, the weight of the node pair is determined based on the edge weight of the edge between the two nodes in the node pair and the heat of the data entries corresponding to the two nodes in the node pair. For example, the weight of the node pair is the weighted sum of the edge weight associated with the node pair and the sum of the heats corresponding to the two nodes in the node pair. On this basis, the node connection closeness of the first community where the node is located can be determined based on the weight of each node pair in the first community where the node is located and the heat of the data entry corresponding to each node in each node pair.

[0181] The following illustrates the determination of the graph community of each node by taking the Louvain algorithm (also known as the Fast unfolding algorithm) as an example. The Louvain algorithm is a greedy optimization algorithm for discovering community structure in networks, which divides the communities in the network by maximizing the modularity. The present application can improve the formula for calculating the tightness of the community in the Louvain algorithm, so that the improved calculation formula can combine the heat of the data entries and the edge weight to determine the node contact tightness of the community. For example, based on the Louvain algorithm, the present application can calculate the node contact tightness Q of the first community by the following formula four:

[0182]

[0183] Wherein m is the sum of the edge weights of each edge in the first community. The value of is 0 or 1, if node i and node j belong to the same community before calculating the contact tightness of the first community this time, then The value of is 1, otherwise 0.

[0184] W ij is the weight of the node pair composed of node i and node j, which can be calculated by the following formula five:

[0185]

[0186] Wherein A ij is the edge weight of the edge between node i and node j, w i is the heat corresponding to node i (i.e. the heat of the data entry corresponding to node i), w j is the heat corresponding to node j. Node i and node j can be any two different nodes in the first community.

[0187] S503, optimizing the first community of each node to obtain the second community of each node.

[0188] Wherein the node contact tightness of the second community of the node is higher than the node contact tightness of the first community of the node.

[0189] Among them, optimizing the first community where the node is located can be to move the node from the current first community to other communities, and recalculate the node connection density of the newly joined community of the node. If the node connection density of the newly joined node is lower than the node connection density corresponding to the first community where the node was originally located, the community joined by the node will be replaced, and the operation of recalculating the node connection density of the newly joined community of the node will be returned until the node connection density corresponding to the newly joined community of the node exceeds the node connection density of the first community where the node was originally located, and the newly joined community of the node will be used as the second community where the node is located.

[0190] It is understandable that after determining the second community where the node is located, the second community where the node is located can be used as the current first community, and step S503 is repeated until the second communities where all nodes are located remain unchanged. In this case, the optimization of the first community where each node is located is completed. At this time, the second community where the node is located is the graph community to which the node ultimately belongs.

[0191] For example, the Louvain algorithm may be used to optimize the first community in which each node is located, so as to ultimately determine a suitable second community for each node. Details will not be repeated here.

[0192] S504: For each node, determine the second community where the node is located as the graph community to which the node belongs, and obtain at least one graph community.

[0193] certainly, Figure 5 This is just an example of an implementation method. If the graph community to which each node belongs is determined by other methods, it is also applicable to this embodiment without any specific limitation.

[0194] It is understandable that since the probability of archived data entries being accessed is relatively low, in order to reduce the amount of data required to process subsequent knowledge graph updates and the next data entry archiving, this application also compresses and encodes the attribute data of the nodes in the knowledge graph corresponding to the archived data entries to reduce the amount of data in the knowledge graph.

[0195] Corresponding to a data processing method provided in this application, this application also provides a data processing device. Figure 6 A schematic diagram of the structure of a data processing device provided by the present application is shown, and the device includes:

[0196] A graph obtaining unit 601 is configured to obtain a knowledge graph corresponding to a plurality of data entries, wherein the knowledge graph includes a plurality of nodes, each node representing an entity determined in a data entry, and the attributes of the node include at least dynamic attributes, and the dynamic attributes of the node are used to represent accessed information of the data entry corresponding to the node;

[0197] A popularity determination unit 602 is configured to determine, for a target node in the knowledge graph, the popularity of a target data entry corresponding to the target node based on the dynamic attributes of the target node and the association relationship between the target node and other nodes in the knowledge graph;

[0198] A category determination unit 603 is configured to determine a heat attribute category of the target data entry based on the heat of the target data entry, wherein the heat attribute category is used to indicate whether the target data entry is hot data or cold data;

[0199] The data storage unit 604 is configured to store the target data entry into a target storage area corresponding to the heat attribute category of the target data entry.

[0200] In a possible implementation, the atlas obtaining unit includes at least one of the following:

[0201] A graph construction subunit, configured to construct a knowledge graph corresponding to the plurality of data entries to be stored based on the plurality of data entries;

[0202] A node adding subunit is used to respond to the new data entry currently obtained and add the node corresponding to the new data entry in the constructed knowledge graph to obtain an updated knowledge graph corresponding to the multiple data entries;

[0203] The graph updating subunit is used to update the dynamic attributes of the node corresponding to the data entry in the knowledge graph in response to the update of the accessed information of the data entry corresponding to the node in the knowledge graph, so as to obtain an updated knowledge graph.

[0204] In another possible implementation, the heat determination unit includes:

[0205] A first popularity determination subunit is configured to determine, for a target node in the knowledge graph whose popularity has not yet been determined or whose dynamic attributes have been updated, the popularity of the target data entry corresponding to the target node, based on the dynamic attributes of the target node and the association between the target node and other nodes in the knowledge graph, if the time between the generation time of the target data entry corresponding to the target node and the current moment is greater than a set time, wherein the association includes at least one of the degree and centrality of the target node;

[0206] The second heat determination sub-unit is used to determine the heat of the target data entry corresponding to the target node based on the heat corresponding to at least one neighbor node of the target node in the knowledge graph and the common access time between the target node and the neighbor node, if the generation time of the target data entry corresponding to the target node is not greater than the set time from the current time, the heat corresponding to the neighbor node is the heat of the data entry corresponding to the neighbor node, and the common access time is the time when the target data entry corresponding to the target node and the data entry corresponding to the neighbor node were last jointly accessed.

[0207] In yet another possible implementation, the second heat determination subunit includes:

[0208] a neighbor determination subunit, configured to determine at least one neighbor node of the target node from the knowledge graph if the time between the generation time of the target data entry corresponding to the target node and the current moment is not greater than a set time;

[0209] a weighted determination subunit, configured to determine, for each neighbor node, a weighted weight corresponding to the neighbor node based on a time interval between a common access time of the target node and the neighbor node and a current time;

[0210] The heat determination subunit is used to determine the heat of the target data entry corresponding to the target node based on the heat and weighted weight corresponding to each of the neighboring nodes.

[0211] In yet another possible implementation, the neighbor determination subunit includes:

[0212] a neighbor sampling subunit, configured to sample at least two neighbor nodes from a set of neighbor nodes belonging to the target node in the knowledge graph, wherein the at least two neighbor nodes include at least one neighbor node whose corresponding data entry belongs to hot data and at least one neighbor node whose corresponding data entry belongs to cold data;

[0213] The heat determination subunit includes:

[0214] a subgraph determination subunit, configured to determine, from the knowledge graph, a target number of sub-knowledge graphs extending from the target node through the at least two neighboring nodes;

[0215] The heat aggregation sub-unit is used to perform weighted aggregation on the heat corresponding to each neighbor node in the sub-knowledge graph and the heat corresponding to each layer of child nodes of the neighbor node based on the heat and weighted weight corresponding to each neighbor node in the sub-knowledge graph, as well as the heat corresponding to each layer of child nodes of the neighbor node, in combination with the graph neural network algorithm, to obtain the heat of the target data entry corresponding to the target node.

[0216] In yet another possible implementation, the data storage unit includes:

[0217] a first storage subunit, configured to store the target data entry into a target storage area corresponding to a heat attribute category of the target data entry if the target data entry is a data entry that has not been stored;

[0218] The second storage subunit is used to migrate the target data entry from the current storage area to a target storage area that matches the heat attribute category of the target data entry if the target data entry belongs to a stored data entry and the current storage area where the target data entry is currently located does not match the heat attribute category of the target data entry.

[0219] In another possible implementation, Figure 7 FIG. 6 is a diagram showing another structural diagram of a data processing device provided by the present application. In addition to a graph obtaining unit 601, a popularity determining unit 602, a category determining unit 603, and a data storage unit 604, the device further includes:

[0220] An edge weight determination unit 605 is configured to determine edge weights of edges between nodes in the knowledge graph based on associations between nodes in the knowledge graph if there is at least one to-be-archived node in the knowledge graph whose corresponding data entry meets an archiving condition.

[0221] A community determination unit 606 is configured to divide the knowledge graph into at least one graph community based on the popularity of data entries corresponding to each node in the knowledge graph and the edge weights of edges between different nodes, each graph community including at least one node;

[0222] The archiving processing unit 607 is used to determine the target graph community where the node to be archived is located, determine at least one data entry to be archived from the data entries corresponding to each node in the target graph community, and archive the at least one data entry to be archived.

[0223] In a possible implementation, the community determination unit includes:

[0224] An initial determination subunit, configured to determine an initial community to which each node in the knowledge graph belongs, with the initial community to which the node belongs being the first community to which the node currently resides;

[0225] a closeness determination subunit, configured to determine, for each node, a node connection closeness corresponding to the first community where the node is located based on the popularity of data entries corresponding to each node in the first community where the node is located and the edge weights of the edges between the nodes;

[0226] a community optimization subunit, configured to optimize the first community where each node is located to obtain a second community where each node is located, wherein the node connection density corresponding to the second community where the node is located is higher than the node connection density corresponding to the first community where the node is located;

[0227] The community determination subunit is configured to determine the second community where the node is located as the graph community to which the node belongs, thereby obtaining at least one graph community.

[0228] In yet another possible implementation, the compactness determination subunit includes:

[0229] a node pair processing subunit, configured to determine, for a node pair consisting of any two nodes in the first community where the node is located, a weight of the node pair based on an edge weight of an edge between the two nodes in the node pair and the popularity of data entries corresponding to the two nodes in the node pair;

[0230] The density determination subunit is configured to determine the node connection density of the first community where the node is located based on the weights of each node pair in the first community where the node is located and the popularity of the data entry corresponding to each node in each node pair.

[0231] An electronic device is also provided in an embodiment of the present application. Figure 8 , which shows a schematic diagram of the composition structure of the electronic device, the electronic device at least includes a processor 801 and a memory 802;

[0232] The processor 801 is configured to execute the data processing method described in any one of the above embodiments;

[0233] The memory 802 is used to store programs required by the processor to perform operations.

[0234] It is understandable that the electronic device may further include a display unit 803 and an input unit 804 .

[0235] Of course, the electronic device may also have Figure 8 There is no limitation to more or fewer components.

[0236] A computer program product is also provided in an embodiment of the present application, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any data processing method provided in the embodiment of the present application.

[0237] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any data processing method provided in the embodiment of the present application.

[0238] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0239] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.

[0240] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0241] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, training device or data center to another website, computer, training device or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as a training device, a data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

Claims

1. A data processing method, comprising: Obtaining a knowledge graph corresponding to the plurality of data entries, the knowledge graph comprising a plurality of nodes, each node representing an entity determined in a data entry, the attributes of the node comprising at least a dynamic attribute, the dynamic attribute of the node being used to characterize accessed information of the data entry corresponding to the node; For a target node in the knowledge graph, determining the popularity of a target data entry corresponding to the target node based on the dynamic attributes of the target node and the association relationship between the target node and other nodes in the knowledge graph; determining a heat attribute category of the target data entry based on the heat of the target data entry, the heat attribute category being used to indicate whether the target data entry is hot data or cold data; The target data entry is stored in a target storage area corresponding to the heat attribute category of the target data entry.

2. The data processing method according to claim 1, wherein obtaining the knowledge graph corresponding to the plurality of data entries comprises at least one of the following: Based on the multiple data entries to be stored, construct a knowledge graph corresponding to the multiple data entries; In response to the currently obtained new data entry, adding a node corresponding to the new data entry to the constructed knowledge graph to obtain an updated knowledge graph corresponding to the multiple data entries; In response to an update in accessed information of a data entry corresponding to a node in a knowledge graph, the dynamic attributes of the node corresponding to the data entry in the knowledge graph are updated to obtain an updated knowledge graph.

3. The data processing method according to claim 1 or 2, wherein for a target node in the knowledge graph, determining the popularity of a target data entry corresponding to the target node based on the dynamic attributes of the target node and the association relationship between the target node and other nodes in the knowledge graph comprises: For a target node in the knowledge graph whose popularity has not yet been determined or whose dynamic attributes have been updated, if the time between the generation time of the target data entry corresponding to the target node and the current moment is greater than the set time, the popularity of the target data entry corresponding to the target node is determined based on the dynamic attributes of the target node and the association between the target node and other nodes in the knowledge graph, where the association includes at least one of the degree and centrality of the target node; If the generation time of the target data entry corresponding to the target node is not longer than the set time from the current moment, the heat of the target data entry corresponding to the target node is determined based on the heat corresponding to at least one neighbor node of the target node in the knowledge graph and the common access time between the target node and the neighbor node. The heat corresponding to the neighbor node is the heat of the data entry corresponding to the neighbor node, and the common access time is the time when the target data entry corresponding to the target node and the data entry corresponding to the neighbor node were last jointly accessed.

4. The data processing method according to claim 3, wherein determining the popularity of a target data entry corresponding to the target node based on the popularity corresponding to at least one neighbor node of the target node in the knowledge graph and the common access time between the target node and the neighbor node comprises: Determine at least one neighbor node of the target node from the knowledge graph; For each neighbor node, determine a weight corresponding to the neighbor node based on the time between the common access time of the target node and the neighbor node and the current time; Based on the heat and weighted weight corresponding to each of the neighbor nodes, the heat of the target data entry corresponding to the target node is determined.

5. The data processing method according to claim 4, wherein determining at least one neighbor node of the target node from the knowledge graph comprises: Sampling at least two neighbor nodes from a set of neighbor nodes belonging to the target node in the knowledge graph, wherein the at least two neighbor nodes include at least one neighbor node whose corresponding data entry belongs to hot data and at least one neighbor node whose corresponding data entry belongs to cold data; The determining the heat of the target data entry corresponding to the target node based on the heat and weighted weight corresponding to each of the neighboring nodes includes: Determine, from the knowledge graph, a target number of layers of sub-knowledge graphs extending from the target node through the at least two neighboring nodes; Based on the heat and weighted weight corresponding to each neighbor node in the sub-knowledge graph, as well as the heat corresponding to each layer of child nodes of the neighbor node, the heat corresponding to each neighbor node in the sub-knowledge graph and the heat corresponding to each layer of child nodes of the neighbor node are weightedly aggregated in combination with the graph neural network algorithm to obtain the heat of the target data entry corresponding to the target node.

6. The data processing method according to claim 1, further comprising: If there is at least one to-be-archived node in the knowledge graph whose corresponding data entry meets the archiving condition, determining the edge weights of the edges between the nodes in the knowledge graph based on the association relationships between the nodes in the knowledge graph; Divide the knowledge graph into at least one graph community based on the popularity of data entries corresponding to each node in the knowledge graph and the edge weights of edges between different nodes, each graph community including at least one node; Determine a target graph community where the node to be archived is located, determine at least one data entry to be archived from data entries corresponding to each node in the target graph community, and archive the at least one data entry to be archived.

7. The data processing method according to claim 6, wherein the dividing the knowledge graph into at least one graph community based on the popularity of the data entries corresponding to each node in the knowledge graph and the edge weights of the edges between different nodes comprises: Determine the initial community to which each node in the knowledge graph belongs, and take the initial community to which the node belongs as the first community to which the node currently belongs; For each node, determining the node connection closeness corresponding to the first community where the node is located based on the popularity of data entries corresponding to each node in the first community where the node is located and the edge weights of the edges between the nodes; Optimizing the first community where each node is located to obtain the second community where each node is located, wherein the node connection closeness corresponding to the second community where the node is located is higher than the node connection closeness corresponding to the first community where the node is located; The second community where the node is located is determined as the graph community to which the node belongs, to obtain at least one graph community.

8. The data processing method according to claim 7, wherein determining the node connection closeness corresponding to the first community where the node is located based on the popularity of data entries corresponding to each node in the first community where the node is located and the edge weights of the edges between the nodes comprises: For a node pair consisting of any two nodes in the first community where the node is located, determining a weight of the node pair based on an edge weight of an edge between the two nodes in the node pair and the popularity of data entries corresponding to the two nodes in the node pair; The node connection closeness of the first community where the node is located is determined based on the weight of each node pair in the first community where the node is located and the popularity of the data entry corresponding to each node in each node pair.

9. The data processing method according to claim 1, wherein storing the target data entry into a target storage area corresponding to the heat attribute category of the target data entry comprises: If the target data entry is a data entry that has not been stored, storing the target data entry in a target storage area corresponding to the heat attribute category of the target data entry; If the target data entry belongs to a stored data entry and the current storage area where the target data entry is currently located does not match the heat attribute category of the target data entry, the target data entry is migrated from the current storage area to a target storage area that matches the heat attribute category of the target data entry.

10. A data processing device comprising: a graph obtaining unit, configured to obtain a knowledge graph corresponding to a plurality of data entries, wherein the knowledge graph includes a plurality of nodes, each node representing an entity determined in a data entry, and the attributes of the node include at least a dynamic attribute, wherein the dynamic attribute of the node is used to represent accessed information of the data entry corresponding to the node; a popularity determination unit, configured to determine, for a target node in the knowledge graph, the popularity of a target data entry corresponding to the target node based on the dynamic attributes of the target node and the association relationship between the target node and other nodes in the knowledge graph; a category determining unit, configured to determine a heat attribute category of the target data entry based on the heat of the target data entry, wherein the heat attribute category is used to indicate whether the target data entry is hot data or cold data; The data storage unit is configured to store the target data entry into a target storage area corresponding to the heat attribute category of the target data entry.