Data processing method, device, computer readable medium and electronic device
By mining and pruning the relationship graph network core and compressing the device cluster when preset conditions are met, the problems of large computing resources consumption and low data processing efficiency in big data analysis are solved, and more efficient data processing is achieved.
Patent Information
- Application Number
- CN202011626906.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-31
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-12-31
AI Technical Summary
The problems of high computing resources consumption and low data processing efficiency in big data analysis.
By obtaining a graph network that represents the interaction relationship between multiple interactive objects, the device cluster is used to mine core degree, iteratively update the node core degree, and prune the graph network according to the core degree, and compress the device cluster when the network scale meets the preset conditions.
Reduces the consumption of computing resources, improves data processing efficiency, and saves additional time overhead caused by parallel computing.
Smart Images

Figure CN113515672B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to a data processing method, a data processing device, a computer-readable medium, and an electronic device. Background Art
[0002] With the development of computer and network technology, users can establish various interactive relationships with each other on the basis of providing business services on the network platform. For example, users can establish social relationships with other users on the network social platform, and can also establish financial transactions with other users on the network payment platform. On this basis, the network platform will also accumulate a large amount of user data, which includes data related to the user's own attributes when using the network platform, and also includes interactive data generated by establishing interactive relationships between different users.
[0003] Reasonable sorting and mining of user data can enable the network platform to better provide users with convenient and efficient platform services based on user characteristics. However, with the continuous accumulation of user data, the increasingly large data scale will bring increasing data processing pressure, and the network platform will also need to spend more and more computing resources and time for data analysis. Therefore, how to improve the efficiency of big data analysis and reduce related costs is an urgent problem to be solved. Summary of the invention
[0004] The purpose of this application is to provide a data processing method, a data processing device, a computer-readable medium and an electronic device, which can at least to some extent overcome the technical problems existing in big data analysis, such as large consumption of computing resources and low data processing efficiency.
[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.
[0006] According to one aspect of an embodiment of the present application, a data processing method is provided, the method comprising: obtaining a relationship graph network for representing an interaction relationship between multiple interaction objects, the relationship graph network comprising nodes for representing interaction objects and edges for representing interaction relationships; performing coreness mining on the relationship graph network through a device cluster comprising multiple computing devices to iteratively update the node coreness of all or part of the nodes in the relationship graph network; performing pruning processing on the relationship graph network according to the node coreness to remove some nodes and edges in the relationship graph network; when the network scale of the relationship graph network meets a preset network compression condition, compressing the device cluster to remove some computing devices in the device cluster.
[0007] According to one aspect of an embodiment of the present application, a data processing device is provided, which includes: a graph network acquisition module, configured to acquire a relationship graph network for representing interaction relationships between multiple interaction objects, the relationship graph network including nodes for representing interaction objects and edges for representing interaction relationships; a coreness mining module, configured to perform coreness mining on the relationship graph network through a device cluster including multiple computing devices, so as to iteratively update the node coreness of all or part of the nodes in the relationship graph network; a network pruning module, configured to perform pruning processing on the relationship graph network according to the node coreness, so as to remove some nodes and edges in the relationship graph network; a cluster compression module, configured to compress the device cluster to remove some computing devices in the device cluster when the network scale of the relationship graph network meets a preset network compression condition.
[0008] In some embodiments of the present application, based on the above technical solution, the cluster compression module includes: a stand-alone computing unit, configured to select a computing device in the device cluster as a target node for performing stand-alone computing on the relationship graph model, and remove other computing devices except the target node from the device cluster.
[0009] In some embodiments of the present application, based on the above technical solution, the coreness mining module includes: a network segmentation unit, configured to perform segmentation processing on the relationship graph network to obtain a partitioned graph network composed of some nodes and edges in the relationship graph network; a network allocation unit, configured to allocate the partitioned graph network to a device cluster including multiple computing devices to determine the computing device used to perform coreness mining on the partitioned graph network; a partition mining unit, configured to perform coreness mining on the partitioned graph network to iteratively update the node coreness of each node in the relationship graph network.
[0010] In some embodiments of the present application, based on the above technical solution, the partition mining unit includes: a node selection subunit, configured to select a computing node for core degree mining in the current iteration round in the partition graph network, and determine a neighbor node having an adjacency relationship with the computing node; a core degree acquisition subunit, configured to obtain the current node core degrees of the computing node and the neighbor node in the current iteration round; a core degree calculation subunit, configured to determine the temporary node core degree of the computing node based on the current node core degree of the neighbor node, and mark the computing node whose temporary node core degree is less than the current node core degree as an active node; a core degree update subunit, configured to update the current node core degree of the active node based on the temporary node core degree, and determine the active node and the neighbor node having an adjacency relationship with the active node as computing nodes for core degree mining in the next iteration round.
[0011] In some embodiments of the present application, based on the above technical solution, the coreness calculation subunit includes: an h-index calculation subunit, configured to determine the h-index of the computing node based on the current node coreness of the neighboring node, and use the h-index as the temporary node coreness of the computing node, wherein the h-index indicates that the current node coreness of at most h neighboring nodes among all neighboring nodes of the computing node is greater than or equal to h.
[0012] In some embodiments of the present application, based on the above technical solution, the h-index calculation subunit includes: a node sorting subunit, configured to sort all neighbor nodes of the computing node in order from high to low according to the coreness of the current node, and assign an arrangement number starting with 0 to each of the neighbor nodes; a node screening subunit, configured to compare the arrangement number of each neighbor node and the current node coreness, respectively, to screen neighbor nodes whose arrangement numbers are greater than or equal to the current node coreness according to the comparison results; an h-index determination subunit, configured to determine the current node coreness of the neighbor node with the smallest arrangement number among the screened neighbor nodes as the h-index of the computing node.
[0013] In some embodiments of the present application, based on the above technical solution, the node selection subunit includes: an identifier reading subunit, configured to read the node identifier of the node to be updated from the first storage space, the node to be updated includes the active node that updated the node coreness in the previous iteration round and the neighbor node having an adjacency relationship with the active node; an identifier selection subunit, configured to select a computing node for coreness mining in the current iteration round in the partitioned graph network according to the node identifier of the node to be updated.
[0014] In some embodiments of the present application, based on the above technical solution, the device also includes: a coreness writing module, configured to write the updated current node coreness of the active node into a second storage space, and the second storage space is used to store the node coreness of all nodes in the relationship graph network; an identifier writing module, configured to obtain the node identifier of the active node and the neighboring nodes of the active node, and write the node identifier into a third storage space, and the third storage space is used to store the computing nodes for coreness mining in the next iteration round; a space covering module, configured to overwrite the data in the first storage space with the third storage space and reset the third storage space after completing the mining of the cores of all partitioned graph networks in the current iteration round.
[0015] In some embodiments of the present application, based on the above technical solution, the coreness acquisition subunit includes: a coreness reading subunit, configured to read the current node coreness of the computing node and the neighbor node in the current iteration round from a second storage space, and the second storage space is used to store the node coreness of all nodes in the relationship graph network.
[0016] In some embodiments of the present application, based on the above technical solution, the network pruning module includes: a minimum coreness acquisition unit, configured to obtain the minimum coreness of active nodes in the current iteration round and the minimum coreness of active nodes in the previous iteration round; a convergence node screening unit, configured to screen the convergence nodes in the relationship graph network according to the minimum coreness of the active nodes in the previous iteration round if the minimum coreness of the active nodes in the current iteration round is greater than the minimum coreness of the active nodes in the previous iteration round, the convergence nodes are nodes whose node coreness is less than or equal to the minimum coreness of the active nodes in the previous iteration round; a convergence node removal unit, configured to remove the convergence nodes and the edges connected to the convergence nodes from the relationship graph network.
[0017] In some embodiments of the present application, based on the above technical solution, the network compression condition includes that the number of edges in the relationship graph network is less than a preset number threshold.
[0018] According to one aspect of an embodiment of the present application, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the data processing method in the above technical solution is implemented.
[0019] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the data processing method in the above technical solution by executing the executable instructions.
[0020] According to one aspect of the embodiments of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the data processing method in the above technical solution.
[0021] In the technical solution provided in the embodiment of the present application, a relationship graph network is established based on the business data involving the interactive relationship between interactive objects. By utilizing the structural characteristics and sparsity of the relationship graph network, distributed computing can be first performed through the device cluster to perform core degree mining in different regions. With the continuous iterative update of the node core degree, the relationship graph network is pruned to "prune" the nodes and corresponding edges that have been iteratively converged, so that the relationship graph network is continuously compressed and reduced with the iterative update of the node core degree, thereby reducing the consumption of computing resources. On this basis, when the relationship graph network is compressed to an appropriate size, the device cluster can be further compressed, which can not only release a large amount of computing resources, but also save additional time overhead such as data distribution caused by parallel computing, thereby improving data processing efficiency.
[0022] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 The following is a block diagram of the architecture of a data processing system using the technical solution of the present application.
[0025] Figure 2 A flow chart showing the steps of a data processing method in one embodiment of the present application is shown.
[0026] Figure 3 A flowchart of the method steps for coreness mining based on distributed computing in one embodiment of the present application is shown.
[0027] Figure 4 A flowchart of the steps of performing coreness mining on a partition graph network in one embodiment of the present application is shown.
[0028] Figure 5 A flowchart of the steps of selecting a computing node in one embodiment of the present application is shown.
[0029] Figure 6 A flowchart of the steps of determining the h-index of a computing node in one embodiment of the present application is shown.
[0030] Figure 7 A flowchart of the steps of summarizing the node coreness mining results of a partitioned graph network in one embodiment of the present application is shown.
[0031] Figure 8 A schematic diagram of the process of compressing and pruning a relationship graph network based on iterative updates of node coreness in one embodiment of the present application is shown.
[0032] Fig. 9 The overall architecture and processing flow chart of k-core mining in an application scenario in an embodiment of the present application are shown.
[0033] Fig.10 The structural block diagram of the data processing device provided in the embodiment of the present application is schematically shown.
[0034] Fig.11 The structure block diagram of a computer system suitable for implementing an electronic device of an embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0035] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more comprehensive and complete and fully convey the concept of the example embodiments to those skilled in the art.
[0036] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0037] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0038] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0039] Figure 1 The following is a block diagram of the architecture of a data processing system using the technical solution of the present application.
[0040] like Figure 1 As shown, data processing system 100 may include terminal device 110 , network 120 , and server 130 .
[0041] The terminal device 110 may include various electronic devices such as smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart watches, smart glasses, and vehicle-mounted terminals. The terminal device 110 may be installed with clients of various application programs such as video application clients, music application clients, social application clients, and payment application clients, so that users can obtain corresponding application services based on the client of the application program.
[0042] The server 130 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The network 120 may be a communication medium of various connection types that can provide a communication link between the terminal device 110 and the server 130, such as a wired communication link or a wireless communication link.
[0043] According to the implementation requirements, the system architecture in the embodiment of the present application can have any number of terminal devices, networks and servers. For example, the server 130 can be a server group composed of multiple server devices. In addition, the technical solution provided in the embodiment of the present application can be applied to the terminal device 110, can also be applied to the server 130, or can be implemented by the terminal device 110 and the server 130 together, and the present application does not make any special restrictions on this.
[0044] For example, when a user uses a social application on the terminal device 110, he can send messages to other platform users on the network social platform or conduct network social behaviors such as voice conversations and video conversations. Based on this process, social relationships can be established between different platform users, and corresponding social business data will be generated on the network social platform. For another example, when a user uses a payment application on the terminal device 110, he can make payments or receive payments from other platform users on the network payment platform. Based on this process, payment relationships can be established between different platform users, and corresponding payment business data will be generated on the network payment platform.
[0045] After collecting relevant user data such as social business data or payment business data, the embodiment of the present application can construct a graph network model based on the interactive relationship in the user data, and perform data mining on the graph network model to obtain the business attributes of the user in the interactive relationship. Taking the payment application scenario as an example, in a graph network with merchants and consumers, the node represents the merchant or consumer, and the edge represents the payment relationship between the two nodes. Generally speaking, the merchant node is more in the center of the network. Therefore, the core degree (core value) of the node can be used as a topological feature and input into the downstream machine learning task, so as to mine the business model and identify whether the node in the graph network model is a merchant or a consumer. In addition, in the risk control scenario of the payment business, it is possible to detect whether a node (or edge) has abnormal behavior based on the data mining of the graph network model, so as to perform detection tasks of abnormal behaviors such as illegal credit intermediation, cashing out, multiple loans, gambling, etc.
[0046] In order to improve the efficiency of big data analysis and mining, the embodiments of the present application can use cloud technology for distributed computing.
[0047] Cloud technology refers to a hosting technology that unifies hardware, software, network and other resources within a wide area network or local area network to achieve data computing, storage, processing and sharing. Cloud technology involves network technology, information technology, integration technology, management platform technology, application technology, etc. applied in the cloud computing business model, which can form a resource pool and be used on demand, flexibly and conveniently. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the rapid development and application of the Internet industry, each item may have its own identification mark in the future, and all need to be transmitted to the background system for logical processing. Data of different levels will be processed separately, and all kinds of industry data require strong system backing support, which can only be achieved through cloud computing.
[0048] Cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computers, allowing various application systems to obtain computing power, storage space and information services as needed. The network that provides resources is called a "cloud". From the user's perspective, the resources in the "cloud" are infinitely scalable and can be obtained at any time, used on demand, expanded at any time, and paid for by use.
[0049] As a provider of basic cloud computing capabilities, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service) platform) will be established, and various types of virtual resources will be deployed in the resource pool for external customers to choose to use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.
[0050] According to the logical function division, the PaaS (Platform as a Service) layer can be deployed on the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on the PaaS layer. SaaS can also be deployed directly on IaaS. PaaS is a platform for software operation, such as databases, web containers, etc. SaaS is a variety of business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.
[0051] Big data refers to a collection of data that cannot be captured, managed, and processed by conventional software tools within a certain time frame. It is a massive, high-growth, and diverse information asset that requires new processing models to have stronger decision-making power, insight discovery, and process optimization capabilities. With the advent of the cloud era, big data has also attracted more and more attention. Big data requires special technologies to effectively process large amounts of data within a tolerable time frame. Technologies applicable to big data include large-scale parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems.
[0052] Artificial intelligence cloud service is generally referred to as AIaaS (AI as a Service). This is the service mode of a mainstream artificial intelligence platform. Specifically, the AIaaS platform will split several common AI services and provide independent or packaged services in the cloud. This service model is similar to opening an AI theme mall: all developers can access and use one or more artificial intelligence services provided by the platform through API interfaces. Some senior developers can also use the AI framework and AI infrastructure provided by the platform to deploy and operate their own cloud artificial intelligence services.
[0053] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0054] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0055] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, drones, robots, smart medical care, smart customer service, etc. I believe that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0056] The following describes in detail the technical solutions such as the data processing method, data processing device, computer-readable medium, and electronic device provided in the embodiments of the present application in combination with specific implementation methods.
[0057] Figure 2 A flowchart of a data processing method in one embodiment of the present application is shown. The data processing method can be Figure 1 The terminal device 110 shown in FIG. Figure 1 The method may be executed on the server 130 shown in the figure, or may be executed jointly by the terminal device 110 and the server 130. Figure 2 As shown, the data processing method may mainly include the following steps S210 to S240.
[0058] Step S210: Obtain a relationship graph network for representing the interaction relationship between multiple interaction objects, where the relationship graph network includes nodes for representing the interaction objects and edges for representing the interaction relationship.
[0059] Step S220: performing coreness mining on the relationship graph network through a device cluster including multiple computing devices to iteratively update the node coreness of all or part of the nodes in the relationship graph network.
[0060] Step S230: Pruning the relationship graph network according to the node coreness to remove some nodes and edges in the relationship graph network.
[0061] Step S240: When the network scale of the relationship graph network meets the preset network compression condition, the device cluster is compressed to remove some computing devices in the device cluster.
[0062] In the data processing method provided in the embodiment of the present application, a relationship graph network is established based on the business data involving the interactive relationship between interactive objects. By utilizing the structural characteristics and sparsity of the relationship graph network, distributed computing can be first performed through the device cluster to perform core degree mining in different regions. With the continuous iterative update of the node core degree, the relationship graph network is pruned to "prune" the nodes and corresponding edges that have been iteratively converged, so that the relationship graph network is continuously compressed and reduced with the iterative update of the node core degree, thereby reducing the consumption of computing resources. On this basis, when the relationship graph network is compressed to a suitable size, the device cluster can be further compressed, which can not only release a large amount of computing resources, but also save additional time overhead such as data distribution caused by parallel computing, thereby improving data processing efficiency.
[0063] The following is a detailed description of each method step of the data processing method in the above embodiment.
[0064] In step S210, a relationship graph network for representing the interaction relationship between multiple interaction objects is obtained, where the relationship graph network includes nodes for representing the interaction objects and edges for representing the interaction relationship.
[0065] The interactive objects may be various user objects that conduct business interactions on the business platform. For example, in an online payment scenario involving commodity transactions, the interactive objects may include consumers who initiate online payments and merchants who receive payments. The interactive relationship between the interactive objects may be an online payment relationship established between consumers and merchants based on payment events.
[0066] In an embodiment of the present application, by collecting business data generated by business transactions between multiple interactive objects, multiple interactive objects and the interactive relationships between the interactive objects can be extracted therefrom, thereby establishing a relationship graph network composed of nodes (Node) and edges (Edge), wherein each node can represent an interactive object, and the edge connecting two interactive objects represents that an interactive relationship has been established between the two interactive objects.
[0067] In step S220, coreness mining is performed on the relationship graph network through a device cluster including multiple computing devices to iteratively update the node coreness of all or part of the nodes in the relationship graph network.
[0068] Node coreness is a parameter used to measure the importance of each node in a graph network. In an embodiment of the present application, the node coreness of a node can be represented by the number of cores (coreness) of each node determined when performing k-core decomposition on the graph network. The k-core of a graph refers to the remaining subgraph after repeatedly removing nodes whose degree is less than or equal to k. Among them, the degree of a node is equal to the number of neighbor nodes that have a direct adjacency relationship with the node. In general, the degree of a node can also reflect the importance of a node in a local area of a graph network to a certain extent, and by mining the number of cores of a node, the importance of a node can be better measured globally.
[0069] If a node exists in k-core and is removed from (k+1)-core, then the number of cores of this node is k. k-core mining is an algorithm to calculate the number of cores of all nodes in a graph. For example, the original graph network is a graph with 0 cores; 1 core is a graph with all isolated points removed; 2 cores means first removing all nodes with degree less than 2, and then removing points with degree less than 2 from the remaining graph, and so on, until it cannot be removed; 3 cores means first removing all points with degree less than 3, and then removing points with degree less than 3 from the remaining graph, and so on, until it cannot be removed... The number of cores of a node is defined as the order of the largest core in which the node is located. For example, if a node is at most in 5 cores but not in 6 cores, then the number of cores of this node is 5.
[0070] Figure 3 A flowchart of the method steps for core degree mining based on distributed computing in one embodiment of the present application is shown. Figure 3 As shown, based on the above embodiments, step S220 of performing coreness mining on the relationship graph network through a device cluster including multiple computing devices to iteratively update the node coreness of all or part of the nodes in the relationship graph network may include the following steps S310 to S330.
[0071] Step S310: segment the relationship graph network to obtain a partitioned graph network consisting of some nodes and edges in the relationship graph network.
[0072] A relationship graph network with a larger network scale can be segmented to obtain multiple relatively smaller partition graph networks. In one embodiment of the present application, the method for segmenting a relationship graph network may include: first, selecting multiple segmentation center points in the relationship graph network according to a preset number of segmentations, and then using the segmentation center points as clustering centers, clustering all nodes in the relationship graph network to assign each node to a segmentation center point that is closest to it, and finally segmenting the relationship graph network into multiple partition graph networks according to the clustering results of the nodes. The segmentation center points can be nodes selected in the relationship graph network according to preset rules or randomly selected nodes.
[0073] In one embodiment of the present application, a certain overlapping area can be retained between two adjacent partitioned graph networks, and a portion of nodes and edges can be shared in the overlapping area, thereby generating a certain computational redundancy and improving the reliability of coreness mining for each partitioned graph network.
[0074] Step S320: Allocate the partitioned graph network to a device cluster including a plurality of computing devices to determine a computing device for performing coreness mining on the partitioned graph network.
[0075] By assigning multiple partition graph networks to different computing devices, distributed computing of coreness mining can be realized through a device cluster composed of computing devices, thereby improving data processing efficiency.
[0076] In one embodiment of the present application, when the relationship graph network is segmented, the same number of partition graph networks can be obtained according to the number of computing devices available in the device cluster. For example, if the device cluster performing distributed computing includes M computing devices, the relationship graph network can be segmented into M partition graph networks.
[0077] In another embodiment of the present application, the relationship graph network can also be divided into several partition graph networks of similar size according to the computing power of a single computing device, and then each partition graph network is assigned to the same number of computing devices. For example, the relationship graph network includes N nodes, and the relationship graph network can be divided into N / T partition graph networks, where T is the number of nodes of a single partition graph network determined according to the computing power of a single computing device. When the relationship graph network is large in scale and the number of partition graph networks is large, the number of nodes contained in each partition graph network is basically equal to the number of nodes. After the relationship graph network is divided, N / T computing devices are selected from the device cluster, and a partition graph network is assigned to each computing device respectively. When the number of devices in the device cluster is less than N / T, multiple partition graph networks can be assigned to some or all of the computing devices according to the computing power and working status of the computing devices.
[0078] Step S330: Perform coreness mining on the partition graph network to iteratively update the node coreness of each node in the relationship graph network.
[0079] In one embodiment of the present application, the node coreness of each node in the relationship graph network may be initialized and assigned a value according to a preset rule, and then the node coreness of each node may be iteratively updated in each iteration round.
[0080] In some optional implementations, the node coreness can be initialized according to the node degree. Specifically, in the relationship graph network, the number of neighbor nodes with adjacency to each node is obtained respectively, and then the node coreness of each node is initialized according to the number of neighbor nodes. The degree of a node represents the number of connections between a node and its neighbor nodes. In some other implementations, the weight information can also be determined in combination with the node's own attributes, and then the node coreness is initialized and assigned according to the node degree and weight information.
[0081] Figure 4 FIG. 1 is a flowchart showing the steps of performing core degree mining on a partition graph network in one embodiment of the present application. Figure 4 As shown, based on the above embodiments, the coreness mining of the partition graph network in step S330 to iteratively update the node coreness of each node in the relationship graph network may include the following steps S410 to S440.
[0082] Step S410: Select a computing node for core degree mining in the current iteration round in the partition graph network, and determine a neighboring node having an adjacency relationship with the computing node.
[0083] In the first iteration round after the node coreness is initialized, all nodes in the partition graph network can be determined as computing nodes. The computing nodes are the nodes that need to perform coreness mining calculations in the current iteration round. Based on the mining results, it can be determined whether the node coreness of each node needs to be updated.
[0084] In each iteration of coreness mining, the computing nodes that need to be mined in the current iteration can be determined based on the coreness mining results of the previous iteration and the update results of the node coreness. Some or all of these computing nodes will update the node coreness in the current iteration. Other nodes except computing nodes will not be mined in the current iteration, and naturally will not update the node coreness.
[0085] The neighbor nodes in the embodiment of the present application refer to other nodes that have a direct connection relationship with a node. Since the node coreness of each node will be affected by its neighbor nodes, as the iteration continues, the nodes whose node coreness is not updated in the current iteration round may also be selected as calculation nodes in the subsequent iteration process.
[0086] Figure 5 A flowchart of the steps of selecting a computing node in one embodiment of the present application is shown. Figure 5 As shown, the step S410 of selecting computing nodes for core degree mining in the partition graph network in the current iteration round may include the following steps S510 to S520.
[0087] Step S510: Reading node identifiers of nodes to be updated from the first storage space, where the nodes to be updated include active nodes whose node coreness is updated in the previous iteration round and neighbor nodes having an adjacency relationship with the active nodes.
[0088] Since the partition graph networks that make up the relationship graph network are processed in a distributed manner on different computing devices, and in the edge areas of two adjacent partition graph networks, two adjacent nodes will be divided into different partition graph networks, but the node coreness of the two will still affect each other. Therefore, in order to maintain the synchronization and consistency of the node coreness update in the partition graph network while performing distributed computing, the embodiment of the present application allocates a first storage space in the system to store the node identifiers of all nodes to be updated in the relationship graph network.
[0089] In an iteration, when a node in a partitioned graph network updates its node coreness according to the coreness mining result, the node can be marked as an active node. The active node and its neighboring nodes are nodes to be updated, and the node identifiers of the nodes to be updated can be written into the first storage space.
[0090] Step S520: Selecting a computing node for core degree mining in the current iteration round in the partition graph network according to the node identifier of the node to be updated.
[0091] When an iteration round begins, each computing device may read the node identifier of the node to be updated from the first storage space, thereby selecting a computing node for core degree mining in the current iteration round in the partitioned graph network to which the computing device is assigned.
[0092] By executing steps S510 to S520 as described above, the node identifiers of all nodes to be updated in the relationship graph network can be summarized through the first storage space after each iteration round, and when a new iteration round begins, data can be distributed to different computing devices so that each computing device selects a computing node in the partitioned graph network it maintains.
[0093] Step S420: Obtain the current node coreness of the computing node and its neighboring nodes in the current iteration round.
[0094] The embodiment of the present application can monitor and update the node coreness of the node in real time according to the coreness mining result in each iteration round. The current node coreness of each node in the current iteration round is the latest node coreness determined after the previous iteration round.
[0095] In an optional implementation, the embodiment of the present application may allocate a second storage space in the system to store the node coreness of all nodes in the relationship graph network. When a computing device needs to perform coreness mining and updating based on existing coreness data, the current node coreness of the computing node and neighboring nodes in the current iteration round can be read from the second storage space.
[0096] Step S430: Determine the temporary node coreness of the computing node according to the current node coreness of the neighboring node, and mark the computing node whose temporary node coreness is less than the current node coreness as an active node.
[0097] Taking coreness as an example, in the related technology of this application, a recursive pruning method can be used to mine the coreness of a graph network based on the definition of k-core. Specifically, starting from k=1, nodes with degrees less than or equal to k and their connecting edges can be continuously removed from the graph until the degrees of all remaining nodes in the graph are greater than k. Recursive pruning is similar to "peeling an onion", and the core value of all nodes peeled off in the kth round is k. However, since this method calculates the number of cores by gradually shrinking the graph network as a whole from the outside to the inside, this method can only use centralized computing to process the entire graph network data in serial, and it is difficult to apply distributed parallel processing. When faced with ultra-large-scale (tens of billions / hundreds of billions) relationship chain networks, there are problems such as long computing time and poor computing performance.
[0098] In order to overcome this problem, in one embodiment of the present application, an iterative method based on the h-indicator can be used to perform coreness mining. Specifically, the embodiment of the present application can determine the h-index of the computing node based on the current node coreness of the neighboring nodes, and use the h-index as the temporary node coreness of the computing node. The h-index indicates that the current node coreness of at most h neighboring nodes among all neighboring nodes of the computing node is greater than or equal to h.
[0099] For example, a computing node has five neighbor nodes, and the current node coreness of these five neighbor nodes are 2, 3, 4, 5, and 6. According to the order of node coreness from small to large, among the five neighbor nodes of the computing node, 5 neighbor nodes have current node coreness greater than or equal to 1, 5 neighbor nodes have current node coreness greater than or equal to 2, 4 neighbor nodes have current node coreness greater than or equal to 3, 3 neighbor nodes have current node coreness greater than or equal to 4, 2 neighbor nodes have current node coreness greater than or equal to 5, and 1 neighbor node has current node coreness greater than or equal to 6. It can be seen that among all the neighbor nodes of the computing node, at most 3 neighbor nodes have current node coreness greater than or equal to 3, so the h index of the computing node is 3, and it can be determined that the temporary node coreness of the computing node is 3.
[0100] Figure 6 FIG. 1 is a flowchart showing the steps of determining the h-index of a computing node in one embodiment of the present application. Figure 6 As shown, based on the above embodiments, the method for determining the h-index of a computing node according to the current node coreness of neighboring nodes may include the following steps S610 to S630.
[0101] Step S610: Sort all neighbor nodes of the computing node in descending order of the coreness of the current node, and assign a sequence number starting with 0 to each neighbor node.
[0102] Step S620: respectively compare the ranking sequence number of each neighbor node with the coreness of the current node, and select neighbor nodes whose ranking sequence number is greater than or equal to the coreness of the current node according to the comparison result.
[0103] Step S630: among the neighbor nodes screened out, the current node coreness of the neighbor node with the smallest ranking number is determined as the h-index of the calculation node.
[0104] The embodiment of the present application can quickly and efficiently determine the h-index of the computing node by sorting and screening, which is particularly suitable for situations where the number of computing nodes is large.
[0105] Step S440: updating the current node coreness of the active node according to the temporary node coreness, and determining the active node and the neighboring nodes having an adjacency relationship with the active node as computing nodes for coreness mining in the next iteration round.
[0106] After obtaining the temporary node coreness of the active node, the temporary node coreness can be compared with the current node coreness of the active node. If the temporary node coreness is smaller than the current node coreness, the current node coreness can be replaced with the temporary node coreness. If the two are the same, it means that the computing node does not need to be updated in the current iteration round.
[0107] In one embodiment of the present application, after updating the current node coreness of the active node based on the temporary node coreness, the overall update results in the relationship graph network can be summarized based on the update results of the node coreness in each partition graph network, thereby providing a coreness mining basis for the next iteration round.
[0108] Figure 7 A flowchart of the steps for summarizing the node coreness mining results of a partitioned graph network in one embodiment of the present application is shown. Figure 7 As shown, based on the above embodiments, the method for summarizing the node coreness mining results of each partition graph network may include the following steps S710 to S730.
[0109] Step S710: writing the updated current node coreness of the active node into the second storage space, where the second storage space is used to store the node coreness of all nodes in the relationship graph network.
[0110] Step S720: Obtain node identifiers of the active node and the neighboring nodes of the active node, and write the node identifiers into a third storage space, where the third storage space is used to store computing nodes for core degree mining in the next iteration round.
[0111] Step S730: After completing the mining of the cores of all partition graph networks in the current iteration round, the third storage space overwrites the data in the first storage space and resets the third storage space.
[0112] The embodiment of the present application configures a third storage space, and implements aggregation and distribution of node coreness mining results of a partitioned graph network based on updates and resets of the third storage space in each iteration round, thereby ensuring stability and reliability of data processing while utilizing distributed computing to improve data processing efficiency.
[0113] In step S230, the relationship graph network is pruned according to the node coreness to remove some nodes and edges in the relationship graph network.
[0114] With the mining and iterative updating of node coreness, the nodes and edges in the relationship graph network will gradually reach a converged and stable state, and the node coreness will not be updated in the subsequent iteration process, nor will it affect the coreness mining results of other nodes. For these converged nodes, they can be pruned and removed to reduce the data size of the relationship graph network and the partition graph network.
[0115] In an optional embodiment, the embodiment of the present application can obtain the minimum coreness of the active nodes in the current iteration round and the minimum coreness of the active nodes in the previous iteration round; if the minimum coreness of the active nodes in the current iteration round is greater than the minimum coreness of the active nodes in the previous iteration round, then the convergence nodes in the relationship graph network are filtered according to the minimum coreness of the active nodes in the previous iteration round, and the convergence nodes are nodes whose node coreness is less than or equal to the minimum coreness of the active nodes in the previous iteration round; the convergence nodes and the edges connected to the convergence nodes are removed from the relationship graph network.
[0116] Figure 8 A schematic diagram of the process of compressing and pruning a relationship graph network based on iterative updates of node coreness in one embodiment of the present application is shown.
[0117] The key to the compression pruning method is to analyze the changes in the core value of the node in each iteration. Indicates the core value of node v in the tth iteration, minCore (t) Indicates the minimum core value of the nodes whose core values are updated in the tth iteration.
[0118]
[0119] When a node's core value is updated, its updated core value must be smaller than the original core value. According to the rule that the node's core value decreases with each iteration, when minCore (t) >minCore (t-1) When , it indicates that all core values are less than or equal to minCore (t-1) The nodes have converged and will not be updated in the future. According to the k-core mining feature, nodes with smaller core values do not affect the iteration of nodes with larger core values. Therefore, the nodes that have converged in each round of iteration and their corresponding edges can be "cut off", so that the relationship graph network is gradually compressed and reduced as the iteration proceeds.
[0120] like Figure 8 As shown, according to the initialized core value, the initial minimum core value can be determined as minCore (0)= 1. After the first round of iterations, the core values of some nodes are updated. Among the nodes with updated core values, the minimum core value is minCore (1) = 1. After the second round of iteration, the core values of some nodes are updated. Among the nodes with updated core values, the minimum core value is minCore (2) =1.
[0121] Due to minCore (2) >minCore (1) , which can trigger the pruning of the relationship graph network and remove the nodes with core value 1, thereby achieving the purpose of compressing the relationship graph network.
[0122] In step S240, when the network scale of the relationship graph network meets the preset network compression condition, the device cluster is compressed to remove some computing devices in the device cluster.
[0123] As the scale of the relationship graph network continues to shrink, the computing resources required for node coreness mining are gradually reduced. At this time, some computing resources can be released as the iteration proceeds to reduce resource overhead.
[0124] In one embodiment of the present application, the network compression condition may include that the number of edges in the relationship graph network is less than a preset number threshold. The device cluster may be compressed by re-segmenting the relationship graph network according to the network size of the relationship graph network after pruning to obtain a reduced number of partition graph networks, and calling a relatively small number of computing devices based on the reduced number of partition graph networks.
[0125] In one embodiment of the present application, when the network scale of the compressed relationship graph network meets a certain condition, a computing device can be selected from the device cluster as the target node for single-machine computing of the relationship graph model, and other computing devices except the target node can be removed from the device cluster. In this way, the distributed computing mode based on multiple computing devices can be transformed into a centralized computing mode based on a single computing device.
[0126] It should be noted that although the steps of the method in the embodiment of the present application are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.
[0127] Based on the introduction of the data processing method in the above embodiments, it can be known that the data processing method provided in the embodiments of the present application relates to a method for k-core mining based on the idea of compression pruning. In some optional implementations, the method can be based on the iterative update of the h-index, and the graph network compression pruning can be automatically performed when the specified conditions are met. The method flow of the data processing method provided in the embodiments of the present application in an application scenario may include the following steps.
[0128] (1) For each node v in the graph network G(V,E), use the node degree to initialize its core value. where deg(v) represents the node degree, that is, the number of neighbors of the node. Initialize minCore with the minimum node degree, that is,
[0129] (2) Set the numMsgs parameter to represent the number of nodes whose core values change in each iteration, and initialize numMsgs with zero.
[0130] (3) For each node in G(V,E), the h-index (i.e., h-index value) is calculated based on the core values of its neighboring nodes, which is used as the core value of this iteration. Here N(v) represents the set of neighbor nodes of node v. When the core value of the node is updated, numMsgs is increased by 1, and the minimum core value minCore of the node updated in this round is calculated. (t) ,
[0131] (4) Determine whether numMsgs is 0. When numMsgs is 0, it indicates that the core values of all nodes are no longer updated and the iteration stops; otherwise, execute step (5).
[0132] (5) Determine minCore (t) >minCore (t-1) Is it true? If so, execute the compression pruning strategy: save the core value less than or equal to minCore (t-1) Nodes and corresponding core values are removed from the G(V,E) iteration graph to obtain the compressed subgraph G′(V,E). Continue to iterate steps 3-5 on G′(V,E); if minCore is not satisfied (t) >minCore (t-1) , continue to iterate steps 3-5 on the original image.
[0133] For k-core mining of large-scale graph networks, the above iterative steps are first carried out in a distributed parallel computing manner. When the scale of the compressed subgraph G′(V,E) meets the given conditions (for example, the number of edges is less than 30 million), the distributed computing can be converted to a single-machine computing mode. The single-machine computing mode can not only release a large amount of computing resources, but also save the additional time overhead such as data distribution brought by parallel computing. Especially for graph networks with long chain structures, the later stages of iteration usually focus on the update of long chain nodes, and it is more appropriate to use the single-machine computing mode at this time.
[0134] The k-core mining algorithm in the embodiment of the present application can implement distributed computing on the Spark on Angel platform. Among them, Spark is a fast and general computing engine designed for large-scale data processing, and Angel is a high-performance distributed machine learning platform designed and developed based on the parameter server (PS) concept. The Spark on Angel platform is a high-performance distributed computing platform that combines Angel's powerful parameter server function with Spark's large-scale data processing capabilities, supporting traditional machine learning, deep learning, and various graph algorithms.
[0135] Fig. 9 The overall architecture and processing flow chart of k-core mining in an application scenario of the present application embodiment are shown. Fig. 9 As shown in the figure, under the drive of Spark Driver, each Executor is responsible for storing the adjacency table partition data (that is, the network data of the partitioned graph network Graph Partion), calculating the h-index value and performing compression and pruning operations, and the AngelParameter Server is responsible for storing and updating the node core value, that is, Fig. 9 In order to utilize the sparsity of k-core mining to speed up iterative convergence, the nodes that need to be calculated for this round of iteration and the next round of iteration are stored on PS at the same time, which are Fig. 9 The ReadMessage vector and WriteMessage vector in . The nodes that are updated in this round of iteration are called active nodes. According to the property that the core value of a node is determined by its neighboring nodes, the change of the core value of an active node will affect the core value of its neighboring nodes. Therefore, its neighboring nodes should be calculated in the next round of iteration. Therefore, WriteMessage stores the neighboring nodes of the active node in this round of iteration in real time.
[0136] The Executor and PS process data in the following interactive manner in each iteration.
[0137] (1) Initialize minCore on Executor (t) =minCore (t-1) At the same time, two vector spaces, changedCore and keys2calc, are opened for this round of iteration, which are used to store the updated nodes in this round of iteration and the nodes that need to be calculated in the next round of iteration respectively.
[0138] (2) Pull the nodes that need to be calculated in this round of iteration (hereinafter referred to as calculation nodes) from the ReadMessage of the PS. If it is the first iteration, all nodes are pulled.
[0139] (3) Determine all nodes involved in the calculation in this round of iteration (the calculation node and its corresponding neighbors) from the calculation nodes obtained in step 2, and pull the corresponding core values from the coreness of the PS.
[0140] (4) For each node v in the computational nodes, calculate the h-index value of the core value of its neighboring node as the new core value of the node if Will Write to changedCore and set the core value of node v to be greater than minCore (t-1) The neighbor nodes are written into keys2calc to determine minCore (t) ,
[0141] (5) Use changedCore to update the coreness vector on PS, and use keys2calc to update the WriteMessage vector on PS.
[0142] Finally, when all partition data have completed a round of iteration, on the PS, ReadMessage is replaced with WriteMessage, and WriteMessage is reset to prepare for the next round of PS reading and writing. (t) After that, determine minCore (t) >minCore (t-1) If true, the above compression and pruning method is performed on all data partitions.
[0143] The k-core mining method based on compression concept provided in the embodiment of the present application can solve the problems of high resource overhead and long time consumption caused by k-core mining in ultra-large-scale networks. According to the iterative characteristics of k-core mining, a real-time compression method is designed, which can release some computing resources as the iteration proceeds; the k-core mining performance is improved by combining the advantages of distributed parallel computing and single-machine computing; the k-core mining method based on compression concept is implemented on the Spark on Angel high-performance graph computing platform, which can support ultra-large-scale networks with tens of billions / hundreds of billions of edges, with low resource overhead and high performance.
[0144] The following introduces an apparatus embodiment of the present application, which can be used to execute the data processing method in the above-mentioned embodiment of the present application. Fig.10 The structure block diagram of the data processing device provided in the embodiment of the present application is schematically shown. Fig.10 As shown, the data processing device 1000 may mainly include: a graph network acquisition module 1010, configured to acquire a relationship graph network for representing the interaction relationship between multiple interaction objects, wherein the relationship graph network includes nodes for representing the interaction objects and edges for representing the interaction relationship; a coreness mining module 1020, configured to perform coreness mining on the relationship graph network through a device cluster including multiple computing devices to iteratively update the node coreness of all or part of the nodes in the relationship graph network; a network pruning module 1030, configured to perform pruning processing on the relationship graph network according to the node coreness to remove some nodes and edges in the relationship graph network; a cluster compression module 1040, configured to compress the device cluster to remove some computing devices in the device cluster when the network scale of the relationship graph network meets a preset network compression condition.
[0145] In some embodiments of the present application, based on the above embodiments, the cluster compression module 1040 includes: a stand-alone computing unit, configured to select a computing device in the device cluster as a target node for performing stand-alone computing on the relationship graph model, and remove other computing devices except the target node from the device cluster.
[0146] In some embodiments of the present application, based on the above embodiments, the coreness mining module 1020 includes: a network segmentation unit, configured to perform segmentation processing on the relationship graph network to obtain a partitioned graph network composed of some nodes and edges in the relationship graph network; a network allocation unit, configured to allocate the partitioned graph network to a device cluster including multiple computing devices to determine the computing device used to perform coreness mining on the partitioned graph network; a partition mining unit, configured to perform coreness mining on the partitioned graph network to iteratively update the node coreness of each node in the relationship graph network.
[0147] In some embodiments of the present application, based on the above embodiments, the partition mining unit includes: a node selection subunit, configured to select a computing node for core degree mining in the current iteration round in the partition graph network, and determine a neighbor node having an adjacency relationship with the computing node; a core degree acquisition subunit, configured to obtain the current node core degrees of the computing node and the neighbor node in the current iteration round; a core degree calculation subunit, configured to determine the temporary node core degree of the computing node based on the current node core degree of the neighbor node, and mark the computing node whose temporary node core degree is less than the current node core degree as an active node; a core degree update subunit, configured to update the current node core degree of the active node based on the temporary node core degree, and determine the active node and the neighbor node having an adjacency relationship with the active node as computing nodes for core degree mining in the next iteration round.
[0148] In some embodiments of the present application, based on the above embodiments, the coreness calculation subunit includes: an h-index calculation subunit, configured to determine the h-index of the computing node based on the current node coreness of the neighboring nodes, and use the h-index as the temporary node coreness of the computing node, wherein the h-index indicates that the current node coreness of at most h neighboring nodes among all neighboring nodes of the computing node is greater than or equal to h.
[0149] In some embodiments of the present application, based on the above embodiments, the h-index calculation subunit includes: a node sorting subunit, configured to sort all neighbor nodes of the computing node in order from high to low according to the coreness of the current node, and assign an arrangement number starting with 0 to each of the neighbor nodes; a node screening subunit, configured to compare the arrangement number of each neighbor node and the current node coreness, respectively, to screen neighbor nodes whose arrangement numbers are greater than or equal to the current node coreness according to the comparison results; an h-index determination subunit, configured to determine the current node coreness of the neighbor node with the smallest arrangement number among the screened neighbor nodes as the h-index of the computing node.
[0150] In some embodiments of the present application, based on the above embodiments, the node selection subunit includes: an identifier reading subunit, configured to read the node identifier of the node to be updated from the first storage space, the node to be updated includes the active node whose node coreness was updated in the previous iteration round and the neighbor node having an adjacency relationship with the active node; an identifier selection subunit, configured to select a computing node for coreness mining in the current iteration round in the partitioned graph network according to the node identifier of the node to be updated.
[0151] In some embodiments of the present application, based on the above embodiments, the data processing device also includes: a coreness writing module, configured to write the updated current node coreness of the active node into a second storage space, and the second storage space is used to store the node coreness of all nodes in the relationship graph network; an identifier writing module, configured to obtain the node identifier of the active node and the neighboring nodes of the active node, and write the node identifier into a third storage space, and the third storage space is used to store the computing nodes for coreness mining in the next iteration round; a space covering module, configured to overwrite the data in the first storage space with the third storage space and reset the third storage space after completing the mining of the cores of all partitioned graph networks in the current iteration round.
[0152] In some embodiments of the present application, based on the above embodiments, the coreness acquisition subunit includes: a coreness reading subunit, configured to read the current node coreness of the computing node and the neighboring node in the current iteration round from a second storage space, and the second storage space is used to store the node coreness of all nodes in the relationship graph network.
[0153] In some embodiments of the present application, based on the above embodiments, the network pruning module 1030 includes: a minimum coreness acquisition unit, configured to acquire the minimum coreness of active nodes in the current iteration round and the minimum coreness of active nodes in the previous iteration round; a convergence node screening unit, configured to screen convergence nodes in the relationship graph network according to the minimum coreness of active nodes in the previous iteration round if the minimum coreness of active nodes in the current iteration round is greater than the minimum coreness of active nodes in the previous iteration round, wherein the convergence node is a node whose node coreness is less than or equal to the minimum coreness of active nodes in the previous iteration round; and a convergence node removal unit, configured to remove the convergence node and the edges connected to the convergence node from the relationship graph network.
[0154] In some embodiments of the present application, based on the above embodiments, the network compression condition includes that the number of edges in the relationship graph network is less than a preset number threshold.
[0155] The specific details of the data processing device provided in each embodiment of the present application have been described in detail in the corresponding method embodiments and will not be repeated here.
[0156] Fig.11 The structure block diagram of a computer system for implementing an electronic device according to an embodiment of the present application is schematically shown.
[0157] It should be noted that Fig.11The computer system 1100 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0158] like Fig.11 As shown, the computer system 1100 includes a central processing unit 1101 (CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1102 (ROM) or the program loaded from the storage part 1108 to the random access memory 1103 (RAM). Various programs and data required for system operation are also stored in the random access memory 1103. The central processing unit 1101, the read-only memory 1102 and the random access memory 1103 are connected to each other through a bus 1104. An input / output interface 1105 (Input / Output interface, i.e., I / O interface) is also connected to the bus 1104.
[0159] The following components are connected to the input / output interface 1105: an input section 1106 including a keyboard, a mouse, etc.; an output section 1107 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1108 including a hard disk, etc.; and a communication section 1109 including a network interface card such as a LAN card, a modem, etc. The communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the input / output interface 1105 as needed. A removable medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1110 as needed so that a computer program read therefrom is installed into the storage section 1108 as needed.
[0160] In particular, according to an embodiment of the present application, the process described in each method flow chart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, and the computer program contains a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1109, and / or installed from the removable medium 1111. When the computer program is executed by the central processor 1101, various functions defined in the system of the present application are executed.
[0161] It should be noted that the computer-readable medium shown in the embodiment of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by an instruction execution system, device or device or used in combination with it. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0162] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the above-mentioned module, program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0163] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.
[0164] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the implementation methods of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the implementation methods of the present application.
[0165] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary technical means in the art that are not disclosed in the present application.
[0166] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A data processing method, characterized in that: The method comprises: Acquire a relationship graph network for representing interaction relationships between a plurality of interaction objects, wherein the relationship graph network includes nodes for representing interaction objects and edges for representing interaction relationships; Performing coreness mining on the relationship graph network through a device cluster including a plurality of computing devices to iteratively update the node coreness of all or part of the nodes in the relationship graph network; Pruning the convergent nodes in the relationship graph network according to the node coreness to remove some nodes and edges in the relationship graph network; the convergent nodes are nodes whose node coreness is no longer updated during the iteration process; When the network scale of the relationship graph network meets a preset network compression condition, the device cluster is compressed to remove some computing devices in the device cluster.
2. The data processing method according to claim 1, characterized in that: The compressing the device cluster to remove some computing devices in the device cluster includes: A computing device is selected from the device cluster as a target node for performing single-machine computing on the relationship graph network, and other computing devices except the target node are removed from the device cluster.
3. The data processing method according to claim 1, characterized in that: The performing coreness mining on the relationship graph network by using a device cluster including a plurality of computing devices to iteratively update the node coreness of each node in the relationship graph network includes: Segmenting the relationship graph network to obtain a partition graph network consisting of some nodes and edges in the relationship graph network; Distributing the partitioned graph network to a device cluster including a plurality of computing devices to determine a computing device for performing coreness mining on the partitioned graph network; Coreness mining is performed on the partition graph network to iteratively update the node coreness of each node in the relationship graph network.
4. The data processing method according to claim 3, characterized in that: The performing core degree mining on the partition graph network to iteratively update the node core degree of each node in the relationship graph network includes: Selecting a computing node for core degree mining in the current iteration round in the partition graph network, and determining a neighboring node having an adjacency relationship with the computing node; Obtaining the current node coreness of the computing node and the neighboring node in the current iteration round; Determine the temporary node coreness of the computing node according to the current node coreness of the neighboring node, and mark the computing node whose temporary node coreness is less than the current node coreness as an active node; The current node coreness of the active node is updated according to the temporary node coreness, and the active node and neighbor nodes having an adjacency relationship with the active node are determined as computing nodes for coreness mining in the next iteration round.
5. The data processing method according to claim 4, characterized in that: The determining the temporary node coreness of the computing node according to the current node coreness of the neighboring node includes: The h-index of the computing node is determined according to the current node coreness of the neighboring nodes, and the h-index is used as the temporary node coreness of the computing node. The h-index indicates that among all the neighboring nodes of the computing node, at most h neighboring nodes have current node coreness greater than or equal to h.
6. The data processing method according to claim 5, characterized in that: The determining the h-index of the computing node according to the current node coreness of the neighboring node includes: Sort all neighbor nodes of the computing node in descending order of the coreness of the current node, and assign a sequence number starting with 0 to each neighbor node; Compare the ranking numbers of each neighbor node with the coreness of the current node respectively, and select neighbor nodes whose ranking numbers are greater than or equal to the coreness of the current node according to the comparison results; Among the neighbor nodes screened out, the current node coreness of the neighbor node with the smallest ranking number is determined as the h-index of the computing node.
7. The data processing method according to claim 4, characterized in that: The selecting of computing nodes for core degree mining in the current iteration round in the partition graph network includes: Reading node identifiers of nodes to be updated from the first storage space, the nodes to be updated include active nodes whose node coreness is updated in the previous iteration round and neighbor nodes having an adjacency relationship with the active nodes; A computing node for core degree mining in the current iteration round is selected in the partition graph network according to the node identifier of the node to be updated.
8. The data processing method according to claim 7, characterized in that: After updating the current node coreness of the active node according to the temporary node coreness, the method further includes: Writing the updated current node coreness of the active node into a second storage space, where the second storage space is used to store the node coreness of all nodes in the relationship graph network; Obtaining node identifiers of the active node and neighboring nodes of the active node, and writing the node identifiers into a third storage space, where the third storage space is used to store computing nodes for core degree mining in the next iteration round; After completing the mining of the cores of all partition graph networks in the current iteration round, the third storage space overwrites the data in the first storage space and resets the third storage space.
9. The data processing method according to claim 4, characterized in that: The obtaining of the current node coreness of the computing node and the neighboring node in the current iteration round includes: The current node coreness of the computing node and the neighboring node in the current iteration round is read from a second storage space, where the second storage space is used to store the node coreness of all nodes in the relationship graph network.
10. The data processing method according to claim 1, characterized in that: The pruning of the relationship graph network according to the node coreness to remove some nodes and edges in the relationship graph network includes: Obtain the minimum coreness of active nodes in the current iteration round and the minimum coreness of active nodes in the previous iteration round; If the minimum coreness of the active nodes in the current iteration round is greater than the minimum coreness of the active nodes in the previous iteration round, then the convergence nodes in the relationship graph network are screened according to the minimum coreness of the active nodes in the previous iteration round, and the convergence nodes are nodes whose node coreness is less than or equal to the minimum coreness of the active nodes in the previous iteration round; The convergence node and the edges connected to the convergence node are removed from the relationship graph network.
11. The data processing method according to claim 1, characterized in that: The network compression condition includes that the number of edges in the relationship graph network is less than a preset number threshold.
12. The data processing method according to claim 1, characterized in that: Before performing core degree mining on the relationship graph network by using a device cluster including a plurality of computing devices, the method further includes: In the relationship graph network, respectively obtain the number of neighbor nodes having an adjacency relationship with each node; The node coreness of each node is initialized according to the number of nodes.
13. A data processing device, characterized in that: include: A graph network acquisition module is configured to acquire a relationship graph network for representing interaction relationships between a plurality of interaction objects, wherein the relationship graph network includes nodes for representing interaction objects and edges for representing interaction relationships; A coreness mining module is configured to perform coreness mining on the relationship graph network through a device cluster including a plurality of computing devices, so as to iteratively update the node coreness of all or part of the nodes in the relationship graph network; A network pruning module is configured to prune the convergent nodes in the relationship graph network according to the node coreness, so as to remove some nodes and edges in the relationship graph network; The converged node is a node whose node coreness is no longer updated during the iteration process; The cluster compression module is configured to perform compression processing on the device cluster to remove some computing devices in the device cluster when the network scale of the relationship graph network meets a preset network compression condition.
14. A computer-readable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the data processing method according to any one of claims 1 to 12 is implemented.
15. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to perform the data processing method according to any one of claims 1 to 12 by executing the executable instructions.
16. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the data processing method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Systems and methods for finding star structures as communities in networks
CN102726010A
Methods for the graphical representation of genomic sequence data
US20160342737A1
Community discovery method, device, server and computer storage medium
US20190179615A1