A data processing method and related device

By selecting high-quality nodes using a generative flow model and updating the task network using a loss function, the problem of high training time and computational cost in active learning is solved, achieving fast and efficient data processing.

CN115618065BActive Publication Date: 2026-06-02HUAWEI TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2022-09-28
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing active learning methods require significant time and computational resources to train models, making it difficult to quickly select high-quality samples for labeling while maintaining accuracy.

Method used

By generating a flow model, a subset of nodes are selected from multiple nodes. The task network is updated using a loss function. By combining the similarity and correlation of node data, high-quality nodes are quickly selected for training.

Benefits of technology

While ensuring accuracy, we can quickly select high-quality nodes, reduce training time and computing power costs, and improve data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115618065B_ABST
    Figure CN115618065B_ABST
Patent Text Reader

Abstract

The application discloses a data processing method, and relates to the field of artificial intelligence, and the method comprises the following steps: according to the data of a plurality of nodes, selecting part of nodes from the plurality of nodes by generating a flow model; and according to the data of the part of nodes, obtaining a loss, wherein the loss is used for updating the generated flow model; wherein the loss is related to one of the following: a positive contribution of updating a task network when training the task network according to the data of the part of nodes, a similarity of a joint distribution between the data of the part of nodes and the data of the plurality of nodes, and an association degree of the data of the part of nodes to a target node or a target subnetwork. According to the application, the part of nodes can be selected from the plurality of nodes by generating the flow model, and the generated flow model can gradually have the ability of selecting high-quality nodes from the plurality of nodes based on the construction of the loss, so that the high-quality nodes can be quickly selected from the plurality of nodes under the premise of ensuring the accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and more particularly to a data processing method and related equipment. Background Technology

[0002] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0003] Active learning is a machine learning or artificial intelligence method that actively selects the most valuable samples for annotation. Its goal is to train the model to achieve the best possible performance using as few high-quality sample annotations as possible. In other words, active learning methods can improve the gain of both samples and annotations, maximizing model performance within a limited annotation budget. It is a solution that improves data efficiency from the perspective of samples, and therefore is applied to tasks with high annotation costs and difficulties.

[0004] In existing technologies, the active learning problem is equivalent to selecting and adding a set of labeled samples such that the maximum distance between other samples and the labeled sample set is minimized; this is essentially the k-center set coverage problem. However, each iteration requires training the model to convergence on the new labeled dataset in preparation for the next round of queries, resulting in excessive time and computational overhead.

[0005] Therefore, there is an urgent need for a method that can quickly perform active learning while ensuring accuracy. Summary of the Invention

[0006] This application provides a data processing method that can quickly select high-quality nodes from multiple nodes while ensuring accuracy.

[0007] In a first aspect, this application provides a data processing method, the method comprising: selecting a subset of nodes from the plurality of nodes based on data from a plurality of nodes by generating a flow model; obtaining a loss based on the data from the subset of nodes, the loss being used to update the generating flow model; wherein the loss is related to one of the following: the positive contribution of the task network update during training based on the data from the subset of nodes, the similarity of the joint distribution between the data from the subset of nodes and the data from the plurality of nodes, and the degree of correlation of the data from the subset of nodes to a target node or a target subnet, wherein the target node is one of the plurality of nodes, and the target subnet is composed of a subset of the plurality of nodes.

[0008] In this embodiment of the application, by generating a flow model, some nodes can be selected from multiple nodes, and based on the construction of loss, the generating flow model can gradually acquire the ability to select high-quality nodes from multiple nodes, thereby enabling the rapid selection of high-quality nodes from multiple nodes while ensuring accuracy.

[0009] In one possible implementation, the task network is a graph neural network, and the plurality of nodes are nodes in the graph information.

[0010] In one possible implementation, the node's data includes information about the node and information about the connections between the node and other nodes.

[0011] In one possible implementation, the positive contribution is positively correlated with the processing accuracy of the updated task network obtained after training the task network with data from the partial nodes.

[0012] In one possible implementation, the degree of association is related to the accuracy of predicting the label of the target node or the node in the target subnet through the task network based on the data of the partial nodes; or, the degree of association is related to the similarity between the partial nodes and the corresponding ground truth, where the ground truth consists of multiple nodes.

[0013] In one possible implementation, the partial nodes include intermediate nodes and termination nodes; the loss specifically includes a first loss corresponding to the intermediate node and a second loss corresponding to the termination node; the first loss is the difference between the first input flow and the output flow of the intermediate node; the second loss is the difference between the second input flow and the reward value of the termination node; the reward value is related to one of the following: the positive contribution to the update of the task network when training the task network based on the data of the partial nodes, or the similarity between the data of the partial nodes and the joint distribution of the data of the multiple nodes.

[0014] In one possible implementation, the task network is used to predict the labels of the partial nodes based on the data of the partial nodes.

[0015] In one possible implementation, the step of processing data from multiple nodes by generating a stream model and selecting a subset of nodes from the multiple nodes includes: processing data from multiple nodes by generating a stream model, sequentially selecting nodes from the multiple nodes until the number of selected nodes reaches a preset value, thereby obtaining the subset of nodes.

[0016] In one possible implementation, the node is a chip, and the node's data is a local fragment of the chip or fault information; or,

[0017] The node is a node in the communication network, and the node's data includes key performance indicators (KPIs), operational data, or alarm information of the network element.

[0018] The degree of correlation can be causal relationship;

[0019] In this chip, a local segment can be a local area on the chip surface, and multiple local segments can be multiple local areas on the chip surface. Any two local segments have the same area size and external contour shape. The same area size can be understood as the area of ​​the region containing the local segment being the same, and the same external contour shape can be understood as the external contour shape of the region containing the local segment being the same, such as both being squares or rectangles with the same aspect ratio. In one possible implementation, the area of ​​each local segment is within a preset range; the area of ​​each local segment cannot be too large or too small. The area of ​​the local segment can be related to the size of the chip; the larger the chip, the larger the area of ​​the local segment. For example, the area of ​​the local segment and the area of ​​the chip can maintain a certain ratio. The area of ​​the local segment can also be related to the spacing length between the basic units on the chip. For example, the side length of the local segment can be set to a preset multiple of the spacing length between the basic units (such as copper-padded polygon areas on the chip), such as 3 times, 4 times, 5 times, etc. Each local segment may include arranged devices and / or connecting lines between devices. In this embodiment, a local segment may specifically be image information of each local segment or other information that can express the arrangement of devices or the structure of connecting lines on the local segment. Based on this information, the structural features of the local segment can be uniquely determined.

[0020] The chip's fault information may include the number of times each local segment appears in the diagnostic report; or the probability that each local segment causes the faulty chip to fail.

[0021] Key performance indicators (KPIs) can be used to measure the operational status of network elements in a communication network. Typically, anomaly detection equipment collects observation data for each KPI at different times.

[0022] In one possible implementation, a node can be a user, and the node's data can be the attribute information of the item and the attribute information of the user.

[0023] The degree of correlation can be causal relationship;

[0024] In one possible implementation, the user's attribute information includes at least one of the following: gender, age, occupation, income, hobbies, and education level.

[0025] In one possible implementation, the item's attribute information includes at least one of the following: item name, developer, installation package size, category, and rating.

[0026] The user's attribute information can be attributes related to the user's preferences, including at least one of gender, age, occupation, income, hobbies, and education level. Gender can be male or female, age can be a number between 0 and 100, occupation can be teacher, programmer, chef, etc., hobbies can be basketball, tennis, running, etc., and education level can be primary school, junior high school, high school, university, etc. This application does not limit the specific type of user attribute information.

[0027] The items can be physical or virtual, such as apps, audio / video files, web pages, and news articles. The attribute information of the items can be at least one of the following: item name, developer, installation package size, category, and rating. For example, if an item is an application, the category can be chat, parkour, office, etc., and the rating can be a score or review of the item. This application does not limit the specific type of attribute information of the items.

[0028] Among them, the operation type can be the type of behavior operation performed by the user on the item. On online platforms and applications, users often have a variety of interaction forms with the item (that is, there are multiple operation types), such as browsing, clicking, adding to cart, and purchasing in e-commerce platforms.

[0029] It should be understood that the causal relationships between multiple variables ultimately obtained through the generation flow model can include a causal relationship between at least one node and the user's operation type.

[0030] Secondly, this application provides a data processing method, the method comprising:

[0031] Based on data from multiple nodes, a flow model is generated to select a subset of nodes from the multiple nodes;

[0032] The data from a subset of nodes is used as training samples for the task network; or,

[0033] The selected nodes are those nodes that best represent the characteristics of the plurality of nodes; or,

[0034] The partial nodes are the nodes with the highest correlation to the target node or target subnet among the plurality of nodes; the target node is one of the plurality of nodes, and the target subnet is composed of a partial number of the plurality of nodes; the target node is the starting node when the generation flow model selects the partial nodes, and a node in the target subnet is the starting node when the generation flow model selects the partial nodes.

[0035] In one possible implementation, the task network is a graph neural network, and the plurality of nodes are nodes in the graph information.

[0036] In one possible implementation, the node's data includes information about the node and information about the connections between the node and other nodes.

[0037] In one possible implementation, the task network is used to predict the labels of the partial nodes based on the data of the partial nodes.

[0038] In one possible implementation, the step of selecting a subset of nodes from the multiple nodes by generating a flow model based on data from multiple nodes includes:

[0039] By generating a stream model, data from multiple nodes is processed, and nodes are selected sequentially from the multiple nodes until the number of selected nodes reaches a preset value, thus obtaining the partial nodes.

[0040] Thirdly, this application provides a data processing apparatus, the apparatus comprising:

[0041] The processing module is used to select a subset of nodes from the multiple nodes by generating a flow model based on the data from the multiple nodes;

[0042] The update module is used to obtain a loss based on the data of the partial nodes, and the loss is used to update the generated flow model; wherein the loss is related to one of the following:

[0043] The positive contribution of the task network to the update of the task network when training the task network based on the data of the partial nodes, the similarity of the joint distribution between the data of the partial nodes and the data of the multiple nodes, and the degree of correlation of the data of the partial nodes with the target node or the target subnet, wherein the target node is one of the multiple nodes and the target subnet is composed of some of the multiple nodes.

[0044] In one possible implementation, the task network is a graph neural network, and the plurality of nodes are nodes in the graph information.

[0045] In one possible implementation, the node's data includes information about the node and information about the connections between the node and other nodes.

[0046] In one possible implementation, the positive contribution is positively correlated with the processing accuracy of the updated task network obtained after training the task network with data from the partial nodes.

[0047] In one possible implementation, the degree of correlation is related to the accuracy of predicting the labels of the target node or nodes in the target subnet using the task network based on the data from the partial nodes; or,

[0048] The degree of association is related to the similarity between the partial nodes and their corresponding truth values, where the truth values ​​are multiple nodes.

[0049] In one possible implementation, the partial nodes include intermediate nodes and termination nodes; the loss specifically includes a first loss corresponding to the intermediate node and a second loss corresponding to the termination node;

[0050] The first loss is the difference between the first input flow and the first output flow of the intermediate node;

[0051] The second loss is the difference between the second input flow and the reward value of the termination node; the reward value is related to one of the following:

[0052] The positive contribution of the task network to the update of the task network when training the task network based on the data of the partial nodes, or the similarity of the joint distribution between the data of the partial nodes and the data of the multiple nodes.

[0053] In one possible implementation, the task network is used to predict the labels of the partial nodes based on the data of the partial nodes.

[0054] In one possible implementation, the processing module is specifically used for:

[0055] By generating a stream model, data from multiple nodes is processed, and nodes are selected sequentially from the multiple nodes until the number of selected nodes reaches a preset value, thus obtaining the partial nodes.

[0056] Fourthly, this application provides a data processing apparatus, the apparatus comprising:

[0057] The processing module is used to select a subset of nodes from the multiple nodes by generating a flow model based on the data from the multiple nodes;

[0058] The data from a subset of nodes is used as training samples for the task network; or,

[0059] The selected nodes are those nodes that best represent the characteristics of the plurality of nodes; or,

[0060] The partial nodes are the nodes with the highest correlation to the target node or target subnet among the plurality of nodes; the target node is one of the plurality of nodes, and the target subnet is composed of a partial number of the plurality of nodes; the target node is the starting node when the generation flow model selects the partial nodes, and a node in the target subnet is the starting node when the generation flow model selects the partial nodes.

[0061] In one possible implementation, the task network is a graph neural network, and the plurality of nodes are nodes in the graph information.

[0062] In one possible implementation, the node's data includes information about the node and information about the connections between the node and other nodes.

[0063] In one possible implementation, the task network is used to predict the labels of the partial nodes based on the data of the partial nodes.

[0064] In one possible implementation, the processing module is specifically used for:

[0065] By generating a stream model, data from multiple nodes is processed, and nodes are selected sequentially from the multiple nodes until the number of selected nodes reaches a preset value, thus obtaining the partial nodes.

[0066] Fifthly, embodiments of this application provide a data processing apparatus, which may include a memory, a processor, and a bus system, wherein the memory is used to store a program, and the processor is used to execute the program in the memory to perform the methods described in the first aspect and any optional methods thereof, or the methods described in the second aspect and any optional methods thereof.

[0067] Sixthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the first aspect and any optional method thereof, or the second aspect and any optional method thereof.

[0068] In a seventh aspect, embodiments of this application provide a computer program product including instructions that, when run on a computer, cause the computer to perform the first aspect and any optional method thereof, or the second aspect and any optional method thereof.

[0069] Eighthly, this application provides a chip system including a processor for supporting a data processing device in implementing some or all of the functions involved in the foregoing aspects, such as transmitting or processing data involved in the foregoing methods; or, information. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for the execution device or training device. This chip system may be composed of chips or may include chips and other discrete devices. Attached Figure Description

[0070] Figure 1 This is a schematic diagram of an application architecture;

[0071] Figure 2 This is a schematic diagram of an application architecture;

[0072] Figure 3 This is a schematic diagram of an application architecture;

[0073] Figure 4 This is a schematic diagram of an application architecture;

[0074] Figure 5 This is a schematic diagram of an application architecture;

[0075] Figure 6 This is a schematic diagram of the generation process of a generative flow model;

[0076] Figure 7 An illustration of the cloud services provided in the embodiments of this application;

[0077] Figure 8 An illustration of an embodiment of the data processing method provided in this application;

[0078] Figure 9 This is a schematic diagram of active learning using a graph neural network in an embodiment of this application;

[0079] Figure 10 This is a schematic diagram of an active domain adaptive method in an embodiment of this application;

[0080] Figure 11This is a schematic diagram of an association relationship identification in an embodiment of this application;

[0081] Figure 12 This is a schematic diagram of a software architecture in an embodiment of this application;

[0082] Figure 13 This application provides an embodiment of a data processing apparatus.

[0083] Figure 14 A schematic diagram of the structure of the execution device provided in the embodiments of this application;

[0084] Figure 15 This is a schematic diagram of a server structure provided in an embodiment of this application;

[0085] Figure 16 This is a schematic diagram of a chip structure provided in an embodiment of this application. Detailed Implementation

[0086] The embodiments of the present invention will now be described with reference to the accompanying drawings. The terminology used in the embodiments section is for illustrative purposes only and is not intended to limit the scope of the invention.

[0087] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0088] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0089] The terms “substantially,” “about,” and similar terms used herein are used as approximations rather than as terms of degree, and are intended to take into account the inherent biases of measurements or calculations known to those skilled in the art. Furthermore, the use of “may” in describing embodiments of the invention refers to “one or more possible embodiments.” The terms “use,” “using,” and “used” used herein are to be considered synonymous with the terms “utilize,” “utilizing,” and “utilized,” respectively. Additionally, the term “exemplary” is intended to refer to an instance or illustration.

[0090] First, we introduce the application scenarios of this application. This application can be applied to, but is not limited to, active learning of graph neural networks, active domain adaptation of generative flow networks, and graph neural network interpreters based on generative flow networks. These scenarios will be described in detail below:

[0091] 1. Active learning in graph neural networks;

[0092] Active learning is a machine learning or artificial intelligence method that actively selects the most valuable samples for annotation. Its goal is to achieve the best possible model performance using as few high-quality samples as possible. In other words, active learning methods can improve the gain of samples and annotations, maximizing model performance within a limited annotation budget. It is a solution to improve data efficiency from the perspective of samples, and therefore it is applied to tasks with high annotation costs and difficulties.

[0093] For graph neural networks, training them based on a large number of training samples requires a significant amount of time. However, training sets often include both high-quality and low-quality data. Even with a small amount of data, a high-quality graph neural network model can still be trained quickly based on high-quality data.

[0094] In this embodiment of the application, a generative flow model is trained, which can select some high-quality samples from multiple training samples for use in training the graph neural network.

[0095] Specifically, the generative flow model can be used to predict labels for each state node by employing a graph neural network. For a set of input datasets, where b state nodes are pre-defined, the generative flow model network will generate a series of directed acyclic graphs (DAGs) with b state nodes. The loss is calculated by the neural network in the generative flow model and a pre-defined reward function to optimize the network, thereby learning a strategy that makes it more likely that DAGs with high reward values ​​will be sampled and generated. The DAG with the largest reward value is the generative flow relationship that best fits the dataset.

[0096] 2. Active Domain Adaptation of Generative Flow Networks

[0097] Similar to graph neural network active learning, both aim to train a generative flow model. The generative flow model can select some high-quality samples from multiple training samples. The difference lies in the training method and the purpose of the selected samples.

[0098] 3. Identification of the correlation between nodes

[0099] Graph data consists of multiple nodes, and there may be certain relationships between different nodes or between subgraphs (which may also contain multiple nodes), such as causal relationships. A generative flow model can be used to identify which nodes in the graph data have a high degree of correlation with a specified node, or vice versa.

[0100] Taking correlation identification as an example of a task related to generating flow networks, in one possible implementation, fault identification in various scenarios can be performed by implementing tasks related to generating flow networks. For example, it can be used for fault identification in communication networks, systemic defect identification in chips, fault node identification in computer transactions, and mechanical fault identification, etc. These will be explained in detail below:

[0101] In one possible implementation, fault identification in the communication network can be achieved by implementing tasks related to generating stream networks.

[0102] Key performance indicators (KPIs) in communication networks are used to measure the operational status of network elements. Typically, anomaly detection devices collect observation data for each KPI at different times. If the observed KPI data is abnormal, it indicates that the operational status of a network element in the communication network is abnormal. Network maintenance engineers need to investigate the cause of the abnormal KPI to troubleshoot the problem.

[0103] In one possible implementation, the causal relationship between certain abnormal KPIs can be determined based on data processing methods. For example, the root cause of the abnormality of the first KPI is due to the abnormality of the second KPI. Thus, network operation and maintenance engineers can identify the faulty network element based on the second KPI to eliminate the fault.

[0104] In one possible implementation, KPI is equivalent to the variable in the embodiments of this application, and the observation data of KPI at different times is equivalent to the data of the variable in the embodiments of this application. The causal relationship between KPI can be obtained through the data processing method in the embodiments of this application.

[0105] In one possible implementation, systematic defects in a chip can be identified by implementing tasks related to generating flow networks.

[0106] With the development of electronic product functions and the expansion of application areas, chips, as the core components of electronic products, have become an indispensable part of people's lives. Chip production is mainly divided into two parts: layout design and manufacturing. Layout design typically includes multi-layered circuit functional design, while manufacturing includes processes such as production, packaging, and testing. When the same chip design is used with different manufacturing processes, some circuit structures that are normal under the original process may have defects, resulting in a lower-than-expected chip yield. This type of circuit structure with design defects due to process changes is called a systemic defect.

[0107] The presence of systemic defects increases the likelihood of chip circuit failure. Chips with malfunctioning circuits cannot function properly, leading to a decrease in chip yield. A decline in yield increases production costs and may even cause products to miss their sales window. Therefore, identifying the root causes of systemic defects is crucial for improving product yield. To identify systemic defects, the chip's design structure can be analyzed to determine the types of local segments on the chip that may lead to potential chip failure.

[0108] In one possible implementation, the type of local segment that may cause chip failure can be determined by using data processing methods based on images of various segments on the chip.

[0109] In one possible implementation, fault node identification in computer transactions can be achieved by implementing tasks related to generating flow networks.

[0110] With the development of computer technology, the number of transactions that computer devices can handle has increased rapidly. Furthermore, computer devices can execute the same transaction multiple times a day to meet the needs of a large number of users. To improve the performance of transaction execution, it is necessary to analyze transaction issues in order to better execute transactions.

[0111] Currently, the transaction analysis process typically involves: during transaction execution, real-time collection of execution records; extraction of information about each node called by the transaction from these records, including the node's name, call duration, status code, and call relationships between different nodes; and then displaying this information on the interface. Based on this information, causal identification methods can be used to determine if any issues exist with these nodes, ultimately identifying the node causing the transaction problem.

[0112] In one possible implementation, mechanical fault identification can be achieved by implementing tasks related to generating flow networks.

[0113] For machining systems, if the causal relationship between each attribute and whether the product is qualified has been determined, the attributes that have the greatest impact on unqualified products can be adjusted first based on the causal relationship found.

[0114] In one possible implementation, for a power transmission system, if the intermediate voltage at each transmission device, the operating state of the transmission system, and the causal relationship between current and power loss have been determined, then the variables with the greatest impact on power loss can be prioritized for adjustment based on the found causal relationships. This approach can improve the performance of the power transmission system.

[0115] In one possible implementation, causal identification related to user behavior in the recommendation field can be performed using data processing methods.

[0116] In one possible implementation, user operation logs can be obtained, which may include user actions on items, item attribute information, and user attribute information. The causal relationship between each attribute information and the user's actions can be determined by implementing tasks related to the generation flow network.

[0117] In this application embodiment, the functions corresponding to the above three scenarios can be referred to as tasks related to the generation of streaming networks.

[0118] In terms of product implementation, the embodiments of this application can be applied to applications (or other types of computer program products) that implement tasks related to generating streaming networks, cloud services provided by cloud-side servers that are related to implementing tasks related to generating streaming networks, etc.

[0119] The following sections will introduce the application program for implementing tasks related to the generation of streaming networks in the embodiments of this application, focusing on both the functional architecture and the product architecture that implements the functions.

[0120] Reference Figure 1 , Figure 1 This is a schematic diagram of the functional architecture of the application that implements tasks related to generating streaming networks in this embodiment of the application:

[0121] In one possible implementation, embodiments of this application include a system capable of automatically identifying high-quality data from multiple input data sets (e.g., an application implementing tasks related to generating streaming networks). Figure 1 As shown, the application 102 that implements the task of generating stream network-related data can receive multiple data 101 inputs (e.g., data from multiple nodes in graph data) and output high-quality data 103 from the multiple data. The application 102 that implements the task of generating stream network-related data can be executed on at least one computer system (for example) and includes computer code that, when executed by one or more computers, causes the computers to perform data processing methods described herein.

[0122] In one possible implementation, the application that performs tasks related to generating streaming networks can run on an end-user device or on a cloud-based server.

[0123] For example, the terminal device may have an application installed that performs tasks related to generating streaming networks, including data input, data processing (e.g., the data processing method in the embodiments of this application), and data output actions that can be performed by the terminal device.

[0124] For example, the terminal device can have a client application installed that performs tasks related to generating streaming networks. The actions of data input and data output can be performed by the terminal device, while the actions of data processing (such as the data processing method in the embodiments of this application) can be performed by the cloud-side server. That is, the terminal device can transmit the data required for data processing (such as the data processing method in the embodiments of this application) to the cloud-side server. After the cloud-side server completes the data processing action, it can return the data processing result to the terminal device on the terminal side, and the terminal device can output based on the processing result.

[0125] The following describes the entity architecture of the application that runs and implements tasks related to generating streaming networks in the embodiments of this application.

[0126] Reference Figure 2 , Figure 2 This is a schematic diagram of the entity architecture of the application that runs and implements tasks related to the generation of streaming networks in the embodiments of this application:

[0127] See Figure 2 , Figure 2 A schematic diagram of a system architecture is shown. The system may include a terminal 100 and a server 200. The server 200 may include one or more servers (…). Figure 2(The example includes a server), where server 200 can provide programs for one or more terminals to perform tasks related to generating streaming networks.

[0128] The terminal 100 may have an application installed to perform tasks related to generating streaming networks, or a webpage related to performing tasks related to generating streaming networks may be opened. The application and webpage may provide an interface for performing tasks related to generating streaming networks. The terminal 100 may receive relevant data (multiple data) input by the user on the interface for performing tasks related to generating streaming networks, and send the data to the server 200. The server 200 may obtain a processing result (high-quality data among the multiple data) based on the received data, and return the processing result to the terminal 100.

[0129] It should be understood that in some optional implementations, the terminal 100 can also complete the action of obtaining the data processing result based on the received data on its own, without the need for the server to cooperate. This application embodiment is not limited to this.

[0130] The following description Figure 2 The product form of the mid-terminal 100;

[0131] The terminal 100 in this application embodiment can be a mobile phone, tablet computer, wearable device, vehicle device, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc., and this application embodiment does not impose any restrictions on it.

[0132] Figure 3 A schematic diagram of an optional hardware structure for terminal 100 is shown.

[0133] refer to Figure 3 As shown, the terminal 100 may include a radio frequency unit 110, a memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160 (optional), a speaker 161 (optional), a microphone 162 (optional), a processor 170, an external interface 180, a power supply 190, and other components. Those skilled in the art will understand that... Figure 3 These are merely examples of terminals or multi-functional devices and do not constitute a limitation on terminals or multi-functional devices. They may include more or fewer components than shown in the illustration, or combine certain components, or use different components.

[0134] The input unit 130 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the portable multi-functional device. Specifically, the input unit 130 may include a touchscreen 131 (optional) and / or other input devices 132. The touchscreen 131 can collect touch operations performed by the user on or near it (such as operations performed by the user using fingers, knuckles, styluses, or any suitable object on or near the touchscreen), and drive the corresponding connection devices according to a pre-set program. The touchscreen can detect the user's touch actions, convert the touch actions into touch signals and send them to the processor 170, and can receive and execute commands sent by the processor 170; the touch signal includes at least touch point coordinate information. The touchscreen 131 can provide an input interface and an output interface between the terminal 100 and the user. In addition, various types of touchscreens, such as resistive, capacitive, infrared, and surface acoustic wave, can be used to implement the touchscreen. Besides the touchscreen 131, the input unit 130 may also include other input devices. Specifically, other input devices 132 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.

[0135] Other input devices 132 can receive parameters related to the tasks of generating the streaming network, such as multiple data in the embodiments of this application.

[0136] The display unit 140 can be used to display information input by the user or information provided to the user, various menus of the terminal 100, interactive interfaces, file display, and / or playback of any multimedia file. In this embodiment, the display unit 140 can be used to display the interface of an application that performs tasks related to generating streaming networks, high-quality data from multiple datasets, etc.

[0137] The memory 120 can be used to store instructions and data. The memory 120 may primarily include an instruction storage area and a data storage area. The data storage area can store various types of data, such as multimedia files and text. The instruction storage area can store software units such as operating systems, applications, and instructions required for at least one function, or subsets or extended sets thereof. It may also include non-volatile random access memory. It provides the processor 170 with hardware, software, and data resources for managing the computing device, supporting control software and applications. It is also used for storing multimedia files, as well as storing running programs and applications.

[0138] The processor 170 is the control center of the terminal 100. It connects various parts of the terminal 100 via various interfaces and lines. By running or executing instructions stored in the memory 120 and calling data stored in the memory 120, it performs various functions of the terminal 100 and processes data, thereby providing overall control of the terminal device. Optionally, the processor 170 may include one or more processing units; preferably, the processor 170 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 170. In some embodiments, the processor and memory can be implemented on a single chip; in some embodiments, they can also be implemented separately on independent chips. The processor 170 can also be used to generate corresponding operation control signals, send them to the corresponding components of the computing device, read and process data in the software, especially read and process data and programs in the memory 120, so that the various functional modules therein perform corresponding functions, thereby controlling the corresponding components to act according to the instructions.

[0139] The memory 120 can be used to store software code related to the data processing method, and the processor 170 can execute the steps of the chip's data processing method, and can also schedule other units (such as the above-mentioned input unit 130 and display unit 140) to achieve the corresponding functions.

[0140] The radio frequency unit 110 (optional) can be used for receiving and transmitting signals during information transmission or calls. For example, it can receive downlink information from the base station and process it for the processor 170; additionally, it can transmit uplink data to the base station. Typically, the RF circuit includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. Furthermore, the radio frequency unit 110 can also communicate wirelessly with network devices and other devices. This wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0141] In this embodiment of the application, the radio frequency unit 110 can send data to the server 200 and receive the task results related to the implementation of the streaming network sent by the server 200.

[0142] It should be understood that the radio frequency unit 110 is optional and can be replaced with other communication interfaces, such as a network port.

[0143] The terminal 100 also includes a power supply 190 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 170 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.

[0144] Terminal 100 also includes an external interface 180, which can be a standard Micro USB interface or a multi-pin connector, which can be used to connect terminal 100 to other devices for communication or to connect a charger to charge terminal 100.

[0145] Although not shown, terminal 100 may also include a flash, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with various functions, etc., which will not be described in detail here. Some or all of the methods described below can be applied to, for example... Figure 3 In the terminal 100 shown.

[0146] The following description Figure 2 The product form of the mid-range server 200;

[0147] Figure 4 A structural diagram of a server 200 is provided, as follows: Figure 4 As shown, server 200 includes bus 201, processor 202, communication interface 203, and memory 204. Processor 202, memory 204, and communication interface 203 communicate with each other via bus 201.

[0148] Bus 201 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0149] The processor 202 can be any one or more of the following processors: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP), or digital signal processor (DSP).

[0150] Memory 204 may include volatile memory, such as random access memory (RAM). Memory 204 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0151] The memory 204 can be used to store software code related to the data processing method, and the processor 202 can execute the steps of the chip's data processing method, and can also schedule other units to achieve corresponding functions.

[0152] It should be understood that the aforementioned terminal 100 and server 200 can be centralized or distributed devices. The processors (e.g., processor 170 and processor 202) in the aforementioned terminal 100 and server 200 can be hardware circuits (such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), general-purpose processors, digital signal processors (DSPs), microprocessors or microcontrollers, etc.) or combinations of these hardware circuits. For example, the processor can be a hardware system with instruction execution capabilities, such as a CPU or DSP, or a hardware system without instruction execution capabilities, such as an ASIC or FPGA, or a combination of the aforementioned hardware systems without instruction execution capabilities and hardware systems with instruction execution capabilities.

[0153] It should be understood that the data processing methods in the embodiments of this application involve AI-related operations. When performing AI operations, the instruction execution architecture of the terminal device and the server is not limited to... Figure 3 as well as Figure 4 The processor and memory architecture shown below. Figure 5 The system architecture provided in the embodiments of this application will be described in detail.

[0154] Figure 5 This is a schematic diagram of the system architecture provided for an embodiment of this application. For example... Figure 5 As shown, the system architecture 500 includes an execution device 510, a training device 520, a database 530, a client device 540, a data storage system 550, and a data acquisition system 560.

[0155] The execution device 510 includes a calculation module 511, an I / O interface 512, a preprocessing module 513, and a preprocessing module 514. The calculation module 511 may include a target model / rule 501, while the preprocessing modules 513 and 514 are optional.

[0156] The execution device 510 can be a terminal device or a server that runs the application that performs tasks related to generating streaming networks.

[0157] The data acquisition device 560 is used to collect training samples. Training samples can include information about I / O units, bump information, and the total number of connections, etc. After collecting the training samples, the data acquisition device 560 stores these training samples in the database 530.

[0158] The training device 520 can use the database 530 or training samples (such as multiple data in the embodiments of this application) from the client device 540 to train the neural network (such as the generative flow model in the embodiments of this application) to obtain the target model / rule 501 and high-quality data from the multiple data.

[0159] It should be noted that in practical applications, the training samples maintained in database 530 may not all come from the data acquisition device 560; they may also be received from other devices (e.g., from client device 540). Furthermore, it should be noted that training device 520 may not necessarily train the target model / rule 501 entirely based on the training samples maintained in database 530; it may also obtain training samples from the cloud or other sources for model training. The above description should not be construed as limiting the embodiments of this application.

[0160] Optionally, the target model / rule 501 trained using training device 520 can be applied to different systems or devices, such as... Figure 5 The execution device 510 shown can be a terminal, such as a mobile phone terminal, tablet computer, laptop computer, augmented reality (AR) / virtual reality (VR) device, vehicle terminal, etc., or it can be a server, etc.

[0161] Specifically, the training device 520 can transmit the trained model or causal recognition results to the execution device 510.

[0162] exist Figure 5 In the execution device 510, an input / output (I / O) interface 512 is configured for data interaction with external devices. Users can input data (such as multiple data or multiple variables in this embodiment) into the I / O interface 512 through the client device 540.

[0163] Preprocessing modules 513 and 514 are used to preprocess the input data received from the I / O interface 512. It should be understood that preprocessing modules 513 and 514 may be absent, or only one preprocessing module may be used. When preprocessing modules 513 and 514 are absent, the calculation module 511 can be used directly to process the input data.

[0164] During the preprocessing of input data by the execution device 510, or during the calculation module 511 of the execution device 510 performing calculations and other related processes, the execution device 510 can call data, code, etc. in the data storage system 550 for corresponding processing, or store the data, instructions, etc. obtained from the corresponding processing into the data storage system 550.

[0165] Finally, I / O interface 512 provides the processing results (e.g., high-quality data from multiple data sets) to client device 540, thereby providing them to the user.

[0166] exist Figure 5 In the illustrated scenario, the user can manually provide input data, which can be done through the interface provided by I / O interface 512. Alternatively, the client device 540 can automatically send input data to I / O interface 512. If user authorization is required for the client device 540 to automatically send input data, the user can set the corresponding permissions in the client device 540. The user can view the output results of the execution device 510 on the client device 540, which can be presented in various forms such as display, sound, or animation. The client device 540 can also act as a data acquisition terminal, collecting the input data and output results of the input I / O interface 512 as shown in the figure, and storing them as new sample data in database 530. Alternatively, data can be collected directly from the I / O interface 512 without going through the client device 540, using the input data and output results of the input I / O interface 512 as shown in the figure, and storing them as new sample data in database 530.

[0167] It is worth noting that, Figure 5 This is merely a schematic diagram of a system architecture provided in an embodiment of this application. The positional relationships between the devices, components, modules, etc., shown in the diagram do not constitute any limitation. For example, in Figure 5 In this context, the data storage system 550 is an external storage device relative to the execution device 510. However, in other cases, the data storage system 550 may also be placed within the execution device 510. It should be understood that the aforementioned execution device 510 may be deployed within the client device 540.

[0168] From the training side of the model:

[0169] In this embodiment of the application, the training device 520 can access the memory ( Figure 5 (Not shown in the diagram, but can be integrated into the training device 520 or deployed separately from the training device 520) The code stored in the diagram can be used to implement the steps related to model training in the embodiments of this application.

[0170] In this embodiment of the application, the training device 520 may include hardware circuits (such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), general-purpose processors, digital signal processors (DSPs), microprocessors or microcontrollers, etc.) or combinations of these hardware circuits. For example, the training device 520 may be a hardware system with instruction execution capabilities, such as a CPU or DSP, or a hardware system without instruction execution capabilities, such as an ASIC or FPGA, or a combination of the aforementioned hardware systems without instruction execution capabilities and hardware systems with instruction execution capabilities.

[0171] It should be understood that the training device 520 can be a combination of a hardware system without the function of executing instructions and a hardware system with the function of executing instructions. Some steps related to the training of the neutralization model provided in the embodiments of this application can also be implemented by the hardware system in the training device 520 without the function of executing instructions, which is not limited here.

[0172] II. Cloud services provided by the server related to the functions implemented by the streaming model:

[0173] In one possible implementation, the server can provide services related to the implementation of the streaming model to the client through an application programming interface (API).

[0174] In this process, the terminal device can send relevant parameters (such as multiple data) to the server through the API provided by the cloud. The server can obtain the processing result (such as high-quality data among multiple data) based on the received parameters and return the processing result to the terminal.

[0175] The description of the terminal and server can be found in the above embodiments, and will not be repeated here.

[0176] like Figure 6 This illustrates the process of using a cloud service provided by a cloud platform that implements related functions of a generative flow model.

[0177] 1. Activate and purchase content moderation services.

[0178] 2. Users can download the software development kit (SDK) corresponding to the content moderation service. Cloud platforms usually provide multiple development versions of the SDK for users to choose from according to their development environment needs, such as JAVA version SDK, Python version SDK, PHP version SDK, Android version SDK, etc.

[0179] 3. After downloading the corresponding version of the SDK to their local machine according to their needs, users can import the SDK project into their local development environment, configure and debug it in the local development environment, and develop other functions in the local development environment, thus forming an application that integrates the capabilities related to the implementation of the generation flow model.

[0180] 4. When applications implementing functions related to the generation stream model need to perform these functions, they can trigger API calls to those functions. When an application triggers a function related to the generation stream model, it initiates an API request to the running instance of the service implementing that function in the cloud environment. This API request carries multiple data points, which are then processed by the running instance in the cloud environment to obtain the processing result.

[0181] 5. The cloud environment returns the processing result to the application, thus completing a service call.

[0182] Since the embodiments of this application involve a large number of neural network applications, for ease of understanding, the relevant terms and concepts such as neural networks involved in the embodiments of this application will be introduced below.

[0183] (1) Neural Network

[0184] A neural network can be composed of neural units, which can be defined as a computational unit that takes xs (i.e., input data) and an intercept of 1 as input. The output of this computational unit can be:

[0185]

[0186] Where s = 1, 2, ..., n, where n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer, and the activation function can be the sigmoid function. A neural network is a network formed by connecting multiple of the above-mentioned individual neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, which can be a region composed of several neural units.

[0187] (2) Deep Neural Networks

[0188] Deep Neural Networks (DNNs), also known as multilayer neural networks, can be understood as neural networks with many hidden layers, though there's no specific metric for "many." DNNs can be categorized into three layers based on their position: input layers, hidden layers, and output layers. Generally, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. All layers are fully connected, meaning that any neuron in the i-th layer is connected to any neuron in the (i+1)-th layer. Although DNNs appear complex, the operation of each layer is actually quite simple, resembling a linear relationship as follows: in, It is the input vector. It is the output vector. α is the offset vector, W is the weight matrix (also called coefficients), and α() is the activation function. Each layer is simply an adjustment of the input vector. The output vector is obtained through such a simple operation. Because DNNs have many layers, the coefficients W and the offset vector... The number of these parameters is therefore quite large. The definitions of these parameters in a DNN are as follows: Taking the coefficient W as an example: Assuming a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as... The superscript 3 represents the layer number where coefficient W resides, while the subscript corresponds to the output third layer index 2 and the input second layer index 4. In summary, the coefficients from the k-th neuron in layer L-1 to the j-th neuron in layer L are defined as follows: It's important to note that the input layer does not have a W parameter. In deep neural networks, more hidden layers allow the network to better represent complex real-world situations. Theoretically, the more parameters a model has, the higher its complexity and "capacity," meaning it can perform more complex learning tasks. Training a deep neural network is essentially the process of learning the weight matrix, with the ultimate goal of obtaining the weight matrix of all layers in the trained deep neural network (a weight matrix formed by the vectors W from many layers).

[0189] (3) Graph:

[0190] A graph is a data structure that includes at least one vertex and at least one edge. In some scenarios, vertices in a graph can be mapped to entities, and edges can be mapped to relationships between entities. A graph can be directed or undirected. Of course, a graph can also include other data besides vertices and edges, such as vertex labels and edge labels.

[0191] (4) Loss Function

[0192] In training a deep neural network, to ensure the output closely approximates the desired predicted value, we compare the network's prediction with the target value. Based on the difference, we update the weight vector of each layer (usually pre-configuring parameters before the initial update). For example, if the prediction is too high, the weight vector is adjusted to predict a lower value. This adjustment continues until the deep neural network predicts the target value or a value very close to it. Therefore, we need to predefine "how to compare the difference between the predicted and target values," which is the loss function or objective function. These are important equations used to measure the difference between the predicted and target values. Taking the loss function as an example, a higher output value (loss) indicates a greater difference, and training the deep neural network becomes a process of minimizing this loss.

[0193] (5) Backpropagation algorithm

[0194] Convolutional neural networks can employ backpropagation (BP) to correct the parameters in the initial super-resolution model during training, thereby reducing the reconstruction error loss. Specifically, forward propagation of the input signal to the output generates an error loss; this error loss information is then propagated back to update the parameters in the initial super-resolution model, leading to convergence of the error loss. The backpropagation algorithm is an error-loss-driven backpropagation process aimed at obtaining the optimal parameters of the super-resolution model, such as the weight matrix.

[0195] (6) Causal relationship

[0196] A causal relationship between a pair of variables (for example, variables A and B) can be understood as variable A causing variable B; that is, variable A is the dependent variable of variable B, and variable B is the effect variable of variable A. Specifically, all other things being equal, a change in variable A will lead to a change in variable B.

[0197] (7) Variables

[0198] A variable can be a feature of data. For example, a variable can be a feature dimension of image data, such as a semantic element in the image, like the ear area or glasses area in a portrait image, or a pixel channel in the image, such as the R channel, G channel, and B channel. Another example is that a variable can be a local segment of a chip, and the data of the variable can be an image of that local segment. Yet another example is that a variable can be a feature dimension of text data, such as a root cause of a fault. Variables can also be a feature dimension of audio data, a feature dimension of video data, and so on.

[0199] (8) Reinforcement learning

[0200] Reinforcement learning is one of the paradigms and methodologies of machine learning, used to describe and solve the problem of how an agent learns strategies to maximize rewards or achieve specific goals during its interaction with the environment.

[0201] (9) Generative Flow Model

[0202] Generative flow models (or flow-based generative flow models) are flow models that sample and construct composite structures using sequential decision-making, where the probability of generating a structure is proportional to its reward value. Generative flow models construct composite structures using sequential decision-making, employing a directed acyclic graph (DAG) to build the flow model. This means each state node has multiple parent nodes, unlike tree structures where each state node has only one parent node. The model has a unique initial node and multiple terminal nodes. The model starts sampling from the initial node, generating action sequences, transitioning between states, until reaching the terminal nodes, at which point sampling ends. The terminal node corresponds to the generated composite structure.

[0203] The initial node can include output flow, intermediate nodes can include input flow and output flow (or reward value), and the terminal node can include input flow and reward value. Imagine this flow model as a water pipe; the flow rate at the initial node is the total inflow of the model, and the sum of the flow rates at all terminal nodes (e.g., the total outflow) is the total outflow of the model. For each intermediate node, inflow equals outflow. The inflow and outflow values ​​of each node are predicted using a neural network. By optimizing the flow matching constraint as the objective function, the model can learn a policy (i.e., optimize the neural network) such that the probability of sampling and generating composite structures is proportional to their reward value; structures with higher reward values ​​are more likely to be sampled. In this way, the generative flow model can sample a series of structures with high reward values.

[0204] An exemplary diagram of the generative flow model can be found by referring to Figure 7 ,like Figure 7 As shown, where s i Represents the state, x j This represents a composite structure. s0 is the initial node, s5, s8, and s... 11 s 13 s 16 These are termination nodes because they correspond to the composite structures x5, x8, x 11 x 13 x 16 .

[0205] (10) Bayesian network structure learning

[0206] Bayesian network structure learning is about finding a causal network structure that best matches the training sample set, given a set of data samples.

[0207] (11) Directed Acyclic Graph

[0208] In graph theory, if a directed graph cannot be returned to any vertex by any number of edges, then the graph is a directed acyclic graph.

[0209] (12) Transitive closure

[0210] In graph theory, the transitive closure C describes whether a node can be reached from another node via a directed arrow. If there is a valid directed path from node A to node B, then the position (B,A) is marked as 1 in the adjacency matrix.

[0211] (13) Adjacency Matrix

[0212] An adjacency matrix is ​​a square matrix that represents the adjacency relationship between vertices. The value is 0 or 1, where 0 indicates no direct relationship and 1 indicates a relationship. The adjacency matrix of a directed acyclic graph cannot contain symmetric or diagonal positions with 1.

[0213] Active learning is a machine learning or artificial intelligence method that actively selects the most valuable samples for annotation. Its goal is to train the model to achieve the best possible performance using as few high-quality sample annotations as possible. In other words, active learning methods can improve the gain of both samples and annotations, maximizing model performance within a limited annotation budget. It is a solution that improves data efficiency from the perspective of samples, and therefore is applied to tasks with high annotation costs and difficulties.

[0214] In existing technologies, the active learning problem is equivalent to selecting and adding a set of labeled samples such that the maximum distance between other samples and the labeled sample set is minimized; this is essentially the k-center set coverage problem. However, each iteration requires training the model to convergence on the new labeled dataset in preparation for the next round of queries, resulting in excessive time and computational overhead.

[0215] Therefore, there is an urgent need for a method that can quickly perform active learning while ensuring accuracy.

[0216] To solve the above problems, refer to Figure 8 , Figure 8 This is a flowchart illustrating a data processing method provided in an embodiment of this application, such as... Figure 8 As shown in the embodiment of this application, a data processing method includes:

[0217] 801. Based on the data from multiple nodes, a flow model is generated to select some nodes from the multiple nodes.

[0218] 802. Based on the data of the aforementioned partial nodes, a loss is obtained, and the loss is used to update the generated flow model; wherein the loss is related to one of the following:

[0219] The positive contribution of the task network update to the training of the task network based on the data of the partial nodes, the similarity of the joint distribution between the data of the partial nodes and the data of the multiple nodes, and the degree of correlation of the data of the partial nodes with the target node or the target subnet, wherein the target node is one of the multiple nodes and the target subnet is composed of some of the multiple nodes.

[0220] In one possible implementation, the multiple nodes can be nodes in the graph data, and the data of the multiple nodes can include node information and connection relationship information between the nodes and other nodes.

[0221] In one possible implementation, the node is a chip, and the node's data is a local fragment of the chip or fault information; or,

[0222] The node is a node in the communication network, and the node's data includes key performance indicators (KPIs), operational data, or alarm information of the network element.

[0223] The degree of correlation can be causal relationship;

[0224] In this chip, a local segment can be a local area on the chip surface, and multiple local segments can be multiple local areas on the chip surface. Any two local segments have the same area size and external contour shape. The same area size can be understood as the area of ​​the region containing the local segment being the same, and the same external contour shape can be understood as the external contour shape of the region containing the local segment being the same, such as both being squares or rectangles with the same aspect ratio. In one possible implementation, the area of ​​each local segment is within a preset range; the area of ​​each local segment cannot be too large or too small. The area of ​​the local segment can be related to the size of the chip; the larger the chip, the larger the area of ​​the local segment. For example, the area of ​​the local segment and the area of ​​the chip can maintain a certain ratio. The area of ​​the local segment can also be related to the spacing length between the basic units on the chip. For example, the side length of the local segment can be set to a preset multiple of the spacing length between the basic units (such as copper-padded polygon areas on the chip), such as 3 times, 4 times, 5 times, etc. Each local segment may include arranged devices and / or connecting lines between devices. In this embodiment, a local segment may specifically be image information of each local segment or other information that can express the arrangement of devices or the structure of connecting lines on the local segment. Based on this information, the structural features of the local segment can be uniquely determined.

[0225] The chip's fault information may include the number of times each local segment appears in the diagnostic report; or the probability that each local segment causes the faulty chip to fail.

[0226] Key performance indicators (KPIs) can be used to measure the operational status of network elements in a communication network. Typically, anomaly detection equipment collects observation data for each KPI at different times.

[0227] In one possible implementation, a node can be a user, and the node's data can be the attribute information of the item and the attribute information of the user.

[0228] The degree of correlation can be causal relationship;

[0229] In one possible implementation, the user's attribute information includes at least one of the following: gender, age, occupation, income, hobbies, and education level.

[0230] In one possible implementation, the item's attribute information includes at least one of the following: item name, developer, installation package size, category, and rating.

[0231] The user's attribute information can be attributes related to the user's preferences, including at least one of gender, age, occupation, income, hobbies, and education level. Gender can be male or female, age can be a number between 0 and 100, occupation can be teacher, programmer, chef, etc., hobbies can be basketball, tennis, running, etc., and education level can be primary school, junior high school, high school, university, etc. This application does not limit the specific type of user attribute information.

[0232] The items can be physical or virtual, such as apps, audio / video files, web pages, and news articles. The attribute information of the items can be at least one of the following: item name, developer, installation package size, category, and rating. For example, if an item is an application, the category can be chat, parkour, office, etc., and the rating can be a score or review of the item. This application does not limit the specific type of attribute information of the items.

[0233] Among them, the operation type can be the type of behavior operation performed by the user on the item. On online platforms and applications, users often have a variety of interaction forms with the item (that is, there are multiple operation types), such as browsing, clicking, adding to cart, and purchasing in e-commerce platforms.

[0234] It should be understood that the causal relationships between multiple variables ultimately obtained through the generation flow model can include a causal relationship between at least one node and the user's operation type.

[0235] Taking nodes in graph data as an example, the data of a node can include features in multiple dimensions, such as at least one of the following dimensions:

[0236] Feature 1: N(v) represents the set of all neighboring nodes of node v, and the hyperparameter α is used for scaling.

[0237] Feature 2: It represents the uncertainty of n nodes, and the entropy is calculated using the predicted labels of the GNN;

[0238] Features 3 and 4: Calculate the forward KL divergence and backward KL divergence based on the probability of the current node's predicted label and the probability of the neighboring nodes' predicted labels;

[0239]

[0240]

[0241]

[0242]

[0243] Feature 5: Used to indicate whether a node has been tagged.

[0244] In one possible implementation, a streaming model can be generated to process data from multiple nodes and select a subset of nodes from those nodes. For example, a streaming model can be generated to process data from multiple nodes and select nodes sequentially from those nodes until the number of selected nodes reaches a preset value, thus obtaining the subset of nodes.

[0245] For active learning of graph neural networks based on the generative flow model, a loss can be obtained based on the data of the partial nodes, and the loss is used to update the generative flow model; wherein, the loss is related to the positive contribution of the task network update when training the task network based on the data of the partial nodes.

[0246] In other words, a subset of nodes can be selected by the generative flow model, and the data of these nodes can be used as training samples for the task network (e.g., a graph neural network). The positive impact of the data from these nodes on the accuracy of the task network during training can be used as a loss. This loss can be used to update the generative flow model. For example, this positive impact can be used as the reward value of the terminating node among the nodes selected by the generative flow model.

[0247] In one possible implementation, the task network is used to predict the labels of the selected nodes based on the data from those nodes. The positive impact on model accuracy can be represented by the model's accuracy in label prediction (or, more accurately, classification accuracy).

[0248] For example, given a sequence of node labels of length b, calculate the reward; the reward is estimated by calculating the classification accuracy of the GNN. Here, the reward is defined as the trajectory of node labels of length b, used as the validation set.

[0249]

[0250] After calculating the output of the GNN model, the accuracy of the model output is then calculated, and this accuracy is used as the reward.

[0251] Calculate the flow of all parent nodes and all child nodes, where inflow equals outflow, use the trajectory balancing loss function, and then update the policy network; continue until the end of the epoch.

[0252]

[0253] In one possible implementation, the partial nodes include intermediate nodes and termination nodes; the loss specifically includes a first loss corresponding to the intermediate node and a second loss corresponding to the termination node; the first loss is the difference between the first input flow and the output flow of the intermediate node; the second loss is the difference between the second input flow and the reward value of the termination node; the reward value is related to one of the following: the positive contribution to the update of the task network when training the task network based on the data of the partial nodes, or the similarity between the data of the partial nodes and the joint distribution of the data of the multiple nodes.

[0254] The following example uses a GNN as the task network to illustrate active learning of a graph neural network based on a generative flow model:

[0255] All parent nodes are labeled, and their state matrices are calculated. This continues until all b nodes are labeled. A classification model based on a generative flow model is trained using labeled samples. The goal of active learning in a graph neural network is to train action probabilities that follow a policy of true reward allocation, specifically modeled as... Figure 9 The left side shows a directed acyclic graph (DAG). Dots correspond to states, and edges between dots represent state transitions. The initial state is represented by white dots, intermediate states by light gray dots, and the final state by dark gray dots. Given a state, the policy network calculates the unnormalized action probability in each layer, corresponding to gray dots on the gray edges, which we can consider as water pipes. The number of gray dots represents the water flow, equivalent to the action probability. After layer calculation, each label set is evaluated according to the reward function. The training process of the policy network is as follows... Figure 9 As shown in the right half.

[0256] For example, during algorithm execution, multiple epochs of step size can be looped. During the loop budget(b) times, b nodes can be selected and labeled (ground truth). Given a batch of graph data, its state is passed through a policy network (e.g., a graph convolutional network GCN) to output the action selection probabilities for all graph nodes. The action with the highest probability is selected, and node v is chosen. i Mark its feature 5 as 1 and transition to the next state. Update the classification model of the next-step GNN based on the newly labeled node. Find a GNN of type 1 based on the current node until it converges.

[0257] The following describes the process of a generative flow model selecting nodes sequentially from multiple nodes:

[0258] In one possible implementation, the generative flow model may include a first neural network, a second neural network, a third neural network, etc.

[0259] In one possible implementation, the first, second, and third neural networks can be multi-layer perceptrons (MLPs) or convolutional models. An MLP is a feedforward artificial neural network model that maps multiple input datasets to a single output dataset.

[0260] In one possible implementation, the generative stream model may also include an embedding layer that can embed data from multiple nodes to obtain corresponding embedded representations.

[0261] In this process, the first neural network selects the next node based on the data of the nodes already selected in one iteration. For example, the "data of the nodes already selected" can be the first information. The first neural network can predict the second information (the second information has one or more selected nodes compared to the first information) based on the first information (for example, the first information (or the embedded representation corresponding to the first information) is input into the first neural network).

[0262] As can be seen, the generation flow model in this application embodiment does not directly sample some nodes from multiple nodes, but generates multiple nodes sequentially in each iteration through a sequence generation method. Each generation is based on the generation result of the previous generation, and the number of selected nodes increases continuously as the node selection process progresses.

[0263] For active domain adaptation in generative flow networks, the node data can be referred to the introduction of node data in active learning above, and will not be repeated here.

[0264] Similar to active learning, generative flow models can select a subset of nodes from multiple nodes. However, the data from these subset of nodes does not need to be used to train the task network. Instead, the loss is constructed directly based on the data from the selected subset of nodes and the relationships between the multiple data points. These relationships can be related to the similarity of the joint distribution between the data from the subset of nodes and the data from the multiple nodes.

[0265] like Figure 10 As shown, in implementing active domain adaptation based on a generative flow network, the source domain data and unlabeled target domain graph data are first selected and input into GflowNet. The generative flow policy network (MLP) selects an action, namely, selecting an unlabeled target domain data. After selection, the sample label is obtained, and then the state node is updated to the next step. The input flow is matched with the output flow. The loop continues until the number of labeled samples meets the budget, which indicates the end of the generative flow. At this time, the flow reward is calculated. The similarity between the source domain graph data and the target domain graph data is calculated as the generative flow trajectory reward. Then, the loss is calculated using the trajectory balancing loss, and the policy network is updated. We can then train a classification GNN model using labeled source domain target graph data until it converges.

[0266] For example, the specific steps may include the following:

[0267] Step 1: Prepare graph data for the source and target domains, and use this data to pre-train the graph neural network model.

[0268] Step 1.1: The state of the data is represented by the similarity relationship between the graph data of the source and target domains, and whether the target domain data is labeled, where Z is the target representation matrix, and μ... t The mixed source domain matrix is ​​represented as follows:

[0269]

[0270] Step 2: Input into GFlowNet, use the policy network to select unlabeled target domain data. After selecting this node, use the GNN classification network to classify the sample label, mark the already labeled state change, and update its next state.

[0271] Step 3: Repeat step 2 to continue labeling the other unlabeled data until the number of labeled samples reaches the budget number, and then generate a complete trajectory.

[0272] Step 4: Use the conditional distribution similarity between the labeled target domain data and all target domain data as the GFlowNet reward:

[0273]

[0274] Step 5: Here, we use trajectory balancing loss, which, for a single state node, makes all input streams equal to the output stream, to calculate the GFlowNet loss:

[0275]

[0276] Step 6: Update the policy network of GFlowNet using this loss.

[0277] Step 7: The outer layer then trains the graph neural network (GNN) using the already labeled target domain data until it converges. Repeat the above steps until the training is complete.

[0278] For identifying the correlation between nodes, the node data can be obtained by extracting features from the original information of the nodes (such as the attribute information of the nodes mentioned above, which is data that has not been processed by the feature extraction network). Optionally, the feature extraction network can be a GNN.

[0279] For example, a starting node can be sampled and input into the policy network to select a neighboring node as the action. Graph data consists of multiple nodes, and there may be relationships between different nodes or between subgraphs (which may also contain multiple nodes), such as causal relationships. A generative flow model can be used to identify which nodes in the graph data have a high correlation with a specified node, or which nodes in the graph data have a high correlation with a specified subgraph. The starting node can be either the specified node or a node in the specified subgraph.

[0280] In one possible implementation, the two values ​​can be compared with their original feature vector x. i Connect the nodes and enhance their features to obtain the initial feature representation x. t :

[0281]

[0282] Considering the relationships between nodes in graph-structured data, combining information from its neighbors is crucial for each node. Here, nonlinear transformation and information propagation are separated, as shown in the following formula:

[0283]

[0284] Output via MLP:

[0285]

[0286] Finally, the softmax function is used to calculate the action selection:

[0287]

[0288] When constructing the loss, the degree of association may be related to the accuracy of predicting the label of the target node or the node in the target subnet through the task network based on the data of the partial nodes; or, the degree of association may be related to the similarity between the partial nodes and the corresponding ground truth, where the ground truth consists of multiple nodes.

[0289] For example, the following steps may be included:

[0290] The GNN model is used to predict labels for data after state transitions.

[0291] Based on the predicted and true labels, the cross-entropy function is used instead of the conditional entropy function, and the reward function is defined as follows:

[0292]

[0293] The sampling nodes are repeated until a stopping condition is triggered. Condition one is to stop the loop when the selected number of nodes is the budgeted amount. Condition two is to use a self-attention mechanism, as shown in the following formula:

[0294]

[0295] Finally, a method similar to time difference is used to optimize the target flow matching function.

[0296]

[0297] A specific process can be as follows: Figure 11 As shown.

[0298] The following describes a software architecture illustration from an embodiment of this application, such as... Figure 12 As shown, an architecture is proposed that uses a graph neural network (GNN) in a generative flow model to predict labels for each state node. For a set of input datasets, where b state nodes are pre-defined, the generative flow model network will generate a series of directed acyclic graphs (DAGs) with b state nodes. The neural network in GFlowNets and the pre-defined reward function jointly calculate the flowmatching loss to optimize the network, thereby learning a strategy that makes it more likely that DAGs with high reward values ​​will be sampled and generated. The DAG with the largest reward value is the generative flow relationship required to best fit the dataset.

[0299] In this embodiment of the application, there is a directed acyclic graph constraint: the generation flow model itself is a directed acyclic graph structure, that is, a state node has multiple parent nodes, such as... Figure 7 In the middle s3, there are two parent nodes, s1 and s2. However, the state node s in the generated flow model... i A graph can be either a cyclic graph or an acyclic graph, which can represent a graph data structure.

[0300] This application also provides a data processing method, which can be based on the above... Figure 8 The model inference process performed by the corresponding embodiment's generated flow model includes:

[0301] Based on data from multiple nodes, a flow model is generated to select a subset of nodes from the multiple nodes;

[0302] The data from a subset of nodes is used as training samples for the task network; or,

[0303] The selected nodes are those nodes that best represent the characteristics of the plurality of nodes; or,

[0304] The partial nodes are the nodes with the highest correlation to the target node or target subnet among the plurality of nodes; the target node is one of the plurality of nodes, and the target subnet is composed of a partial number of the plurality of nodes; the target node is the starting node when the generation flow model selects the partial nodes, and a node in the target subnet is the starting node when the generation flow model selects the partial nodes.

[0305] In one possible implementation, the task network is a graph neural network, and the plurality of nodes are nodes in the graph information.

[0306] In one possible implementation, the node's data includes information about the node and information about the connections between the node and other nodes.

[0307] In one possible implementation, the task network is used to predict the labels of the partial nodes based on the data of the partial nodes.

[0308] In one possible implementation, the step of selecting a subset of nodes from the multiple nodes by generating a flow model based on data from multiple nodes includes:

[0309] By generating a stream model, data from multiple nodes is processed, and nodes are selected sequentially from the multiple nodes until the number of selected nodes reaches a preset value, thus obtaining the partial nodes.

[0310] Reference Figure 13 , Figure 13A schematic diagram of a data processing apparatus provided in this application embodiment, wherein the apparatus 1300 includes:

[0311] Processing module 1301 is used to select a subset of nodes from the multiple nodes by generating a flow model based on the data from multiple nodes;

[0312] The specific description of the processing module 1301 can be found in the description of step 801 in the above embodiments, and will not be repeated here.

[0313] Update module 1302 is used to obtain a loss based on the data of the partial nodes, and the loss is used to update the generated flow model; wherein the loss is related to one of the following:

[0314] The positive contribution of the task network to the update of the task network when training the task network based on the data of the partial nodes, the similarity of the joint distribution between the data of the partial nodes and the data of the multiple nodes, and the degree of correlation of the data of the partial nodes with the target node or the target subnet, wherein the target node is one of the multiple nodes and the target subnet is composed of some of the multiple nodes.

[0315] The specific description of the update module 1302 can be found in the description of step 802 in the above embodiment, and will not be repeated here.

[0316] In one possible implementation, the task network is a graph neural network, and the plurality of nodes are nodes in the graph information.

[0317] In one possible implementation, the node's data includes information about the node and information about the connections between the node and other nodes.

[0318] In one possible implementation, the positive contribution is positively correlated with the processing accuracy of the updated task network obtained after training the task network with data from the partial nodes.

[0319] In one possible implementation, the degree of correlation is related to the accuracy of predicting the labels of the target node or nodes in the target subnet using the task network based on the data from the partial nodes; or,

[0320] The degree of association is related to the similarity between the partial nodes and their corresponding truth values, where the truth values ​​are multiple nodes.

[0321] In one possible implementation, the partial nodes include intermediate nodes and termination nodes; the loss specifically includes a first loss corresponding to the intermediate node and a second loss corresponding to the termination node;

[0322] The first loss is the difference between the first input flow and the first output flow of the intermediate node;

[0323] The second loss is the difference between the second input flow and the reward value of the termination node; the reward value is related to one of the following:

[0324] The positive contribution of the task network to the update of the task network when training the task network based on the data of the partial nodes, or the similarity of the joint distribution between the data of the partial nodes and the data of the multiple nodes.

[0325] In one possible implementation, the task network is used to predict the labels of the partial nodes based on the data of the partial nodes.

[0326] In one possible implementation, the processing module is specifically used for:

[0327] By generating a stream model, data from multiple nodes is processed, and nodes are selected sequentially from the multiple nodes until the number of selected nodes reaches a preset value, thus obtaining the partial nodes.

[0328] This application embodiment also provides a data processing apparatus, the apparatus comprising:

[0329] The processing module is used to select a subset of nodes from the multiple nodes by generating a flow model based on the data from the multiple nodes;

[0330] The data from a subset of nodes is used as training samples for the task network; or,

[0331] The selected nodes are those nodes that best represent the characteristics of the plurality of nodes; or,

[0332] The partial nodes are the nodes with the highest correlation to the target node or target subnet among the plurality of nodes; the target node is one of the plurality of nodes, and the target subnet is composed of a partial number of the plurality of nodes; the target node is the starting node when the generation flow model selects the partial nodes, and a node in the target subnet is the starting node when the generation flow model selects the partial nodes.

[0333] In one possible implementation, the task network is a graph neural network, and the plurality of nodes are nodes in the graph information.

[0334] In one possible implementation, the node's data includes information about the node and information about the connections between the node and other nodes.

[0335] In one possible implementation, the task network is used to predict the labels of the partial nodes based on the data of the partial nodes.

[0336] In one possible implementation, the processing module is specifically used for:

[0337] By generating a stream model, data from multiple nodes is processed, and nodes are selected sequentially from the multiple nodes until the number of selected nodes reaches a preset value, thus obtaining the partial nodes.

[0338] The following describes an execution device provided in an embodiment of this application. Please refer to [link / reference]. Figure 14 , Figure 14 This is a schematic diagram of an execution device provided in an embodiment of this application. The execution device 1400 can specifically be a mobile phone, tablet, laptop, smart wearable device, etc., and is not limited thereto. Specifically, the execution device 1400 includes: a receiver 1401, a transmitter 1402, a processor 1403, and a memory 1404 (wherein the execution device 1400 may have one or more processors 1403). Figure 14 (Taking a processor as an example), processor 1403 may include application processor 14031 and communication processor 14032. In some embodiments of this application, receiver 1401, transmitter 1402, processor 1403 and memory 1404 may be connected via a bus or other means.

[0339] Memory 1404 may include read-only memory and random access memory, and provides instructions and data to processor 1403. A portion of memory 1404 may also include non-volatile random access memory (NVRAM). Memory 1404 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.

[0340] Processor 1403 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together through a bus system, which may include not only the data bus, but also power buses, control buses, and status signal buses. However, for clarity, all buses are referred to as the bus system in the diagram.

[0341] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 1403. The processor 1403 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 1403 or by instructions in software form. The processor 1403 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1403 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 1404. Processor 1403 reads the information in memory 1404 and, in conjunction with its hardware, completes the steps of the above method.

[0342] Receiver 1401 can be used to receive input digital or character information, and to generate signal inputs related to the settings and function control of the execution device. Transmitter 1402 can be used to output digital or character information; transmitter 1402 can also be used to send instructions to the disk group to modify the data in the disk group.

[0343] In one embodiment of this application, the processor 1403 is configured to execute according to Figure 8 The model inference steps of the generated stream model obtained by the data processing method in the corresponding embodiment.

[0344] This application also provides a server; please refer to [link / reference]. Figure 15 , Figure 15This is a schematic diagram of a server structure provided in an embodiment of this application. Specifically, server 1500 is implemented by one or more servers. Server 1500 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1515 (e.g., one or more processors) and memory 1532, and one or more storage media 1530 (e.g., one or more mass storage devices) for storing application programs 1542 or data 1544. The memory 1532 and storage media 1530 can be temporary or persistent storage. The program stored in storage media 1530 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the server. Furthermore, the CPU 1515 may be configured to communicate with storage media 1530 and execute the series of instruction operations in storage media 1530 on server 1500.

[0345] Server 1500 may also include one or more power supplies 1510, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558; or, one or more operating systems 1541, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0346] In this embodiment, the central processing unit 1515 is used to execute... Figure 8 The steps of the data processing method in the corresponding embodiment.

[0347] This application also provides a computer program product including computer-readable instructions, which, when run on a computer, causes the computer to perform steps as performed by the aforementioned execution device, or causes the computer to perform steps as performed by the aforementioned training device.

[0348] This application also provides a computer-readable storage medium storing a program for signal processing, which, when run on a computer, causes the computer to perform steps as performed by the aforementioned execution device, or causes the computer to perform steps as performed by the aforementioned training device.

[0349] The execution device, training device, or terminal device provided in this application embodiment can specifically be a chip. The chip includes a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, pins, or circuits. The processing unit can execute computer execution instructions stored in the storage unit to cause the chip within the execution device to execute the data processing method described in the above embodiments, or to cause the chip within the training device to execute the steps related to model training in the above embodiments. Optionally, the storage unit can be a storage unit within the chip, such as a register or cache. Alternatively, the storage unit can be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).

[0350] For details, please refer to Figure 16 , Figure 16 This is a schematic diagram of a chip provided in an embodiment of this application. The chip can be represented as a neural network processor (NPU) 1600. The NPU 1600 is mounted as a coprocessor on the host CPU, and tasks are assigned by the host CPU. The core part of the NPU is the arithmetic circuit 1603, which is controlled by the controller 1604 to extract matrix data from the memory and perform multiplication operations.

[0351] In some implementations, the arithmetic circuit 1603 internally includes multiple processing engines (PEs). In some implementations, the arithmetic circuit 1603 is a two-dimensional pulsating array. The arithmetic circuit 1603 can also be a one-dimensional pulsating array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1603 is a general-purpose matrix processor.

[0352] For example, suppose we have an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from the weight memory 1602 and caches it in each PE of the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from the input memory 1601 and performs matrix operations with matrix B. The partial result or the final result of the obtained matrix is ​​stored in the accumulator 1608.

[0353] Unified memory 1606 is used to store input and output data. Weight data is directly transferred to weight memory 1602 via Direct Memory Access Controller (DMAC) 1605. Input data is also transferred to unified memory 1606 via DMAC.

[0354] BIU stands for Bus Interface Unit, which is used for interaction between the AXI bus and the DMAC and the Instruction Fetch Buffer (IFB) 1609.

[0355] The Bus Interface Unit (BIU) 1610 is used by the instruction fetch memory 1609 to fetch instructions from external memory, and also by the memory access controller 1605 to fetch the original data of the input matrix A or the weight matrix B from external memory.

[0356] The DMAC is mainly used to move input data from external memory DDR to unified memory 1606, or to weight data to weight memory 1602, or to input data to input memory 1601.

[0357] The vector computation unit 1607 includes multiple arithmetic processing units that further process the output of the computation circuit as needed, such as vector multiplication, vector addition, exponential operations, logarithmic operations, size comparisons, etc. It is mainly used for computation in non-convolutional / fully connected layers of neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.

[0358] In some implementations, the vector computation unit 1607 can store the processed output vector in the unified memory 1606. For example, the vector computation unit 1607 can apply a linear function, or a nonlinear function, to the output of the computation circuit 1603, such as linear interpolation of feature planes extracted by a convolutional layer, or, for example, a vector of accumulated values, to generate activation values. In some implementations, the vector computation unit 1607 generates normalized values, pixel-level summed values, or both. In some implementations, the processed output vector can be used as activation input to the computation circuit 1603, for example, for use in subsequent layers of the neural network.

[0359] The instruction fetch buffer 1609 connected to the controller 1604 is used to store the instructions used by the controller 1604.

[0360] The unified memory 1606, input memory 1601, weight memory 1602, and instruction fetch memory 1609 are all on-chip memories. External memory is proprietary to this NPU hardware architecture.

[0361] The processor mentioned above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of the above program.

[0362] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0363] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0364] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0365] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

Claims

1. A data processing method, characterized in that, The method includes: Based on data from multiple nodes, a flow model is generated to select a subset of nodes from the multiple nodes; Based on the data from the aforementioned nodes, a loss is obtained, which is used to update the generated flow model; wherein, the loss is related to one of the following: The positive contribution of the task network update to the training of the task network based on the data of the partial nodes, the similarity of the joint distribution between the data of the partial nodes and the data of the multiple nodes, and the degree of correlation of the data of the partial nodes to the target node or the target subnet, wherein the target node is one of the multiple nodes and the target subnet is composed of some of the multiple nodes; The node is a chip, and the data of the node is a partial fragment of the chip or fault information; or, the node is a node of a communication network, and the data of the node is the key performance indicators (KPIs) of the network element, operating data, or alarm information; or, the node is a user, and the data of the node is the attribute information of the item and the attribute information of the user. The nodes include intermediate nodes and termination nodes; the losses specifically include a first loss corresponding to the intermediate nodes and a second loss corresponding to the termination nodes. The first loss is the difference between the first input flow and the first output flow of the intermediate node; The second loss is the difference between the second input flow and the reward value of the termination node; the reward value is related to one of the following: The positive contribution of the task network to the update of the task network when training the task network based on the data of the partial nodes, or the similarity of the joint distribution between the data of the partial nodes and the data of the multiple nodes.

2. The method according to claim 1, characterized in that, The task network is a graph neural network, and the multiple nodes are nodes in the graph information.

3. The method according to claim 1, characterized in that, The data of the node includes the node's information and the connection relationship information between the node and other nodes.

4. The method according to claim 1, characterized in that, The positive contribution is positively correlated with the processing accuracy of the updated task network obtained after training the task network with data from the aforementioned nodes.

5. The method according to claim 1, characterized in that, The degree of correlation is related to the accuracy of predicting the labels of the target node or nodes in the target subnet using the task network based on the data from the partial nodes; or, The degree of association is related to the similarity between the partial nodes and their corresponding truth values, where the truth values ​​are multiple nodes.

6. The method according to claim 1, characterized in that, The task network is used to predict the labels of the selected nodes based on the data from those nodes.

7. The method according to any one of claims 1 to 6, characterized in that, The process of generating a streaming model, processing data from multiple nodes, and selecting a subset of nodes from those nodes includes: By generating a stream model, data from multiple nodes is processed, and nodes are selected sequentially from the multiple nodes until the number of selected nodes reaches a preset value, thus obtaining the partial nodes.

8. A data processing apparatus, characterized in that, The device includes: The processing module is used to select a subset of nodes from the multiple nodes by generating a flow model based on the data from the multiple nodes; The update module is used to obtain a loss based on the data of the partial nodes, and the loss is used to update the generated flow model; wherein the loss is related to one of the following: The positive contribution of the task network update to the training of the task network based on the data of the partial nodes, the similarity of the joint distribution between the data of the partial nodes and the data of the multiple nodes, and the degree of correlation of the data of the partial nodes to the target node or the target subnet, wherein the target node is one of the multiple nodes and the target subnet is composed of some of the multiple nodes; The node is a chip, and the data of the node is a partial fragment of the chip or fault information; or, the node is a node of a communication network, and the data of the node is the key performance indicators (KPIs) of the network element, operating data, or alarm information; or, the node is a user, and the data of the node is the attribute information of the item and the attribute information of the user. The nodes include intermediate nodes and termination nodes; the losses specifically include a first loss corresponding to the intermediate nodes and a second loss corresponding to the termination nodes. The first loss is the difference between the first input flow and the first output flow of the intermediate node; The second loss is the difference between the second input flow and the reward value of the termination node; the reward value is related to one of the following: The positive contribution of the task network to the update of the task network when training the task network based on the data of the partial nodes, or the similarity of the joint distribution between the data of the partial nodes and the data of the multiple nodes.

9. The apparatus according to claim 8, characterized in that, The task network is a graph neural network, and the multiple nodes are nodes in the graph information.

10. The apparatus according to claim 8, characterized in that, The data of the node includes the node's information and the connection relationship information between the node and other nodes.

11. The apparatus according to claim 8, characterized in that, The positive contribution is positively correlated with the processing accuracy of the updated task network obtained after training the task network with data from the aforementioned nodes.

12. The apparatus according to claim 8, characterized in that, The degree of correlation is related to the accuracy of predicting the labels of the target node or nodes in the target subnet using the task network based on the data from the partial nodes; or, The degree of association is related to the similarity between the partial nodes and their corresponding truth values, where the truth values ​​are multiple nodes.

13. The apparatus according to claim 8, characterized in that, The task network is used to predict the labels of the selected nodes based on the data from those nodes.

14. The apparatus according to any one of claims 8 to 13, characterized in that, The processing module is specifically used for: By generating a stream model, data from multiple nodes is processed, and nodes are selected sequentially from the multiple nodes until the number of selected nodes reaches a preset value, thus obtaining the partial nodes.

15. A data processing apparatus, characterized in that, The device includes a memory and a processor; the memory stores code, and the processor is configured to retrieve the code and perform the method as described in any one of claims 1 to 7.

16. A computer-readable storage medium, characterized in that, Includes computer-readable instructions that, when executed on a computer device, cause the computer device to perform the method of any one of claims 1 to 7.

17. A computer program product, characterized in that, Includes computer-readable instructions that, when executed on a computer device, cause the computer device to perform the method as described in any one of claims 1 to 7.