Method and apparatus for managing data transmission across multiple regions, device and medium
By processing data transmission within a region and providing aggregated messages when crossing regions, a cross-regional transmission graph is generated, which solves the problem of high resource overhead in cross-regional data transmission management and achieves efficient and secure data lineage management.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2024-11-01
- Publication Date
- 2026-05-07
AI Technical Summary
Existing data transmission management technologies suffer from high resource consumption and low management efficiency when spanning multiple regions, especially lacking effective technical solutions in online lineage management.
By processing data transmission within a region and providing aggregated messages during cross-regional transmission, the aggregation device handles only cross-regional traffic, reducing data volume and managing transmission in a distributed manner, generating a cross-regional transmission graph.
It effectively reduced data transmission volume, balanced the workload of computing devices in different regions, improved the efficiency and security of data lineage management, and reduced network bandwidth and storage costs.
Smart Images

Figure CN2024129457_07052026_PF_FP_ABST
Abstract
Description
Method, apparatus, device and medium for managing data transmission across multiple regions TECHNICAL FIELD
[0001] Exemplary implementations of the present disclosure generally relate to data management, and in particular, to a method, apparatus, device and computer-readable storage medium for managing data transmission across multiple regions. BACKGROUND
[0002] In a computer environment, various types of data can flow between different locations and / or processing processes to achieve an intended task. Data lineage refers to the relationships of data in its entire life cycle, such as the origin, flow and transformation of data. Data lineage records the data propagation path and data processing process of data from collection, transmission, storage, analysis. Through data lineage, the origin, history and credibility of data can be better understood, and data can be processed more safely and efficiently. However, the data transmission process can span multiple regions (e.g., multiple data centers), which results in a process of managing data transmission across multiple regions needing to handle a huge number of data records, thereby generating a huge resource overhead. The performance of existing data transmission management technical solutions is not satisfactory, and it is desirable to improve the management efficiency of data lineage.
[0003] SUMMARY
[0004] In a first aspect of the present disclosure, a method for managing data transmission is provided. In the method, a first aggregation message from a first computing device is received, the first aggregation message indicating that data is transmitted from a first service of a first group of services to a second service of a second group of services, the first group of services being provided by a first group of devices located in a first region, and the second group of services being provided by a second group of devices located in a second region. Based on the first aggregation message, a transmission graph indicating the transmission of data between the plurality of services is determined.
[0005] In a second aspect of the present disclosure, an apparatus for managing data transmission is provided. The apparatus comprises: a receiving module configured to receive a first aggregation message from a first computing device, the first aggregation message indicating that data is transmitted from a first service of a first group of services to a second service of a second group of services, the first group of services being provided by a first group of devices located in a first region, and the second group of services being provided by a second group of devices located in a second region; and a determining module configured to determine, based on the first aggregation message, a transmission graph indicating the transmission of data between the plurality of services.
[0006] In a third aspect of the disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the electronic device to perform the method according to the first aspect of the disclosure.
[0007] In a fourth aspect of the disclosure, a computer-readable storage medium is provided, having stored thereon a computer program which, when executed by a processor, causes the processor to implement the method according to the first aspect of the disclosure.
[0008] In a fifth aspect of the disclosure, a computer program product is provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to the first aspect of the disclosure.
[0009] It is to be understood that the particulars shown herein are by way of example and for purposes of illustrative discussion of the various embodiments of the present disclosure only and are not intended to limit the scope of the present disclosure to the particular embodiment illustrated. Other BRIEF DESCRIPTION OF DRAWINGS
[0010] In the following detailed description, reference will be made to the accompanying drawings, of which:
[0011] FIG. 1 shows a block diagram of data transmission according to one example implementation of the present disclosure;
[0012] FIG. 2 shows a block diagram for managing data transmission according to some implementations of the present disclosure;
[0013] FIG. 3 shows a block diagram of online lineage according to some implementations of the present disclosure;
[0014] FIG. 4 shows a block diagram of data lineage according to some implementations of the present disclosure;
[0015] FIG. 5 shows a block diagram of a transmission graph of aggregated data according to some implementations of the present disclosure;
[0016] FIG. 6 shows a block diagram of a transmission graph of data according to some implementations of the present disclosure;
[0017] FIG. 7 shows a block diagram of a transmission graph of data according to some implementations of the present disclosure;
[0018] FIG. 8 shows a block diagram of a transmission graph of data across multiple regions according to some implementations of the present disclosure;
[0019] Figure 9 shows a block diagram for managing data transmission according to some implementations of this disclosure;
[0020] Figure 10 shows a flowchart of a method for managing data transmission according to some implementations of this disclosure;
[0021] Figure 11 shows a block diagram of an apparatus for managing data transmission according to some implementations of the present disclosure; and
[0022] Figure 12 shows a block diagram of a device capable of implementing various implementations of the present disclosure. Detailed Implementation
[0023] Implementations of this disclosure will now be described in more detail with reference to the accompanying drawings. While some implementations of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. Rather, these implementations are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and implementations of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0024] In the description of the implementation methods disclosed herein, the term "comprising" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the relationships between various data. For example, the aforementioned relationships can be obtained based on various currently known and / or future-developed technical solutions.
[0025] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0026] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.
[0027] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0028] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.
[0029] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0030] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of subsequent actions performed in response to such event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the occurrence of the event or the fulfillment of the condition.
[0031] Example Environment
[0032] First, the application environment of some implementation methods according to this disclosure is described. In the field of data transmission management, distributed tracing technology has been proposed to manage and analyze request paths and performance in distributed systems. This technical solution can help developers and maintenance personnel understand the flow of requests between various services, identify performance bottlenecks and fault points, etc. Table 1 shows the explanation of key terms in data transmission management.
[0033] Table 1. Terminology Explanation
[0034] Referring to Figure 1, which illustrates a block diagram 100 of data transmission according to an exemplary implementation of this disclosure, the specific meanings of the terms are described. The left side of Figure 1 shows a service call chain 102, and the right side shows the corresponding span of this chain. Nodes A, B, C, and D in the figure correspond to multiple services A, B, C, and D, respectively, and the edges represent the call relationships between services. As shown in Figure 1, the request corresponds to a globally unique TraceID, which is used to associate each sub-call with the initial request to form a complete chain. SpanIDs correspond to 0, 1, 2, and 1.1 on the left side of Figure 1 to represent specific sub-calls. ParentSpanID represents the parent span; for example, if the SpanID of service B calling service D is 1.1, then the ParentSpanID is the SpanID of service A calling service B (SpanID = 1), thus representing two adjacent calls. The right side of Figure 1 shows the corresponding time periods for each span.
[0035] The following describes how link management works. (1) Request enters the system: When a request enters the system, a TraceID and a root SpanID are generated. The root span represents the origin of the request (e.g., an API gateway or frontend service, etc.). (2) Request propagates between services: When a service processes a request, it generates a new SpanID and sets the ParentSpanID to the caller's SpanID. The TraceID and SpanID are propagated between services via HTTP headers or other protocols. (3) Data collection and storage: Services send the generated span data to the management system for collection, storage, and indexing. (4) Link data analysis and visualization: The management system provides query and visualization interfaces that allow users to view the complete link of a request and detailed information about each span. Users can analyze information such as the request path, latency, and errors to identify performance bottlenecks and fault points.
[0036] When processing a request, the service can generate new spans. Table 2 below shows the data structure of each span associated with the request. This data structure can include multiple rows, each storing a record corresponding to a span.
[0037] Table 2 Data Structure of Span
[0038] According to some implementations of this disclosure, the span may also include other information: start_time and end_time represent the start and end times of the span, respectively, to help identify performance bottlenecks and failure points. service represents the name and / or identifier of the service, and idc represents the region where the service is located (e.g., Internet Data Center (IDC)). tag represents metadata information related to the span; different services may have different metadata. Metadata may include, for example, the entry point http_path and http_method for HTTP services; and SQL information during MySQL service execution. Optionally and / or additionally, user-defined information may also be stored. It should be understood that Table 2 above is merely illustrative and may contain more, fewer, and / or different fields. For example, span data may further include method / interface information to construct method / interface-level call relationships.
[0039] In the context of this disclosure, a call (e.g., service B calling service D) can generate two spans: a client-type span at service B and a service-type span at service D. The data structure described above can be used to generate transport graphs that describe complex data transfer relationships.
[0040] In the context of this disclosure, data lineage can be used to describe data transmission. Specifically, data lineage represents the relationships of data throughout its lifecycle, including its source, flow, and transformation. Data lineage records the data propagation path and processing procedures from acquisition, transmission, storage, and analysis. Through data lineage, we can better understand the source, history, and reliability of data, and process data more securely and efficiently.
[0041] The data lineage lifecycle can include multiple stages. For example, the client-side stage includes the collection, storage, and transmission of data on the client-side. This stage may involve security information, such as sensitive data like phone numbers, and is typically a single line of data. This stage's lineage can be called the client lineage or mobile lineage. The online stage includes the entire computation and transmission process from the gateway to the storage of data in an online database (MySQL / Redis / ...), and the security information involved is usually a single line of data. This stage's lineage can be called the online lineage. The offline stage includes the process of data moving from the online database to the offline database and performing various computational tasks between the offline databases, and the security information involved is usually batch data. This stage's lineage is called the offline lineage.
[0042] The complete lifecycle of data can include, for example, the data moving from end nodes (such as apps and web pages) through gateway nodes to backend online services, then being stored in a database (such as MySQL, possibly with intermediate caching such as Redis), and finally being queried by the backend online services. Furthermore, the data stored in the database can be synchronized to offline databases (such as Hive) via certain synchronization nodes for data analysis (or used as training data, etc.), and finally presented to users or fed back to the online services to improve processing capabilities.
[0043] Data transmission management has various application scenarios. Firstly, data lineage plays a crucial role in data management (e.g., security management). It helps organizations understand the source and destination of data, trace the history of data changes, ensure data quality and reliability, and improve the efficiency and effectiveness of data management. Data management may face various challenges, such as invisible data flow, uncontrollable risks, and unknown impacts of data updates. Data lineage management can address these issues. Through data lineage, users can identify the source and evolution of data to ensure data quality and consistency. Furthermore, regarding data security requirements, data lineage can assist in identifying the propagation of secure data, tracking data usage records, and processing procedures.
[0044] However, data transmission may span multiple regions, leading to a massive amount of data records to be processed and resulting in significant resource overhead when managing data transmission across multiple regions based on data lineage. Existing data transmission management technologies have unsatisfactory performance, thus necessitating improvements in data lineage management efficiency.
[0045] Overview of data transfer management
[0046] To at least partially address the shortcomings of the prior art, a method for managing data transmission is proposed according to an exemplary implementation of this disclosure. In summary, to manage data transmission across multiple regions, data transmission within a region can be processed within that region, and an aggregation message can be provided when cross-region data transmission is determined to occur. For example, this aggregation message can be provided to an aggregation device that only relates to cross-region data transmission and not to data transmission within the region. This allows the aggregation device to handle only traffic related to cross-region transmissions, while non-cross-region data traffic is handled by computing devices within the region. In this way, data transmission volume can be reduced and data transmission can be managed in a distributed manner, thereby balancing the workload of individual devices.
[0047] Referring to Figure 2, which describes an outline of an exemplary implementation of the present disclosure, Figure 2 illustrates a block diagram 200 for managing data transmission according to some implementations of the present disclosure. As shown in Figure 2, multiple services may include a first set of services 210 and a second set of services 220. The first set of services may be provided by a first set of devices located in a first region (e.g., data center 1), and the second set of services may be provided by a second set of devices located in a second region (e.g., data center 2).
[0048] According to some implementations of this disclosure, an aggregation device can be provided, which can be located at any location, such as in a first region, a second region, or a location other than the first and second regions. The methods of this disclosure can be performed at the aggregation device. Specifically, a first aggregation message can be received from a first computing device, indicating that data is being transferred from a first service of a first group of services to a second service of a second group of services. For example, first aggregation message 230 can indicate that data is being transferred from service 214 to service 222. Further, a transmission graph representing the data transmission among the multiple services can be determined based on the first aggregation message 230.
[0049] As shown in Figure 2, edge 250 in transmission graph 240 can be determined based on the first aggregation message 230, subgraph 216 can be predetermined (e.g., determined by a first computing device in a first region), and subgraph 226 can be predetermined (e.g., determined by a second computing device in a second region). In this way, transmission graph 240 spanning multiple regions can be determined in a distributed manner, thereby balancing the workload of the various computing devices in the multiple regions.
[0050] Detailed process of managing data transmission
[0051] According to some implementations of this disclosure, online lineage and offline lineage are two important processes in data transmission management. For offline lineage, various solutions have been proposed, most of which are based on SQL parsing. However, for online lineage, no satisfactory technical solution exists. This disclosure mainly deals with cross-region data transmission in online lineage. Referring to Figure 3, which describes an overview of online lineage, Figure 3 shows a block diagram 300 of online lineage according to some implementations of this disclosure. As shown in Figure 3, legend 310 represents a regular node, and legend 320 represents a master node. Online lineage describes the calling relationships between services. For example, service B can call service A, service A can call service E, service A can call service D, and service D can insert data into service E (e.g., a storage service).
[0052] Typically, data lineage is described in the form of a graph, which can be called a lineage graph. The points in the graph are the nodes that the lineage focuses on, and the edges represent the relationships between two nodes, i.e., the lineage relationships. Online lineage focuses on the actual call relationships of traffic and has the concept of a master node. One node corresponds to one lineage graph, and this node is the master node in that lineage graph. See Figure 4 for more details, which shows a block diagram 400 of data lineage according to some implementations of this disclosure. As shown in the figure, node A can represent a service, with multiple services connected upstream and downstream. There can be four call paths: ① to ④. However, some paths do not actually exist.
[0053] Legend 410 represents a regular node, and Legend 420 represents a master node. The path BAD at the bottom of Figure 4 shows the lineage with node A as the master node. All four edges (①, ②, ③, ④) in the graph have actual traffic (path: ①→③, ②→④), meaning that for a master node in an online lineage graph, all other nodes have actual traffic relationships with the master node. Path CAE shows the lineage with node C as the master node. For master node C, there is a link ②→④, and its lineage graph will not contain B and D. Similarly, for master node B, there is a link ①→③, and its lineage graph will not contain C and E.
[0054] For online lineages, the number of nodes corresponds one-to-one with the number of lineage graphs; that is, the number of lineage graphs equals the number of nodes. When querying any node X in an online lineage, it's possible that node Y does not exist in X's lineage graph; this type of lineage can be classified as asymmetric lineage.
[0055] It should be understood that a TraceID corresponds to a service call chain, which is not equivalent to the lineage graph of online data. The lineage graph derived from the service call chain requires certain transformations and calculations. See Figure 5 for further details, which shows a block diagram 500 of the transport graph of aggregated data according to some implementations of this disclosure.
[0056] Figure 5 (right side) shows the links originating from node A, each link corresponding to a unique TraceID. However, these links do not represent the complete link originating from node A. Due to the different environments of each request (e.g., different input parameters leading to different call branches; rate limiting, circuit breaking, etc. causing partial call termination, etc.), multiple different links exist. Multiple different links can be aggregated to obtain the complete link originating from node A. During the aggregation process, a graph can be constructed based on the span data involved by different TraceIDs. As shown on the right side of Figure 5, link graphs corresponding to TraceID-1, TraceID-2, TraceID-3, and TraceID-4 can be constructed respectively. This graph can include a set of nodes and edges, and these nodes and edges can be deduplicated to obtain a complete aggregated link graph (as shown on the left side of Figure 5).
[0057] Based on some implementations of this disclosure, a similar approach can be used to manage data transmission across multiple regions, thereby generating a cross-regional link graph. For example, each link graph can include multiple nodes, and the reachable path of each node in this link is the lineage graph with this lineage point as the master node. All lineage graphs of a node can be aggregated to form a complete graph, which is the complete lineage graph of this node. The example in Figure 5 is relatively simple; the lineage subgraphs of all lineage points in the links corresponding to all TraceIDs on the right are the same. For example, for TraceID-1, the lineage subgraph of A is A->B->C, the lineage subgraph of B is A->B->C, and the lineage subgraph of C is also A->B->C.
[0058] In the following description, further details are given with reference to Figure 6, which shows a block diagram 600 of a data transfer graph according to some implementations of this disclosure. As shown on the left side of Figure 6, the lineage subgraph of node A is the entire left graph, the lineage subgraph of node C is A->B->C->D, and the lineage subgraph of node E is A->B->E->F->D. The same applies to other nodes: for each node in the graph, the reachable path in this graph is the lineage subgraph of that node.
[0059] For a given link, the lineage subgraphs corresponding to lineage nodes may be different. For example, a node may have more than one entry point for a call. As shown in Figure 6 above, node B's request entry point also includes node G. The lineage subgraph G->C->E of node B can be constructed using the right-hand diagram. Finally, the complete lineage graph of node B is an integrated graph that aggregates all call links passing through node B.
[0060] For large enterprises, data transmission spans vast distances and the structure of link data is complex; a single data center may generate terabytes of data per hour. Due to the massive data volume, span data similar to that shown in Table 2 is typically stored in offline databases, with a data partition generated every hour. In this case, link data is not stored long-term; it is generally only stored for a predetermined time window (e.g., a few hours, a few days, etc.) to perform link analysis. Typically, link data is stored within the data center and processed by computing devices within the data center. Therefore, the calculated lineage information is limited to within the data center. If a cross-regional request (e.g., across multiple data centers) occurs, it will not be reflected in the data lineage, resulting in incomplete data lineage.
[0061] Aggregation technologies have been proposed. For example, link data from all data centers can be synchronized to a single data center, and then a lineage graph can be constructed using the method described above. However, data transmission across data centers results in extremely high bandwidth and storage costs. Furthermore, due to factors such as data security, there is a potential risk of data leakage.
[0062] Unlike existing technical solutions, this disclosure provides a technical solution for managing data transmission across multiple regions. See Figure 7 for further details. Figure 7 shows a block diagram 700 of data transmission according to some implementations of this disclosure. As shown in Figure 7, blank nodes represent nodes in data center 1, and shaded nodes represent nodes in data center 2. Nodes E and F can be added to the link shown on the left side of Figure 1. In this case, nodes A, B, C, and D correspond to services in data center 1, and nodes E and F correspond to services in data center 2. Table 3 shows the span data in data center 2.
[0063] Table 3 Examples of span data
[0064] In cross-regional data transmission, if node C in data center 1 calls node E in data center 2, due to the established network connection, the existence of E can be observed on the C side; however, the downstream node F of E cannot be observed. At this point, a complete picture of the link is not available on either side. To achieve cross-regional data transmission management, the C->E information can be retained at data center 1 to form new span data. Assuming Table 2 shows the span data at data center 1, extended span data as shown in Table 4 can be formed based on Table 2. The last row of Table 4 includes the new span data across regions.
[0065] Table 4. Expansion span data for Data Center 1
[0066] Similarly, node C can be observed at node E in the data center, but relevant data from other upstream nodes of node C cannot be obtained. In this case, C->E can be stored at data center 2 to form new span data. Based on Table 3, the extended span data shown in Table 5 can be formed. The first row of Table 5 includes the new span data spanning different regions.
[0067] Table 5. Expansion span data for Data Center 2
[0068] It should be understood that even if data spans multiple data centers, the same request uses the same TraceID. In this case, all span data corresponding to the same TraceID can be identified. By aggregating the found span data, a complete link can be formed. Subgraphs of the link can be computed at different locations, thereby balancing the computational load. See Figure 8 for further details, which shows a block diagram 800 of a data transfer graph across multiple regions according to some implementations of this disclosure. As shown in Figure 8, node C in data center 1 can call node E in data center 2, and node E in data center 2 can receive calls from node C in data center 1. This constitutes cross-region data transfer.
[0069] According to some implementation methods of this disclosure, it is possible to base it on triples.<trace_id,subgraph,need_idcs> The aggregation message can be constructed in a specific way. For example, for a first region, the first aggregation message may include: an identifier of the data, a first transport subgraph indicating that the data is being transferred between the first group of services, and an identifier of a second region associated with the data transfer. Specifically, `trace_id` can represent the identifier of the data, i.e., the link that initiated the data transfer request. `subgraph` can represent the first transport subgraph of the data being transferred between the first group of services, for example, the transport subgraph of data in data center 1. `need_idcs` can represent the identifier of a second region associated with the data transfer, for example, the identifier of another region involved in cross-region data transfer.
[0070] In the example above, node C in data center 1 calls node E in data center 2, where another region is data center 2. In other words, another data center can represent a different partition from the region corresponding to the subgraph among multiple regions. Using some implementations of this disclosure, the information required to compute the cross-regional transport graph can be clearly described, thereby reducing the overhead of data transmission and data storage.
[0071] For a TraceID, the call chain graph within each IDC can be determined, and it can also be seen which other IDCs this graph has call relationships with, thus obtaining the "data center of the call chain subgraph to be supplemented". The above only illustrates the case of two data centers; alternative and / or additional locations can transmit data across more data centers. Assuming a TraceID exists in three data centers, for example, data center 1 -> data center 2 -> data center 3, the following data can be generated:
[0072] On the data center 1 side: This includes the TraceID, the corresponding subgraph 1 of the data center 1 link graph, and the list of IDCs to be supplemented (need_idcs = [data center 2]).<trace_id,subgraph1,need_idcs> Triplet.
[0073] On the data center 2 side: contains TraceID, the corresponding link subgraph2 on the data center 2 side, and data center need_idcs = [data center 1, data center 3] for information to be supplemented.
[0074] On the data center 3 side: contains TraceID, the corresponding link subgraph3 on the data center 3 side, and data center need_idcs = [data center 2].
[0075] According to some implementations of this disclosure, the first aggregated message can be generated by a first computing device, which can be deployed in a first area and can be a physical device and / or a virtual device. The first computing device can be a server in data center 1, etc. Specifically, the first computing device can receive a first plurality of records representing data transmissions, indicating the paths through which data is transmitted between multiple services. Here, the first plurality of records may be, for example, the extended span data of data center 1 shown in Table 4 above. A first transmission subgraph can be determined based on the first plurality of records in the manner described above.
[0076] In the context of this disclosure, multiple records can be continuously received, and spans from each data record can be continuously added to the current transport subgraph to generate the latest transport subgraph. According to some implementations of this disclosure, an initial transport subgraph associated with a first group of services can be obtained. In response to determining that the initial transport subgraph does not include nodes corresponding to services in the records, edges in the initial transport records are updated based on paths in the records, and nodes in the initial transport records are updated using services to determine the first transport subgraph.
[0077] In the initial stage, the initial transport subgraph can be empty. At this point, nodes and edges can be continuously added to the subgraph based on each record. For example, assuming the initial transport subgraph is empty, each record in Table 4 can be accessed one by one. Based on the first record, node A can be added to the subgraph; based on the second record, node B can be added to the subgraph, and an edge can be added between node A and node B. Further, based on the third record, node D can be added to the subgraph, and an edge can be added between node B and node D, and so on.
[0078] According to some implementations of this disclosure, in response to determining that the initial transport subgraph includes nodes corresponding to services in the records, the edges in the initial transport subgraph are updated based on the paths in the records to determine the first transport subgraph. Continuing the example above, assuming that nodes A and B already exist in the initial transport subgraph, the edge between nodes A and B can be updated based on the second record. Assuming that the edge includes the attribute "number" to represent the number of data transfers between the two nodes, the number can be incremented by one. Using some implementations of this disclosure, transport subgraphs related to each region can be determined separately within multiple regions. In this way, data transfer can be managed in a distributed manner, thereby balancing the workload of various computing devices.
[0079] According to some implementations of this disclosure, in response to determining that a first record among a first plurality of records indicates that data is transmitted from a first service to a second service, a first aggregation message is generated. Specifically, the first computing device can process each record in Table 4 one by one. When processing the last record, it can be determined that the record indicates cross-regional data transmission, that is, data is transmitted from node C in data center 1 to node E in data center 2. At this time, an aggregation message <123,subgraph01,idc02> can be generated. Here, "123" represents the TraceID of the data, subgraph01 represents the transmission subgraph of data center 1, and idc02 indicates that the cross-regional data transmission involves data center 2. According to some implementations of this disclosure, the data shown in Table 4 can be encapsulated to generate the aggregation message.
[0080] According to some implementations of this disclosure, the aggregation device can determine the transmission graph based on the aggregation message. Specifically, a first node in the transmission graph can be determined based on a first service, a second node in the transmission graph can be determined based on a second service, and an edge between the first node and the second node can be determined. Utilizing some implementations of this disclosure, the aggregation device can process only data transmissions crossing two region boundaries and generate edges crossing region boundaries, thereby reducing workload. Specifically, the aggregation device can determine node C corresponding to service C in data center 1 and node E corresponding to service E in data center 2 from subgraph01 based on the aggregation message, and add an edge between node C and node E.
[0081] According to some implementations of this disclosure, during the determination of the transport graph, the transport graph and the first transport subgraph can be aggregated to update the transport graph. Continuing the example above, the graph including C->E can be added to subgraph01 to determine the transport graph including the internal transport of data center 1 and the cross-area transport across data centers 1 and 2.
[0082] It should be understood that although the above description of processing within one region uses data center 1 as an example, records in other data centers can be processed in a similar manner. Alternatively and / or additionally, a second computing device in a second region (e.g., data center 2) can operate in a manner similar to the first computing device to manage data transfers within data center 02, as well as data transfers across data centers.
[0083] According to some implementations of this disclosure, the second computing device may receive a second plurality of records representing data transmission, the second plurality of records indicating the path of data transmission between multiple services (e.g., extended span data of data center 2 as shown in Table 5). Further, the second computing device may determine a second transmission subgraph based on the second plurality of records. In response to determining that a second record in the second plurality of records indicates that data is transmitted from a first service to a second service, a second aggregation message is generated.
[0084] Specifically, when processing the first record, it can be determined that the record indicates cross-regional data transmission, that is, the data received by node E in data center 2 comes from node C in data center 1. At this time, an aggregation message <123,subgraph02,idc01> can be generated. Here, "123" represents the TraceID of the data, subgraph02 represents the transmission subgraph of data center 2, and idc01 indicates that the cross-regional data transmission involves data center 1. According to some implementations of this disclosure, the data shown in Table 5 can be encapsulated to generate the aggregation message.
[0085] Generally, the number of cross-IDC call links is much lower than the number of non-cross-IDC call links, and the calculated link subgraph only contains node and edge information, without the need to transmit other irrelevant information. In this way, the amount of aggregated messages related to cross-regional transmission is far less than the original data volume. This method can significantly reduce the network bandwidth and storage space requirements for managing cross-regional data transmission.
[0086] According to some implementations of this disclosure, aggregated messages from multiple regions can be synchronized to an aggregation device. This aggregation device can be a physical device and / or a virtual device located anywhere. For example, to further reduce data transmission volume, the aggregation device can be located within a data center. In this case, triplet data can be synchronized to the same IDC for processing, while other non-cross-IDC links are computed in their local IDCs. This significantly reduces the bandwidth required for cross-IDC management data transmission and eliminates the need to store copies of the original link data. Furthermore, cross-IDC synchronization only transmits the computed transmission subgraph, which, since it only involves nodes and edges, complies with data security specifications and does not involve any sensitive information. This avoids potential security risks associated with cross-region data transmission.
[0087] According to some implementations of this disclosure, the aggregation device can receive aggregation messages from multiple regions. Specifically, in response to determining that a second aggregation message has been received from a second computing device within a predetermined time range, the transmission graph is updated based on the second aggregation message, which indicates that data has been transmitted from a first service to a second service. Continuing the example above, the aggregation device can receive aggregation messages <123,subgraph01,idc02> from the first computing device and <123,subgraph02,idc01> from the second computing device. These two messages actually represent the cross-regional transmission of the same data. Compared to receiving only a single aggregation message from a single computing device, receiving two aggregation messages indicates higher reliability, thereby improving the accuracy of determining the transmission graph.
[0088] According to some implementations of this disclosure, the second aggregation message further includes: a second transport subgraph representing data being transferred between the second group of services; and updating the transport graph further includes: aggregating the transport graph and the second transport subgraph to update the transport graph. The aggregation device can further aggregate the second transport subgraph, whereby the transport graph may include an internal transport subgraph in data center 1, a cross-region transport subgraph (including node C->E), and an internal transport subgraph in data center 2. Using some implementations of this disclosure, a complete transport graph spanning multiple data centers can be generated.
[0089] Alternate and / or additional locations can aggregate cross-regional transport subgraphs (including node C->E) into subgraph02. This allows for the determination of a transport graph that includes internal transports within data center 2, as well as cross-regional transports spanning data centers 1 and 2. As time progresses, the aggregation device can continuously receive data transmissions triggered by other requests, thereby generating a transport graph that includes more paths.
[0090] According to some implementations of this disclosure, received aggregated messages can be stored in a message queue. During the process of determining a transmission graph based on a first aggregated message, in response to determining that no second aggregated message has been received from the second computing device within a predetermined time range, the transmission graph is determined based on the first aggregated message. Using some implementations of this disclosure, it is not necessary to determine the transmission graph only after receiving two aggregated messages describing the same data transmission; instead, the transmission graph can be determined based on a single aggregated message. In this way, received aggregated messages can be processed at any time, avoiding message blocking. After the aggregated message has been processed, the first aggregated message can be removed from the message queue, thereby reducing the data storage overhead of the aggregation device.
[0091] It should be understood that while message aggregation can significantly reduce the amount of data to be transmitted, it can generate massive amounts of data during operation, resulting in the aggregation device receiving a large number of aggregation messages. For example, this could reach tens of thousands of messages per second, or even more. It should also be understood that aggregation messages are intermediate data and do not require long-term storage. In this case, after finding the subgraph of all regions corresponding to the TraceID, the aggregation message can be deleted, thus avoiding prolonged occupation of the aggregation device's storage space.
[0092] According to some implementations of this disclosure, the first aggregated message and the second aggregated message can be generated at as close a time as possible, so that the aggregation device can continuously process the two aggregated messages describing the same data transmission. Specifically, the difference between the first generation time of the first record and the second generation time of the second record satisfies a predetermined threshold, and the first aggregated message and the second aggregated message are stored in a message queue in chronological order.
[0093] Specifically, the above issue involves data alignment. Request calls (even across data centers) typically complete within seconds. The timestamps of the same TraceID corresponding to span data in different data centers are extremely close, and the span data will be stored in data partitions of the same time period in different data centers, or in adjacent data partitions. When processing link data, data partition alignment can be considered, that is, ensuring that the lineage calculation programs in each data center are processing data partitions of the same time period as much as possible. In this way, it can be ensured that the triples from each data center arrive at the aggregation device at similar times, thus achieving fast processing.
[0094] It should be understood that the time of data partition generation for each IDC can differ, potentially by hours. In this case, optimizing the generation of data partitions across each IDC can improve timeliness. Furthermore, lineage calculation can be performed using the time corresponding to the last generated data partition from N IDCs. Assuming Hive is used to store the records, the partition to be executed each time can be specified in HSQL: WHERE date = "${date-XX}" AND hour = "${hour-YY}", thus specifying the latest partition to consume based on XX days and YY hours prior to the current execution time.
[0095] The individual steps of managing cross-region data transfer have been described separately. The entire management process is described below with reference to Figure 9, which describes the aggregation device deployed in data center 1. Figure 9 shows a block diagram 900 for managing data transfer according to some implementations of this disclosure. As shown in Figure 9, span data in storage 1 can be written to MQ1 (i.e., message queue 1) via data synchronization task 810. Here, the span data is arranged according to trace_id. Since the span data within each partition has already been sorted according to trace_id in the synchronization task, the data written to MQ1 is also sorted according to trace_id, and thus the subsequent triples are also sorted. In this way, it can be ensured that all triple information with the same trace_id arrives at the aggregation device and is processed within the closest possible time.
[0096] Data written from storage 1 to the message queue (MQ) follows the partition key = trace_id. This ensures that all spans of data with the same trace_id are written to the same partition. Subsequently, all spans of data with the same trace_id are processed by the same lineage computing service instance in the aggregation device, thereby ensuring that the complete link subgraph of trace_id within the current data center is determined within a single service instance.
[0097] Alternatively and / or additionally, span data after passing from MQ1 to the lineage calculation service may undergo anomaly processing and correction (not shown in Figure 9). For example, rate limiting, anomaly filtering, and correction may be performed. Since some span data may contain dirty data, preprocessing can eliminate noise and improve the accuracy of the generated link graph. Subsequently, all span data corresponding to a trace_id will be processed by the same service instance in chronological order. When a new trace_id is identified within this partition, the list of span data with the old trace_id (e.g., ...) can be...<trace_id,spanlist> Packaging, and then generating<trace_id,subgraph,need_idcs> Triplet.
[0098] It should be understood that regardless of whether `need_idcs` is empty (i.e., regardless of whether there is cross-IDC behavior), lineage calculation is performed within the current IDC, thus ensuring complete lineage within the IDC. If cross-region behavior exists, at box 812, the triple data can be sent to the message queue (MQ) of the aggregated IDC (corresponding to MQ-2). This MQ must also maintain the order of messages within the partition, and data written to the MQ must also adhere to the partition key = trace_id. This ensures that triples with the same trace_id are processed by the same service instance.
[0099] When a service instance processes a triplet, it can check if another triplet for that trace_id exists in the cache. If not, the triplet is stored in the cache (cross-IDC processing will only begin if the need_idcs value in the triplet is not empty). The update time corresponding to the trace_id can be recorded.<trace_id,update_time> If the corresponding triple is found in the cache, the contents of the two triples can be retrieved.<subgraph,need_idcs> To aggregate in order to form new<trace_id,subgraph,need_idcs> Triples. This can determine if a new triple is missing information about other data centers (IDCs). If not, the link graph is complete and can be generated. If still missing, subsequent data can be processed until a predetermined time threshold is reached.
[0100] Aggregation equipment can be queried periodically.<trace_id,update_time> If the list shows that `now - update_time` > `lifetime`, meaning `update_time` hasn't been updated within a reasonable timeframe, data synchronization might fail due to exceptional circumstances. Triples in the cache that don't include all links can be processed.
[0101] Ultimately, the full data lineage of an IDC and all cross-IDC data lines can be stored in the aggregation device. Other edge IDCs (e.g., Data Center 2) contain data lines within the current IDC. By aggregating the data lines from all IDCs, the full data lineage of all IDCs can be obtained. According to some implementations of this disclosure, although a failure of a service instance may lead to the loss of some temporary information in the cache, and the reallocation of service instances may result in incomplete collection of data corresponding to trace_id, etc., the data lineage is aggregated from numerous link subgraphs, and the loss of some single-link data has no impact on generating a correct global graph.
[0102] By utilizing the exemplary implementation of this disclosure, the aggregation device handles only traffic related to inter-zone transmissions, while non-inter-zone data traffic is handled by computing devices within the zone. In this way, data transmission can be managed in a distributed manner, thereby balancing the workload of each device.
[0103] Example process
[0104] Figure 10 illustrates a flowchart of a method 1000 for managing data transmission according to some implementations of this disclosure. At block 1010, a first aggregation message is received from a first computing device, indicating that data is being transmitted from a first service of a first group of services in a plurality of services to a second service of a second group of services in a plurality of services, the first group of services being provided by a first group of devices located in a first region, and the second group of services being provided by a second group of devices located in a second region. At block 1020, based on the first aggregation message, a transmission graph representing the data transmission among the plurality of services is determined.
[0105] According to some implementations of this disclosure, the first aggregation message includes: an identifier of the data, a first transport subgraph indicating that the data is transmitted between the first group of services, and an identifier of a second region associated with the data transmission; and determining the transport graph further includes: aggregating the transport graph and the first transport subgraph to update the transport graph.
[0106] According to some implementations of this disclosure, the method further includes: in response to determining that a second aggregated message has been received from a second computing device within a predetermined time range, updating a transmission graph based on the second aggregated message, wherein the second aggregated message indicates that data has been transmitted from a first service to a second service.
[0107] According to some implementations of this disclosure, the second aggregation message further includes: a second transport subgraph representing data being transferred between the second group of services; and updating the transport graph further includes: aggregating the transport graph and the second transport subgraph to update the transport graph.
[0108] According to some implementations of this disclosure, the first aggregate message is generated by the first device through the following steps: receiving a first plurality of records representing data transmission, the first plurality of records indicating the path through which data is transmitted between a plurality of services; determining a first transmission subgraph based on the first plurality of records; and generating the first aggregate message in response to determining that the first record in the first plurality of records indicates that data is transmitted from a first service to a second service.
[0109] According to some implementations of this disclosure, the second aggregation message is generated by the second device through the following steps: receiving a second plurality of records representing data transmission, the second plurality of records indicating the path through which data is transmitted between a plurality of services; determining a second transmission subgraph based on the second plurality of records; and generating the second aggregation message in response to determining that a second record in the second plurality of records indicates that data is transmitted from a first service to a second service.
[0110] According to some implementations of this disclosure, the difference between the first generation time of the first record and the second generation time of the second record satisfies a predetermined threshold, and the first aggregated message and the second aggregated message are stored in the message queue in chronological order.
[0111] According to some implementations of this disclosure, the first transport subgraph is generated based on the following steps: obtaining an initial transport subgraph associated with a first group of services; in response to determining that the initial transport subgraph includes nodes corresponding to services in the record, updating the edges in the initial transport subgraph based on the paths in the record to determine the first transport subgraph.
[0112] According to some implementations of this disclosure, the first transport subgraph is generated based on the following steps: in response to determining that the initial transport subgraph does not include nodes corresponding to services in the record, the edges in the initial transport record are updated based on the paths in the record, and the nodes in the initial transport record are updated using services, to determine the first transport subgraph.
[0113] According to some implementations of this disclosure, determining a transmission graph based on a first aggregated message includes: in response to determining that no second aggregated message has been received from a second computing device within a predetermined time range, determining a transmission graph based on the first aggregated message; and removing the first aggregated message.
[0114] According to some implementations of this disclosure, determining the transmission graph includes: determining a first node in the transmission graph based on a first service; determining a second node in the transmission graph based on a second service; and determining the edge between the first node and the second node.
[0115] Example devices and equipment
[0116] Figure 11 shows a block diagram of an apparatus 1100 for managing data transmission according to some implementations of the present disclosure. The apparatus includes: a receiving module 1110 configured to receive a first aggregation message from a first computing device, the first aggregation message indicating that data is being transmitted from a first service of a first group of services in a plurality of services to a second service of a second group of services in a plurality of services, the first group of services being provided by a first group of devices located in a first region, and the second group of services being provided by a second group of devices located in a second region; and a determining module 1120 configured to determine, based on the first aggregation message, a transmission graph representing the data transmission among the plurality of services.
[0117] According to some implementations of this disclosure, the first aggregated message includes: an identifier of the data, a first transport subgraph indicating that the data is transmitted between the first group of services, and an identifier of a second region associated with the data transmission; and the determining module is further configured to determine that the transport graph further includes: an aggregated transport graph and the first transport subgraph, in order to update the transport graph.
[0118] According to some implementations of this disclosure, the apparatus further includes a processing module configured to: in response to determining that a second aggregated message has been received from a second computing device within a predetermined time range, update a transmission graph based on the second aggregated message, the second aggregated message indicating that data has been transmitted from a first service to a second service.
[0119] According to some implementations of this disclosure, the second aggregation message further includes: a second transport subgraph representing data being transferred between the second group of services; and the processing module is further configured to: aggregate the transport graph and the second transport subgraph to update the transport graph.
[0120] According to some implementations of this disclosure, the first aggregate message is generated by the first device through the following steps: receiving a first plurality of records representing data transmission, the first plurality of records indicating the path through which data is transmitted between a plurality of services; determining a first transmission subgraph based on the first plurality of records; and generating the first aggregate message in response to determining that the first record in the first plurality of records indicates that data is transmitted from a first service to a second service.
[0121] According to some implementations of this disclosure, the second aggregation message is generated by the second device through the following steps: receiving a second plurality of records representing data transmission, the second plurality of records indicating the path through which data is transmitted between a plurality of services; determining a second transmission subgraph based on the second plurality of records; and generating the second aggregation message in response to determining that a second record in the second plurality of records indicates that data is transmitted from a first service to a second service.
[0122] According to some implementations of this disclosure, the difference between the first generation time of the first record and the second generation time of the second record satisfies a predetermined threshold, and the first aggregated message and the second aggregated message are stored in the message queue in chronological order.
[0123] According to some implementations of this disclosure, the first transport subgraph is generated based on the following steps: obtaining an initial transport subgraph associated with a first group of services; in response to determining that the initial transport subgraph includes nodes corresponding to services in the record, updating the edges in the initial transport subgraph based on the paths in the record to determine the first transport subgraph.
[0124] According to some implementations of this disclosure, the first transport subgraph is generated based on the following steps: in response to determining that the initial transport subgraph does not include nodes corresponding to services in the record, the edges in the initial transport record are updated based on the paths in the record, and the nodes in the initial transport record are updated using services, to determine the first transport subgraph.
[0125] According to some implementations of this disclosure, the determining module is further configured to: determine a transmission graph based on the first aggregated message in response to determining that no second aggregated message has been received from the second computing device within a predetermined time range; and remove the first aggregated message.
[0126] According to some implementations of this disclosure, the determining module is further configured to: determine a first node in the transmission graph based on a first service; determine a second node in the transmission graph based on a second service; and determine an edge between the first node and the second node.
[0127] Figure 11 shows a block diagram of a device 1100 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 1100 shown in Figure 11 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The computing device 1100 shown in Figure 11 can be used to implement the methods described above.
[0128] As shown in Figure 11, computing device 1100 is in the form of a general-purpose computing device. Components of computing device 1100 may include, but are not limited to, one or more processors or processing units 1110, memory 1120, storage devices 1130, one or more communication units 1140, one or more input devices 1150, and one or more output devices 1160. Processing unit 1110 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 1120. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 1100.
[0129] Computing device 1100 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 1100, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 1120 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 1130 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within computing device 1100.
[0130] The computing device 1100 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG11, disk drives for reading or writing from removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading or writing from removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 1120 may include a computer program product 1125 having one or more program modules configured to perform various methods or actions of various implementations of this disclosure.
[0131] The communication unit 1140 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 1100 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 1100 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.
[0132] Input device 1150 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1160 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 1100 can also communicate as needed with one or more external devices (not shown) via communication unit 1140. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with computing device 1100, or with any device that enables computing device 1100 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interfaces (not shown).
[0133] According to exemplary implementations of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.
[0134] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0135] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0136] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0137] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0138] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for managing data transmission across multiple regions, comprising: Receive a first aggregation message from a first computing device, the first aggregation message indicating that data is transmitted from a first service of a first group of services in a plurality of services to a second service of a second group of services in a plurality of services, the first group of services being provided by a first group of devices located in a first region, and the second group of services being provided by a second group of devices located in a second region; as well as Based on the first aggregated message, a transmission graph representing the data being transmitted between the multiple services is determined.
2. The method according to claim 1, wherein the first aggregation message comprises: The data identifier, a first transport subgraph indicating that the data is transmitted between the first group of services, and an identifier of a second region associated with the data transmission; as well as Determining the transmission graph further includes: aggregating the transmission graph and the first transmission subgraph to update the transmission graph.
3. The method according to claim 2, further comprising: In response to determining that a second aggregated message has been received from a second computing device within a predetermined time range, the transmission graph is updated based on the second aggregated message, which indicates that the data has been transmitted from the first service to the second service.
4. The method of claim 3, wherein the second aggregation message further comprises: This represents a second transport subgraph indicating the data being transmitted between the second group of services; Updating the transmission graph further includes: aggregating the transmission graph and the second transmission subgraph to update the transmission graph.
5. The method of claim 3, wherein the first aggregation message is generated by the first device through the following steps: Receive a first plurality of records representing the transmission of the data, the first plurality of records indicating the path through which the data is transmitted between the plurality of services; Based on the first plurality of records, the first transmission subgraph is determined; as well as In response to determining that a first record among the first plurality of records indicates that the data is transferred from the first service to the second service, the first aggregation message is generated.
6. The method of claim 5, wherein the second aggregation message is generated by the second device through the following steps: Receive a second plurality of records representing the transmission of the data, the second plurality of records indicating the path through which the data was transmitted between the plurality of services; Based on the second plurality of records, the second transmission subgraph is determined; as well as In response to determining that a second record among the second plurality of records indicates that the data was transferred from the first service to the second service, a second aggregate message is generated.
7. The method according to claim 6, wherein the difference between the first generation time of the first record and the second generation time of the second record satisfies a predetermined threshold, and the first aggregated message and the second aggregated message are stored in a message queue in chronological order.
8. The method of claim 2, wherein the first transmission subgraph is generated based on the following steps: Obtain the initial transport subgraph associated with the first group of services; In response to determining that the initial transport subgraph includes nodes corresponding to services in the record, the edges in the initial transport subgraph are updated based on the paths in the record to determine the first transport subgraph.
9. The method of claim 8, wherein the first transport subgraph is generated based on the following steps: in response to determining that the initial transport subgraph does not include nodes corresponding to the service in the record, updating the edges in the initial transport record based on the paths in the record, and updating the nodes in the initial transport record using the service, to determine the first transport subgraph.
10. The method of claim 1, wherein determining the transport graph based on the first aggregated message comprises: In response to determining that no second aggregated message has been received from the second computing device within a predetermined time range, the transmission map is determined based on the first aggregated message; as well as Remove the first aggregated message.
11. The method of claim 1, wherein determining the transmission map comprises: The first node in the transmission graph is determined based on the first service; The second node in the transmission graph is determined based on the second service; as well as Determine the edge between the first node and the second node.
12. An apparatus for managing data transmission across multiple regions, comprising: A receiving module is configured to receive a first aggregated message from a first computing device, the first aggregated message indicating that data is transmitted from a first service of a first group of services in a plurality of services to a second service of a second group of services in a plurality of services, the first group of services being provided by a first group of devices located in a first region, and the second group of services being provided by a second group of devices located in a second region. as well as A determination module is configured to determine, based on the first aggregation message, a transmission graph representing the data being transmitted between the plurality of services.
13. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 11.
14. A computer-readable storage medium having a computer program stored thereon, the computer program causing the processor to implement the method according to any one of claims 1 to 11 when executed by a processor.
15. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Call link data processing method and device
CN111464352A
Micro-service link topology processing method and device and readable storage medium
CN113810234A
Message push link tracking method and system, electronic equipment and storage medium
CN117527519A
Tracking Application Utilization of Microservices
US20200021505A1