Heterogeneous computing system and cache consistency maintenance method, device, equipment, and medium
By abstracting the cache consistency protocol of heterogeneous computing systems into metadata index, multicast synchronization and hybrid response protocols, the most appropriate protocol is dynamically selected based on the task load characteristics, the problem of cache inconsistency in heterogeneous computing systems is solved, and efficient cache consistency maintenance and performance optimization are achieved.
Patent Information
- Application Number
- CN202510866086.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-26
AI Technical Summary
When multiple computing units hold data copies at the same time through private caches, heterogeneous computing systems lead to cache inconsistency problems. The existing technology cannot accurately adapt to the most suitable consistency protocol, resulting in poor performance.
The cache consistency protocol is abstracted into metadata index protocol, multicast synchronization protocol and hybrid response protocol. The most matching protocol is dynamically selected based on the task load feature value of heterogeneous computing systems, and the uncertainty of load characteristics is quantified through entropy weights to achieve multi-dimensional load awareness and efficient consistency maintenance.
It realizes efficient cache consistency maintenance of heterogeneous computing systems in different load scenarios, ensures continuous and efficient operation of system performance, reduces read and write delays and network loads, and improves bandwidth utilization.
Smart Images

Figure CN120353614B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of heterogeneous computing, and in particular to a heterogeneous computing system and a cache consistency maintenance method, device, electronic device, non-volatile storage medium, and computer program product thereof. Background Art
[0002] Heterogeneous computing systems integrate computing units with different architectures. When multiple computing units simultaneously hold data copies through private caches, modifications to the data by any computing unit will cause logical faults in the global storage space, that is, cache inconsistency problems will occur.
[0003] In order to solve the problem of cache consistency being unable to adapt to changing workloads when maintained through static consistency protocols, related technologies dynamically select protocols by setting simple read-write ratio thresholds. However, this fails to accurately adapt to the most appropriate consistency protocol, resulting in poor performance of heterogeneous computing systems. Summary of the Invention
[0004] The present invention provides a heterogeneous computing system and its cache consistency maintenance method, device, electronic device, non-volatile storage medium, and computer program product. Based on the dynamic characteristics of task load, it can accurately determine the consistency protocol that best matches the current task load of the heterogeneous computing system, realize effective switching of consistency maintenance protocols, and ensure that the heterogeneous computing system maintains high performance.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0006] In one aspect, the present invention provides a cache consistency maintenance method, comprising:
[0007] The cache consistency protocol of heterogeneous computing systems is redefined as metadata indexing protocol, multicast synchronization protocol and hybrid response protocol.
[0008] According to the entropy weights of different feature dimensions determined by the feature values of each data block in the heterogeneous computing system, the protocol that best matches the current task load operation status is determined from the metadata indexing protocol, multicast synchronization protocol and hybrid response protocol; the feature dimensions include at least read and write operation features and data heat features.
[0009] Among them, the metadata indexing protocol is a protocol in which each computing unit includes directory information for storing local metadata, and the target data blocks that meet the condition that the write frequency is greater than the preset frequency threshold have central directory information; the multicast synchronization protocol is a protocol in which each computing unit performs corresponding operations based on the monitoring information of the multicast channel, and the multicast range is dynamically adjusted according to the cache hit rate; the hybrid response protocol is a fusion protocol of the directory protocol and the broadcast protocol.
[0010] Another aspect of the present invention provides a cache consistency maintenance device, comprising:
[0011] The protocol abstraction module is used to redefine the cache consistency protocol of heterogeneous computing systems into a metadata indexing protocol, a multicast synchronization protocol, and a hybrid response protocol; among them, the metadata indexing protocol is a protocol in which each computing unit includes directory information for storing local metadata, and the target data blocks that meet the condition that the write frequency is greater than the preset frequency threshold have central directory information; the multicast synchronization protocol is a protocol in which each computing unit performs corresponding operations based on the monitoring information of the multicast channel, and the multicast range is dynamically adjusted according to the cache hit rate; the hybrid response protocol is a fusion protocol of the directory protocol and the broadcast protocol.
[0012] The protocol dynamic selection module is used to determine the protocol that best matches the current task load operation status from the metadata indexing protocol, multicast synchronization protocol and hybrid response protocol based on the entropy weights of different feature dimensions determined by the feature values of each data block in the heterogeneous computing system; the feature dimensions include at least read and write operation features and data heat features.
[0013] The present invention also provides an electronic device comprising a memory and a processor, wherein the processor is configured to implement the steps of any of the above-mentioned cache consistency maintenance methods when executing a computer program stored in the memory.
[0014] The present invention also provides a non-volatile storage medium having a computer program stored thereon, which implements the steps of any of the above-mentioned cache consistency maintenance methods when executed by a processor.
[0015] The present invention also provides a computer program product, comprising a computer program / instruction, which implements the steps of any of the above cache consistency maintenance methods when executed by a processor.
[0016] Finally, the present invention also provides a heterogeneous computing system, which includes at least a first computing node, a second computing node and a cache consistency controller, wherein the cache consistency controller is connected to the first computing node and the second computing node, and the first computing node and the second computing node include at least two computing units; wherein the cache consistency controller is used to implement the steps of any of the above-mentioned cache consistency maintenance methods when executing a computer program.
[0017] The advantage of the technical solution provided by the present invention is that the cache consistency protocol is abstracted into three modes: metadata indexing protocol, multicast synchronization protocol and hybrid response protocol, which correspond to high-write scenarios, high-read scenarios and dynamic balance scenarios respectively. The entropy weights of different feature dimensions are determined according to the characteristic values of the data blocks of the heterogeneous computing system in the current scenario. The uncertainty of the load characteristics is quantified by the entropy weights. The lower the entropy value, the higher the discrimination of the feature dimension for protocol selection. According to the entropy weights of the read and write operation characteristics, it is possible to perceive whether the current scenario is a high-read load scenario or a high-write load scenario. The entropy weights of the data heat characteristics can be used to dynamically evaluate the hot spot distribution characteristics and time sensitivity of the data blocks, thereby realizing multi-dimensional load perception of the heterogeneous computing system, and being able to accurately adapt to the consistency protocol that best matches the current task load operation status of the heterogeneous computing system, thereby realizing dynamic cache consistency of the heterogeneous computing system and ensuring continuous and efficient maintenance process, thereby ensuring efficient operation of the heterogeneous computing system.
[0018] In addition, the present invention also provides corresponding implementation devices, electronic devices, non-volatile storage media, computer program products and heterogeneous computing systems for the cache consistency maintenance method, further making the method more practical, and the devices, electronic devices, non-volatile storage media, computer program products and heterogeneous computing systems have corresponding advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 A schematic diagram of a hardware composition framework applicable to the cache consistency maintenance method provided by the present invention;
[0021] Figure 2 A schematic diagram of a process flow of a cache consistency maintenance method provided by the present invention;
[0022] Figure 3 A schematic diagram of a framework of a homogeneous computing system in an exemplary application scenario;
[0023] Figure 4 A schematic diagram of a framework of a heterogeneous computing system in an exemplary application scenario;
[0024] Figure 5 A structural framework diagram of an exemplary embodiment of the cache consistency maintenance device provided by the present invention;
[0025] Figure 6A structural diagram of an exemplary embodiment of an electronic device provided by the present invention;
[0026] Figure 7 A structural diagram of an exemplary embodiment of a heterogeneous computing device provided by the present invention;
[0027] Figure 8 A topological diagram of an exemplary embodiment of the intra-cluster communication architecture provided by the present invention;
[0028] Figure 9 A topological diagram of another exemplary embodiment of the intra-cluster communication architecture provided by the present invention;
[0029] Figure 10 A topological diagram of an exemplary embodiment of the communication architecture of heterogeneous computing devices provided by the present invention. DETAILED DESCRIPTION
[0030] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. The terms "first," "second," "third," "fourth," etc. in the specification and the accompanying drawings are used to distinguish different objects rather than to describe a specific order. Furthermore, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. The term "exemplary" means "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior or better than other embodiments.
[0031] With the rapid iterative development of multimodal artificial intelligence applications, the computing paradigm has undergone structural changes. The demand for massive data processing and the expansion of the scale of neural networks have formed a two-way driving force. Computing hardware systems with diversified architectures have been born, and heterogeneous computing systems that integrate multiple computing units have been widely used in intelligent computing scenarios.
[0032] Heterogeneous computing systems leverage complementary advantages by integrating computing units with different architectures. For example, general-purpose CPUs (Central Processing Units) execute logic control and task scheduling, while parallel computing units such as GPUs (Graphics Processing Units) and Tensor Processing Units (TPUs) perform intensive tensor operations. FPGAs (Field Programmable Gate Arrays) are reconfigured to adapt to dynamic algorithm requirements, while ASICs (Application-Specific Integrated Circuits) are customized for specific scenarios to achieve optimal energy efficiency. Through dynamic resource allocation, these systems can meet the differentiated needs of scenarios like image recognition and semantic understanding, while achieving a precise balance between power consumption and performance based on workload. With the growing demand for collaboration between edge computing and the cloud, processor-level interconnect technology and the construction of virtualized resource pools enable the dynamic allocation of distributed computing nodes, leading to widespread application in areas such as ultra-large-scale pre-trained model deployment and real-time video analysis. Taking multi-heterogeneous computing systems as an example, these systems integrate a variety of heterogeneous processors, accelerators, and memory components. These systems are typically composed of multiple computing units with significant differences in performance, power consumption, programming models, and memory consistency models. This creates a "memory wall effect" where the data throughput demands of the computing cores far exceed the memory bandwidth supply. To avoid severe performance imbalances in multi-heterogeneous computing systems, these systems undergo corresponding optimizations at the physical, logical, and control layers. For example, at the physical layer, a tiered storage topology consisting of register files, shared caches, and global memory is constructed; at the logical layer, an asymmetric memory access model based on the NUMA (Non-Uniform Memory Access) architecture is implemented; and at the control layer, intelligent prefetching algorithms and dynamic bandwidth allocation mechanisms are deployed. However, due to the huge differences in semantics and communication patterns between the memory architectures of heterogeneous computing systems, heterogeneous computing systems have the problem of inconsistent cache data status: when multiple computing units simultaneously hold data copies through private caches, any modification of the data by any computing unit will cause a logical fault in the global storage space, that is, data inconsistency between the caches of multiple computing units or inconsistency between the cache and memory data. This is especially true in concurrent execution scenarios. For example, when a GPU streaming multiprocessor updates shared parameters, if it fails to synchronize to the CPU's last-level cache in time, it will cause deviations in the algorithm calculation results, and even be serious enough to cause system-level execution errors.
[0033] To ensure the stable operation of heterogeneous computing systems, related technologies achieve cache coherence through protocol stack innovation and hardware co-design. At the software protocol layer, these technologies maintain cache coherence in heterogeneous computing systems in read-intensive, write-infrequent workloads using a bus snooping protocol, and in scenarios with frequent data updates using a directory protocol. In this bus snooping protocol, the private cache of each processing unit in the heterogeneous computing system continuously monitors bus transactions. Upon detecting a coherence request broadcast from the bus, predefined response logic is triggered based on the local cache line state. This leverages the physical properties of the bus medium to ensure transaction atomicity and ordering, simplifying the state transition logic into a deterministic finite state machine. Cache state synchronization can be achieved with minimal effort through a global broadcast communication model. However, as multi-core processors scale, bus loads increase. The fact that only one CPU is allowed to occupy the bus at a time leads to frequent blocking of memory access requests, significantly increasing memory access latency and potentially wasting resources. Furthermore, the bus snooping protocol has poor scalability, making it difficult to adapt to the needs of large-scale processor systems. In other words, the bus snooping protocol can effectively maintain cache coherence in small-scale heterogeneous computing systems with minimal effort. The target protocol achieves cache coherence through a global public directory that records the state information of global cache lines. This global public directory includes the coherence state and a list of replica owners. In the directory protocol, all coherence messages are forwarded through the directory structure. By querying the replica owner list, efficient point-to-point message transmission is achieved, reducing network communication overhead. Furthermore, the directory protocol effectively manages replicas of shared data. By establishing a directory to track the storage location and status of data, multiple processors can conveniently access shared data, improving data availability and access efficiency. Large-scale multi-core processor systems can effectively coordinate cache interactions in a multi-core environment through centralized directory management, achieving efficient and reliable cache coherence maintenance. However, storing and updating the directory consumes resources, which may affect overall system performance. Furthermore, the directory protocol may introduce additional communication latency, slowing data access, making it unsuitable for large-scale heterogeneous computing systems with high real-time requirements and limited resources.
[0034] Regarding cache consistency maintenance, a related technique uses a snooping update strategy to maintain cache consistency across heterogeneous computing systems. This technique uses a translator module to convert the internal memory consistency model of a cluster into a unified consistency protocol based on the C11 memory consistency model. The consistency controller module, acting as a consistency protocol manager, receives messages from each translator and uniformly reorders and routes them. This leverages the loose ordering behavior of heterogeneous clusters to achieve effective compatibility between multiple heterogeneous memory consistency models. Another related technique uses a directory-based heterogeneous consistency protocol. This technique designs a composite consistency protocol fusion method for heterogeneous clusters. This method adheres to a global composite consistency model while preserving the consistency models of the original clusters. It then uses proxy caching to fuse the consistency models of different clusters, ensuring consistency across the entire system. Furthermore, another related technique designs a litmus test scheme for composite consistency models, using formal testing tools to verify that heterogeneous protocols adhere to the composite consistency model. These three approaches incur additional overhead to maintain consistency, which may impact system performance. Another related technology uses CXL (Compute Express Link, a high-speed interconnection standard) to unify memory semantics at the protocol level to achieve atomic operation support across device caches. However, this method remains at the technical planning and standards level, and its industrialization is relatively lagging behind.
[0035] Although related technologies can alleviate memory access latency to a certain extent, some problems still exist: updating the protocol requires a fixed global directory, it is impossible to dynamically optimize the protocol strategy, it has poor flexibility, and poor performance in complex load scenarios. In addition, protocol synthesis relies on static cluster configuration, lacks runtime adaptability, and cannot dynamically adjust the protocol according to the system status, resulting in less than ideal performance in different load scenarios and an inability to effectively adapt to diverse heterogeneous architectures. In order to adapt to the dynamic characteristics of task loads and achieve flexible switching of consistency maintenance protocols, protocols are usually selected through simple read-write ratio thresholds. This method cannot accurately capture complex cross-device data dependencies and is difficult to achieve effective protocol optimization.
[0036] In view of this, the present invention abstracts the cache consistency protocol into three modes: metadata indexing protocol, multicast synchronization protocol and hybrid response protocol, which correspond to high-write scenarios, high-read scenarios and dynamic balance scenarios respectively. The entropy weights of different feature dimensions are determined according to the characteristic values of the data blocks of the heterogeneous computing system in the current scenario. The discrimination of the protocol selection by the entropy weight is used to perceive whether the current scenario is a high-read load scenario or a high-write load scenario, and the consistency protocol that best matches the current task load operation status of the heterogeneous computing system can be accurately adapted. In combination with the specific application environment architecture or the specific hardware architecture on which the execution of the cache consistency maintenance method depends, the specific application environment architecture or the specific hardware architecture is described here. The following is combined with Figure 1 Some possible application scenarios involved in the technical solution of the present invention are introduced by way of example, which may include the following:
[0037] In this embodiment, a heterogeneous computing system includes multiple servers 1, each server 1 serving as a computing node of the heterogeneous computing system. Each server includes at least two computing units 10. In addition to its own CPU, it also includes at least one built-in computing unit of a different type from the CPU architecture, such as a GPU, FPGA, or ASIC. The communication structure of the heterogeneous computing system is as follows: all computing units of all computing nodes are divided according to computing unit type, and computing units of the same computing unit type are divided into clusters. The computing units in the same cluster can be connected via a tree topology, and one computing unit is selected as the master computing unit. Different clusters interact through the master computing unit, and the clusters are connected via the master computing unit using a ring topology.
[0038] Given the complex characteristics of heterogeneous computing systems, which span protocols, hierarchies, and time and space, and in order to overcome the difficulties in maintaining consistency and the high protocol complexity of heterogeneous computing systems in related technologies, achieve dynamic cache consistency in heterogeneous computing systems, and ensure a continuous and efficient maintenance process, the present invention proposes a dynamic consistency maintenance method based on a hybrid of bus snooping and directory protocols. A cache consistency controller containing computer program code implementing the following method steps can be built into any of the aforementioned servers or cloud servers, and a data collection component can be built into each server to collect data such as the number of read and write operations and read and write times. The cache consistency controller redefines the cache consistency protocol of the heterogeneous computing system into a metadata indexing protocol, a multicast synchronization protocol, and a hybrid response protocol. The metadata indexing protocol is a protocol in which each computing unit includes directory information storing local metadata, and target data blocks that meet the condition that the write frequency exceeds a preset frequency threshold have central directory information. The multicast synchronization protocol is a protocol in which each computing unit performs corresponding operations based on the monitoring information of the multicast channel, and the multicast range is dynamically adjusted according to the cache hit rate. The hybrid response protocol is a fusion protocol of the directory protocol and the broadcast protocol. According to the entropy weights of different feature dimensions determined by the characteristic values of each data block of the heterogeneous computing system, the protocol that best matches the current task load operation status is determined based on the entropy weights, thereby achieving effective switching of the consistency maintenance protocol based on the dynamic characteristics of the task load, ensuring that the heterogeneous computing system maintains high performance.
[0039] It should be noted that the above application scenarios are merely provided to facilitate understanding of the concepts and principles of the present invention, and the embodiments of the present invention are not limited in this respect. On the contrary, the embodiments of the present invention can be applied to any applicable scenario. After introducing the technical solution of the present invention, various non-limiting embodiments of the present invention are described in detail below with reference to the accompanying drawings and specific embodiments.
[0040] Computing units of different architectures in heterogeneous computing systems exhibit fundamental differences in communication patterns and consistency models. Traditional CPUs, driven by their inherent need for strict memory ordering and low-latency access, employ invalidation mechanisms triggered by read and write operations, such as the modified exclusive shared invalid (MESI) protocol. This protocol ensures data consistency by strictly enforcing the Single Writer Multiple Reader (SWMR) invariant. Massively parallel architectures such as GPUs employ a more relaxed consistency model. Their non-SWMR self-invalidation protocols allow for asynchronous maintenance of cache state between computing cores. This approach, which improves data throughput by relaxing consistency constraints, creates a semantic gap with CPU communication mechanisms. The incompatibility of protocol specifications between non-computing units increases the complexity of cross-architecture cache state synchronization, making it more difficult to maintain cache consistency protocols for heterogeneous computing systems while balancing strict consistency guarantees with parallel computing efficiency.
[0041] Furthermore, the multi-layered heterogeneity, differences in communication mechanisms, and asymmetric concurrency control models of heterogeneous computing systems complicate data conflicts and deadlocks. First, CPUs employ a strongly consistent memory model, relying on sophisticated cache coherence protocols to maintain data visibility. Accelerators like GPUs employ a weakly consistent model, achieving data coordination through batch synchronization or explicit barriers. This fragmented memory model makes it difficult to establish globally unified data synchronization logic for cross-device access. For example, atomic operations on shared variables performed by the CPU may not be correctly perceived by the GPU, leading to implicit data contention. Second, the heterogeneity of concurrency control mechanisms exacerbates the risk of conflicts. CPUs rely on fine-grained synchronization primitives such as locks and semaphores, while GPUs employ thread group synchronization (such as line constraint synchronization) or global barriers within the SIMT (Single Instruction Multiple Threads) architecture. The interaction between the two can result in asymmetric wait relationships. For example, if a CPU thread is preempted by a GPU kernel while holding a resource, the GPU thread group may deadlock while waiting for the CPU to release resources. Furthermore, heterogeneous computing systems often use non-uniform memory architectures. The latency differences in cross-level data transmission render traditional deadlock detection algorithms ineffective, and resource competition in inter-device communication links may trigger distributed deadlocks. In addition, the fragmentation of debugging tools makes it difficult to trace the source of problems. The CPU-GPU debugging interfaces and performance counters are incompatible, making it difficult to capture abnormal states in cross-device dependency chains. In summary, heterogeneous computing systems exhibit complex characteristics across protocols, levels, time and space. Traditional single-architecture solutions cannot be directly migrated. This embodiment provides a method for achieving dynamic cache consistency in heterogeneous computing systems and ensuring continuous and efficient maintenance processes. Please refer to [1]. Figure 2 , Figure 2A schematic flow chart of a cache consistency maintenance method provided in this embodiment may include the following contents:
[0042] S201: Redefine the cache consistency protocol of the heterogeneous computing system as a metadata indexing protocol, a multicast synchronization protocol, and a hybrid response protocol.
[0043] Among them, the heterogeneous computing system can be a multi-heterogeneous computing system, a non-multi-heterogeneous computing system, or any type of heterogeneous computing system, as long as it meets the definition of a heterogeneous computing system. The present invention does not impose any restrictions on this. Heterogeneous computing systems and homogeneous computing systems have cache consistency differences. Homogeneous computing systems are systems composed of computing units using the same type of instruction set and architecture. For example, a homogeneous computing system represented by a multi-core CPU, such as Figure 3 As shown, homogeneous computing systems rely on a shared memory model to achieve uniform memory address sharing and utilize cache coherence protocols such as MESI and MOESI (Modified Exclusive Shared Invalid Owned) to maintain cache data consistency across different processor cores. Homogeneous computing systems utilize a hardware-supported bus snooping coherence protocol to ensure cache data consistency: each core's cache controller monitors all memory access requests on the bus. When a core (e.g., core 1) modifies cached data, it sends a modification request to the bus. Upon detecting this request, the cache controllers of other cores update or invalidate their own cached copies as needed. This cache coherence maintenance method is simple to implement and has low hardware overhead. However, as the number of cores increases, bus bandwidth becomes a performance bottleneck, making it unsuitable for large-scale multi-core computing systems. When a homogeneous computing system processes a read request from a core, the cache coherence maintenance process includes the following: a local cache check phase: if a valid copy exists in the local cache, the data is returned directly; if the local cache misses, a read request is sent to the bus. Bus broadcast domain monitoring phase: The bus broadcasts the read request to all cores, and the cache controllers of other cores check the status of their local caches. Data acquisition phase: The main memory or the cache holding the latest data returns the data through the bus, the core that originally initiated the request receives the data, marks the cache line status, and completes all operations of the read request. For heterogeneous computing systems, such as Figure 4As shown in the figure, each computing node of a heterogeneous computing system, such as computing node A and computing node B, has a local coherence controller (LCC) and protocol converter deployed within each computing unit. The system level relies on a global coherence controller (GCC) to achieve cross-device collaboration. The cache coherence maintenance process of the heterogeneous computing system is as follows: Local request processing phase: When the LCC receives a coherence request, it first completes the request parsing and response within the computing node to which it belongs. When local resources cannot meet the request conditions, such as cache line status conflict or data missing, the request forwarding mechanism is triggered. Global coordination phase: Unfinished local requests are routed to the GCC, which distributes the requests to the LCC of the target computing node according to the topology protocol and cross-device routing rules. Cross-device response phase: After the target computing node completes request processing, it generates a coherence response message and returns it to the GCC. The GCC forwards the response to the original requester LCC based on the transaction identifier, completing the closed-loop transaction processing. Therefore, cache coherence in heterogeneous computing systems requires a GCC interface and protocol converter functionality. The GCC interface must support instruction set mapping for heterogeneous coherence protocols and achieve multi-protocol compatibility through a standardized transaction interface. The protocol converter must convert local and global protocol states, such as converting local protocol primitives into a global transaction format, performing semantic parsing and local state updates on global response messages, maintaining the protocol conversion state machine, and ensuring atomicity during the conversion process. Therefore, the bus-based coherence maintenance protocol for homogeneous computing systems has high communication overhead and is suitable for small-scale clusters, while the directory protocol for heterogeneous computing systems has strong scalability, low communication overhead, and is suitable for large-scale clusters. Both cache coherence protocols have their own advantages and disadvantages. Therefore, this step combines the respective advantages of the bus snooping protocol and the directory protocol. Based on different read-write ratios, that is, the task load and characteristics of heterogeneous computing systems, the cache coherence protocol is abstracted to obtain three protocol modes suitable for different task load scenarios. As shown in Table 1, each protocol mode can be flexibly switched and coordinated to complete the cache coherence maintenance process for heterogeneous computing systems.
[0044] The metadata indexing protocol is implemented in a directory manner. Based on the directory protocol, it features a lightweight directory tree and a hierarchical directory structure (not a global directory). Each computing unit maintains its own metadata, such as cache status and version number. Specifically, each computing unit stores directory information of its local metadata in its own directory. Local metadata refers to the computing unit's own metadata. This directory information can be stored in a directory table, which records the status of each cache line in the computing unit's local private cache, whether read or write data hits the cache, and version numbers. Central directory information can be deployed in one of the computing units, the highest-performing computing unit, or a user-specified computing unit, such as the computing unit where a data block with a high write frequency resides. Specifically, target data blocks with a write frequency greater than a preset frequency threshold have central directory information. The central directory information at least records the correspondence between different directory information and the corresponding computing unit, as well as the data storage address. This information reflects the global cache status. The directory information of each computing unit and the central directory information constitute the lightweight directory tree of the metadata indexing protocol. The metadata indexing protocol uses a lightweight directory tree to query the directory only when necessary, avoiding broadcast storms. The hierarchical directory also reduces global lock contention, reducing write latency by 15%-20%. Compared to traditional broadcast protocols, it eliminates broadcast traffic and provides more stable network loads. The multicast synchronization protocol is suitable for scenarios with high read loads. By maximizing cache hit rates, the local cache directly responds to read requests, reducing metadata queries and metadata verification latency, thereby effectively reducing read latency. In this protocol, each computing unit performs operations based on information monitored on the multicast channel, and the multicast range is dynamically adjusted based on the cache hit rate. In other words, the multicast synchronization protocol uses probabilistic multicasting, with the probability of the multicast range being dynamically adjusted based on the computing unit's cache hit rate. Furthermore, it employs listener-based cache verification, meaning that computing units within the multicast range monitor the multicast channel and execute corresponding operations upon receiving update notifications. Compared to traditional directory protocols, this protocol eliminates directory queries for read operations, resulting in lower latency, reducing read latency by 25%-30%. Compared to traditional all-broadcast protocols, probabilistic multicast reduces redundant traffic, lowering network load by 10% to 15%, optimizing bandwidth, and increasing bandwidth utilization by 20%. The Hybrid Response Protocol is a fusion of the directory and broadcast protocols. It dynamically balances mixed read and write loads and adapts to rapid load changes. Compared to traditional directory protocols, this dynamic fusion of directory and broadcast mechanisms avoids the limitations of a single protocol. Compared to traditional broadcast protocols, it supports on-demand switching and offers greater compatibility.
[0045] Table 1 Comparison of three protocols and traditional protocols
[0046]
[0047] S202: According to the entropy weights of different feature dimensions determined by the feature values of each data block of the heterogeneous computing system, determine the protocol that best matches the current task load operation status from the metadata indexing protocol, the multicast synchronization protocol, and the hybrid response protocol.
[0048] Among them, the characteristic values at least include the relevant values of the read and write operation characteristics and the relevant values of the data heat characteristics. Correspondingly, the characteristic dimensions at least include the read and write operation characteristics and the data heat characteristics. By weighting the multi-dimensional characteristics, the deviation of the single characteristic index is avoided and the adaptability to the complex load of the heterogeneous computing system is enhanced. Those skilled in the art can measure the degree of difference between the read and write operation characteristics and the data heat characteristics in the metadata indexing protocol, the multicast synchronization protocol and the hybrid response protocol according to any objective weighting method based on the information entropy theory, and assign different weights to different characteristics. The lower the entropy value, the higher the discrimination of the protocol selection of the characteristic dimension, that is, the size of the entropy weight of different characteristics can distinguish different task load states. Since the three protocols are divided based on different task load states, after determining the entropy weights corresponding to different characteristic dimensions, the optimal consistency protocol is the protocol corresponding to the corresponding task load state. If the cache consistency protocol currently used by the heterogeneous computing system is the same as the optimal cache consistency protocol selected, no protocol adjustment is required. If the cache consistency protocol currently used by the heterogeneous computing system is different from the optimal cache consistency protocol selected, switch to the optimal cache consistency protocol selected in this step to achieve effective dynamic adjustment of the cache consistency protocol mode.
[0049] In the technical solution provided in this embodiment, the cache consistency protocol is abstracted into three modes: metadata indexing protocol, multicast synchronization protocol and hybrid response protocol, which correspond to high-write scenarios, high-read scenarios and dynamic balance scenarios respectively. The entropy weights of different feature dimensions are determined according to the characteristic values of the data blocks of the heterogeneous computing system in the current scenario. The uncertainty of the load characteristics is quantified by the entropy weights. The lower the entropy value, the higher the discrimination of the feature dimension for protocol selection. According to the entropy weights of the read and write operation characteristics, it is possible to perceive whether the current scenario is a high-read load scenario or a high-write load scenario. The entropy weights of the data heat characteristics can be dynamically evaluated. The hot spot distribution characteristics and time sensitivity of the data blocks can be realized, and multi-dimensional load perception of the heterogeneous computing system can be realized. It can accurately adapt to the consistency protocol that best matches the current task load operation status of the heterogeneous computing system, realize dynamic cache consistency of the heterogeneous computing system, and ensure that the maintenance process is continuously efficient, thereby ensuring efficient operation of the heterogeneous computing system.
[0050] In the above embodiment, there is no limitation on how to determine the entropy weights of different feature dimensions. Based on the above embodiment, the present invention also provides an implementation process for calculating the entropy weights of different feature dimensions based on the feature values of each data block of the heterogeneous computing system, which may include: a process for calculating the information entropy value of the current feature: for each data block, based on the ratio of the original feature value of the current feature of the current data block and the sum of the original feature values of the current features of all data blocks of the heterogeneous computing system, determine the feature probability distribution information of the current data block in the current feature; based on the feature probability distribution information of the current features of each data block of the heterogeneous computing system, determine the information entropy value of the current feature; based on the proportional relationship between the information entropy value of the current feature and the information entropy values of all features, determine the entropy weight of each feature.
[0051] Among them, the features can be, for example, read operation features, write operation features, access frequency features, and local attenuation factor features. The current feature can be any feature that requires entropy weight calculation. The current data block refers to a data block among all data blocks in the heterogeneous computing system. The original feature value is the specific value of the current feature in the data block at the current moment. For example, if the current feature is a read operation feature, and the number of read operations of data block A at the current moment is 8, then the original feature value of the read operation feature of the data block is 8. Feature probability distribution information is used to measure the normalized proportion of a specified data block in a specified dimension. The feature probability distribution information can ensure that the total proportion of all data blocks in the heterogeneous computing system in the same feature dimension is 1. For example, the current data block Feature probability distribution information at the current feature j It can be expressed as:
[0052] ;
[0053] in, For data blocks The original feature value at the current feature j, n represents the total number of data blocks.
[0054] For example, an exemplary method for calculating the information entropy value of each feature is to calculate the product of the feature probability distribution information of the current data block in the current feature and its logarithmic function value, and use the negative value of the sum of the product values corresponding to each data block of the heterogeneous computing system as the information entropy value of the current feature. As an efficient calculation method, the information entropy value calculation relationship can be pre-stored and the information entropy value of each feature can be calculated in sequence by calling the information entropy value calculation relationship. The information entropy value calculation relationship can be expressed as:
[0055] ;
[0056] in, is the information entropy value of feature dimension j.
[0057] For example, an exemplary method for calculating the entropy weight of the current feature is: respectively calculate the difference between the information entropy value of each feature and the target value; and take the ratio of the difference corresponding to the current feature to the sum of the differences corresponding to other features as the entropy weight of the current feature. As an efficient calculation method, the entropy weight calculation relationship can be pre-stored, and the entropy weight of each feature can be calculated in turn by calling the entropy weight calculation relationship. The entropy weight calculation relationship can be expressed as:
[0058] ;
[0059] in, is the information entropy value of feature dimension j, M is the total number of features, and m is one of the M features.
[0060] As can be seen from the above, this embodiment first calculates the information entropy value of the feature, and uses the feature probability distribution information to ensure that the total proportion of all data blocks in the heterogeneous computing system on the same feature dimension is 1. Finally, the entropy weight of the feature is determined based on the ratio of the information entropy values of different features. This normalized weight can reflect the contribution of each feature to protocol selection. The lower the entropy value, the higher the information purity and the greater the weight. This can more accurately quantify the uncertainty of the feature and improve the discrimination of the feature dimension on protocol selection.
[0061] The above embodiment does not impose any limitation on how to select a protocol based on the entropy weight. Based on the above embodiment, the present invention further defines the corresponding selection method between the entropy weights of different features and the three protocols from multiple perspectives, which may include the following:
[0062] If the entropy weight of the read operation feature at the current moment is greater than the entropy weights of other feature dimensions, the multicast synchronization protocol is selected as the protocol that best matches the task load running state of the heterogeneous computing system at the current moment; if the entropy weight of the write operation feature at the current moment is greater than the entropy weights of other feature dimensions, the metadata indexing protocol is selected as the protocol that best matches the task load running state of the heterogeneous computing system at the current moment; if the entropy weight of the data heat feature at the current moment is greater than the entropy weights of other feature dimensions, then according to the data heat change trend, the protocol that best matches the task load running state at the current moment is determined from the metadata indexing protocol, multicast synchronization protocol and hybrid response protocol.
[0063] In this embodiment, the feature dimensions include read operation features, write operation features, and data heat features. The read operation feature can be represented by the number of read operations, reflecting the read intensity of the data block, and the write operation feature can be represented by the number of write operations, reflecting the write intensity of the data block. After determining the entropy weights of the read operation feature, write operation feature, and data heat feature according to the above method or the method described in the related art, the entropy weights of each feature are compared. If the entropy weight of the read operation feature is large, it indicates that the current heterogeneous computing system has a large read operation. In order to maximize system performance, a multicast synchronization protocol for high read load scenarios can be selected as the protocol that best matches the current task load state of the heterogeneous computing system. If the entropy weight of the write operation feature is large, it indicates that the current heterogeneous computing system has a large write operation. In order to maximize system performance, a metadata indexing protocol for write and read load scenarios can be selected as the protocol that best matches the current task load state of the heterogeneous computing system. As for the data heat feature, high-frequency data heat will trigger a more aggressive caching strategy. If the data heat decays rapidly, protocol switching should be suppressed to avoid over-optimization of cold data. The selection can be based on the data heat trend.
[0064] Exemplarily, this embodiment may use access frequency characteristics and locality attenuation factor characteristics to represent data heat characteristics. Access frequency characteristics may be represented based on statistical data access frequency values (times / second). Access frequency characteristics can measure data heat. High frequency may trigger more aggressive caching strategies. The locality attenuation factor quantifies the stability of access patterns through the coefficient of variation of time intervals. It can accurately measure the time sensitivity of data access through the ratio of the standard deviation to the mean. The larger the value, the faster the data heat decays. The locality attenuation factor can reduce cold data misjudgments, avoid protocol switching overhead for occasional access, enhance dynamic adaptability, and improve the robustness of the cache consistency mechanism in load fluctuation scenarios. If the entropy weight of the access frequency feature is greater than the entropy weight of the locality attenuation factor feature, the metadata indexing protocol is selected as the protocol that best matches the task load operating state of the heterogeneous computing system at the current moment. If the entropy weight of the locality attenuation factor feature is greater than the entropy weight of the access frequency feature, when selecting a cache consistency protocol based on the locality attenuation factor feature, the locality attenuation factor feature can be measured by the access time distribution feature of the data block. In other words, the protocol that best matches the task load operating state at the current moment can be determined based on the access time distribution feature. If the access time distribution feature determines that the data heat is stable, the multicast synchronization protocol can be selected. If the access time distribution feature determines that the access fluctuates drastically, protocol switching should be suppressed. If the access time distribution feature determines that the access fluctuates to a certain extent but the overall heat remains moderate, the hybrid response protocol can be used.
[0065] From the above, it can be seen that this embodiment integrates the read-write ratio, access frequency, and access time distribution through entropy weight, more accurately understands the dynamic characteristics of task load, improves the accuracy of protocol selection, and realizes flexible switching of consistency maintenance protocols.
[0066] For example, in this embodiment, the local attenuation factor feature can be used alone to represent the data heat feature. The local attenuation factor feature can be represented by the access time distribution feature. A plurality of access time interval values are determined according to the multiple consecutive access moments of the current data block. If each access time interval value meets the preset same similarity condition, such as being the same or not much different, such as three intervals 、 , , then the multicast synchronization protocol is selected as the protocol that best matches the task load operation status of the heterogeneous computing system at the current moment; if the values of each access time interval meet the preset post-burst access restriction conditions, such as there is a large deviation between the values of one access interval and the other access intervals, such as , , , then generate the protocol switching suppression prompt information; if the access time interval values meet the preset medium fluctuation access conditions, the difference between the access interval values does not fluctuate much, such as , , , then the hybrid response protocol is selected as the protocol that best matches the task load operation status of the heterogeneous computing system at the current moment.
[0067] Furthermore, to accurately measure the local attenuation factor characteristic, this embodiment also provides a quantitative representation of the local attenuation factor characteristic. This method obtains at least four consecutive access times for each data block and calculates the access time interval between each two adjacent access times. The ratio of the mean and standard deviation of each access time interval is used as the local attenuation factor characteristic. As an efficient implementation method, the local attenuation factor characteristic calculation relationship can be pre-stored and then called to calculate the local attenuation factor characteristic of the i-th data block. The local attenuation factor characteristic calculation relationship can be expressed as:
[0068] in, , .
[0069] in, is the local attenuation factor characteristic, the time interval variable Represents a data block The time interval between the last three consecutive visits to Indicates the time difference between the first and second visits, Indicates the time difference between the second and third visits, The physical meaning of this formula is to reflect the access time distribution characteristics of the data block. For example, when , which is called a short interval, indicating that the data is frequently accessed and has high heat; when , which is called a long interval, indicating that the data may enter a cold state. 、 as well as Represents three time intervals, Represents three time intervals 、 as well as The standard deviation of Represents three time intervals 、 as well as When the standard deviation is large, the time intervals vary greatly, indicating an unstable access pattern; conversely, a low standard deviation indicates uniform time intervals and a regular access pattern.
[0070] After the above embodiment quantifies the local attenuation factor feature, the corresponding protocol can be further selected by a simple numerical comparison: if the value of the local attenuation factor feature is approximately equal to 0, approximately equal to, or approximately equal to, that is, a value in the vicinity of 0±a smaller value, where the smaller value is a number not exceeding 0.1, then the multicast synchronization protocol is selected as the protocol that best matches the task load operation state of the heterogeneous computing system at the current moment; if the value of the local attenuation factor feature is greater than 1, a protocol switching inhibition prompt message is generated; if the value of the local attenuation factor feature is approximately equal to 1, equivalent to, approximately equal to, that is, a value in the vicinity of 1±a smaller value, where the smaller value is a number not exceeding 0.1, then the hybrid response protocol is selected as the protocol that best matches the task load operation state of the heterogeneous computing system at the current moment.
[0071] In this embodiment, when , indicating highly uniform intervals, such as periodic access. For example, the time interval: 、 , , , ,but , data heat is stable, data heat is stable. This indicates that the access interval fluctuates dramatically, such as a burst of access followed by a long period of inactivity. The time interval is: , , , , data heat decays rapidly, protocol switching should be suppressed to avoid over-optimization of cold data. This indicates that the access interval fluctuates to a certain extent, but the overall popularity remains moderate and the popularity decays slowly. For example, the time interval: , , , .
[0072] As can be seen from the above, this embodiment uses the locality attenuation factor to quantify the volatility of access time, distinguish between short-term and long-term hot and cold data, and realize dynamic data heat perception. In the case of rapid decay of data heat, it tends to maintain the current protocol, suppress the cold data switching overhead, avoid excessive response to occasional access, and prioritize the allocation of efficient protocols to stable hot data, thereby improving cache hit rate and resource utilization.
[0073] Compared to the protocol selection method provided in the above embodiment, in order to cope with sudden load changes and long-term fluctuations in heterogeneous computing systems and reduce performance jitter in heterogeneous computing systems, the present invention also provides another protocol selection method, which may include the following:
[0074] Based on historical threshold information and attenuation adjustment information, the thresholds at different moments in the time window of the current moment are calculated, and the maximum and minimum values are selected from them. Based on the maximum and minimum values, the existing maximum and minimum thresholds corresponding to the current moment are updated to obtain the maximum new threshold and the minimum new threshold; according to the entropy weights of different features and their corresponding adjustment factors, the protocol matching degree value is determined; according to the numerical relationship between the protocol matching degree value and the maximum new threshold and the minimum new threshold, the protocol that best matches the current task load operation status is selected from the metadata indexing protocol, multicast synchronization protocol and hybrid response protocol.
[0075] In this embodiment, the historical threshold information is the value data of the maximum threshold and the minimum threshold in the specified time period before the current moment. For example, the historical threshold information can be the average value of the sum of all thresholds in this period. The attenuation adjustment information can be a fixed attenuation constant value, or a value that can be flexibly adjusted according to actual conditions, or a value that changes linearly with time. This does not affect the implementation of the present invention. In addition, the maximum and minimum values can be further adjusted by the decrease rate of the delay threshold, such as reducing the decrease value from 5 to 4.8 to obtain the maximum and minimum thresholds. The adjustment factor can be set for different features according to actual conditions and experience. The protocol matching degree value is a value generated by integrating the entropy weights of different features and the corresponding adjustment coefficients according to the operation rules.
[0076] Exemplarily, this embodiment also provides a method for adaptively adjusting the maximum new threshold using a sliding window. and the minimum new threshold The determination method is, t represents the moment value in the current time window T, according to Calculate the threshold at different times within the time window , represents the historical average threshold, Indicates the decay constant (default is 0.1) and selects the maximum and minimum values as the maximum new threshold and the minimum new threshold .
[0077] Exemplarily, an exemplary method for calculating the protocol matching degree value may be: determining the read weight data based on the entropy weight of the read operation feature, the read adjustment factor, the number of read operations, and the number of write operations; determining the write weight data based on the entropy weight of the write operation feature, the write adjustment factor, the number of read operations, and the number of write operations; determining the access distribution data based on the entropy weight of the locality attenuation factor feature, the numerical value of the locality attenuation factor feature, and the heat adjustment factor; and determining the protocol matching degree value based on the read weight data, the write weight data, and the access distribution data. As an efficient calculation method, with feature dimensions including read operation features, write operation features, and locality attenuation factor features, a protocol matching degree value calculation formula may be pre-stored, and the protocol matching degree value of the i-th data block Bi may be calculated by calling the protocol matching degree value calculation formula. The protocol matching degree value calculation formula may be expressed as:
[0078] ;
[0079] in, Indicates the protocol matching degree value of the i-th data block Bi, Indicates the number of read operations of the i-th data block Bi, Indicates the number of write operations of the i-th data block Bi, represents the local attenuation factor eigenvalue of the i-th data block Bi, are the entropy weights of the read operation feature, write operation feature, and locality attenuation factor feature, The default values of 1.2, 0.8, and 0.5 are the respective adjustment factors. By using the above adjustment factors or entropy weights as penalty items, the protocol switching for access time-dispersed data is suppressed, reducing the false positive of cold data.
[0080] Exemplarily, the implementation method of determining the protocol by comparing the protocol matching degree value with the threshold size may be: if the protocol matching degree value is greater than the maximum new threshold, the protocol that best matches the task load operation status at the current moment is the multicast synchronization protocol; if the protocol matching degree value is less than the minimum new threshold, the protocol that best matches the task load operation status at the current moment is the metadata indexing protocol; if the protocol matching degree value is greater than or equal to the minimum new threshold and less than or equal to the maximum new threshold, the protocol that best matches the task load operation status at the current moment is the hybrid response protocol.
[0081] For example, when , enable the multicast synchronization protocol suitable for high-read low-frequency scenarios, and reduce metadata query overhead through multicast updates. , enabling the metadata indexing protocol for high-write locality scenarios, mitigating the risk of broadcast storms through a central directory. For other values, a hybrid response protocol is used, dynamically fusing two mechanisms. Writes use the metadata indexing protocol to maintain directory status, while reads probabilistically choose between local caching and metadata verification based on real-time load.
[0082] As can be seen from the above, this embodiment can dynamically select protocols through adaptive thresholds and gradient smoothing, and dynamically adjust thresholds to avoid overfitting caused by static thresholds. Under sudden loads, such as instantaneous write-intensive tasks, it can quickly converge to the metadata indexing protocol and reduce directory query latency. In stable read scenarios, the multicast synchronization protocol is maintained for a long time to reduce metadata overhead, thereby achieving flexible switching of consistency maintenance protocols based on the dynamic characteristics of task loads, responding to load mutations and long-term fluctuations, and reducing performance jitter.
[0083] Based on the above embodiments, in order to further reduce the write operation latency, this embodiment also provides an implementation method for dynamic directory compression, which may include the following contents: if the cache consistency protocol of the heterogeneous computing system is switched to the metadata indexing protocol, the access frequency of each data block of the heterogeneous computing system at the current moment is counted; the directory entries of the target data blocks below the preset access frequency threshold are compressed and stored.
[0084] The metadata indexing protocol of this embodiment uses a lightweight directory tree to query the directory only when necessary, avoiding broadcast storms. At the same time, global lock contention is reduced through hierarchical directories. To further optimize the protocol, dynamic directory compression can also be performed: directory entries of cold data with low access frequency are compressed and stored. A preset access frequency threshold can be set in advance. Directory entries below the preset access frequency threshold are considered to have low access frequency, reducing memory usage. Compared with the traditional directory protocol directory hierarchical compression, metadata overhead is reduced by 30%, and write operation latency is reduced by 15%~20%, effectively reducing write operation latency in high write load scenarios and improving system performance.
[0085] Based on the above embodiments, this embodiment also provides request processing based on a multicast synchronization protocol, which may include the following contents: if the cache consistency protocol of the heterogeneous computing system is switched to a multicast synchronization protocol, when the target computing unit receives a read operation request by listening to the multicast channel, the corresponding target data is read from the private local cache of the target computing unit; after reading the target data, if the cache status of the private local cache is in a non-dirty state, it is marked as expired and an asynchronous refresh operation is performed; if the cache status of the private local cache is in a dirty state, the conflict detection task is triggered.
[0086] The multicast synchronization protocol is suitable for high-read load scenarios. By maximizing cache hit rates, the local cache directly responds to read requests, reducing metadata queries and metadata verification latency, thereby effectively reducing read latency. The multicast synchronization protocol requires each computing unit to perform corresponding operations based on the information monitored on the multicast channel, and the multicast range is dynamically adjusted based on the cache hit rate. In other words, the multicast synchronization protocol uses probabilistic multicasting, with the probability of dynamically adjusting the multicast range based on the computing unit's cache hit rate. Secondly, it employs listener-based cache verification, meaning that computing units within the multicast range monitor the multicast channel and perform corresponding operations upon receiving update notifications. For example, if the local cache is in the "not dirty" state, meaning the data in the cache line has not been modified, it is directly marked as expired and asynchronously refreshed. If the local cache is in the "dirty" state, meaning the data in the cache line has been modified but not yet written back to main memory, conflict detection mechanisms, such as version number comparison, are triggered to prioritize the latest version. Compared to traditional directory protocols, this multicast synchronization protocol eliminates directory lookups for read operations, resulting in lower latency, reducing read latency by 25% to 30%. Compared to traditional full-broadcast protocols, probabilistic multicast reduces redundant traffic, lowering network load by 10% to 15%, optimizing bandwidth, and improving bandwidth utilization by 20%.
[0087] Based on the above embodiment, this embodiment further provides request processing based on a hybrid response protocol, which may include the following:
[0088] If the cache consistency protocol of the heterogeneous computing system is switched to a hybrid response protocol, when a write operation request is received, the consistency of the private local cache of each computing unit of the heterogeneous computing system is maintained according to the metadata indexing protocol during the processing of the write operation request; when a read operation request is received, if the task load of the heterogeneous computing system is greater than the preset load threshold, the consistency of the private local cache of each computing unit of the heterogeneous computing system is maintained according to the metadata indexing protocol during the processing of the read operation request; if the task load of the heterogeneous computing system is less than or equal to the preset load threshold, the consistency of the private local cache of each computing unit of the heterogeneous computing system is maintained according to the multicast synchronization protocol during the processing of the read operation request.
[0089] In this embodiment, the hybrid response protocol is a fusion of a directory protocol and a broadcast protocol. This protocol decouples read and write strategies: write operations utilize the metadata indexing protocol's directory lock mechanism to ensure consistency. For read operations, a policy is dynamically selected based on real-time load. If the current load is low, the local cache of the multicast synchronization protocol is used to directly respond. If the load is high, metadata verification is triggered to avoid multicast congestion. Load levels can be measured using pre-set thresholds. A state migration engine maintains a protocol state machine for data blocks, which records the cache status of all data blocks, including modified, exclusive, dirty, invalid, and shared. This cache state ensures data validity and, in turn, data consistency, ensuring consistent copies of the same data across different local private caches in heterogeneous computing systems. This protocol dynamically balances mixed read and write loads, adapting to rapid load changes. The state migration engine prevents cache invalidation during protocol switching, ensuring smooth transitions and avoiding performance fluctuations. Compared to traditional directory protocols, this dynamic fusion of directory and broadcast mechanisms avoids the limitations of a single protocol. Compared to traditional broadcast protocols, it supports on-demand switching and offers greater compatibility.
[0090] It should be noted that there is no strict order in which the steps in the present invention are performed. As long as they conform to a logical order, the steps can be performed simultaneously or in a predetermined order. Figure 2 This is just a schematic and does not mean that this is the only execution order.
[0091] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. The present invention also provides a corresponding device for the cache consistency maintenance method, which further makes the method more practical. Among them, the device can be described from the perspective of functional modules and hardware. The cache consistency maintenance device provided by the present invention is introduced below. The following description will introduce the functions of each program module of this embodiment. The cache consistency maintenance device described below and the cache consistency maintenance method described above can be referenced to each other.
[0092] From the perspective of functional modules, see Figure 5 , Figure 5 This is a structural diagram of a cache consistency maintenance device provided in this embodiment in a specific implementation manner. The device may include:
[0093] The protocol abstraction module 501 is used to redefine the cache consistency protocol of the heterogeneous computing system into a metadata indexing protocol, a multicast synchronization protocol, and a hybrid response protocol; wherein the metadata indexing protocol is a protocol in which each computing unit includes directory information for storing local metadata, and the target data block that meets the condition that the write frequency is greater than a preset frequency threshold has central directory information; the multicast synchronization protocol is a protocol in which each computing unit performs corresponding operations based on the monitoring information of the multicast channel, and the multicast range is dynamically adjusted according to the cache hit rate; the hybrid response protocol is a fusion protocol of the directory protocol and the broadcast protocol.
[0094] The protocol dynamic selection module 502 is used to determine the protocol that best matches the current task load operation status from the metadata indexing protocol, multicast synchronization protocol and hybrid response protocol based on the entropy weights of different feature dimensions determined by the feature values of each data block in the heterogeneous computing system; the feature dimensions include at least read and write operation features and data heat features.
[0095] Exemplarily, in some implementations of this embodiment, the above-mentioned protocol dynamic selection module 502 can also be used for: if the entropy weight of the read operation feature at the current moment is greater than the entropy weight of other feature dimensions, then the multicast synchronization protocol is selected as the protocol that best matches the task load running state of the heterogeneous computing system at the current moment; if the entropy weight of the write operation feature at the current moment is greater than the entropy weight of other feature dimensions, then the metadata indexing protocol is selected as the protocol that best matches the task load running state of the heterogeneous computing system at the current moment; if the entropy weight of the data heat feature at the current moment is greater than the entropy weight of other feature dimensions, then according to the data heat change trend, the protocol that best matches the task load running state at the current moment is determined from the metadata indexing protocol, the multicast synchronization protocol and the hybrid response protocol.
[0096] As an illustrative implementation of the above embodiment, the above-mentioned protocol dynamic selection module 502 can also be further used for: for the case where the data heat characteristics include access frequency characteristics and local attenuation factor characteristics, the local attenuation factor characteristics are determined by the access time distribution characteristics of the data block. If the entropy weight of the access frequency characteristics is greater than the entropy weight of the local attenuation factor characteristics, the metadata index protocol is selected as the protocol that best matches the task load running state of the heterogeneous computing system at the current moment; if the entropy weight of the local attenuation factor characteristics is greater than the entropy weight of the access frequency characteristics, the local attenuation factor characteristics are determined according to the access time distribution characteristics of the data block, and according to the access time distribution characteristics, the protocol that best matches the task load running state at the current moment is determined.
[0097] As another illustrative implementation of the above embodiment, the above-mentioned protocol dynamic selection module 502 can also be further used to: determine multiple access time interval values based on multiple consecutive access moments of the current data block; if each access time interval value meets the preset same similarity condition, then select the multicast synchronization protocol as the protocol that best matches the task load operation status of the heterogeneous computing system at the current moment; if each access time interval value meets the preset post-burst access restriction condition, then generate protocol switching inhibition prompt information; if each access time interval value meets the preset medium fluctuation access condition, then select the hybrid response protocol as the protocol that best matches the task load operation status of the heterogeneous computing system at the current moment.
[0098] As another illustrative implementation of the above embodiment, the above protocol dynamic selection module 502 can be further used to: obtain at least four consecutive access times of each data block, and calculate the access time interval value between each two adjacent access times; and use the ratio of the mean and standard deviation of each access time interval value as a local attenuation factor feature.
[0099] As an illustrative implementation of the above embodiment, the above protocol dynamic selection module 502 can be further used to: if the value of the local attenuation factor characteristic is approximately equal to 0, then select the multicast synchronization protocol as the protocol that best matches the task load operation status of the heterogeneous computing system at the current moment; if the value of the local attenuation factor characteristic is greater than 1, then generate protocol switching inhibition prompt information; if the value of the local attenuation factor characteristic is approximately equal to 1, then select the hybrid response protocol as the protocol that best matches the task load operation status of the heterogeneous computing system at the current moment.
[0100] Exemplarily, in some other implementations of this embodiment, the above-mentioned protocol dynamic selection module 502 can also be used for: the calculation process of the information entropy value of the current feature: for each data block, based on the ratio of the original feature value of the current feature of the current data block and the sum of the original feature values of the current features of all data blocks of the heterogeneous computing system, determine the feature probability distribution information of the current feature of the current data block; based on the feature probability distribution information of the current features of each data block of the heterogeneous computing system, determine the information entropy value of the current feature; based on the proportional relationship between the information entropy value of the current feature and the information entropy values of all features, determine the entropy weight of each feature.
[0101] As an illustrative implementation method of the above embodiment, the above protocol dynamic selection module 502 can also be further used to: calculate the product of the feature probability distribution information of the current data block in the current feature and its logarithmic function value; and use the negative value of the sum of the product values corresponding to each data block of the heterogeneous computing system as the information entropy value of the current feature.
[0102] As another illustrative implementation method of the above embodiment, the above protocol dynamic selection module 502 can also be further used to: calculate the difference between the information entropy value of each feature and the target value respectively; and use the ratio of the difference corresponding to the current feature to the sum of the differences corresponding to other features as the entropy weight of the current feature.
[0103] Illustratively, in some other implementations of this embodiment, the above-mentioned protocol abstraction module 501 can also be used for: if the cache consistency protocol of the heterogeneous computing system is switched to the metadata indexing protocol, counting the access frequency of each data block of the heterogeneous computing system at the current moment; compressing and storing the directory entries of the target data blocks below the preset access frequency threshold.
[0104] Illustratively, in some other implementations of this embodiment, the above-mentioned protocol abstraction module 501 can also be used for: if the cache consistency protocol of the heterogeneous computing system is switched to a multicast synchronization protocol, when the target computing unit receives a read operation request by monitoring the multicast channel, the corresponding target data is read from the private local cache of the target computing unit; after reading the target data, if the cache status of the private local cache is in a non-dirty state, it is marked as expired and an asynchronous refresh operation is performed; if the cache status of the private local cache is in a dirty state, the conflict detection task is triggered.
[0105] Illustratively, in some other implementations of this embodiment, the above-mentioned protocol abstraction module 501 can also be used for: if the cache consistency protocol of the heterogeneous computing system is switched to a hybrid response protocol, when a write operation request is received, then in the process of processing the write operation request, the consistency of the private local cache of each computing unit of the heterogeneous computing system is maintained according to the metadata indexing protocol; when a read operation request is received, if the task load of the heterogeneous computing system is greater than the preset load threshold, then in the process of processing the read operation request, the consistency of the private local cache of each computing unit of the heterogeneous computing system is maintained according to the metadata indexing protocol; if the task load of the heterogeneous computing system is less than or equal to the preset load threshold, then in the process of processing the read operation request, the consistency of the private local cache of each computing unit of the heterogeneous computing system is maintained according to the multicast synchronization protocol.
[0106] Exemplarily, in some other implementations of this embodiment, the above-mentioned protocol dynamic selection module 502 can also be used to: calculate the thresholds at different times within the time window of the current moment based on historical threshold information and attenuation adjustment information, and select the maximum and minimum values therefrom; update the existing maximum threshold at the current moment according to the maximum value, and update the existing minimum threshold at the current moment according to the minimum value, so as to update the existing maximum threshold and minimum threshold at the current moment and obtain the maximum new threshold and the minimum new threshold; determine the protocol matching degree value according to the entropy weights of different features and their corresponding adjustment factors; select the protocol that best matches the task load running status at the current moment from the metadata indexing protocol, the multicast synchronization protocol and the hybrid response protocol according to the numerical relationship between the protocol matching degree value and the maximum new threshold and the minimum new threshold.
[0107] As an illustrative implementation method of the above embodiment, the above-mentioned protocol dynamic selection module 502 can also be further used to: determine the read proportion data based on the entropy weight, read adjustment factor, number of read operations and number of write operations of the read operation feature; determine the write proportion data based on the entropy weight, write adjustment factor, number of read operations and number of write operations of the write operation feature; determine the access distribution data based on the entropy weight of the local attenuation factor feature, the numerical value of the local attenuation factor feature and the heat adjustment factor; determine the protocol matching degree value based on the read proportion data, write proportion data and access distribution data.
[0108] As another illustrative implementation method of the above embodiment, the above-mentioned protocol dynamic selection module 502 can also be further used for: if the protocol matching degree value is greater than the maximum new threshold, the protocol that best matches the task load running state at the current moment is the multicast synchronization protocol; if the protocol matching degree value is less than the minimum new threshold, the protocol that best matches the task load running state at the current moment is the metadata indexing protocol; if the protocol matching degree value is greater than or equal to the minimum new threshold and less than or equal to the maximum new threshold, the protocol that best matches the task load running state at the current moment is the hybrid response protocol.
[0109] The cache consistency maintenance device mentioned above is described from the perspective of functional modules. Furthermore, the present invention also provides an electronic device, which is described from the perspective of hardware. Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present invention in one implementation manner. The electronic device includes a memory 601 and a processor 602. The memory 601 stores a computer program, and the processor 602 is configured to run the computer program to perform the steps of any of the above cache consistency maintenance method embodiments.
[0110] An embodiment of the present application further provides a non-volatile storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above cache consistency maintenance method embodiments when running.
[0111] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0112] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above cache consistency maintenance method embodiments are implemented.
[0113] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned cache consistency maintenance method embodiments are implemented.
[0114] Finally, the present invention also provides a heterogeneous computing system, see Figure 7 The heterogeneous computing system includes at least a first computing node 701, a second computing node 702 and a cache consistency controller 703. The cache consistency controller 703 is connected to the first computing node 701 and the second computing node 702. The first computing node 701 and the second computing node 702 include at least two computing units. The cache consistency controller 703 is used to implement the cache consistency maintenance method described in any of the above cache consistency maintenance method embodiments when executing a computer program. It can dynamically optimize the cache consistency protocol, significantly improve the performance and reliability of the heterogeneous computing system, and provide efficient technical support for intelligent driving, cloud computing, edge computing and other fields.
[0115] Furthermore, to overcome the difficulties in maintaining consistency in heterogeneous computing systems, high protocol complexity, and poor scalability in related technologies, the present invention, based on a dynamic consistency maintenance method that combines bus snooping with a directory protocol, also makes corresponding improvements to the communication structure of each computing unit within the system. This embodiment defines the connection method for each computing unit or node in a heterogeneous computing system and constructs a communication network cluster, which may include the following:
[0116] In heterogeneous computing systems, information exchange is inevitable between computing nodes and computing units. The network topology between computing nodes and computing units, as well as the method of information transmission and synchronization, impacts communication performance. Currently, the common communication architecture in related technologies adopts a centralized approach: the node with the highest computing performance is selected as the central processing node, and the remaining heterogeneous computing nodes are designated as edge nodes. All edge nodes communicate independently with the central processing node, and edge nodes do not communicate directly with each other. Although simple to implement, as the number of computing nodes or computing units expands, the central node must handle an exponential increase in communication requests, leading to increased contention for communication bandwidth, which can easily cause system throughput to decrease and transmission latency to increase. Furthermore, in scenarios with limited physical bandwidth, concurrent access from multiple nodes can easily trigger packet collisions and queue overflows, resulting in the loss of critical gradient synchronization requests or response timeouts. Furthermore, the risk of a single point of failure in the central node can directly lead to a global system crash, hindering mission continuity and disrupting normal business operations.
[0117] In view of this, this embodiment divides the computing units of all computing nodes in a heterogeneous computing system. This embodiment uses two computing nodes as an example. The computing units of the first computing node and the second computing node are divided into multiple clusters according to the computing unit type. The computing units in the same cluster are of the same type and include a master computing unit. Different clusters are connected in a ring topology via their respective master computing units. To avoid description, the leader node selected from the cluster is defined as the master computing unit. Since the computing units in the same cluster are of the same type and have similar performance, a node can be randomly selected as the master computing unit, and the other nodes in the cluster are slave computing units. The master and slave terms are used for distinction only and do not restrict the master-slave relationship between the two.
[0118] As can be seen from the above, the heterogeneous computing system of this embodiment adopts a simple and highly scalable ring topology for inter-cluster communication. Data is transmitted in a fixed order on the ring, without involving large-scale data exchange and synchronization operations, resulting in low communication overhead. Furthermore, the ring topology is well adapted to heterogeneous computing systems of varying sizes. By organizing computing nodes into a ring structure and leveraging ring communication and aggregation operations, it can be applied to artificial intelligence tasks. It can also efficiently achieve global updates of model parameters, accelerate the training process, and improve the efficiency of parallel computing.
[0119] In the above embodiment, computing units of the same type are divided into a cluster, and multiple groups are connected using a ring topology. To further improve communication performance, the present invention also compares the impact of different network topologies on overall performance for computing units within each cluster, and determines the optimal intra-cluster topology, which may include the following:
[0120] In this embodiment, the communication bandwidth of all computing units in the heterogeneous computing system is 2, and if any two nodes are connected, it is recorded as 1 communication bandwidth. The following are used in the computing group: Figure 8 The tree topology shown is similar to Figure 9 In the star topology shown, there is a delay in sending information from the edge computing unit to the local head computing unit, which is also the main computing unit.
[0121] For a tree topology: At time 1, computing units B and C send their respective information to computing unit A, while computing units F and G send their information to computing unit E. This operation can be completed within one unit of time. At time 2, computing units A and E can simultaneously send their information to computing unit 1. Therefore, this operation can be completed within one unit of time. In summary, in a tree topology, it takes two units of time for the edge computing unit to send information to the local head node computing unit. For a star topology: At time 1, considering the total communication bandwidth of 2, to avoid all computing units sending information simultaneously, a rule for information transmission is established. Without loss of generality, information is sent sequentially according to the computing unit number. First, computing units A and B simultaneously send their respective information to computing unit 1, while the other computing units remain static. Therefore, this operation can be completed within one unit of time. At time 2, computing units C and E simultaneously send their information to computing unit 1, while the other computing units remain static. Therefore, the above operations can be completed within 1 unit of time; at time 3, computing unit G sends its information to computing unit 1 and completes the task within 1 unit of time. It can be seen that the star topology requires 3 units of time to complete the transmission of information from the edge computing unit to the local head computing unit. Therefore, under the conditions of the same number of edge computing units and limited bandwidth, the communication latency of the tree topology is more advantageous than that of the star topology. In addition, the tree topology has the following advantages: in terms of cross-domain communication and information transmission, the tree topology can be used to organize and manage nodes and data in cross-domain scenarios, providing an effective information exchange and transmission mechanism; in terms of fault tolerance and redundancy, the tree topology can increase the system's fault tolerance through redundancy and backup. When a computing unit fails, the tree topology can automatically route traffic to other available computing units to ensure system stability and availability; in terms of cross-domain data consistency and synchronization, data consistency and synchronization are important issues in cross-domain scenarios. The tree topology can be used to synchronize and update data between different domains to ensure data consistency. In terms of scalability and maintainability, the hierarchical structure of the tree topology allows for easy addition of new computing units, enabling expansion of heterogeneous computing systems.
[0122] Based on this, in an embodiment of the present invention, the heterogeneous computing system includes at least a first cluster and a second cluster; the first cluster includes at least a first master computing unit and multiple first slave computing units, and the second cluster includes at least a second master computing unit and multiple second slave computing units; the first cluster and the second cluster communicate through the interconnected first master computing unit and the second master computing unit; the first master computing unit and each first slave computing unit are connected according to a tree topology structure, and the second master computing unit and each second slave computing unit are connected according to a tree topology structure.
[0123] In order to make the technical solution of the present invention more clearly understood by those skilled in the art, the present invention also provides an exemplary communication architecture of a heterogeneous computing system, such as Figure 10 As shown, the following may be included: the computing units of the heterogeneous computing system are divided into five clusters, with the main computing units of each cluster being computing unit 1, computing unit 2, computing unit 3, computing unit 4, and computing unit 5. Computing unit 1 in group 1, computing unit 2 in group 2, computing unit 3 in group 3, computing unit 4 in group 4, and computing units in group 5 are connected in a ring topology. Computing nodes in the same group are connected in a multi-layer tree topology, and a main computing unit is selected. For example, if computing units A through G in group 1 are arranged in multiple layers using a binary tree, and computing unit 1 is selected as the leader node, node communication congestion and latency issues can be effectively resolved.
[0124] As can be seen from the above, the heterogeneous computing system of this embodiment adopts a system communication architecture of "local tree structure" + "global ring structure". Through the decoupling design of physical links and logical paths, hierarchical routing optimization is achieved through the tree topology structure to reduce the cross-level communication load; at the same time, redundant links are built with the help of the ring topology to achieve load balancing on the basis of improving the robustness of the topology. It not only avoids the traffic bottleneck and single point failure risk of the central node, but also significantly improves the system's anti-congestion ability through the multi-path transmission mechanism, thereby achieving an optimal balance between data communication efficiency and system reliability. In addition, the internal consistency maintenance protocol of the cluster composed of homogeneous devices in the cluster is easy to implement, and there is no problem of protocol conversion and translation, so the additional performance overhead of the protocol can be reduced; each group selects a main computing unit, which is responsible for receiving and sending global information, significantly reducing the topological complexity and communication traffic under the full interconnection mode. Furthermore, the ring connection between groups is easy to expand, and the protocol synchronization only needs to consider adjacent computing units without considering other computing units, reducing the difficulty of consistency maintenance. In summary, the heterogeneous computing system topology design module combines ring and tree networks to fully leverage their respective strengths and meet the system requirements for various computing scenarios. By properly designing and configuring the network structure, the system's flexibility, scalability, security, performance, and efficiency can be improved, promoting communication, collaboration, data consistency, and load balancing among heterogeneous nodes.
[0125] The above is a detailed introduction to a heterogeneous computing system and its cache consistency maintenance method, device, electronic device, non-volatile storage medium, and computer program product provided by the present invention. The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other. Whether the units and algorithm steps of each example described in each disclosed embodiment are executed in electronic hardware or computer software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, and such implementation should not be considered to exceed the scope of the present invention. Without departing from the principles of the present invention, the present invention can also be improved and modified in a number of ways, and these improvements and modifications also fall within the scope of protection of the present invention.
Claims
1. A cache consistency maintenance method, characterized in that: include: Redefine the cache coherence protocol of heterogeneous computing systems into a metadata indexing protocol, a multicast synchronization protocol, and a hybrid response protocol; Determine, based on entropy weights of different feature dimensions determined by feature values of each data block of the heterogeneous computing system, a protocol that best matches the current task load operation state from among the metadata indexing protocol, the multicast synchronization protocol, and the hybrid response protocol; the feature dimensions include at least read and write operation features and data heat features; The metadata indexing protocol is a protocol in which each computing unit includes directory information for storing local metadata, and target data blocks that meet the condition that the write frequency is greater than a preset frequency threshold have central directory information; the multicast synchronization protocol is a protocol in which each computing unit performs corresponding operations based on the monitoring information of the multicast channel, and the multicast range is dynamically adjusted according to the cache hit rate; the hybrid response protocol is a fusion protocol of the directory protocol and the broadcast protocol; The step of determining the protocol that best matches the current task load running state from among the metadata indexing protocol, the multicast synchronization protocol, and the hybrid response protocol includes: If the entropy weight of the read operation feature at the current moment is greater than the entropy weights of other feature dimensions, the multicast synchronization protocol is selected as the protocol that best matches the task load running state of the heterogeneous computing system at the current moment; if the entropy weight of the write operation feature at the current moment is greater than the entropy weights of other feature dimensions, the metadata indexing protocol is selected as the protocol that best matches the task load running state of the heterogeneous computing system at the current moment; if the entropy weight of the data heat feature at the current moment is greater than the entropy weights of other feature dimensions, then according to the data heat change trend, the protocol that best matches the task load running state at the current moment is determined from the metadata indexing protocol, the multicast synchronization protocol and the hybrid response protocol.
2. The cache consistency maintenance method according to claim 1, wherein: The data heat characteristics include access frequency characteristics and locality attenuation factor characteristics. The locality attenuation factor characteristics are determined based on the access time distribution characteristics of the data block. Based on the data heat change trend, the protocol that best matches the current task load operation status is determined from the metadata indexing protocol, the multicast synchronization protocol, and the hybrid response protocol, including: If the entropy weight of the access frequency feature is greater than the entropy weight of the locality attenuation factor feature, selecting the metadata indexing protocol as the protocol that best matches the task load running state of the heterogeneous computing system at the current moment; If the entropy weight of the local attenuation factor feature is greater than the entropy weight of the access frequency feature, a protocol that best matches the task load running state at the current moment is determined based on the access time distribution feature.
3. The cache consistency maintenance method according to claim 2, wherein: Determining a protocol that best matches the current task load running state based on the access time distribution characteristics includes: Determine multiple access time interval values according to multiple consecutive access moments of the current data block; If the access time interval values satisfy the preset same similarity condition, selecting the multicast synchronization protocol as the protocol that best matches the task load running state of the heterogeneous computing system at the current moment; If the access time interval values meet the preset post-burst access restriction condition, a protocol switching inhibition prompt message is generated; If the access time interval values meet the preset medium fluctuation access condition, the hybrid response protocol is selected as the protocol that best matches the task load running state of the heterogeneous computing system at the current moment.
4. The cache consistency maintenance method according to claim 2, wherein: The locality attenuation factor feature is determined according to the access time distribution feature of the data block, including: Obtain at least four consecutive access times of each data block, and calculate the access time interval between every two adjacent access times; The ratio of the mean value and the standard deviation of each access time interval value is used as the local attenuation factor feature.
5. The cache consistency maintenance method according to claim 4, characterized in that: The local attenuation factor characteristic is the standard deviation of each access time interval value divided by the corresponding mean. According to the access time distribution characteristic, determining the protocol that best matches the current task load running state includes: If the value of the locality attenuation factor characteristic is approximately equal to 0, selecting the multicast synchronization protocol as the protocol that best matches the task load running state of the heterogeneous computing system at the current moment; If the value of the local attenuation factor feature is greater than 1, generating protocol switching inhibition prompt information; If the value of the local attenuation factor characteristic is approximately equal to 1, the hybrid response protocol is selected as the protocol that best matches the task load running state of the heterogeneous computing system at the current moment.
6. The cache consistency maintenance method according to claim 1, wherein: The entropy weights of different feature dimensions determined based on the feature values of each data block in the heterogeneous computing system include: The information entropy value calculation process for the current feature is as follows: for each data block, based on the ratio of the original feature value of the current feature of the current data block to the sum of the original feature values of the current features of all data blocks in the heterogeneous computing system, determining the feature probability distribution information of the current feature of the current data block; and determining the information entropy value of the current feature based on the feature probability distribution information of the current features of each data block in the heterogeneous computing system. The entropy weight of each feature is determined based on the proportional relationship between the information entropy value of the current feature and the information entropy values of all features.
7. The cache consistency maintenance method according to claim 6, characterized in that: Determining, based on feature probability distribution information of current features of each data block of the heterogeneous computing system, an information entropy value of the current feature, including: Calculate the product of the feature probability distribution information of the current data block at the current feature and its logarithmic function value; The negative value of the sum of the product values corresponding to the data blocks of the heterogeneous computing system is used as the information entropy value of the current feature.
8. The cache consistency maintenance method according to claim 6, wherein: The entropy weight of each feature is determined based on the proportional relationship between the information entropy value of the current feature and the information entropy value of all features, including: Calculate the difference between the information entropy value of each feature and the target value respectively; The ratio of the difference value corresponding to the current feature to the sum of the difference values corresponding to other features is used as the entropy weight of the current feature.
9. The cache consistency maintenance method according to claim 1, wherein: After determining the protocol that best matches the current task load running state from among the metadata indexing protocol, the multicast synchronization protocol, and the hybrid response protocol, the method further includes: If the cache coherence protocol of the heterogeneous computing system is switched to the metadata indexing protocol, counting the access frequency of each data block of the heterogeneous computing system at the current moment; The directory entries of the target data blocks whose access frequency is lower than a preset threshold are compressed and stored.
10. The cache consistency maintenance method according to claim 1, wherein: After determining the protocol that best matches the current task load running state from among the metadata indexing protocol, the multicast synchronization protocol, and the hybrid response protocol, the method further includes: If the cache coherence protocol of the heterogeneous computing system is switched to the multicast synchronization protocol, when the target computing unit receives a read operation request by monitoring the multicast channel, the corresponding target data is read from the private local cache of the target computing unit; After reading the target data, if the cache state of the private local cache is not dirty, it is marked as expired and an asynchronous refresh operation is performed; if the cache state of the private local cache is dirty, a conflict detection task is triggered.
11. The cache consistency maintenance method according to claim 1, wherein: After determining the protocol that best matches the current task load running state from among the metadata indexing protocol, the multicast synchronization protocol, and the hybrid response protocol, the method further includes: If the cache coherence protocol of the heterogeneous computing system is switched to the hybrid response protocol, when a write operation request is received, the coherence of the private local cache of each computing unit of the heterogeneous computing system is maintained according to the metadata indexing protocol during processing of the write operation request; When a read operation request is received, if the task load of the heterogeneous computing system is greater than a preset load threshold, then in the process of processing the read operation request, the consistency of the private local caches of the heterogeneous computing system is maintained according to the metadata indexing protocol; if the task load of the heterogeneous computing system is less than or equal to the preset load threshold, then in the process of processing the read operation request, the consistency of the private local caches of the computing units of the heterogeneous computing system is maintained according to the multicast synchronization protocol.
12. The cache consistency maintenance method according to any one of claims 1 to 11, characterized in that: Determining, based on entropy weights of different feature dimensions determined by feature values of each data block of the heterogeneous computing system, a protocol that best matches the current task load running state from among the metadata indexing protocol, the multicast synchronization protocol, and the hybrid response protocol, including: Based on the historical threshold information and attenuation adjustment information, the thresholds at different times within the time window of the current moment are calculated, and the maximum and minimum values are selected from them; The maximum threshold value currently existing at the moment is updated according to the maximum value, and the minimum threshold value currently existing at the moment is updated according to the minimum value, so as to obtain a new maximum threshold value and a new minimum threshold value; Determine the protocol matching degree value based on the entropy weights of different features and their corresponding adjustment factors; According to the numerical relationship between the protocol matching degree value and the maximum new threshold and the minimum new threshold, the protocol that best matches the task load running status at the current moment is selected from the metadata indexing protocol, the multicast synchronization protocol and the hybrid response protocol.
13. The cache consistency maintenance method according to claim 12, characterized in that: The protocol matching degree value is determined based on the entropy weights of different features and their corresponding adjustment factors, including: Determine the read proportion data based on the entropy weight of the read operation feature, the read adjustment factor, the number of read operations, and the number of write operations; Determine write proportion data based on the entropy weight of the write operation feature, the write adjustment factor, the number of read operations, and the number of write operations; Determining access distribution data according to the entropy weight of the locality attenuation factor feature, the value of the locality attenuation factor feature, and the heat adjustment factor; A protocol matching degree value is determined according to the read weight data, the write weight data, and the access distribution data.
14. The cache consistency maintenance method according to claim 12, wherein: Selecting, according to a numerical relationship between the protocol matching degree value and the maximum new threshold and the minimum new threshold, a protocol that best matches the current task load running state from among the metadata indexing protocol, the multicast synchronization protocol, and the hybrid response protocol, including: If the protocol matching degree value is greater than the maximum new threshold, the protocol that best matches the task load running state at the current moment is the multicast synchronization protocol; If the protocol matching degree value is less than the minimum new threshold, the protocol that best matches the task load running state at the current moment is the metadata index protocol; If the protocol matching degree value is greater than or equal to the minimum new threshold and less than or equal to the maximum new threshold, the protocol that best matches the task load running state at the current moment is the hybrid response protocol.
15. A cache consistency maintenance device, characterized in that: include: A protocol abstraction module is used to redefine the cache coherence protocol of heterogeneous computing systems into a metadata indexing protocol, a multicast synchronization protocol, and a hybrid response protocol. The metadata indexing protocol is a protocol in which each computing unit includes directory information for storing local metadata, and target data blocks that meet the condition that the write frequency exceeds a preset frequency threshold have central directory information. The multicast synchronization protocol is a protocol in which each computing unit performs corresponding operations based on monitoring information of the multicast channel, and the multicast range is dynamically adjusted according to the cache hit rate. The hybrid response protocol is a fusion protocol of the directory protocol and the broadcast protocol. A protocol dynamic selection module is configured to determine, based on the entropy weights of different feature dimensions determined by the feature values of each data block of the heterogeneous computing system, a protocol that best matches the current task load operation state from the metadata indexing protocol, the multicast synchronization protocol, and the hybrid response protocol; the feature dimensions include at least read and write operation features and data heat features; Among them, the protocol dynamic selection module is further used to: if the entropy weight of the read operation feature at the current moment is greater than the entropy weight of other feature dimensions, then select the multicast synchronization protocol as the protocol that best matches the task load running state of the heterogeneous computing system at the current moment; if the entropy weight of the write operation feature at the current moment is greater than the entropy weight of other feature dimensions, then select the metadata indexing protocol as the protocol that best matches the task load running state of the heterogeneous computing system at the current moment; if the entropy weight of the data heat feature at the current moment is greater than the entropy weight of other feature dimensions, then according to the data heat change trend, determine the protocol that best matches the task load running state at the current moment from the metadata indexing protocol, the multicast synchronization protocol and the hybrid response protocol.
16. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the cache consistency maintenance method according to any one of claims 1 to 14 when executing the computer program.
17. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a computer program, which, when executed by a processor, implements the steps of the cache consistency maintenance method according to any one of claims 1 to 14.
18. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the cache consistency maintenance method according to any one of claims 1 to 14 are implemented.
19. A heterogeneous computing system, characterized in that: At least comprising a first computing node, a second computing node and a cache consistency controller, wherein the cache consistency controller is connected to the first computing node and the second computing node, and the first computing node and the second computing node comprise at least two computing units; The cache coherence controller is configured to implement the steps of the cache coherence maintenance method according to any one of claims 1 to 14 when executing a computer program.
20. The heterogeneous computing system according to claim 19, wherein: The computing units of the first computing node and the computing units of the second computing node are divided into a plurality of clusters according to the computing unit type; The computing units in the same cluster are of the same type and include a main computing unit. Different clusters are connected through their respective main computing units in a ring topology.
21. The heterogeneous computing system according to claim 20, wherein: including at least a first cluster and a second cluster; The first cluster includes at least a first master computing unit and a plurality of first slave computing units, and the second cluster includes at least a second master computing unit and a plurality of second slave computing units; The first cluster and the second cluster communicate via the interconnected first main computing unit and the second main computing unit; The first master computing unit is connected to each first slave computing unit in a tree topology structure, and the second master computing unit is connected to each second slave computing unit in a tree topology structure.
Citation Information
Patent Citations
Cache management method and system of server
CN117290259A
Highly extensible cache coherency protocol directory design structure
CN117827697A