Inter-die coherence processing system, method and apparatus, device, and medium
By setting a consistency judgment module and signature table in the chip to determine the target chip, the problem of low consistency efficiency across multiple Die in the prior art is solved, and chip performance and storage resource utilization are improved.
Patent Information
- Application Number
- PCT/CN2024/097816
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-03
- Filing Date
- 2024-06-06
- Publication Date
- 2025-05-08
AI Technical Summary
When processing chip consistency, the prior art requires judgment across multiple dies, resulting in too low processing efficiency and large calculation amount, resulting in a reduced chip performance.
A chip consistency processing system is designed, including at least one first chip and a chipset. The first chip is provided with a consistency judgment module, and a signature table that stores signature information of all the second chips, for determining the target chip that needs to be synchronized.
By avoiding broadcasting chip consistency requests to all other chips, bandwidth consumption across chip transmission is reduced, performance of multi-core processors is improved, and utilization of on-chip storage resources is improved.
Smart Images

Figure CN2024097816_08052025_PF_FP_ABST
Abstract
Description
Chip consistency processing system, method, device, equipment and medium thereof
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese Patent Application No. 2023114524798 filed in China on November 3, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure relates to the field of chip technology, and in particular to a chip consistency processing system, and its method, apparatus, device, storage medium, computer program product, and computer program. Background Art
[0004] The rapid development of applications such as artificial intelligence, high-performance computing, and networking has led to an increasing demand for high-performance processors. Multi-core processors are the development direction of high-performance processors, and the trend is towards an increasing number of processor cores on a chip, as well as a growing variety of processor cores, including CPU (Central Processing Unit) cores, AI (Artificial Intelligence) cores, GPU (Graphics Processing Unit) cores, DPU (Data Processing Unit) cores, DSP (Digital Signal Processor) cores, and other accelerator cores. However, in related technologies, due to the large cache capacity on each die in a chiplet system, the bandwidth for cross-die access is limited, and the latency of cross-die access is much greater than the latency of intra-die access. As a result, when processing chip consistency, judgments need to be made across multiple dies, resulting in low processing efficiency and a large amount of computation, which reduces chip performance.
[0005] Summary of the Invention
[0006] Embodiments of the present disclosure provide a chip consistency processing system, and a method, apparatus, device, storage medium, computer program product, and computer program thereof.
[0007] In a first aspect, an embodiment of the present disclosure provides a chip consistency processing system, comprising at least one first chip and a chipset, wherein:
[0008] The chipset includes a plurality of second chips;
[0009] The first chip is provided with a consistency judgment module, which stores a signature table of signature information of all second chips and is used to determine the target chip that needs to be synchronized using the signature information and the received consistency request; and
[0010] The consistency judgment module is further used to:
[0011] Based on the consistency request, determining a data address corresponding to the consistency request;
[0012] Performing a hash calculation on the data address to obtain a hash value corresponding to the data address;
[0013] Mapping the hash value into the signature table to obtain a hash value mapping result for each chip; and
[0014] Based on the hash value mapping result, the chip containing the data address is determined to be the target chip.
[0015] In some embodiments, the second chip is one or more of the following:
[0016] Central processing unit chip CPU Die, graphics processing unit chip GPU Die, artificial intelligence chip AIDie, data processor chip DPU Die, digital signal processor chip DSP Die and accelerator chip Accelerator Die.
[0017] In some embodiments, the first chip is an input / output chip (IO Die).
[0018] The chip consistency processing system provided by the embodiment of the present disclosure establishes a first chip and sets a consistency judgment module in the first chip to determine the target chip that needs to be synchronized in the consistency request, thereby avoiding broadcasting the chip consistency request to all other chips, reducing the bandwidth consumption during cross-chip transmission, thereby improving the performance of the multi-core processor and improving the utilization of storage resources on the chip.
[0019] In a second aspect, an embodiment of the present disclosure provides a chip consistency processing method, which is applied to the first chip of the system according to the embodiment of the first aspect above, wherein the method includes:
[0020] Obtaining a consistency request sent by any chip, where the consistency request is used to request synchronization of data of the chip sending the consistency request and multiple second chips;
[0021] Based on the consistency request, determining the target chip that needs to be synchronized; and
[0022] The consistency request is sent to the target chip, so that the target chip synchronizes data according to the consistency request.
[0023] In some embodiments, determining a target chip that needs to be synchronized based on a coherence request includes:
[0024] Based on the consistency request, determining a data address corresponding to the consistency request; and
[0025] From the pre-stored signature table, the chip containing the data address is determined to be the target chip.
[0026] In some embodiments, determining a chip containing a data address as a target chip from a pre-stored signature table includes:
[0027] Perform hash calculation on the data address to obtain the hash value corresponding to the data address;
[0028] Map the hash value in the signature table to obtain the hash value mapping result for each chip; and
[0029] Based on the hash value mapping result, the chip containing the data address is determined to be the target chip.
[0030] In some embodiments, determining a chip containing a data address as a target chip from a pre-stored signature table includes:
[0031] Determining the chip set containing the data address from a pre-stored signature table; and
[0032] The data address is sent to the relay chip corresponding to the chipset, so that the relay chip determines the target chip containing the data address in the chipset according to the data address. The relay chip is used to determine the target chip among a preset number of second chips according to the data address.
[0033] In some embodiments, the signature table is a Bloom filter or an improved Counting Bloom filter.
[0034] The present disclosure provides a chip consistency processing method, including:
[0035] First, obtain the consistency request sent by any chip, then determine the target chip that needs to be synchronized based on the consistency request, and finally send the consistency request to the target chip so that the target chip synchronizes data according to the consistency request. This solves the technical problem that when processing chip consistency, judgments need to be made across multiple Dies, which results in low processing efficiency and a large amount of calculation, resulting in low chip performance utilization. The method provided by the embodiment of the present disclosure uniformly determines the addresses where consistency requests exist on other chips, and then sends the consistency request to the chip corresponding to the address where the consistency request exists, avoiding broadcasting the chip's consistency request to all other chips, reducing bandwidth consumption during cross-chip transmission, thereby improving the performance of multi-core processors and improving the utilization of storage resources on the chip.
[0036] In a third aspect, an embodiment of the present disclosure provides a first chip of a system based on the embodiment of the first aspect above, implementing a chip consistency processing device, the device comprising:
[0037] An acquiring unit, configured to acquire a consistency request sent by any chip, wherein the consistency request is used to request synchronization of data of the chip sending the consistency request and a plurality of second chips;
[0038] a determination unit, configured to determine a target chip that needs to be synchronized based on the consistency request; and
[0039] The sending unit is used to send a consistency request to the target chip, so that the target chip synchronizes data according to the consistency request.
[0040] In some embodiments, the determining unit is specifically configured to:
[0041] Based on the consistency request, determining a data address corresponding to the consistency request; and
[0042] From the pre-stored signature table, the chip containing the data address is determined to be the target chip.
[0043] In some embodiments, the determining unit is specifically configured to:
[0044] Perform hash calculation on the data address to obtain the hash value corresponding to the data address;
[0045] Map the hash value in the signature table to obtain the hash value mapping result for each chip; and
[0046] Based on the hash value mapping result, the chip containing the data address is determined to be the target chip.
[0047] In some embodiments, the determining unit is specifically configured to:
[0048] Determining the chip set containing the data address from a pre-stored signature table; and
[0049] The data address is sent to the relay chip corresponding to the chipset, so that the relay chip determines the target chip containing the data address in the chipset according to the data address. The relay chip is used to determine the target chip among a preset number of second chips according to the data address.
[0050] In some embodiments, the signature table is a Bloom filter or an improved Counting Bloom filter.
[0051] In a fourth aspect, an embodiment of the present disclosure provides an electronic device, including:
[0052] Memory;
[0053] processor; and
[0054] computer programs;
[0055] The computer program is stored in the memory and is configured to be executed by the processor to implement the chip consistency processing method as described in the second embodiment above.
[0056] In a fifth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the chip consistency processing method of the embodiment of the second aspect are implemented.
[0057] In a sixth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the chip consistency processing method of the embodiment of the second aspect above.
[0058] In a seventh aspect, an embodiment of the present disclosure provides a computer program, comprising computer program code, which, when executed on a computer, enables the computer to execute the steps of the chip consistency processing method of the second aspect embodiment. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0061] FIG1 is a schematic diagram of the structure of a chip consistency system provided by an embodiment of the present disclosure;
[0062] FIG2 is a schematic diagram of the structure of a multi-core processor provided by an embodiment of the present disclosure;
[0063] FIG3 is a schematic diagram of the structure of a 64-core processor provided by an embodiment of the present disclosure;
[0064] FIG4 is a schematic diagram of the structure of a CPU Die provided by an embodiment of the present disclosure;
[0065] FIG5 is a schematic diagram of a flow chart of a chip consistency processing method provided by an embodiment of the present disclosure;
[0066] FIG6 is a schematic diagram of a specific flow chart of a chip consistency processing method provided by an embodiment of the present disclosure;
[0067] FIG7 is a schematic diagram of a Bloom Filter determination principle provided by an embodiment of the present disclosure;
[0068] FIG8 is a schematic diagram of a specific flow chart of another chip consistency processing method provided by an embodiment of the present disclosure;
[0069] FIG9 is a schematic structural diagram of a chip consistency processing device provided by an embodiment of the present disclosure;
[0070] FIG10 is a schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0071] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features therein can be combined with each other in the absence of conflict.
[0072] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0073] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0074] 1. In the embodiments of the present disclosure, the term "and / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0075] 2. In the embodiments of the present disclosure, the term "CPU" refers to the Central Processing Unit (CPU), which serves as the computing and control core of a computer system and is the final execution unit for information processing and program execution.
[0076] 3. The term "GPU" in this disclosure refers to a graphics processing unit (GPU), also known as a display core, visual processor, or display chip. It is a microprocessor that performs image and graphics-related operations. In fields such as scientific computing and deep learning, GPUs can also be used as computing accelerators, significantly improving computing speed and efficiency.
[0077] 4. The term "Die" in the embodiments of the present disclosure refers to a chip. In an integrated circuit, a die is a small piece of semiconductor material on which a given functional circuit is manufactured. In the present disclosure, a chip refers to a die, which is the smallest unit of a whole piece of semiconductor material chip.
[0078] 5. In the embodiments of the present disclosure, the term "chiplet" refers to a pre-fabricated chip with specific functions that can be combined and integrated, also known as a core particle.
[0079] The rapid development of applications such as artificial intelligence, high-performance computing, and networking is driving increasing demand for the computing power of high-performance processors. Multi-core processors are the development direction of high-performance processors, and the trend is towards an increasing number of processor cores on a chip, as well as a growing variety of core types, including CPU (Central Processing Unit) cores, AI (Artificial Intelligence) cores, GPU (Graphics Processing Unit) cores, DPU (Data Processing Unit) cores, DSP (Digital Signal Processor) cores, and other accelerator cores. This is accompanied by a rapid increase in the design scale and complexity of high-performance processors, resulting in larger chip areas, increased manufacturing difficulties, and increased losses due to reduced yields.
[0080] Through chiplet design, ultra-large chips can be cut into independent small chips according to different functional modules, manufactured separately, and then the dies are interconnected on the package through die-to-die connections to form a single chip. This not only effectively improves yield but also reduces the cost increase caused by defective rates. Moreover, each die in the chiplet system can be manufactured separately using the most suitable process and then assembled into a system chip through advanced packaging. It does not need to be manufactured on a single wafer using the same advanced process. Each die can also be independently upgraded, which can significantly reduce chip manufacturing costs.
[0081] The cache capacity on each die in a chiplet system is large, the bandwidth for cross-die access is limited, and the latency for cross-die access is much greater than the latency for intra-die access. Therefore, how to efficiently achieve multi-core cache coherence in a chiplet system and improve storage resource utilization is a key issue currently receiving widespread attention and research in academia and industry.
[0082] FIG1 is a schematic diagram of the structure of a chip consistency system provided by an embodiment of the first aspect of the present disclosure. As shown in FIG1 , the system includes at least one first chip 101 and a chipset 102. The chipset 102 includes multiple second chips. The first chip 101 is provided with a consistency judgment module. The module stores a signature table containing signature information of each second chip, and is used to use the signature information and the received consistency request to determine the chip that needs to be synchronized in the second chip. The number of first chips 101 can be one or more, and can be set as needed. In some embodiments, the first chip and the second chip can be one or more of a central processing unit chip CPU Die, a graphics processing unit chip GPU Die, an artificial intelligence chip AIDie, a data processing chip DPU Die, a digital signal processor chip DSP Die, and an accelerator chip Accelerator Die. The accelerator includes processors in other fields, other functional accelerators, and other application scenario accelerators. For example, a network processor, an image signal processor, a display processor, an encryption and decryption accelerator, an audio accelerator, a video codec accelerator, various mathematical operation accelerators, a convolution operation accelerator, a matrix operation accelerator, etc.
[0083] In some embodiments, the first chip is an IO Die. For example, in the multi-core processor shown in FIG2 , N (N is a positive integer) processor chips (processor Die) and one input / output chip are included. The N processor chips are the second chip in the chipset, and the input / output chip is the first chip.
[0084] In one example, as shown in Figure 3, in a 64-core processor, the processor die is the CPU die, with a total of eight CPU dies, meaning N is eight. The processor core is the CPU core, and each CPU die, as shown in Figure 4, contains eight processor cores. The first-level cache (L1Cache) and second-level cache (L2Cache) are private caches for the processor cores, and the third-level cache (L3Cache) is a shared last-level cache (LLC). The first-level cache, second-level cache, and third-level cache are all on the CPU die. The IO die of a multi-core processor includes the CPU die's signature component and peripheral interfaces. The peripheral interfaces include memory interfaces, high-speed serial interfaces (PCIE interfaces), and low-speed input and output interfaces (low-speed IO interfaces). The CPU die and IO die are interconnected via a die-to-die interface. All coherence requests from the processor die first enter the IO die for determination, and the IO die receives coherence requests from other dies.
[0085] FIG5 is a flow chart of a chip consistency processing method provided by an embodiment of the second aspect of the present disclosure, which specifically includes the following steps S501 to S503 as shown in FIG5 .
[0086] S501: Obtain a consistency request sent by any chip.
[0087] In some embodiments, a consistency request is first obtained from any chip, which is used to request data synchronization between multiple chips. In the disclosed embodiments, this can be performed by an independent input / output chip (IO Die) or by any other chip in a multi-core processor.
[0088] S502: Determine a target chip that needs to be synchronized based on the consistency request.
[0089] In some embodiments, first, based on the consistency request, the data address corresponding to the consistency request is determined, and then the chip containing the data address is determined to be the target chip from a pre-stored signature table. Specifically, when confirming the target chip, the data address is first hashed to obtain a hash value corresponding to the data address, and then the hash value is mapped in the signature table to obtain a hash value mapping result for each chip, and finally, based on the hash value mapping result, the chip containing the data address is determined to be the target chip. In another embodiment, based on the confirmation of the target chip by multiple chips, hierarchical processing can be performed, and the first processed chip determines in which chipset the target chip is, and then the designated processing chip (relay chip) in the chipset makes further judgments. Specifically, the chipset containing the data address is determined from a pre-stored signature table, and then the data address is sent to the relay chip corresponding to the chipset, so that the relay chip determines the target chip containing the data address in the chipset according to the data address.
[0090] S503: Send a consistency request to the target chip, so that the target chip synchronizes data according to the consistency request.
[0091] In some embodiments, after determining the target chip that needs to perform data synchronization, the consistency request is sent to the target chip to enable the target chip to synchronize data.
[0092] FIG6 is a schematic diagram of a specific flow of a chip consistency processing method provided by an embodiment of the second aspect of the present disclosure, which specifically includes the following steps S601 to S604 as shown in FIG6 .
[0093] S601: Obtain a consistency request sent by any chip.
[0094] In some embodiments, a consistency request initiated by any chip is first obtained, where the consistency request is used to request synchronization of data of multiple chips.
[0095] S602: Based on the consistency request, determine the data address corresponding to the consistency request.
[0096] In some embodiments, the consistency request is parsed to obtain a data address corresponding to the consistency request.
[0097] S603: Determine from a pre-stored signature table that the chip containing the data address is the target chip.
[0098] In some embodiments, by querying the signature mechanism of each processor Die, it is determined whether the address of the consistency request issued by the processor Die is on other processor Dies, and on which processor Dies. If the signature mechanism determines that the address of the consistency request exists on other processor Dies, the consistency request is sent to the processor Die corresponding to the address of the consistency request. This not only reduces the table entries recorded by the IO Die, but also reduces the access to the processor Die, and reduces the bandwidth load of the consistency transmission between Dies, thereby improving the performance of the multi-core processor and reducing power consumption and area.
[0099] In some embodiments, as shown in Figure 7, a consistency request from a first processor is sent to a Bloom filter signature module of a second processor in an I / O chip for verification. The module determines whether the consistency request address is in the second processor. If the second processor is in the second processor, the consistency request is sent to the second processor for processing; otherwise, it is not sent to the second processor. Similarly, the first processor sends its consistency request address to the signature modules of each processor in the I / O chip for verification. If a hit is detected in a specific processor, the consistency request is sent to the corresponding processor for processing; otherwise, it is not sent to the corresponding processor. When performing signature verification, a Bloom filter or an improved counting bloom filter (CBF) can be used. Of course, a combination of Bloom filters and CBFs can also be used. If the signature verification result is yes, the element is not necessarily in the set; however, if the result is no, the element is definitely not in the set. Therefore, an element determined by the Bloom filter to be not in a processor is definitely not in the corresponding processor. Therefore, the Bloom filter signature mechanism can save space, filter out most unnecessary broadcasts to processors, and ensure correctness.
[0100] Specifically, when using Bloom Filter, since the main data structure of Bloom Filter is a bit array consisting of m binary bits, all bits are initially set to 0. At the same time, Bloom Filter also requires k hash functions, which map input elements to k positions in the bit array, and each position is marked as 1. The specific hash function can be any hash algorithm, such as MD5, SHA1, etc., which is not limited in the embodiment of the present disclosure. In response to an element being inserted into the Bloom filter, it is necessary to map the element to k positions in the bit array through k hash functions, and set the values of these positions to 1. In response to querying whether an element is in the Bloom filter, it is only necessary to map the element to k positions in the bit array through k hash functions, and check whether the values of these positions are all 1. If they are all 1, the element may exist in the set, otherwise the element definitely does not exist in the set.
[0101] When using CBF, the main feature of CBF is that it does not only store 0 or 1 in the bit array, but uses an L-bit counter to store the number of times each element appears. Specifically, CBF uses a bit array C of m elements and k hash functions, each hash function can map an element to a position in the bit array. For an element e, the corresponding CBF calculation method is as follows: add 1 to the counters of all positions in C mapped by the hash function. In response to the counter value exceeding a threshold (usually 3), the element is marked as existing in the CBF. In response to querying whether an element is in the CBF, it is necessary to query whether the values of the counters at all positions mapped by the hash function are greater than 0. If they are all greater than 0, it is considered that the element may exist in the CBF. In response to deleting an element, it is necessary to decrement 1 from the counters at all positions mapped by the hash function in the CBF. If the counter at a certain position is reduced to 0, it is considered that the element no longer exists in the CBF.
[0102] S604: Send the consistency request to the target chip, so that the target chip synchronizes data according to the consistency request.
[0103] In some embodiments, after determining the target chip that needs to perform data synchronization, the consistency request is sent to the target chip to enable the target chip to synchronize data.
[0104] FIG8 is a schematic diagram of a specific flow chart of another chip consistency processing method provided by an embodiment of the second aspect of the present disclosure, which specifically includes the following steps S801 to S804 as shown in FIG8 .
[0105] S801. Obtain a consistency request sent by any chip.
[0106] S802: Based on the consistency request, determine the data address corresponding to the consistency request.
[0107] Step S801 and step S802 are the same as step S601 and step S602 and are not described again here.
[0108] S803: Determine from a pre-stored signature table that the chip containing the data address is the target chip.
[0109] In some embodiments, in response to judgments by multiple IO Dies, the first IO Die first determines the Die group corresponding to the data address, and then sends the data address to the relay chip of the Die group. The relay chip can also be another IO Die or another chip, and the embodiment of the present disclosure does not limit this. Otherwise, it will not be sent to the corresponding Die group, and the relay chip of the Die group will calculate the specific processor Die to which it needs to be sent. The two judgments can use the same or different hash value algorithms, and the embodiment of the present disclosure does not limit this.
[0110] S804: Send a consistency request to the target chip, so that the target chip synchronizes data according to the consistency request.
[0111] In some embodiments, after determining the target chip that needs to perform data synchronization, the consistency request is sent to the target chip to enable the target chip to synchronize data.
[0112] FIG9 is a schematic diagram of the structure of a chip consistency processing device provided in an embodiment of the third aspect of the present disclosure. The chip consistency processing device 900 provided in an embodiment of the present disclosure can execute the processing flow provided in the above chip consistency processing method embodiment. As shown in FIG9 , the chip consistency processing device 900 includes an acquisition unit 901, a determination unit 902, and a sending unit 903, wherein:
[0113] An acquiring unit 901 is configured to acquire a consistency request sent by any chip, where the consistency request is used to request synchronization of data of multiple chips;
[0114] A determining unit 902 is configured to determine a target chip that needs to be synchronized based on the consistency request;
[0115] The sending unit 903 is configured to send a consistency request to a target chip, so that the target chip synchronizes data according to the consistency request.
[0116] In some embodiments, the determining unit 902 is specifically configured to:
[0117] Based on the consistency request, determining a data address corresponding to the consistency request; and
[0118] From the pre-stored signature table, the chip containing the data address is determined to be the target chip.
[0119] In some embodiments, the determining unit 902 is specifically configured to:
[0120] Perform hash calculation on the data address to obtain the hash value corresponding to the data address;
[0121] Map the hash value in the signature table to obtain the hash value mapping result for each chip; and
[0122] Based on the hash value mapping result, the chip containing the data address is determined to be the target chip.
[0123] In some embodiments, the determining unit 902 is specifically configured to:
[0124] Determining the chip set containing the data address from a pre-stored signature table; and
[0125] The data address is sent to the relay chip corresponding to the chipset, so that the relay chip determines the target chip containing the data address in the chipset according to the data address. The relay chip is used to determine the target chip among a preset number of second chips according to the data address.
[0126] In some embodiments, the signature table is a Bloom filter or an improved Counting Bloom filter.
[0127] The chip consistency processing device of the embodiment shown in FIG9 can be used to execute the technical solution of the above method embodiment. Its implementation principle and technical effects are similar and will not be described in detail here.
[0128] In addition, the chip consistency processing method and apparatus described in conjunction with Figures 1 to 9 can be implemented by an electronic device. Figure 10 shows a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the fourth aspect of the present disclosure.
[0129] As shown in Figure 10, the electronic device 1000 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1008 into the random access memory (RAM) 1003 to implement the chip consistency processing method of the embodiment described in the embodiment of the present disclosure. In the RAM 1003, various programs and data required for the operation of the electronic device 1000 are also stored. The processing device 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0130] In a fifth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the chip consistency processing method of the embodiment of the second aspect are implemented.
[0131] In a sixth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the chip consistency processing method of the embodiment of the second aspect above.
[0132] In a seventh aspect, an embodiment of the present disclosure provides a computer program, comprising computer program code, which, when executed on a computer, enables the computer to execute the steps of the chip consistency processing method of the second aspect embodiment.
[0133] It should be noted that the explanations of the chip consistency processing method and apparatus in the aforementioned embodiments are also applicable to the computer-readable storage medium, computer program product, and computer program in the embodiments of the present disclosure, and will not be repeated here.
[0134] All embodiments of the present disclosure may be implemented individually or in combination with other embodiments, and are all considered to be within the scope of protection claimed by the present disclosure.
[0135] Typically, the following devices may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device 1000 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 10 illustrates the electronic device 1000 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0136] In some embodiments, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart, thereby implementing the voice control method as described above. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1009, or installed from the storage device 1008, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0137] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In the embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0138] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0139] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0140] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:
[0141] Get the consistency request sent by any chip. The consistency request is used to request data synchronization of multiple chips.
[0142] Based on the consistency request, determining the target chip that needs to be synchronized; and
[0143] The consistency request is sent to the target chip, so that the target chip synchronizes data according to the consistency request.
[0144] In some embodiments, when the above one or more programs are executed by the electronic device, the electronic device may also execute other steps described in the above embodiments.
[0145] Computer program code for performing the operations of the embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0146] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0147] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.
[0148] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0149] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0150] The chip consistency processing system provided by the embodiment of the present disclosure establishes a first chip and sets a consistency judgment module in the first chip to determine the target chip that needs to be synchronized in the consistency request, thereby avoiding broadcasting the chip consistency request to all other chips, reducing the bandwidth consumption during cross-chip transmission, thereby improving the performance of the multi-core processor and improving the utilization of storage resources on the chip.
[0151] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0152] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the function specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0153] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0155] Although the preferred embodiments of the present disclosure have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present disclosure.
[0156] Obviously, those skilled in the art may make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure is intended to include these modifications and variations.
Claims
1. A chip consistency processing system, wherein the system comprises at least one first chip and a chipset, wherein: The chipset includes a plurality of second chips; The first chip is provided with a consistency judgment module, the consistency judgment module stores a signature table of signature information of all second chips, and is used to determine a target chip that needs to be synchronized by using the signature information and the received consistency request; and The consistency judgment module is further used for: Based on the consistency request, determining a data address corresponding to the consistency request; Performing a hash calculation on the data address to obtain a hash value corresponding to the data address; Mapping the hash value into the signature table to obtain a hash value mapping result for each chip; and Based on the hash value mapping result, a chip containing the data address is determined to be a target chip.
2. The system according to claim 1, wherein the second chip is one or more of the following: Central processing unit chip CPU Die, graphics processing unit chip GPU Die, artificial intelligence chip AIDie, data processor chip DPU Die, digital signal processor chip DSP Die and accelerator chip Accelerator Die.
3. The system according to claim 1 or 2, wherein the first chip is an input-output chip IO Die.
4. A chip consistency processing method, applied to the first chip of the system according to any one of claims 1 to 3, wherein the method comprises: Obtaining a consistency request sent by any chip, wherein the consistency request is used to request synchronization of data of the chip sending the consistency request and a plurality of second chips; Based on the consistency request, determining a target chip that needs to be synchronized; and The consistency request is sent to the target chip, so that the target chip synchronizes data according to the consistency request.
5. The method according to claim 4, wherein the determining the target chip that needs to be synchronized based on the consistency request comprises: Based on the consistency request, determining a data address corresponding to the consistency request; and From a pre-stored signature table, a chip containing the data address is determined to be a target chip.
6. The method according to claim 5, wherein determining the chip containing the data address as the target chip from the pre-stored signature table comprises: Performing a hash calculation on the data address to obtain a hash value corresponding to the data address; Mapping the hash value into the signature table to obtain a hash value mapping result for each chip; and Based on the hash value mapping result, a chip containing the data address is determined to be a target chip.
7. The method according to claim 5, wherein determining the chip containing the data address as the target chip from the pre-stored signature table comprises: Determine the chipset containing the data address from a pre-stored signature table; and The data address is sent to the relay chip corresponding to the chipset, so that the relay chip determines the target chip containing the data address in the chipset according to the data address, and the relay chip is used to determine the target chip in a preset number of second chips according to the data address. 8 . The method according to claim 5 , wherein the signature table is a Bloom filter or an improved Counting Bloom filter.
9. Based on the first chip of the system according to any one of claims 1 to 3, a chip consistency processing device is implemented, wherein the device comprises: An acquisition unit, used for acquiring a consistency request sent by any chip, wherein the consistency request is used for requesting data of the chip that sends the consistency request synchronously with a plurality of second chips; A determination unit, configured to determine a target chip that needs to be synchronized based on the consistency request; and The sending unit is used to send the consistency request to the target chip, so that the target chip synchronizes data according to the consistency request.
10. The device according to claim 9, wherein the determining unit is specifically configured to: Based on the consistency request, determining a data address corresponding to the consistency request; and From a pre-stored signature table, a chip containing the data address is determined to be a target chip.
11. The device according to claim 10, wherein the determining unit is specifically configured to: Performing a hash calculation on the data address to obtain a hash value corresponding to the data address; Mapping the hash value into the signature table to obtain a hash value mapping result for each chip; and Based on the hash value mapping result, a chip containing the data address is determined to be a target chip.
12. The device according to claim 10, wherein the determining unit is specifically configured to: Determining a chipset containing the data address from a pre-stored signature table; and The data address is sent to the relay chip corresponding to the chipset, so that the relay chip determines the target chip containing the data address in the chipset according to the data address, and the relay chip is used to determine the target chip in a preset number of second chips according to the data address. 13 . The device according to claim 10 , wherein the signature table is a Bloom filter or an improved Counting Bloom filter.
14. An electronic device comprising: Memory; processor; and Computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the chip consistency processing method according to any one of claims 4 to 8. 15 . A computer-readable storage medium having computer program instructions stored thereon, wherein when the computer program instructions are executed by a processor, the chip consistency processing method according to claim 4 is implemented. 16 . A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the chip consistency processing method according to claim 4 is implemented. 17 . A computer program, comprising computer program codes, which, when executed on a computer, enable the computer to execute the chip consistency processing method according to claim 4 .
Citation Information
Patent Citations
Layering system for achieving caching consistency protocol and method thereof
CN103440223A
Task processing method, chip, multi-chip module, electronic equipment and storage medium
CN116339944A
Chip consistency processing system, method and device thereof, equipment and medium
CN117170986A
Signature generation by a data processing device
US11461175B1
Accelerated recovery for snooped addresses in a coherent attached processor proxy
US20140201468A1
Cited By
Memory allocation and data processing method and device based on dual-Die cross storage
CN120371723A
Inter-DIE cache consistency processing method and data sharing acceleration device
CN121435862A