A method, apparatus, storage medium, and electronic device for determining a chip structure

By obtaining cache and device attribute information, analyzing the chip structure, building a structural diagram and using neural network models, the problem of complex chip optimization of the entire machine manufacturer is solved, and more efficient performance tuning and verification is achieved.

CN119903019BActive Publication Date: 2025-07-08INSPUR (SHANDONG) COMPUTER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510386445.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-08
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

When using complex chips, the entire machine manufacturers lack the details of the internal microarchitecture of the chip, which makes it difficult to optimize targetedly and affect the performance of chips.

Method used

By obtaining cache attribute information and device attribute information, analyzing cache attribute information and kernel characteristics, building a chip structure diagram, and automatically generating logical architecture diagrams using neural network models to help designers understand the internal composition and layout of the chip.

Benefits of technology

It realizes that in the absence of transparent information on chip microarchitecture, it can be more precisely tuned and verified, and improves chip performance and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119903019B_ABST
    Figure CN119903019B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a method, apparatus, storage medium, and electronic device for determining a chip structure, relating to the field of computer technology. The method is applied to an electronic device and includes: obtaining cache attribute information and / or device attribute information; determining chip structure information according to the cache attribute information and / or device attribute information; and constructing a chip structure diagram according to the chip structure information. In this way, by analyzing the cache attribute information and / or device attribute information, the hardware configuration of the chip can be quickly understood, and a structure diagram can be automatically generated based on the hardware configuration, enabling designers to more intuitively understand the internal composition and layout of the chip, so as to perform more accurate optimization and verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, an apparatus, a storage medium, and an electronic device for determining a chip structure. Background Art

[0002] With the continuous progress of chip technologies, especially in the context of the growing high-performance computing requirements in artificial intelligence, big data processing, virtual reality, etc., the design and structure of chips have become increasingly complex. To meet the increasing computing power requirements, chip manufacturers have continuously introduced new architecture innovations in the design, such as multi-core processing, heterogeneous computing, integrating more functional units on the chip, etc. Although the introduction of these technologies has greatly improved the computing power of the processor, it has also made the chip structure more complex.

[0003] However, due to various considerations, chip manufacturers usually do not disclose the internal processes and technical details of the chips. This situation causes difficulties for the complete machine manufacturers in scheduling and optimizing the processes and running processes after using the chips. The lack of sufficient micro-architecture details makes it difficult for the complete machine manufacturers to carry out targeted optimizations according to the actual situation, affecting the maximum performance of the chips. Therefore, a method for analyzing and dissecting the chip micro-architecture is needed. Summary of the Invention

[0004] The present disclosure provides a method, an apparatus, a storage medium, and an electronic device for determining a chip structure to at least solve the above technical problems existing in the prior art.

[0005] The technical solution of the embodiments of the present disclosure is implemented as follows:

[0006] In a first aspect, an embodiment of the present disclosure provides a method for determining a chip structure. The method is applied to an electronic device and includes:

[0007] Obtain cache attribute information and / or device attribute information;

[0008] Determine chip structure information according to the cache attribute information and / or the device attribute information;

[0009] Construct a chip structure diagram according to the chip structure information.

[0010] In the above solution, the determining the chip structure information according to the cache attribute information and / or the device attribute information includes:

[0011] Analyze the cache attribute information and / or the device attribute information to determine cache characteristics and / or core characteristics;

[0012] Analyze the cache characteristics and / or the core characteristics to obtain the chip structure information.

[0013] In the above solution, the cache attribute information includes: cache levels and / or attribute information of each cache level;

[0014] Analyzing the cache attribute information to determine cache characteristics includes:

[0015] Determining information about the first-level instruction cache, first-level data cache, second-level cache, and third-level cache according to the cache levels and / or attribute information of each cache level;

[0016] Among them, the information includes at least one of the following: cache size, number of cache blocks, cache block size.

[0017] In the above solution, the device attribute information includes: slot-related information;

[0018] Analyzing the device attribute information to determine core characteristics includes:

[0019] Determining a first characteristic according to the slot-related information; the first characteristic includes:

[0020] Number of slots, number of processor cores included in each slot, and whether each of the processor cores supports hyper-threading technology.

[0021] In the above solution, the device attribute information includes: non-uniform memory access (NUMA) node information;

[0022] Analyzing the device attribute information to determine core characteristics includes:

[0023] Determining a second characteristic according to the NUMA node information; the second characteristic includes:

[0024] Number of NUMA nodes, number of processor cores included in each NUMA node, and distribution of processor cores within each NUMA node.

[0025] In the above solution, analyzing the cache attribute information and / or the device attribute information to determine cache characteristics and / or core characteristics includes:

[0026] Analyzing the cache attribute information and / or the device attribute information to determine a third characteristic; the third characteristic includes:

[0027] Cache size corresponding to each processor core in each cache level;

[0028] Cache write policies supported by each cache level;

[0029] Whether the NUMA node supports hyper-threading technology.

[0030] In the above solution, the analysis of the cache feature and / or core feature to obtain chip structure information includes:

[0031] Analyze the cache feature and / or core feature based on the first rule to determine the second structure information, where the second structure information includes: the situation of the cache shared among processor cores.

[0032] In the above solution, the analysis of the cache feature and / or core feature to obtain chip structure information includes:

[0033] Analyze the second feature based on the first rule to determine the first structure information, where the first structure information includes: the number of processor groups (CCX, CPU Complex) in the NUMA node, and the number of processor cores in each CCX.

[0034] In the above solution, constructing a chip structure diagram according to the chip structure information includes:

[0035] Use a drawing model to construct a chip structure diagram according to the chip structure information.

[0036] In the above solution, the method further includes:

[0037] Obtain a training data set; the training data set includes at least one sample chip structure information and the logical architecture diagram of the sample chip structure corresponding to each sample chip structure information;

[0038] Train a preset neural network model according to the training data set to obtain a trained neural network model as the drawing model.

[0039] In the above solution, the obtaining of the cache attribute information and / or device attribute information includes:

[0040] Use a test tool to call at least one query instruction, and obtain the cache attribute information and / or device attribute information of the electronic device according to the query instruction.

[0041] In a second aspect, an embodiment of the present disclosure provides a chip structure determination device, which is applied to an electronic device, and the device includes:

[0042] An acquisition module, configured to acquire cache attribute information and / or device attribute information;

[0043] A first processing module, configured to determine chip structure information according to the cache attribute information and / or device attribute information;

[0044] A second processing module, configured to construct a chip structure diagram according to the chip structure information.

[0045] In a third aspect, embodiments of the present disclosure provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein,

[0046] the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any one of the chip structure determination methods.

[0047] In a fourth aspect, embodiments of the present disclosure provide a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the chip structure determination method according to any one of the above.

[0048] Embodiments of the present disclosure have the following beneficial effects:

[0049] By applying the chip structure determination method, device, storage medium, and electronic device provided by the embodiments of the present disclosure, cache attribute information and / or device attribute information are obtained; chip structure information is determined according to the cache attribute information and / or device attribute information; and a chip structure diagram is constructed according to the chip structure information. In this way, by analyzing the cache attribute information and / or device attribute information, the hardware configuration of the chip can be quickly understood, and a structure diagram is automatically generated based on the hardware configuration, enabling designers to more intuitively understand the internal composition and layout of the chip, so as to perform more accurate optimization and verification.

[0050] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a flowchart of a chip structure determination method provided by an embodiment of the present disclosure;

[0052] Figure 2 is a schematic diagram of a chip structure provided by an application embodiment of the present disclosure;

[0053] Figure 3 is a schematic diagram of the structure of a chip structure determination device provided by an embodiment of the present disclosure;

[0054] Figure 4 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] To make the objectives, features, and advantages of the present disclosure more apparent and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present disclosure with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present disclosure.

[0056] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.

[0057] If similar descriptions such as "first / second" appear in the application documents, the following explanation is added. In the following description, the terms "first / second / third" only distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0059] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are explained. The nouns and terms involved in the embodiments of the present disclosure are subject to the following explanations.

[0060] CCX (CPU Complex): Refers to a small set composed of multiple cores in a central processing unit (CPU) chip. Each CCX contains several cores internally. It is a functional module in the CPU architecture for processing multi-threaded tasks and improving computing performance. Generally, a CCX contains one or more cores, and multiple CCXs can be combined to form a more powerful multi-core processor.

[0061] CCD (Core Chiplet Die): Refers to a complete Die. It can be understood that one CCD is one Die.

[0062] Die: Refers to a small chip that is a component of the CPU and is a physical packaging unit. A complete CPU chip is composed of multiple Dies, and each Die independently bears a certain computing power. Each Die can contain multiple functional modules or computing cores and is the smallest physical unit of the CPU chip.

[0063] A CPU chip usually consists of multiple CCDs, and each CCD is an independent Die that contains multiple CCXs. Each CCD is connected to other CCDs through interconnection technology to jointly form a powerful multi-core computing unit.

[0064] In modern high-performance CPUs, multiple Dies are integrated into one chip to improve overall performance and efficiency. Each Die is called a CCD, each CCD contains one or more CCXs, and each CCX contains one or more cores. Through this hierarchical structure, the CPU can flexibly allocate computing resources according to the needs of tasks to achieve a balance between high performance and low power consumption.

[0065] With the continuous increase in computing requirements, especially in the field of high-performance computing, higher demands are placed on the core count, performance, and efficiency of processors. For example, AMD has adopted innovative packaging technologies and architecture designs in its Zen architecture to address these challenges and promote the improvement of its multi-core processing capabilities. The Zen1 architecture first introduced the MCM (Multi-Chip Module) packaging scheme, abandoning the traditional monolithic design, that is, packaging multiple small chips (Dies) into one CPU. This design concept aims to increase the core count and achieve high-concurrency processing capabilities by combining multiple Dies.

[0066] However, although the MCM packaging has made progress in terms of core count, there are significant deficiencies in the communication efficiency between multiple Dies. Due to using off-chip buses for communication between Dies, communication latency and bandwidth bottlenecks have become important factors affecting overall performance, especially in some application scenarios that require frequent access to multiple cores, and the performance of the MCM architecture is not satisfactory. Therefore, although the MCM technology has advantages in increasing the core count and enhancing production flexibility, the communication performance bottlenecks it brings still limit the further improvement of processor performance in some cases.

[0067] To solve this problem, AMD has made relatively significant architectural adjustments in the Zen2 architecture, especially introducing the separation design of the IO Die. Specifically, AMD has separated the IO (Input / Output) function from the core chip (CoreDie) and designed a dedicated independent IO Die. By using the previous-generation process to manufacture the IO Die, AMD can improve the yield rate and reduce costs. At the same time, the core chip can adopt the latest process to operate with higher performance and lower power consumption. In this way, not only is the production efficiency of the Die improved, but the overall cost of chip manufacturing is also effectively reduced. At the same time, the independent design of the IO Die also solves the communication bottleneck problem in the previous MCM architecture, greatly improving the data transfer efficiency between Core Dies, and thus enhancing the overall performance of the CPU.

[0068] However, with the continuous development of the above technologies, although the chip performance has been improved, the structural complexity of the chip has also increased because it involves the coordination of different functional modules (Core Die, IO Die), different manufacturing processes, and optimized transmission paths. The chip is no longer a simple single integrated structure but consists of multiple small modules that need to work together efficiently to achieve stronger performance and lower costs.

[0069] In practical applications, due to various considerations, chip manufacturers usually do not disclose the internal process and technical details of the chip. This situation makes it difficult for OEMs to optimize the scheduling of processes and running processes after using the chip. The lack of sufficient microarchitecture details makes it difficult for OEMs to optimize according to the actual situation, affecting the maximum performance of the chip. Therefore, a method for analyzing and dissecting the chip microarchitecture is needed.

[0070] Based on this, the embodiments of the present disclosure provide a method, device, storage medium, and electronic device for determining a chip structure, obtaining cache attribute information and / or device attribute information; determining chip structure information according to the cache attribute information and / or device attribute information; and constructing a chip structure diagram according to the chip structure information. In this way, by analyzing the cache attribute information and / or device attribute information, the hardware configuration of the chip can be quickly understood, and a structure diagram can be automatically generated based on the hardware configuration, so as to analyze and dissect the chip microarchitecture, enabling designers to more intuitively understand the internal composition and layout of the chip, and thus perform more accurate tuning and verification.

[0071] Figure 1 The flowchart of a method for determining a chip structure provided by the embodiments of the present disclosure is as follows Figure 1 As shown, the method is applied to an electronic device, and the method for determining a chip structure includes:

[0072] Step 101: Obtain cache attribute information and / or device attribute information;

[0073] Step 102: Determine chip structure information based on the cache attribute information and / or device attribute information;

[0074] Step 103: Construct a chip structure diagram according to the chip structure information.

[0075] Here, the cache attribute information may refer to: specific information about the cache part in the chip, such as the size of the cache, type (Level 1 (L1) cache, Level 2 (L2) cache, etc.), and hierarchical structure of the cache.

[0076] The device attribute information may refer to: detailed information about other hardware or devices used by the chip, such as the number of processor cores and supported interfaces.

[0077] Based on the obtained cache attribute information and / or device attribute information, infer the chip structure information. For example, the size and hierarchy of the cache may affect the architecture design of the chip, and device attributes (such as the number of cores and interface types) also help determine how different modules of the chip are interconnected. After determining the chip structure information, visualize this information to construct a chip structure diagram, which shows the layout, connection method, and hierarchical relationship of each component inside the chip.

[0078] In this way, since designers can more intuitively understand the internal composition and layout of the chip, it can help designers understand the internal architecture of the chip and how each component works together, thereby enabling more accurate optimization and verification.

[0079] In some embodiments, the determining the chip structure information based on the cache attribute information and / or device attribute information includes:

[0080] Analyze the cache attribute information and / or the device attribute information to determine cache characteristics and / or core characteristics;

[0081] Analyze the cache characteristics and / or core characteristics to obtain chip structure information.

[0082] Here, cache characteristics refer to specific characteristics of the cache, such as cache size, cache level (L1, L2, etc.), and cache access speed.

[0083] Core characteristics refer to characteristics of the processor core, such as the performance of each core and the number of supported threads.

[0084] By analyzing cache characteristics (such as cache size, levels, etc.) and processor core characteristics (such as the number of cores, etc.), the overall structure and performance of the chip can be further inferred. For example, by analyzing the relationship between the number of processor cores and the cache levels, it can be deduced whether the chip supports multi - cores and whether the cache is shared, etc.

[0085] Thus, by analyzing the information of cache attributes and device characteristics, the overall architecture design of the chip can be determined and inferred, including cache configuration, processor core characteristics, etc. Ultimately, this information helps to understand how the chip works and how to optimize performance.

[0086] In some embodiments, the cache attribute information includes: the number of cache levels and / or the attribute information of each level of cache;

[0087] The analysis of the cache attribute information to determine cache characteristics includes:

[0088] According to the number of cache levels and / or the attribute information of each level of cache, determine the information of the first - level instruction cache, the first - level data cache, the second - level cache, and the third - level cache;

[0089] Wherein, the information includes at least one of the following: cache size, number of cache blocks, cache block size.

[0090] Here, the number of cache levels: refers to different levels of cache. A chip can have multiple levels of cache (such as L1, L2, L3). L1 is usually the smallest but the fastest cache, while L3 is usually the largest but the slowest.

[0091] Cache attribute: refers to various attributes of the cache, such as cache size.

[0092] In an example, the obtained cache attribute information includes but is not limited to the following:

[0093] LEVEL1_ICACHE_SIZE: 32768;

[0094] LEVEL1_ICACHE_ASSOC: 8;

[0095] LEVEL1_ICACHE_LINESIZE: 64;

[0096] LEVEL1_DCACHE_SIZE: 32768;

[0097] LEVEL1_DCACHE_ASSOC: 8;

[0098] LEVEL1_DCACHE_LINESIZE: 64;

[0099] LEVEL2_CACHE_SIZE: 524288;

[0100] LEVEL2_CACHE_ASSOC: 8;

[0101] LEVEL2_CACHE_LINESIZE: 64;

[0102] LEVEL3_CACHE_SIZE: 268435356;

[0103] LEVEL3_CACHE_ASSOC: 64;

[0104] LEVEL3_CACHE_LINESIZE: 64;

[0105] LEVEL4_CACHE_SIZE: 0;

[0106] LEVEL4_CACHE_ASSOC: 0;

[0107] LEVEL4_CACHE_LINESIZE: 0.

[0108] Among them, ASSOC represents Associativity, which refers to the associativity of the cache, that is, the organization method of the cache, indicating how the cache is partitioned and stores data. The cache can be divided into multiple cache sets, and each set contains multiple cache lines. The associativity determines the number of cache lines within each set.

[0109] LINESIZE refers to the size of the cache line, that is, the amount of data that each cache line can store. The cache line is the basic unit of the cache, and the amount of data loaded from memory into the cache each time is one cache line.

[0110] ICACHE represents the instruction cache, and DCACHE represents the data cache.

[0111] LEVEL1, LEVEL2, LEVEL3, and LEVEL4 respectively represent the first-level cache, second-level cache, third-level cache, and fourth-level cache.

[0112] Based on the above content for analysis, the information of the first-level instruction cache, first-level data cache, second-level cache, and third-level cache as shown in Table 1 below can be obtained.

[0113] Table 1

[0114]

[0115] Here, considering the detailed information of each level of cache (such as cache size, number of cache blocks, cache block size) helps to reveal the chip's architecture design. For example, the configurations of the first-level instruction cache, first-level data cache, second-level cache, and third-level cache are often closely related to the chip's architecture and design goals. By analyzing the cache attribute information, some characteristics of the chip's structure can be inferred.

[0116] In some embodiments, the device attribute information includes: socket-related information;

[0117] Analyze the device attribute information to determine core characteristics, including:

[0118] According to the socket-related information, determine the first characteristic; the first characteristic includes:

[0119] The number of sockets, the number of processor cores included in each socket, and whether each processor core supports hyper-threading technology.

[0120] Here, the socket-related information can reflect the number of sockets in the device, the hardware included in each socket (such as processors, memory modules, etc.).

[0121] In one example, the obtained socket-related information includes but is not limited to the following:

[0122] Socket count: 2;

[0123] Socket Designation: CPU0;

[0124] L1 Cache Handle: 0x0006;

[0125] L2 Cache Handle: 0x0007;

[0126] L3 Cache Handle: 0x0008;

[0127] Core Count: 64;

[0128] Core Enabled: 64;

[0129] Thread Count: 128.

[0130] Among them, "Socket count" represents the number of sockets; "Socket Designation" represents the socket identifier; "CacheHandle" represents the cache handle, which is usually used as a unique identifier in the system. "Core Count" represents the number of cores. As mentioned above, this processor has 64 cores. "Core Enabled" represents the number of enabled cores. As mentioned above, all 64 cores are enabled. "ThreadCount" represents the number of threads. As mentioned above, it means that the processor supports 128 threads, usually because the hyper-threading technology is enabled, and each core can handle two threads.

[0131] The socket-related information may further include the following information:

[0132] Characteristics:

[0133] 64-bit capable, indicating support for the 64-bit architecture;

[0134] Multi-Core, indicating a multi-core processor, meaning that this processor has multiple cores (64 cores);

[0135] Hardware Thread, indicating support for hardware threads;

[0136] Execute Protection, indicating execution protection, which helps prevent certain types of malicious attacks.

[0137] Enhanced Virtualization, indicating enhanced virtualization support.

[0138] Based on the above socket-related information, the required information can be analyzed, such as: the number of processor cores included in a single socket (64 as above), and whether each of the processor cores supports the hyper-threading technology (supported as above).

[0139] Here, considering that the number of sockets helps to understand the processor scalability, and the number of processor cores in each socket reveals the core scale of the chip. Further analyzing whether the cores support the hyper-threading technology can infer the design of the chip in terms of parallel processing ability and multi-task execution efficiency. By inferring these characteristics, it provides a basis for constructing the chip structure diagram.

[0140] In some embodiments, the device attribute information includes: Non-Uniform Memory Access (NUMA) node information;

[0141] Analyze the device attribute information to determine the core characteristics, including:

[0142] Based on the NUMA node information, determine the second characteristic; the second characteristic includes:

[0143] The number of NUMA nodes, the number of processor cores contained in each NUMA node, and the distribution of processor cores within each NUMA node.

[0144] Here, NUMA is a memory architecture where each processor core or processor cluster has its own local memory, and other memories are shared among different processors. This architecture helps improve the system performance, but the memory access speed may vary between different nodes.

[0145] Here, according to the obtained NUMA node information, understand the number and configuration of different NUMA nodes, and determine the second feature, such as:

[0146] The number of NUMA nodes: How many NUMA nodes are there in total in the chip;

[0147] The number of processor cores contained in each NUMA node;

[0148] The distribution of processor cores within each NUMA node: How the processor cores are distributed within the NUMA node, whether evenly or with other special configurations.

[0149] In one example, the NUMA node information may include but is not limited to the following:

[0150] NUMA node 0 CPU: 0 - 15, 128 - 143;

[0151] NUMA node 1 CPU: 16 - 31, 144 - 159;

[0152] NUMA node 2 CPU: 32 - 47, 160 - 175;

[0153] NUMA node 3 CPU: 48 - 63, 176 - 191;

[0154] NUMA node 4 CPU: 64 - 79, 192 - 207;

[0155] NUMA node 5 CPU: 80 - 95, 208 - 223;

[0156] NUMA node 6 CPU: 96 - 111, 224 - 239;

[0157] NUMA node 7 CPU: 112 - 127, 240 - 255.

[0158] Based on the above content, the second feature can be analyzed and obtained, including:

[0159] The number of NUMA nodes is 8, and each node contains multiple CPU cores. The processor cores of each node are identified in the format of CPU:X - Y; each NUMA node has 32 CPU cores, including two different core number ranges (for example, NUMA node 0 has CPU 0 - 15 and CPU 128 - 143).

[0160] Here, considering that the number of NUMA nodes reveals the memory and processor distribution architecture, and the number and distribution of processor cores within each node help understand how the chip optimizes data access and parallel computing. With this information, the performance, memory access efficiency, and load balancing strategy of the chip in a multi - processor environment can be inferred, and thus the chip's structure can be speculated.

[0161] In some embodiments, the analyzing the cache attribute information and / or the device attribute information to determine the cache characteristics and / or core characteristics includes:[[]]END]

[0162] Analyzing the cache attribute information and / or the device attribute information to determine a third characteristic; the third characteristic includes:[[]]END]

[0163] The cache size corresponding to each processor core in each level of cache;

[0164] The cache write policies supported by each level of cache;

[0165] Whether the NUMA node supports hyper - threading technology.

[0166] Here, after obtaining the above first characteristic and second characteristic, the first characteristic, the second characteristic, and other information can be combined for further speculation to obtain the third characteristic; for example, the other information obtained can include: the number of threads, the attribute information of the cache, etc.; among them, the attribute information of the cache such as the information about the L1 cache is as follows:

[0167] Socket Designation: L1 - Cache: indicates that the cache type is L1 cache.

[0168] Configuration: Enabled, Not Socketed, Level 1: indicates that the L1 cache is enabled, but has no socket and belongs to the first - level cache.

[0169] Operational Mode: Write Back: indicates that the operation mode is write - back; of course, it may also be other combinations such as write - back and write - through, update - on - demand, write - through, etc.

[0170] Location: Internal: indicates that the cache is located inside the processor.

[0171] Installed Size: 4 MB: It indicates that the installed cache size is 4 MB.

[0172] Maximum Size: 4 MB: It indicates that the maximum supported cache size is 4 MB.

[0173] Supported SRAM Types: Pipeline Burst: It indicates that the supported SRAM type is Pipeline Burst, a memory access mode.

[0174] Installed SRAM Type: Pipeline Burst: It indicates that the installed SRAM type is also Pipeline Burst.

[0175] Speed: 1 ns: It indicates that the cache access speed is 1 nanosecond (ns).

[0176] Error Correction Type: Multi-bit ECC: It indicates that multi-bit ECC (Error Correction Code) technology is used, which can detect and correct errors in multiple bits.

[0177] System Type: Unified: It indicates that the cache system type is unified, that is, the data and instruction caches are shared.

[0178] The same applies to other L2 caches, L3 caches, etc., and no examples will be given here one by one.

[0179] Through the above information, the third feature of the chip can be determined. For example, the sizes of each level of cache reveal the data access capabilities of each core, and the cache write policy reflects the design strategy of the chip in terms of data consistency and performance optimization. At the same time, analyzing whether the NUMA node supports hyper-threading technology can infer the performance scheduling method of the chip in multi-core and multi-thread parallel computing. Combining these features, the overall architecture of the chip, the cache management method, and the cooperative working mode of multi-processors can be inferred.

[0180] In some embodiments, analyzing the cache features and / or core features to obtain chip structure information includes:

[0181] Analyzing the second feature based on the first rule to determine the first structure information, where the first structure information includes: the number of processor groups CCX in the NUMA node, and the number of processor cores in each CCX.

[0182] Here, the first rule may be a rule for performing feature integration and analysis, which can indicate how to combine the above various features to determine the chip structure information.

[0183] Here, NUMA node: refers to a memory access architecture. Each NUMA node has a set of CPUs and memory, and there is a certain latency between the processor and the memory.

[0184] There are multiple processor groups (CCX, Core Complex) within each NUMA node. They are a small unit composed of multiple cores, and each CCX contains several processor cores.

[0185] Here, it is also possible to infer the NUMA node situation based on detecting the access latency of each CPU to different NUMA nodes and combining the range of access latency. Suppose the latency of CPU0 - 7 is in a certain range interval (for example, the latency distribution of CPU0 - 7 is 106.7, 106.7, 108.2, 108.3, 106.9, 106.2, 108.2, 108.8), the latency of CPU8 - 15 is in another range interval (110.3, 110.3, 110.3, 110.8, 110.7, 110.5, 111.0, 111.8), the latency of CPU16 - 23 is in a range interval (204.3, 204.5, 204.3, 204.8, 204.7, 204.5, 206.0, 206.8), and so on; through the above various features, it can be analyzed that, for example, NUMA0 node (including 0 - 15 CPUs), the latency of CPU 0 - 7 is in one range, 8 - 15 is in another range, and the conclusion that 8 cores are a whole can be drawn.

[0186] In some embodiments, analyzing the cache features and / or core features to obtain the chip structure information includes:

[0187] Analyzing the cache features and / or core features based on the first rule to determine the second structure information, where the second structure information includes: the situation of shared cache between processor cores.

[0188] In some embodiments, the method may further include: obtaining access test information and combining the access test information to determine the second structure information.

[0189] Here, based on the cache features and / or core features obtained above, information such as the cache hierarchy and the size of each level of cache can be known; and information about the processor cores, such as the number of cores and the configuration of each core.

[0190] In one example, it is also possible to make inferences in combination with the transmission latency of the inter-core access cache to determine whether the cache is shared. For example, the following transmission latencies are detected:

[0191] Core0->Core2 26.3ns;

[0192] Core0->Core1 27.9ns;

[0193] Core0->Core3 27.6ns;

[0194] Core0->Core4 97.9ns.

[0195] In this way, it can be determined that Core4 does not share the L3 cache with Core0, Core1, Core2, and Core3, and further determine that each CCX contains 4 processor cores.

[0196] In another example, the data of Core0 accessing the L2 cache of other cores is obtained as shown in Table 2 below.

[0197] Table 2

[0198]

[0199] Based on Table 2 above, it can be further determined that 4 cores share one L3 cache and do not share it with other cores. The CCX also accesses other CCXs through the IO Die and does not support direct data transmission between CCXs. One NUMA node contains 2 CCDs, and each CCD contains 2 CCXs, rather than one NUMA node containing 4 independent CCXs.

[0200] It should be noted that the above examples are only for illustrating that various cache attribute information, access test information, device attribute information, etc. can be obtained, and the chip structure information can be determined by analyzing and extracting the required features. There is no limitation on the specific first rule adopted. The first rule can actually include various information, which defines methods for various feature extraction and information integration and inference.

[0201] Here, by combining the analysis of cache features and core features, the chip structure is speculated. Among them, analyzing the number of processor groups (CCXs) in the NUMA node and the number of processor cores in each CCX helps to understand the multi-processor architecture of the chip and its internal distribution. Knowing the number of cores in each CCX under the NUMA node can speculate the parallel computing ability of the chip and the coordination between multiple cores.

[0202] Next, by analyzing the cache characteristics and core characteristics, it is also possible to determine the shared cache situation among the processor cores. The shared cache (such as the L3 cache) is typically used for sharing data among multiple cores. Understanding how the processor cores share the cache, as well as the size, hierarchy, and distribution of the cache, helps to infer the design concept of the chip in terms of optimizing data access, reducing bottlenecks, and improving the performance of multi-core processing.

[0203] In addition, by combining the access test information to confirm the second structural information, the accuracy of the analysis is further verified. The access test information is usually used to evaluate the access pattern, latency, and bandwidth of the cache. Through testing, the actual performance of each level of cache can be obtained, such as the cache hit rate, access latency, and data transfer rate, etc., so as to more accurately infer the performance characteristics and potential bottlenecks of the chip.

[0204] Based on the analysis of the above first rule, the required chip structure information can be analyzed so as to use the chip structure information for drawing the chip structure diagram.

[0205] In some embodiments, constructing a chip structure diagram according to the chip structure information includes:

[0206] Using a drawing model, constructing a chip structure diagram according to the chip structure information.

[0207] Here, the drawing model is used to present the hardware architecture of the chip as a graphical chip structure diagram in a visual way according to the chip structure information.

[0208] For example, using the above method to determine the chip structure information as follows:

[0209] (1) A single socket contains 4 NUMA nodes;

[0210] (2) Each NUMA node contains 2 CCDs;

[0211] (3) Each CCD contains two CCXs, and each CCX contains 4 processor cores (Core);

[0212] (4) Each Core contained within the CCX contains a private L1 instruction Cache, L1 data Cache, private L2 Cache, and shared L3 Cache.

[0213] Combining the above information, a chip structure diagram as shown in Figure 2 can be drawn, as shown in Figure 2As shown, Socket-0 is drawn, which includes NUMA Node-0, NUMA Node-1, NUMA Node-2, and NUMA Node-3; and specifically presents the structures and associations of CCD-0, CCD-1, CCX-0, CCX-1, Core 0, Core 1, Core 2, Core 3, and L1 instruction cache (L1 Icache), L1 data cache (L1 Dcache), L2 cache (L2 cache), and L3 cache (L3 cache) in NUMA Node 0. It should be understood that NUMA Node 1, NUMA Node 2, and NUMA Node 3 can also have a similar structure to NUMA Node 0.

[0214] Here, using the drawing model to build a chip structure diagram can present the unknown chip hardware architecture in an intuitive and clear way, helping technicians better understand the various components of the chip and their interrelationships. Through the graphical chip structure diagram, the chip's NUMA nodes, processor groups (CCX), cores, cache levels (such as L1, L2, L3 cache) and their connections can be displayed, making the chip's design structure more transparent.

[0215] Furthermore, technicians can formulate corresponding core binding methods based on the chip structure diagram when conducting program testing, and determine which cores the program should be bound to for optimal performance. Even if the internal architecture is kept opaque to the outside world by the chip manufacturer, the system manufacturer can still obtain the chip framework required for application tuning and implement the maximum performance tuning method, thereby promoting more efficient chip development and application.

[0216] In some embodiments, the method further comprises:

[0217] Acquire a training data set; the training data set includes at least one sample chip structure information and a logical architecture diagram of the sample chip structure corresponding to each sample chip structure information;

[0218] A preset neural network model is trained according to the training data set to obtain a trained neural network model as the drawing model.

[0219] Here, the drawing model can be trained based on a neural network model, for example, a convolutional neural network (CNN) or a graph neural network (GNN) can be used, which can process data with spatial or graph structures.

[0220] In practical applications, multiple chip structure information samples can be collected, and each sample includes the detailed hardware architecture of the chip (such as hierarchical information of NUMA nodes, CCDs, CCXs, cores, and caches, etc.). For each chip structure sample, a corresponding logical architecture diagram is provided, which represents how the levels and components of the chip structure are organized and connected, similar to the topology diagram of the chip.

[0221] Using this data (chip structure information and its corresponding logical architecture diagram), a neural network model is trained. The neural network model learns the relationship between the chip structure and its logical architecture diagram, and masters how to generate or predict the corresponding logical architecture diagram from the input chip structure information.

[0222] Through the training process, the neural network can gradually optimize its parameters, and finally obtain a model that can automatically generate the corresponding architecture diagram according to the new chip structure information. The trained neural network model can be used to draw the logical architecture diagram of the chip, that is, when new chip structure information is obtained as input, the neural network can automatically generate the corresponding logical architecture diagram based on the existing training experience.

[0223] In this way, the drawing model can automatically generate the corresponding logical architecture diagram according to the chip structure information, avoiding the cumbersome process and errors in the traditional manual drawing. This method can accelerate the design iteration, improve the design accuracy, especially when dealing with complex structures, it can effectively reduce human errors. At the same time, the automated process enables designers to quickly verify and adjust the architecture, shortening the R & D cycle. By reducing manual intervention, this method can also significantly save design costs, especially in large-scale and high-complexity chip designs.

[0224] In some embodiments, the method further includes at least one of the following:

[0225] Receiving the tuning test result, where the tuning test result is the operation result after process debugging based on the drawn chip structure diagram; feeding back the operation result to the drawing model to optimize the drawing model;

[0226] Receiving at least one other sample chip structure information and the logical architecture diagram of the sample chip structure corresponding to each sample chip structure information, and updating the drawing model according to the at least one other sample chip structure information and the logical architecture diagram of the sample chip structure corresponding to each sample chip structure information.

[0227] Specifically, a closed-loop feedback mechanism is designed, and the results of the running process (such as performance, power consumption, throughput, etc.) are used as feedback for optimizing the drawing model. Through this mechanism, the drawing model can adjust the strategy according to the differences in each running result, and thus continuously optimize.

[0228] Here, considering that the obtained operation results may not meet the expectations after the process configuration is optimized according to the chip structure diagram, that is, the operation results that should be achieved theoretically by applying this configuration to the chip do not appear. Here, it is considered that there may be deviations in the drawing of the storage chip result diagram. Therefore, a method of feedback mechanism for optimizing the drawing model is proposed.

[0229] For example, if the running process performs poorly (such as slow running speed, unbalanced resource utilization, etc.), it indicates that there may be room for optimization in the chip structure diagram. Or, the current chip structure diagram error causes the existing operation results not to meet the requirements. For another example, if the power consumption is too high, it is necessary to check whether there is excessive computing or communication overhead, and the design can be further adjusted to reduce the power consumption.

[0230] Here, different weights are set for the tuning test results of different samples. For example, some tuning test results come from known or verified chip structures. Then these tuning test results are usually more accurate and reliable. Set higher weights for these test results to ensure that more reliance is placed on these test data with higher certainty during adjustment, thereby accelerating the optimization process and reducing potential errors.

[0231] For some chip structures determined based on speculation or incomplete data, their accuracy may not be as high as that of directly verified structures. Therefore, the tuning test results of these chip structures may be more uncertain or noisy. If overly relying on these speculative results, it may cause the optimization process to deviate from the best path or overfit certain specific assumptions. Therefore, reducing the weights of these results helps to avoid the model making wrong adjustments due to these inaccurate data.

[0232] In this way, continuously adjusting the drawing model according to the actual operation results can effectively improve the drawing accuracy.

[0233] In some embodiments, the obtaining of the cache attribute information and / or device attribute information includes:

[0234] Using a test tool to call at least one query instruction, and obtaining the cache attribute information and / or device attribute information of the electronic device according to the query instruction.

[0235] Here, a technician can use a test tool to send at least one query instruction to the electronic device, and obtain the cache attribute information of the device and / or other device attribute information through this instruction.

[0236] Among them, the test tool: refers to a dedicated software or hardware tool for testing or analyzing an electronic device. It can be used to detect the performance, hardware status, etc. of the device.

[0237] Query instruction: This is a command or instruction through which requests can be sent to an electronic device to obtain relevant information.

[0238] Here, the electronic device can have a human-computer interaction interface, and various query instructions can be called through test software to obtain various cache attribute information and / or device attribute information. Multiple query instructions can be called by themselves in this way, or can be selectively called by the user according to their own needs, and can be specifically used based on the actual situation, which is not limited here.

[0239] In this way, technicians use the test tool to send query instructions to conveniently and quickly obtain information about the device's cache and other hardware attributes.

[0240] In some embodiments, the test tool and the drawing model can be integrated into a certain hardware or software.

[0241] For example, at the software level, the test tool and the drawing model can be implemented as part of a software package. For instance, some performance monitoring tools (such as CPU-Z, GPU-Z) can integrate performance testing and graphical display functions. In this case, the test tool will collect hardware data, and the function of the drawing model will process these data and display them on the user interface.

[0242] For another example, at the hardware level, certain dedicated hardware (such as embedded systems, instrument test equipment, etc.) can integrate the corresponding test tools and drawing functions. There may be an embedded processor on the hardware platform, which is used to collect various information and display the drawing results of the drawing model through a display interface or an external device. Of course, the collected information can also be processed by a program.

[0243] By integrating the test tool and the drawing model, data can be automatically collected and visualized images can be generated. This integration improves efficiency, avoids switching between multiple tools, and enables users to complete testing, data analysis, and result display on one platform. Through graphical display, users can intuitively understand various information, features, chip structures, etc.

[0244] Through the method provided by the embodiments of the present disclosure, the structure of the chip can be inferred without knowing the micro-architecture of the chip, thereby helping the whole machine manufacturer to optimize CPU scheduling. The whole machine manufacturer can formulate a core binding strategy according to the inferred chip structure, determine which cores the program runs on with the best performance during program testing, and thus avoid the program from migrating frequently or competing for shared resources (such as caches) between multiple cores, and can more efficiently utilize the core resources of the processor, improving the overall system efficiency; and, reducing the uncertainty during program operation can enhance the stability and predictability of the program, and avoid performance fluctuations caused by core switching or unbalanced load and other beneficial effects.

[0245] In this way, even if chip manufacturers maintain the opacity of the microarchitecture, the overall machine manufacturers can still optimize the performance of applications, which helps to improve performance, optimize resource utilization, reduce energy consumption, and can perform effective program optimization in the absence of transparent information about the chip microarchitecture.

[0246] Figure 3 The structural schematic diagram of a chip structure determination device provided by an embodiment of the present disclosure; as Figure 3 shown, the device is applied to an electronic device, and the device includes:

[0247] An acquisition module, configured to acquire cache attribute information and / or device attribute information;

[0248] A first processing module, configured to determine chip structure information according to the cache attribute information and / or the device attribute information;

[0249] A second processing module, configured to construct a chip structure diagram according to the chip structure information.

[0250] In some embodiments, the first processing module is configured to analyze the cache attribute information and / or the device attribute information to determine cache characteristics and / or core characteristics;

[0251] Analyze the cache characteristics and / or core characteristics to obtain chip structure information.

[0252] In some embodiments, the cache attribute information includes: the number of cache levels and / or the attribute information of each level of cache;

[0253] The first processing module is configured to determine the information of the first-level instruction cache, the information of the first-level data cache, the information of the second-level cache, and the information of the third-level cache according to the number of cache levels and / or the attribute information of each level of cache;

[0254] Wherein, the information includes at least one of the following: cache size, number of cache blocks, cache block size.

[0255] In some embodiments, the device attribute information includes: slot-related information;

[0256] The first processing module is configured to determine a first feature according to the slot-related information; the first feature includes:

[0257] The number of slots, the number of processor cores included in each slot, and whether each processor core supports hyper-threading technology.

[0258] In some embodiments, the device attribute information includes: non-uniform memory access (NUMA) node information;

[0259] The first processing module is configured to determine a second feature according to the NUMA node information; the second feature includes:

[0260] The number of NUMA nodes, the number of processor cores included in each NUMA node, and the distribution of processor cores within each NUMA node.

[0261] In some embodiments, the first processing module is configured to analyze the cache attribute information and / or the device attribute information to determine a third feature; the third feature includes:

[0262] The cache size corresponding to each processor core in each level of cache;

[0263] The cache write policies supported by each level of cache;

[0264] Whether the NUMA node supports hyper-threading technology.

[0265] In some embodiments, the first processing module is configured to analyze the second feature based on a first rule to determine first structure information, where the first structure information includes: the number of CCXs of the NUMA node, and the number of processor cores of each CCX.

[0266] In some embodiments, the first processing module is configured to analyze the cache feature and / or the core feature based on a first rule to determine second structure information, where the second structure information includes: the situation of shared cache among processor cores.

[0267] In some embodiments, the second processing module is configured to use a drawing model to construct a chip structure diagram according to the chip structure information.

[0268] In some embodiments, the second processing module is further configured to obtain a training data set; the training data set includes at least one sample chip structure information and the logical architecture diagram of the sample chip structure corresponding to each sample chip structure information;

[0269] Train a preset neural network model according to the training data set to obtain a trained neural network model as the drawing model.

[0270] In some embodiments, the obtaining module is configured to use a test tool to call at least one query instruction, and obtain the cache attribute information and / or the device attribute information of the electronic device according to the query instruction.

[0271] It can be understood that when the chip structure determination device provided in the above embodiments implements the corresponding chip structure determination method, the above processing can be allocated to different program modules as needed to complete all or part of the processing described above. In addition, the device provided in the above embodiments and the embodiments of the corresponding method belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.

[0272] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the chip structure determination method.

[0273] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, where the executable instructions are stored. When the executable instructions are executed by a processor, the processor will be caused to execute the chip structure determination method provided in the embodiments of the present application.

[0274] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.

[0275] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, and may be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0276] As an example, the executable instructions may or may not correspond to a file in the file system, and may be stored as part of a file that stores other programs or data. For example, they may be stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (for example, files that store one or more modules, subroutines, or code portions).

[0277] As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed at multiple locations and interconnected by a communication network.

[0278] Figure 4 A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure; as Figure 4 shown, the electronic device 40 includes: a processor 401 and a memory 402 communicatively connected to the processor 401; the memory 402 stores instructions executable by the processor 401. When the instructions are executed by the processor 401, the processor 401 is enabled to execute:

[0279] Obtain cache attribute information and / or device attribute information;

[0280] Determine chip structure information according to the cache attribute information and / or device attribute information;

[0281] Construct a chip structure diagram according to the chip structure information.

[0282] The electronic device provided in the above embodiment and the embodiment of the corresponding chip structure determination method belong to the same concept. For the specific implementation process, please refer to the method embodiment and will not be elaborated here.

[0283] In practical applications, the electronic device 40 may further include: at least one network interface 403. Each component in the electronic device 40 is coupled together through a bus system 404. It can be understood that the bus system 404 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 404 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 4 all kinds of buses are labeled as the bus system 404. Among them, the number of the processors 401 may be at least one, and the number of the memories 402 may be at least one. The network interface 403 is used for the communication between the electronic device 40 and other devices in a wired or wireless manner.

[0284] The memory 402 in the embodiment of the present disclosure is used to store various types of data to support the operation of the electronic device 40.

[0285] The method disclosed in the above embodiments of the present disclosure can be applied to or implemented by a processor 401. The processor 401 may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 401 or instructions in the form of software. The above-mentioned processor 401 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 401 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present disclosure, it can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, and this storage medium is located in the memory 402. The processor 401 reads the information in the memory 402 and combines its hardware to complete the steps of the foregoing chip structure determination method.

[0286] In some embodiments, the electronic device 40 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontroller units (MCUs), microprocessors, or other electronic components for performing the foregoing method.

[0287] It should be understood that various forms of the processes shown above can be used, reordering, adding, or deleting steps. For example, the steps recited in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitation is made herein.

[0288] In the above description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and they can be combined with each other without conflict.

[0289] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terms used in this disclosure are for the purpose of describing embodiments of this disclosure only and are not intended to limit this disclosure.

[0290] It should be understood that in various embodiments of this disclosure, the magnitude of the sequence number of each implementation process does not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this disclosure.

[0291] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this disclosure, "a plurality of" means two or more unless otherwise specifically defined.

[0292] As described above, the above are only specific embodiments of this disclosure, but the protection scope of this disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by this disclosure can easily think of changes or substitutions, which should all be covered by the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be subject to the protection scope of the claims.

Claims

1. A method for determining a chip structure, characterized in that, The method is applied to an electronic device, and the method includes: Obtain cache attribute information and / or device attribute information; Determine chip structure information according to the cache attribute information and / or device attribute information; Construct a chip structure diagram according to the chip structure information; The determining the chip structure information according to the cache attribute information and / or device attribute information includes: Analyze the cache attribute information and / or the device attribute information to determine cache characteristics and / or core characteristics; Analyze the cache characteristics and / or core characteristics to obtain chip structure information; The device attribute information includes: slot-related information; Analyzing the device attribute information to determine core characteristics includes: Determine a first feature according to the slot-related information; the first feature includes: The number of slots, the number of processor cores included in each slot, and whether each of the processor cores supports hyper-threading technology.

2. The method according to claim 1, wherein The cache attribute information includes: the cache level and / or the attribute information of each level of cache; Analyzing the cache attribute information to determine cache characteristics includes: Determine the information of the level-1 instruction cache, the level-1 data cache, the level-2 cache, and the level-3 cache according to the cache level and / or the attribute information of each level of cache; Wherein, the information includes at least one of the following: cache size, number of cache blocks, and cache block size.

3. The method according to claim 1, wherein The device attribute information further includes: non-uniform memory access (NUMA) node information; Analyzing the device attribute information to determine core characteristics further includes: Determine a second feature according to the NUMA node information; the second feature includes: The number of NUMA nodes, the number of processor cores included in each NUMA node, and the distribution of processor cores within each NUMA node.

4. The method according to claim 1, wherein The analyzing the cache attribute information and / or the device attribute information to determine cache characteristics and / or core characteristics includes: Analyze the cache attribute information and / or the device attribute information to determine a third feature; the third feature includes: The cache size corresponding to each processor core in each level of cache; The cache write policies supported by each level of cache; Whether the NUMA node supports hyper-threading technology.

5. The method according to claim 3, characterized in that, The analyzing the cache characteristics and / or core characteristics to obtain chip structure information includes: Analyze the second feature based on a first rule to determine first structure information, the first structure information includes: the number of processor groups (CCXs) in the NUMA node, and the number of processor cores in each CCX.

6. The method according to claim 1, characterized in that, Analyzing the cache characteristics and / or core characteristics to obtain chip structure information includes: Analyze the cache characteristics and / or core characteristics based on a first rule to determine second structure information, the second structure information includes: the situation of shared cache among processor cores.

7. The method according to claim 1, wherein Constructing a chip structure diagram according to the chip structure information includes: Use a drawing model to construct a chip structure diagram according to the chip structure information.

8. The method according to claim 7, wherein The method further includes: Obtain a training data set; the training data set includes at least one sample chip structure information and the logical architecture diagram of the sample chip structure corresponding to each sample chip structure information; Train a preset neural network model according to the training data set to obtain a trained neural network model as the drawing model.

9. The method according to claim 1, characterized in that, The obtaining of the cache attribute information and / or device attribute information includes: Use a test tool to call at least one query instruction, and obtain the cache attribute information and / or device attribute information of the electronic device according to the query instruction.

10. A chip structure determination device, characterized in that The device is applied to an electronic device, and the device includes: An obtaining module, configured to obtain cache attribute information and / or device attribute information; A first processing module, configured to determine chip structure information according to the cache attribute information and / or device attribute information; A second processing module, configured to construct a chip structure diagram according to the chip structure information; The second processing module is configured to analyze the cache attribute information and / or the device attribute information to determine cache features and / or core features; analyze the cache features and / or core features to obtain chip structure information; The device attribute information includes: slot-related information; the second processing module is configured to determine a first feature according to the slot-related information; the first feature includes: the number of slots, the number of processor cores included in each slot, and whether each of the processor cores supports hyper-threading technology.

11. The device according to claim 10, characterized in that, The obtaining module is configured to use a test tool to call at least one query instruction, and obtain the cache attribute information and / or device attribute information of the electronic device according to the query instruction; The second processing module is configured to use the drawing model to construct a chip structure diagram according to the chip structure information.

12. An electronic device, characterized in that, Includes: At least one processor; And a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 9.

13. A non-transitory computer-readable storage medium storing computer instructions, characterized in that The computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Core particle system memory controller layout optimization method

    CN118332999A

  • Coarse-grained reconfigurable chip mapping method and device based on pre-scheduling

    CN119537305A