Chip layout determination method and related device
Patent Information
- Application Number
- CN202611274065.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-21
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]相关技术中,针对大模型硬件芯片的布线版图确定通常采用基于固定规则或均匀划分方式进行硬件区域分配和布线设计,导致芯片版图中的硬件区域配置与实际权重数据特性之间存在偏差;此外,相关技术中在进行权重信息与芯片物理结构映射时,布线过程难以直接适配大模型权重分布特性,影响芯片布线版图确定的准确性和资源利用效果
[0023]综上,本申请基于目标大模型权重矩阵中各权重值的出现概率对芯片版图区域进行划分,由于各硬件区域面积与对应权重值的出现概率呈正相关,因此出现频率较高的权重值能够获得更大的物理实现区域,而出现频率较低的权重值仅占用较小区域,由此,芯片版图中的硬件资源分配能够与模型权重的数据分布特征相匹配,相比于固定划分方式,可以减少因权重分布不均导致的区域利用率不足问题,使有限的版图空间得到更加合理的分配,能够提高芯片布线资源与硬件区域的利用效率;通过基于权重矩阵以及模型输入激活信号对应物理管脚信息建立接入映射关系表,将权重矩阵中的各输入维度与对应硬件区域建立关联关系,使输入激活信号能够按照权重对应关系被准确引导至目标硬件区域,使得该映射过程直接利用权重矩阵的数据特征确定信号接入路径,能够避免输入信号与权重区域之间无规则连接造成的布线复杂化问题,能够降低布线连接关系的不确定性,提高后续布线处理的准确性和可实现性;基于所述接入映射关系表,仅在芯片版图的目标金属层内执行布线处理,使权重相关的连接关系能够按照预先确定的映射关系转换为物理金属连接,从而形成与目标大模型权重对应的目标布线版图,使得布线过程由权重矩阵直接驱动,并限定在目标金属层范围内进行,无需依赖额外的权重存储映射过程即可实现权重信息到物理版图连接关系的转换,同时便于对权重对应的布线结构进行独立调整,提高芯片版图生成过程的灵活性。综上所述,本申请提供的芯片布线版图确定方法通过基于大模型权重分布特征进行版图区域划分、输入映射及目标金属层布线,能够实现权重信息到芯片物理连接结构的直接转换,从而提高版图资源利用效率、降低布线复杂度并提升芯片版图生成的灵活性。
Smart Images

Figure CN122819142A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of chip manufacturing technology, and more specifically, to a method and related equipment for determining chip wiring layout. Background Technology
[0002] With the rapid development of artificial intelligence technology, the parameter scale of large-scale neural network models (such as large language models) continues to grow, placing higher demands on chip computing power, storage capacity, and energy efficiency. As a crucial foundation for supporting efficient inference of large models, chip routing layout determination technology needs to rationally plan the internal hardware resources and metal interconnect structures of the chip based on the computing architecture and data processing requirements to achieve an effective mapping between model parameters and physical hardware structure. Therefore, the research and application of chip routing layout determination technology for large model application scenarios is particularly important.
[0003] In related technologies, the determination of the wiring layout for large-model hardware chips typically employs fixed rules or uniform partitioning methods for hardware region allocation and wiring design. This leads to discrepancies between the hardware region configuration in the chip layout and the actual weight data characteristics. Furthermore, when mapping weight information to the chip's physical structure, the wiring process struggles to directly adapt to the weight distribution characteristics of large models, affecting the accuracy of the chip wiring layout determination and resource utilization. In other words, related technologies suffer from low hardware resource utilization and difficulty in fully adapting to the weight distribution characteristics of large models during the chip wiring layout determination process. Summary of the Invention
[0004] In the summary section of this application, the relevant technical solutions are described in general terms, and a series of simplified concepts are introduced. These concepts will be further elaborated in the detailed embodiments section. This summary section should not be construed as limiting the key or essential technical features of the claimed solutions, nor is it intended to limit the scope of protection of the claimed solutions.
[0005] The chip routing layout determination method and related equipment provided in this application can perform layout area division, input mapping and target metal layer routing based on the weight distribution characteristics of a large model. This enables the direct conversion of weight information to the chip physical connection structure, thereby improving the efficiency of layout resource utilization, reducing routing complexity and enhancing the flexibility of chip layout generation.
[0006] In a first aspect, this application provides a method for determining chip wiring layout, comprising: dividing a plurality of hardware regions in the chip layout corresponding one-to-one with each of the weight values in the weight matrix of a target large model, wherein the area of each hardware region is positively correlated with the probability of occurrence of the corresponding weight value; mapping each input dimension of the weight matrix to at least one corresponding hardware region based on the weight matrix and the physical pin information corresponding to the model input activation signal of the target large model, thereby obtaining an access mapping table; and performing wiring processing in the target metal layer of the chip layout based on the access mapping table to obtain the target wiring layout of the chip layout.
[0007] In some implementations, the step of mapping each input dimension of the weight matrix to at least one corresponding hardware region based on the physical pin information corresponding to the model input activation signal of the target large model to obtain an access mapping table includes: abstracting and generating multiple virtual access points based on the physical pin information, wherein each virtual access point corresponds one-to-one with each input dimension of the weight matrix; and mapping each virtual access point to at least one corresponding hardware region based on the weight matrix to obtain the access mapping table.
[0008] In some implementations, the step of performing routing processing within the target metal layer of the chip layout based on the access mapping table to obtain the target routing layout of the chip layout includes: based on a comprehensive routing cost function, performing a routing process from each of the virtual access points to the corresponding hardware regions in the access mapping table within the target metal layer of the chip layout to obtain an initial routing layout of the chip layout; generating pin access points for each bit counter unit in the initial routing layout based on the total number of target input lines in each hardware region of the initial routing layout and the upper limit of the input capability of a single bit counter unit to obtain an intermediate routing layout, wherein the total number of target input lines is the statistical number of virtual access points mapped to the hardware region, and the upper limit of the input capability of a single bit counter unit is the maximum number of access channels supported by a single bit counter unit determined based on preset timing constraints and physical design rules; and performing a preset verification optimization process on the intermediate routing layout to determine the verified intermediate routing layout as the target routing layout.
[0009] In some implementations, before executing a routing process from each of the virtual access points to the corresponding hardware regions in the access mapping table within the target metal layer of the chip layout based on the comprehensive routing cost function, to obtain the initial routing layout of the chip layout, the chip routing layout determination method further includes: obtaining routing loss assessment data of the target metal layer, wherein the routing loss assessment data includes resistance value per unit length, capacitance value per unit area, and via parasitic parameter value; obtaining power distribution network layout information of the chip layout, wherein the power distribution network layout information includes power line location coordinates, power pad coordinates, and power strip distribution data; and constructing the comprehensive routing cost function based on the routing loss assessment data and the power distribution network layout information, wherein the comprehensive routing cost function is obtained by weighted summation of routing resistance cost term, routing capacitance cost term, design rule violation cost term, and power distribution network conflict cost term.
[0010] In some implementations, the initial routing layout of the chip layout is obtained by executing a routing process from each of the virtual access points to the corresponding hardware regions in the access mapping table within the target metal layer of the chip layout based on the comprehensive routing cost function. This includes: generating multiple candidate routing paths within the target metal layer using a preset three-dimensional routing algorithm, with each of the virtual access points as the routing start point and the corresponding hardware region in the access mapping table as the routing end point; determining the cost value of each candidate routing path based on the comprehensive routing cost function; determining the candidate routing path with the minimum cost value among the multiple candidate routing paths as the target routing path; and generating the initial routing layout based on the target routing path.
[0011] In some implementations, the step of generating pin access points for each bit counting unit in the initial routing layout based on the total number of target input lines in each hardware region of the initial routing layout and the upper limit of the input capability of a single bit counting unit to obtain an intermediate routing layout includes: for each hardware region, if the total number of target input lines is less than or equal to the upper limit of the input capability, then a single bit counting unit is configured in the corresponding hardware region, and multiple input pins are generated in a number equal to the total number of target input lines; if the total number of target input lines is greater than the upper limit of the input capability, then based on the ratio of the total number of target input lines to the upper limit of the input capability, the corresponding hardware region is divided into multiple hardware sub-regions, and a bit counting unit and corresponding multiple input pins are configured for each hardware sub-region; the configured bit counting units and input pins are integrated into the initial routing layout to obtain the intermediate routing layout.
[0012] In some implementations, the step of performing a preset verification and optimization process on the intermediate routing layout, and determining the verified intermediate routing layout as the target routing layout, includes: performing a design rule check on the intermediate routing layout to obtain a first verification result; and performing a layout and schematic check on the intermediate routing layout. Figure 1 The first verification is performed to obtain a second verification result; the second verification result is performed to perform a power distribution network compatibility verification on the intermediate wiring layout to obtain a third verification result; when the first verification result, the second verification result, and the third verification result are all passed, the intermediate wiring layout is determined as the target wiring layout; otherwise, the wiring path and / or bit counting unit layout of the intermediate wiring layout are adjusted until all verifications are passed.
[0013] In some implementations, performing power distribution network compatibility verification on the intermediate routing layout to obtain a third verification result includes: determining the network voltage drop and network current density of the power distribution network based on the power distribution network layout information in the intermediate routing layout; and determining that the third verification result is passed when the network voltage drop is less than a preset voltage drop threshold and the network current density is less than a preset current density threshold.
[0014] In some implementations, mapping the virtual access points to at least one corresponding hardware region based on the weight matrix to obtain the access mapping table includes: traversing the weight matrix; for each target input dimension, determining the target weight value corresponding to each output dimension; if the target weight value is a non-disconnect value, establishing a mapping relationship between the virtual access point corresponding to the target input dimension and the hardware region corresponding to the target weight value; and generating the access mapping table based on all the mapping relationships.
[0015] In some implementations, before dividing the chip layout into multiple hardware regions corresponding one-to-one with each weight value in the weight matrix based on the probability of occurrence of each weight value in the target large model, the chip routing layout determination method further includes: performing quantization encoding processing on the weight matrix based on a preset number of quantization bits to obtain a discretized set of quantized weight values, wherein the preset number of quantization bits is 1 bit, 2 bits, 4 bits, or 8 bits; and determining the probability of occurrence of each weight value in the weight matrix based on the frequency of occurrence of each quantized weight value in the set of quantized weight values.
[0016] In some implementations, the probability of occurrence of each weight value in the weight matrix based on the target large model is used to divide the chip layout into multiple hardware regions corresponding one-to-one with each weight value. This includes: for each weight value, determining the target area of the hardware region corresponding to the weight value based on the probability of occurrence of the weight value and the total physical area of the bit counting array used for weight encoding in the chip layout; and dividing the chip layout into regions according to each target area to obtain multiple hardware regions corresponding one-to-one with each weight value.
[0017] In some implementations, before performing routing processing within the target metal layer of the chip layout based on the access mapping table, the chip routing layout determination method further includes: obtaining a routing hierarchy configuration file imported from a user interface; parsing the routing hierarchy configuration file to obtain protection rule constraints for the target metal layer and the bottom fixed layer for weight encoding, wherein the bottom fixed layer includes metal layers for power connections, clock connections, and basic device connections; and constructing inter-layer isolation logic between the target metal layer and the bottom fixed layer based on the protection rule constraints to restrict routing processing operations from being performed within the target metal layer.
[0018] Secondly, this application also provides a chip wiring layout determination device, comprising: a region division unit, used to divide a plurality of hardware regions in the chip layout corresponding one-to-one with each of the weight values based on the occurrence probability of each weight value in the weight matrix of the target large model, wherein the area of each hardware region is positively correlated with the occurrence probability of the corresponding weight value; a mapping establishment unit, used to map each input dimension of the weight matrix to at least one corresponding hardware region based on the weight matrix and the physical pin information corresponding to the model input activation signal of the target large model, to obtain an access mapping relationship table; and a wiring processing unit, used to perform wiring processing in the target metal layer of the chip layout based on the access mapping relationship table, to obtain the target wiring layout of the chip layout.
[0019] Thirdly, this application also provides an electronic device, including: a memory and a processor, wherein the processor is configured to implement the steps of the chip wiring layout determination method described in the first aspect when executing a computer program stored in the memory.
[0020] Fourthly, this application also provides a chip whose wiring layout is determined by the chip wiring layout determination method provided in the embodiments of this application.
[0021] Fifthly, this application also provides a computer-readable storage medium storing computer-executable instructions or a computer program, wherein when the computer-executable instructions or the computer program are executed by a processor, the steps of the chip wiring layout determination method described in the first aspect are implemented.
[0022] Sixthly, this application also provides a computer program product, including a computer program or computer executable instructions, wherein when the computer program or computer executable instructions are executed by a processor, the steps of the chip wiring layout determination method provided in the embodiments of this application are implemented.
[0023] In summary, this application divides the chip layout area based on the probability of occurrence of each weight value in the target large model weight matrix. Since the area of each hardware region is positively correlated with the probability of occurrence of the corresponding weight value, weight values with higher occurrence frequencies can obtain a larger physical implementation area, while weight values with lower occurrence frequencies occupy only a smaller area. Thus, the allocation of hardware resources in the chip layout can match the data distribution characteristics of the model weights. Compared with a fixed division method, this can reduce the problem of insufficient area utilization caused by uneven weight distribution, and make the limited layout space more rationally allocated, thereby improving the utilization efficiency of chip wiring resources and hardware areas. By establishing an access mapping relationship table based on the weight matrix and the physical pin information corresponding to the model input activation signals, the input dimensions in the weight matrix are associated with the corresponding hardware areas, so that the input activation signals can be mapped according to the weight correspondence. The method accurately guides the signal to the target hardware region, enabling the mapping process to directly determine the signal access path using the data characteristics of the weight matrix. This avoids the wiring complexity caused by irregular connections between the input signal and the weight region, reduces the uncertainty of wiring connections, and improves the accuracy and feasibility of subsequent wiring processing. Based on the access mapping table, wiring processing is performed only within the target metal layer of the chip layout. This allows weight-related connections to be converted into physical metal connections according to a pre-determined mapping relationship, forming a target wiring layout corresponding to the target large model weights. The wiring process is directly driven by the weight matrix and confined to the target metal layer, eliminating the need for additional weight storage mapping processes to achieve the conversion of weight information to physical layout connections. It also facilitates independent adjustment of the wiring structure corresponding to the weights, improving the flexibility of the chip layout generation process. In summary, the chip wiring layout determination method provided in this application, through layout region division, input mapping, and target metal layer wiring based on the large model weight distribution characteristics, can achieve direct conversion of weight information to the chip's physical connection structure, thereby improving layout resource utilization efficiency, reducing wiring complexity, and enhancing the flexibility of chip layout generation. Attached Figure Description
[0024] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit this specification. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 A schematic flowchart illustrating a chip wiring layout determination method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the composition structure of a chip wiring layout determination device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the composition structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0025] The terms used in the specification, claims, and drawings of this application, such as "first," "second," "third," "fourth," etc. (if any), are used to distinguish similar objects and not to describe a specific order or sequence. Therefore, it is to be understood that these terms can be used interchangeably where appropriate, allowing the described embodiments to be used in different orders, unless specifically required by the illustrations or description. Furthermore, the terms "is" and "has," and any variations thereof, are intended to cover, non-exclusively, all possible constituent elements. For example, a process, method, system, product, or apparatus comprising several steps or units is not necessarily limited to the steps or units explicitly listed, but may also include other steps or units not explicitly listed, or steps or units inherent to the process, method, product, or apparatus.
[0026] In this application, a "module" or "unit" refers to a computer program or part of a computer program that has a specific function and works in conjunction with other related parts to achieve a predetermined goal. These modules or units can be implemented by software, hardware (e.g., processing circuitry or memory), or a combination of both. One or more processors or memories can implement one or more modules or units. Furthermore, each module or unit can also be part of a larger module or unit.
[0027] The technical solutions of this application will be described in detail below with reference to the accompanying drawings of the embodiments. It should be noted that the described embodiments are only a part of this application, and not all embodiments. In the following description, the "some embodiments" mentioned are only a subset of all possible embodiments, which may be the same or different subsets, and different embodiments can be combined with each other without conflict.
[0028] Figure 1 This is a schematic flowchart illustrating a chip wiring layout determination method provided in an embodiment of this application. For example, see [link to example]. Figure 1The chip wiring layout determination method provided in this application embodiment may include the following steps 101 to 103: Step 101: Based on the probability of occurrence of each weight value in the weight matrix of the target large model, divide the chip layout into multiple hardware regions that correspond one-to-one with each weight value. The area of each hardware region is positively correlated with the probability of occurrence of the corresponding weight value.
[0029] In some examples, the target large model is an artificial intelligence language model to be deployed to a hardware acceleration chip for inference operations. The target large model can be obtained through various channels, including downloading publicly available pre-trained models from open-source model libraries, training custom business models through deep learning training platforms, and importing enterprise-developed specialized adaptation models. For example, the target large model may include the Generative Pre-trained Transformer (GPT) series of models, the Large Language Model Meta AI (LLaMA) series of models, etc., which can adapt to different scales of inference needs from tens of billions to trillions of parameters.
[0030] The weight matrix is the parameter carrier of the target large model. It refers to a two-dimensional numerical set that represents the connection strength between input and output neurons in a network layer. The row dimension of the weight matrix corresponds to the output dimension of the model, and the column dimension corresponds to the input dimension of the model. Each element in the matrix represents the connection weight value between the corresponding input node and the output node. The weight matrix can be extracted from the parameter file of the target large model to extract the weight parameters of a specific network layer and organize them into a standard two-dimensional matrix structure according to the correspondence between the output dimension and the input dimension. Taking a single-layer fully connected network as an example, if the model input dimension is 4096 and the output dimension is 4096, the corresponding weight matrix contains 4096 rows and 4096 columns, totaling 16,777,216 independent weight elements.
[0031] The probability of each weight value is the basis for hardware resource allocation. It refers to the proportion of the frequency of each type of discrete weight value in the entire weight matrix out of the total number of elements, and is used to reflect the density of the distribution of different weight values in the model. First, the weight matrix is quantized to obtain a set of discrete weight values. Then, the total number of occurrences of each type of weight value is counted. The occurrence frequency of a single type of weight value is divided by the total number of elements in the weight matrix to calculate the probability of occurrence of the corresponding weight value. Taking a 1-bit quantized weight matrix as an example, if the matrix contains a total of 16 million elements, of which 9.6 million elements are in the on state and 6.4 million elements are in the off state, then the probability of occurrence of the on state weight value is 0.6 and the probability of occurrence of the off state weight value is 0.4.
[0032] Chip layout is the final output carrier of chip physical design. It refers to the physical design file generated in the back-end design stage of integrated circuits. It is used to characterize the geometry, position coordinates and connection relationships of semiconductor devices and via structures inside the chip. It is the direct basis for chip tape-out production. Chip layout can be output through application-specific integrated circuit (ASIC) back-end design tools, including complete physical design information such as chip device layout, metal layer structure, and power supply network. For example, chip layout can cover the entire process layer structure from the bottom active device layer to the top metal layer and can be directly delivered to wafer manufacturers for tape-out production.
[0033] The multiple hardware regions corresponding one-to-one with each weight value are the physical space carriers of weight encoding. They refer to independent physical blocks divided within the bit counting array of the chip layout. Each hardware region uniquely corresponds to a type of quantized weight value, and the corresponding weight routing resources and computing units are deployed within the region. The total number of such hardware regions can be determined based on the number of categories of quantized weight values, and then the area and coordinate range of each region can be allocated based on the probability of occurrence of each type of weight value. Taking a 4-bit quantized weight scenario as an example, after quantization, there are a total of 16 discrete weight values, which correspond to 16 independent hardware regions, each region corresponding to a weight value from 0 to 15.
[0034] For example, the total physical area of the bit counting array used for weight encoding in the chip layout can be extracted first. Then, based on the probability of occurrence of each type of weight value after quantization, the target area of the hardware region corresponding to each weight value is calculated. The target area is the product of the total physical area of the bit counting array and the probability of occurrence of the corresponding weight value. After the area calculation is completed, the hardware regions are arranged sequentially within the bit counting array range of the chip layout according to the weight values. Isolation spacing conforming to process design rules is reserved between adjacent regions to ensure that the physical boundaries of each region do not overlap. For high-frequency weight values with a high probability of occurrence, a larger hardware region is allocated, configuring more routing channels and computing unit resources. For low-frequency weight values with a low probability of occurrence, a smaller hardware region is allocated to control the occupation of redundant hardware resources. After all regions are divided, the weight value identifier, region coordinate range, and region area parameters corresponding to each hardware region are recorded to form weighted region division layout data, providing a physical space basis for subsequent weight mapping and routing processing.
[0035] By implementing step 101, the chip layout area is divided based on the probability of occurrence of each weight value in the target large model weight matrix, and the area of each hardware region is positively correlated with the probability of occurrence of the corresponding weight value. This allows different weight values to occupy matching physical resources according to their actual distribution characteristics. For weight values with higher occurrence probabilities, larger hardware regions can be allocated to meet their higher connectivity requirements. For weight values with lower occurrence probabilities, the resource occupation of the corresponding regions can be reduced, thereby avoiding the resource waste caused by uneven weight distribution in traditional fixed region division methods, improving the utilization rate of chip layout space and the rationality of hardware resource allocation.
[0036] Step 102: Based on the weight matrix and the physical pin information corresponding to the model input activation signal of the target large model, map each input dimension of the weight matrix to at least one corresponding hardware region to obtain the access mapping table.
[0037] In some examples, the model input activation signal is the feature activation signal input to the current computational layer when the target large model performs inference operations. It is the input data source for the bit counting statistics unit to perform counting operations. The value of the model input activation signal can be dynamically updated as the inference input content changes. It works in conjunction with the weight encoding structure fixed on the chip through metal connections to complete the multiplication, accumulation, and addition operations in the large model operation. The model input activation signal can be extracted from the network structure definition file of the target large model. The number and bit width of all input activation signals can be determined according to the input feature scale of the current computational layer. Taking the Transformer network of a large language model as an example, the input feature vector of the feedforward neural network layer is the model input activation signal of that layer. If the input feature dimension of the current layer is 4096, there are 4096 independent model input activation signals, each signal corresponding to the value of one feature dimension.
[0038] Physical pin information is a set of parameters for the physical ports used to access model input activation signals in the chip layout. It includes information such as the physical coordinates of each port, the metal layer to which it belongs, the signal name, and the port direction. It serves as the starting point for physical routing. Physical pin information can be extracted from the port constraint file of the application-specific integrated circuit back-end design, or the pin parameters corresponding to the input activation signals can be filtered from the chip input / output port layout planning file. Taking a large model acceleration chip as an example, its bit counter array has a dedicated signal port group on the input side, corresponding to the physical pins of 4096 input activation signals. The physical pin information includes the two-dimensional coordinates of each pin in the chip layout, the metal layer number to which it is connected, and the signal transmission bit width parameters.
[0039] Each input dimension of the weight matrix is an independent data channel divided along the column dimension of the weight matrix. Each input dimension corresponds to a column of data in the weight matrix and corresponds one-to-one with a model input activation signal, representing the weight connection relationship between that signal and all output dimensions. Each input dimension of the weight matrix can be obtained by performing dimension parsing on the weight matrix, with the column index as the unique identifier of the input dimension. The total number of input dimensions is exactly equal to the number of columns in the weight matrix. For example, for a weight matrix with a row and column size of 4096 by 4096, each input dimension corresponds to the 0th to the 4095th column of the matrix, totaling 4096 independent input dimensions. Each input dimension corresponds to the weight connection set of a model input activation signal.
[0040] The access mapping table is a structured data table that records the relationship between each input dimension and its corresponding target hardware region. It clarifies the hardware region that each input activation signal needs to reach through routing, and serves as guidance data for subsequent target metal layer routing. The access mapping table generates a standardized mapping data table by traversing all columns of the weight matrix, matching the hardware region corresponding to the valid weight value in each column, and summarizing all relationships. For example, in a 1-bit quantized weight scenario, the weight values are divided into two categories: on and off, corresponding to two independent hardware regions. The access mapping table only records the hardware region identifier to which the on-state weight belongs for each input dimension, and does not establish a mapping relationship for the off-state weight to avoid invalid routing.
[0041] For example, firstly, all physical pin information used to access the model input activation signal can be extracted from the chip layout to establish a one-to-one correspondence between each input dimension and its corresponding physical pin. Then, the weight matrix is traversed column by column. For each input dimension, all weight values contained in that column are extracted, the valid weight values are filtered out and matched with the corresponding hardware regions to establish a mapping relationship between the input dimension and the corresponding hardware region. For invalid weight values that are in the off state, no corresponding mapping relationship is established to avoid redundant routing resource occupation. After completing the traversal and matching of all input dimensions, the correspondence between all input dimensions and hardware regions is structured and organized to generate an access mapping relationship table containing input dimension identifiers, corresponding hardware region identifiers, and associated physical pin information, providing clear path endpoints and target region basis for subsequent routing processing in the target metal layer.
[0042] By implementing step 102, a clear connection relationship is established between the model input activation signal and the physical region corresponding to the weight. Since this mapping process determines the correspondence between the input dimension and the hardware region based on the data characteristics of the weight matrix, it can reduce the disordered connection between the input signal and the hardware region, reduce the uncertainty of the connection relationship in the subsequent wiring process, make the wiring path planning more accurate, and improve the feasibility of the chip physical implementation process.
[0043] Step 103: Based on the access mapping table, perform routing processing within the target metal layer of the chip layout to obtain the target routing layout of the chip layout.
[0044] In some examples, the target metal layer is a dedicated metal layer in the chip's multi-layer metal interconnect structure, designated to carry the weight encoding routing of large models. This layer directly represents the weight values of the large model through the connectivity and disconnection of the metal lines, realizing a weight-as-connection in-memory computing architecture and serving as the physical carrier of metal embedding technology. The target metal layer can be pre-specified by the designer based on the chip process node, the electrical characteristics of the metal layer, and the weight deployment requirements. Strict hierarchical constraints are set during the routing process to ensure that all weight-related routing operations are completed within the target metal layer. For example, in the design of a large model acceleration chip at a 12-nanometer process node, the sixth metal layer at the top can be selected as the target metal layer. This layer has a larger metal thickness and lower parasitic resistance, making it suitable for long-distance transmission of high-frequency input activation signals. The bottom metal layers are used for basic device interconnection, power distribution, and clock tree distribution, and do not participate in weight encoding routing.
[0045] The target routing layout is the final physical design file that meets the consistency requirements of process design rules and weight logic after weighted coding routing processing. This layout completely records the geometry, position coordinates, and connection relationships of all weighted coding metal lines within the target metal layer, corresponding to the effective connections of the large model weight matrix. It is the delivery data for chip back-end design and tape-out manufacturing. The target routing layout can be generated through an automated routing process, using the access mapping table as the routing basis, completing the path layout of all effective connections within the target metal layer, and outputting a standard layout data format that can be directly imported into application-specific integrated circuit back-end design tools for subsequent design work.
[0046] For example, firstly, the layer number and routable area of the target metal layer can be confirmed, the basic framework data of the chip layout and the design rules of the corresponding process node can be loaded, and the process constraint parameters such as the minimum line width and minimum line spacing within the target metal layer can be clarified. Then, the access mapping table can be imported, the physical pin coordinates and target hardware area identifiers corresponding to each input dimension can be extracted, and the starting coordinates and ending area boundaries of each routing can be established one by one. After starting the automated routing process, the metal lines are laid out along the plane direction of the target metal layer to the corresponding hardware area, starting from the coordinates of each input pin. During the routing process, the process design rules must be strictly followed to ensure that the spacing between adjacent metal lines meets the minimum spacing requirement and the width of the metal lines meets the minimum line width requirement. At the same time, it must be ensured that all metal lines are within the target metal layer and do not extend to other metal layers. For multiple input metal lines mapped to the same hardware area, they are arranged and accessed sequentially along the access side boundary of the hardware area to avoid routing congestion at the area entrance. After all effective connections are laid out, all routing geometry data and location information within the target metal layer are organized, and a complete layout data file is output to form the target routing layout of the chip layout, providing a routing basis for the subsequent chip physical design process.
[0047] By implementing step 103, the connection relationship corresponding to the weight matrix can be directly converted into the physical metal connection structure in the chip layout, thereby forming the target wiring layout corresponding to the target large model weight. Since the wiring process is directly driven by the weight mapping relationship and is limited to the specified metal layer, the association conversion between weight information and chip physical structure can be realized, reducing the additional weight storage mapping process, while improving the flexibility of wiring structure adjustment and reducing the complexity of chip layout generation.
[0048] In summary, this application's embodiments divide the chip layout area based on the occurrence probability of each weight value in the target large model weight matrix. Since the area of each hardware region is positively correlated with the occurrence probability of the corresponding weight value, weight values with higher occurrence frequencies can obtain a larger physical implementation area, while weight values with lower occurrence frequencies occupy only a smaller area. Thus, the allocation of hardware resources in the chip layout can match the data distribution characteristics of the model weights. Compared with a fixed division method, it can reduce the problem of insufficient area utilization caused by uneven weight distribution, making the limited layout space more rationally allocated and improving the utilization efficiency of chip wiring resources and hardware areas. By establishing an access mapping relationship table based on the weight matrix and the physical pin information corresponding to the model input activation signals, the input dimensions in the weight matrix are associated with the corresponding hardware areas, so that the input activation signals can be mapped according to the weight correspondence. The signal is accurately guided to the target hardware region, allowing the mapping process to directly determine the signal access path using the data characteristics of the weight matrix. This avoids the wiring complexity caused by irregular connections between the input signal and the weight region, reduces the uncertainty of wiring connections, and improves the accuracy and feasibility of subsequent wiring processing. Based on the access mapping table, wiring processing is performed only within the target metal layer of the chip layout, enabling weight-related connections to be converted into physical metal connections according to a pre-determined mapping relationship. This forms a target wiring layout corresponding to the target large model weights, allowing the wiring process to be directly driven by the weight matrix and confined to the target metal layer. It eliminates the need for additional weight storage mapping processes to achieve the conversion of weight information to physical layout connections, while also facilitating independent adjustments to the wiring structure corresponding to the weights, thus improving the flexibility of the chip layout generation process. In summary, the chip wiring layout determination method provided in this application, through layout region division, input mapping, and target metal layer wiring based on the large model weight distribution characteristics, can achieve direct conversion of weight information to the chip's physical connection structure, thereby improving layout resource utilization efficiency, reducing wiring complexity, and enhancing the flexibility of chip layout generation.
[0049] In some embodiments, step 102 may include: abstracting and generating multiple virtual access points based on physical pin information, wherein each virtual access point corresponds one-to-one with each input dimension of the weight matrix; and mapping the virtual access points to at least one corresponding hardware region based on the weight matrix to obtain an access mapping table.
[0050] In some examples, multiple virtual access points are standardized sets of routing start nodes formed by logically abstracting the physical pins corresponding to the model input activation signals. Each virtual access point uniquely corresponds to one input dimension of the weight matrix and one model input activation signal. This unifies scattered physical pins into logical access nodes of the same specification, eliminating the interference caused by the scattered distribution of physical pins on routing planning and providing a regular starting benchmark for subsequent unified routing. During implementation, the physical pin information corresponding to all input activation signals can be read, a unique virtual access point identifier can be assigned to each input dimension, and the virtual access point identifier can be matched with the corresponding physical pin. Coordinate information and signal attributes are bound together to complete the generation and assembly of all virtual access points. Physical pin information can be extracted from the port constraint file of the application-specific integrated circuit backend design to ensure the accurate one-to-one correspondence between virtual access points and physical ports. Taking the feedforward network layer of a large language model with an input dimension of 4096 as an example, 4096 independent virtual access points are generated. Each virtual access point corresponds to one input dimension and one physical input pin. All virtual access points adopt a unified numbering rule and data format. Subsequent mapping and routing planning are all performed based on virtual access points as the basic unit, without the need to directly process scattered physical pin parameters.
[0051] Based on the weight matrix, the process of mapping virtual access points to at least one corresponding hardware region to obtain the access mapping table can be achieved by traversing all the data in the weight matrix column by column. For each virtual access point in a column, all weight values contained in that column are extracted, valid weight values are filtered and matched with the corresponding hardware region to establish the association between the virtual access point and the corresponding hardware region. After summarizing all the associations, a standardized access mapping table is generated. For example, in a 1-bit quantized weight deployment scenario, the weight values are divided into two categories: on-state and off-state, corresponding to two independent hardware regions within the chip layout. During the traversal, if the weight column corresponding to a virtual access point contains an on-state weight, then the virtual access point is mapped to the hardware region corresponding to the on-state weight. No mapping relationship is established for off-state weights to avoid invalid routing. The final generated access mapping table contains the identifiers of all virtual access points, the identifiers of the corresponding target hardware regions, and the bound physical pin information.
[0052] For example, firstly, the extracted physical pin information can be imported, and the physical pins can be sorted according to the input dimensions of the weight matrix. A consecutive virtual access point number can be assigned to each input dimension. Simultaneously, the coordinate data and signal width parameters of the corresponding physical pins can be bound to the corresponding numbered virtual access points, forming a complete set of virtual access points. After generating the virtual access points, the quantized data of the weight matrix is loaded, and traversal processing is performed column by column, one by one, based on the input dimensions. For each virtual access point's corresponding weight column, all weight values in the column are read sequentially, and the non-disconnected valid weight values are identified. Based on the weight values and the hardware area... A one-to-one correspondence is established to match the hardware area identifier corresponding to the valid weight value, and a mapping relationship between the current virtual access point and the hardware area is established. If the weight column corresponding to the same virtual access point contains multiple types of valid weight values, a mapping relationship between the virtual access point and the hardware areas corresponding to the multiple weights is established respectively. After all virtual access points have been traversed, all mapping relationships are sorted according to the number order of the virtual access points, and an access mapping relationship table containing the virtual access point identifier, the corresponding hardware area identifier, and the bound physical pin coordinates is generated. All data is stored in a unified structured format and can be directly imported into the subsequent cabling process as the basis for path planning.
[0053] Through the implementation of the above embodiments, a unified and clear connection relationship is formed between the input activation signal and the weight region. Compared with the method of directly wiring based on physical pins, the embodiments of this application use virtual access points to abstractly manage the input signals, which can reduce the complexity of the input signal connection relationship and enable the subsequent wiring process to be executed according to the determined mapping relationship. This can improve the efficiency of wiring planning and reduce the waste of wiring resources caused by the uncertainty of the connection relationship.
[0054] In some embodiments, step 103 may include: based on a comprehensive routing cost function, performing a routing process from each virtual access point to the corresponding hardware region in the access mapping table within the target metal layer of the chip layout to obtain an initial routing layout; generating pin access points for each bit counter in the initial routing layout based on the total number of target input lines in each hardware region and the upper limit of the input capability of a single bit counter, to obtain an intermediate routing layout, wherein the total number of target input lines is the statistical number of virtual access points mapped to the hardware region, and the upper limit of the input capability of a single bit counter is the maximum number of access channels supported by a single bit counter determined based on preset timing constraints and physical design rules; performing a preset verification optimization process on the intermediate routing layout, and determining the verified intermediate routing layout as the target routing layout.
[0055] In some examples, the structured routing cost function is a mathematical calculation function used to quantitatively evaluate the overall performance of routing paths. It serves as the basis for selecting the optimal path during automatic routing and can integrate multi-dimensional chip design constraints to quantitatively rank the merits of different candidate paths, ensuring that the final selected routing path matches the design goals. It can be constructed by selecting corresponding evaluation dimensions and setting weight coefficients for each dimension based on the chip's timing performance requirements, process design rules, and routing optimization goals, using a weighted summation method. The weight coefficients can be flexibly adjusted according to design priorities. For example, in large-model weighted coding routing scenarios, the structured routing cost function can incorporate evaluation dimensions such as routing length and process compliance. When the design priority leans towards timing performance, the weight ratio of dimensions related to signal delay can be increased, guiding the routing engine to prioritize routing paths with lower delays.
[0056] The routing process from each virtual access point to the corresponding hardware area in the access mapping table is an automated process that uses virtual access points as the starting point and the corresponding hardware area as the ending point to complete the metal interconnection within the routerable range of the target metal layer. It is the execution step that transforms logical mapping relationships into physical metal interconnections. This routing process is executed by an automated routing engine, using the access mapping table as input, to plan and generate continuous metal interconnection paths within the planar space of the target metal layer. For large model computing layers containing thousands of input signals, the routing process can sequentially process the interconnection requirements of each virtual access point, guiding each input activation signal from its corresponding virtual access point location to the target hardware area, completing the physical deployment of all effective weighted connections. The initial routing layout is the intermediate layout file after automatic routing processing, before bit counting unit pin configuration and verification optimization. It contains the path geometry information of all metal interconnections within the target metal layer and serves as the foundation for subsequent pin configuration and verification optimization.
[0057] The total number of target input lines is the total number of input metal lines connected to a single hardware area, reflecting the scale of the input signals in that hardware area and serving as the basis for configuring the bit counting unit and the number of pins. The total number of target input lines can be obtained by counting the total number of virtual access points mapped to the corresponding hardware area. Each virtual access point mapped to that area corresponds to an independent input metal line. If a hardware area corresponding to a certain weight value has a total of 256 virtual access points mapped, then the total number of target input lines in that hardware area is 256, which corresponds to the need to connect 256 independent input activation signals.
[0058] The bit counting unit is a computing unit in a metal-embedded in-memory computing architecture. It is used to perform bit counting statistics on the multiple input activation signals, which is equivalent to the multiplication and accumulation calculation in large model operations. It is the hardware carrier for realizing weight as connection and connection as computation. The upper limit of the input capability of a single bit counter unit is the maximum number of input signal channels that a single bit counter unit can stably support. It is a parameter that measures the computational carrying capacity of the bit counter unit and determines the scale of input signals that a single unit can process. The upper limit of the input capability of a single bit counter unit is the maximum number of access channels that a single bit counter unit can support, determined based on preset timing constraints and physical design rules. It is a clear limitation on the basis for determining the upper limit of the input capability, indicating that the upper limit is not a fixed design value, but a reasonable maximum value derived by combining performance requirements and process constraints. The preset timing constraints specify the maximum delay requirements for signal transmission and computation processing, and the physical design rules specify process limitations such as the area of the unit and the pin spacing. The two together constrain the maximum number of channels that a single unit can access, ensuring that the unit can operate stably at the rated operating frequency. For example, a bit counter unit designed based on a certain 12-nanometer process node has an upper limit of 256 input channels, that is, it can simultaneously access a maximum of 256 input activation signals and complete stable bit counting operations.
[0059] Pin access points are physical port nodes on a bit counter unit used to connect to input metal lines. They serve as the connection interface between the input metal lines and the internal computing circuitry of the bit counter unit. Each pin access point corresponds to an independent input signal channel. Pin access points are generated based on the number of input channels and physical layout of the bit counter unit. Corresponding access nodes are arranged on the interface side of the unit, and the coordinate position and signal attributes of each access point are clearly defined. For example, a bit counter unit with a single input capability of up to 256 corresponds to 256 pin access points, which are evenly arranged along the input edge of the unit. Each access point can be connected to an input metal line to guide external input signals into the counting circuitry inside the unit.
[0060] The intermediate routing layout is an intermediate layout file after the bit counting unit layout and pin access point configuration are completed. It is the product of the initial routing layout after pin configuration optimization. It already contains complete metal interconnects and computing unit layout, and can directly enter the subsequent verification and optimization stage. Based on the initial routing layout, the bit counting unit deployment and pin access point generation of each hardware area can be completed. After integrating the unit physical information and pin information into the original layout, the intermediate routing layout is obtained.
[0061] The process of performing a preset verification and optimization process on the intermediate routing layout and determining the intermediate routing layout that passes the verification as the target routing layout can be carried out by a dedicated backend verification tool. Preset verification items can be executed, and the layout can be iteratively optimized based on the verification results. After all verifications pass, the final target routing layout is output. The preset verification and optimization process may include verification items such as process compliance checks and logic consistency checks. The design quality of the intermediate routing layout is verified item by item. If there are any violations, the corresponding routing paths or cell layouts are adjusted until all verification items pass. The layout at this point is the final target routing layout.
[0062] For example, firstly, a preset integrated cabling cost function and target metal layer design rule constraints can be loaded, the access mapping table and coordinate information of each virtual access point can be imported, and the automated cabling process can be started. The cabling engine takes each virtual access point as the starting point and its mapped hardware area as the ending target, searches for cabling paths that meet the cost function requirements within the target metal layer, and outputs the initial cabling layout after completing the deployment of all valid connections. Then, it traverses each hardware area, counts the total number of target input lines corresponding to each hardware area, and configures a corresponding number of bit counting units for each hardware area in combination with the upper limit of the input capability of a single bit counting unit, and generates pin access points that match the input line scale. The bit counting units and pin information are integrated into the initial cabling layout to obtain the intermediate cabling layout. Finally, a preset verification and optimization process is started to perform multi-dimensional compliance and functional verification on the intermediate cabling layout. If there are problems that do not meet the design requirements, the cabling path or the layout position of the bit counting units are adjusted accordingly. After the adjustment is completed, the verification is performed again until all verification items meet the design requirements, and finally a qualified target cabling layout is output.
[0063] Through the implementation of the above embodiments, the routing path from the virtual access point to the corresponding hardware area is optimized based on the comprehensive routing cost function, and a corresponding number of pin access points are generated according to the actual number of input lines in the hardware area and the upper limit of the input capability of the bit counting unit, so that the physical access capability of the bit counting unit can match the actual weight connection requirements. By performing a verification and optimization process on the generated intermediate routing layout, it can be further ensured that the routing result meets the chip manufacturing requirements, and the problems of hardware resource redundancy or insufficient access capability caused by a fixed number of pins can be avoided, thereby improving the resource utilization of the bit counting unit and the manufacturability of the chip routing layout.
[0064] In some embodiments, before obtaining the initial routing layout of the chip layout by executing a routing process from each virtual access point to the corresponding hardware area in the access mapping table within the target metal layer of the chip layout based on the comprehensive routing cost function, the chip routing layout determination method may further include: obtaining routing loss assessment data of the target metal layer, wherein the routing loss assessment data may include resistance value per unit length, capacitance value per unit area, and via parasitic parameter value; obtaining power distribution network layout information of the chip layout, wherein the power distribution network layout information may include power line location coordinates, power pad coordinates, and power strip distribution data; and constructing a comprehensive routing cost function based on the routing loss assessment data and the power distribution network layout information, wherein the comprehensive routing cost function is obtained by weighted summation of routing resistance cost term, routing capacitance cost term, design rule violation cost term, and power distribution network conflict cost term.
[0065] In some examples, wiring loss assessment data is a set of physical parameters used to quantify the signal transmission loss of metal wiring within a target metal layer. It encompasses the parasitic electrical parameters of the metal trace itself and the interlayer connection structure, serving as the foundational data for calculating wiring signal delay and transmission loss. Wiring loss assessment data can be directly extracted from the process design suite of the corresponding process node, or obtained through simulation calculations of the target metal layer structure using dedicated parasitic parameter extraction tools. Wiring loss assessment data can include resistance per unit length, capacitance per unit area, and via parasitic parameter values. Resistance per unit length is the parasitic resistance value per unit length of metal trace within the target metal layer, a parameter that measures wiring resistance loss. Its value is directly related to the resistivity and thickness of the metal layer material. For example, if the resistance per unit length of the sixth metal layer at a certain process node is 0.05 ohms per micrometer, then a 1000-micrometer-long metal trace in this layer has a total parasitic resistance of 50 ohms. The unit area capacitance is the parasitic capacitance value formed between a unit area of metal traces within a target metal layer and the surrounding metal structure and substrate. It is a parameter affecting signal transmission delay and power consumption, and its value is directly related to the thickness of the interlayer dielectric and the metal area. For example, the unit area capacitance value of the sixth metal layer at a certain process node is 0.8 femtofarads per square micrometer, meaning that the total parasitic capacitance of a metal trace in this layer with an area of 1000 square micrometers is 800 femtofarads. The via parasitic parameter value is the parasitic resistance and capacitance value of via structures used to connect different metal layers. Vias are connection nodes for interlayer signal transmission, and their parasitic parameters introduce additional signal transmission losses. For example, a standard-sized via connecting the fifth and sixth metal layers has a parasitic resistance of 5 ohms and a parasitic capacitance of 0.2 femtofarads. Each additional interlayer via connection introduces a corresponding value of parasitic loss.
[0066] Power distribution network (PDN) layout information is a complete set of physical parameters for the power distribution network (PDN) in the chip layout. It characterizes the layout location and electrical constraints of the power supply network and serves as a reference for routing avoidance and ensuring power supply stability. PDN layout information can be fully exported from the power planning file of the chip back-end design, including the geometric layout information and electrical constraint parameters of the power supply network. PDN layout information can include power line location coordinates, power pad coordinates, and power strip distribution data. Power line location coordinates are the physical coordinate information of the metal traces used to transmit power supply voltage in the chip layout. They include the start coordinates, end coordinates, trace width parameters, and the metal layer information to which the power line belongs. They are the transmission carrier of the power supply network. For example, multiple horizontal global power lines are arranged in the top metal layer of the chip. Each power line corresponds to specific start and end coordinates and trace width parameters. Signal routing must maintain a safe distance from these power lines in accordance with regulations. Power supply pad coordinates are the physical location coordinates of the pad structures on the chip used for external power input. They are the input nodes of the power supply network, and the external power supply is connected to the internal power supply network of the chip through the pads. For example, dozens of power input pads are evenly distributed around the chip, each with a unique center coordinate and size parameters. During routing, it is necessary to avoid overlapping with the pad structures or insufficient spacing. Power supply strip distribution data is the layout and distribution information of wide power supply metal strips used for regional power supply in the chip. It includes the position, width, coverage area, and current carrying capacity parameters of the power supply strips, which are used to provide stable power supply with high current and low voltage drop to local circuit areas. For example, multiple vertical wide power supply strips are arranged in the bit counting array area. The width, spacing, and coverage area of each power supply strip are recorded in the power supply strip distribution data, which is an important reference for regional routing avoidance.
[0067] The process of constructing a comprehensive cabling cost function based on cabling loss assessment data and power distribution network layout information is a processing flow that integrates multiple constraints such as electrical loss, process compliance, and power supply compatibility to generate a mathematical function that can quantify the merits of cabling paths. It is a prerequisite for achieving high-quality automated cabling. This process extracts the calculation rules for four types of cost items, sets the corresponding weight coefficients for each cost item, and combines them into a complete comprehensive cabling cost function through weighted summation. For high-frequency large model acceleration chips, the weight coefficients of the cabling resistance cost item and the power distribution network conflict cost item can be increased to construct a comprehensive cabling cost function that prioritizes timing performance and power supply stability, adapting to the high-frequency operation requirements of large model inference. The four types of cost items include cabling resistance cost item, cabling capacitance cost item, design rule violation cost item, and power distribution network conflict cost item.
[0068] The routing resistance cost term is an evaluation component in the synthetic routing cost function used to quantify the total resistance loss of routing. Its value is positively correlated with the total routing length and the resistance per unit length. The higher the resistance loss, the larger the corresponding cost term value, which is used to guide routing to prioritize paths with lower resistance loss. This cost term can be calculated by multiplying the total length of the candidate routing path by the resistance per unit length to obtain the total resistance, and then multiplying by the corresponding weighting coefficient. For example, if the total length of a candidate routing path is 1000 micrometers, the resistance per unit length is 0.05 ohms, and the total resistance is 50 ohms, the routing resistance cost term value of that path can be obtained by multiplying by the corresponding weighting coefficient.
[0069] The cabling capacitance cost term is an evaluation component in the synthetic cabling cost function used to quantify the total capacitance loss of cabling. Its value is positively correlated with the total cabling area and the capacitance per unit area. The higher the capacitance value, the larger the corresponding cost term value, which is used to guide cabling to prioritize paths with lower capacitance delay. This cost term is calculated by multiplying the total area of the candidate cabling path by the capacitance per unit area to obtain the total capacitance, and then multiplying by the corresponding weighting coefficient. For example, if the total area of a candidate cabling path is 2000 square micrometers, the capacitance per unit area is 0.8 femtofarads, and the total capacitance is 1600 femtofarads, the cabling capacitance cost term value of that path can be obtained by multiplying by the corresponding weighting coefficient.
[0070] The design rule violation cost item is an evaluation component in the comprehensive routing cost function used to quantify the degree to which routing violates process design rules. It covers various scenarios such as trace width violations, trace spacing violations, and via size violations. The more violations there are and the more severe the violation, the larger the corresponding cost item value. It is used to constrain routing to meet the chip fabrication process requirements. This cost item is calculated by counting the number of design rule violations in candidate routing paths, assigning a value based on the severity of the violation, and then multiplying it by the corresponding weight coefficient. For example, if a candidate routing path has 2 trace spacing violations, 1 via size violation, and a total of 3 process violations, multiplying it by the corresponding weight coefficient will give the design rule violation cost item value for that path.
[0071] The power distribution network conflict cost term is an evaluation component in the structured cabling cost function used to quantify the degree of conflict between cabling and the power distribution network. It covers conflict scenarios such as line overlap and spacing less than the safe value. The more frequent and severe the conflicts, the larger the corresponding cost term value, which is used to guide cabling to actively avoid the power distribution network structure. This cost term is calculated by counting the number of conflicts between candidate cabling paths and the power distribution network structure, assigning a value based on the severity of the conflict, and then multiplying it by the corresponding weight coefficient. For example, if a candidate cabling path has one overlap and one insufficient spacing with a power line, accumulating two power distribution network conflicts, multiplying it by the corresponding weight coefficient will give the value of the power distribution network conflict cost term for that path.
[0072] For example, firstly, wiring loss assessment data for the target metal layer can be extracted from the process design suite of the corresponding process node. The unit length resistance value, unit area capacitance value, and parasitic parameter values of the corresponding interlayer vias are read and organized into a standardized parasitic parameter dataset. Then, the chip power distribution network layout information is imported, and the power line location coordinates, power pad coordinates, and power strip distribution data are extracted to establish a complete set of power network obstacle coordinates, clarifying the boundary range and safety distance requirements for various power supply structures. After completing the basic parameter extraction, the wiring resistance cost item, wiring capacitance cost item, design rule violation cost item, and other related factors are determined. The calculation rules for power distribution network conflict cost items take into account the timing performance requirements, process manufacturing requirements, and power supply reliability requirements of the chip. Weight coefficients are set for the four cost items. When timing priority is high, the weight of resistor and capacitor cost items is increased; when power supply reliability priority is high, the weight of power distribution network conflict cost items is increased. Finally, the four cost items are integrated into a unified synthetic routing cost function through weighted summation. This function is configured into the automatic routing engine to provide a unified quantitative evaluation basis for subsequent routing path search, ensuring that the final generated routing path takes into account electrical performance, process compliance, and power supply compatibility.
[0073] Through the implementation of the above embodiments, before routing, routing loss assessment data of the target metal layer and chip power distribution network layout information are obtained. Based on factors such as routing resistance, routing capacitance, design rule violations, and power distribution network conflicts, a comprehensive routing cost function is constructed, enabling the routing path selection process to simultaneously consider signal transmission loss, process constraints, and power supply network influence. Compared with routing based solely on geometric distance, the embodiments of this application can reduce the routing quality degradation caused by ignoring parasitic parameters or power distribution network constraints, and improve the reliability of the routing layout.
[0074] In some embodiments, the aforementioned routing process based on the comprehensive routing cost function, executing a routing flow from each virtual access point to the corresponding hardware region in the access mapping table within the target metal layer of the chip layout to obtain the initial routing layout of the chip layout, may include: using each virtual access point as the routing start point and the corresponding hardware region in the access mapping table as the routing end point, generating multiple candidate routing paths within the target metal layer using a preset three-dimensional routing algorithm; determining the cost value of each candidate routing path based on the comprehensive routing cost function; determining the candidate routing path with the lowest cost value among the multiple candidate routing paths as the target routing path; and generating the initial routing layout based on the target routing path.
[0075] In some examples, the preset 3D routing algorithm is a computational algorithm that is pre-integrated into the automated routing engine and can automatically search for feasible routing paths in the 3D space of the chip's metal interconnect structure. It can simultaneously take into account planar routing planning, inter-layer connection configuration, and obstacle avoidance, and serves as the computational carrier for generating multiple sets of candidate routing schemes. The preset 3D routing algorithm can be obtained by integrating a mature commercial routing engine kernel, or it can be custom-developed based on classic routing algorithms such as maze search and path expansion. Before use, search constraints and iteration parameters can be configured according to the target metal layer range and process design rules.
[0076] Candidate routing paths are multiple optional metal connection paths that meet basic connectivity requirements, generated by the 3D routing algorithm in the specified routing start and end areas. Each path contains complete routing coordinates, line width parameters, and layer affiliation information, serving as alternatives for subsequent cost evaluation and optimal selection. Candidate routing paths can be generated by a preset 3D routing algorithm through multi-directional path expansion search. A single routing requirement can generate 3 to 8 candidate paths with different directions, covering different detour schemes and resource consumption. For example, for a routing requirement from a virtual access point to the corresponding hardware area, the preset 3D routing algorithm can generate 4 candidate routing paths, which reach the target hardware area by detouring from above, going straight below, and meandering in the middle, respectively. The total length, number of turns, and distance to surrounding obstacles of each path are different.
[0077] Cost value is a quantitative value calculated based on the structured cabling cost function. It is used to comprehensively measure the electrical loss, process compliance, and power supply compatibility of a single cabling path. The lower the cost value, the better the overall performance of the path, and it serves as the basis for selecting the optimal cabling path. The cost value can be calculated by substituting parameters such as the total length, total area, number of vias, number of design rule violations, and number of power distribution network conflicts of a single candidate cabling path into the structured cabling cost function and then performing a weighted summation. For example, a candidate cabling path with a total length of 1200 micrometers, no design rule violations, and no power distribution network conflicts has a cost value of 12.5 calculated by substituting these parameters into the structured cabling cost function. Another candidate path has a total length of 900 micrometers but has one power distribution network spacing violation, resulting in a cost value of 15.2. Therefore, the former has better overall performance.
[0078] The target routing path is the routing path with the lowest cost selected from multiple candidate routing paths with the same start and end points. It is the routing scheme with the best overall performance under this routing requirement and is used to finally draw it into the routing layout. The target routing path is selected by comparing the cost of all candidate routing paths under the same routing requirement. If there are multiple paths with the same cost, the path with shorter routing length and fewer turns can be further prioritized. For example, if a certain routing requirement generates 4 candidate paths with corresponding cost values of 12.5, 15.2, 13.8 and 14.1, the path with the lowest cost value of 12.5 is determined as the target routing path for this routing requirement.
[0079] The process of generating an initial routing layout based on the target routing path is an integration process that transforms the optimal path scheme into standard physical layout data. It integrates the target routing paths corresponding to all routing requirements into the layout framework of the target metal layer to form a complete initial routing layout file. This process is executed by the layout generation module, which converts the target routing paths corresponding to all virtual access points into geometric data that meets the process requirements and writes them into the layout database of the target metal layer to form an initial routing layout containing all valid weighted connections.
[0080] For example, the configuration and import of the comprehensive cabling cost function can be completed first, clarifying the cabling range and obstacle set of the target metal layer, and loading the mapping data of all virtual access points and corresponding hardware areas in the access mapping table. The cabling engine processes the cabling requirements of each virtual access point one by one, setting the physical pin coordinates corresponding to the virtual access point as the cabling start point, setting the access side boundary of the corresponding hardware area as the cabling end point area, calling the preset three-dimensional cabling algorithm and limiting the search space to only cover the target metal layer, and generating multiple candidate cabling paths for each cabling requirement through multi-directional path expansion. After the candidate path generation is completed, the trace length, trace area, design rule violation and power distribution network conflict of each candidate path are extracted in sequence, and substituted into the comprehensive cabling cost function to calculate the cost value of each path. The cost values of all candidate paths under the same cabling requirement are compared, and the path with the smallest cost value is selected as the target cabling path for the requirement. After the target cabling paths of all cabling requirements are filtered, all target cabling paths are converted into standard format layout geometry data and uniformly integrated into the layout framework of the target metal layer to generate a complete initial cabling layout, providing basic layout data for subsequent bit counting unit pin configuration.
[0081] Through the implementation of the above embodiments, virtual access points are used as the starting point of cabling, corresponding hardware areas are used as the ending point of cabling, and multiple candidate cabling paths are generated using a three-dimensional cabling algorithm. Then, the path with the lowest cost is selected as the target cabling path based on the comprehensive cabling cost function. This enables the cabling process to select a more suitable connection path according to the actual layout environment, which can reduce path redundancy caused by fixed cabling methods, improve the utilization efficiency of target metal layer cabling resources, and improve the matching degree between the cabling structure and the large model weight mapping relationship.
[0082] In some embodiments, the aforementioned method of generating pin access points for each bit counting unit in the initial routing layout based on the total number of target input lines in each hardware region of the initial routing layout and the upper limit of the input capability of a single bit counting unit to obtain an intermediate routing layout may include: for each hardware region, if the total number of target input lines is less than or equal to the upper limit of the input capability, then a single bit counting unit is configured in the corresponding hardware region, and multiple input pins are generated equal to the total number of target input lines; if the total number of target input lines is greater than the upper limit of the input capability, then based on the ratio of the total number of target input lines to the upper limit of the input capability, the corresponding hardware region is divided into multiple hardware sub-regions, and a bit counting unit and multiple corresponding input pins are configured for each hardware sub-region; the configured bit counting units and input pins are integrated into the initial routing layout to obtain the intermediate routing layout.
[0083] In some examples, all hardware regions can be traversed sequentially, and the total number of target input lines for each region can be extracted one by one. This number is then compared with a pre-defined upper limit for the input capability of a single bit counter unit. When the total number of target input lines is less than or equal to the upper limit, a single bit counter unit is deployed in that hardware region, and an input pin count exactly equal to the total number of target input lines is generated, without any additional redundancy. When it is detected that the total number of target input lines in a hardware region exceeds the upper limit for the input capability of a single unit, the total number of target input lines is divided by the upper limit, and the result is rounded up to obtain the required total number of bit counter units and the total number of hardware sub-regions. Subsequently, the original hardware region is evenly divided into a corresponding number of hardware sub-regions, and each sub-region is allocated a number of input lines not exceeding the upper limit for the input capability. Each sub-region is then configured with a bit counter unit and a corresponding number of input pins. For example, if the upper limit for the input capability of a single bit counter unit is set to 256 channels, and the total number of target input lines for a hardware region corresponding to a certain low-frequency weight is 172, then... If the value is less than 256, a single bit counter unit is configured in this hardware area, generating 172 input pins simultaneously. It is not necessary to fully configure 256 pins; the hardware routing and unit resources corresponding to the excess pins can be directly released. The total number of target input lines in the hardware area corresponding to a certain high-frequency weight is 620, which is greater than 256. The ratio of 620 to 256 is calculated and rounded up to 3. The hardware area is then divided into three hardware sub-areas, allocated with 207, 207, and 206 input lines respectively. Each sub-area is configured with a bit counter unit, generating the corresponding number of input pins. The input line size of all sub-areas does not exceed the capacity limit of a single unit. The physical layout data of the bit counter unit can be retrieved from the ASIC standard cell library. The unit is precisely placed in the designated position of the corresponding hardware area or sub-area. The physical coordinates of each input pin are aligned with the ends of the corresponding input metal lines. The metal connection relationships and layer data within the layout are updated synchronously. After completing the integration and verification of all units and pins, a complete intermediate routing layout is output.
[0084] For example, the initial wiring layout data and the target input line count parameters corresponding to each hardware region can be loaded first. The pre-set upper limit value of the input capability of a single-pin counter unit can be retrieved, and all hardware regions can be traversed sequentially according to the weighted region numbering order. For a single hardware region, its target input line count is read and compared with the upper limit of the input capability. If the target input line count is less than or equal to the upper limit of the input capability, the single-pin counter unit is placed at the center of the hardware region, and input pins equal to the target input line count are evenly distributed along the access side edge of the hardware region, ensuring the pin direction is consistent with the wiring direction of the input metal lines, reducing the number of wiring bends and vias, and lowering parasitic losses. If the target input line count is greater than the upper limit of the input capability, the ratio of the target input line count to the upper limit of the input capability is calculated. The ratio is calculated and rounded up to obtain the number of sub-regions. The original hardware area is then evenly divided into the corresponding number of hardware sub-regions along the wiring direction. A balanced input line scale is allocated to each sub-region to ensure that the number of input lines in each sub-region does not exceed the input capacity limit. Subsequently, a bit counting unit is placed at the center of each sub-region, and the corresponding number of input pins are generated along the access edge of each sub-region. After completing the bit counting unit configuration and pin generation for all hardware areas, the physical layout of all bit counting units, the geometric data of the input pins, and the original metal wiring data are integrated. The connection relationship between each input metal line and the corresponding pin is verified. After confirming that there are no open circuits, short circuits, or other connection problems, a complete intermediate wiring layout is generated, providing a complete layout data foundation for subsequent verification and optimization processes.
[0085] Through the implementation of the above embodiments, when the number of input lines does not exceed the input capacity of a single bit counting unit, only the corresponding number of input pins are configured to avoid additional hardware resource occupation; when the number of input lines exceeds the input capacity, the access capacity is expanded by splitting the hardware area and configuring multiple bit counting units, which can make the configuration scale of the bit counting unit match the actual weight connection requirements, reduce hardware resource waste, and improve the resource utilization efficiency of the chip wiring layout.
[0086] In some embodiments, the aforementioned execution of a preset verification and optimization process on the intermediate routing layout, determining the verified intermediate routing layout as the target routing layout, may include: performing a design rule check on the intermediate routing layout to obtain a first verification result; and performing a layout and schematic check on the intermediate routing layout. Figure 1 The first verification is performed to obtain the second verification result; the second verification result is obtained by performing power distribution network compatibility verification on the intermediate routing layout; when the first verification result, the second verification result, and the third verification result are all passed, the intermediate routing layout is determined as the target routing layout; otherwise, the routing path and / or bit counting cell layout of the intermediate routing layout are adjusted until all verifications are passed.
[0087] In some examples, the first verification result is to perform a design rule check on the intermediate routing layout. The compliance judgment result output by RuleCheck (DRC) is used to characterize whether the physical geometry of the intermediate wiring layout conforms to the wafer manufacturing design rules of the corresponding process node. It is a verification indicator to ensure that the layout can be successfully fabricated. The design rule check module in the application-specific integrated circuit back-end design tool can be called to import the process design suite rule file of the corresponding process node and check all traces, vias, and pin structures of the target metal layer in the intermediate wiring layout item by item. The check items include process constraints such as minimum line width, minimum line spacing, via size, and interlayer overlap area. Finally, the judgment result of pass or fail is output, and the location and type of all violations are marked. For example, when performing a design rule check on the intermediate wiring layout of a certain 12-nanometer process node, if the line width and line spacing of all metal traces meet the process requirements and there are no via size violations or interlayer alignment errors, the first verification result is pass. If two line spacings are found to be less than the minimum allowable value and one via size is found to be non-compliant, the first verification result is fail, and the specific coordinate information of the three violations is output simultaneously.
[0088] The second verification result is the execution of layout and schematic on the intermediate routing layout. Figure 1 The logic consistency judgment result output after Layout Versus Schematic (LVS) is used to characterize whether the metal connection relationship of the intermediate routing layout completely matches the circuit schematic corresponding to the weight matrix of the large model. It is a verification indicator to ensure the correctness of the layout logic function. The consistency verification module in the application-specific integrated circuit back-end design tool can be called to first extract the physical connection netlist from the intermediate routing layout, and then compare it item by item with the schematic source netlist generated by the weight matrix of the large model to check for logical errors such as open circuits, short circuits, and connection misalignments. Finally, the judgment result of pass or fail is output, and the location and type of all logical errors are marked. For example, when performing consistency verification on a weighted routing layout with 1-bit quantization, if all metal conducting connections in the layout correspond one-to-one with the conducting weights of the weight matrix, and there are no additional short circuits or missing open circuits, the second verification result is pass. If an open circuit is detected in one weighted connection that should be conducting, and an unexpected short circuit is detected in one adjacent trace, the second verification result is fail, and the corresponding locations of the two errors are output simultaneously.
[0089] The third verification result is the power supply compatibility judgment conclusion output after performing power distribution network compatibility verification on the intermediate routing layout. It is used to characterize whether there is a conflict between the routing structure and the chip's power distribution network, and whether the operating performance of the power supply system meets the design requirements. It is a verification indicator to ensure the stable and reliable power supply of the chip. The complete layout information of the chip's power distribution network can be imported to check whether the spacing between the metal traces and power supply structures such as power lines, power pads, and power strips in the intermediate routing layout meets the safety requirements. At the same time, it evaluates the overall performance of the power supply system after the introduction of the routing, and finally outputs a pass or fail judgment result, while marking the location and impact of all conflict items. For example, if a power distribution network compatibility verification is performed on an intermediate routing layout, and all signal traces maintain a compliant spacing with the power supply structure and do not have a negative impact on the performance of the power supply network, then the third verification result is pass. If two traces are found to have a spacing of less than the safety threshold with the power lines, and there is a risk of signal crosstalk and power supply interference, then the third verification result is fail, and the coordinate information of the conflict location is output simultaneously.
[0090] For example, the intermediate routing layout file to be verified can be loaded first, and the design rule file of the corresponding process node, the schematic source netlist file corresponding to the large model weight matrix, and the chip power distribution network layout information can be imported simultaneously to complete the parameter configuration of the verification environment. Then, three verification processes can be started in sequence. First, the design rule check can be performed, traversing all metal structures and via structures in the target metal layer, verifying all process constraints, and outputting the first verification result and the corresponding violation details. Then, the layout and schematic check can be performed. Figure 1 For consistency verification, the physical connection netlist of the layout is extracted and compared node by node with the source netlist of the schematic diagram. The second verification result and the corresponding logic error details are output. Finally, the power distribution network compatibility verification is performed to check the safe distance between signal traces and various power supply structures and evaluate the overall power supply performance. The third verification result and the corresponding conflict details are output. After all three verifications are completed, the results of the three verifications are summarized for unified judgment. If all three results are passed, the current intermediate routing layout is directly determined as the target routing layout. If there are any failed verification items, the routing paths or bit counting cell layout positions of the relevant areas are adjusted according to the type and specific location of the violation. After the violation is fixed, all three verifications are re-executed. After multiple rounds of iterative optimization, the target routing layout that meets all design requirements is finally output.
[0091] Through the implementation of the above embodiments, design rule checks and layout and schematic comparisons are performed on the intermediate routing layout. Figure 1The system performs consistency verification and power distribution network compatibility verification, and adjusts the routing path and / or bit counting cell layout when verification fails, so that the final target routing layout meets the requirements of manufacturing rules, connection correctness and power supply reliability. This can reduce the chip production risks caused by routing errors, layout inconsistencies or power supply conflicts, and improve the reliability of the chip routing layout determination process.
[0092] In some embodiments, the aforementioned power distribution network compatibility verification of the intermediate routing layout to obtain a third verification result may include: determining the network voltage drop and network current density of the power distribution network based on the power distribution network layout information in the intermediate routing layout; and determining that the third verification result is passed when the network voltage drop is less than a preset voltage drop threshold and the network current density is less than a preset current density threshold.
[0093] In some examples, network voltage drop and network current density are two quantitative indicators for measuring the power supply performance of a power distribution network. Network voltage drop refers to the voltage drop caused by the parasitic resistance of the path when the supply current flows through the metal traces, vias, and other conductive structures of the power distribution network. Its value directly determines the actual supply voltage level of the internal functional circuits of the chip and is a parameter for evaluating power supply stability. Network current density refers to the magnitude of the current passing through a unit cross-sectional area of the power supply metal trace. It characterizes the current load of the power supply metal structure and is a parameter for predicting the risk of electromigration failure. These two parameters... The system can import the complete power distribution network structure of the intermediate routing layout, the parasitic parameters of the target metal layer routing, and the operating power consumption model of the bit counting array. Steady-state power supply simulation calculations can be performed using a dedicated power integrity simulation tool to obtain the maximum network voltage drop and the network current density distribution of each power supply segment in the global range. For example, for a 12-nanometer process large-scale accelerated chip with a core nominal voltage of 0.8 volts, after importing the intermediate routing layout and corresponding power consumption data, simulation can be performed to obtain a maximum network voltage drop of 32 millivolts in the bit counting array region and a maximum network current density of 4.2 megaamperes per square meter in the core power supply band.
[0094] The preset voltage drop threshold is the maximum voltage drop value allowed by the power distribution network in advance. It is a hard constraint indicator to ensure the normal operation of the internal circuits of the chip. Its value setting is directly related to the timing margin and functional stability of the chip. The preset voltage drop threshold can be determined by combining the chip's nominal operating voltage, process node characteristics and the timing redundancy of the internal circuit. General integrated circuit design specifications usually control this threshold within 5% of the nominal supply voltage. For example, for a large model acceleration chip with a nominal core voltage of 0.8 volts, the preset voltage drop threshold can be set to 40 millivolts, corresponding to 5% of the nominal voltage. If the actual network voltage drop exceeds this threshold, it will lead to insufficient power supply to the circuit, causing timing violations or even functional failure.
[0095] The preset current density threshold is the maximum allowable current density value for power supply metal traces. It is used to avoid electromigration failure under long-term high current load on metal traces and ensure the long-term operating life of the chip. The preset current density threshold can be extracted from the process design kit (PDK) of the corresponding process node. Different metal layers and different trace widths correspond to different threshold parameters, and the wafer manufacturer provides the official limit value based on the process reliability test results. For example, the preset current density threshold of the top wide power supply metal layer of a certain 12-nanometer process node is 6 megaamperes per square meter. If the actual current density exceeds this threshold for a long time, it will cause directional migration of metal atoms, resulting in trace open circuit or performance degradation, and shortening the chip's lifespan.
[0096] The simulation results in a global maximum network voltage drop, which is lower than the preset voltage drop threshold. Simultaneously, the simulation results in a maximum network current density, which is lower than the preset current density threshold, are compared. The third verification result is considered passed only if both indicators are below their respective thresholds. If either indicator exceeds its threshold, the third verification result is considered failed, and the specific location and magnitude of the exceeding indicator are output. For example, after simulating a power distribution network for a certain intermediate cabling layout, if the maximum network voltage drop is 32 millivolts, which is lower than the preset threshold of 40 millivolts, and the maximum network current density is 4.2 megaamperes per square meter, which is lower than the preset threshold of 6 megaamperes per square meter, the third verification result is considered passed. If the simulation results in a maximum network voltage drop of 45 millivolts, which exceeds the preset threshold, the third verification result is directly considered failed.
[0097] For example, firstly, the complete power distribution network layout information corresponding to the intermediate routing layout can be imported, and the geometric parameters and material properties of power lines, power pads, and power strips can be extracted. Simultaneously, the power consumption model of the bit counting array and the parasitic parameter library of the target metal layer can be imported to complete the parameter configuration of the power integrity simulation environment. Then, a steady-state power supply simulation calculation can be started, and the voltage and current distribution of the entire power distribution network can be solved using simulation tools. The maximum network voltage drop value within the bit counting array area, as well as the peak network current density of each power strip and power line segment, can be extracted. Finally, the maximum network voltage drop obtained from the simulation can be compared with a pre-set voltage drop threshold. The maximum network current density is compared with a pre-set current density threshold. If both values are less than the corresponding threshold, the power distribution network compatibility verification is deemed successful, and the third verification result is output as a pass. If any indicator exceeds the corresponding threshold, the specific location of the out-of-standard area is located, and the signal wiring path in the corresponding area is adjusted to reduce interference with the power supply structure, or the width of the local power supply band is optimized to reduce the current density. After the adjustment is completed, the power supply simulation and indicator comparison are re-executed. After multiple rounds of iterative optimization until both indicators meet the threshold requirements, the third verification result as a pass is finally output.
[0098] By implementing the above embodiments, based on the power distribution network information in the intermediate wiring layout, the network voltage drop and network current density are detected, and the verification is determined to be passed when the preset voltage drop threshold and current density threshold are met. This enables the chip wiring layout to take power supply stability requirements into account during the generation process, which can avoid the power supply performance degradation caused by the mutual influence between weight-related wiring and power distribution network, and improve the stability of the chip during operation.
[0099] In some embodiments, the aforementioned mapping of virtual access points to at least one corresponding hardware region based on the weight matrix to obtain an access mapping table may include: traversing the weight matrix, determining the target weight value between the target input dimension and each output dimension for each target input dimension; if the target weight value is a non-disconnect value, establishing a mapping relationship between the virtual access point corresponding to the target input dimension and the hardware region corresponding to the target weight value; and generating an access mapping table based on all mapping relationships.
[0100] In some examples, the target input dimension is a single input dimension currently being processed during the weight matrix traversal. It corresponds to a column of data in the weight matrix, representing an independent model input activation signal, and is the basic processing unit for establishing the mapping relationship. The target input dimension can be traversed column by column in the weight matrix according to a preset column index order. The currently selected column for processing is the target input dimension, identified by a unique column index number. For example, for a weight matrix with a total of 4096 columns, the traversal is performed sequentially from column 0 to column 4095. When the column with index 128 is reached, that column is the current target input dimension, corresponding to the 128th model input activation signal.
[0101] The target weight values corresponding to the target input dimension and each output dimension are the weight values at the intersection positions of the target input dimension and each output dimension row in the corresponding column. They represent the connection strength and connection status between the input activation signal and each output, and are the basis for determining whether physical wiring is needed. You can locate the matrix column corresponding to the target input dimension and read the weight elements of each row in the column in sequence. Each element is the target weight value under the corresponding output dimension.
[0102] Non-disconnect values are the set of weight values in the quantization weights that represent a conductive metal connection requiring the corresponding physical metal wire to be laid, corresponding to the disconnect values that represent a disconnected metal connection requiring no metal wire to be laid. They are the criteria for selecting valid wiring requirements. Non-disconnect values can be predefined according to preset weight quantization rules. In single-bit quantization scenarios, the weight value representing the conductive state is the non-disconnect value. In multi-bit quantization scenarios, all non-zero quantization weight values belong to the set of non-disconnect values. For example, in a 1-bit quantization weight encoding scheme, a weight with a value of 1 represents a conductive metal wire and belongs to the non-disconnect value; a weight with a value of 0 represents a disconnected metal wire, does not belong to the non-disconnect value, and does not require the corresponding physical wiring to be laid. When the target weight value is determined to be a non-disconnect value, the system first matches the hardware area identifier corresponding to the target weight value, then associates and binds the virtual access point identifier corresponding to the target input dimension with the hardware area identifier to form a valid mapping record. Finally, all valid mapping records are summarized and organized to form a structured data table, outputting mapping data in a unified format as input for subsequent cabling processes. For example, if the virtual access point number corresponding to the target input dimension is 128, the current target weight value is 1, and the corresponding hardware area is A, then a mapping relationship is established between virtual access point 128 and hardware area A, recording that the input signal needs to be cabled to hardware area A. Alternatively, after completing a full traversal and mapping determination of 4096 input dimensions, all valid mapping entries are summarized, sorted by virtual access point number from smallest to largest, and an access mapping relationship table containing virtual access point identifiers, corresponding hardware area identifiers, and bound physical pin coordinates is generated, which can be directly imported into the cabling engine.
[0103] For example, the first step is to load the quantized weight matrix data, confirm the one-to-one correspondence between each quantized weight value and the hardware region, and extract the binding information between all virtual access points and input dimensions. The weight matrix is then traversed column by column in ascending order of column index. Each column is treated as the current target input dimension, and the target weight values corresponding to all rows within that column are extracted. Each target weight value is then checked to determine if it is a non-disconnect value. If the target weight value is non-disconnect, the corresponding hardware region is matched, and a mapping relationship is established between the virtual access point and the hardware region corresponding to the current target input dimension, generating a mapping record. If the target weight value is disconnect, no mapping relationship is generated, and the weight item is skipped to avoid invalid routing resource usage. If the same target input dimension corresponds to multiple different non-disconnect weight values, a mapping relationship is established between the virtual access point and each corresponding hardware region. After completing the traversal and mapping determination of all input dimensions, all valid mapping records are summarized, organized in order according to the virtual access point number, and the physical pin coordinate information bound to each virtual access point is supplemented. Finally, a complete and standardized access mapping relationship table is generated, providing accurate endpoint mapping basis for subsequent automatic routing within the target metal layer.
[0104] By implementing the above embodiments, the weight matrix is traversed to determine the weight value corresponding to each input dimension, and a mapping relationship between the virtual access point and the corresponding hardware area is established only for non-disconnect values. This enables the access mapping relationship table to accurately reflect the actual effective weight connection relationship, avoids additional routing processing for invalid connections, reduces unnecessary routing resource occupation, and makes the metal connection structure in the chip layout more consistent with the actual weight data characteristics, thereby improving routing efficiency.
[0105] In some embodiments, before step 101, the aforementioned chip wiring layout determination method may further include: performing quantization encoding processing on the weight matrix based on a preset number of quantization bits to obtain a discretized set of quantization weight values, wherein the preset number of quantization bits is 1 bit, 2 bits, 4 bits or 8 bits; and determining the probability of occurrence of each weight value in the weight matrix based on the frequency of occurrence of each quantization weight value in the set of quantization weight values.
[0106] In some examples, the preset quantization bit count is a pre-defined bit width parameter for the weight values, which determines the total number of values and the precision level after weight discretization. It can be flexibly selected according to the precision requirements of large models and hardware resource overhead. The available levels include 1 bit, 2 bits, 4 bits, and 8 bits. The preset quantization bit count can be determined comprehensively based on the precision requirements of model deployment and chip hardware resource budget, and is passed to the quantization processing module through a configuration file. Different bit counts correspond to different weight densities and calculation precisions. For example, 1-bit quantization can be selected for inference scenarios with extreme weight density, trading the lowest precision for the highest hardware deployment density; 4-bit quantization can be selected for general inference scenarios that balance accuracy and energy efficiency, achieving a high weight deployment density while ensuring model inference performance.
[0107] Quantization encoding is the process of mapping continuous floating-point weight values in the original weight matrix to discrete integer values with a finite number of quantization bits, adapting software weights to the physical characteristics of metal continuity. Symmetric or asymmetric quantization strategies can be used. First, the quantization threshold and quantization step size of the weight values are determined. Then, each floating-point weight is mapped to the nearest discrete quantization level, completing the numerical conversion of the entire matrix. For example, for a 1-bit quantization scenario, the quantization encoding process divides the original floating-point weights into two categories based on the zero-value threshold: positive weights are mapped to the value 1 representing metal continuity, and negative weights and zero values are mapped to the value 0 representing metal disconnection, ultimately resulting in a weight matrix containing only these two types of discrete values. For a 4-bit quantization scenario, the quantization encoding process maps the original floating-point weights to integers in the range of 0 to 15, with each bit corresponding to an independent metal wiring and bit counting unit. Finally, the complete dot product result is obtained by accumulating the weights by bit.
[0108] The discretized quantization weight value set is a numerical set consisting of all non-repeating discrete weight values in the weight matrix after quantization encoding. The total number of elements in the set is 2 raised to the power of the preset quantization bit number, and each value uniquely corresponds to a type of hardware region. The discretized quantization weight value set can be obtained by extracting all unique weight values from the quantized weight matrix, removing duplicates, and sorting them in ascending order. For example, when the preset quantization bit number is 1 bit, the discretized quantization weight value set contains two elements, corresponding to the off state and the on state, respectively. When the preset quantization bit number is 4 bits, the discretized quantization weight value set contains sixteen elements, with values ranging from 0 to 15, and each value corresponds to an independent weight level and hardware region.
[0109] The quantized weight matrix can be iterated through, and the total number of occurrences of each quantized weight value can be counted one by one. Then, the occurrence frequency of a single type of weight value is divided by the total number of elements in the weight matrix to obtain the occurrence probability of the corresponding weight value. The sum of the occurrence probabilities of all weight values is 1. For example, a 1-bit quantized weight matrix contains a total of 16 million elements. The on-state weight with a value of 1 occurs 9.6 million times, and the off-state weight with a value of 0 occurs 6.4 million times. After dividing by the total number of elements, the occurrence probability of the on-state weight is 0.6, and the occurrence probability of the off-state weight is 0.4.
[0110] For example, the process begins by reading the quantization configuration file to obtain the pre-set quantization bit count parameter, confirming the corresponding number of quantization levels and quantization algorithm parameters. Then, the original weight matrix of the target large model is loaded, the matrix's dimensionality information and all floating-point weight values are extracted, the quantization threshold and quantization step size are determined according to the preset quantization algorithm, and point-by-point quantization encoding is performed on all weight elements, converting continuous floating-point weights into discrete integer values of corresponding bit widths, resulting in a complete quantization weight matrix and a set of discretized quantization weight values. After quantization encoding is completed, the entire quantization weight matrix is traversed, and the total frequency of each quantization weight value within the set is counted sequentially, recording the statistical results corresponding to each weight value. The frequency of each quantization weight value is divided by the total number of elements in the weight matrix to calculate the probability of occurrence of each weight value, which is then organized into a weight probability distribution table. All probability data are bound one-to-one with the corresponding weight value, providing accurate data support for probability-based hardware region partitioning in subsequent steps.
[0111] Through the implementation of the above embodiments, before dividing the hardware region according to the weight probability, the target large model weight matrix can first be quantized and encoded with a preset number of bits, and the weight probability can be determined based on the frequency of occurrence of the quantized weight values, so that the continuous weight data is converted into a discrete data form suitable for chip physical mapping; by supporting different quantization bit number configurations, the weight encoding precision can be adjusted according to the chip design requirements, so that the subsequent hardware region division can be based on a clear weight category, thereby improving the accuracy of weight distribution analysis and layout resource allocation.
[0112] In some embodiments, the aforementioned step 101 may include: for each weight value, determining the target area of the hardware region corresponding to the weight value based on the probability of occurrence of the weight value and the total physical area of the bit counting array used for weight encoding in the chip layout; and dividing the chip layout into regions according to each target area to obtain multiple hardware regions that correspond one-to-one with each weight value.
[0113] In some examples, the total physical area of the bit counting array used for weight encoding is the total area of the entire physical region within the chip layout specifically reserved for deploying weight encoding metal wiring and bit counting calculation units. It is the total spatial boundary and area benchmark of all hardware regions, and its range is predetermined by the overall chip layout plan. It does not include the area occupied by other functional modules such as power distribution networks, clock trees, and input / output interfaces. The coordinates of the four sides of the bit counting array region can be extracted from the layout plan file of the application-specific integrated circuit back-end design, and the rectangular area value of the region can be calculated by the coordinate difference. For example, in a large model acceleration chip of a certain 12-nanometer process node, the total horizontal width of the bit counting array is 2000 micrometers and the total vertical height is 1000 micrometers. Then the total physical area of the bit counting array used for weight encoding is 2,000,000 square micrometers.
[0114] The target area is the physical area that should be allocated to the hardware region corresponding to a single weight value. It is the direct dimensional basis for dividing the physical region, and its value is directly related to the probability of occurrence of the corresponding weight value. High-frequency weights correspond to larger target area areas, while low-frequency weights correspond to smaller target area areas. The target area value corresponding to a single weight value can be calculated by multiplying the probability of occurrence of a single weight value by the total physical area of the bit counting array. For example, if the total physical area of the bit counting array is 2,000,000 square micrometers and the probability of occurrence of a certain weight value is 0.3, then the target area corresponding to that weight value is 600,000 square micrometers. When using a 4-bit quantization scheme, there are 16 categories of quantization weight values, and the probability of occurrence of each category of weight value is different. By multiplying the probability of occurrence of each category of weight by the total physical area of the bit counting array, 16 different target area values can be obtained. The area value corresponding to a high-frequency weight is greater than the area value corresponding to a low-frequency weight.
[0115] The process of dividing the chip layout according to the area of each target region to obtain multiple hardware regions corresponding to each weight value involves arranging the hardware regions corresponding to each weight value sequentially within the total area of the bit counter array according to a preset arrangement rule. Isolation spacing that conforms to process design rules is reserved between adjacent regions. The length and width dimensions of each region are adjusted to match the corresponding target region area. Finally, the four-sided coordinate boundaries of each hardware region are determined, forming a set of one-to-one corresponding hardware regions. For example, using a horizontal sequential arrangement rule, the hardware regions corresponding to 16 weight values are arranged sequentially along the horizontal direction of the bit counter array, with all regions maintaining a consistent vertical height. The horizontal width of each region is adjusted to match the corresponding target region area, and a 1-micron isolation spacing is reserved between adjacent regions. This results in 16 hardware regions with clearly defined coordinate ranges, each corresponding to a weight value.
[0116] For example, firstly, the coordinates of the four sides of the bit counter array can be extracted from the chip layout planning file to calculate the total physical area of the bit counter array used for weight encoding. It is confirmed that no other fixed functional modules occupy this area, and it can be entirely used for weight region division. Then, the various weight values and their corresponding probability data obtained from quantization statistics are loaded. Each type of weight value is traversed in ascending order of its numerical value, and the probability of occurrence of a single type of weight is multiplied by the total physical area of the bit counter array to calculate the target area corresponding to each type of weight value. The difference between the sum of all target area areas and the total physical area of the bit counter array is verified to be within the allowable error range of the process. After completing the area calculation, a horizontally uniform height arrangement is adopted. The bit counting array is divided into regions, ensuring that the vertical height of all hardware regions is consistent with the total height of the bit counting array. The corresponding horizontal width is calculated in reverse based on the target area of each region. The hardware regions are arranged from left to right according to the numerical order of their weight values, with isolation spacing between adjacent regions that meets process requirements to avoid interference between the wiring of different weight regions. After all regions are arranged, the weight value identifier, four-sided coordinate range, and actual area parameters of each hardware region are recorded to form complete weight region division layout data. This ensures that each hardware region uniquely corresponds to a type of weight value, and that the region area is positively correlated with the probability of weight occurrence, providing a precise physical space basis for subsequent input mapping and wiring processing.
[0117] Through the implementation of the above embodiments, the target area of the hardware region corresponding to each weight value is calculated based on the occurrence probability of each weight value and the total physical area of the bit counting array used for weight encoding. The layout area is then divided according to the target area, so that the size of the hardware region can quantitatively match the weight distribution characteristics. Compared with the method of uniformly dividing the layout area, the fineness of the layout area allocation can be further improved, so that the limited chip area is given priority to the hardware resources corresponding to high-frequency weights, thereby improving the overall hardware utilization efficiency.
[0118] In some embodiments, before performing routing processing within the target metal layer of the chip layout based on the aforementioned access mapping table, the chip routing layout determination method may further include: obtaining a routing hierarchy configuration file imported from a user interface; parsing the routing hierarchy configuration file to obtain protection rule constraints for the target metal layer and the underlying fixed layer for weight encoding, wherein the underlying fixed layer may include a metal layer for power connections, clock connections, and basic device connections; and constructing inter-layer isolation logic between the target metal layer and the underlying fixed layer based on the protection rule constraints to restrict routing processing operations from being performed within the target metal layer.
[0119] In some examples, the user interface is a human-computer interaction channel provided by the routing design tool for designers to configure parameters and import files. It is used to receive routing hierarchy constraints and configuration information input by the user and serves as the input carrier for implementing configurable routing hierarchies. The user interface can be integrated into the graphical interface of the application-specific integrated circuit (ASIC) back-end design tool, or it can provide a command-line script interface for users to call in batches, supporting the uploading of configuration files, parameter modification, and rule validation. For example, designers can upload pre-written routing hierarchy configuration files through the graphical user interface. The interface has a built-in file format validation function that can automatically detect the completeness and legality of the configuration content, ensuring that the configuration parameters can be correctly recognized by the routing engine.
[0120] Routing hierarchy configuration files are structured configuration files pre-written by designers to clearly define the target metal layer number corresponding to the weight encoding, the range of the bottom fixed layer, and the corresponding protection rules. They serve as the data source for routing hierarchy constraints. Routing hierarchy configuration files can be written according to the overall metal layer planning and functional division of the chip and imported into the routing system through a user interface. The file content can include fields such as a list of metal layer numbers, layer functional attributes, and operation permission rules. For example, a routing hierarchy configuration file may explicitly designate the sixth metal layer as the dedicated layer for weight encoding, and the first to fifth metal layers as the bottom fixed functional layers. It may also stipulate that the original traces and device connections within the fixed layers must not be modified or occupied by weighted routing operations.
[0121] The target metal layer used for weight encoding is a metal layer specifically designated by the routing hierarchy configuration file to carry the metal routing for weight encoding of large models. It is the physical medium carrying the weight connection relationship, and all weight-related signal routing is limited to this layer. The routing hierarchy configuration file can be parsed to extract the metal layer number marked with the weight encoding function and set it as the only valid layer for routing operations. For example, the configuration file specifies the sixth metal layer at the top as the target metal layer for weight encoding. This layer has a large metal thickness and low parasitic resistance, which is suitable for long-distance transmission of high-frequency input activation signals and meets the high-performance computing requirements of large model inference.
[0122] The protection rules constraints of the bottom fixed layer are a set of protective constraint rules defined in the routing hierarchy configuration file for the bottom fixed functional layers. They are used to limit the restricted areas of routing operations and ensure that fixed lines such as power connections, clock connections, and basic device connections are not encroached upon or damaged by weighted routing. The routing hierarchy configuration file can be parsed to extract the range of metal layers marked as fixed functions and the corresponding prohibited operation types, boundary restricted areas, and other constraint parameters to obtain the protection rules constraints. For example, the protection rules constraints can clearly define the first to fifth metal layers as the bottom fixed layers, which are responsible for power distribution, clock tree distribution, and basic device interconnection, respectively. Weighted routing operations are prohibited from generating any traces in the above metal layers and from changing the connection relationships and geometric parameters of the original metal lines.
[0123] The process of constructing interlayer isolation logic between the target metal layer and the underlying fixed layer based on protection rule constraints is the core processing flow that transforms the rules in the configuration file into a hierarchical constraint mechanism that the routing engine can execute. By establishing hierarchical permission barriers, it ensures that all weighted routing operations are strictly limited to the target metal layer and will not intrude into the underlying fixed layer. For example, after the interlayer isolation logic is constructed, the routing engine will automatically filter the routerable resources of the first to fifth metal layers when performing path search. All candidate routing paths will only be generated within the sixth metal layer and will not extend downward to the underlying fixed layer, nor will they damage the existing power and clock lines in the fixed layer.
[0124] For example, the user interface of the cabling design tool can be activated first to receive the cabling hierarchy configuration file imported by the designer. The file format and parameter validity are initially verified to confirm the completeness and validity of the configuration content. Then, the configuration parsing module is called to read the entire contents of the cabling hierarchy configuration file, extracting the target metal layer number used for weight encoding. Simultaneously, the layer range and corresponding protection rule constraints of the bottom fixed layers are extracted, clarifying the functional attributes and prohibited operation types of each fixed layer. After parameter extraction, the cabling engine constructs inter-layer isolation logic based on the protection rule constraints, marking the bottom fixed layers as unprotected in the internal cabling resource space. The routing restriction shields all trace and via resources on the fixed layers and sets layer boundaries for routing operations, limiting path searching and trace generation to the target metal layer only. After the interlayer isolation logic is built, all weight-encoding related metal traces will be strictly restricted to the target metal layer during subsequent routing processes. This will not cause any changes to the power supply, clock, and basic device connections of the underlying fixed layers. When the model weights are updated, only the traces within the target metal layer need to be regenerated. The underlying fixed layers can be fully reused without redesign, effectively shortening the weight iteration cycle and reducing the cost of the fabrication mask.
[0125] By implementing the above embodiments, a user-configured routing layer configuration file is obtained, and based on the target metal layer information and the underlying fixed layer protection rules therein, interlayer isolation logic is established between the target metal layer and the underlying fixed layer. This ensures that weight-encoded routing is executed only within the specified target metal layer, while avoiding impact on the fixed metal layers used for power connections, clock connections, and basic device connections. This enables independent adjustment of weight-related routing while ensuring the stability of the chip's basic functional structure, improving the flexibility of chip routing layout updates, and reducing resource consumption caused by redesigning all metal layers.
[0126] Furthermore, as an implementation of the foregoing method embodiments, this application also provides a chip wiring layout determination apparatus for implementing the foregoing method embodiments. This apparatus embodiment corresponds to the foregoing method embodiments. For ease of reading, this chip wiring layout determination apparatus embodiment will not repeat the details of the foregoing method embodiments one by one, but it should be understood that the apparatus in this application embodiment can correspondingly implement all the contents of the foregoing method embodiments. For example... Figure 2As shown, the chip wiring layout determination device 20 includes: a region division unit 201, a mapping establishment unit 202, and a wiring processing unit 203. The region division unit 201 is used to divide the chip layout into multiple hardware regions corresponding to each weight value based on the probability of occurrence of each weight value in the weight matrix of the target large model. The area of each hardware region is positively correlated with the probability of occurrence of the corresponding weight value. The mapping establishment unit 202 is used to map each input dimension of the weight matrix to at least one corresponding hardware region based on the weight matrix and the physical pin information corresponding to the model input activation signal of the target large model, thereby obtaining an access mapping table. The wiring processing unit 203 is used to perform wiring processing within the target metal layer of the chip layout based on the access mapping table, thereby obtaining the target wiring layout of the chip layout.
[0127] In some embodiments, the mapping establishment unit 202 is further configured to abstract and generate multiple virtual access points based on physical pin information, wherein each virtual access point corresponds one-to-one with each input dimension of the weight matrix; based on the weight matrix, the virtual access points are mapped to at least one corresponding hardware region to obtain an access mapping relationship table.
[0128] In some embodiments, the routing processing unit 203 is further configured to, based on the comprehensive routing cost function, execute a routing process from each virtual access point to the corresponding hardware region in the access mapping table within the target metal layer of the chip layout to obtain an initial routing layout of the chip layout; based on the total number of target input lines in each hardware region of the initial routing layout and the upper limit of the input capability of a single bit counter unit, generate pin access points for each bit counter unit in the initial routing layout to obtain an intermediate routing layout, wherein the total number of target input lines is the statistical number of virtual access points mapped to the hardware region, and the upper limit of the input capability of a single bit counter unit is the maximum number of access channels supported by a single bit counter unit determined based on preset timing constraints and physical design rules; execute a preset verification optimization process on the intermediate routing layout, and determine the intermediate routing layout that passes the verification as the target routing layout.
[0129] In some embodiments, the routing processing unit 203 is further configured to acquire routing loss assessment data of the target metal layer, wherein the routing loss assessment data includes resistance value per unit length, capacitance value per unit area, and via parasitic parameter value; acquire power distribution network layout information of the chip layout, wherein the power distribution network layout information includes power line location coordinates, power pad coordinates, and power strip distribution data; and construct a comprehensive routing cost function based on the routing loss assessment data and the power distribution network layout information, wherein the comprehensive routing cost function is obtained by weighted summation of routing resistance cost term, routing capacitance cost term, design rule violation cost term, and power distribution network conflict cost term.
[0130] In some embodiments, the cabling processing unit 203 is further configured to take each virtual access point as the cabling starting point and the corresponding hardware area in the access mapping table as the cabling ending point, generate multiple candidate cabling paths in the target metal layer using a preset three-dimensional cabling algorithm; determine the cost value of each candidate cabling path based on the comprehensive cabling cost function; determine the candidate cabling path with the minimum cost value among the multiple candidate cabling paths as the target cabling path; and generate an initial cabling layout based on the target cabling path.
[0131] In some embodiments, the routing processing unit 203 is further configured to, for each hardware region, if the total number of target input lines is less than or equal to the upper limit of input capability, configure a single bit counting unit in the corresponding hardware region and generate multiple input pins equal to the total number of target input lines; if the total number of target input lines is greater than the upper limit of input capability, divide the corresponding hardware region into multiple hardware sub-regions based on the ratio of the total number of target input lines to the upper limit of input capability, and configure a bit counting unit and multiple corresponding input pins for each hardware sub-region; integrate the configured bit counting units and input pins into the initial routing layout to obtain an intermediate routing layout.
[0132] In some embodiments, the routing processing unit 203 is further configured to perform design rule checks on the intermediate routing layout to obtain a first verification result; and to perform layout and schematic analysis on the intermediate routing layout. Figure 1 The first verification is performed to obtain the second verification result; the second verification result is obtained by performing power distribution network compatibility verification on the intermediate routing layout; when the first verification result, the second verification result, and the third verification result are all passed, the intermediate routing layout is determined as the target routing layout; otherwise, the routing path and / or bit counting cell layout of the intermediate routing layout are adjusted until all verifications are passed.
[0133] In some embodiments, the routing processing unit 203 is further configured to determine the network voltage drop and network current density of the power distribution network based on the power distribution network layout information in the intermediate routing layout; when the network voltage drop is less than a preset voltage drop threshold and the network current density is less than a preset current density threshold, the third verification result is determined to be passed.
[0134] In some embodiments, the mapping establishment unit 202 is further configured to traverse the weight matrix, and for each target input dimension, determine the target weight value corresponding to the target input dimension and each output dimension; if the target weight value is a non-disconnect value, then establish a mapping relationship between the virtual access point corresponding to the target input dimension and the hardware area corresponding to the target weight value; and generate an access mapping relationship table based on all mapping relationships.
[0135] In some embodiments, the region partitioning unit 201 is further configured to perform quantization encoding processing on the weight matrix based on a preset number of quantization bits to obtain a discrete set of quantized weight values, wherein the preset number of quantization bits is 1 bit, 2 bits, 4 bits or 8 bits; and determine the probability of occurrence of each weight value in the weight matrix based on the frequency of occurrence of each quantized weight value in the set of quantized weight values.
[0136] In some embodiments, the region division unit 201 is further configured to, for each weight value, determine the target area of the hardware region corresponding to the weight value based on the probability of occurrence of the weight value and the total physical area of the bit counting array used for weight encoding in the chip layout; and divide the chip layout according to each target area to obtain multiple hardware regions that correspond one-to-one with each weight value.
[0137] In some embodiments, the chip routing layout determination device 20 further includes a constraint unit for obtaining a routing hierarchy configuration file imported from a user interface; parsing the routing hierarchy configuration file to obtain protection rule constraints for the target metal layer and the bottom fixed layer for weight encoding, wherein the bottom fixed layer includes a metal layer for power connection, clock connection and basic device connection; and constructing inter-layer isolation logic between the target metal layer and the bottom fixed layer based on the protection rule constraints to restrict routing processing operations to be performed within the target metal layer.
[0138] This application also provides a computer-readable storage medium storing computer-executable instructions or a computer program, which, when executed by a processor, will cause the processor to perform any step of the chip wiring layout determination method provided in this application.
[0139] In some embodiments, the computer-readable storage medium may be a random access memory (RAM), a read-only memory (ROM), flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); or it may be a variety of devices that include one or any combination of the above-mentioned memories.
[0140] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0141] In some embodiments, computer-executable instructions may, but do not necessarily, correspond to files in a file system, and may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0142] In some embodiments, computer-executable instructions may be deployed to execute on an electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0143] like Figure 3 As shown, this application also provides an electronic device 30, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, it implements any step of the chip wiring layout determination method described above.
[0144] This application also provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium and executes the computer program or computer-executable instructions, causing the electronic device to perform any step of the chip wiring layout determination method described above.
[0145] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for determining chip wiring layout, characterized in that, include: Based on the probability of occurrence of each weight value in the weight matrix of the target large model, multiple hardware regions corresponding one-to-one with each weight value are divided in the chip layout. The area of each hardware region is positively correlated with the probability of occurrence of the corresponding weight value. Based on the weight matrix and the physical pin information corresponding to the model input activation signal of the target large model, each input dimension of the weight matrix is mapped to at least one corresponding hardware region to obtain an access mapping table. Based on the access mapping table, routing is performed within the target metal layer of the chip layout to obtain the target routing layout of the chip layout.
2. The chip wiring layout determination method according to claim 1, characterized in that, The physical pin information corresponding to the model input activation signals of the target large model, based on the weight matrix and the model input activation signals of the target large model, maps each input dimension of the weight matrix to at least one corresponding hardware region to obtain an access mapping table, including: Based on the physical pin information, multiple virtual access points are abstracted and generated, wherein each virtual access point corresponds one-to-one with each input dimension of the weight matrix; Based on the weight matrix, the virtual access points are mapped to at least one corresponding hardware region to obtain the access mapping table.
3. The chip wiring layout determination method according to claim 2, characterized in that, The step of performing routing processing within the target metal layer of the chip layout based on the access mapping table to obtain the target routing layout of the chip layout includes: Based on the comprehensive routing cost function, within the target metal layer of the chip layout, a routing process is executed from each of the virtual access points to the corresponding hardware regions in the access mapping table to obtain the initial routing layout of the chip layout. Based on the total number of target input lines in each hardware area of the initial routing layout and the upper limit of the input capability of a single bit counter unit, pin access points are generated for each bit counter unit in the initial routing layout to obtain an intermediate routing layout. The total number of target input lines is the statistical number of virtual access points mapped to the hardware area, and the upper limit of the input capability of a single bit counter unit is the maximum number of access channels supported by a single bit counter unit, determined based on preset timing constraints and physical design rules. A preset verification and optimization process is performed on the intermediate routing layout, and the intermediate routing layout that passes the verification is determined as the target routing layout.
4. The chip wiring layout determination method according to claim 3, characterized in that, Before determining the initial routing layout of the chip layout by performing a routing process from each virtual access point to the corresponding hardware region in the access mapping table within the target metal layer of the chip layout based on the comprehensive routing cost function, the chip routing layout determination method further includes: Obtain wiring loss assessment data for the target metal layer, wherein the wiring loss assessment data includes resistance per unit length, capacitance per unit area, and via parasitic parameter values; Obtain the power distribution network layout information of the chip layout, wherein the power distribution network layout information includes power line position coordinates, power pad coordinates, and power strip distribution data; Based on the cabling loss assessment data and the power distribution network layout information, the comprehensive cabling cost function is constructed, wherein the comprehensive cabling cost function is obtained by weighted summation of cabling resistance cost, cabling capacitance cost, design rule violation cost, and power distribution network conflict cost.
5. The chip wiring layout determination method according to claim 3, characterized in that, Based on the comprehensive routing cost function, within the target metal layer of the chip layout, a routing process is executed from each of the virtual access points to the corresponding hardware regions in the access mapping table to obtain the initial routing layout of the chip layout, including: Using each virtual access point as the starting point of the cabling and the corresponding hardware area in the access mapping table as the ending point of the cabling, multiple candidate cabling paths are generated in the target metal layer by a preset three-dimensional cabling algorithm. Based on the comprehensive cabling cost function, the cost of each candidate cabling path is determined; The candidate routing path with the lowest cost among the multiple candidate routing paths is determined as the target routing path; The initial routing layout is generated based on the target routing path.
6. The chip wiring layout determination method according to claim 3, characterized in that, The intermediate routing layout is obtained by generating pin access points for each bit counting unit in the initial routing layout based on the total number of target input lines in each hardware area of the initial routing layout and the upper limit of the input capability of a single bit counting unit, including: For each hardware region, if the total number of target input lines is less than or equal to the upper limit of input capability, a single bit counting unit is configured in the corresponding hardware region, and multiple input pins are generated in a number equal to the total number of target input lines. If the total number of target input lines is greater than the upper limit of input capability, then based on the ratio of the total number of target input lines to the upper limit of input capability, the corresponding hardware area is divided into multiple hardware sub-regions, and each hardware sub-region is configured with a bit counting unit and multiple corresponding input pins. The configured bit counting units and input pins are integrated into the initial routing layout to obtain the intermediate routing layout.
7. The chip wiring layout determination method according to claim 3, characterized in that, The step of performing a preset verification and optimization process on the intermediate routing layout, and determining the verified intermediate routing layout as the target routing layout, includes: A design rule check is performed on the intermediate routing layout to obtain the first verification result; Perform a layout and schematic consistency verification on the intermediate wiring layout to obtain a second verification result; A power distribution network compatibility verification was performed on the intermediate wiring layout to obtain a third verification result. When the first verification result, the second verification result, and the third verification result are all passed, the intermediate routing layout is determined as the target routing layout; otherwise, the routing path and / or bit counting cell layout of the intermediate routing layout are adjusted until all verifications are passed.
8. The chip wiring layout determination method according to claim 7, characterized in that, The power distribution network compatibility verification performed on the intermediate cabling layout yields a third verification result, including: Based on the power distribution network layout information in the intermediate wiring layout, the network voltage drop and network current density of the power distribution network are determined. When the network voltage drop is less than a preset voltage drop threshold and the network current density is less than a preset current density threshold, the third verification result is determined to be passed.
9. The chip wiring layout determination method according to claim 2, characterized in that, The step of mapping the virtual access points to at least one corresponding hardware region based on the weight matrix to obtain the access mapping table includes: Traverse the weight matrix and, for each target input dimension, determine the corresponding target weight value between the target input dimension and each output dimension; If the target weight value is a non-disconnect value, then a mapping relationship is established between the virtual access point corresponding to the target input dimension and the hardware area corresponding to the target weight value; Based on all the mapping relationships, the access mapping relationship table is generated.
10. The chip wiring layout determination method according to claim 1, characterized in that, Before determining the chip routing layout by considering the probability of occurrence of each weight value in the weight matrix based on the target large model and dividing the chip layout into multiple hardware regions corresponding one-to-one with each weight value, the chip routing layout determination method further includes: The weight matrix is quantized and encoded based on a preset number of quantization bits to obtain a set of discrete quantized weight values, wherein the preset number of quantization bits is 1 bit, 2 bits, 4 bits or 8 bits. Based on the frequency of occurrence of each quantized weight value in the set of quantized weight values, the probability of occurrence of each weight value in the weight matrix is determined.
11. The chip wiring layout determination method according to claim 1, characterized in that, The probability of occurrence of each weight value in the weight matrix based on the target large model is used to divide the chip layout into multiple hardware regions corresponding one-to-one with each of the weight values, including: For each weight value, the target area of the hardware region corresponding to the weight value is determined based on the probability of occurrence of the weight value and the total physical area of the bit counting array used for weight encoding in the chip layout. Based on the area of each target region, the chip layout is divided into regions to obtain multiple hardware regions that correspond one-to-one with each weight value.
12. The chip wiring layout determination method according to claim 1, characterized in that, Before performing routing processing within the target metal layer of the chip layout based on the access mapping table, the chip routing layout determination method further includes: Retrieve the wiring hierarchy configuration file imported from the user interface; Parse the routing hierarchy configuration file to obtain the protection rule constraints for the target metal layer and the bottom fixed layer used for weight encoding, wherein the bottom fixed layer includes a metal layer for power connection, clock connection and basic device connection; Based on the protection rule constraints, interlayer isolation logic is constructed between the target metal layer and the underlying fixed layer to restrict wiring operations from being performed within the target metal layer.
13. A chip wiring layout determination device, characterized in that, include: The region partitioning unit is used to divide the chip layout into multiple hardware regions corresponding to each weight value based on the probability of occurrence of each weight value in the weight matrix of the target large model. The area of each hardware region is positively correlated with the probability of occurrence of the corresponding weight value. The mapping establishment unit is used to map each input dimension of the weight matrix to at least one corresponding hardware region based on the weight matrix and the physical pin information corresponding to the model input activation signal of the target large model, so as to obtain an access mapping relationship table. The routing processing unit is used to perform routing processing in the target metal layer of the chip layout based on the access mapping table to obtain the target routing layout of the chip layout.
14. A chip, characterized in that, The chip wiring layout is determined by the chip wiring layout determination method according to any one of claims 1 to 12.
15. An electronic device comprising: A memory and a processor, characterized in that the processor, when executing a computer program stored in the memory, implements the steps of the chip wiring layout determination method as described in any one of claims 1 to 12.
16. A computer-readable storage medium having stored thereon computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or the computer program are executed by a processor, the steps of the chip wiring layout determination method as described in any one of claims 1 to 12 are implemented.