Arithmetic unit chip configuration method, computing subsystem, and intelligent computing platform

By dynamically evaluating the optimization model and optimizing chip parameters through cross-mutation operations, the problem of balancing low energy consumption and high performance in traditional methods is solved, and efficient optimization of chip design and energy consumption reduction are achieved.

WO2025189504A1PCT designated stage Publication Date: 2025-09-18GUANGDONG QINZHI SCIENCE & TECHNOLOGY RESEARCH INSTITUTE

Patent Information

Application Number
PCT/CN2024/083731
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-15
Filing Date
2024-03-26
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Traditional high-energy consumption equipment cannot meet environmental protection needs and it is difficult to achieve a balance between low energy consumption and high performance. Manual adjustment methods cannot fully consider the mutual influence of multiple factors, resulting in increased chip design complexity and power consumption.

Method used

A dynamic evaluation optimization model is used to generate the initial chip population, and chip parameters are optimized through cross-mutation operations and iterative cycles to ensure a balance between performance and power consumption, avoid local optimal solutions, and achieve global optimal solutions.

Benefits of technology

It improves chip optimization efficiency, reduces power consumption, ensures high performance while reducing design complexity and energy consumption, and achieves a balance between low energy consumption and high performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024083731_18092025_PF_FP_ABST
    Figure CN2024083731_18092025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the field of data processing, and in particular to an arithmetic unit chip configuration method, a computing subsystem, and an intelligent computing platform. The method comprises: in response to a configuration instruction for a target chip, generating an initial chip population corresponding to the target chip; performing strategy optimization on the initial chip population by using a dynamic evaluation and optimization model, so as to obtain an optimal initial individual in the initial chip population; performing crossover and mutation operations on the optimal initial individual so as to obtain an offspring chip population of the target chip; and performing iterative loop on the offspring chip population until an offspring individual meeting a chip optimization objective is selected, or until a preset iteration stop condition is met, and configuring chip parameters of the target chip on the basis of the optimal offspring individual in a chip population of a last generation. The method can improve the chip optimization efficiency, reduce the power consumption of arithmetic units, optimize the energy efficiency of arithmetic units, achieve the balance of low-energy-consumption and high-performance for arithmetic units, and improve the operation efficiency of devices.
Need to check novelty before this filing date? Find Prior Art

Description

Arithmetic unit chip setting method, computing subsystem and intelligent computing platform Technical Field

[0001] The present application relates to the field of data processing, and in particular to an arithmetic unit chip setting method, a computing subsystem, and an intelligent computing platform. Background Art

[0002] At present, in order to improve the popularity of intelligent applications in various industries and fields, it is urgent to build an intelligent computing platform to assist the construction of intelligent supercomputing centers, provide a foundation for the construction of artificial intelligence platforms for scientific research, industry, and urban services, and further realize talent gathering, industrial upgrading, and development through intelligent computing platforms.

[0003] With the increasing application of computationally intensive tasks like artificial intelligence and deep learning, traditional high-energy-consuming equipment can no longer meet environmental protection needs. Building low-energy computing units can help reduce the power consumption of equipment and the energy consumption of the entire data center, thereby saving energy, reducing carbon emissions, and promoting green and sustainable development.

[0004] In related technologies, the chip design of low-energy computing units involves optimizing multiple parameters, such as performance, power consumption, and heat dissipation. This optimization process requires considering trade-offs between different objectives, making design optimization very challenging. Traditional manual adjustment methods often fail to fully account for the interplay of numerous factors, making it difficult to obtain optimal parameters. Furthermore, as chip functionality demands increase, circuit complexity also increases, leading to increased power consumption. This power consumption not only affects the device's battery life but also increases the need for heat dissipation. Complex heat dissipation systems also complicate the design.

[0005] Therefore, how to improve chip optimization efficiency and ensure a balance between low energy consumption and high performance is a technical problem that needs to be solved urgently.

[0006] Summary of the Invention

[0007] The present application provides an arithmetic unit chip setting method, a computing subsystem and an intelligent computing platform to improve chip optimization efficiency, reduce the power consumption of the arithmetic unit, optimize energy efficiency, achieve a balance between low energy consumption and high performance for the arithmetic unit, and improve the operating efficiency of the equipment.

[0008] In a first aspect, the present application provides a method for configuring an arithmetic unit chip, the method comprising:

[0009] In response to a setting instruction of a target chip, an initial chip population corresponding to the target chip is generated; wherein each initial individual in the initial chip population is used to represent an initial setting state of the target chip, and each initial setting state corresponds to a set of initial chip parameters of the target chip;

[0010] Using a dynamic evaluation optimization model, the strategy of the initial chip population is optimized to obtain the optimal initial individual in the initial chip population;

[0011] Performing a crossover mutation operation on the optimal initial individual to obtain a population of progeny chips of the target chip; wherein each progeny individual in the population of progeny chips is used to represent an iterative setting state of the target chip, and each iterative setting state corresponds to a set of iterative chip parameters of the target chip;

[0012] An iterative loop is performed on the descendant chip population until a descendant individual that meets the chip optimization goal is selected, or when a preset iteration stop condition is reached, the chip parameters of the target chip are set based on the optimal descendant individual in the last generation chip population; wherein the chip parameters include at least chip structure, logic unit, and circuit wiring structure.

[0013] In a second aspect, an embodiment of the present application provides a computing subsystem, the system comprising:

[0014] a generating unit configured to generate an initial chip population corresponding to the target chip in response to a setting instruction of the target chip; wherein each initial individual in the initial chip population is used to represent an initial setting state of the target chip, and each initial setting state corresponds to a set of initial chip parameters of the target chip;

[0015] an optimization unit configured to perform strategy optimization on the initial chip population using a dynamic evaluation optimization model to obtain an optimal initial individual in the initial chip population;

[0016] An iterative setting unit is configured to perform a crossover mutation operation on the optimal initial individual to obtain a population of offspring chips of the target chip; wherein each offspring individual in the offspring chip population is used to represent an iterative setting state of the target chip, and each iterative setting state corresponds to a set of iterative chip parameters of the target chip; an iterative loop is performed on the offspring chip population until an offspring individual that meets the chip optimization target is selected, or when a preset iteration stop condition is reached, the chip parameters of the target chip are set based on the optimal offspring individual in the last generation chip population; wherein the chip parameters include at least chip structure, logic unit, and circuit wiring structure.

[0017] In a third aspect, an embodiment of the present application provides a computing device, the computing device comprising:

[0018] at least one processor, memory, and input-output unit;

[0019] The memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the arithmetic unit chip setting method of the first aspect.

[0020] In a fourth aspect, a computer-readable storage medium is provided, which includes instructions. When the instructions are executed on a computer, the computer executes the arithmetic unit chip setting method of the first aspect.

[0021] In the technical solution provided by the embodiment of the present application, first, in response to the setting instruction of the target chip, an initial chip population corresponding to the target chip is generated. Each initial individual in the initial chip population is used to represent an initial setting state of the target chip, and each initial setting state corresponds to a set of initial chip parameters of the target chip. Then, a dynamic evaluation optimization model is used to perform strategy optimization on the initial chip population to obtain the optimal initial individual in the initial chip population. Then, a crossover mutation operation is performed on the optimal initial individual to obtain a child chip population of the target chip. Each child individual in the child chip population is used to represent an iterative setting state of the target chip, and each iterative setting state corresponds to a set of iterative chip parameters of the target chip. Thus, diversity and exploration are introduced into the chip setting scheme through the crossover mutation operation, and more possible solutions are found in the search space, which helps to avoid the phenomenon of easily falling into local optimal solutions in related technologies, helps to obtain the global optimal solution, further improves chip performance, and ensures a balance between performance and power consumption. Finally, an iterative loop is executed on the descendant chip population until a descendant individual that meets the chip optimization target is selected, or when a preset iteration stop condition is reached. Based on the optimal descendant individual from the last generation of chips, the chip parameters of the target chip are set. This process iteratively optimizes the descendant chip population, gradually improving chip performance until a solution that meets the optimization target is found, further improving chip performance and ensuring a balance between performance and power consumption. Chip parameters include at least the chip structure, logic units, and circuit wiring structure.

[0022] The technical solution of this application provides an automated and intelligent arithmetic unit chip setting process, which realizes parameter optimization of the target chip by learning and optimizing the chip population through dynamic evaluation optimization model, as well as cross-mutation and iteration of the chip population, reducing the workload and time consumption of manual design in related technologies, effectively avoiding the problem of falling into local optimal solutions, helping to obtain the global optimal solution, further improving chip performance and chip optimization efficiency, and ensuring a balance between low energy consumption and high performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0024] FIG1 is a schematic flow chart of a method for configuring an arithmetic unit chip according to an embodiment of the present application;

[0025] FIG2 is a schematic diagram showing the principle of an iterative cycle method according to an embodiment of the present application;

[0026] FIG3 is a schematic diagram of the principle of a dynamic evaluation optimization model according to an embodiment of the present application;

[0027] FIG4 is a schematic diagram of a principle of a pre-configuration layer according to an embodiment of the present application;

[0028] FIG5 is a schematic diagram showing the principle of a simulation parameter layer according to an embodiment of the present application;

[0029] FIG6 is a schematic diagram showing the principle of an evaluation strategy network layer according to an embodiment of the present application;

[0030] FIG7 is a schematic diagram of the structure of a computing subsystem according to an embodiment of the present application;

[0031] FIG8 is a schematic structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0032] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application pertains. The terms used herein in the specification of this application are for the purpose of describing specific embodiments only and are not intended to limit this application.

[0034] At present, in order to improve the popularity of intelligent applications in various industries and fields, it is urgent to build an intelligent computing platform to assist the construction of intelligent supercomputing centers, provide a foundation for the construction of artificial intelligence platforms for scientific research, industry, and urban services, and further realize talent gathering, industrial upgrading, and development through intelligent computing platforms.

[0035] With the increasing application of computationally intensive tasks like artificial intelligence and deep learning, traditional high-energy-consuming equipment can no longer meet environmental protection needs. Building low-energy computing units can help reduce the power consumption of equipment and the energy consumption of the entire data center, thereby saving energy, reducing carbon emissions, and promoting green and sustainable development.

[0036] In related technologies, the chip design of low-energy computing units involves optimizing multiple parameters, such as performance, power consumption, and heat dissipation. This optimization process requires considering trade-offs between different objectives, making design optimization very challenging. Traditional manual adjustment methods often fail to fully account for the interplay of numerous factors, making it difficult to obtain optimal parameters. Furthermore, as chip functionality demands increase, circuit complexity also increases, leading to increased power consumption. This power consumption not only affects the device's battery life but also increases the need for heat dissipation. Complex heat dissipation systems also complicate the design.

[0037] Therefore, how to improve chip optimization efficiency and ensure a balance between low energy consumption and high performance is a technical problem that needs to be solved urgently.

[0038] To solve at least one of the above technical problems, an embodiment of the present application provides an arithmetic unit chip setting method, a computing subsystem, and an intelligent computing platform.

[0039] Specifically, in the arithmetic unit chip setting scheme, first, in response to the setting instruction of the target chip, an initial chip population corresponding to the target chip is generated. Each initial individual in the initial chip population is used to represent an initial setting state of the target chip, and each initial setting state corresponds to a set of initial chip parameters of the target chip. Then, a dynamic evaluation optimization model is used to perform strategy optimization on the initial chip population to obtain the optimal initial individual in the initial chip population. Next, a crossover and mutation operation is performed on the optimal initial individual to obtain a child chip population of the target chip. Each child individual in the child chip population is used to represent an iterative setting state of the target chip, and each iterative setting state corresponds to a set of iterative chip parameters of the target chip. Thus, the crossover and mutation operation introduces diversity and exploration into the chip setting scheme, and searches for more possible solutions in the search space, which helps to avoid the phenomenon of easily falling into local optimal solutions in related technologies, helps to obtain the global optimal solution, further improves chip performance, and ensures a balance between performance and power consumption. Finally, an iterative loop is executed on the descendant chip population until a descendant individual that meets the chip optimization goals is selected, or when a preset iteration stop condition is reached. Based on the optimal descendant individual in the final generation of chips, the chip parameters of the target chip are set. This process gradually optimizes chip performance and converges to a more optimal solution, thereby further improving chip performance and ensuring a balance between performance and power consumption. Chip parameters include at least the chip structure, logic units, and circuit wiring structure.

[0040] The arithmetic unit chip setting solution provides an automated and intelligent arithmetic unit chip setting process. Through the dynamic evaluation optimization model, the chip population is learned and optimized, and the chip population is cross-mutated and iterated, the parameter optimization of the target chip is achieved. This reduces the workload and time consumption of manual design in related technologies, effectively avoids falling into the problem of local optimal solutions, helps to obtain the global optimal solution, further improves chip performance and chip optimization efficiency, and ensures a balance between low energy consumption and high performance.

[0041] The arithmetic unit chip configuration scheme provided in the embodiments of the present application can be executed by an electronic device, which can be a server, server cluster, or cloud server. The electronic device can also be a terminal device such as a mobile phone, computer, tablet computer, wearable device, or dedicated device (such as a dedicated terminal device with an arithmetic unit chip configuration system). In an optional embodiment, the electronic device can be installed with a service program for executing the arithmetic unit chip configuration scheme.

[0042] FIG1 is a schematic diagram of a method for configuring an arithmetic unit chip according to an embodiment of the present application. As shown in FIG1 , the method includes the following steps:

[0043] 101 , in response to a setting instruction of a target chip, generate an initial chip population corresponding to the target chip.

[0044] In the embodiment of the present application, each initial individual in the initial chip population is used to represent an initial setting state of the target chip, and each initial setting state corresponds to a set of initial chip parameters of the target chip. Specifically, if a target chip is to be designed, such as an embedded chip for smart home control, then each initial individual may include the following set of initial chip parameters:

[0045] Processor type: such as ARM Cortex-M4 or RISC-V architecture

[0046] Memory size: such as 4KB, 8KB or 16KB

[0047] Storage capacity: such as 128MB, 256MB, or 512MB of flash memory

[0048] Communication interface: such as Wi-Fi, Bluetooth, Zigbee, etc.

[0049] Power consumption level: such as low power consumption, medium power consumption, high power consumption

[0050] These parameter combinations in the initial individuals represent different possible initialization settings for the target chip. For each initial individual, these initial chip parameters collectively determine the characteristics and performance of the chip represented by that initial individual. During the optimization process, these parameters can be generated randomly or set based on prior knowledge or experience to ensure that the individuals in the population cover the diversity of the target chip design space and provide sufficient search space for subsequent optimization algorithms. By performing fitness assessments and genetic manipulations on the parameters in the initial individuals, the individuals in the initial chip population can be continuously improved and optimized until the chip design solution that best suits the target requirements is found. Therefore, the initialization setting represented by each initial individual is an exploration and representation of the target chip design space.

[0051] For example, suppose you want to design a target chip, such as a low-power computing unit chip, with the goal of achieving a balance between performance, power consumption, and area, and being able to perform various tasks such as image processing and artificial intelligence. In this example, an initial chip population is generated based on this goal. Each initial individual represents an initial setting state of the target chip and contains a set of initial chip parameters, as shown below:

[0052] 1. The initial chip parameters of individual 1 are as follows:

[0053] -Processor type: ARM Cortex-A78

[0054] -GPU: Mali-G78

[0055] -Memory: 8GB LPDDR5

[0056] - Storage: 256GB UFS 3.1

[0057] -Process technology: 7nm

[0058] -Performance: Single-core performance is strong, suitable for image processing tasks

[0059] 2. The initial chip parameters of individual 2 are as follows:

[0060] -Processor type: Qualcomm Snapdragon 888

[0061] -GPU: Adreno 660

[0062] -Memory: 12GB LPDDR5

[0063] - Storage: 512GB UFS 3.1

[0064] -Process technology: 5nm

[0065] -Performance: Excellent multi-core performance, suitable for artificial intelligence computing

[0066] 3. The initial chip parameters of individual 3 are as follows:

[0067] -Processor type: Apple A15 Bionic

[0068] -GPU: Apple GPU

[0069] -Memory: 16GB LPDDR5

[0070] -Storage: 1TB NVMe

[0071] -Process technology: 5nm

[0072] -Performance: Excellent overall performance, suitable for multitasking

[0073] By generating an initial chip population with different parameter combinations, different design schemes can be tried during the optimization process. By evaluating the fitness parameters of each individual, individuals with higher fitness are selected for subsequent genetic operations, and continuous evolution and optimization are carried out to ultimately obtain a more targeted and better design scheme to meet the required chip design goals.

[0074] In another example, a preset mechanism can be used to randomly generate an initial chip population corresponding to a target chip. Suppose you want to design a target chip, an embedded chip for an IoT device. You can use a preset mechanism to randomly generate an initial chip population corresponding to the target chip. The specific steps are as follows:

[0075] First, determine the size of the initial chip population: Based on design requirements and computing resources, determine the number of individuals in the initial chip population, for example, 100 individuals. Second, determine the range of chip parameters: For each parameter of the target chip, such as processor type, memory, power consumption, etc., determine its value range. For example, the processor type can be ARM Cortex-M4 or RISC-V; the memory can be 4KB or 8KB, etc. Next, randomly generate initial individuals: Based on the determined parameter range, randomly generate initial individuals within the population size. Each individual represents an initial setting state of the target chip. For example, for the processor type, randomly select ARM Cortex-M4 or RISC-V; for the memory, randomly select 4KB or 8KB. Repeat step 3 until the population size is reached: loop through step 3 to generate the specified number of initial individuals to form the initial chip population.

[0076] This pre-set mechanism randomly generates a set of initial individuals with different parameter combinations, representing different initial settings for the target chip. Subsequent iterations of the evolutionary algorithm optimize and improve these individuals to gradually approach and exceed the design requirements of the target chip.

[0077] 102 , using a dynamic evaluation optimization model to perform strategy optimization on the initial chip population to obtain an optimal initial individual in the initial chip population.

[0078] In this embodiment, a dynamic evaluation optimization model can be used to optimize the strategy of an initial chip population to obtain the optimal initial individual in the population. First, the problem needs to be clearly defined. For example, consider designing an embedded chip based on a balance between performance, power consumption, and area. This problem can be formalized as a dynamic evaluation optimization problem, where the chip design parameters are considered the actions of the agent and the performance evaluation metrics of the target chip are considered the reward signal. Furthermore, a simulation environment is constructed to simulate the chip design and evaluation process. This environment can generate a target chip based on specific chip parameters and calculate its performance, power consumption, and area metrics. A dynamic evaluation optimization algorithm, such as a deep dynamic evaluation optimization network, can then be used to build a model that learns and generates optimal chip design strategies. This model receives the state of the environment (i.e., the current chip design parameters) as input and outputs the next action (i.e., the new design parameter settings). Finally, the dynamic evaluation optimization algorithm can be used for iterative training, optimizing the strategy model through interaction with the environment. In each iteration, an initial individual is selected as the current chip design parameters, and the strategy model is used to generate new chip parameter settings. The generated chip parameters are then applied to the simulation environment, the chip performance metrics are calculated, and a reward signal is obtained. The parameters of the policy model can be updated based on the reward signal so that the model can better generate excellent design strategies.

[0079] Through multiple iterative training, the performance of the strategy model can be gradually improved, and the optimal initial individual can be found, that is, the chip design parameter setting with the best performance under given goals and constraints.

[0080] For example, assuming the goal is to design an embedded chip for an IoT device, a dynamic evaluation optimization model can be used to optimize the chip's processor type, memory size, storage capacity, and other settings. The dynamic evaluation optimization model can generate a target chip based on the chip parameters and optimize the chip parameters using chip performance evaluation metrics (such as power consumption and processing speed) as reward signals. Through multiple iterative training, the optimal initial individual for a given objective can be found—that is, the chip design parameter settings with the best performance and fitness. This results in the best initial individual from the initial chip population that meets the requirements and has the best performance.

[0081] 103 , performing a crossover mutation operation on the optimal initial individual to obtain a population of progeny chips of the target chip.

[0082] Each child chip individual in the child chip population is used to represent an iterative setting state of the target chip, and each iterative setting state corresponds to a set of iterative chip parameters of the target chip.

[0083] The iterative setup state refers to the state or configuration of a specific system, tool, or model during an iterative process. In the context of chip design or optimization, the iterative setup state refers to the parameters, configuration, or settings used during each iteration, which represent a specific state of the system, tool, or model during the iteration. For example, in the process of optimizing a chip design, each iteration uses a specific set of chip parameter settings to generate and evaluate the chip design. This set of chip parameter settings constitutes the setup state for that iteration. By continuously adjusting the parameter settings and iterating, the chip design can be gradually optimized, and a different setup state can be obtained in each iteration. Therefore, the iterative setup state refers to the specific parameters, configuration, or settings used by the system, tool, or model in each iteration of the optimization process, corresponding to the system state and design solution at that iteration stage.

[0084] Iterative chip parameters refer to chip design parameters that are adjusted and optimized during the iterative process of chip design or optimization. In each iteration, these chip parameters are continuously modified and adjusted to achieve better performance, power consumption, area, and other metrics, gradually approaching or achieving the design goals. For example, if you are optimizing the design of an embedded chip, chip parameters may include processor type, frequency, cache size, memory capacity, hardware accelerator usage, and so on. In each iteration, we adjust the values ​​or configurations of these chip parameters to generate a new chip design and conduct a performance evaluation. Based on the evaluation results, we obtain feedback and adjust the chip parameters for the next round of iteration.

[0085] By continuously adjusting and iterating chip parameters, the chip design can be gradually optimized during the iterative process, improving its performance, reducing power consumption, and shrinking its area, ultimately resulting in a target chip that meets the design requirements. Therefore, iterative chip parameters are the chip design parameters that need to be adjusted and optimized at each iteration during the chip design optimization process, enabling incremental improvement and optimization of the chip design.

[0086] In this embodiment, a crossover mutation operation is performed on the optimal initial individual to obtain a population of offspring chips of the target chip. Each offspring individual in the offspring chip population will represent an iterative setting state of the target chip, that is, a set of iterative chip parameters. First, through a crossover operation, two individuals are randomly selected from the optimal initial individual, and two new offspring individuals are generated by exchanging some of their chip parameters. The crossover operation can be performed using different strategies, such as single-point crossover, multi-point crossover, or uniform crossover. Specifically, a crossover point can be randomly selected, and the chip parameters of the two individuals are exchanged at the crossover point, thereby generating two new offspring individuals.

[0087] Mutation introduces random changes or alterations into offspring individuals. Mutation generates new offspring individuals by changing certain chip parameters within an individual. For example, one or more chip parameters can be randomly selected and subjected to small random changes, such as adding or subtracting values ​​within a small variation range, to generate new offspring individuals.

[0088] Through crossover and mutation operations, new variations and diversity can be introduced based on the optimal initial individual, thus generating a population of daughter chips. Each daughter individual in the daughter chip population represents the target chip's settings after one iteration, namely, a set of iterative chip parameters. These parameters can be inherited from the parent individual or modified through crossover and mutation operations. In this way, chip parameter settings can be continuously optimized and improved during each iteration, gradually approaching and exceeding the design requirements of the target chip.

[0089] It should be noted that in practical applications, when performing crossover and mutation operations, it is necessary to ensure that the generated offspring chip population still meets the requirements of the target chip design based on the needs and constraints of the specific problem.

[0090] 104 , performing an iterative loop on the progeny chip population until a progeny individual that meets the chip optimization target is selected, or when a preset iteration stop condition is reached, setting chip parameters of the target chip based on the optimal progeny individual in the last generation chip population.

[0091] In the embodiments of this application, the preset iteration stopping condition is a termination condition set in the optimization algorithm. Once this condition is met, the optimization algorithm will stop iterating. This condition can be when the number of iterations reaches a preset maximum value, when the change in a chip performance indicator is less than a certain threshold, or when a user-defined optimization goal is achieved. In actual applications, the user will select this condition based on the specific situation.

[0092] Chip optimization objectives are the goals to be achieved during the iterative process of an optimization algorithm, including but not limited to improving performance, reducing power consumption, reducing area, and improving stability. In chip design, the specific optimization objectives vary depending on the specific application scenario and requirements of the chip. For example, for an embedded system chip, the optimization goal might be improving performance while reducing power consumption; for a communications chip, the optimization goal might be increasing communication speed and stability. In the optimization algorithm, chip optimization objectives are typically converted into specific performance indicators or objective functions, which are used to measure and benchmark the performance of each individual chip.

[0093] Therefore, in the optimization algorithm, the selection of preset iteration stop conditions and the setting of chip optimization goals are very important, as they directly affect the convergence of the optimization results and the final optimization effect.

[0094] In particular, for the ALU chip, possible chip optimization goals include but are not limited to the following:

[0095] First, improve computing speed: Arithmetic unit chips are primarily used to perform various computations, and computing speed is often one of the most important metrics in practical applications. Therefore, optimization goals can include increasing computing speed, employing more efficient algorithms, or improving efficiency through methods such as increasing parallelism.

[0096] Second, reducing power consumption: With the widespread adoption of mobile devices and embedded systems, reducing power consumption has become a crucial optimization goal for computing chips. By employing low-power design techniques to reduce both static and dynamic power consumption, the chip's energy-saving performance can be improved, extending device battery life.

[0097] Third, improve stability: In certain high-precision computing scenarios, the stability of the ALU chip is a key optimization goal. To achieve this goal, various methods can be used to improve the chip's accuracy and stability, such as adding error-correcting codes, adopting redundant designs, and improving power supply stability.

[0098] Fourth, area reduction: For miniaturized arithmetic unit chips, area is often a key optimization goal. By compressing the logic structure, reducing the number of flip-flops, and optimizing internal wiring, chip area can be reduced, lowering manufacturing costs and increasing integration.

[0099] In summary, when optimizing an ALU chip, appropriate optimization targets can be selected and set according to specific needs to improve the chip's performance, stability, energy-saving performance, cost, and other indicators.

[0100] Chip parameters include at least chip structure, logic units, and circuit wiring structure. Specifically, chip parameters are crucial factors in determining chip design and performance. Optimizing and improving these parameters can effectively enhance chip performance and functionality. Chip parameters encompass many aspects, such as chip structure, logic units, and circuit wiring structure. For example, chip structure refers to the basic physical structure and hierarchy of the chip, including overall chip size, hierarchy, and the layout of each functional unit. For example, the structural parameters of an embedded chip may include the chip's total area, hierarchy, and the arrangement of functional modules. For example, logic units are the basic building blocks used to implement various logical functions within a chip, including logic gates, sequential units, and memory cells. When designing a chip, it is necessary to rationally select the type and number of logic units and appropriately optimize and adjust them. For example, the logic unit parameters of a digital signal processor (DSP) chip may include the number and type of arithmetic units (ALUs), as well as the use of accelerators. For example, circuit wiring structure refers to the interconnection structure and wiring layout between various logic devices within the chip. These parameters, including wiring density, line width, line spacing, and the number of metal layers, directly impact signal transmission speed and power consumption. For example, the circuit wiring parameters of a high-speed communication chip may include the transmission distance of each channel, the layout and routing of the circuit board, etc.

[0101] As an optional embodiment, in 104, performing an iterative loop on the daughter chip population, as shown in FIG2 , can be implemented as follows:

[0102] 201, obtaining the fitness parameter of each offspring individual in the offspring chip population;

[0103] 202 , continue to perform a crossover mutation operation on the optimal offspring individual with the highest fitness parameter to obtain a next-generation chip population of the target chip;

[0104] 203, looping to obtain the fitness parameter of each offspring individual in the next-generation chip population, and continuing to perform crossover mutation operations on the optimal offspring individual with the highest fitness parameter to obtain the next-generation chip population, until an offspring individual that meets the chip optimization goal is selected, or when a preset iteration stop condition is reached, the iterative calculation is stopped.

[0105] In step 201, the fitness parameter of each offspring individual in the offspring chip population is obtained, which helps to evaluate the performance of each offspring individual in the chip design process. In an embodiment of the present application, the fitness parameter can be an evaluation index for chip performance, power consumption, area, etc., which is used to quantify the quality of chip design. That is, the fitness parameter is an evaluation index used to measure the adaptability or degree of advantage of an individual in an evolutionary algorithm or optimization problem. During the optimization process, each individual is assigned a fitness parameter value to reflect its performance in the solution space. The fitness parameter is usually defined based on the specific objectives or constraints of the problem, and can be a single indicator or a combination of multiple indicators.

[0106] In practical applications, fitness parameters are used to assess the performance of individuals in the current chip design process. By calculating the fitness parameters of individuals, their performance can be quantitatively compared to determine which individuals are superior or more suitable for the target solution space. In evolutionary algorithms, fitness parameters serve as the basis for selection operations, determining which individuals are replicated or passed on to the next generation. Generally, individuals with higher fitness have a higher chance of being selected as parents, where their genotypes can be combined with other individuals to produce superior offspring. The fitness parameter is the driving factor in the optimization process, guiding the optimization algorithm towards a more optimal solution. By performing operations such as mutation and crossover on individuals with lower fitness, the optimization algorithm generates new individuals to find better solutions, thereby continuously improving the value of the fitness parameter during the optimization process.

[0107] By continuing to perform crossover and mutation operations on the optimal offspring individual with the highest fitness parameter in steps 202 and 203, continuous optimization of the current optimal solution can be achieved. In this way, even if the current optimal solution has been found, the dynamic process of improving and optimizing it is still maintained in the next generation chip population.

[0108] The fitness parameter evaluation metrics include at least one of the following: the target chip's clock frequency, logic delay, static power consumption, dynamic power consumption, die area, on-chip memory area, and error rate. Clock frequency refers to the number of clock cycles a chip operates in and directly affects its operating speed. In some high-performance applications, increasing clock frequency can significantly improve chip efficiency and computing speed. Logic delay refers to the delay time a chip experiences during logical calculations or signal transmission. Lower logic delay generally improves chip computing speed and response speed. Static power consumption is the fixed power consumption of a chip during operation and does not vary with operating conditions. Lower static power consumption helps reduce the chip's overall power consumption and improves energy efficiency, which is particularly important for resource-limited applications such as mobile devices and wireless sensors. Dynamic power consumption refers to the power consumed by a chip during logic switching and signal transmission. Lower dynamic power consumption can reduce chip heat generation and power consumption, thereby improving battery life. Die area refers to the physical space occupied by a chip and is directly related to chip manufacturing cost and integration level. A smaller die area can reduce manufacturing cost and increase chip integration level. On-chip memory area refers to the area of ​​memory used to store data and instructions within a chip. A smaller on-chip memory area can reduce chip size and improve memory access efficiency and capacity. The error rate refers to the probability of a chip error occurring during operation. A lower error rate improves chip reliability and stability, which is particularly important in applications requiring high computational accuracy.

[0109] It is worth noting that in chip optimization, appropriate fitness parameters can be selected according to specific application requirements and optimization goals to evaluate the quality of individual chips, and appropriate chip individuals can be selected through optimization methods such as genetic algorithms and optimization algorithms to achieve optimization goals.

[0110] For example, suppose that within the population of offspring chips, individual A is evaluated and found to have the optimal fitness parameter, meaning it possesses the best performance within the current design space. By continuously performing crossover and mutation operations on individual A in steps 202 and 203, individual A can be further improved and optimized, maintaining its leading position in the next generation of chip populations and potentially achieving even higher fitness parameters. This iterative cycle continuously optimizes the offspring population, ultimately achieving the optimal offspring individual that meets the chip optimization objectives under given stopping conditions.

[0111] Further optionally, the progeny chip population P is calculated using the following formula: (t+1) The fitness parameter f(x) of the offspring individual x is

[0112] Wherein, (t+1) is the progeny chip population P(t+1) The number of iterations, n is the total number of evaluation indicators, Indicates the fitness parameter score of the offspring individual x on the i-th evaluation index, F i (x) represents the actual score of offspring individual x on the i-th evaluation index, represents the minimum reference value of the i-th evaluation index, represents the maximum reference value of the i-th evaluation index, w i Represents the weight coefficient of the i-th evaluation index.

[0113] In the embodiment of the present application, the parameter optimization of the target chip is achieved by learning and optimizing the chip population through dynamic evaluation optimization model, as well as cross-mutation and iteration of the chip population, which reduces the workload and time consumption of manual design in related technologies, effectively avoids the problem of falling into the local optimal solution, helps to obtain the global optimal solution, further improves the chip performance and chip optimization efficiency, and ensures a balance between low energy consumption and high performance.

[0114] In the above or following embodiments, it is assumed that the dynamic evaluation optimization model includes at least: a pre-configuration layer, a simulation parameter layer, an evaluation strategy network layer, a value function network layer, and an output layer.

[0115] As an example, each part will be introduced below. In the dynamic evaluation optimization model, the pre-configuration layer, simulation parameter layer, evaluation strategy network layer, value function network layer and output layer play different roles, which are described in detail as follows:

[0116] The preconfiguration layer maps the initial configuration state of each individual in the initial chip population to the simulation environment state space to obtain the simulation environment state parameters for each individual. The preconfiguration layer helps map the parameter configuration of the initial chip individuals to the simulation environment for subsequent simulation and evaluation.

[0117] The simulation parameter layer calculates the simulated chip operation status of each initial individual by running the simulated environment state parameters in the simulated environment state space. This layer uses the simulated environment to simulate the chip operation status and helps evaluate the performance of each individual in the virtual environment.

[0118] The evaluation strategy network layer simulates chip operation and uses a pre-set reward feedback module to cyclically calculate the cumulative reward return value for each initial individual. The cumulative reward return value includes chip performance evaluation value, static power consumption evaluation value, dynamic power consumption evaluation value, area evaluation value, etc., which are used to evaluate the quality of each individual.

[0119] The value function network layer uses a simulated state value function to update the reward feedback module's strategy based on the cumulative reward return value of each initial individual. This layer uses the value function to evaluate and update each individual, thereby optimizing and selecting the optimal strategy.

[0120] The output layer selects the optimal individual from the initial chip population with the highest cumulative reward value based on the cumulative reward value obtained in the last cycle. The output layer outputs the optimization results and selects the initial individual that performs best in the simulation environment as the optimal solution.

[0121] Through the dynamic evaluation optimization model of the above hierarchical structure, the initial chip population can be optimized through operation and evaluation in a simulation environment, thereby obtaining the optimal initial individuals to meet the pre-set optimization goals and evaluation indicators.

[0122] Based on the above hypothetical structure, in step 102, a dynamic evaluation optimization model is used to perform strategy optimization on the initial chip population to obtain the optimal initial individual in the initial chip population, as shown in FIG3 , including the following steps:

[0123] 301, mapping the initial setting state of each initial individual in the initial chip population to the simulation environment state space through the pre-configuration layer to obtain the simulation environment state parameters of each initial individual;

[0124] 302, running the simulation environment state parameters in the simulation environment state space through the simulation parameter layer to obtain the simulation chip operation status of each initial individual;

[0125] 303, through the evaluation strategy network layer, based on the operation of the simulated chip, using a preset reward feedback module, cyclically calculate the cumulative reward return value of each initial individual until a set number of cycles is reached or a preset convergence condition is met;

[0126] 304, using the value function network layer, based on the cumulative reward return value of each initial individual, using the simulated state value function to update the strategy of the reward feedback module;

[0127] 305 , through the output layer, according to the cumulative reward return value obtained in the last cycle, select the initial individual with the highest cumulative reward return value from the initial chip population as the optimal initial individual.

[0128] In an embodiment of the present application, the cumulative reward return value is an indicator used to evaluate each individual in the dynamic evaluation optimization model. These indicators involve various aspects of the chip, including at least: chip performance evaluation value, static power consumption evaluation value, dynamic power consumption evaluation value, and area evaluation value.

[0129] Among them, the chip performance evaluation value refers to the efficiency that the chip can achieve when performing tasks, such as operating speed, the amount of data that can be processed, etc. Chip performance is usually one of the important indicators in chip design. The static power consumption evaluation value refers to the power consumption of the chip when it is idle. Static power consumption directly affects the total power consumption and power consumption of the chip. Therefore, static power consumption is also one of the indicators often considered when optimizing the chip. The dynamic power consumption evaluation value refers to the power consumed by the chip when running tasks. The dynamic power consumption of the chip reflects the relationship between chip performance and power consumption, and is usually one of the important indicators in chip design. The area evaluation value refers to the size of the physical space (area) occupied by the chip. The area of ​​the chip is directly related to the manufacturing cost and integration level of the chip.

[0130] Evaluating each individual design using the four metrics above helps measure the strengths and weaknesses of different designs and optimize the optimal design. Other factors, such as chip design reliability, design stability, and fault tolerance, can also be considered during the evaluation process to more comprehensively assess the strengths and weaknesses of each individual design.

[0131] Through steps 301 to 304, each individual in the initial chip population can be comprehensively evaluated and optimized, so that the best-performing individual can be finally selected as the optimal initial individual, providing better design solutions and results for chip design, achieving more effective power consumption prediction and optimization, and thus improving the system's energy efficiency performance and user experience.

[0132] As an optional embodiment, in 301, the initial setting state of each initial individual in the initial chip population is mapped to the simulation environment state space through the pre-configuration layer to obtain the simulation environment state parameters of each initial individual, as shown in FIG4 , further comprising the following steps:

[0133] 401, performing state encoding on the initialization setting state of each initial individual to obtain an initialization setting state vector corresponding to each initial individual in the simulation environment state space;

[0134] 402 , mapping each initialized state vector into the simulation environment state space for parameter simulation processing to obtain the simulation environment state parameters of each initial individual.

[0135] In steps 401 and 402, first, the initialization state of each initial individual is state-encoded to obtain the initialization state vector corresponding to each individual in the simulated environment state space. The effect of this step is to convert the initialization state of each individual into a specific vector representation, which is convenient for processing and analysis in the simulated environment state space. Then, each initialization state vector is mapped into the simulated environment state space for parameter simulation processing to obtain the simulated environment state parameters of each initial individual. The effect of this step is to convert the encoded initialization state vector into the parameter value in the simulated environment, thereby accurately simulating the state of each individual in the simulated environment.

[0136] In the embodiment of the present application, the initialization setting state vector s of the initial individual m is m The acquisition process of s is expressed as the following formula: m =f(e m )=[pm , E_stastic(m), E_dynamic(m), A(m)];

[0137] Among them, m represents the mth initial individual, e m Represents the state vector obtained after encoding the mth initial individual, s m represents the state vector of the simulated environment corresponding to the mth initial individual, f(e m ) represents the state vector e m Obtain the corresponding initialization setting state vector s m The calculation process function, p m represents the performance index of the mth initial individual, E_stastic(m) represents the static power consumption of the mth initial individual, E_dynamic(m) represents the dynamic power consumption of the mth initial individual, and A(m) represents the area of ​​the mth initial individual.

[0138] The combination of these two steps converts the initial setup state of each individual in the initial chip population into a parameter representation in the simulation environment state space, providing accurate input for subsequent simulation and evaluation. This process transforms the abstract initial setup state of the initial individuals into specific, actionable simulation environment state parameters, providing a foundation and convenience for subsequent evaluation and policy optimization, helping to more accurately assess the performance of each individual and promote policy optimization.

[0139] Optionally, the simulation parameter layer includes at least a simulator and a monitoring simulator. A simulator is a computer program used to simulate the chip's operating state. Its primary function is to convert simulation environment state parameters into input parameters required for simulator operation. The simulator then simulates chip operation using the input parameters to obtain the chip's operating status for each initial individual. These statuses include indicators such as operating speed, processing power, and performance, which are used to calculate the cumulative reward return value.

[0140] A monitoring simulator is a computer program used to monitor chip performance and extract key performance indicators. Its primary function is to monitor the chip's operating status for each initial individual output by the simulator and extract key performance indicators. Depending on the type of indicators to be extracted, the monitoring simulator can monitor and extract a variety of metrics, such as clock frequency, logic delay, static power consumption, and dynamic power consumption.

[0141] By introducing the simulation parameter layer, the operation of each initial individual chip can be simulated and monitored in a virtual environment, outputting data on chip performance and key indicators. This helps evaluate the performance and anti-interference capabilities of each individual chip, and provides important data support and feedback for subsequent optimization solutions. In practical applications, the simulation parameter layer can also be expanded to other specific simulation and monitoring modules to meet specific design and evaluation requirements.

[0142] In the above step 302, the simulation environment state parameters are run in the simulation environment state space through the simulation parameter layer to obtain the simulated chip operation status of each initial individual, as shown in FIG5 , further comprising the following steps:

[0143] 501, converting the simulation environment state parameters of each initial individual into the input parameters required for the simulator to run;

[0144] 502, running a simulator to perform chip operation simulation on the input parameters to obtain a chip operation state of each initial individual;

[0145] 503 , using a monitoring simulator to monitor the chip operation status of each initial individual to obtain chip operation monitoring data of each initial individual.

[0146] The chip operation monitoring data includes at least: clock frequency, logic delay, static power consumption, dynamic power consumption, chip area, on-chip memory area, and error rate. Similar to the above description, it will not be repeated here.

[0147] For example, assume there is an initial chip instance with some initial setup parameters, such as voltage, clock frequency, and router settings. These parameters can be converted into input parameters required for simulator operation. In step 501, the simulation environment parameters of each initial instance are converted into the input parameters required for simulator operation. These parameters may include digital-to-analog converter (DAC) settings, transmission protocol settings, hardware accelerator settings, etc. These parameters are converted into an input format that the simulator can understand to facilitate chip operation simulation. Then, in step 502, the simulator is run to simulate chip operation using the input parameters. In the simulator, the input parameters are used to simulate the chip's operating state, such as the transmission characteristics, logic operations, and signal processing of the simulated circuit. Through the simulator, chip operation status data corresponding to each initial instance can be obtained, such as power consumption, timing, and signal strength. Next, in step 503, the chip operation status of each initial instance is monitored using a monitoring simulator. The monitoring simulator can monitor the chip's status in real time during operation and record monitoring data such as clock frequency, logic delay, static power consumption, and dynamic power consumption. This data provides detailed information about the operation of the initial instance chip.

[0148] Combining the above steps, we can complete the complete process from initial individual parameter setting to simulator operation and then monitoring simulator data, thereby obtaining chip operating status and performance data for each initial individual. This data is invaluable for subsequent evaluation and optimization, helping to determine the optimal initial individual and optimize the design solution.

[0149] In the above step 303, the strategy network layer is evaluated, based on the operation of the simulated chip, and a preset reward feedback module is used to cyclically calculate the cumulative reward return value of each initial individual until the set number of cycles is reached or the preset convergence condition is met. As shown in FIG6 , the following steps are further included:

[0150] 601 , based on the chip operation monitoring data of each initial individual, calculate the instant reward return value of each initial individual through the reward feedback module.

[0151] The instant reward value is an evaluation result based on chip operating status data. By analyzing key data such as the chip's performance and power consumption indicators, the instant reward value can be calculated according to pre-set reward rules. Furthermore, the instant reward value may include at least a performance indicator reward and a power consumption indicator reward.

[0152] Taking the design of an image processor as an example, we will describe how to calculate an immediate reward based on chip operation monitoring data. Assume there is an initial individual whose design parameters include clock frequency (Fclk), algorithm complexity (C), and resource usage (A). In steps 501, 502, and 503, the chip operation monitoring data for this chip under these given design parameters has been obtained. Furthermore, the immediate reward for this individual can be calculated using a pre-defined reward feedback module as follows:

[0153] Performance indicator reward: Assuming that the performance indicators of the image processor (such as image clarity and color reproduction) are expected to be as high as possible, a performance indicator reward rule can be set. For example, if the processing speed (SP) is higher than the preset threshold (SPt), a performance indicator reward of 0.8 is given; if the algorithm complexity (C) is higher than the preset threshold (Ct), a performance indicator reward of 0.5 is given.

[0154] Assuming that the processing speed of the initial individual is 100fps, the algorithm complexity is 2, and SPt and Ct are 90fps and 3 respectively, then the performance indicator reward value of this individual is 0.8 points.

[0155] Power consumption indicator reward: Assuming that we want the power consumption of the image processor to be as low as possible, we can set a power consumption indicator reward rule. For example, if the power consumption (P) is lower than the preset threshold (Pt), a power consumption indicator reward of 0.5 is given.

[0156] Assuming the initial individual's chip power consumption is 1.5W and Pt is 2W, the individual's power consumption indicator reward value is 0.5 points. Based on the above calculation rules, the individual's immediate reward return value is 0.8 + 0.5 = 1.3 points. Through the above calculation process, we can obtain the immediate reward return value of each initial individual, which can be used for subsequent cumulative reward return value calculation and optimization decision making.

[0157] 602 , cumulatively calculating the instantaneous reward return value of each initial individual and the historical cumulative reward return value obtained in the previous iteration cycle to obtain the cumulative reward return value of each initial individual.

[0158] In step 602, the immediate reward return value of each initial individual is cumulatively calculated with the historical cumulative reward return value obtained in the previous iteration cycle to obtain the cumulative reward return value of each initial individual. This cumulative calculation can accumulate the impact of the immediate reward return value, thereby better considering the overall performance of the individual. This provides more accurate reference data for subsequent optimization and decision-making.

[0159] Further optionally, the cumulative reward return value can be calculated by multiplying the product of the discount factor and the historical cumulative reward return value with the sum of the immediate reward return value. Furthermore, the cumulative reward return value is the product of the discount factor and the historical cumulative reward return value with the sum of the immediate reward return value. In this way, the introduction of the discount factor can take into account the importance of the immediate reward return value and the correlation between the previous and next moments when calculating the cumulative reward return value, thereby more reasonably evaluating the cumulative reward return value of each initial individual.

[0160] Through the above steps, the cumulative reward value of each initial individual can be calculated repeatedly to evaluate the individual's performance and advantages and disadvantages, thereby guiding the next step of optimization and adjustment. Through continuous iteration and optimization, the overall performance and efficiency of the chip can be gradually improved.

[0161] In the above embodiment, in the following set of equations, the first equation represents the calculation of the reward prediction error, that is, the error between the current reward value and the predicted reward value. The second equation represents the use of the gradient descent method to update the parameters of the function approximation based on the current reward prediction error, thereby achieving the effect of updating the simulated state value function. That is, the parameter update formula of the simulated state value function is expressed as: δ=R(s)+γ*V(s′)-V(s); Δθ=α*ΔθV(s)*δ;

[0162] Among them, δ represents the reward prediction deviation, which represents the error between the actual immediate reward return value and the predicted immediate reward return value, R(s) represents the immediate reward return value obtained by taking a certain parameter under the state variable s of the simulated state value function, γ represents the discount factor, V(s) represents the predicted state value parameter under the state variable s of the simulated state value function, V(s′) represents the predicted state value parameter under another state variable s′ of the simulated state value function, α is the learning rate, θ represents the parameter approximated by the simulated state value function, ΔθV(s) represents the calculation result of the gradient solution of the predicted state value parameter under the parameter θ, Δθ represents the function approximation parameter, and Δθ is used to update the parameters of the simulated state value function based on the reward prediction deviation.

[0163] In summary, through the above steps, each individual in the initial chip population can be comprehensively evaluated and optimized, so that the best-performing individual can be finally selected as the optimal initial individual, providing better design solutions and results for chip design, achieving more effective power consumption prediction and optimization, and thus improving the system's energy efficiency performance and user experience.

[0164] In another embodiment of the present application, a computing subsystem is provided. As shown in FIG7 , the computing subsystem includes the following units:

[0165] a generating unit configured to generate an initial chip population corresponding to the target chip in response to a setting instruction of the target chip; wherein each initial individual in the initial chip population is used to represent an initial setting state of the target chip, and each initial setting state corresponds to a set of initial chip parameters of the target chip;

[0166] an optimization unit configured to perform strategy optimization on the initial chip population using a dynamic evaluation optimization model to obtain an optimal initial individual in the initial chip population;

[0167] An iterative setting unit is configured to perform a crossover mutation operation on the optimal initial individual to obtain a population of offspring chips of the target chip; wherein each offspring individual in the offspring chip population is used to represent an iterative setting state of the target chip, and each iterative setting state corresponds to a set of iterative chip parameters of the target chip; an iterative loop is performed on the offspring chip population until an offspring individual that meets the chip optimization target is selected, or when a preset iteration stop condition is reached, the chip parameters of the target chip are set based on the optimal offspring individual in the last generation chip population; wherein the chip parameters include at least chip structure, logic unit, and circuit wiring structure.

[0168] Further optionally, the iteration setting unit performs an iterative loop on the descendant chip population, and is specifically configured to:

[0169] Obtaining the fitness parameter of each offspring individual in the offspring chip population;

[0170] Continue to perform crossover and mutation operations on the optimal offspring individuals with the highest fitness parameters to obtain a next generation chip population of the target chip;

[0171] The steps of obtaining the fitness parameter of each offspring individual in the next-generation chip population and continuing to perform crossover mutation operations on the optimal offspring individual with the highest fitness parameter to obtain the next-generation chip population are executed cyclically until an offspring individual that meets the chip optimization goal is selected or a preset iteration stopping condition is reached, and the iterative calculation is stopped.

[0172] Further optionally, the progeny chip population P is calculated using the following formula: (t+1) The fitness parameter f(x) of the offspring individual x is

[0173] Wherein, (t+1) is the progeny chip population P (t+1) The number of iterations, n is the total number of evaluation indicators, Indicates the fitness parameter score of the offspring individual x on the i-th evaluation index, F i(x) represents the actual score of offspring individual x on the i-th evaluation index, represents the minimum reference value of the i-th evaluation index, represents the maximum reference value of the i-th evaluation index, w i Represents the weight coefficient of the i-th evaluation index;

[0174] The evaluation index of the fitness parameter includes at least one of the following: clock frequency, logic delay, static power consumption, dynamic power consumption, chip area, on-chip memory area, and error rate of the target chip.

[0175] Further optionally, the dynamic evaluation optimization model includes at least: a pre-configuration layer, a simulation parameter layer, an evaluation strategy network layer, a value function network layer, and an output layer;

[0176] The optimization unit uses a dynamic evaluation optimization model to perform strategy optimization on the initial chip population to obtain the optimal initial individual in the initial chip population. Specifically, the optimization unit is configured as follows:

[0177] Mapping the initial setting state of each initial individual in the initial chip population to the simulation environment state space through the pre-configuration layer to obtain the simulation environment state parameters of each initial individual;

[0178] Running the simulation environment state parameters in the simulation environment state space through the simulation parameter layer to obtain the simulation chip operation status of each initial individual;

[0179] Through the evaluation strategy network layer, based on the operation of the simulated chip, a preset reward feedback module is used to cyclically calculate the cumulative reward return value of each initial individual until a set number of cycles is reached or a preset convergence condition is met; the cumulative reward return value includes at least: chip performance evaluation value, static power consumption evaluation value, dynamic power consumption evaluation value, and area evaluation value;

[0180] Through the value function network layer, based on the cumulative reward return value of each initial individual, the reward feedback module is updated using the simulated state value function;

[0181] Through the output layer, according to the cumulative reward return value obtained in the last cycle, the initial individual with the highest cumulative reward return value is selected from the initial chip population as the optimal initial individual.

[0182] Further optionally, the optimization unit maps the initial setting state of each initial individual in the initial chip population to the simulation environment state space through the pre-configuration layer to obtain the simulation environment state parameters of each initial individual, and is specifically configured as follows:

[0183] Performing state encoding on the initial setting state of each initial individual to obtain the initial setting state vector corresponding to each initial individual in the state space of the simulation environment;

[0184] Mapping each initialization setting state vector into the simulation environment state space for parameter simulation processing to obtain the simulation environment state parameters of each initial individual;

[0185] Among them, the initial setting state vector s of the initial individual m is m The acquisition process of s is expressed as the following formula: m =f(e m )=[p m ,E_stastic)m),E_dynamic(m),A(m))];

[0186] Among them, m represents the mth initial individual, e m Represents the state vector obtained after encoding the mth initial individual, s m represents the state vector of the simulated environment corresponding to the mth initial individual, f(e m ) represents the state vector e m Obtain the corresponding initialization setting state vector s m The calculation process function, p m represents the performance index of the mth initial individual, E_static(m) represents the static power consumption of the mth initial individual, E_dynamic(m) represents the dynamic power consumption of the mth initial individual, and A(m) represents the area of ​​the mth initial individual.

[0187] Further optionally, the simulation parameter layer includes at least: a simulator and a monitoring simulator; the optimization unit operates the simulation environment state parameters in the simulation environment state space through the simulation parameter layer to obtain the simulation chip operation status of each initial individual, and is specifically configured as follows:

[0188] Convert the simulation environment state parameters of each initial individual into the input parameters required for the simulator to run;

[0189] The operation simulator performs chip operation simulation on the input parameters to obtain the chip operation state of each initial individual;

[0190] Using a monitoring simulator to monitor the chip operation status of each initial individual to obtain chip operation monitoring data of each initial individual;

[0191] The chip operation monitoring data includes at least: clock frequency, logic delay, static power consumption, dynamic power consumption, chip area, on-chip memory area, and error rate.

[0192] Further optionally, the optimization unit, through the evaluation strategy network layer, based on the operation of the simulated chip, uses a preset reward feedback module to cyclically calculate the cumulative reward return value of each initial individual until a set number of cycles is reached or a preset convergence condition is met, and is specifically configured as follows:

[0193] Based on the chip operation monitoring data of each initial individual, the reward feedback module calculates the immediate reward return value of each initial individual; wherein the immediate reward return value includes at least: performance indicator reward and power consumption indicator reward;

[0194] The instant reward return value of each initial individual and the historical cumulative reward return value obtained in the previous iteration cycle are accumulated and calculated to obtain the cumulative reward return value of each initial individual; the cumulative reward return value is the product of the discount factor and the historical cumulative reward return value, and the sum of the instant reward return value.

[0195] Further optionally, the parameter update formula of the simulation state value function is expressed as: δ=R(s)+γ*V(s′)-V(s); Δθ=α*ΔθV(s)*δ;

[0196] Among them, δ represents the reward prediction deviation, which represents the error between the actual immediate reward return value and the predicted immediate reward return value, R(s) represents the immediate reward return value obtained by taking a certain parameter under the state variable s of the simulated state value function, γ represents the discount factor, V(s) represents the predicted state value parameter under the state variable s of the simulated state value function, V(s′) represents the predicted state value parameter under another state variable s′ of the simulated state value function, α is the learning rate, θ represents the parameter approximated by the simulated state value function, ΔθV(s) represents the calculation result of the gradient solution of the predicted state value parameter under the parameter θ, Δθ represents the function approximation parameter, and Δθ is used to update the parameters of the simulated state value function based on the reward prediction deviation.

[0197] In the embodiments of the present application, the chip optimization efficiency can be improved, the power consumption of the operator can be reduced, the energy efficiency can be optimized, the balance between low energy consumption and high performance can be achieved for the operator, and the operating efficiency of the equipment can be improved.

[0198] In another embodiment of the present application, an intelligent computing platform is provided, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0199] Memory for storing computer programs;

[0200] The processor is used to implement the arithmetic unit chip setting method described in the method embodiment when executing the program stored in the memory.

[0201] The communication bus 1140 mentioned in the electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus 1140 can be divided into an address bus, a data bus, a control bus, etc.

[0202] For example, let's assume we need to build a large-scale, autonomous, and controllable intelligent computing platform based on specialized neural network chips. This platform will provide the hardware foundation for the research, development, and construction of intelligent computing platforms. This intelligent computing platform will also provide the hardware foundation for the construction of an intelligent supercomputing center. This center will serve as an artificial intelligence platform for scientific research, industry, and urban development, thereby attracting talent and developing industries.

[0203] Specifically, the intelligent computing platform primarily consists of five components: an intelligent hardware platform, an intelligent computing cloud operating system, application environment development, a big data platform, and an intelligent application PaaS platform. Based on intelligent computing theory, the intelligent hardware platform integrates deep learning chips, AI smart accelerator cards, and distributed servers to provide the foundational hardware support for the entire supercomputing platform and related derivative platforms. Its primary components include the intelligent computing subsystem, the network switching subsystem, the data storage subsystem, and the support and management subsystem.

[0204] Further optionally, the intelligent computing subsystem is the hardware module responsible for computing, mainly from the construction of low-energy arithmetic units, sparse memory access DMA (Direct Memory Access), deep learning processor cache structure, deep learning storage consistency, artificial intelligence processor card design, and dedicated servers equipped with intelligent processing cards.

[0205] The embodiment of the present application provides an arithmetic unit chip configuration method for constructing a low-energy arithmetic unit.

[0206] For ease of representation, FIG8 shows only one thick line, but this does not mean that there is only one bus or one type of bus.

[0207] The communication interface 1120 is used for communication between the electronic device and other devices.

[0208] The memory 1130 may include a random access memory (RAM) or a non-volatile memory (non-volatile memory), such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.

[0209] The above-mentioned processor 1110 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0210] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed, can implement the steps that can be performed by the electronic device in the above method embodiment.

Claims

1. A method for configuring an arithmetic unit chip, characterized in that: include: In response to a setting instruction of a target chip, an initial chip population corresponding to the target chip is generated; wherein each initial individual in the initial chip population is used to represent an initial setting state of the target chip, and each initial setting state corresponds to a set of initial chip parameters of the target chip; Using a dynamic evaluation optimization model, the strategy of the initial chip population is optimized to obtain the optimal initial individual in the initial chip population; Performing a crossover mutation operation on the optimal initial individual to obtain a population of progeny chips of the target chip; wherein each progeny individual in the population of progeny chips is used to represent an iterative setting state of the target chip, and each iterative setting state corresponds to a set of iterative chip parameters of the target chip; An iterative loop is performed on the descendant chip population until a descendant individual that meets the chip optimization goal is selected, or when a preset iteration stop condition is reached, the chip parameters of the target chip are set based on the optimal descendant individual in the last generation chip population; wherein the chip parameters include at least chip structure, logic unit, and circuit wiring structure.

2. The method for configuring an arithmetic unit chip according to claim 1, wherein: The performing of an iterative loop on the daughter chip population includes: Obtaining the fitness parameter of each offspring individual in the offspring chip population; Continue to perform crossover and mutation operations on the optimal offspring individuals with the highest fitness parameters to obtain a next generation chip population of the target chip; The steps of obtaining the fitness parameter of each offspring individual in the next-generation chip population and continuing to perform crossover mutation operations on the optimal offspring individual with the highest fitness parameter to obtain the next-generation chip population are executed cyclically until an offspring individual that meets the chip optimization goal is selected or a preset iteration stopping condition is reached, and the iterative calculation is stopped.

3. The method for configuring an arithmetic unit chip according to claim 2, wherein: The progeny chip population P is calculated using the following formula: (t+1) The fitness parameter f(x) of the offspring individual x is Wherein, (t+1) is the progeny chip population P (t+1) The number of iterations, n is the total number of evaluation indicators, Indicates the fitness parameter score of the offspring individual x on the i-th evaluation index, F i (x) represents the offspring individual x in the ith The actual score on the evaluation index, represents the minimum reference value of the i-th evaluation index, represents the maximum reference value of the i-th evaluation index, w i Represents the weight coefficient of the i-th evaluation index; The evaluation index of the fitness parameter includes at least one of the following: clock frequency, logic delay, static power consumption, dynamic power consumption, chip area, on-chip memory area, and error rate of the target chip.

4. The method for configuring an arithmetic unit chip according to claim 1, wherein: The dynamic evaluation optimization model at least includes: a pre-configuration layer, a simulation parameter layer, an evaluation strategy network layer, a value function network layer, and an output layer; The method of using a dynamic evaluation optimization model to perform strategy optimization on the initial chip population to obtain the optimal initial individual in the initial chip population includes: Mapping the initial setting state of each initial individual in the initial chip population to the simulation environment state space through the pre-configuration layer to obtain the simulation environment state parameters of each initial individual; Running the simulation environment state parameters in the simulation environment state space through the simulation parameter layer to obtain the simulation chip operation status of each initial individual; Through the evaluation strategy network layer, based on the operation of the simulated chip, a preset reward feedback module is used to cyclically calculate the cumulative reward return value of each initial individual until a set number of cycles is reached or a preset convergence condition is met; the cumulative reward return value includes at least: chip performance evaluation value, static power consumption evaluation value, dynamic power consumption evaluation value, and area evaluation value; Through the value function network layer, based on the cumulative reward return value of each initial individual, the reward feedback module is updated using the simulated state value function; Through the output layer, according to the cumulative reward return value obtained in the last cycle, the initial individual with the highest cumulative reward return value is selected from the initial chip population as the optimal initial individual.

5. The method for configuring an arithmetic unit chip according to claim 4, wherein: The initialization setting state of each initial individual in the initial chip population is mapped to the simulation environment state space through the pre-configuration layer to obtain the simulation environment state parameters of each initial individual, including: Performing state encoding on the initial setting state of each initial individual to obtain the initial setting state vector corresponding to each initial individual in the state space of the simulation environment; Mapping each initialization setting state vector to the simulation environment state space for parameter simulation processing to obtain the simulation environment state parameters of each initial individual; Among them, the initialization setting state vector s of the initial individual m m The acquisition process is expressed as the following formula: s m =f(e m )=[p m ,E_stastic(m),E_dynamic(m),A(m)]; Among them, m represents the mth initial individual, e m Represents the state vector obtained after encoding the mth initial individual, s m represents the state vector of the simulated environment corresponding to the mth initial individual, f(e m ) represents the state vector e m Obtain the corresponding initialization setting state vector s m The calculation process function, p m represents the performance index of the mth initial individual, E_stastic(m) represents the static power consumption of the mth initial individual, E_dynamic represents the dynamic power consumption of the mth initial individual, and A(m) represents the area of ​​the mth initial individual.

6. The method for configuring an arithmetic unit chip according to claim 4, wherein: The simulation parameter layer includes at least: simulator and monitoring simulator; The step of running the simulation environment state parameters in the simulation environment state space through the simulation parameter layer to obtain the simulation chip operation status of each initial individual includes: Convert the simulation environment state parameters of each initial individual into the input parameters required for the simulator to run; The operation simulator performs chip operation simulation on the input parameters to obtain the chip operation state of each initial individual; Using a monitoring simulator to monitor the chip operation status of each initial individual to obtain chip operation monitoring data of each initial individual; The chip operation monitoring data includes at least: clock frequency, logic delay, static power consumption, dynamic power consumption, chip area, on-chip memory area, and error rate.

7. The method for configuring an arithmetic unit chip according to claim 6, wherein: The evaluation strategy network layer, based on the operation of the simulated chip, uses a preset reward feedback module to cyclically calculate the cumulative reward return value of each initial individual until the set number of cycles is reached or the preset convergence condition is met, including: Based on the chip operation monitoring data of each initial individual, the reward feedback module calculates the immediate reward return value of each initial individual; wherein the immediate reward return value includes at least: performance indicator reward and power consumption indicator reward; The instant reward return value of each initial individual and the historical cumulative reward return value obtained in the previous iteration cycle are accumulated and calculated to obtain the cumulative reward return value of each initial individual; the cumulative reward return value is the product of the discount factor and the historical cumulative reward return value, and the sum of the instant reward return value.

8. The method for configuring an arithmetic unit chip according to claim 7, wherein: The parameter update formula of the simulation state value function is expressed as: δ=R(s)+γ*V(s′)-V(s); Δθ=α*ΔθV(s)*δ; Wherein, δ represents the reward prediction deviation, which represents the error between the actual instant reward return value and the predicted instant reward return value, R(s) represents the instant reward return value obtained by taking a certain parameter under the state variable s of the simulated state value function, γ represents the discount factor, V(s) represents the predicted state value parameter under the state variable s of the simulated state value function, and V(s′) represents the predicted state value parameter under another state variable s′ of the simulated state value function. value parameter, α is the learning rate, θ represents the parameter of the simulated state value function approximation, ΔθV(s) represents the calculation result of the gradient solution of the predicted state value parameter under the parameter θ, Δθ represents the function approximation parameter, and Δθ is used to update the parameters of the simulated state value function based on the reward prediction deviation.

9. A computing subsystem, characterized in that: The computing subsystem includes: a generating unit configured to generate an initial chip population corresponding to the target chip in response to a setting instruction of the target chip; Each initial individual in the initial chip population is used to represent an initial setting state of the target chip, and each initial setting state corresponds to a set of initial chip parameters of the target chip; an optimization unit configured to perform strategy optimization on the initial chip population using a dynamic evaluation optimization model to obtain an optimal initial individual in the initial chip population; An iterative setting unit is configured to perform a crossover mutation operation on the optimal initial individual to obtain a population of offspring chips of the target chip; wherein each offspring individual in the offspring chip population is used to represent an iterative setting state of the target chip, and each iterative setting state corresponds to a set of iterative chip parameters of the target chip; an iterative loop is performed on the offspring chip population until an offspring individual that meets the chip optimization target is selected, or when a preset iteration stop condition is reached, the chip parameters of the target chip are set based on the optimal offspring individual in the last generation chip population; wherein the chip parameters include at least chip structure, logic unit, and circuit wiring structure.

10. An intelligent computing platform, characterized in that: The intelligent computing platform includes: at least one processor, memory, and input-output unit; The memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the method for setting an arithmetic unit chip according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Swarm intelligence method for consensus active learning

    CN111160511A

  • Integrated circuit design method based on population optimization algorithm

    CN112765933A

  • Chip circuit design method based on double-layer multi-objective optimization

    CN117454824A

  • Multi-objective chip circuit parameter optimization design method

    CN117556775A

  • Apparatus and method for automated reward shaping

    US20240046154A1

Cited By

  • Method and device for optimizing sealing structure of extra-high pressure wellhead equipment

    CN120874481A

  • Digital integrated circuit TMR reinforcement decision method based on multi-objective parameter optimization

    CN122113811A

  • Memory management method and electronic equipment

    CN122285462A

  • A method and system for optimizing energy efficiency of high-power-consuming equipment based on cloud-edge-device collaboration

    CN122571068A