Chip screening method and system based on deep Q learning
Through the chip screening method based on deep Q learning, chips suitable for parallel connection are screened, which solves the problem of uneven current distribution in semiconductor chips when parallel connection, and improves the stability and reliability of the system.
Patent Information
- Application Number
- CN202510342196.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When semiconductor chips work in parallel, due to differences in device parameters, internal parasitic characteristics and temperature distribution, uneven current distribution will result in problems such as uneven junction temperature, overheating and device aging, which will affect the stability and reliability of the module.
The chip screening method based on deep Q learning is adopted, and chips suitable for parallel connection are screened from multiple chips by obtaining chip parameter information, by using the network model of deep Q learning, to reduce parameter differences and reduce current junction temperature imbalance.
It effectively reduces the parameter differences between parallel devices, reduces the current junction temperature imbalance caused by chip differences, and improves the performance and life of the chip system.
Smart Images

Figure CN120354912A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of semiconductor chips, and particularly relates to a chip screening method and system based on deep Q-learning. Background Art
[0002] Semiconductors are widely used in the field of power electronics due to their superior high power density and high efficiency characteristics. However, when multiple semiconductor chips operate in parallel, due to differences in device parameters, internal parasitic characteristics, and temperature distribution, it is easy to cause uneven current distribution, which in turn leads to problems such as uneven junction temperature, overheating, and device aging, affecting the stability and reliability of the module. Therefore, how to reduce current imbalance and temperature difference to improve the performance and lifespan of the chip system has become an urgent technical problem to be solved. Summary of the Invention
[0003] The purpose of the present invention is to provide a chip screening method and system based on deep Q-learning to solve the deficiencies in the prior art. It proposes a chip screening strategy based on deep Q-learning, and the selected chips can reduce the parameter differences between parallel devices and effectively reduce the uneven current and junction temperature caused by chip differences.
[0004] An embodiment of the present application provides a chip screening method based on deep Q-learning, and the method includes: Obtain n chips to be processed and parameter information of the chips to be processed; Based on the parameter information of the chips to be processed, use a network model based on deep Q-learning to determine m parallel chips from the n chips to be processed.
[0005] Optionally, based on the parameter information of the chips to be processed, using a network model based on deep Q-learning to determine m parallel chips from the n chips to be processed includes: Step 1: Obtain an m-dimensional vector state s and use a greedy strategy to determine an action a; wherein, the greedy strategy includes randomly determining a swap action with a probability; Step 2: Initialize a network model based on deep Q-learning and an experience replay pool; wherein, initializing the network model based on deep Q-learning includes initializing a neural network Q-Network and a target network ; Step 3: Execute the calculation of the reward function value for the state s and execute the action a with a probability to determine an experience sample and store it in the experience replay pool; wherein, represents the reward function, represents the updated state; Step 4: According to the experience sample , calculate the target Q-value and the current Q-value; Step 5: According to the target Q-value and the current Q-value, use the minimized loss function to update the weights of the neural network Q-Network; Step 6: Repeat the above Steps 1 to 5 until a preset stop condition is reached to output m parallel chips.
[0006] Optionally, the reward function satisfies:
[0007] where represents the state corresponding reward value, represents the variance operation, represents the performance parameter of the chip.
[0008] Optionally, calculating the target Q-value and the current Q-value according to the experience sample includes: Calculate the target Q-value through the following formula:
[0009] where represents the target Q-value, represents the discount factor, respectively represent the updated state, the updated action, and the weights of the updated neural network Q-Network; Calculate the current Q-value through the following formula:
[0010] where represents the current Q-value, represents the weights of the neural network Q-Network.
[0011] Optionally, the weights of the target network are determined based on the weights of the neural network Q-Network, and in the initialization stage, the weights of the target network are equal to the weights of the neural network Q-Network.
[0012] Optionally, updating the weights of the neural network Q-Network using the minimized loss function according to the target Q-value and the current Q-value includes: Update the weights of the neural network Q-Network in the following manner:
[0013] Among them, N represents the number of experience samples randomly sampled from the experience replay pool.
[0014] Optionally, the chip to be processed includes: a power semiconductor chip.
[0015] Another embodiment of the present application provides a chip screening system based on deep Q-learning, and the system includes: An acquisition module, configured to acquire n chips to be processed and parameter information of the chips to be processed; A determination module, configured to determine m parallel chips from the n chips to be processed based on the parameter information of the chips to be processed by using a network model based on deep Q-learning.
[0016] Another embodiment of the present application provides a storage medium, in which a computer program is stored, and wherein the computer program is configured to implement the method described in any one of the above when running.
[0017] Another embodiment of the present application provides an electronic device, including a memory and a processor, a computer program is stored in the memory, and the processor is configured to run the computer program to implement the method described in any one of the above.
[0018] Compared with the prior art, the present invention first acquires n chips to be processed and parameter information of the chips to be processed; based on the parameter information of the chips to be processed, m parallel chips are determined from the n chips to be processed by using a network model based on deep Q-learning. It proposes a chip screening strategy based on deep Q-learning, and the screened chips can reduce the parameter differences between parallel devices, effectively reducing the uneven current junction temperature phenomenon caused by chip differences. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a hardware structure block diagram of a computer terminal for a chip screening method based on deep Q-learning provided by an embodiment of the present invention; Figure 2 It is a flow chart of a chip screening method based on deep Q-learning provided by an embodiment of the present invention; Figure 3 It is a flow chart of determining final parallel chips by a network model based on deep Q-learning provided by an embodiment of the present invention; Figure 4 It is a structure diagram of a chip screening system based on deep Q-learning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention, and cannot be construed as a limitation to the present invention.
[0021] An embodiment of the present invention first provides a chip screening method based on deep Q - learning. This method can be applied to electronic devices, such as computer terminals, specifically ordinary computers, tablets, etc.
[0022] The following takes running on a computer terminal as an example to elaborate on it in detail. Figure 1 It is a hardware structure block diagram of a computer terminal for a chip screening method based on deep Q - learning provided by an embodiment of the present invention. As Figure 1 shown, this computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a non - volatile storage medium and an internal memory.
[0023] The non - volatile storage medium can store an operating system and a computer program. This computer program includes program instructions. When the program instructions are executed, the processor can execute any chip screening method based on deep Q - learning.
[0024] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0025] The internal memory provides an environment for the operation of the computer program in the non - volatile storage medium. When this computer program is executed by the processor, the processor can execute any chip screening method based on deep Q - learning.
[0026] This network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 1 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0027] It should be understood that the processor can be a central processing unit (CPU), and this processor can also be other general - purpose processors, digital signal processors (DSPs), application - specific integrated circuits (ASICs), field - programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general - purpose processor can be a microprocessor or this processor can also be any conventional processor, etc.
[0028] See Figure 2 , Figure 2 which is a schematic flowchart of a chip screening method based on deep Q - learning provided by an embodiment of the present invention, and may include the following steps: S201: Obtain n chips to be processed and parameter information of the chips to be processed.
[0029] Specifically, first obtain the number n of chips to be processed. Starting from the first chip to be processed, read the device parameters of the th chip . The screening method in this application mainly solves the problem of uneven current sharing of parallel chips. The parameter information of the chips with more prominent influence may include threshold voltage Vth, transconductance gfs, on - resistance Rds, etc. Other parameters can also be selected according to actual experimental requirements, and the selected parameters can also include any one or more of the above.
[0030] S202: Based on the parameter information of the chips to be processed, use a network model based on deep Q - learning to determine m parallel chips from the n chips to be processed.
[0031] Specifically, see Figure 3 , Figure 3 which is a schematic flowchart of determining the final parallel chips using a network model based on deep Q - learning provided by an embodiment of the present invention. It includes initializing the state s, action a, and the network. Subsequently, calculate the reward function for the state s, perform the action a probabilistically to obtain the updated state s' and store it in the experience replay pool, calculate the target Q - value and the current Q - value; update the weights using the minimized loss function, and then determine whether the stop condition is reached. If so, output the optimal solution; otherwise, re - execute the steps of initializing the state s, action a, and the network and subsequent steps until the stop condition is reached to determine m parallel chips.
[0032] Exemplarily, based on the parameter information of the chips to be processed, using a network model based on deep Q - learning to determine m parallel chips from the n chips to be processed includes: Step 1: Obtain the m - dimensional vector state s and use the greedy strategy to determine the action a; where the greedy strategy includes randomly determining the swap action with a probability of .
[0033] First, randomly select m chips to be processed to form the m - dimensional vector state s, where each dimension of the vector state corresponds to the performance parameter of the chip , where represents the j - th parameter of the i - th chip.
[0034] Among them, the action a can be determined using the greedy strategy, that is, with a probability of Randomly select an unselected chip to be processed to replace one of the currently selected chips; or select the swap operation with the largest Q value predicted by the deep Q-learning network model with a probability of ( ). Here, the Q value is the result of the training of the deep Q-learning network model later, and each experience sample has a corresponding Q value.
[0035] Step 2: Initialize the deep Q-learning-based network model and the experience replay pool; among them, initializing the deep Q-learning-based network model includes initializing the neural network Q-Network and the target network ; Among them, the structure of the deep Q-learning-based network model can be divided into an input layer, which is used to form a matrix of the parameters of the n chips to be processed input; for non-image inputs, the hidden layer can adopt a fully connected layer, and its output layer is the Q value, providing a Q value for each possible action.
[0036] The neural network Q-Network is a method to approximate the Q-value function in Q-learning through a deep neural network. The neural network is used to approximate the Q-value function to avoid storing the Q value in a table. The target network is exactly the same as the neural network Q-Network in terms of the number of layers, the number of neurons in each layer, etc. The experience replay pool is used to store the results of each action for calculating the Q value.
[0037] Step 3: Calculate the reward function value for the state s, execute the action a with probability, and determine the experience sample and store it in the experience replay pool; among them, represents the reward function, represents the updated state; Specifically, the variance between performance parameters is used to measure the performance proximity between parallel chips. The higher the performance proximity, the smaller the variance and the higher the reward; each state s and the next action a can obtain an experience sample and store it in the experience replay pool; every time an interaction is performed (a complete process of state, action, reward, and next state transition).
[0038] In an alternative embodiment, the reward function may satisfy:
[0039] Among them, represents the state corresponding reward value, represents the variance operation, represents the performance parameter of the chip.
[0040] It should be noted that the state is the state vector of the selected m parallel chips, which will be randomly selected in the initial state. Each time, one of the chip parameters within the state s is replaced, and each state is stored in the experience replay pool for the next sampling.
[0041] Step 4: According to the said experience sample , calculate the target Q value and the current Q value; Exemplarily, the calculating the target Q value and the current Q value according to the said experience sample may include: Calculate the target Q value through the following formula:
[0042] where, represents the target Q value, represents the discount factor, respectively represent the updated state, the updated action, and the weights of the updated neural network Q-Network; Calculate the current Q value through the following formula:
[0043] where, represents the current Q value, represents the weights of the neural network Q-Network.
[0044] Among them, the weights of the said target network are determined according to the weights of the neural network Q-Network, and in the initialization stage, the weights of the said target network are equal to the weights of the neural network Q-Network.
[0045] Step 5: According to the target Q value and the current Q value, use the minimization loss function to update the weights of the neural network Q-Network.
[0046] Specifically, the updating the weights of the neural network Q-Network according to the target Q value and the current Q value, using the minimization loss function includes: Update the weights of the neural network Q-Network in the following way:
[0047] where, N represents the number of experience samples randomly sampled from the experience replay pool.
[0048] It should be noted that the weight values of the target network can be set, that is, updated from the neural network Q-Network every certain number of steps to ensure stable training. And the weights can be parameters to be learned, which are random in the initial state.
[0049] Step 6: Repeat the above Steps 1 to 5 until a preset stop condition is reached, so as to output m parallel chips.
[0050] Specifically, execute the above Steps 1 to 5, that is, repeatedly perform the above calculations of state s, action a, reward function, target Q value, current Q value, and weight update, and select an appropriate probability to select the next action a until the stop condition is reached. The stop condition can be reaching the maximum number of iterations, reaching a satisfactory solution, or the Q value converging, etc. When the stop condition is reached, output the last state s, and output m parallel chips.
[0051] Repeat the steps of using the network model based on deep Q-learning to determine m parallel chips from n chips to be processed until the m parallel chips with the optimal performance are finally determined.
[0052] The network model based on deep Q-learning in this application belongs to the reinforcement learning algorithm, approximates the Q value through a neural network, and learns based on environmental interaction and experience replay. The deep Q-learning algorithm is applicable to dynamic decision-making and complex environments, especially applicable to optimization problems with a large state space and continuous actions. And it is applicable to complex dynamic optimization problems, especially policy optimization and continuous decision-making problems. Through a neural network, a non-linear mapping relationship can be approximated, which is suitable for large-scale optimization problems. It can autonomously learn in the environment and is applicable to adaptive optimization scenarios.
[0053] It should be noted that the above chips to be processed may include power semiconductor chips. Power semiconductor chips are the core devices for realizing electric energy conversion and circuit control in the electronic industrial chain, mainly used for functions such as power conversion, amplification, switching, and protection. Their core role is similar to that of a power heart and is widely used in fields such as energy, transportation, industry, and consumer electronics.
[0054] Compared with the prior art, the present invention first obtains n chips to be processed and the parameter information of the chips to be processed; based on the parameter information of the chips to be processed, uses a network model based on deep Q-learning to determine m parallel chips from the n chips to be processed. It proposes a chip screening strategy based on deep Q-learning, and the screened chips can reduce the parameter differences between parallel devices, effectively reducing the uneven current junction temperature phenomenon caused by chip differences.
[0055] Another embodiment of this application provides a chip screening system based on deep Q-learning, as Figure 3Schematic diagram of a chip screening system based on deep Q - learning. The system includes: An acquisition module 401, configured to acquire n chips to be processed and parameter information of the chips to be processed; A determination module 402, configured to determine m parallel chips from the n chips to be processed based on the parameter information of the chips to be processed by using a network model based on deep Q - learning.
[0056] Compared with the prior art, the present invention first acquires n chips to be processed and parameter information of the chips to be processed; based on the parameter information of the chips to be processed, m parallel chips are determined from the n chips to be processed by using a network model based on deep Q - learning. It proposes a chip screening strategy based on deep Q - learning, and the screened chips can reduce the parameter differences between parallel devices, effectively reducing the uneven current - junction temperature phenomenon caused by chip differences.
[0057] An embodiment of the present invention also provides a storage medium, in which a computer program is stored. Wherein, the computer program is configured to implement the steps in the above - mentioned method embodiment when running.
[0058] Specifically, in this embodiment, the above - mentioned storage medium can be configured to store a computer program for performing the following steps: S201: Acquire n chips to be processed and parameter information of the chips to be processed; S202: Based on the parameter information of the chips to be processed, determine m parallel chips from the n chips to be processed by using a network model based on deep Q - learning.
[0059] Specifically, in this embodiment, the above - mentioned storage medium may include, but is not limited to: various media that can store computer programs such as USB flash drives, read - only memories (abbreviated as ROM), random access memories (abbreviated as RAM), mobile hard disks, magnetic disks, or optical discs.
[0060] Compared with the prior art, the present invention first acquires n chips to be processed and parameter information of the chips to be processed; based on the parameter information of the chips to be processed, m parallel chips are determined from the n chips to be processed by using a network model based on deep Q - learning. It proposes a chip screening strategy based on deep Q - learning, and the screened chips can reduce the parameter differences between parallel devices, effectively reducing the uneven current - junction temperature phenomenon caused by chip differences.
[0061] An embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in the above method embodiment.
[0062] Specifically, the above electronic device may further include a transmission device and an input / output device. Among them, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0063] Specifically, in this embodiment, the above processor may be configured to execute the following steps through a computer program: S201: Obtain n chips to be processed and parameter information of the chips to be processed; S202: Based on the parameter information of the chips to be processed, use a network model based on deep Q-learning to determine m parallel chips from the n chips to be processed.
[0064] Compared with the prior art, the present invention first obtains n chips to be processed and parameter information of the chips to be processed; based on the parameter information of the chips to be processed, uses a network model based on deep Q-learning to determine m parallel chips from the n chips to be processed. It proposes a chip screening strategy based on deep Q-learning, and the screened chips can reduce the parameter differences between parallel devices, effectively reducing the phenomenon of uneven current junction temperature caused by chip differences.
[0065] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0066] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0067] In several embodiments provided by the present invention, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above unit division is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0068] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0069] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0070] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in various embodiments of the present invention. And the aforementioned memory includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0071] The embodiments of the present invention have been introduced in detail above. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A chip screening method based on deep Q-learning, characterized in that The method includes: Obtaining n chips to be processed and parameter information of the chips to be processed; Based on the parameter information of the chips to be processed, using a network model based on deep Q-learning, determining m parallel-connected chips from the n chips to be processed.
2. The method according to claim 1, wherein The determining, based on the parameter information of the chips to be processed, m parallel-connected chips from the n chips to be processed by using a network model based on deep Q-learning includes: Step 1: Obtain the m-dimensional vector state s and use the greedy policy to determine the action a; where, the greedy policy includes randomly determining a swap action with probability; Step 2: Initialize the network model based on deep Q-learning and the experience replay pool; among them, initializing the network model based on deep Q-learning includes initializing the neural network Q-Network and the target network ; Step 3: Calculate the reward function value for the state s, and perform the action a with probability to determine the experience sample and store it in the experience replay pool; where represents the reward function, represents the updated state; Step 4: According to the empirical sample , calculate the target Q value and the current Q value; Step 5: Update the weights of the neural network Q-Network using the minimized loss function based on the target Q value and the current Q value Step 6: Repeatedly execute the above steps 1 to 5 until a preset stop condition is reached to output m parallel-connected chips.
3. The method according to claim 2, wherein The reward function satisfies: Among them, represents the state corresponding reward value, represents the variance operation, represents the performance parameters of the chip.
4. The method according to claim 3, wherein Said according to the empirical sample , calculating a target Q value and a current Q value, including: Calculating the target Q value through the following formula: Among them, represents the target Q value, represents the discount factor, respectively represent the updated state, the updated action, and the weights of the updated neural network Q-Network; Calculating the current Q value through the following formula: Among them, represents the current Q value, represents the weights of the neural network Q-Network.
5. The method according to claim 4, characterized in that, The target network weights are determined based on the weights of the neural network Q-Network, and in the initialization phase, the weights of the target network are equal to the weights of the neural network Q-Network.
6. The method according to claim 5, characterized in that, Using the minimized loss function based on the target Q value and the current Q value Update the weights of the neural network Q-Network, including: Updating the weights of the neural network Q-Network in the following manner: where N represents the number of empirical samples randomly sampled from the experience replay pool.
7. The method according to any one of claims 1 to 6, characterized in that The chips to be processed include: power semiconductor chips.
8. A chip screening system based on deep Q-learning, characterized in that, The system includes: An obtaining module, configured to obtain n chips to be processed and parameter information of the chips to be processed; A determining module, configured to determine m parallel-connected chips from the n chips to be processed based on the parameter information of the chips to be processed by using a network model based on deep Q-learning.
9. A storage medium, characterized in that, A computer program is stored in the storage medium, wherein the computer program is configured to implement the method according to any one of claims 1 to 7 when running.
10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and system for screening parallel chips in power semiconductor module
CN107957541A
Display information combination determination method and device, storage medium and electronic equipment
CN111401594A
Batch classification method for semiconductor power chips
CN113298126A
Chip design method and terminal
CN116187265A
Chip screening method and device, computer readable storage medium and computer equipment
CN116727297A