Reinforced learning-driven core particle resource optimization configuration method and system and medium
By constructing a core library and defining the state space, action space, and reward function, and using reinforcement learning algorithms to optimize core configuration, the problem of search space explosion in core resource optimization configuration is solved, thereby improving radar detection and spectrum sensing performance.
Patent Information
- Application Number
- CN202510792803.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-11-04
AI Technical Summary
Existing technologies have failed to effectively apply reinforcement learning algorithms to solve the problem of optimal allocation of core resources, leading to search space explosion and computational challenges when the complexity of the core system increases.
A core library is constructed, defining the state space, action space, and reward function. Core configuration is optimized through reinforcement learning algorithms. Combining the task probability model and the link output signal-to-noise ratio model, an optimization objective function is constructed to achieve the optimal configuration.
By optimizing the configuration of chips in the radio frequency link through reinforcement learning, the target performance of radar detection or spectrum sensing can be improved, thereby increasing detection accuracy and reliability.
Smart Images

Figure CN120892175A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of resource configuration, and particularly relates to a reinforcement learning driven core particle resource optimization configuration method, system and medium. BACKGROUND
[0002] As a new technology, the resource configuration problem of core particle technology is still in its infancy in the academic field. Current research on core particle optimization mainly focuses on physical design, packaging interconnection and interface standards, while theoretical research in the direction of resource allocation, task scheduling and system-level optimization is still limited. The core particle resource configuration problem is essentially a resource allocation optimization problem, which has been extensively studied in the academic field, especially in the field of wireless communication. Spectrum resource allocation and interference resource allocation have formed a relatively mature theoretical system. These studies mainly focus on achieving optimal allocation under limited resource constraints to maximize system efficiency or improve allocation fairness.
[0003] Reinforcement learning is a machine learning method based on the continuous interaction between agents and the environment, aiming to maximize long-term rewards by optimizing strategies to achieve autonomous decision-making. However, existing research has not applied reinforcement learning algorithms to core particle resource optimization configuration problems, so there is an urgent need for a core particle resource optimization configuration method to solve the search space explosion and computational challenge faced by core particle combination as the complexity of core particle systems increases. SUMMARY
[0004] Therefore, it is necessary to propose a reinforcement learning driven core particle resource optimization configuration, system, medium and device to solve the above problems.
[0005] A reinforcement learning driven core particle resource optimization configuration method, the method comprising:
[0006] A core particle library containing core particle types is constructed, the core particle types including filter core particles, low noise amplifier core particles, mixer core particles and variable gain amplifier core particles, each core particle is described by a multi-dimensional parameter matrix of intrinsic characteristic parameters, the intrinsic characteristic parameters including frequency range, gain, noise coefficient and power consumption;
[0007] An evaluation index is determined according to the performance requirements of a target task, and a task probability model is constructed based on the evaluation index, the target task including radar detection or spectrum sensing;
[0008] Based on the superheterodyne receiver framework, a 7-stage radio frequency link model with 7 core particle modules in cascade is constructed according to the selectable core particle types in the core particle library, the 7-stage radio frequency link model is simplified to a 5-stage radio frequency link model through equivalent conversion, and a link output signal-to-noise ratio model is constructed based on the 5-stage radio frequency link model;
[0009] Determine an optimization objective function according to the task probability model and the link output signal-to-noise ratio model, combine the task working frequency constraint and the core particle library resource constraint, and construct a core particle resource optimization configuration model;
[0010] Define a state space, an action space, and a reward function;
[0011] Determine an optimal core particle configuration combination of the core particle resource optimization configuration model based on the state space, the action space, and the reward function.
[0012] The performance requirement of the target task is determined to determine an evaluation index, and a task probability model is constructed based on the evaluation index, the target task includes radar detection or spectrum sensing, and specifically includes:
[0013] If the target task is radar detection, the radar target detection probability is selected as the evaluation index;
[0014] The radar target detection probability is modeled by using a matched filter algorithm and a binary hypothesis test, and a probability model of the radar detection task is constructed, and the task probability model of the radar detection task is:
[0015]
[0016] Wherein, P d is a task probability model, Q(·) and Q -1 (·) are Gaussian Q functions and inverse functions thereof, N is a sampling number, P f is a false alarm probability, and SNR is a link model output signal-to-noise ratio;
[0017] If the target task is spectrum sensing, the spectrum sensing detection probability is selected as the evaluation index;
[0018] The spectrum sensing detection probability is modeled by using an energy detection method, and a probability model of the spectrum sensing task is constructed, and the task probability model of the spectrum sensing task is:
[0019]
[0020] Wherein, P d is a task probability model, Q(·) and Q -1 (·) are Gaussian Q functions and inverse functions thereof, N is a sampling number, P f is a false alarm probability, and SNR is a link model output signal-to-noise ratio.
[0021] The 7-stage radio frequency link model is constructed based on the superheterodyne receiver framework, and the 7-stage radio frequency link model includes, in sequence according to a signal transmission path, a first die, a second die, a third die, a fourth die, a fifth die, a sixth die and a seventh die; wherein the first die is a filter die, the second die is a low-noise amplifier die, the third die is a mixer die, the fourth die is a variable gain amplifier die, the fifth die is a mixer die, the sixth die is a filter die, and the seventh die is a low-noise amplifier die.
[0022] The 7-stage radio frequency link model is constructed based on the superheterodyne receiver framework, and the 7-stage radio frequency link model includes, in sequence according to a signal transmission path, a first die, a second die, a third die, a fourth die, a fifth die, a sixth die and a seventh die; wherein the first die is a filter die, the second die is a low-noise amplifier die, the third die is a mixer die, the fourth die is a variable gain amplifier die, the fifth die is a mixer die, the sixth die is a filter die, and the seventh die is a low-noise amplifier die.
[0023] The noise of the filter die in the 7-stage radio frequency link model is converted to the low-noise amplifier die by equivalent conversion to obtain a 5-stage radio frequency link model, and the 5-stage radio frequency link model includes, in sequence according to a signal transmission path, an equivalent first die, a second die, a third die, a fourth die and a fifth die; wherein the equivalent first die is a low-noise amplifier die, the equivalent second die is a mixer die, the equivalent third die is a variable gain amplifier die, the equivalent fourth die is a mixer die, and the equivalent fifth die is a low-noise amplifier die.
[0024] The total noise coefficient of the 5-stage radio frequency link model is determined.
[0025] The link output signal-to-noise ratio model is constructed according to the signal power and the noise coefficient received by the 5-stage radio frequency link model, and the link output signal-to-noise ratio model is:
[0026]
[0027] wherein P r is the signal power received by the receiving antenna, F is the noise coefficient of the link model, SNR is the output signal-to-noise ratio of the link model, and N0 is the noise power of the link model.
[0028] 4. The core die resource optimization configuration method driven by the reinforcement learning according to claim 3, wherein the link output signal-to-noise ratio model is constructed according to the signal power and the noise coefficient received by the 5-stage radio frequency link model, and the link output signal-to-noise ratio model is:
[0029] If the target task is radar detection, the signal power received by the 5-stage radio frequency link model is a first signal power, and the link output signal-to-noise ratio model of the radar detection is constructed according to the first signal power and the noise coefficient.
[0030] If the target task is spectrum sensing, the signal power received by the 5-level radio frequency link model is a second signal power, and a link output signal-to-noise ratio model for spectrum sensing is constructed according to the second signal power and a noise coefficient.
[0031] The optimization objective function is determined according to the task probability model and the link output signal-to-noise ratio model, and a chip resource optimization configuration model is constructed in combination with a task working frequency constraint and a chip library resource constraint, and specifically includes:
[0032] The link output signal-to-noise ratio model is substituted into the same task probability model of the target task, it is determined that the detection probability and the total noise coefficient are in a linear inverse ratio relationship, and then an optimization objective function is constructed according to the total noise coefficient, and minimizing the total noise coefficient is taken as an optimization objective;
[0033] The chip resource optimization configuration model is constructed according to the optimization objective function, the task working frequency constraint and the chip library resource constraint, and the chip resource optimization configuration model is:
[0034] Min: F
[0035]
[0036] Wherein, F is a noise coefficient of a link model, node i is a chip at an i-th position in the link model, node i .FreRange is a frequency range of the i-th position chip, RangeMin and RangeMax are respectively a minimum value and a maximum value of a working frequency of the task, Node is each chip in the link model, and Arr is a selectable chip resource.
[0037] The state space, the action space and the reward function are defined, and specifically include:
[0038] The state space is determined according to intrinsic characteristic parameters of the chip as a state, and the state space is:
[0039] S t ={Loss1, NF1, G1, NF2, G2, NF3, G3, NF4, G4, Loss2, NF5, G5};
[0040] Wherein, Loss1 is the insertion loss of the first core particle, NF1 is the noise figure of the first core particle, G1 is the gain of the first core particle, NF2 is the noise figure of the second core particle, G2 is the gain of the second core particle, NF3 is the noise figure of the third core particle, G3 is the gain of the third core particle, NF4 is the noise figure of the fourth core particle, G4 is the gain of the fourth core particle, Loss2 is the insertion loss of the second core particle, NF5 is the noise figure of the fifth core particle, and G5 is the gain of the fifth core particle.
[0041] According to the combination of the optional intrinsic feature parameters corresponding to the currently selected core particle as an action, an action space is determined, and the action space is:
[0042]
[0043] Wherein, n t is the number of combinations of optional intrinsic feature parameters corresponding to the core particle selected at time step t.
[0044] The contribution of the device corresponding to the core particle to the noise figure is taken as a reward, and a reward function is determined, and the reward function is:
[0045]
[0046] Wherein, r t is the reward function of time step t, node1Loss is the insertion loss of the first core particle, node2NF is the noise figure of the first core particle, node3NF is the noise figure of the third core particle, node2G is the gain of the second core particle, node j G is the gain of the core particle at the jth position, node t NF is the noise figure of the core particle at the tth position, node6Loss is the insertion loss of the core particle at the sixth position, and node7NF is the noise figure of the core particle at the seventh position.
[0047] Wherein, based on the state space, the action space and the reward function, an optimal core particle configuration combination of the core particle resource optimization configuration model is determined, and specifically includes:
[0048] An experience replay buffer is initialized, and the experience replay buffer is used to store states, actions, rewards and next states;
[0049] An e-greedy strategy is used for action selection, random exploration is performed with a probability of e at the current state, and the current action is selected with a probability of 1-e, to obtain a current reward and a next state, and the current state, action, reward value and next state are stored in the experience replay buffer;
[0050] A deep learning network is set up, data is randomly sampled from an experience replay buffer to train the deep learning network, the trained deep learning network is deployed on a core particle resource optimization configuration model, and when a reward converges to a maximum value, a corresponding current action is an optimal core particle configuration combination.
[0051] A core particle resource optimization configuration system driven by reinforcement learning, the system comprising:
[0052] A core particle library construction module is configured to construct a core particle library containing core particle types, the core particle types including filter core particles, low noise amplifier core particles, mixer core particles and variable gain amplifier core particles, each core particle being described by a multi-dimensional parameter matrix of intrinsic characteristic parameters, the intrinsic characteristic parameters including a frequency range, a gain, a noise figure and a power consumption;
[0053] A task probability model construction module is configured to determine evaluation indexes according to performance requirements of target tasks, and construct a task probability model based on the evaluation indexes, the target tasks including radar detection or spectrum sensing;
[0054] An output signal-to-noise ratio model construction module is configured to construct a 7-stage radio frequency link model of 7 core particle modules in cascade based on a superheterodyne receiver framework according to selectable core particle types in the core particle library, simplify the 7-stage radio frequency link model into a 5-stage radio frequency link model through equivalent conversion, and construct a link output signal-to-noise ratio model based on the 5-stage radio frequency link model;
[0055] A core particle resource optimization configuration model construction is configured to determine an optimization objective function according to the task probability model and the link output signal-to-noise ratio model, and construct a core particle resource optimization configuration model in combination with task working frequency constraints and core particle library resource constraints;
[0056] A definition module is configured to define a state space, an action space and a reward function;
[0057] An optimal core particle configuration combination determination module is configured to determine an optimal core particle configuration combination of the core particle resource optimization configuration model based on the state space, the action space and the reward function.
[0058] A computer readable storage medium storing a computer program, the computer program being executed by a processor to cause the processor to perform the steps of the method as described above.
[0059] A computer device comprising a memory and a processor, the memory storing a computer program, the computer program being executed by the processor to cause the processor to perform the steps of the method as described above.
[0060] The embodiments of the present application have the following beneficial effects:
[0061] The application optimizes the core particle configuration in the radio frequency link through reinforcement learning, thereby improving the target performance of radar detection or spectrum sensing. First, a core particle library containing multiple core particle types is constructed, and each core particle is described in detail by a multi-dimensional parameter matrix, such as frequency range, gain, noise coefficient and power consumption, to provide a basis for optimization configuration. Second, the evaluation index is determined according to the specific task requirement, and the task probability model is established based on the index. Then, based on the superheterodyne receiver framework, seven selectable core particle types are selected to construct a cascaded radio frequency link model, and the complexity of optimization is reduced through a simplified model, so that the output signal-to-noise ratio is easier to analyze. Next, the optimization objective function is constructed by combining the task working frequency and the core particle resource constraint, aiming to make the link output reach the optimal signal-to-noise ratio. By defining the state space, action space and reward function, the reinforcement learning algorithm is used to find the best core particle configuration combination, and finally the optimal resource configuration and performance improvement under specific task conditions are realized, which significantly improves the detection accuracy and reliability of radar and spectrum sensing. BRIEF DESCRIPTION OF DRAWINGS
[0062] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0063] Among them:
[0064] Figure 1 The flowchart of an embodiment of the core particle resource optimization configuration method driven by reinforcement learning provided by the present application is shown.
[0065] Figure 2 The flowchart of another embodiment of the core particle resource optimization configuration method driven by reinforcement learning provided by the present application is shown.
[0066] Figure 3 The structure diagram of the deep Q network provided by the present application is shown.
[0067] Figure 4 The structure diagram of an embodiment of the core particle resource optimization configuration system driven by reinforcement learning provided by the present application is shown.
[0068] Figure 5 The structure diagram of an embodiment of the medium provided by the present application is shown. DETAILED DESCRIPTION
[0069] With reference to the accompanying drawings: the technical solutions in the embodiments of the present application will be described clearly and completely, obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of the present application.
[0070] As shown in Figure 1 , Figure 1 The flowchart of an embodiment of a core particle resource optimization configuration method driven by reinforcement learning provided by the present application. The core particle resource optimization configuration method driven by reinforcement learning, the method comprises:
[0071] S101: a core particle library containing core particle types is constructed, the core particle types include filter core particles, low noise amplifier core particles, mixer core particles and variable gain amplifier core particles, each core particle is described by a multi-dimensional parameter matrix of intrinsic characteristic parameters, the intrinsic characteristic parameters include frequency range, gain, noise coefficient and power consumption.
[0072] Exemplarily, the constructed core particle library contains various core particle types, such as low noise amplifier core particles, variable gain core particles, filter core particles, mixer core particles, etc., a multi-dimensional parameter matrix is established for each core particle category, and through a characteristic parameter orthogonalization design method, configurable core particle options with different intrinsic characteristic parameters (including frequency range, gain, noise coefficient, power consumption, etc.) are provided to meet the differentiated needs of heterogeneous integrated systems for core particle performance.
[0073] S102: determining an evaluation index according to the performance requirement of a target task, and constructing a task probability model based on the evaluation index, the target task includes radar detection or spectrum sensing.
[0074] Exemplarily, two typical application scenarios in the field of wireless communication and signal processing, radar detection and spectrum sensing, are taken as the two types of driving tasks for research. For the radar detection task, the radar target detection probability is selected as the evaluation index, and then the matched filtering algorithm and binary hypothesis testing are used to model the target detection probability, so that the task probability model of the radar detection task can be obtained as:
[0075]
[0076] Where, P d is the task probability model, Q(·) and Q -1 (·) are Gaussian Q function and its inverse function, N is the sampling number, P f is the false alarm probability, and SNR is the output signal-to-noise ratio of the link model.
[0077] For the spectrum sensing task, the detection probability is selected as the evaluation index, the detection probability is modeled based on the energy detection method, and the task probability model of the spectrum sensing task is:
[0078]
[0079] where P d is the task probability model, Q(·) and Q -1 (·) are Gaussian Q function and its inverse function, N is the sampling number, P f is the false alarm probability, and SNR is the output signal-to-noise ratio of the link model.
[0080] S103: Based on the superheterodyne receiver framework, a 7-stage radio frequency link model cascaded by 7 core grain modules is constructed according to the selectable core grain types in the core grain library, the 7-stage radio frequency link model is simplified into a 5-stage radio frequency link model through equivalent conversion, and a link output signal-to-noise ratio model is constructed based on the 5-stage radio frequency link model.
[0081] Illustratively, both radar detection and spectrum sensing select a superheterodyne receiver as the task framework for signal processing. Based on the superheterodyne receiver framework, a 7-stage radio frequency link model cascaded by 7 core grain modules is constructed, and the 7-stage link model includes, in order according to the signal transmission path: a first core grain, a second core grain, a third core grain, a fourth core grain, a fifth core grain, a sixth core grain, and a seventh core grain; wherein the first core grain is a filter core grain, the second core grain is a low-noise amplifier core grain, the third core grain is a mixer core grain, the fourth core grain is a variable gain amplifier core grain, the fifth core grain is a mixer core grain, the sixth core grain is a filter core grain, and the seventh core grain is a low-noise amplifier core grain. A total of m=7 core grains are included, represented as Node=[node1,node2,...,node m ], wherein four types of core grains, filter core grains, low-noise amplifier core grains, mixer core grains, and variable gain amplifier core grains, are included, and for each type of core grain, a plurality of core grains with different intrinsic characteristic parameters are available for selection in the core grain library.
[0082] The filter core grains, low-noise amplifier core grains, mixer core grains, and variable gain amplifier core grains in the core grain library are extracted and denoted as Arr={FLT,LNA,MIX,VGA}.
[0083] FLT={Flt1,Flt2,...,Flt n}
[0084] Flt=<id,FreRange,Loss,Pc>
[0085] Wherein, n is the number of filter chip in the chip library, set filter chip Flt as a triple, Flt.id is the chip number of the filter chip, Flt.FreRange is the frequency range supported by the filter chip, Flt.Loss and Flt.Pc are the insertion loss and power consumption of the filter chip respectively.
[0086] LNA = {Lna1, Lna2,..., Lna n}
[0087] Lna = <id, FreRange, G, NF, Pc>
[0088] Wherein, n is the number of low noise amplifier chip in the chip library, set low noise amplifier chip Lna as a quadruple, Lna.id is the chip number of the low noise amplifier chip, Lna.FreRange is the frequency range supported by the low noise amplifier chip, Lna.G, Lna.NF and Lna.Pc are the gain, noise figure and power consumption of the low noise amplifier chip respectively.
[0089] The mixer chip MIX can be expressed as follows:
[0090] MIX = {Mix1, Mix2,..., Mix n}
[0091] Mix = <id, FreRange, G, NF, Pc>
[0092] Wherein, n is the number of mixer chip in the chip library, Mix.id is the chip number of the mixer chip, Mix.FreRange is the frequency range supported by the mixer chip, Mix.G is the gain of the mixer chip, Mix.G is the noise figure of the mixer chip, Mix.Pc is the power consumption of the mixer chip.
[0093] The variable gain amplifier chip VGA can be expressed as follows:
[0094] VGA = {Vga1, Vga2,..., Vga n}
[0095] Vga = <id, FreRange, G, NF, Pc>
[0096] Wherein, n is the number of variable gain amplifier chip in the chip library, Vga.id is the chip number of the variable gain amplifier chip, Vga.FreRange is the frequency range supported by the variable gain amplifier chip, Vga.G is the gain of the variable gain amplifier chip, Vga.G is the noise figure of the variable gain amplifier chip, Vga.Pc is the power consumption of the variable gain amplifier chip.
[0097] Further, by using the equivalent principle of the radio frequency front-end multi-stage module cascade network, the noise of the filter core particle is converted to the active device behind it, that is, the gain of the active device minus the insertion loss of the filter core particle, the noise figure of the active device assumes the insertion loss of the filter core particle, the noise of the filter core particle in the 7-stage radio frequency link model is converted to the low noise amplifier core particle through equivalent conversion, and a 5-stage radio frequency link model is obtained. The 5-stage radio frequency link model includes, in order according to the signal transmission path: the equivalent first core particle, the second core particle, the third core particle, the fourth core particle and the fifth core particle, wherein the equivalent first core particle is a low noise amplifier core particle, the equivalent second core particle is a mixer core particle, the equivalent third core particle is a variable gain amplifier core particle, the equivalent fourth core particle is a mixer core particle, and the equivalent fifth core particle is a low noise amplifier core particle.
[0098] The total noise figure of the 5-stage radio frequency link model is determined, as shown in the following formula:
[0099]
[0100] Wherein, node2NF is the noise figure of the second core particle, node1Loss is the insertion loss of the first core particle, node2G is the gain of the second core particle, node3NF is the noise figure of the third core particle, node3Loss is the insertion loss of the third core particle, node3G is the gain of the third core particle, node5NF is the noise figure of the fifth core particle, node7NF is the noise figure of the seventh core particle, node6Loss is the insertion loss of the sixth core particle, node4G is the gain of the fourth core particle, node5G is the gain of the fifth core particle, and F is the noise figure of the equivalent link model.
[0101] According to the signal power and the noise figure received by the 5-stage radio frequency link model, a link output signal-to-noise ratio model is constructed, and the link output signal-to-noise ratio model is:
[0102]
[0103] Wherein, P r is the signal power received by the receiving antenna, F is the noise figure of the link model, SNR is the output signal-to-noise ratio of the link model, and N0 is the noise power of the link model.
[0104] The noise power of the link model is:
[0105] N0=KT0B n (F-1)
[0106] Wherein, K is the Boltzmann constant, K=1.38×10 -23 J / K, T0 is the effective noise temperature, the room temperature is 290K, and Bn is the noise figure of the F link model.
[0107] Specifically, the target task is determined to be radar detection, and it is assumed that the transmitting power of the monostatic radar is P t , the transmitting antenna of the radar radiates electromagnetic waves to space using an omnidirectional antenna, and then the first signal power received by the receiving antenna in the 5-level radio frequency link model is:
[0108]
[0109] wherein P r is the signal power received by the receiving antenna, P t is the transmitting power of the monostatic radar, G t is the gain of the transmitting antenna, G r is the gain of the receiving antenna, sigma is the scattering cross-section area of the target, lambda is the signal wavelength, R1 is the distance between the transmitting antenna and the target, and R2 is the distance between the receiving antenna and the target. Generally, the transmitting antenna and the receiving antenna of the monostatic pulse radar are shared, and thus G t =G r =G, R1=R2=R, and then:
[0110]
[0111] Further, the link output signal-to-noise ratio model of radar detection is constructed according to the first signal power and the noise figure.
[0112] The target task is determined to be spectrum sensing, and for the spectrum sensing task, only the free space propagation loss is considered, and then the second signal power received by the receiving antenna of the spectrum sensing in the 5-level radio frequency link model is:
[0113]
[0114] wherein P r is the signal power received by the receiving antenna, P t is the radiation power of the radiation source, G t is the gain of the transmitting antenna, G r is the gain of the receiving antenna, and lambda is the signal wavelength.
[0115] Further, the link output signal-to-noise ratio model of spectrum sensing is constructed according to the second signal power and the noise figure.
[0116] S104: An optimization target function is determined according to the task probability model and the link output signal-to-noise ratio model, and a core particle resource optimization configuration model is constructed in combination with the task working frequency constraint and the core particle library resource constraint.
[0117] Exemplarily, for different tasks, the core particle resources are optimized to realize optimization of task performance, that is, for radar detection and spectrum sensing, the optimization target is to find a group of core particles from the core particle library to maximize the detection probability. The detection probability of radar detection and spectrum sensing increases with the increase of signal-to-noise ratio, so the signal-to-noise ratio is a key factor affecting the detection probability. The link output signal-to-noise ratio model is substituted into the same task probability model of the target task, and it is determined that the detection probability is inversely proportional to the total noise coefficient, and the output signal-to-noise ratio of the link model decreases with the increase of the total noise coefficient. Therefore, the core particle resource optimization configuration problem can be modeled as finding a group of core particles from the core particle library to maximize the output signal-to-noise ratio of the system link model, that is, the total noise coefficient is minimized, and the total noise coefficient is minimized as the optimization target. According to the total noise coefficient, the optimization target function is constructed as follows:
[0118] Min:F
[0119] Wherein, F is the noise coefficient of the link model.
[0120] Further, the constraint conditions of the core particle resource optimization configuration problem are extracted, and the core particle resource optimization configuration model of the core particle resource optimization configuration problem is constructed in combination with the above optimization target function;
[0121] Specifically, since the working frequencies of radar detection and spectrum sensing are 4GHz-6GHz and 0.1GHz-6GHz respectively, the frequency range of each position core particle in the link model must meet the working frequency, that is:
[0122]
[0123] Wherein, node i is the i-th core particle in the link model, node i .FreRange is the frequency range of the i-th core particle, and RangeMin and RangeMax are the minimum and maximum values of the working frequency of the task respectively.
[0124] The core particle resources supported by the core particle library are selected in the optimization process, that is:
[0125]
[0126] Wherein, Node is each core particle in the link model, and Arr is the available core particle resource.
[0127] Therefore, the core particle resource optimization configuration model is as follows:
[0128] Min:F
[0129]
[0130] where F is the noise figure of the link model, node i is the i-th core particle in the link model, node i .FreRange is the frequency range of the i-th core particle, RangeMin and RangeMax are the minimum and maximum values of the operating frequency of the task, Node is each core particle in the link model, and Arr is the selectable core particle resource.
[0131] S105: Define the state space, action space, and reward function.
[0132] Exemplarily, the state vector represents the intrinsic characteristic parameters corresponding to the currently selected core particle, to reflect the historical selection of the entire link. As can be seen from the above, the link contains 5 core particles, and the intrinsic characteristic parameters include gain, noise figure, and insertion loss. Therefore, the definition of the state space is as follows: the initial state S0 is empty, and the state vector is gradually filled as new devices and their parameters are selected at each step. The state space is:
[0133] S t ={Loss1, NF1, G1, NF2, G2, NF3, G3, NF4, G4, Loss2, NF5, G5};
[0134] where Loss1 is the insertion loss of the first core particle, NF1 is the noise figure of the first core particle, G1 is the gain of the first core particle, NF2 is the noise figure of the second core particle, G2 is the gain of the second core particle, NF3 is the noise figure of the third core particle, G3 is the gain of the third core particle, NF4 is the noise figure of the fourth core particle, G4 is the gain of the fourth core particle, Loss2 is the insertion loss of the second core particle, NF5 is the noise figure of the fifth core particle, and G5 is the gain of the fifth core particle.
[0135] According to the combination of the selectable intrinsic characteristic parameters corresponding to the currently selected core particle as the action, the action space is determined
[0136] At each time step t, the agent needs to select a device from the predefined device set and read the corresponding gain, insertion loss, and noise figure from the core particle library. Since the discrete values of the selectable gain, insertion loss, and noise figure of each device are different, the size of the action space at each time step is dynamically changing, depending on the selected device and the discrete value range of the gain and noise figure supported by the device. According to the combination of the selectable intrinsic characteristic parameters corresponding to the currently selected core particle as the action, the action space is determined, and the action space is:
[0137]
[0138] where n tThe number of combinations of optional intrinsic feature parameters corresponding to the core particle selected at time step t. This number depends on the number of core particle device categories in the core particle library currently selected.
[0139] The design basis is the contribution of each device to the noise figure, according to the objective function analysis, if the current state is s t , then the reward function r t at the tth moment is defined as:
[0140]
[0141] Where r t is the reward function at time step t, node1Loss is the insertion loss of the first core particle, node2NF is the noise figure of the first core particle, node3NF is the noise figure of the third core particle, node2G is the gain of the second core particle, node j G is the gain of the jth core particle, node t NF is the noise figure of the core particle at the tth position, node6Loss is the insertion loss of the sixth core particle, and node7NF is the noise figure of the seventh core particle.
[0142] S106: Based on the state space, action space and reward function, determine the optimal core particle configuration combination of the core particle resource optimization configuration model.
[0143] Exemplarily, an experience replay buffer is initialized, which is used to store states, actions, rewards and next states; an e-greedy strategy is used for action selection, random exploration is performed at a probability of e in the current state, and the current action is selected at a probability of 1-e, to obtain the current reward and the next state, and the current state, action, reward value and next state are stored in the experience replay buffer; a deep learning network is set, data is randomly sampled from the experience replay buffer to train the deep learning network, and the trained deep learning network is deployed to the core particle resource optimization configuration model, when the reward converges to the maximum value, the corresponding current action is the optimal core particle configuration combination, so that the detection probability of radar detection and spectrum sensing is maximized.
[0144] It can be known from the above description that the core particle configuration in the radio frequency link is optimized by reinforcement learning, so that the target performance of radar detection or spectrum sensing is improved. First, a core particle library containing multiple core particle types is constructed, and each core particle is described in detail by a multi-dimensional parameter matrix, such as frequency range, gain, noise coefficient and power consumption, to provide a basis for optimization configuration. Secondly, the evaluation index is determined according to the specific task requirement, and the task probability model is established based on the evaluation index. Then, based on the superheterodyne receiver framework, 7 kinds of selectable core particle types are selected to construct a cascaded radio frequency link model, and the complexity of optimization is reduced by simplifying the model, so that the output signal-to-noise ratio is easier to analyze. Then, the optimization objective function is constructed by combining the task working frequency and the core particle resource constraint, aiming to make the link output reach the optimal signal-to-noise ratio. By defining the state space, action space and reward function, the reinforcement learning algorithm is used to find the best core particle configuration combination, and finally the optimal resource configuration and performance improvement under specific task conditions are realized, which significantly improves the detection accuracy and reliability of radar and spectrum sensing.
[0145] As shown in Figure 2 , Figure 2 The flowchart of another embodiment of the core particle resource optimization configuration method driven by reinforcement learning provided by the application is shown. The core particle resource optimization configuration method driven by reinforcement learning comprises the following steps:
[0146] S201: Construct a core particle library containing core particle types, the core particle types including filter core particles, low noise amplifier core particles, mixer core particles and variable gain amplifier core particles, each core particle being described by a multi-dimensional parameter matrix, the intrinsic characteristic parameters including frequency range, gain, noise coefficient and power consumption.
[0147] S202: Determine the evaluation index according to the performance requirement of the target task, and construct the task probability model based on the evaluation index, the target task including radar detection or spectrum sensing.
[0148] S203: Based on the superheterodyne receiver framework, a 7-stage radio frequency link model cascaded by 7 core particle modules is constructed according to the selectable core particle types in the core particle library, the 7-stage radio frequency link model is simplified into a 5-stage radio frequency link model by equivalent conversion, and a link output signal-to-noise ratio model is constructed based on the 5-stage radio frequency link model.
[0149] S204: Determine the optimization objective function according to the task probability model and the link output signal-to-noise ratio model, and construct the core particle resource optimization configuration model by combining the task working frequency constraint and the core particle library resource constraint.
[0150] S205: Define the state space, action space and reward function.
[0151] Exemplarily, the state space is determined according to intrinsic characteristic parameters of the core particles as states, and the state space is:
[0152] S t ={Loss1,NF1,G1,NF2,G2,NF3,G3,NF4,G4,Loss2,NF5,G5};
[0153] Wherein, Loss1 is the insertion loss of the first core particle, NF1 is the noise figure of the first core particle, G1 is the gain of the first core particle, NF2 is the noise figure of the second core particle, G2 is the gain of the second core particle, NF3 is the noise figure of the third core particle, G3 is the gain of the third core particle, NF4 is the noise figure of the fourth core particle, G4 is the gain of the fourth core particle, Loss2 is the insertion loss of the second core particle, NF5 is the noise figure of the fifth core particle, and G5 is the gain of the fifth core particle.
[0154] The action space is determined according to the combination of the optional intrinsic characteristic parameters corresponding to the currently selected core particle as an action, and the action space is:
[0155]
[0156] Wherein, n t is the number of combinations of optional intrinsic characteristic parameters corresponding to the core particle selected at time step t.
[0157] The reward function is determined by taking the contribution of the device corresponding to the core particle to the noise figure as the reward, and the reward function is:
[0158]
[0159] Wherein, r t is the reward function at time step t, node1Loss is the insertion loss of the first core particle, node2NF is the noise figure of the first core particle, node3NF is the noise figure of the third core particle, node2G is the gain of the second core particle, node j G is the gain of the core particle at the jth position, node t NF is the noise figure of the core particle at the tth position, node6Loss is the insertion loss of the core particle at the sixth position, and node7NF is the noise figure of the core particle at the seventh position.
[0160] S206: Initialize the experience replay buffer, and the experience replay buffer is used to store states, actions, rewards and next states.
[0161] Exemplarily, an experience replay buffer D is initialized for storing tuples of state (s), action (a), reward (r) and new state (s'), if the size of D exceeds a preset capacity N, some old experiences are randomly removed. The update formula of the experience replay buffer is: D <- D U {(s t ,a t ,r t ,s t+1 )}.
[0162] S207: An action is selected by using an epsilon-greedy strategy, random exploration is performed with a probability of epsilon in the current state, the current action is selected with a probability of 1-epsilon, a current reward and a next state are obtained, and the current state, action, reward value and next state are stored in the experience replay buffer.
[0163] Exemplarily, an action is selected by using an epsilon-greedy strategy, an action a t is selected in the state s t at time step t, a reward r t and a next state s t+1 are obtained. This strategy balances exploration and experience utilization by performing random exploration with a probability of epsilon and selecting the current optimal strategy with a probability of 1-epsilon, and the formula is expressed as:
[0164] ε = max (ε min , ε max -k*l)
[0165] Wherein, l represents the number of iterations, k represents the decay coefficient, the algorithm can control the reduction rate of exploration behavior by adjusting k, while ensuring that the algorithm still retains a certain exploration probability even at a higher iteration number. In the link building process, the agent selects an action according to the current state and executes it, observes the new state (s') and obtains the reward (r).
[0166] Further, an experience replay mechanism is used to store tuples of state (s), action (a), reward (r) and new state (s') to reduce data correlation and improve learning efficiency. The specific steps are as follows: ① By monitoring the selected core grain and its intrinsic characteristic parameter values (gain, noise figure and insertion loss) at the current position of the link in real time, these parameters are used to construct the current state description. ② Using a deep neural network, the intrinsic characteristic parameter values are adjusted by evaluating possible actions based on the current state to achieve the optimization goal. ③ The effect of the action taken by the agent is evaluated by the reward function r t . ④ The state of the environment is updated by the action of the agent, and the changes in the intrinsic characteristic parameter values of the selected core grain at the current position of the link are monitored, and these updated parameters are used to describe the new state.
[0167] S208: Set up a deep learning network, randomly sample data from the experience replay buffer to train the deep learning network, deploy the trained deep learning network to the chip resource optimization configuration model, and when the reward converges to the maximum value, the corresponding current action is the optimal chip configuration combination.
[0168] Illustratively, two identical deep neural networks are set up, one for generating the current Q value (Q network) and the other for generating the target Q value (target network) to stabilize the training process. A batch of experiences is randomly extracted from the experience replay buffer for training the deep neural network.
[0169] Specifically, a relatively stable Q value estimate is provided by the target network to facilitate the calculation of the target y value.
[0170] ① Calculate the target y value: according to the obtained immediate reward r and the next state s', use the target network to calculate the maximum Q' value, i.e. y = r + γmax a′ Q'(s', a', θ - ). Where γ is the discount factor, which is a key parameter used to balance immediate rewards and future rewards.
[0171] ② Calculate the loss using the mean square error loss function: L(θ) = E[(y t -Q(s t ,a t ; θ)) 2 ]
[0172] Further, gradient descent method is performed to update the Q network parameters of the current step: Update the parameters of the target Q network every C steps: θ - = θ
[0173] If the current step is step S202, step == 2, then pass through the self-attention mechanism: input through the first fully connected layer fc1(s), calculate the query (query), key (key), value (value), and thus calculate the attention weight and weighted average value, output the attention weighted feature, and pass the attention weighted feature to the subsequent network layer.
[0174] Further, adjust the hyperparameters of the deep neural network, including learning rate (α), discount factor (γ), initial exploration rate (ε max ), minimum exploration rate (ε min ) and experience replay buffer size.
[0175] Further, deploy the trained deep learning network to the chip resource optimization configuration model, and when the reward converges to the maximum value, the corresponding current action is the optimal chip configuration combination.
[0176] As can be known from the above description, the present application solves the multi-step decision problem by constructing a deep learning network (deep Q network) as an approximation of the Q function. Each step has an independent Q network and a target network, both of which are composed of three fully connected layers. The deep Q network structure with self-attention mechanism is as shown in Figure 3 . Figure 3 The structure diagram of the deep Q network provided by the present application is shown in step S202, which uses a Q network with a self-attention mechanism, and the ordinary Q network is used in the remaining steps.
[0177] In the training process, the agent stores the state, action, reward and next state tuple using the experience replay mechanism, and trains through batch sampling. In order to optimize the learning process, the algorithm uses the concept of target network, which periodically copies the parameters of the current network to the target network, thereby reducing the volatility in learning. Each time the agent calculates the error between the current Q network and the target network, and optimizes the network parameters through backpropagation to maximize the future cumulative reward. In addition, by adjusting the epsilon value through the epsilon-greedy strategy, the agent gradually reduces the exploration behavior and increases the dependence on the current optimal strategy. In order to improve the learning efficiency, the code also introduces a step-by-step training mechanism, which executes decision-making in each step, and each step of the Q network is independently trained and weighted sampled through the experience pool to ensure that important experiences are paid more attention in the update process. In this way, the deep neural network (deep Q network) can gradually improve the efficiency of resource allocation decision-making, and ultimately achieve efficient performance in complex tasks.
[0178] As shown in Figure 4 , Figure 4 The structure diagram of an embodiment of a core particle resource optimization configuration system driven by the reinforcement learning provided by the present application is shown. The core particle resource optimization configuration system 10 driven by the reinforcement learning includes:
[0179] The core particle library construction module 11 is used to construct a core particle library containing core particle types, and the core particle types include filter core particles, low noise amplifier core particles, mixer core particles and variable gain amplifier core particles. Each core particle is described by a multi-dimensional parameter matrix to describe the intrinsic characteristic parameters, and the intrinsic characteristic parameters include frequency range, gain, noise coefficient and power consumption.
[0180] The task probability model construction module 12 is used to determine the evaluation index according to the performance requirement of the target task, and construct the task probability model based on the evaluation index. The target task includes radar detection or spectrum sensing.
[0181] The output signal-to-noise ratio model construction module 13 is configured to construct a 7-stage radio frequency link model with 7 chip modules cascaded based on a superheterodyne receiver framework according to the selectable chip types in the chip library, simplify the 7-stage radio frequency link model into a 5-stage radio frequency link model through equivalent conversion, and construct a link output signal-to-noise ratio model based on the 5-stage radio frequency link model.
[0182] The chip resource optimal configuration model construction 14 is configured to determine an optimization objective function according to the task probability model and the link output signal-to-noise ratio model, and construct a chip resource optimal configuration model in combination with a task working frequency constraint and a chip library resource constraint.
[0183] The definition module 15 is configured to define a state space, an action space, and a reward function.
[0184] The optimal chip configuration combination determination module 16 is configured to determine an optimal chip configuration combination of the chip resource optimal configuration model based on the state space, the action space, and the reward function.
[0185] Exemplarily, in the chip library construction module 11, a chip library containing chip types is constructed, the chip types including filter chips, low-noise amplifier chips, mixer chips, and variable gain amplifier chips, each chip being described by a multi-dimensional parameter matrix for intrinsic characteristic parameters, the intrinsic characteristic parameters including a frequency range, a gain, a noise figure, and a power consumption. In the task probability model construction module 12, it is determined that the target task is radar detection, and a radar target detection probability is selected as an evaluation index; a matched filter algorithm and a binary hypothesis test are adopted to model the radar target detection probability, and a probability model of the radar detection task is constructed, the task probability model of the radar detection task being:
[0186]
[0187] wherein P d is the task probability model, Q(·) and Q -1 (·) are Gaussian Q functions and their inverse functions, N is a sampling number, P f is a false alarm probability, and SNR is a link model output signal-to-noise ratio.
[0188] It is determined that the target task is spectrum sensing, and a spectrum sensing detection probability is selected as an evaluation index; an energy detection method is adopted to model the spectrum sensing detection probability, and a probability model of the spectrum sensing task is constructed, the task probability model of the spectrum sensing task being:
[0189]
[0190] wherein P d is the task probability model, Q(·) and Q -1 (·) are Gaussian Q functions and their inverse functions, N is a sampling number, P fFalse alarm probability, SNR is the output signal-to-noise ratio of the link model.
[0191] In the output signal-to-noise ratio model construction module 13, a 7-stage radio frequency link model is constructed based on the superheterodyne receiver framework, and the 7-stage link model includes, in order according to the signal transmission path: a first kernel, a second kernel, a third kernel, a fourth kernel, a fifth kernel, a sixth kernel, and a seventh kernel. The first kernel is a filter kernel, the second kernel is a low-noise amplifier kernel, the third kernel is a mixer kernel, the fourth kernel is a variable gain amplifier kernel, the fifth kernel is a mixer kernel, the sixth kernel is a filter kernel, and the seventh kernel is a low-noise amplifier kernel. The noise of the filter kernel in the 7-stage radio frequency link model is converted to the low-noise amplifier kernel through equivalent conversion to obtain a 5-stage radio frequency link model. The 5-stage radio frequency link model includes, in order according to the signal transmission path: an equivalent first kernel, a second kernel, a third kernel, a fourth kernel, and a fifth kernel. The equivalent first kernel is a low-noise amplifier kernel, the equivalent second kernel is a mixer kernel, the equivalent third kernel is a variable gain amplifier kernel, the equivalent fourth kernel is a mixer kernel, and the equivalent fifth kernel is a low-noise amplifier kernel. The total noise coefficient of the 5-stage radio frequency link model is determined. The link output signal-to-noise ratio model is constructed according to the signal power and the noise coefficient received by the 5-stage radio frequency link model, and the link output signal-to-noise ratio model is:
[0192]
[0193] wherein, P r is the signal power received by the receiving antenna, F is the noise coefficient of the link model, SNR is the output signal-to-noise ratio of the link model, and N0 is the noise power of the link model.
[0194] In the kernel resource optimization configuration model construction 14, the link output signal-to-noise ratio model is substituted into the same task probability model of the target task to determine that the detection probability and the total noise coefficient are in a linear inverse proportion relationship. Then, an optimization objective function is constructed according to the total noise coefficient, and the minimization of the total noise coefficient is taken as the optimization objective.
[0195] According to the optimization objective function, the task working frequency constraint, and the kernel library resource constraint, a kernel resource optimization configuration model is constructed, and the kernel resource optimization configuration model is:
[0196] Min: F
[0197]
[0198] wherein, F is the noise coefficient of the link model, node i is the kernel at the i-th position in the link model, node i.FreRange represents the frequency range of the i-th core, RangeMin and RangeMax represent the minimum and maximum operating frequencies of the task, Node represents each core in the link model, and Arr represents the available core resources.
[0199] In definition module 15, the state space, action space, and reward function are defined. In optimal particle configuration combination determination module 16, an experience replay buffer is initialized to store the state, action, reward, and next state. An ε-greedy strategy is used for action selection. In the current state, random exploration is performed with probability ε, and the current action is selected with probability 1-ε to obtain the current reward and the next state. The current state, action, reward value, and next state are stored in the experience replay buffer. A deep learning network is set up, and data is randomly sampled from the experience replay buffer to train the deep learning network. The trained deep learning network is deployed on the particle resource optimization configuration model. When the reward converges to the maximum value, the corresponding current action is the optimal particle configuration combination.
[0200] like Figure 5 As shown, Figure 5 This is a schematic diagram of the structure of an embodiment of the medium provided by the present invention. The medium 20 stores at least one computer program 21, which is executed by a processor to perform the following... Figure 1 and Figure 2 The method shown is detailed above and will not be repeated here. In one embodiment, the storage medium 20 can be a storage chip, hard disk, portable hard disk, USB flash drive, optical disk, or other read / write storage device, or it can be a server, etc.
[0201] Furthermore, the processes depicted in the accompanying drawings do not necessarily have to be performed in the specific or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0202] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer-readable storage media are basically similar to the method embodiments, and therefore described more simply; relevant parts can be referred to the descriptions of the method embodiments.
[0203] The apparatus, device, nonvolatile computer readable storage medium and method provided by the embodiments of the present specification are corresponding, therefore, the apparatus, device, nonvolatile computer storage medium also has similar beneficial technical effects as the corresponding method, since the beneficial technical effects of the method have been described in detail above, therefore, the beneficial technical effects of the corresponding apparatus, device, nonvolatile computer storage medium will not be described here.
[0204] The system, apparatus, module or unit illustrated by the above embodiments can be specifically implemented by a computer chip or entity, or by a product with certain functions. A typical implementation device is a computer. Specifically, the computer may, for example, be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0205] For the convenience of description, the above apparatus is described as various units respectively by functions. Of course, the functions of each unit can be implemented in the same or multiple software and / or hardware in the implementation of the present specification. Those skilled in the art should understand that the embodiments of the present specification can be provided as a method, a system, or a computer program product. Therefore, the embodiments of the present specification can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0206] The present specification is described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to the embodiments of the present specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The apparatus that implements the functions specified in one block or multiple blocks.
[0207] These computer program instructions can also be stored in a computer readable storage medium that can guide the computer or other programmable data processing devices to work in a specific way, so that the instructions stored in the computer readable storage medium produce a manufactured product including instruction apparatus, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1one or more processes and / or blocks Figure 1 the function(s) specified in the block or blocks.
[0208] These computer program instructions can also be loaded into computer or other programmable data processing devices to cause a series of operational steps to be performed on the computer or other programmable devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable devices provide steps for implementing the functions of the flow Figure 1 one or more processes and / or blocks Figure 1 the function(s) specified in the block or blocks.
[0209] In one typical configuration, the computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0210] The memory can include non-persistent memory and / or volatile memory, such as random access memory (RAM) about which the computer stores information about the operating environment. This memory is an example of computer readable media. The memory can also include non-volatile memory, such as read only memory (ROM), electrically programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), flash memory, or other memory technologies, CD-ROM, digital versatile disks (DVD), or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information for access by a computer.
[0211] Computer readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for storage of information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically programmable read only memory (EEPROM), flash memory or other memory technologies, compact disc read only memory (CD-ROM), digital versatile disks (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. According to the definition herein, computer readable media does not include transitory media, such as modulated data signals and carrier waves.
[0212] It should also be noted that the terms "comprising", "containing", or any other variant thereof, are intended to encompass a non-exclusive inclusion, such that a process, method, article or apparatus that comprises a list of elements does not include only those elements recited, but can also include other elements not expressly listed or inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.
[0213] The specification can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in local and remote computer storage media including memory storage devices.
[0214] The various embodiments in the specification are described in progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, the system embodiments are described simply because they are basically similar to the method embodiments, and the relevant parts can be referred to the description of the method embodiments.
[0215] The above only describes the preferred embodiments of the present application, and of course cannot limit the scope of the present application. Any equivalent changes made according to the claims of the present application are still within the scope of the present application.
Claims
1. A reinforcement learning-driven method for optimizing the allocation of chip resources, characterized in that, The method includes: A chip library containing chip types is constructed. The chip types include filter chips, low-noise amplifier chips, mixer chips, and variable gain amplifier chips. Each chip is described by an intrinsic characteristic parameter matrix, which includes frequency range, gain, noise figure, and power consumption. Evaluation metrics are determined based on the performance requirements of the target task, and a task probability model is constructed based on the evaluation metrics. The target task includes radar detection or spectrum sensing. Based on the superheterodyne receiver framework, a 7-level RF link model consisting of 7 cascaded RF modules is constructed according to the selectable RF types in the RF module library. The 7-level RF link model is simplified to a 5-level RF link model through equivalent transformation, and a link output signal-to-noise ratio model is constructed based on the 5-level RF link model. The objective function is determined based on the task probability model and the link output signal-to-noise ratio model. The core resource optimization configuration model is then constructed by combining the task operating frequency constraint and the core resource library constraint. Define the state space, action space, and reward function; Based on the state space, action space, and reward function, the optimal core configuration combination of the core resource optimization configuration model is determined.
2. The reinforcement learning-driven chip resource optimization allocation method according to claim 1, characterized in that, The process involves determining evaluation metrics based on the performance requirements of the target task and constructing a task probability model based on these metrics. The target task includes radar detection or spectrum sensing, specifically including: If the target task is determined to be radar detection, then the radar target detection probability is selected as the evaluation index. The probability of radar target detection is modeled using a matched filtering algorithm and a binary hypothesis test, thus constructing a probabilistic model for the radar detection task. The probabilistic model for the radar detection task is as follows: Among them, P d For the task probability model, Q(·) and Q -1 (·) represents the Gaussian Q-function and its inverse function, N is the number of samples, and P is the inverse function. f , where SNR is the false alarm probability and SNR is the signal-to-noise ratio output by the link model. If the target task is determined to be spectrum sensing, then the spectrum sensing detection probability is selected as the evaluation index. The probability of spectrum sensing detection is modeled using the energy detection method, and a probability model for the spectrum sensing task is constructed. The probability model for the spectrum sensing task is as follows: Among them, P d For the task probability model, Q(·) and Q -1 (·) represents the Gaussian Q-function and its inverse function, N is the number of samples, and P is the inverse function. f is the false alarm probability, and SNR is the signal-to-noise ratio output by the link model.
3. The reinforcement learning-driven chip resource optimization allocation method according to claim 2, characterized in that, The superheterodyne receiver framework constructs a 7-level RF link model consisting of 7 cascaded RF modules based on selectable RF module types from the RF module library. This 7-level RF link model is simplified to a 5-level RF link model through equivalent transformation. A link output signal-to-noise ratio (SNR) model is then constructed based on this 5-level RF link model, specifically including: A 7-level RF link model is constructed based on a superheterodyne receiver framework, consisting of 7 cascaded core modules. The 7-level link model includes, in sequence according to the signal transmission path: first core, second core, third core, fourth core, fifth core, sixth core, and seventh core; wherein, the first core is a filter core, the second core is a low-noise amplifier core, the third core is a mixer core, the fourth core is a variable gain amplifier core, the fifth core is a mixer core, the sixth core is a filter core, and the seventh core is a low-noise amplifier core. By converting the noise of the filter core in the 7-level RF link model to the low-noise amplifier core through equivalent transformation, a 5-level RF link model is obtained. The 5-level RF link model includes, in sequence according to the signal transmission path, the equivalent first core, second core, third core, fourth core, and fifth core. The equivalent first core is a low-noise amplifier core, the equivalent second core is a mixer core, the equivalent third core is a variable gain amplifier core, the equivalent fourth core is a mixer core, and the equivalent fifth core is a low-noise amplifier core. Determine the total noise figure of the 5-level RF link model; Based on the signal power and noise figure received from the 5-level RF link model, a link output signal-to-noise ratio (SNR) model is constructed. The link output SNR model is as follows: Among them, P r F is the signal power received by the receiving antenna, SNR is the noise figure of the link model, N0 is the output signal-to-noise ratio of the link model, and N0 is the noise power of the link model.
4. The reinforcement learning-driven chip resource optimization allocation method according to claim 3, characterized in that, The step of constructing the link output signal-to-noise ratio model based on the signal power and noise figure received from the 5-level RF link model specifically includes: If the target task is determined to be radar detection, then the signal power received by the 5-level radio frequency link model is the first signal power, and a link output signal-to-noise ratio model for radar detection is constructed based on the first signal power and the noise figure. If the target task is determined to be spectrum sensing, then the signal power received by the 5-level RF link model is the second signal power, and a spectrum sensing link output signal-to-noise ratio model is constructed based on the second signal power and the noise figure.
5. The reinforcement learning-driven chip resource optimization allocation method according to claim 4, characterized in that, The step involves determining the optimization objective function based on the task probability model and the link output signal-to-noise ratio model, and constructing a core resource optimization configuration model by combining task operating frequency constraints and core resource constraints. Specifically, this includes: Substituting the link output signal-to-noise ratio model into the task probability model with the same target task, it is determined that the detection probability and the total noise coefficient are linearly inversely proportional. Then, an optimization objective function is constructed based on the total noise coefficient, with minimizing the total noise coefficient as the optimization objective. Based on the aforementioned objective function, task operating frequency constraints, and core resource constraints, a core resource optimization configuration model is constructed. The core resource optimization configuration model is as follows: M in:F Where F is the noise figure of the link model, and node i For the core at position i in the link model, node i .FreRange represents the frequency range of the i-th core, RangeMin and RangeMax represent the minimum and maximum operating frequencies of the task, Node represents each core in the link model, and Arr represents the available core resources.
6. The reinforcement learning-driven chip resource optimization allocation method according to claim 5, characterized in that, The definition of the state space, action space, and reward function specifically includes: The state space is determined based on the intrinsic characteristic parameters of the core particles as states, and the state space is as follows: S t ={Loss1,NF1,G1,NF2,G2,NF3,G3,NF4,G4,Loss2,NF5,G5}; Where Loss1 is the insertion loss of the first core, NF1 is the noise figure of the first core, G1 is the gain of the first core, NF2 is the noise figure of the second core, G2 is the gain of the second core, NF3 is the noise figure of the third core, G3 is the gain of the third core, NF4 is the noise figure of the fourth core, G4 is the gain of the fourth core, Loss2 is the insertion loss of the second core, NF5 is the noise figure of the fifth core, and G5 is the gain of the fifth core. The action space is determined based on the combination of selectable intrinsic characteristic parameters corresponding to the currently selected core particle as the action, wherein the action space is: Where, n t The number of possible combinations of intrinsic characteristic parameters corresponding to the core selected at time step t; The reward function is determined by using the contribution of the corresponding device to the noise figure as the reward. The reward function is as follows: Where, r t Let node1Loss be the reward function at time step t, node2NF be the insertion loss of the first core, node3NF be the noise figure of the first core, node2G be the noise figure of the third core, and node2G be the gain of the second core. j G represents the gain of the core at position j in the link model, node t NF is the noise figure of the core at position t, node6Loss is the insertion loss of the core at position 6, and node7NF is the noise figure of the core at position 7.
7. The reinforcement learning-driven chip resource optimization allocation method according to claim 6, characterized in that, Based on the state space, action space, and reward function, the optimal core configuration combination of the core resource optimization configuration model is determined, specifically including: Initialize the experience replay buffer, which is used to store the state, action, reward, and next state; An ε-greedy strategy is adopted for action selection. In the current state, random exploration is performed with probability ε, and the current action is selected with probability 1-ε to obtain the current reward and the next state. The current state, action, reward value and the next state are stored in the experience replay buffer. Set up a deep learning network, randomly sample data from the experience replay buffer to train the deep learning network, and deploy the trained deep learning network onto the core resource optimization configuration model. When the reward converges to the maximum value, the corresponding current action is the optimal core configuration combination.
8. A reinforcement learning-driven chip resource optimization allocation system, characterized in that, The system includes: The chip library construction module is used to build a chip library containing chip types, including filter chips, low-noise amplifier chips, mixer chips, and variable gain amplifier chips. Each chip is described by a multi-dimensional parameter matrix, which includes intrinsic characteristic parameters such as frequency range, gain, noise figure, and power consumption. The task probability model construction module is used to determine the evaluation index according to the performance requirements of the target task, and to construct the task probability model based on the evaluation index. The target task includes radar detection or spectrum sensing. The output signal-to-noise ratio model construction module is used to construct a 7-level RF link model consisting of 7 cascaded RF modules based on the superheterodyne receiver framework and the selectable RF module types in the RF module library. The 7-level RF link model is simplified to a 5-level RF link model through equivalent transformation, and the link output signal-to-noise ratio model is constructed based on the 5-level RF link model. The core resource optimization configuration model is constructed to determine the optimization objective function based on the task probability model and the link output signal-to-noise ratio model, and to construct the core resource optimization configuration model in combination with the task operating frequency constraint and the core resource library constraint. The definition module is used to define the state space, action space, and reward function. The optimal core configuration combination determination module is used to determine the optimal core configuration combination of the core resource optimization configuration model based on the state space, action space, and reward function.
9. A computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the method as claimed in any one of claims 1 to 7.
10. A computer device comprising a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method as claimed in any one of claims 1 to 7.