Method for optimizing robustness of microwave device through reinforcement learning of Fourier descriptor reward function

Through Fourier descriptive sub-reward function and deep reinforcement learning framework, the robustness of microwave devices under the differences in manufacturing processes and environment is solved, and the low-cost and high-performance design of a variety of microwave devices is realized, breaking through device types and process limitations.

CN120579417APending Publication Date: 2025-09-02GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510335928.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The robust design methods of existing microwave devices cannot be adapted to different types of devices, and are costly and difficult to maintain stable performance under different manufacturing processes and environments.

Method used

The Fourier descriptor is used as the reward function, combined with the deep reinforcement learning framework, a robust optimization method is designed, and the robustness adjustable optimization of multiple microwave devices is achieved by adjusting the weight of the reward function, breaking through the limitations of manufacturing process differences.

Benefits of technology

It realizes the stable operation of a variety of microwave devices under different manufacturing conditions, reduces costs and improves performance, and has good generalization capabilities, without the need to redesign the reward function for different devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579417A_ABST
    Figure CN120579417A_ABST
Patent Text Reader

Abstract

The invention discloses a method for optimizing the robustness of a microwave device through reinforcement learning of a Fourier descriptor reward function. The method comprises the following steps: firstly, carrying out matrix processing on structural parameters of the microwave device; secondly, establishing an environment based on a deep reinforcement learning framework, adding disturbance to matrix parameters, and inputting the matrix parameters as a state space; thirdly, designing interaction between the intelligent agent and the environment, and designing a reward function in combination with a Fourier descriptor; then, by adjusting the proportion of the reward function, the robustness of different microwave devices can be optimized without redesign; and finally, analyzing and evaluating the robustness of the device by utilizing Monte Carlo (MC). Experiments show that the robustness of the device is adjusted along with the weight change of the reward function. The method has the characteristics of strong generalization ability and wide applicability, breaks through the limitation of the traditional single device design, can effectively cope with the influence of the manufacturing process tolerance and the device type difference on the performance, and provides a brand new solution for the optimization design of low-cost and high-performance microwave devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of microwave device design and relates to a method for optimizing the robustness of microwave devices through reinforcement learning of a Fourier descriptor reward function. Background Art

[0002] Microwave devices are key components in electronic communications, primarily including antennas and filters. With the rapid development of 5G and 6G communication technologies, higher demands are being placed on microwave devices, and requirements for their performance and manufacturing processes are also continuously increasing. However, in the microwave device manufacturing process, uncertainties such as manufacturing tolerances, machine errors, and environmental variations exist, which significantly impact device performance. To address these issues, a common approach is to improve manufacturing processes, but this often leads to significant cost increases. Another effective solution is to design robust microwave devices to ensure stable operation despite these uncertainties. However, robustness design methods for a single device are not adaptable to a wide range of device types. To address this, this paper introduces a Fourier descriptor as a reward function and proposes a reinforcement learning method for optimizing the robustness of microwave devices using a Fourier descriptor reward function. The reward function designed in this paper has good generalization capabilities and can be widely applied to the robustness optimization design of a wide range of microwave devices without having to redesign the reward function for each specific microwave device. This method addresses the impact of manufacturing tolerances and environmental variations on microwave device performance, providing a new solution for low-cost, high-performance microwave devices.

[0003] Currently, research on robust design of microwave devices, both domestically and internationally, primarily focuses on finding robust solutions to a single local optimal solution. While this approach can impart a certain degree of robustness to microwave devices, research on the magnitude and control of robustness in microwave devices remains relatively lacking. Furthermore, the application of existing methods is typically limited to a single microwave device and has not been widely applied to the robustness-adjustable optimization design of multiple microwave devices. However, by regulating the robustness of multiple microwave devices, it is possible to effectively overcome differences in manufacturing process levels and ensure stable operation of microwave devices under various manufacturing conditions. This is of great significance for achieving low-cost, high-performance microwave devices. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for optimizing the robustness of microwave devices through reinforcement learning based on a Fourier descriptor reward function. The reward function designed in the present invention has good generalization ability and can be widely applied to the robustness adjustable optimization design of various microwave devices without the need to redesign the reward function for different microwave devices. It can control and adjust the robustness of various microwave devices to break the limitations of manufacturing process differences.

[0005] The technical solution adopted by the present invention is a method for optimizing the robustness of microwave devices through reinforcement learning using a Fourier descriptor reward function, which specifically includes the following steps:

[0006] Step 1, matrix representation of the structural parameters of the microwave device;

[0007] Step 2: Build a reinforcement learning environment under the deep reinforcement learning framework and use the result obtained in step 1 as the input of the state space after perturbation processing;

[0008] Step 3: Design a corresponding intelligent agent under the deep reinforcement learning framework and enable it to interact with the environment constructed in step 2;

[0009] Step 4: Based on step 3, design the reward function for different devices by combining the Fourier descriptor and ;

[0010] Step 5: Based on step 4, for different microwave devices, there is no need to redesign the reward function, and the reward function can be adjusted. The proportion of ,complete the design of microwave devices with different robustness;

[0011] Step 6: Based on the results obtained in step 5, the robustness of the device is evaluated by Monte Carlo (MC) analysis. The analysis results show that the robustness of the designed device increases with The weight changes.

[0012] The present invention is also characterized in that:

[0013] The specific process of step 1 is:

[0014] Assume that a microwave device contains multiple structural parameters, the selected structural parameters are diverse, and the multiple structural parameters selected and the obtained structural parameters are random and regular, and the performance is greatly different from the standard.

[0015] The specific process of step 2 is:

[0016] A reinforcement learning environment is constructed based on the deep reinforcement learning framework. The environment consists of a state space and perturbations. First, the structural parameters of the microwave device are matrixed and used as the input of the state space. Then, perturbations are added to simulate abnormal conditions. The form of the perturbation is shown in formula (1):

[0017] (1)

[0018] In the formula, (u, σ 2 ), where u is the mean (mathematical expectation) of the Gaussian distribution, σ 2is the variance of the Gaussian distribution, takes the selected structural parameter as the mean, and sets the size of the variance to simulate the disturbance.

[0019] The specific process of step 3 is:

[0020] The intelligent agent is designed within a deep reinforcement learning framework and consists of the following modules: an action space, an experience pool, a neural network, an evaluation network, a target network, and a greedy policy. The action space is defined as the increase or decrease of structural parameters. A prioritized experience replay mechanism is used to update network parameters, and a CNN (convolutional neural network) is chosen as the neural network architecture. The evaluation network and the target network share the same architecture. The evaluation network is used to calculate policy selection and update Q-values, thereby more stably estimating the target Q-value. The sampling method for the experience pool is shown in Equation (2).

[0021] (2)

[0022] Where, is the current state, is the currently selected action, For the current reward, For the next state.

[0023] The convolutional neural network is shown in formula (3):

[0024] (3)

[0025] Where, Indicates that the status Take action The obtained Q value, Indicates the input status Row, No. The eigenvalues ​​of the columns, is a convolution kernel, represents the bias term.

[0026] The calculation of Q value is shown in formula (4):

[0027] (4)

[0028] Where, Indicates that the status The optimal value evaluation function when taking action a is: Represents an operator, indicating that the largest value is selected from multiple possible values. represents the policy, which describes the probability distribution of taking different actions in a given state, It is a weight coefficient between 0 and 1, which is usually used to measure the impact of future rewards. This coefficient is called the discount factor.

[0029] The network update is shown in formula (6):

[0030] (5)

[0031] (6)

[0032] Gradient descent is shown in formula (8):

[0033] (7)

[0034] (8)

[0035] The specific process of step 4 is:

[0036] The reward function is obtained by calculating the Fourier descriptor Distance1 based on the S parameter performance of the microwave device and the target curve .

[0037] Reward Function As shown in formula (12):

[0038] (9)

[0039] (10)

[0040] (11)

[0041] (12)

[0042] The Fourier descriptor D in formula (9) is K The calculation is obtained through Fourier transform, k is the index of the Fourier coefficient, T is the period of the curve, C(t) is the coordinate of the curve, is a complex exponential function (sine and cosine waves represented by complex numbers), and the Fourier descriptor is a combination of k at different frequencies. The Fourier descriptor can convert the curve from the time domain to the frequency domain by discretizing and processing these coefficients. Formula (10) and Formula (11) are related, where the Fourier descriptor can be used to compare the similarity between two curves. Suppose we have two curves and , and their Fourier descriptors are and , the method for calculating the similarity between the two can use Euclidean distance. The specific reward function is as shown in formula (12), where the smaller the distance, the more similar the curves are, and the higher the reward; the larger the distance, the more different the curves are, and the lower the reward.

[0043] Reward Function As shown in formula (13):

[0044] (13)

[0045] Formula (13) is obtained by analyzing the S parameter curve and the Fourier descriptor of the S parameter curve after adding interference. , and calculate its robustness.

[0046] The specific process of step 5 is:

[0047] Add the two rewards together, as shown in formula (14):

[0048] (14)

[0049] Where, is the reward function By adjusting the weight parameters, the intelligent agent is guided to complete the design of microwave devices for different task indicators.

[0050] The specific process of step 6 is:

[0051] According to step 5, the microwave device structural parameters obtained by the designed different tasks are analyzed by MC, and the verification results show that as As the weight of increases, the robustness of the designed microwave device also gradually increases. As shown in formula (15):

[0052] (15)

[0053] Where, Indicates the robustness of the device, It represents the sum of Fourier descriptors after MC analysis. The smaller the value, the stronger the robustness of the device.

[0054] The beneficial effects of the present invention are as follows: first, the structural parameters of the microwave device are matrixed; second, a reinforcement learning environment is established, the matrixed device structural parameters are used as state input, and disturbances are added to the input; then, an intelligent agent that interacts with the environment is designed; and third, a new reward function is designed in combination with the Fourier descriptor. The reward function has good generalization ability and can be widely applied to the robustness adjustable optimization design of various microwave devices without the need to rebuild the reward function for different microwave devices. The task results are verified, and the results show that the robustness of microwave devices designed for different tasks is different. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a schematic flow diagram of the method of the present invention

[0056] Figure 2 This is a structural diagram of the low-pass filter of embodiment 1 of the present invention.

[0057] Figure 3 This is a planar structural diagram of the low-pass filter of embodiment 1 of the present invention.

[0058] Figure 4 This is the robustness verification result (w value) under different tasks of Example 1 of the present invention

[0059] Figure 5 This is a robustness curve diagram of the low-pass filter designed in Example 1 of the present invention.

[0060] Figure 6 This is a structural diagram of the high-pass filter of embodiment 2 of the present invention.

[0061] Figure 7 This is a planar structural diagram of a high-pass filter according to embodiment 2 of the present invention.

[0062] Figure 8 This is the robustness verification result (w value) under different tasks of Example 2 of the present invention

[0063] Figure 9 This is a robustness curve diagram of the high-pass filter designed in Example 2 of the present invention. DETAILED DESCRIPTION

[0064] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0065] This paper proposes a reinforcement learning method for optimizing the robustness of microwave devices using a Fourier descriptor reward function, aiming to address the challenges posed by differences in manufacturing process levels. The reward function designed by this method has good generalization capabilities, eliminating the need to reconstruct the reward function for different microwave devices. This method can be widely applied to the robustness optimization design of various microwave devices, breaking through device type limitations and enabling adjustable optimization of multiple devices. Furthermore, it overcomes manufacturing process limitations, achieving low-cost, high-performance microwave devices and ensuring stable operation under different manufacturing conditions.

[0066] Refer to the figure, Figure 1 A flow chart of a method for optimizing the robustness of microwave devices through reinforcement learning using a Fourier descriptor reward function according to the present invention is presented.

[0067] Example 1

[0068] See also Figure 2 As shown, a low-pass filter is selected to use this method for robust design and control of its robustness, specifically including the following steps:

[0069] Step 1, see Figure 3As shown in Figure 1, the low-pass filter contains multiple structural parameters. The initial parameters of the filter are shown in Table 1. The structural parameters are coordinate-based, and the optimized structural parameters are matrixed for optimization design. It is ensured that the selected structural parameters are random, regular, and have performance that is significantly different from the standard.

[0070] Table 1 Original dimensions of low-pass filter (unit / mm)

[0071]

[0072] Step 2: Add disturbance to the matrixed structural parameter coordinates as input to the state space.

[0073] Step 3: Design the corresponding intelligent agent based on the deep reinforcement learning framework. The intelligent agent is divided into: action space, experience pool, neural network, evaluation network and target network, and greedy strategy.

[0074] Step 4: Convert the device’s performance indicators and robustness factors into corresponding reward functions and The reward function is obtained by calculating the Fourier descriptor Distance1 based on the filter's S parameter performance and the target curve. , calculate the S parameter curve and the change of its Fourier descriptor Distance2 after adding interference, and calculate its robustness-related reward function .

[0075] Step 5: Adjust rewards The proportion of the total reward will be rewarded Different task targets are set to 0%, 5%, 10%, and 15% respectively. Under the guidance of the reward function, the task is completed quickly, and different structural parameters are obtained as shown in Table 2, which shows that the present invention still maintains good performance after optimizing different devices.

[0076] Table 2 Structural parameters obtained by solving different tasks (unit / mm)

[0077]

[0078] Step 6: Verify the robustness of the obtained low-pass filter.

[0079] It should be noted that the present invention verifies robustness by first performing Monte Carlo (MC) analysis on the low-pass filter to obtain sample curves. Then, these sample curves and the S parameter performance obtained from the task are calculated one by one using Fourier descriptors. By comparing the sum of the values ​​of each Fourier descriptor, the smaller the value, the stronger the robustness of the low-pass filter. (See Table 3) It can be clearly concluded from Table 3 that when the reward As the proportion of gradually increases, the mean value of the Fourier descriptor after MC analysis gradually decreases, that is, the robustness of the filter gradually increases. 21 Comparative analysis of the robustness of the filter is shown in Figure 2. Figure 4 .

[0080] Table 3 MC analysis changes of low-pass filters obtained in different tasks

[0081]

[0082] Example 2

[0083] See also Figure 6 As shown, a high-pass filter is selected to use this method for robust design and control of its robustness, specifically including the following steps:

[0084] Step 1, see Figure 7 As shown in Figure 4, the high-pass filter contains multiple structural parameters. The initial parameters of the filter are shown in Table 4. The structural parameters are coordinate-based, and the optimized structural parameters are matrixed for optimization design. It is ensured that the selected structural parameters are random, regular, and have a large difference in performance from the standard.

[0085] Table 4 Original dimensions of high-pass filter (unit / mm)

[0086]

[0087] Step 2: Add disturbance to the matrixed structural parameter coordinates as input to the state space.

[0088] Step 3: Design the corresponding intelligent agent based on the deep reinforcement learning framework. The intelligent agent is divided into: action space, experience pool, neural network, evaluation network and target network, and greedy strategy.

[0089] Step 4: Convert the device’s performance indicators and robustness factors into corresponding reward functions and The reward function is obtained by calculating the Fourier descriptor Distance1 based on the filter's S parameter performance and the target curve. , calculate the S parameter curve and the change of its Fourier descriptor Distance2 after adding interference, and calculate its robustness-related reward function .

[0090] Step 5: Adjust rewards The proportion of the total reward will be rewarded Different task targets are set to 0%, 5%, 10%, and 15% respectively. Under the guidance of the reward function, the task is completed quickly, and different structural parameters are obtained as shown in Table 5, which shows that the present invention still maintains good performance after optimizing different devices.

[0091] Table 5 Structural parameters obtained by solving different tasks (unit / mm)

[0092]

[0093] Step 6: Verify the robustness of the obtained high-pass filter.

[0094] It should be noted that the present invention verifies robustness by first performing Monte Carlo (MC) analysis on the high-pass filter to obtain sample curves. Then, these sample curves and the S-parameter performance obtained from the task are calculated one by one using Fourier descriptors. By comparing the sum of the values ​​of each Fourier descriptor, the smaller the value, the stronger the robustness of the device. (See Table 6) It can be clearly concluded from Table 6 that when the reward As the proportion of gradually increases, the mean value of the Fourier descriptor after MC analysis gradually decreases, that is, the robustness of the filter gradually increases. 21 Comparative analysis of the robustness of the filter is shown in Figure 2. Figure 8 .

[0095] Table 6 MC analysis changes of high-pass filters obtained in different tasks

[0096]

Claims

1. A reinforcement learning method for optimizing the robustness of microwave devices using a Fourier descriptor reward function, characterized in that: The specific steps include: Step 1, matrix representation of the structural parameters of the microwave device; Step 2: Build a reinforcement learning environment under the deep reinforcement learning framework and use the result obtained in step 1 as the input of the state space after perturbation processing; Step 3: Design a corresponding intelligent agent under the deep reinforcement learning framework and enable it to interact with the environment constructed in step 2; Step 4: Based on step 3, design the reward function for different devices by combining the Fourier descriptor and ; Step 5: Based on step 4, for different microwave devices, there is no need to redesign the reward function, and the reward function can be adjusted. The proportion of ,complete the design of microwave devices with different robustness; Step 6: Based on the results obtained in step 5, the robustness of the device is evaluated by Monte Carlo (MC) analysis; The analysis results show that the robustness of the designed device increases with The weight changes.

2. The method for optimizing the robustness of microwave devices through reinforcement learning using a Fourier descriptor reward function according to claim 1, characterized in that: The specific process of step 1 is: Assume that a microwave device contains multiple structural parameters, and the multiple structural parameters are selected and the obtained structural parameters are random and regular, and the performance is greatly different from the standard.

3. The method for optimizing the robustness of microwave devices through reinforcement learning using a Fourier descriptor reward function according to claim 1, characterized in that: The specific process of step 2 is: Based on the deep reinforcement learning framework, a reinforcement learning environment is established. The environment is divided into: state space and disturbance. The structural parameters of the device are matrixed as state space input, and disturbances are added to simulate abnormal conditions. The added disturbance is shown in formula (1): (1) In the formula, record ,in is the mean (mathematical expectation) of the Gaussian distribution, is the variance of the Gaussian distribution, the selected structural parameter is used as the mean, and the size of the variance is set to simulate the disturbance.

4. The method for optimizing the robustness of microwave devices through reinforcement learning using a Fourier descriptor reward function according to claim 1, wherein: The specific process of step 3 is as follows: The intelligent agent is designed based on the deep reinforcement learning framework. The intelligent agent consists of the following modules: action space, experience pool, neural network, evaluation network and target network, and greedy strategy. The increase or decrease of structural parameters is defined as the action space. The priority experience replay mechanism is used to update the network parameters, and CNN (convolutional neural network) is selected as the neural network architecture. The evaluation network and the target network use the same architecture. The evaluation network is used to calculate the strategy selection and update the Q value, so as to estimate the target Q value more stably. The sampling method of the experience pool is shown in formula (2): (2) Where, is the current state, is the currently selected action, For the current reward, For the next state; The convolutional neural network is shown in formula (3): (3) Where, Indicates that the status Take action The obtained Q value, Indicates the input state Row, No. The eigenvalues ​​of the columns, is a convolution kernel, represents the bias term; The calculation of Q value is shown in formula (4): (4) Where, Indicates that the status The optimal value evaluation function when taking action a is: Represents an operator, indicating that the largest value is selected from multiple possible values. represents the policy, which describes the probability distribution of taking different actions in a given state, Is a weight coefficient between 0 and 1, usually used to measure the impact of future rewards. This coefficient is called the discount factor; The network update is shown in formula (6): (5) (6) Gradient descent is shown in formula (8): (7) (8)。 5. The method for optimizing the robustness of microwave devices through reinforcement learning using a Fourier descriptor reward function according to claim 1, wherein: The specific process of step 4 is as follows: Calculate the Fourier descriptor of the S parameter performance curve and target curve of the microwave device Get the reward function ; Reward Function As shown in formula (12): (9) (10) (11) (12) The Fourier descriptor D in formula (9) is K The calculation is obtained through Fourier transform, k is the index of the Fourier coefficient, T is the period of the curve, C(t) is the coordinate of the curve, is a complex exponential function (sine and cosine waves represented by complex numbers), and the Fourier descriptor is a combination of k at different frequencies; the Fourier descriptor can convert the curve from the time domain to the frequency domain by discretizing and processing these coefficients; formula (10) and formula (11) are related, where the Fourier descriptor can be used to compare the similarity between two curves; assuming that we have two curves C1 and C2, whose Fourier descriptors are D1 and D2 respectively, the method for calculating the similarity between the two can use the Euclidean distance; The specific reward function is as shown in formula (12), where the smaller the distance, the more similar the curves are, and the higher the reward; the larger the distance, the more different the curves are, and the lower the reward; Reward Function As shown in formula (13): (13) Formula (13) is obtained by analyzing the S parameter curve and the Fourier descriptor of the S parameter curve after adding interference. , and calculate its robustness.

6. The method for optimizing the robustness of microwave devices through reinforcement learning using a Fourier descriptor reward function according to claim 1, characterized in that: The specific process of step 5 is as follows: Add the two rewards together, as shown in formula (14): (14) Where, is the reward function By adjusting the weight parameters, the intelligent agent is guided to complete the design of microwave devices for different task indicators.

7. The method for optimizing the robustness of microwave devices through reinforcement learning using a Fourier descriptor reward function according to claim 1, characterized in that: The specific process of step 6 is as follows: According to step 5, the microwave device structural parameters obtained by the designed different tasks are analyzed by MC, and the verification results show that as As the weight of increases, the robustness of the designed microwave device also gradually increases; as shown in formula (15): (15) Where, Indicates the robustness of the device, It represents the sum of Fourier descriptors after MC analysis. The smaller the value, the stronger the robustness of the device.