A low cross-polarization antenna designed based on deep learning
By optimizing the design using a H-shaped coupling slot and a vertical metal support differential feeding structure, along with a CNN-Transformer hybrid neural network model, the problem of low efficiency in traditional millimeter-wave antenna design was solved. This resulted in an antenna with low cross-polarization and stable gain, improving design efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUNAN NORMAL UNIVERSITY
- Filing Date
- 2026-02-05
- Publication Date
- 2026-05-26
AI Technical Summary
Traditional millimeter-wave antenna designs are inefficient, difficult to optimize globally, and differential antenna structures are incompatible with monolithic microwave integrated circuits, resulting in poor cross-polarization suppression.
A differential feeding structure with an H-shaped coupling slot and a vertical metal support is adopted. By combining the CNN-Transformer hybrid neural network model and the simulated annealing algorithm, the antenna design is optimized to achieve low cross-polarization and stable gain.
It achieves antenna design with low cross-polarization, wide impedance bandwidth and stable gain, improving design efficiency and accuracy and shortening the design cycle.
Smart Images

Figure CN121642533B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of 5G / 6G millimeter-wave communication and artificial intelligence, specifically to a low cross-polarization antenna based on deep learning-optimized design. Background Technology
[0002] Millimeter-wave technology is key to achieving ultra-high data rates in future mobile communications. Antennas, as core components of millimeter-wave systems, face multiple stringent requirements, including miniaturization, high gain, wide bandwidth, and low cross-polarization. However, traditional single-ended antennas cannot be directly connected to differential circuits, often requiring an external single-ended to differential converter, leading to additional losses, increased circuit size, and reduced antenna impedance bandwidth. Differential antennas, due to their excellent common-mode rejection capability, effectively improve system anti-interference performance and can be directly integrated with differential circuits, making them a better choice. However, traditional differential antennas suffer from a relatively large overall size, making them difficult to integrate with monolithic microwave integrated circuits. Therefore, how to integrate the differential antenna structure is one of the current technical challenges.
[0003] Traditional antenna design methods heavily rely on engineers' experience and parameter scanning and optimization using full-wave electromagnetic simulation software. Faced with complex antenna structures and interdependent performance indicators, this trial-and-error, manual iterative design approach is inefficient, time-consuming, and struggles to find a globally optimal solution. In recent years, deep learning technology has offered new insights for intelligent antenna design. Commonly used convolutional neural network (CNN) models excel at capturing local features, but they are insufficient at modeling the complex global coupling relationships between multiple antenna structural parameters; while Transformer models can establish global dependencies, they are insensitive to local structural details. Therefore, developing an intelligent optimization design method that can integrate local and global information and perform efficient collaborative prediction and optimization of multiple performance indicators has significant theoretical and engineering value.
[0004] To address the challenges of suppressing cross-polarization, low integration, and the inefficiency and difficulty in global optimization of traditional millimeter-wave antenna designs, this invention proposes a low-cross-polarization antenna based on deep learning-based optimization design. This antenna employs a differential feed structure with an H-shaped coupling slot and vertical metal struts, effectively suppressing cross-polarization without requiring an external single-ended to differential converter. It extends the impedance bandwidth and stabilizes the gain through a double-T-shaped matched short-circuit. The antenna achieves miniaturization through an embedded vertically stacked differential excitation architecture. A CNN-Transformer hybrid neural network model is constructed for optimization design, and combined with simulated annealing (SA) algorithm for optimization training, enabling rapid co-prediction and reverse design of S11 parameters and gain indices, thus improving design efficiency and accuracy. Summary of the Invention
[0005] The purpose of this invention is to overcome the above-mentioned problems and provide a low cross-polarization antenna based on deep learning optimization design. This antenna has the advantages of low cross-polarization, stable gain and miniaturization.
[0006] The objective of this invention is achieved through the following technical solution: a low cross-polarization antenna based on deep learning optimization design.
[0007] A low cross-polarization antenna based on deep learning optimization design includes, from top to bottom, a main radiating element, a first dielectric layer, a bonding and curing layer, a coupling ground metal layer, a second dielectric layer, a feed microstrip line, and a vertical metal support; the main radiating element is disposed on the top surface of the first dielectric layer; the coupling ground metal layer is disposed on the top surface of the second dielectric layer; and the feed microstrip line is disposed on the bottom surface of the second dielectric layer.
[0008] Furthermore, the first dielectric layer and the second dielectric layer are both made of RO3003, with a relative permittivity of 3 and a loss tangent of 0.0013. The adhesive curing layer is made of Rogers 4450F, with a relative permittivity of 3.5 and a loss tangent of 0.004.
[0009] Furthermore, the main radiating element includes a cross-shaped groove, a rectangular matching band, and a double-T-shaped matching short circuit. The cross-shaped groove is formed by the perpendicular intersection of a transverse groove and a longitudinal groove. The rectangular matching band is located within the transverse groove of the cross-shaped groove. The double-T-shaped matching short circuit is formed by two identical T-shaped units symmetrically arranged about the rectangular matching band. Each T-shaped unit includes a longitudinal arm and a transverse arm. The coupling grounding metal layer includes an H-shaped coupling slot, which is composed of a transverse main slot and two vertical secondary slots. The feed microstrip line includes a T-shaped equal power divider, which includes an input main branch, two symmetrical output branches, and two matching rectangular blocks. There are two vertical metal pillars, symmetrically distributed about the antenna's geometric center, penetrating the first dielectric layer and the bonding and curing layer. The upper end connects to the double-T-shaped matching short circuit, and the lower end connects to the coupling grounding metal layer, forming a differential feed path. The electrical length of the vertical metal pillar at the 29GHz frequency point is approximately one-quarter of the dielectric wavelength. This design can offset the impedance change caused by the interface between the dielectric layer and the bonding and curing layer during coupling, maximize coupling efficiency, and ensure stable transmission of differential signals.
[0010] Furthermore, the H-shaped coupling groove and the vertical metal support together constitute an embedded differential excitation architecture, matching the differential operating mode of the double T-shaped matching short path.
[0011] Furthermore, with the antenna's geometric center as the central axis, each functional layer and its internal structure are precisely aligned and symmetrically distributed based on this central axis, forming an integrated stacked structure with vertical alignment between layers and in-plane mirror symmetry.
[0012] Furthermore, the structural dimensions of the antenna can be defined by a set of geometric parameters, including but not limited to: the thickness of the first dielectric layer H1, the thickness of the second dielectric layer H2, the thickness of the bonding and curing layer H3; the side length Lsp of the main radiating element; the length L and width W of the transverse slot of the cross groove, and the width W2 of the longitudinal slot; the length L3 and width W3 of the rectangular matching band, and the distance Gap between the rectangular matching band and the transverse slot of the cross groove; the length T and width W1 of the transverse arm of the T-shaped unit of the double T-shaped matching short circuit, the distance Gin between the transverse arm and the rectangular matching band, the distance Gout between the transverse arm and the transverse slot of the cross groove, and the distance T of the double T-shaped matching short circuit. The longitudinal arm length L5 and width S of the shaped unit, the distance lv between the longitudinal arm and the adjacent side of the main radiating element, and the distance G between the longitudinal arm and the longitudinal slot of the cross groove; the diameter dv of the vertical metal pillar and the distance Sv between the two vertical metal pillars; the width Ws of the main slot of the H-shaped coupling slot on the coupling grounding metal layer, the length Ls of the main slot, the width Wi of the secondary slot, the length Li of the half arm of the secondary slot, and the distance Si between the secondary slot and the starting end of the transverse main slot; the length Lf1 and width Wf of the input main branch of the T-type equal power divider in the feeding microstrip line, the length L4 and width Wf1 of the output branch, the distance Wd between the two output branches, and the length Lt and width Wt of the matching rectangular block.
[0013] A low cross-polarization antenna based on deep learning optimization design is proposed. The optimization design includes the following four steps: parametric modeling and dataset construction, construction and training of CNN-Transformer hybrid neural network model, multi-target screening reverse design, and antenna implementation and verification.
[0014] Step 1: Based on the analysis of differential feeding, slot coupling, and multimode resonance mechanism of the low cross-polarization antenna, six key geometric parameters in antenna design are selected to balance model complexity and prediction accuracy. As design variables for the neural network, Ws is the width of the main slot of the H-shaped coupling slot, Ls is the length of the main slot of the H-shaped coupling slot, Wi is the width of the secondary slot of the H-shaped coupling slot, Li is the half-arm length of the secondary slot of the H-shaped coupling slot, S is the longitudinal arm width of the T-shaped unit of the double T-shaped matching short path, and lv is the distance between the longitudinal arm of the T-shaped unit of the double T-shaped matching short path and the adjacent side of the main radiating element. The four parameters of the H-shaped coupling slot mainly affect the total coupling, resonant point, and differential signal generation, while the two parameters of the double T-shaped matching short path mainly affect the low... The parameters were analyzed to determine the resonant point, impedance matching, associated current distribution, and coupling strength. Within the respective sampling ranges of the six parameters, five discrete values for each parameter were combined and sampled, with sampling ranges of {0.2-0.36}, {3.4-4.6}, {0.15-0.3}, {0.8-1.1}, {0.8-1.0}, and {0.2-0.5}, all in millimeters. The parameter space was then scanned using full-wave electromagnetic simulation software to generate the corresponding output response. A 401-dimensional S11 amplitude vector was sampled within the 24-34 GHz frequency band. Gain value vectors at 12 discrete frequency points ,in The response data for parameter S11, It is a 401-dimensional real vector space. The response data is for the gain parameter. It is a 12-dimensional real vector space; all input and output data are processed by max-min normalization and randomly divided into training set, validation set and test set in a ratio of 8:1:1 to form the basic data pairs for model learning.
[0015] Step 2: Build and train the CNN-Transformer hybrid neural network model, and use the simulated annealing algorithm to optimize the model training process.
[0016] The CNN branch consists of two cascaded one-dimensional convolutional modules. Each module performs convolution, layer normalization, ReLU activation, and max pooling operations, specifically designed to capture local, short-range dependencies between input parameters. The specific steps are as follows: The input is a parameter vector with batch size B and dimension 6. ,in Given a B×6 dimensional real vector space, the input is first reshaped into a sequence of vectors. Where 1 represents the added channel dimension. The vector is a B×1×6 dimensional real number vector, which is then processed through two layers of one-dimensional convolution. The first convolution layer uses 32 convolution kernels with a width of 3. The kernels slide on the sequence, focusing on three consecutive parameters at a time to capture their local dependencies. The formula for the first convolution layer is: ,in This is the output of the first layer of one-dimensional convolution. It is the ReLU activation function. Represents one-dimensional convolution. These are the parameters of the first convolutional kernel, and x is the input sequence-like vector. The first layer is the bias term; the second convolutional layer uses 64 convolutional kernels, which are the feature maps output from the previous layer. Further learning leads to more abstract feature combinations, resulting in a 64-dimensional feature map. The formula for the second convolution layer is: ,in This is the output of the second layer of one-dimensional convolution. It is the ReLU activation function. Represents one-dimensional convolution. This is the output of the first layer of one-dimensional convolution. These are the parameters of the second convolutional kernel. This is the second layer bias term; finally, through global max pooling, the final output is a local feature vector: ,in This is the output of the second convolutional layer. `dim` represents the sequence dimension of this operation, taking the maximum value at `dim=2`. It aggregates the features from the 6 positions of each sample into a single 64-dimensional vector. Thus, features characterizing the local structure are obtained.
[0017] The Transformer branch is mainly used to extract the global dependencies between antenna parameters. Its core components include a linear embedding layer, position encoding, a three-layer Transformer encoder, and an attention pooling layer. Each Transformer encoder layer includes a multi-head self-attention mechanism, a feedforward neural network, residual connections, and layer normalization. The specific steps are as follows: Input parameter vector ,in 6 represents the batch size, and 6 represents the number of parameters. Given a B×6 dimensional real vector space, through a learnable linear transformation matrix... Map it to a high-dimensional embedding vector: ,in The embedded feature vector is then followed by fixed sine-cosine position encoding. To preserve parameter order information: ,in For the high-dimensional input matrix after adding position encoding, For fixed-position encoding matrix; subsequently, After deep interaction through a three-layer Transformer encoder, each layer first performs multi-head self-attention calculation, dividing the 64-dimensional features into two 32-dimensional heads for parallel execution, and then inputting a high-dimensional matrix. The query, key, and value matrix is obtained through linear projection: ,in, For the attention head number, Number the Transformer encoder level. For the first The input matrix of the layer encoder, Corresponding high-dimensional input matrix The query, key, and value matrix is obtained through linear projection. Corresponding to the Learnable weights for each attention head; then perform scaled dot product attention calculation so that each parameter can attend to all other parameters, as shown in the formula: ,in This is an attention score matrix, representing the similarity between the query at each position and the keys at all positions. Scaling factor For each head dimension, , This represents the normalization function along the last dimension. For value matrices, The output of the i-th head is then concatenated and linearly transformed to achieve cross-head feature interaction and fusion. The transformation formula is as follows: ,in This represents the output of concatenating the two heads along the feature dimension. To output the linear transformation weight matrix, This represents the multi-head attention output; after layer normalization and residual connections, the self-attention output is input into the feedforward network: ,in To activate the function, a nonlinearity is introduced. The input features are after passing through the attention layer and the normalization layer. For the first and second layer weights, For the first and second layers of offset, For feedforward neural networks to handle input The transformation result is enhanced nonlinearly through a feedforward network, and training stability is ensured through residual connections and layer normalization; after passing through a three-layer encoder, the output is... The attention weights are converged into a global feature vector through learnable attention weights: The final output is a 64-dimensional global feature. The complex global relationships between fully encoded parameters, among which For positional ordinal numbers, For the first Features of each location For the first The attention weights at each position are calculated using the formula: , For learnable query vectors, For summation index variables.
[0018] The 64-dimensional local feature vector output by the CNN branch The 64-dimensional global feature vector output by the Transformer branch The vectors are directly concatenated to form a 128-dimensional fused feature vector. This fused feature vector is then fed into two independent fully connected network prediction heads. One prediction head outputs a 401-dimensional vector, corresponding to the normalized S11 curve, while the other prediction head outputs a 12-dimensional vector, corresponding to the normalized discrete frequency gain value.
[0019] The model training uses a smooth L1 loss function, jointly optimizing the S11 parameters and gain outputs. The loss function is: ,in, Weights for gain prediction loss. To predict the loss weights for parameters S11, set α=0.7 and β=0.3. The total loss of the model is N, where N is the number of training samples. To smooth the L1 loss function, The sample number. For the first The true gain value of each sample. For the first The predicted gain value for each sample. For the first The true values of the S11 parameters for each sample. For the first Predicted S11 parameter values for each sample.
[0020] Furthermore, during the training initialization phase, a simulated annealing algorithm is introduced to adaptively optimize the initial learning rate globally, and the algorithm sets a temperature decay plan. The initial temperature is The attenuation coefficient is The termination temperature is ,in The current temperature. For the updated temperature; at each temperature Below, based on the current optimal learning rate Generate candidate learning rates : ,in It follows a standard normal distribution; a candidate learning rate is used to quickly train for one epoch, the average loss on the training set is calculated, and then the Metropolis criterion is used to decide whether to accept the candidate learning rate. The model and formula define an acceptance probability P: ,in To use candidate learning rates The loss value calculated on the training set after one training cycle. It is the optimal loss value found during the search process. Given the current temperature, when Accept this loss value as the optimal loss value, and accept this learning rate as the optimal learning rate. The algorithm accepts a poor solution with probability P, effectively preventing it from getting trapped in local optima and ensuring that it explores a wider learning rate space at high temperatures. As the temperature decreases, the probability of accepting a poor solution decreases, and the search gradually converges. When the temperature drops to its lowest point, the optimal learning rate is output. and its corresponding model parameters.
[0021] Step 3: In six key geometric parameters Within the range of values, the parameter is discretized with an engineering step size of 0.01 mm to construct a full-factor grid containing 15625 combinations of candidate parameters. The value ranges are {0.2-0.36}, {3.4-4.6}, {0.15-0.3}, {0.8-1.1}, {0.8-1.0}, and {0.2-0.5}. Then, the model is used to predict the response of all combinations in parallel. The post-processing script automatically identifies the continuous frequency bands in each S11 curve that satisfy the S11 parameter being less than -10 dB, calculates the impedance bandwidth, and finally selects the parameter combination that maximizes the impedance bandwidth and simultaneously meets the gain requirements as the optimal antenna design. The multi-target screening conditions include: within the 27.04-30.93 GHz frequency band, the S11 parameter is less than -10 dB; the peak gain is not less than 7.2 dBi, and the in-band gain fluctuation is less than 0.5 dB.
[0022] Step 4: Substitute the optimal structural parameters obtained from the reverse design into the antenna physical model in the 3D electromagnetic simulation software to perform full-wave electromagnetic simulation, and obtain the antenna's S-parameters and gain; compare the simulation results with the prediction results of the deep learning model to verify whether its impedance bandwidth and gain stability meet the preset comprehensive performance index requirements, thereby completing the closed loop from intelligent optimization design to performance verification, confirming the effectiveness of the design method and the superiority of the obtained antenna structure.
[0023] Compared with existing technologies, this invention has the following advantages: 1. The low cross-polarization antenna based on deep learning optimization design of this invention, while being integrated, also possesses the characteristics of low cross-polarization, wide impedance bandwidth, and stable gain. 2. This invention, a low cross-polarization antenna based on deep learning optimization design, adopts deep learning to optimize antenna design, proposing an optimization design process based on a CNN-Transformer hybrid neural network model and simulated annealing optimization. This achieves joint prediction and inverse optimization of antenna S11 parameters and gain characteristics, transforming the traditional experience-based trial-and-error process into an efficient, global data-driven optimization process, shortening the design cycle and improving design quality. Attached Figure Description
[0024] Figure 1 This is a schematic diagram of an explosion structure for a low cross-polarization antenna based on deep learning optimization design according to the present invention.
[0025] Figure 2 This is a schematic diagram of the main radiating element in a low cross-polarization antenna based on deep learning optimization design according to the present invention.
[0026] Figure 3 This is a schematic diagram of the layout of the coupled ground metal layer and the feed microstrip line in a low cross-polarization antenna based on deep learning optimization design according to the present invention.
[0027] Figure 4 This is a flowchart illustrating a deep learning optimization design method for a low cross-polarization antenna based on deep learning optimization design, according to the present invention.
[0028] Figure 5 This is a comparison of the performance prediction results of a low cross-polarization antenna based on deep learning optimization design in the present invention with the full-wave electromagnetic simulation results in a three-dimensional electromagnetic simulation software. (a) shows the comparison of S11 parameters in the training set, (b) shows the comparison of gain in the training set, (c) shows the comparison of S11 parameters in the test set, and (d) shows the comparison of gain in the test set.
[0029] Figure 6 The normalized radiation pattern at 29.26 GHz is shown for a low cross-polarization antenna based on deep learning optimization design according to the present invention. (a) is the E-plane radiation pattern, and (b) is the H-plane radiation pattern. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0031] like Figure 1As shown, this embodiment of a low cross-polarization antenna based on deep learning optimization design includes, from top to bottom, a main radiating element 1, a first dielectric layer 2, a bonding and curing layer 3, a coupling ground metal layer 4, a second dielectric layer 5, a feed microstrip line 7, and two vertical metal pillars 6 penetrating the first dielectric layer 2 and the bonding and curing layer 3. The main radiating element 1 is disposed on the top surface of the first dielectric layer 2; the coupling ground metal layer 4 is disposed on the top surface of the second dielectric layer 5; and the feed microstrip line 7 is disposed on the bottom surface of the second dielectric layer 5. The two vertical metal pillars 6 are symmetrically distributed about the geometric center of the antenna. With the geometric center of the antenna as the central axis, each functional layer and its internal structure are precisely aligned and symmetrically distributed based on this central axis, forming an integrated stacked structure with vertical alignment between layers and in-plane mirror symmetry.
[0032] like Figure 2 As shown, this embodiment of a low cross-polarization antenna based on deep learning optimization design has a cross-shaped groove 8 etched on the main radiating element 1. This groove is formed by the perpendicular intersection of a transverse groove and a longitudinal groove. A rectangular matching band 9 and a double T-shaped matching shorting line 10 are provided in the transverse groove. The double T-shaped matching shorting line 10 is formed by two T-shaped units with identical structures symmetrically arranged about the rectangular matching band 9. Each T-shaped unit includes a longitudinal arm and a transverse arm. Figure 3 As shown, the coupling grounding metal layer 4 is etched with an H-shaped coupling groove 11, which is composed of a horizontal main groove and two vertical secondary grooves; the power feeding microstrip line 7 includes a T-type equal power divider, which includes an input main branch, two symmetrical output branches and two matching rectangular blocks.
[0033] See Figure 1 , 23. The structural dimensions of a low cross-polarization antenna based on deep learning optimization design in this embodiment can be defined by a set of geometric parameters, including but not limited to: the thickness H1 of the first dielectric layer 2, the thickness H2 of the second dielectric layer 5, the thickness H3 of the bonding and curing layer 3; the side length Lsp of the main radiating element 1; the transverse slot length L and width W of the cross groove 8, and the longitudinal slot width W2; the length L3 and width W3 of the rectangular matching band 9, and the spacing Gap between the rectangular matching band 9 and the transverse slot of the cross groove 8; the transverse arm length T and width W1 of the T-shaped unit of the double T-shaped matching short line 10, the spacing Gin between the transverse arm and the rectangular matching band 9, and the spacing Gout between the transverse arm and the transverse slot of the cross groove 8. The longitudinal arm length L5 and width S of the T-shaped unit of the double T-shaped matching short line 10, the distance lv between the longitudinal arm and the adjacent side of the main radiating element 1, and the distance G between the longitudinal arm and the longitudinal slot of the cross groove 8; the diameter dv of the vertical metal pillar 6, and the distance Sv between the two vertical metal pillars 6; the width Ws of the main slot, the length Ls of the main slot, the width Wi of the secondary slot, the length Li of the half arm of the secondary slot, and the distance Si between the secondary slot and the starting end of the transverse main slot on the H-shaped coupling slot 11 on the coupling grounding metal layer 4; the length Lf1 and width Wf of the input main branch, the length L4 and width Wf1 of the output branch, the distance Wd between the two output branches, and the length Lt and width Wt of the matching rectangular block in the T-shaped equal power divider of the feeding microstrip line 7.
[0034] See Figure 1 , 2 3. In this embodiment, a low cross-polarization antenna based on deep learning optimization design, from signal input to electromagnetic wave radiation, the process and principle are as follows.
[0035] The radio frequency signal is first input from the main branch of the underlying T-type equal power divider for power distribution. The branch line length and linewidth of the T-type equal power divider are optimized to ensure that the energy is evenly distributed to the two output branches. The two signals after distribution have equal amplitude and the same phase.
[0036] When the underlying equalization signal is transmitted to the H-shaped coupling slot 11, the alternating current of the microstrip line will generate an alternating electric field within the H-shaped coupling slot 11. The electric field distribution can be optimized by adjusting the parameters of the main slot and the secondary slot, thereby improving the coupling efficiency. Two vertical metal pillars 6 are symmetrically arranged on both sides of the H-shaped coupling slot 11. The vertical metal pillars 6 penetrate the first dielectric layer 2 and the bonding and curing layer 3, with the lower end connected to the coupling ground metal layer 4 and the upper end connected to the double T-shaped matching short circuit 10. The H-shaped coupling slot and the vertical metal pillars together constitute an embedded differential excitation architecture. The alternating electric field within the slot forms a 180° phase difference at the connection between the two vertical metal pillars 6 and the H-shaped coupling slot 11, which is the differential excitation signal, perfectly matching the differential operating mode of the double T-shaped matching short circuit 10. At the same time, the length of the vertical metal pillar 6 is equal to the sum of the thickness H1 of the first dielectric layer 2 and the thickness H3 of the bonding and curing layer 3, and its electrical length at the 29 GHz frequency point is approximately one-quarter of the dielectric wavelength. This design can offset the impedance change during the coupling process, maximizing the coupling efficiency.
[0037] The differential current is conducted through the vertical metal support 6 to the double-T matching short circuit 10. The double-T matching short circuit 10 consists of two identical T-shaped units forming a matching pair, dividing the slot ring into two independent resonant paths. The impedance bandwidths of the two paths are superimposed and complementary, providing sufficient impedance bandwidth for the antenna in the dual-target frequency band to meet the wideband operation requirements. At the same time, the symmetrical structure helps to form differential-mode current and suppress common-mode interference. The current in the double-T matching short circuit 10 exhibits a symmetrical distribution of "strong at both ends and weak in the middle," ensuring that the current oscillation frequency is accurately matched with the resonant frequency band, without parasitic interference. Oscillations are generated; in addition, the double-T-shaped matching short circuit 10, as the main body of the radiating unit, has its own parasitic inductance. The two vertical slots in the H-shaped coupling slot 11 introduce parasitic capacitance, and the two naturally form a series resonant circuit to cancel impedance distortion in a wide frequency band. The good impedance matching performance can reduce gain fluctuations caused by energy reflection. The symmetrical design further suppresses cross-polarization, allowing energy to be concentrated in the common polarization direction for radiation, avoiding useless losses and stabilizing the gain; finally, the alternating current oscillates on the double-T-shaped matching short circuit 10, exciting directional waves for radiation.
[0038] This embodiment presents a low cross-polarization antenna based on deep learning optimization design. The optimization design process is as follows: Figure 4 As shown, it includes the following four steps: parametric modeling and dataset construction, building and training a CNN-Transformer hybrid neural network model, multi-target screening reverse design, and antenna implementation and verification.
[0039] Step 1: The optimization target in this embodiment is Figure 1The low cross-polarization antenna shown uses six key geometric parameters as design variables for the neural network, including the main slot width Ws of the H-shaped coupling slot 11, the main slot length Ls of the H-shaped coupling slot 11, the secondary slot width Wi of the H-shaped coupling slot 11, the half-arm length Li of the secondary slot of the H-shaped coupling slot 11, the longitudinal arm width S of the T-shaped element of the double T-shaped matching short path 10, and the distance lv between the longitudinal arm of the T-shaped element of the double T-shaped matching short path 10 and the adjacent side of the main radiating element 1. Within the respective sampling ranges of the six parameters, five discrete values for each parameter are selected for combined sampling, with sampling ranges of {0.2-0.36}, {3.4-4.6}, {0.15-0.3}, {0.8-1.1}, {0.8-1.0}, and {0.2-0.5}, all in millimeters. Then, the parameter space is scanned using full-wave electromagnetic simulation software to generate the corresponding output response, sampling a 401-dimensional S11 amplitude vector within the 24-34 GHz frequency band. Gain value vectors at 12 discrete frequency points ,in The response data for parameter S11, It is a 401-dimensional real vector space. The response data is for the gain parameter. It is a 12-dimensional real vector space; all input and output data are processed by max-min normalization and randomly divided into training set, validation set and test set in a ratio of 8:1:1 to form the basic data pairs for model learning.
[0040] Step 2: Construct a CNN-Transformer hybrid neural network model and use simulated annealing algorithm to optimize the model training process.
[0041] The main process of the CNN branch is to capture local correlation patterns between parameters through convolution operations. First, the 6-dimensional input vector is reshaped into a format suitable for processing by a 1-dimensional convolutional network. Where B is the batch size, and in this embodiment, B=32. The data is then fed into the first one-dimensional convolutional layer, which uses 32 convolutional kernels with a width of 3 to scan the input sequence, performing convolution calculations and ReLU activation function operations to initially extract combined features between adjacent parameters. Next, the features are fed into the second one-dimensional convolutional layer, which uses 64 convolutional kernels with a width of 3 to further refine and abstract the features from the previous layer. Both convolutional layers maintain the same sequence length. Finally, to compress the 64 features at 6 positions into a fixed-length feature vector, global max pooling is used, retaining only the maximum value in each feature channel of the sequence, ultimately outputting a 64-dimensional feature vector. It centrally characterizes the electromagnetic properties determined by the local proximity relationships of geometric parameters.
[0042] The main process of the Transformer branch is to establish global dependencies between all parameters through a self-attention mechanism. First, a linear embedding layer maps each 6-dimensional input parameter to a high-dimensional embedding vector. Where B is the batch size, and in this embodiment, B=32. Next, a fixed sine-cosine positional encoding is added, enabling the model to perceive the order of parameters in the sequence. ,in For the high-dimensional input matrix after adding position encoding, The sequence is then encoded into a fixed-position matrix. It is then fed into a module consisting of three identical Transformer encoder layers for deep processing. Within each encoder layer, two core sub-processes are executed sequentially: First, multi-head self-attention computation is performed. This model divides the 64-dimensional features into two 32-dimensional heads for parallel execution. This mechanism allows each parameter in the sequence to dynamically evaluate and aggregate information from all other parameters, thus establishing global contextual relationships. The self-attention output, after residual connections and layer normalization, enters the second sub-process: a feedforward neural network. This is a two-layer fully connected network containing a ReLU activation function for non-linear transformation, used to further transform the features of each parameter. The output of this sub-process also undergoes residual connections and layer normalization. After three such stacked layers, a sequence containing six 64-dimensional vectors rich in global interaction information is obtained. Finally, a learnable attention pooling layer is used. This layer automatically learns the importance weights of each positional feature and performs a weighted summation of the entire sequence, thus aggregating all information and outputting a 64-dimensional feature vector. It deeply encodes the complex, long-range mutual coupling relationships between all antenna structural parameters.
[0043] The results from the two branches are integrated to complete the final prediction. The 64-dimensional local feature vector output from the CNN branch is then used. The 64-dimensional global feature vector output by the Transformer branch The features are directly concatenated to form a 128-dimensional fused feature vector, which contains both local fine structure information affecting antenna performance and global system correlation information. Subsequently, this fused feature vector is simultaneously fed into two independently operating prediction heads. The first prediction head is responsible for outputting the S11 parameter characteristics of the antenna throughout the target frequency band. It maps the 128-dimensional features into a 401-dimensional vector, corresponding to the predicted S11 parameters of 401 sampling points in the 24-34 GHz frequency band. The second prediction head is responsible for outputting the antenna gain characteristics. It maps the 128-dimensional features into a 12-dimensional vector, corresponding to the predicted gain values at 12 specified frequency points.
[0044] Model training consists of two stages: the first stage is initial learning rate optimization, which uses simulated annealing for adaptive search. The algorithm sets the temperature update formula as follows: Initial temperature minimum temperature attenuation coefficient ,in The current temperature. For the updated temperature; at each temperature iteration, generate candidate learning rates based on the current optimal learning rate: ,in It follows a standard normal distribution. To achieve the optimal learning rate, Let the candidate learning rate be used; the model is quickly trained for one cycle using the candidate learning rate, and the training loss is calculated. The Metropolis acceptance criterion is used to decide whether to adopt the candidate solution: if the candidate loss... Less than the current optimal loss If so, accept with probability 1; otherwise, accept with probability 2. Accept; when the temperature drops The search stops at the following point, and the initial optimal learning rate is output. The second stage is formal training, using the initial optimal learning rate. Initialize the adaptive moment estimation optimizer (Adam), and use a smoothed L1 loss function to jointly optimize the prediction error of the S11 parameters and gain: ,in, Weights for gain prediction loss. Predict loss weights for parameters S11, α=0.7, β=0.3. The total loss of the model is N, where N is the number of training samples. To smooth the L1 loss function, The sample number. For the first The true gain value of each sample. For the first The predicted gain value for each sample. For the first The true values of the S11 parameters for each sample. For the first The S11 parameter prediction values for each sample; during training, an adaptive learning rate adjuster (ReduceLROnPlateau) is enabled. When the validation set loss no longer decreases within 10 consecutive training epochs, the learning rate is reduced to half of its original value; the entire training process is conducted under the monitoring of the validation set, with a total of 2000 training epochs, and finally the model parameters with the minimum validation loss are saved.
[0045] Step 3: Based on the trained high-precision prediction model, construct a discretized grid for the six key geometric parameters; [The text abruptly ends here, likely due to an incomplete sentence or missing information.] The parameters are discretized within a range of 0.01 mm engineering steps to construct a full-factor grid containing 15625 candidate parameter combinations, with values ranging from {0.2-0.36}, {3.4-4.6}, {0.15-0.3}, {0.8-1.1}, {0.8-1.0}, and {0.2-0.5}, all in millimeters. Then, the trained model is used to perform parallel forward inference on all designs within this grid, batch-producing the predicted S11 curves and gain spectra for each group. Next, a post-processing script automatically identifies continuous frequency bands in each S11 curve that satisfy an S11 parameter less than -10 dB, calculates the impedance bandwidth, and finally selects the parameter combination that maximizes the impedance bandwidth while simultaneously meeting the gain requirements as the optimal antenna design. The multi-target selection criteria include: S11 parameter less than -10 dB in the 27.04-30.93 GHz band; peak gain not less than 7.2 dBi; and in-band gain fluctuation less than 0.5. dB.
[0046] Table 1 shows a set of optimal structural parameters for low cross-polarization antennas obtained through reverse design. The six key geometric parameters marked with * are optimized. This parameter combination is a preferred embodiment obtained by the method of the present invention, but it is not the only limitation of the present invention.
[0047] Table 1. List of optimized low cross-polarization antenna parameters (unit: mm)
[0048] parameter H1 H2 H3 L W Lsp W2 Gout Value 1.524 0.254 0.09 4.2 1 5.2 1.7 0.1 parameter L3 W3 Gap W1 Gin dv Sv G Value 3.5 0.12 0.3 0.1 0.24 0.4 2.1 0.4 parameter Wf Lf1 Wf1 L4 Wd T L5 Si Value 0.6 1.59 0.12 2.33 2.4 2.4 1.8 0.45 parameter Lt Wt Ws* Ls* Wi* Li* S* lv* Value 0.28 0.6 0.31 4.1 0.2 0.96 0.9 0.4
[0049] Step 4: Substitute the optimal structural parameters obtained from the reverse design into the antenna physical model in the 3D electromagnetic simulation software to perform full-wave electromagnetic simulation and obtain the antenna's S-parameters and gain; compare the simulation results with the prediction results of the deep learning model to verify whether its impedance bandwidth and gain stability meet the preset comprehensive performance index requirements.
[0050] Figure 5This paper presents a low cross-polarization antenna based on deep learning optimization design. The performance prediction results on the training and test sets are compared with the full-wave electromagnetic simulation results from 3D electromagnetic simulation software. After 2000 training iterations, the proposed model demonstrates good fitting ability and spectral consistency in the predicted S11 parameters and gain on both the training and test sets. The predicted S11 curve closely matches the actual full-wave electromagnetic simulation curve in terms of resonant point, matching depth, and -10 dB impedance bandwidth. The impedance bandwidth is verified to be 3.89 GHz in the 27.04–30.93 GHz range. The predicted gain almost overlaps with the actual full-wave electromagnetic simulation gain at 12 frequency points, with a peak gain of 7.74 dBi and gain fluctuation of less than 0.5 dB. This verifies that the proposed deep learning model meets the design requirements in terms of prediction accuracy and reliability.
[0051] Figure 6 (a) shows the radiation pattern of a low cross-polarization antenna designed based on deep learning in the E-plane at 29.26 GHz. The main polarization pattern has good symmetry, and the cross-polarization component is suppressed to below -50 dB, achieving an extremely low level. Figure 6 (b) shows its radiation pattern in the H plane at 29.26 GHz, with good symmetry in the main polarization pattern and cross-polarization below -15 dB.
[0052] The above results demonstrate that the proposed low cross-polarization antenna based on deep learning optimization design can be used for antenna optimization design. This method uses a CNN-Transformer hybrid neural network model to predict antenna S11 parameters and gain indicators. Through simulated annealing optimization training and automated reverse screening processes, multi-objective collaborative optimization is performed. The resulting antenna exhibits low cross-polarization, wide impedance bandwidth, stable gain, and high integration in the millimeter-wave band.
[0053] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. It should be noted that the structure, number of layers, number of parameters, etc., of the CNN-Transformer hybrid neural network model described in the embodiments are only a preferred implementation. Those skilled in the art can adjust it according to actual conditions or adopt other neural network architectures capable of fusing local and global features. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A low cross-polarization antenna based on deep learning-optimized design, characterized in that, The antenna employs a deep learning-optimized design, comprising, from top to bottom, a main radiating element, a first dielectric layer, a bonding and curing layer, a coupling ground metal layer, a second dielectric layer, a feed microstrip line, and a vertical metal support. The main radiating element includes a cross-shaped groove, a rectangular matching band, and a double T-shaped matching short line; the cross-shaped groove is formed by the perpendicular intersection of a transverse groove and a longitudinal groove, the rectangular matching band is located in the transverse groove of the cross-shaped groove, and the double T-shaped matching short line is formed by two T-shaped units with identical structures, the two T-shaped units are symmetrically arranged about the rectangular matching band, and each T-shaped unit includes a longitudinal arm and a transverse arm; The coupling grounding metal layer includes an H-shaped coupling groove, which is composed of a horizontal main groove and two vertical secondary grooves. The power-fed microstrip line includes a T-type equal power divider, which comprises an input main branch, two symmetrical output branches, and two matching rectangular blocks; The vertical metal support consists of two pillars, symmetrically distributed about the geometric center of the antenna, penetrating the first dielectric layer and the bonding and curing layer. The upper end is connected to the double-T-shaped matching short circuit of the main radiating element, and the lower end is connected to the coupling ground metal layer. The H-shaped coupling groove and the vertical metal support together form an embedded differential excitation architecture, matching the differential operating mode of the double-T-shaped matching short circuit. The deep learning optimization design includes the following four parts: parametric modeling and dataset construction, construction and training of CNN-Transformer hybrid neural network model, multi-target screening reverse design, and antenna implementation and verification; Based on the electromagnetic characteristics of the antenna, six key geometric parameters affecting the performance of the low cross-polarization antenna were determined, including the main slot width, main slot length, secondary slot width, secondary slot half-arm length, longitudinal arm width of the T-shaped unit of the double T-shaped matching short path, and the distance between the longitudinal arm of the T-shaped unit of the double T-shaped matching short path and the adjacent side of the main radiating element. Random sampling is performed within a preset parameter range to generate multiple sets of structural parameter combinations. Full-wave electromagnetic simulation is then performed in a 3D electromagnetic simulation software to obtain the S11 parameter curves and gain frequency response corresponding to each set of parameters. Frequency point data is extracted from these parameters to construct a dataset of structural parameters and performance response. The CNN-Transformer hybrid neural network model takes the key geometric parameters as input and S11 parameters and gain as output; the model is trained using the constructed dataset, and a simulated annealing algorithm is introduced during the training process to adaptively optimize the initial learning rate in order to obtain a convergent training model with high prediction accuracy. The trained model is used for reverse design to quickly predict the performance of a large number of candidate parameter combinations; the optimal antenna structure parameter combination that meets the requirements of wide bandwidth and stable gain is obtained by screening based on the comprehensive performance index of preset impedance matching bandwidth and gain. The obtained optimal parameter combination is substituted into the antenna physical model in the three-dimensional electromagnetic simulation software to perform full-wave electromagnetic simulation and verify whether its electrical performance meets the predicted indicators.
2. The low cross-polarization antenna based on deep learning optimization design according to claim 1, characterized in that, The parametric modeling and dataset construction involves the following dataset construction: six key geometric parameters are discretely sampled at the millimeter level within the ranges {0.2-0.36}, {3.4-4.6}, {0.15-0.3}, {0.8-1.1}, {0.8-1.0}, and {0.2-0.5}, and combined to form structural parameters. Each set of structural parameters is then used for full-wave electromagnetic simulation to obtain a 401-dimensional S11 amplitude vector and a gain vector at 12 discrete frequency points sampled in the 24-34 GHz frequency band as the output response. After the input and output data are subjected to max-min normalization, they are randomly divided into training, validation, and test sets in a ratio of 8:1:
1.
3. The low cross-polarization antenna based on deep learning optimization design according to claim 1, characterized in that, The CNN-Transformer hybrid neural network model has a model architecture that specifically includes parallel CNN branches and Transformer branches; The CNN branch consists of two sequentially connected one-dimensional convolutional modules. Each module includes a one-dimensional convolutional layer, a ReLU activation function, and a pooling layer to extract local interaction features between input parameters. The first convolutional layer is used to learn local interaction patterns between parameters. The second convolutional layer further fuses these primary patterns to form a higher-level local feature representation. Between convolutional operations, non-linearity is introduced through the ReLU activation function, and pooling operations are used to reduce the feature dimension. Finally, global pooling aggregates the features from all locations into a 64-dimensional feature vector. The Transformer branch includes a linear embedding layer, position encoding, a three-layer Transformer encoder, and an attention pooling layer, used to extract global features between antenna parameters. The Transformer encoder includes a multi-head self-attention mechanism, a feedforward neural network, residual connections, and layer normalization. The input 6-dimensional antenna key parameters are mapped to a 64-dimensional vector through the linear embedding layer to enhance feature representation. Then, fixed position encoding is added to preserve parameter order information. The data undergoes deep interaction through the three-layer Transformer encoder. In each layer, the multi-head self-attention mechanism divides the 64-dimensional features into two 32-dimensional subspaces for parallel computation, enabling each parameter to interact with all other parameters in the sequence, thereby capturing global dependencies. The first encoder performs preliminary multi-head self-attention computation on the parameter sequence embedded with positional information, establishing basic interaction patterns between parameters, and enhances the nonlinear expressive power of features through a feedforward network. The second encoder further captures long-range coupling relationships, deepening the modeling of the mutual influence between different structural parameters. The third encoder performs high-level integration and abstraction of the extracted global features, outputting a feature sequence containing rich contextual information. The feedforward neural network further extracts nonlinear features, and residual connections and layer normalization ensure training stability. Finally, learnable attention pooling weights and aggregates the six positional features output by the encoder into a 64-dimensional global feature vector. The local feature vector output by the CNN branch is concatenated with the global feature vector output by the Transformer branch to obtain a fused feature vector, which is then mapped to the S11 parameter prediction vector and the gain prediction vector by two independent fully connected networks, respectively.
4. The low cross-polarization antenna based on deep learning optimization design according to claim 1, characterized in that, The multi-target screening reverse design is achieved as follows: within the multi-dimensional design space composed of the six key geometric parameters, the design is discretized with an engineering step size of 0.01 mm to construct a full-factor grid containing 15,625 candidate parameter combinations; using the trained model, forward prediction is performed on all parameter combinations in the grid to obtain the corresponding S11 curves and gain curves; the post-processing program automatically identifies continuous frequency bands in each predicted S11 curve that satisfy the S11 parameter less than -10 dB, calculates their impedance bandwidth, and selects the parameter combination that makes both the impedance bandwidth and gain meet the preset requirements as the optimal antenna design.
5. A low cross-polarization antenna based on deep learning optimization design according to claim 1, characterized in that, The comprehensive performance indicators for antenna implementation and verification include: S11 parameter less than -10 dB in the 27.04-30.93 GHz band; peak gain not less than 7.2 dBi; and in-band gain fluctuation less than 0.5 dB.