A dual-function integrated joint transceiver method for MIMO radar communication based on machine learning

Through the machine learning-based joint design method for transmitting and receiving, the waveform design, detector and decoder in the integrated dual-power radar communication system are optimized, which solves the problem of insufficient performance of the existing system and achieves more efficient detection and decoding performance.

CN117254999BActive Publication Date: 2025-05-16UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310831682.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-07
Publication Date
2025-05-16
Estimated Expiration
2043-07-07

AI Technical Summary

Technical Problem

The existing dual-power radar communication integrated system has insufficient performance in waveform design, joint design of detectors and decoders, and it is difficult to effectively improve the detection and decoding performance of the system.

Method used

Using a joint transceiver and reception design method based on machine learning, the design of dual-power SO-OCDM waveform, radar detector and communication decoder are optimized by building a waveform design network, radar detector and communication decoder network, and using a combination of reinforcement learning and supervised learning.

Benefits of technology

The detection and decoding performance of the integrated dual-power radar communication system has been significantly improved. Compared with traditional methods, machine learning-based methods can optimize system performance more effectively and have less dependence on mathematical models such as channels, targets and clutter.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117254999B_ABST
    Figure CN117254999B_ABST
Patent Text Reader

Abstract

This invention discloses a MIMO radar communication dual-power integrated joint receiving and transmitting method based on machine learning, which belongs to the field of radar communication integration technology. By adopting reinforcement learning and supervised learning, the joint design of dual-power SO-OCDM waveform, radar detector and communication decoder is realized. Compared with traditional methods, the method based on machine learning is a method with controllable complexity, which is less dependent on mathematical models of channels, targets and clutter, and can effectively improve system performance. It provides a feasible method for the joint design of receiving and transmitting in the field of dual-power integration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of radar communication integration, and particularly relates to the design of waveforms, detectors and decoders in dual-power radar communication integration. Background Art

[0002] As the number of electronic devices increases, spectrum resources become scarce, which leads to overlapping radar communication operating bands. Dual-function radar communication integration (DFRC) becomes an effective solution. The DFRC system can reduce system consumption and achieve tight integration by transmitting a dual-function waveform with radar communication function.

[0003] In DFRC systems, waveform design is very critical. In reference 1 (LGdOliveira, B.Nuss, MBAlabd, A.Diewald, M.Pauli, and T.Zwick, "Joint radar-communication systems: Modulation schemes and system design," IEEE Transactions on Microwave Theory and Techniques, vol.70, no.3, pp.1521–1551, 2021), orthogonal chirp multiplexing (OCDM) has been studied for detection and decoding tasks in DFRC systems because of its robustness to multipath, high processing gain and high data rate. For multiple-input multiple-output (MIMO) systems, the OCDM signals transmitted by different antennas are orthogonal to each other, and the chirped subcarriers in the same frequency band are orthogonal, forming a spatially orthogonal OCDM (SO-OCDM) signal.

[0004] In order to improve the performance of the DFRC system, the joint design of the waveform and the receiver can be considered. In reference 2 (C. Wen and T. N. Davidson, "Transceiver design for MIMO-DFRC systems," in ICASSP 2023-2023IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5), the joint optimization of the radar receiving filter and the transmitting waveform based on continuous convex optimization is considered. In addition, in reference 3 (JM Mateos-Ramos, J. Song, Y. Wu, C. ,MFKeskin,V.Yajnanarayana,and H.Wymeersch,“End-to-end learning for integrated sensingand communication,”in ICC 2022-IEEE International Conference onCommunications.IEEE,2022,pp.1942–1947)The self-encoding architecture based on supervised learning has also been applied to the joint design of detection, estimation and transmission waveforms of co-located MIMO. The above method has guiding significance for the transceiver design problem of DFRC system. Therefore, in the MIMO DFRC system, considering the target detection of the radar subsystem and the data decoding tasks of the communication subsystem, it is very necessary to jointly design the transmitted SO-OCDM waveform, the receiving end detector and the decoder. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention proposes a joint design method for transmitting and receiving based on machine learning to improve the detection and decoding performance of the integrated system when the tasks of the radar communication subsystem are target detection and data decoding respectively.

[0006] To implement the proposed method, the technical solution of the present invention is: a MIMO radar communication dual-function integrated joint transceiver method based on machine learning, the method comprising:

[0007] Step 1: At the duplex transmitter, construct the mth transmitter SO-OCDM duplex waveform at kT s The signal at time x m [k], m=1,…,M, k=1,…,K;

[0008]

[0009] Among them, k is the index of the sampling point, E m is the transmitted signal energy, P m is the number of SO-OCDM symbols, I m is the number of subcarriers, rect(·) represents the matrix function, p is the index of the SO-OCDM symbol, i is the index of the chirp subcarrier, T o,m is the complete SO-OCDM symbol time, T s is the sampling time, K is the total number of samples; a p,i,m To modulate the data, 4QAM modulation is used, then Signals p,i,m [k] is

[0010]

[0011] Here, the specific expression of the chirped subcarrier without the guard interval is:

[0012]

[0013] In the above formula, for the SO-OCDM signal transmitted by the mth transmitting antenna, T o,m =T m +T GI,m , T m is the duration of a SO-OCDM symbol, T GI,m is the duration of the guard interval GI, f Δ Represents the frequency interval of the waveforms transmitted by two adjacent transmitting antennas to ensure the spatial orthogonality of the signal, that is, Define the modulation data matrix as Among them, a p,i =[a p,i,1 ,…,a p,i,M ] T ;

[0014] Step 2: Integrate the waveform parameters into a parameter vector:

[0015] λ m =[E m ,T m ,T GI,m ,I m ,P m ] T

[0016] Stack the parameter vectors of M different antennas,

[0017]

[0018] Step 3: Construct a waveform design network, including a fully connected neural network FCN and a SO-OCDM waveform generator; the input is the initialized waveform parameters and rewards, and the output is the optimized SO-OCDM waveform;

[0019] Step 4: nth R A radar receiver at time kT s The received signal is

[0020]

[0021] in, is the sum of the M target reflection path signals, is the sum of the signals of M direct paths, for noise; and represent the path fading of the target reflection path and the direct path, respectively. and Indicates the delay of the target reflection path and the direct path;

[0022] The N received in MIMO integration R The signal sampling values ​​of the radar receivers are arranged in sequence to form the received signal y R

[0023]

[0024] in,

[0025] The SO-OCDM dual-power waveform transmitted by the mth transmitter in the target reflection path is kT s The sampling value at the time,

[0026]

[0027]

[0028]

[0029]

[0030]

[0031] Indicates MN R A column vector whose elements are all 1; z R for noise;

[0032] Step 5: Establish hypothesis testing questions based on whether the target exists or not.

[0033]

[0034] Here, H1 and H0 represent the hypothesis of the existence and non-existence of the target, respectively; for the H1 hypothesis, its label is ε R =1, for the H0 hypothesis, its label is ε R =0;

[0035] Step 6: Construct a detector network based on supervised learning, which includes a signal stacking module, FCN and a comparator; the input of the detector network is the radar received signal, and the output is the judgment;

[0036] Step 7: nth C A communication receiver at time kT s The received signal is n C=1,…,N C ;

[0037]

[0038] in, is the direct path signal, is the target reflection path signal, is the communication noise; and represents the path fading of the corresponding path, and is the delay of the corresponding path;

[0039] Step 8: Perform matched filtering on the communication receiving signal. The output of the filter is:

[0040]

[0041] Here, b p,i,m is the output of the i-th chirp subcarrier in the p-th SO-OCDM symbol transmitted by the m-th transmitting antenna after matched filtering, K o,m represents the number of sampling points in one SO-OCDM symbol time, K GI,m Indicates the number of sampling points within the protection interval. Indicates that at the sampling time (k+pK o,m +K GI,m )T s The conjugate value of the received communication signal; the outputs of all filters are arranged in order as a matrix, and the noisy data matrix is ​​obtained as B = [b 1,1 ,b 1,2 ,…b P,I ] T , where b p,i =[b p,i,1 ,…,b p,i,M ] T , and for m=1,...,M, we have P m =P,I m =I; In order to obtain diversity gain, the label of 4QAM modulated data (ε C ) p,i ∈{0, 1, 2, 3}, corresponding to four possible values;

[0042] Step 9: Construct a decoder network based on supervised learning, which includes modules such as matched filter, FCN, Max, and label-to-symbol mapping; the input of the decoder network is the communication received signal, and the output is the demodulated data;

[0043] Step 10: Establish the optimization problem of the joint design of the transmitter and receiver of the dual-function integrated system

[0044]

[0045]

[0046] Under the condition of satisfying the waveform parameter constraints, this optimization problem designs a network trainable parameter set θ for the waveform. T , the radar detector network has a trainable parameter set θ R , the trainable parameter set θ of the communication decoder network C The design and optimization of SO-OCDM waveform parameters λ are used to optimize radar detection performance and communication decoding performance, ρ R and ρ C Represent the weight factors of radar and communication performance respectively, l R and l C They are the performance indicators of radar and communication respectively, and the fully connected network of the waveform design network uses function Indicates that its input is the initialization waveform parameter λ0;

[0047] Step 11: Generate a transmit signal according to the SO-OCDM signal model, and generate receive data and corresponding labels according to the radar and communication receive signal models;

[0048] Step 12: Determine the loss function and parameter update method of the waveform design network, communication decoder network, and radar detector network; during the training process, all networks use Adam as the optimizer and calculate the gradient through forward propagation;

[0049] For the radar detector network, cross entropy is used as the loss function

[0050]

[0051] in, It is the qth R The labels of the training data, It is the qth R The output corresponding to the training data samples;

[0052] For the communication decoder network, the loss function is:

[0053]

[0054] in, It is the qth C The labels of the training data, It is the qth R The output corresponding to the training data samples is represents the probability that the estimated symbol and the label symbol are equal on the i-th chirp subcarrier of the p-th symbol. At the same time, the weighted sum of the radar and communication cross entropy is used as a reward to construct the loss function of the transmitter network :

[0055]

[0056] Among them, λ p According to Gaussian strategy The sampled waveform parameter vector, ρ R and ρ C are the weight factors for radar and communication respectively, [·] indicates the relationship between λ and p The ergodic average of

[0057] Step 13: Update the trainable parameters of the network nodes through back propagation; when training the detector and decoder networks, fix the parameters θ T , for θ R and θ C Update; when training the waveform design network, fix the network parameters θ R and θ C , for θ T Update; the two training steps are iterated alternately until the entire network converges;

[0058] Step 14: Generate test data according to the transmission and reception signal models, input the radar and communication test data into the detector network and the decoder network respectively, and obtain the output of the network; finally, compare the output of the network with the label to evaluate the system performance. The beneficial effect of the present invention is that by adopting reinforcement learning and supervised learning, the joint design of dual-power SO-OCDM waveform, radar detector and communication decoder is realized. Compared with the traditional method, the method based on machine learning is a method with controllable complexity, which is less dependent on the mathematical models of channels, targets and clutter, and can effectively improve the system performance, and provides a feasible method for the joint design problem of transmission and reception in the field of dual-power integration. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 Waveform design network architecture based on reinforcement learning;

[0060] Figure 2 Radar detector network based on supervised learning;

[0061] Figure 3 Communication decoder network based on supervised learning;

[0062] Figure 4 Joint training process of reinforcement learning and supervised learning;

[0063] Figure 5 ROC curves of different waveform parameters for the joint design of the transmitter and receiver and the design of the receiver only;

[0064] Figure 6 SER curves of different waveform parameters for the joint design of the transmitter and receiver and the design of the receiver only. DETAILED DESCRIPTION

[0065] For the convenience of description, the following definitions are first made:

[0066] () T is the transpose, () H is the conjugate transpose, () * represents conjugate, diag{·} represents diagonal matrix, Diag{·} represents block diagonal matrix, It means to find the ergodic average. Represents the gradient operator.

[0067] Step 1: Construct a dual-power integrated transmission waveform based on SO-OCDM.

[0068] Consider a MIMO dual-function integrated system with M transmitting antennas and N R radar receiving antennas and N C communication receiving antennas. Assume that a frame contains P m SO-OCDM symbols, the sampling period is T s , the sample number is k (k = 1, ..., K), and the total number of sampling points is K. The transmitted signal of the mth (m = 1, ..., M) transmitting antenna is

[0069]

[0070] Among them, E m is the energy of the waveform, I m is the number of chirp subcarriers, a p,i,m is the modulated data on the i-th chirp in the p-th symbol, using 4QAM modulation,

[0071] rect(·) represents the rectangle function,

[0072]

[0073] Here, the specific expression of the chirp subcarrier is

[0074]

[0075] In the above formula, for the SO-OCDM signal transmitted by the mth transmitting antenna, the complete symbol time T o,m =T m +T GI,m , Tm is the duration of a SO-OCDM symbol, T GI,m is the protection interval time, f Δ Represents the frequency interval of the waveforms transmitted by two adjacent transmitting antennas to ensure the spatial orthogonality of the signal, that is, For the pth symbol, the orthogonality between the chirped subcarriers can be expressed as Among them, K o,m and K GI,m denote the number of sampling points in a SO-OCDM symbol time and the guard interval time respectively. The modulation slope and bandwidth of the chirp signal can be calculated as and I m / T m The modulation data matrix is ​​defined as Among them, a p,i =[a p,i,1 ,…,a p,i,M ] T .

[0076] Step 2: Equivalent transformation of SO-OCDM waveform optimization problem

[0077] It can be seen from equations (1)-(3) that the SO-OCDM waveform x m [k] is completely determined by the waveform parameters. Therefore, the SO-OCDM waveform optimization problem is equivalent to the optimization problem of the waveform parameter λ. The waveform parameters are integrated into a parameter vector, which is expressed as

[0078] λ m =[E m ,T m ,T G,m ,I m ,P m ] T (4)

[0079] The parameter vector composed of the waveform parameters of all the transmitted waveforms at the transmitter is

[0080]

[0081] Step 3: Build a reinforcement learning-based waveform design network

[0082] Using the waveform optimization method based on reinforcement learning, the waveform design network architecture is as follows Figure 1 The SO-OCDM waveform design is regarded as an action, and the reward fed back by the receiver to the waveform design network is defined as ρ R l R +ρ C l C , ρ R and ρ Care the weight factors for radar and communication, respectively, l R and l C are the performance indicators of radar and communication respectively. The input of the fully connected network FCN is the initialized waveform parameter λ0, and the output is the optimized waveform parameter λ. Finally, the SO-OCDM generator is used to generate the SO-OCDM transmission waveform of M antennas.

[0083] Step 4: Establish the radar receiver's received signal model and its vector form

[0084] For the radar subsystem, the nth R (n R =1,…,N R )The received signal of the radar receiving antenna is

[0085]

[0086] in, and The mnth R The radar path losses for the direct path and the target reflection path are and here, is the target reflection coefficient, d mt is the distance between the mth transmitting antenna and the target, To target n R The distance between the radar receiving antennas, From the mth transmitting antenna to the nth R The distance between the radar receiving antennas. Mathematical symbols and Respectively and The ratio of transmitted energy to received energy when . Assuming the noise term It obeys a zero-mean mixed Gaussian distribution. No. g R The covariance matrix of the Gaussian mixture components is Corresponding to g R The coefficients of the mixed Gaussian components satisfy

[0087] The signal sampling points of all receiving antennas are stacked up to form a signal vector:

[0088]

[0089] in, Indicates MNR A column vector whose elements are all 1s.

[0090] Step 5: For target detection, establish the following hypothesis testing problem

[0091]

[0092] Here, H1 and H0 represent the hypothesis that the target exists and does not exist, respectively. For the H1 hypothesis, its label is ε R =1, for the H0 hypothesis, its label is ε R =0.

[0093] Step 6: Build a supervised learning based radar detector network

[0094] We use supervised learning in deep learning to detect targets, and transform the detection problem into a binary classification problem. For the hypothesis that the target exists, its label is ε R =1; for the hypothesis that the target does not exist, the label is ε R = 0. The input of the detection network FCN is the data y generated by equation (8) R , the output is p R . Compare the input with the threshold Λ and make a decision If p R ≥Λ, then it is determined that there is a target Otherwise, it is judged that there is no target Label it as ε R =0 data y R The corresponding network output p R Sort from large to small and output p R Take a specific value Λ in the equation so that it satisfies the value greater than Λ and all labels are ε R = 0 is the same as the given false alarm probability, then the corresponding Λ is the detection threshold. The detector network structure is as follows Figure 2 shown.

[0095] Step 7: Build a model of the received signal of the communication receiver

[0096] For the communication subsystem, the nth C (n C =1,…,N C The received signal of the communication receiving antenna is

[0097]

[0098] in, and The mnth C The communication path losses for the direct path and the target reflection path are and is the target reflection coefficient, To target n C The distance between the radar receiving antennas, From the mth transmitting antenna to the nth C The distance between the radar receiving antennas. Mathematical symbols and Respectively, the distance and The ratio of transmitted energy to received energy when . Assuming the noise term It obeys a zero-mean mixed Gaussian distribution. No. g C The covariance matrix of the Gaussian mixture components is Corresponding to g C The coefficients of Gaussian mixture components satisfy

[0099] Step 8: Perform matched filtering on the received signal of the mth communication receiving antenna to obtain the data carried on the i-th chirped subcarrier in the p-th SO-OCDM symbol. The output of the corresponding matched filter is

[0100]

[0101] In order to obtain diversity gain, it is assumed that the data transmitted by different transmitting antennas is the same. The noisy data matrix at the receiving end is defined as:

[0102] B=[b 1,1 ,b 1,2 ,…b P,I ] T , where b p,i =[b p,i,1 ,…,b p,i,M ] T

[0103] And for m=1,...,M, we have P m =P,I m =I. Label of 4QAM modulated data (ε C ) p,i ∈{0, 1, 2, 3}, corresponding to four possible values.

[0104] Step 9: Build a decoder network based on supervised learning

[0105] First, generate the data label ε according to the modulation data matrix and modulation order C The input of FCN is matrix B, and the output is the probability p of all labels. CThe Max module takes the probability matrix p C The label corresponding to the element with the maximum probability in each column vector is the estimated label. Finally, the data is decoded by mapping the label to the data. The decoder network structure is as follows: Figure 3 shown.

[0106] Step 10: Establish the optimization problem of the joint design of transmitter and receiver

[0107] For the entire integrated system, the joint optimization problem of the transmit waveform and the receiver can be written as

[0108]

[0109]

[0110] Among them, l R (λ,θ T ,θ R ) reflects the detection performance of the radar, l C (λ,θ T ,θ C ) reflects the decoding performance of communication, θ T ,θ R and θ C are the trainable parameter sets of the waveform design network, detector network, and decoder network, respectively. In order to optimize radar and communication performance simultaneously, the integrated system takes the weighted sum of the two as the reward of the transmitter. Considering the cross entropy as an indicator of the detection and decoding performance, the joint optimization problem can be understood as: Under the condition of satisfying the waveform parameter constraints, the waveform design network, detector network, decoder network and SO-OCDM waveform parameters are designed and optimized to minimize the weighted sum of cross entropy as the optimization goal.

[0111] Step 11: Generate the SO-OCDM transmission waveform x according to equations (1)-(3) and the iteratively optimized waveform parameter λ m [k] Generate receiving end training data based on radar communication receiving signal model and Then perform data preprocessing to obtain the network training data y R and B. At the same time, generate the labels ε corresponding to the radar and communication data R and ε C .

[0112] Step 12: Determine the loss function and parameter update method of the waveform design network, radar detector network, and communication decoder network, and use Adam as the optimizer to update the network parameters.

[0113] When training the FCN of the waveform design network, the hidden layer activation function is RLU, and the network output is normalized at the output layer so that the waveform parameters meet the constraints. When training the waveform design network, its goal is to minimize the reward, that is, the loss function.

[0114]

[0115] in, is the strategy in reinforcement learning, λ p According to the strategy The sampled action is the optimized waveform. When processing continuous action space, the strategy is usually considered as a Gaussian distribution, and λ0 is the mean of the Gaussian distribution.

[0116] Using Q T The gradient of the loss function is estimated using samples.

[0117]

[0118] Among them, for the qth T samples, For reward, According to the strategy The parameter vector obtained by sampling. In the iterative process, the Adam optimizer is used, and the u+1th detector network parameter iterative update method is

[0119]

[0120] in, Represents the loss function About parameter θ T The gradient of The value of η T is the learning rate. When training the FCN of the detector network, the hidden layer activation function is ReLU, and the output layer uses sigmiod as the activation function. R data samples, the loss function is

[0121]

[0122] Suppose there is Q R samples and labels, the estimated value of the loss function is

[0123]

[0124] In the iterative process, considering using the Adam optimizer, the iterative update process of the u+1th detector network parameters is:

[0125]

[0126] in, Represents the loss function About parameter θ R The gradient of The value of η R is the learning rate.

[0127] When training the decoder network, the hidden layer activation function is ELU, and the output layer uses softmax as the activation function. C data samples, the loss function is

[0128]

[0129] in, Corresponding to p C Medium Estimation Label With the true label Equal probability. Consider using Q C The estimated loss value of samples can be obtained

[0130]

[0131] In the iterative process, consider using the Adam optimizer, and the u+1th detector network iterative update process is:

[0132]

[0133] in, Represents the loss function About parameter θ C The gradient of The value of η C is the learning rate. Step 13: The joint training process of the dual-function integrated system.

[0134] The joint optimization involves training the waveform design network, radar detector network, and communication decoder network. When training the detector and decoder, the network parameters θ are fixed. T , for θ R and θ C When training the waveform design network, the fixed network parameters θ R and θ C , for θ T Update to minimize ρ R l R (λ,θ T ,θ R )+ρ C l C (λ,θ T ,θ C ). To explore the parameter space, a random strategy is adopted Used to generate waveform parameter vector λp ,like Figure 4 As shown in the dashed box in the figure. For the continuous waveform parameter space, the Gaussian strategy is adopted, and the waveform parameter vector is sampled from the Gaussian distribution. is the mean of the Gaussian distribution, The reinforcement learning at the source and the supervised learning at the receiver are iterated alternately until the entire network converges.

[0135] Step 14: Test the trained network.

[0136] After training, the optimal network parameter set can be obtained and In the testing phase, according to the waveform parameter vector Generate a batch of test waveforms. Generate test data and labels based on equations (8) and (9). Then input the radar and communication test data into the radar detection network and the communication decoding network respectively to obtain the network output. Finally, compare the network output with the label to evaluate the system performance.

[0137] exist Figure 5 In the figure, the ROC curve shows the false alarm probability P FA With the detection probability P D Changes in the situation. Figure 6 In the figure, the SER curve shows how the symbol error probability changes with the change of the signal-to-noise ratio SCNR. It can be seen from the figure that the joint design of the waveform design network, the detector network and the decoder network significantly improves the system performance compared with the case of only designing the reception. In the case of joint design of transmission and reception, the joint design of energy and symbol time can further improve the system performance compared with the case of only designing the energy or only designing the symbol time. When the number of subcarriers increases, the SER increases and the detection performance deteriorates because more unknown communication modulation data is carried in a SO-OCDM symbol.

Claims

1. A MIMO radar communication dual-function integrated joint transceiver method based on machine learning, the method comprising: Step 1: At the duplex transmitter, construct the mth transmitter SO-OCDM duplex waveform at kT s The signal at time x m [k], m=1,…,M, k=1,…,K; Among them, k is the index of the sampling point, E m is the transmitted signal energy, P m is the number of SO-OCDM symbols, I m is the number of subcarriers, rect(·) represents the matrix function, p is the index of the SO-OCDM symbol, i is the index of the chirp subcarrier, T o,m is the complete SO-OCDM symbol time, T s is the sampling time, K is the total number of samples; a p,i,m To modulate data, 4QAM modulation is used, then Signals p,i,m [k] is Here, the specific expression of the chirped subcarrier without the guard interval is: In the above formula, for the SO-OCDM signal transmitted by the mth transmitting antenna, T o,m =T m +T GI,m , T m is the duration of a SO-OCDM symbol, T GI,m is the duration of the guard interval GI, f Δ Represents the frequency interval of the waveforms transmitted by two adjacent transmitting antennas to ensure the spatial orthogonality of the signal, that is, k=0,m=1,=,M; define the modulation data matrix as Among them, a p,i =[a p,i,1 ,…,a p,i,M ] T ; Step 2: Integrate the waveform parameters into a parameter vector: λ m =[E m ,T m ,T GI,m ,I m ,P m ] T Stack the parameter vectors of M different antennas, Step 3: Construct a waveform design network, including a fully connected neural network FCN and a SO-OCDM waveform generator; the input is the initialized waveform parameters and rewards, and the output is the optimized SO-OCDM waveform; Step 4: nth R A radar receiver at time kT s The received signal is in, is the sum of the M target reflection path signals, is the sum of the signals of M direct paths, for noise; and represent the path fading of the target reflection path and the direct path, respectively. and Indicates the delay of the target reflection path and the direct path; The N received in MIMO integration R The signal sampling values ​​of the radar receivers are arranged in sequence to form the received signal y R in, The SO-OCDM dual-power waveform transmitted by the mth transmitter in the target reflection path is kT s The sampling value at the time, Indicates MN R A column vector whose elements are all 1; z R for noise; Step 5: Establish hypothesis testing questions based on whether the target exists or not. Here, H1 and H0 represent the hypothesis of the existence and non-existence of the target, respectively; for the H1 hypothesis, its label is ε R =1, for the H0 hypothesis, its label is ε R =0; Step 6: Construct a detector network based on supervised learning, which includes a signal stacking module, FCN and a comparator; the input of the detector network is the radar received signal, and the output is the judgment; Step 7: nth C A communication receiver at time kT s The received signal is n C =1,…,N C ; in, is the direct path signal, is the target reflection path signal, is the communication noise; and represents the path fading of the corresponding path, and is the delay of the corresponding path; Step 8: Perform matched filtering on the communication receiving signal. The output of the filter is: Here, b p,i,m is the output of the i-th chirp subcarrier in the p-th SO-OCDM symbol transmitted by the m-th transmitting antenna after matched filtering, K o,m represents the number of sampling points in one SO-OCDM symbol time, K GI,m Indicates the number of sampling points within the protection interval. Indicates that at the sampling time (k+pK o,m +K GI,m )T s The conjugate value of the received communication signal; the outputs of all filters are arranged in order as a matrix, and the noisy data matrix is ​​obtained as B = [b 1,1 ,b 1,2 ,…b P,I ] T , where b p,i =[b p,i,1 ,…,b p,i,M ] T , and for m=1,...,M, we have P m =P,I m =I; In order to obtain diversity gain, the label of 4QAM modulated data Corresponding to four possible values ​​respectively; Step 9: Construct a decoder network based on supervised learning, which includes modules such as matched filter, FCN, Max, and label-to-symbol mapping; the input of the decoder network is the communication received signal, and the output is the demodulated data; Step 10: Establish the optimization problem of the joint design of the transmitter and receiver of the dual-function integrated system s.t.constraint onλ, λ0 Under the condition of satisfying the waveform parameter constraints, this optimization problem is solved by designing the network trainable parameter set θ for the waveform. T , the radar detector network has a trainable parameter set θ R , the trainable parameter set θ of the communication decoder network C The design and optimization of SO-OCDM waveform parameters λ are used to optimize radar detection performance and communication decoding performance, ρ R and ρ C Represent the weight factors of radar and communication performance respectively, l R and l C They are the performance indicators of radar and communication respectively, and the fully connected network of the waveform design network uses function Indicates that its input is the initialization waveform parameter λ0; Step 11: Generate a transmit signal according to the SO-OCDM signal model, and generate receive data and corresponding labels according to the radar and communication receive signal models; Step 12: Determine the loss function and parameter update method of the waveform design network, communication decoder network, and radar detector network; during the training process, all networks use Adam as the optimizer and calculate the gradient through forward propagation; For the radar detector network, cross entropy is used as the loss function in, It is the qth R The labels of the training data, It is the qth R The output corresponding to the training data samples; For the communication decoder network, the loss function is: in, It is the qth C The labels of the training data, It is the qth R The output corresponding to the training data samples is represents the probability that the estimated symbol and the label symbol are equal on the i-th chirp subcarrier of the p-th symbol; At the same time, the weighted sum of radar and communication cross entropy is used as a reward to construct the loss function of the transmitter network Among them, λ p According to Gaussian strategy λ p |λ0 sampling waveform parameter vector, ρ R and ρ C are the weight factors for radar and communication respectively, Represents about λ p The ergodic average of Step 13: Update the trainable parameters of the network nodes through back propagation; when training the detector and decoder networks, fix the parameters θ T , for θ R and θ C Update; when training the waveform design network, fix the network parameters θ R and θ C , for θ T Update; the two training steps are iterated alternately until the entire network converges; Step 14: Generate test data based on the transmission and reception signal models, input the radar and communication test data into the detector network and decoder network respectively, and obtain the output of the network; finally, compare the output of the network with the label to evaluate the system performance.