Generative AI-based dual-base-station communication and sensing integrated system and sensing method

By combining the generative AI model with the Massive-MIMO-OFDM architecture, the problem of poor performance of traditional AI perception models in complex dynamic scenarios is solved, and high-precision and robust communication and perception integration is achieved, which is suitable for the upgrade of cellular wireless networks and other perception systems.

CN120567618APending Publication Date: 2025-08-29XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510657035.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

Traditional AI perception models have poor perceptual performance in complex dynamic scenarios, poor anti-interference ability, insufficient real-time and robustness, and cannot effectively analyze multi-dimensional coupled parameters and frequent parameter updates and complex calculations.

Method used

Using a dual-base station synesthesia integrated system based on generative AI, a time-frequency-space three-dimensional feature decoupling learning from the attention mechanism and embedded three-dimensional features is constructed, combined with the Massive-MIMO-OFDM architecture, the time-frequency-space feature embedding mechanism and the BERT network layer are introduced, and the channel state information reconstruction and perceptual parameter regression tasks are decomposed to realize the integration of communication and perception functions.

Benefits of technology

It significantly improves the accuracy and stability of channel reconstruction, improves the accuracy and robustness of perception parameters, reduces the dependence on the quality of training data, and has the underlying technical support of high accuracy and high robustness. It is suitable for the upgrade of cellular wireless networks and other computing cost-sensitive and perception-sensitive communication and perception integrated systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120567618A_ABST
    Figure CN120567618A_ABST
Patent Text Reader

Abstract

The invention discloses a dual-base-station sensing integrated system based on generative AI and a sensing method, and mainly solves the problems that an existing AI sensing model cannot analyze multi-dimensional coupling parameters, is poor in dynamic environment robustness and is complex in calculation. Comprising a transmitting link and a receiving link, and a reciprocal communication structure is arranged between the two communication base stations; a time-frequency-space feature embedding mechanism and an attention mechanism are introduced into the front end of the model, and meanwhile, a random mask is used for completing data enhancement in a pre-training stage; tasks are divided into a channel state information reconstruction and prediction task and a perception parameter regression task for decoupling, different design ideas and deployment strategies are used for the two tasks, and the whole model does not need to be trained. According to the method, the sensing precision and stability of the system can be effectively improved, the calculation complexity is remarkably reduced, and the method has wide applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of communication technology and further relates to communication perception integration technology, specifically a dual-base station synaesthesia integration system and perception method based on generative AI, which can be used to enhance the perception function of cellular wireless network base station facilities. Background Art

[0002] Communication-perception integration is a technology that deeply integrates communication and perception functions, enabling collaborative operation through shared hardware and spectrum resources. In practice, it utilizes communication signals or specialized waveforms to implement data transmission and perception functions such as target detection and tracking. Because perception parameters are implicitly encoded in the received signal and are subject to various interferences such as dynamic noise, parameter coupling, and multipath, efficient algorithms must be developed to accurately extract perception parameters and optimize inter-awareness collaboration to meet the perception needs of diverse services.

[0003] Traditional perception methods use structures such as multi-layer perceptrons and convolutional neural networks, and are trained using data sets collected in specific environments. These methods not only suffer from unstable performance, poor anti-interference capabilities, and poor real-time performance, but also have structural flaws in modeling long-term temporal dependencies of channel states. The paper [Naoumi S, Bazzi A, Bomfin R, et al. Complex neural network based joint AoA and AoD estimation for bistaticISAC [J]. IEEE Journal of Selected Topics in Signal Processing, 2024, 18(5): 842–856.] proposes a three-layer complex neural network activation MLP scheme to implement perception parameter regression. This scheme achieves relatively accurate perception accuracy while minimizing the number of parameters. However, the model needs to be retrained in each perception cycle, which continuously occupies device computing power and results in poor real-time perception performance. The paper [ALIKHANI S,CHARAN G,ALKHATEEBA.Large wireless model (LWM):A foundation model for wireless channels:abs / 2411.08872[A].2024] proposes a pre-trained model LWM based on the Transformer architecture, which can generate real-time frequency embedding to capture complex patterns in the channel. However, it lacks application optimization in perception and cannot form a complete set of perception methods. Summary of the Invention

[0004] The present invention aims to address the shortcomings of the above-mentioned prior art and propose a dual-base station synaesthesia integrated system and perception method based on generative AI to solve the problem of poor perception performance of traditional AI perception models in complex dynamic scenes. The present invention takes into account the natural similarities between the way natural language processing (NLP) technology handles problems such as missing information generation and prediction and the perception problems in communication perception integration. NLP word embedding can be compared to embedding at different frequencies in the channel state, and sentence embedding can be compared to embedding at different moments in the channel state. Therefore, applying a model construction method similar to NLP can improve the perception performance of the system and achieve collaborative optimization of communication and perception. By decoupling the attention mechanism from the embedding and learning the time-frequency-space three-dimensional features, the defects of traditional AI perception models such as the inability to parse multi-dimensional coupling parameters, poor robustness in dynamic environments, and complex calculations for frequent parameter updates are overcome. The present invention can be used to upgrade existing wireless cellular networks and can also be used in other communication perception integration systems that are sensitive to computational cost and perception accuracy, and has wide applicability.

[0005] To achieve the above objectives, the technical solutions of the present invention include the following:

[0006] A generative AI-based dual-base station synaesthesia integrated system consists of two communication base stations with identical internal structures and equivalent functions, and a dynamic multipath scattering channel connected between the two base stations; wherein the communication base stations include a transmitting link and a receiving link, and the two communication base stations have a reciprocal communication structure;

[0007] The transmission link of the communication base station includes a code mapping module, a spatial mapping module, an inverse fast Fourier transform IFFT module, a cyclic prefix adding CP module, a digital-to-analog conversion module and a radio frequency front-end processing module in sequence; wherein the code mapping module is used to modulate and encode the original communication bit stream, the spatial mapping module is used to map the modulated signal to the multi-antenna transmission structure, the IFFT module is used to convert the frequency domain signal into a time domain signal, the CP module is used to add a cyclic prefix to alleviate multipath interference, the digital-to-analog conversion module is used to convert the time domain discrete signal into an analog signal, and the radio frequency front-end processing module is used to complete up-conversion and radio frequency signal transmission;

[0008] The receiving link of the communication base station includes, in sequence, a radio frequency front-end processing module, an analog-to-digital conversion module, a cyclic prefix removal module, a fast Fourier transform (FFT) module, a de-spatial mapping module, and a communication data recovery module and a perception information extraction module connected in parallel; wherein, the radio frequency front-end processing module is used to down-convert and preliminarily process the received radio frequency signal, the analog-to-digital conversion module is used to convert the analog signal into a digital signal, the cyclic prefix removal module is used to remove the CP prefix in the received signal, the FFT module is used to convert the received signal from the time domain to the frequency domain, the de-spatial mapping module is used to recover single-channel communication data from the multi-antenna receiving structure, the communication data recovery module is used to restore the original communication bit information, and the perception information extraction module is used to extract the perception parameters of the environmental target from the received signal, thereby realizing the integration of communication and perception functions;

[0009] The dynamic multipath scattering channel is connected between two communication base stations to simulate the impact of a complex dynamic multipath environment on the transmitted and received signals. The channel includes a signal input module, a multipath channel module, a multipath merging module, a Gaussian white noise (AWGN) channel module, and a signal output module. The multipath channel module consists of a line-of-sight (LOS) path and N non-line-of-sight (NLOS) paths, where N is a positive integer greater than or equal to 2.

[0010] Furthermore, the two communication base stations mentioned above have a reciprocal communication structure, specifically: let the two communication base stations be base station A and base station B, where base station A includes a first transmitting link and a first receiving link, and base station B includes a second transmitting link and a second receiving link, the first transmitting link and the second receiving link constitute a first communication direction, and the second transmitting link and the first receiving link constitute a second communication direction, thereby realizing reciprocal communication between the two base stations.

[0011] Furthermore, the above-mentioned LOS path includes a free space path loss calculation unit and a path delay calculation unit, which are used to simulate the direct propagation path in free space; each of the NLOS paths includes a scatterer reflection response calculation unit, a Doppler frequency shift calculation unit, a free space path loss calculation unit and a path delay calculation unit, which are used to simulate the complex propagation behavior of the signal after being reflected by multiple scatterers in the environment.

[0012] A method for target positioning perception using a dual-base station synaesthesia integrated system based on generative AI, comprising the following steps:

[0013] (1) Let the two communication base stations be base station A and base station B. Base station A sends a preset preamble code sequence to base station B's second receiving link through a first transmitting link. Base station B receives the preamble code through the second receiving link and extracts pilot symbols to estimate channel state information.

[0014] (2) Base station B performs singular value decomposition on the estimated channel matrix to calculate precoding weights and combining weights; updates the receiving antenna combining weight parameters of the second receiving link based on the combining weights, and simultaneously feeds the precoding weights back to the first receiving link of base station A via the second transmitting link;

[0015] (3) Base station A updates the antenna precoding vector of the first transmission link based on the precoding weight fed back by base station B; and uses the updated antenna precoding vector to send a data frame including a pilot to base station B via the first transmission link;

[0016] (4) Base station B receives the data frame through the second receiving link and extracts the pilot symbols to complete the initial channel state information estimation; then the estimated channel state information is input into the generative AI model to generate the completed reconstructed channel information;

[0017] (5) Base station B executes the following two processing steps in parallel to complete the sensing task:

[0018] Feedback update process: Based on the reconstructed channel information output by the generative AI, the updated precoding and combining weights are solved, and the receiving antenna combination weight parameters of the second receiving link are updated according to the combining weights. At the same time, the precoding weights are fed back to base station A through the second transmitting link. Base station A updates the precoding parameters according to the feedback updated weights and enters the transmission process of the next data frame;

[0019] Communication perception solution process: Store the reconstructed channel information of the generative AI, perform communication data demodulation and perception parameter regression processing based on this information, input the obtained perception parameters into the target positioning module, and obtain the target position information through polar coordinate solution.

[0020] Compared with the prior art, the present invention has the following advantages:

[0021] First, the system of the present invention adopts a massive multiple-input multiple-output orthogonal frequency division multiplexing (Massive-MIMO-OFDM) architecture, and utilizes the high degree of freedom of the architecture in the frequency domain and spatial domain. Under the premise of not affecting the communication data transmission performance, it can implicitly embed perception parameter information in the communication signal, thereby realizing the integration of communication and perception functions. Through the structural optimization design of the pilot data in the system, especially the introduction of diversity strategy in the time domain, the generative AI model can make full use of the time diversity characteristics of the pilot in the process of channel state information (CSI) recovery and reconstruction, thereby significantly improving the accuracy and stability of channel reconstruction. In addition, this structure also facilitates the generative AI to further extract perception parameters of the reconstruction results, such as target position, speed, reflection characteristics, etc., providing high-precision and high-robustness underlying technical support for the communication and perception integrated system.

[0022] Second, since the present invention introduces a time-frequency-space feature embedding mechanism and a BERT network layer structure including an attention mechanism at the front end of the LSTM network layer when constructing a generative AI model, the architecture ensures long-distance modeling and cross-domain modeling capabilities, and can learn the nonlinear coupling characteristics of perception parameters, and has a natural advantage in perception performance over traditional methods; at the same time, the model uses random masks to complete data enhancement in the pre-training stage, ensuring the generalization ability of predicting full signal-to-noise ratio test data when trained with a single signal-to-noise ratio data set. Compared with traditional machine learning methods, it has lower dependence on the quality of training data and stronger system stability.

[0023] Third, the model architecture designed by the present invention is divided into two parts when decoupling tasks: channel state information reconstruction and prediction task and perception parameter regression task; and different design ideas and deployment strategies are used for the two tasks. The former uses a large-scale data set to build a regional perception intelligent agent and completes deployment in the cloud; the latter uses a low-complexity multi-layer perceptron to reduce regression delay and computational complexity, thereby having edge deployment capabilities; therefore, when performing adaptive training locally, only the multi-layer perceptron network with very few parameters needs to be updated, and there is no need to train the entire model. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a schematic diagram of the overall architecture of the system of the present invention;

[0025] Figure 2 It is a block diagram of the base station transmission link of the system of the present invention;

[0026] Figure 3 It is a block diagram of the dynamic multipath scatterer channel model of the system of the present invention;

[0027] Figure 4 is a block diagram of a base station receiving link of the system of the present invention;

[0028] Figure 5 Schematic diagram of the OFDM time-frequency resource grid of the signal transmitted by the system of the present invention;

[0029] Figure 6 It is a schematic diagram of the hierarchical structure of the generative AI in the system of the present invention;

[0030] Figure 7 Schematic diagram of the workflow for target perception using the system of the present invention;

[0031] Figure 8 is a polar coordinate positioning result diagram of the target perception method in the present invention;

[0032] Figure 9 is a performance gain curve of the target perception method of the present invention using generative AI;

[0033] Figure 10 This is a comparison chart of the perceptual performance of the method of the present invention and the existing method. DETAILED DESCRIPTION

[0034] The present invention will be further described below with reference to the accompanying drawings.

[0035] Example 1: Reference Figure 1-6 The present invention proposes a dual-base station synaesthesia integrated system based on generative AI, which is composed of two communication base stations with the same internal structure and equivalent functions, and a dynamic multipath scattering channel connected between the two base stations; wherein the communication base station includes a transmitting link and a receiving link, and the two communication base stations have a reciprocal communication structure;

[0036] The transmission link of the communication base station includes a code mapping module, a spatial mapping module, an inverse fast Fourier transform IFFT module, a cyclic prefix adding CP module, a digital-to-analog conversion module and a radio frequency front-end processing module in sequence; wherein the code mapping module is used to modulate and encode the original communication bit stream, the spatial mapping module is used to map the modulated signal to the multi-antenna transmission structure, the IFFT module is used to convert the frequency domain signal into a time domain signal, the CP module is used to add a cyclic prefix to alleviate multipath interference, the digital-to-analog conversion module is used to convert the time domain discrete signal into an analog signal, and the radio frequency front-end processing module is used to complete up-conversion and radio frequency signal transmission;

[0037] The receiving link of the communication base station includes, in sequence, a radio frequency front-end processing module, an analog-to-digital conversion module, a cyclic prefix removal module, a fast Fourier transform (FFT) module, a de-spatial mapping module, and a communication data recovery module and a perception information extraction module connected in parallel; wherein, the radio frequency front-end processing module is used to down-convert and preliminarily process the received radio frequency signal, the analog-to-digital conversion module is used to convert the analog signal into a digital signal, the cyclic prefix removal module is used to remove the CP prefix in the received signal, the FFT module is used to convert the received signal from the time domain to the frequency domain, the de-spatial mapping module is used to recover single-channel communication data from the multi-antenna receiving structure, the communication data recovery module is used to restore the original communication bit information, and the perception information extraction module is used to extract the perception parameters of the environmental target from the received signal, thereby realizing the integration of communication and perception functions;

[0038] The dynamic multipath scattering channel is connected between two communication base stations to simulate the impact of a complex dynamic multipath environment on the transmitted and received signals. The channel includes a signal input module, a multipath channel module, a multipath merging module, a Gaussian white noise (AWGN) channel module, and a signal output module. The multipath channel module consists of a line-of-sight (LOS) path and N non-line-of-sight (NLOS) paths, where N is a positive integer greater than or equal to 2.

[0039] In this embodiment, the two communication base stations have a reciprocal communication structure. Specifically, the two communication base stations are base station A and base station B, where base station A includes a first transmission link and a first reception link, and base station B includes a second transmission link and a second reception link. The first transmission link and the second reception link constitute a first communication direction, and the second transmission link and the first reception link constitute a second communication direction, thereby realizing reciprocal communication between the two base stations.

[0040] In this embodiment, the above-mentioned LOS path includes a free space path loss calculation unit and a path delay calculation unit, which are used to simulate the direct propagation path in free space; each of the NLOS paths includes a scatterer reflection response calculation unit, a Doppler frequency shift calculation unit, a free space path loss calculation unit, and a path delay calculation unit, which are used to simulate the complex propagation behavior of the signal after being reflected by multiple scatterers in the environment.

[0041] Example 2: Reference Figure 7 The present invention proposes a method for performing target positioning perception according to the system described in Example 1, and the specific implementation steps include the following:

[0042] Step 1) Let the two communication base stations be base station A and base station B respectively. Base station A sends a preset preamble code sequence to the second receiving link of base station B through the first transmitting link; base station B receives the preamble code through the second receiving link and extracts the pilot symbol to estimate the channel state information. The preamble code described in this embodiment is constructed in a frequency diversity manner. In the multiple-input multiple-output MIMO communication scenario, for the signal transmitted by a single antenna, the subcarrier dimension is spaced N by 1 / 4. t A pilot symbol is inserted into the subcarriers, where N t Indicates the number of transmitting antennas, used to reduce interference between multiple antennas.

[0043] Step 2) Base station B performs singular value decomposition on the estimated channel matrix to calculate precoding weights and combining weights; updates the receiving antenna combining weight parameters of the second receiving link based on the combining weights, and simultaneously feeds the precoding weights back to the first receiving link of base station A via the second transmitting link;

[0044] Step 3) Base station A updates the antenna precoding vector of the first transmission link based on the precoding weights fed back by base station B; and uses the updated antenna precoding vector to send a data frame containing a pilot to base station B via the first transmission link. The data frame described in this embodiment is constructed in a two-dimensional OFDM time-frequency resource grid with a dimension of N. sub ×N sym , where N sym Indicates the number of OFDM symbols in a single frame, N subIndicates the number of effective subcarriers excluding the guard band. In the MIMO scenario, the pilot symbols of the data frame are inserted at intervals in both the time and frequency dimensions, that is, in the time dimension, every interval S t Insert a pilot symbol every S in the frequency dimension f Insert a pilot symbol; the S t With S f Satisfy the following constraints:

[0045] where f dmax is the maximum Doppler shift, T sym is the OFDM symbol duration;

[0046] where τ max is the maximum delay spread, and Δf is the subcarrier spacing.

[0047] Step 4) Base station B receives the data frame through the second receiving link and extracts the pilot symbols to complete the initial channel state information estimation; the estimated channel state information is then input into the generative AI model to generate the completed reconstructed channel information. The generative AI model described in this step of this embodiment includes, in sequence, a channel matrix data enhancement module, an embedding coding module, an embedding synthesis module, a BERT network layer, an LSTM network layer, and a channel matrix reconstruction module, which are used to enhance and reconstruct the received channel state information; wherein, the channel matrix data enhancement module is used to apply a dynamically adjustable random mask to the input channel matrix data CSI to improve the robustness of the model to CSI disturbances; the embedding coding module includes feature embedding, frequency domain embedding, and time domain embedding, which are used to extract spatial dimension features, frequency domain position features, and time series features respectively; the embedding synthesis module is used to add the above embedding results and send them to the BERT network layer; the BERT network layer includes multiple Transformer encoder units, which are used to extract the time-frequency semantic information of CSI in a multi-dimensional feature space; the LSTM network layer is used to characterize the dynamic change characteristics of CSI and output the enhanced hidden state tensor to achieve temporal resolution enhancement; the channel matrix reconstruction module is used to restore the enhanced hidden state to a channel estimation matrix of the target size.

[0048] Step 5) Base station B executes the following two processing steps in parallel to complete the sensing task:

[0049] Feedback update process: Based on the reconstructed channel information output by the generative AI, the updated precoding and combining weights are solved, and the receiving antenna combination weight parameters of the second receiving link are updated according to the combining weights. At the same time, the precoding weights are fed back to base station A through the second transmitting link. Base station A updates the precoding parameters according to the feedback updated weights and enters the transmission process of the next data frame;

[0050] Communication perception solution process: Store the reconstructed channel information of the generative AI, perform communication data demodulation and perception parameter regression processing based on this information, input the obtained perception parameters into the target positioning module, and obtain the target position information through polar coordinate solution. The perception parameter regression processing described in this embodiment is specifically implemented by extracting perception parameters from the reconstructed channel information through the perception parameter regression model MLP. This model includes a three-layer perceptron network structure. Each layer gradually compresses the feature dimensions and outputs the perception parameter prediction value. Each layer uses the RELU activation function and outputs the perception result through the terminal fully connected layer.

[0051] Example 3: The overall architecture of the dual-base station communication perception integrated system based on generative AI proposed in this example is the same as that of Example 1. Figure 1-6 , further detailed description of the following core modules that constitute the system:

[0052] 1. Dual-base station communication architecture:

[0053] like Figure 1 As shown, a dual-base station communication perception integrated system based on generative AI of the present invention includes a base station A, a dynamic multipath scattering channel, and a base station B. The base stations A and B have the same internal structure. Base station A includes a first transmitting link and a first receiving link. Base station B includes a second receiving link and a second transmitting link. The transmitting and receiving links of base stations A and B are reciprocal. The dynamic multipath scattering channel is connected to base stations A and B, respectively, to realize the influence of complex dynamic multipath scenarios on the transmitting and receiving signals.

[0054] 2. Signal processing link:

[0055] like Figure 2 As shown, the first and second transmission chains of the base station include six sequential processing steps, namely, code mapping, space mapping, IFFT, CP addition, digital-to-analog conversion, and RF front-end processing.

[0056] like Figure 3 As shown, the dynamic scattering channel includes five steps: signal input, multipath channel, multipath merging, AWGN channel, and signal output. The multipath channel includes one LOS path and N NLOS paths. The LOS path includes two parts: free-space attenuation calculation and path delay calculation. The NLOS path includes four parts: scatterer reflection response calculation, Doppler effect calculation, free-space attenuation calculation, and path delay calculation.

[0057] like Figure 4As shown, the first and second receiving chains of the base station each include seven processing steps: RF front-end processing, analog-to-digital conversion, CP removal, FFT, spatial demapping, communication data recovery, and sensor extraction. The communication data recovery and sensor extraction steps are performed in parallel.

[0058] 3.OFDM resource allocation strategy:

[0059] like Figure 5 As shown, the OFDM resource grid of the signal transmitted by the dual-base station communication perception integrated system based on generative AI of the present invention includes preamble allocation and data frame. The preamble uses frequency diversity in MIMO to prevent interference, and is spaced N on the subcarrier. t Place a pilot symbol, where N t The data frame is constructed using a two-dimensional time-frequency resource grid, whose dimension is defined as N sub ×N sym , where N sym Indicates the number of OFDM symbols in a single frame, N sub is the number of effective subcarriers after deducting the guard band. In the case of MIMO, time diversity and frequency diversity are used to prevent interference, and the symbol time is separated by S t , subcarrier spacing S f Place a pilot symbol. Pilot interval S t and S f The following relationship must be satisfied:

[0060]

[0061] Among them, in the time dimension constraint, f dmax is the maximum Doppler shift, T sym is the total duration of OFDM symbols; in the frequency dimension constraint, τ max is the maximum delay spread, and Δf is the subcarrier spacing.

[0062] 4. Generative AI Perception Model:

[0063] like Figure 6 As shown, the generative AI model of the present invention includes the following hierarchical structure:

[0064] a) Channel matrix data enhancement layer:

[0065] The channel matrix data is a triplet of dimensions (BatchSize, SequenceLength, FeatureDim), where BatchSize is the input batch size, SequenceLength is the number of subcarriers, and FeatureDim represents the real and imaginary parts of the spatial dimension after the channel matrix data is decomposed. These three dimensions are subsequently referred to as "B," "L," and "F."

[0066] The channel matrix data enhancement is to use a dynamically adjustable random mask to randomly set the channel matrix data CSI to zero to enhance stability, and is calculated using the following formula:

[0067] Mask~Bernoulli(rate),CSI masked =Mask⊙CSI (1)

[0068] The channel matrix data CSI and the mask Mask perform a Hadamard product operation, and the mask Mask satisfies the Bernoulli distribution with a mask rate rate.

[0069] b) Feature Embedding, Frequency Domain Embedding, and Time Embedding Layer: After masking, the channel matrix data (CSI) is fed into the embedding layer. This layer is designed to inversely map the signal generation mechanism of the communication system's physical layer, allowing the model to learn the time-frequency characteristics of the channel information. Time domain embedding encodes the relative position of the input matrix in the time series, similar to sentence embedding in natural language processing (NLP). Frequency domain embedding encodes the subcarrier position of the input data, similar to word embedding in NLP.

[0070] The feature embedding module is used to capture the spatial characteristics of the channel matrix data CSI. This module flattens the spatial dimension tensor and complex tensor of the antenna pair and projects them into the model processing space through a fully connected layer, thereby converting them into a learnable spatial correlation matrix to adapt to the model interface. This process is expressed using the following formula:

[0071]

[0072] Frequency domain embedding involves mapping discrete frequency indexes to a continuous space using an embedding matrix of size L×F. For a given sequence length L, the system first generates a one-hot encoding sequence of [0, L-1]. After embedding, the output dimension is transformed to (B, L, F).

[0073] In time-domain embedding, discrete time indices are mapped to continuous space via an embedding matrix of size L×F. For a given sample length P, the system first generates a one-hot encoding sequence of [0, P-1]. After embedding, the output dimension is transformed to (B, L, F).

[0074] The mathematical essence of frequency domain embedding and time domain embedding can be regarded as projecting the frequency domain and time domain features of the channel matrix data CSI into a high-dimensional Hilbert space.

[0075] c) Embedded synthesis layer:

[0076] The embedding synthesis layer adds the output matrices of feature embedding and frequency domain embedding to adapt to the input requirements of BERT. This process is expressed using the following formula:

[0077]

[0078] d) BERT network layer:

[0079] The BERT layer is an encoder network consisting of 6 layers of Transformer encoder units stacked together. Each encoder layer contains:

[0080] (d1) Self-Attention Head MHSA: 8 independent heads, each with a key / value / query vector dimension of 32

[0081] (d2) Feedforward network FFN: two layers of full connection, weight matrices F×3072 and 3072×F

[0082] (d3) Normalization layer Norm: The trainable parameters are 2×F (scaling factor and bias term)

[0083] The encoder output maintains the input dimension (B, L, F), is added to the initial embedding matrix through the residual connection, and then fed into the subsequent network.

[0084] e) Dimension reshaping and embedding synthesis layer:

[0085] The dimension reshaping and embedding synthesis layer adds the hidden state of (B, L, F) to the time embedding, then replicates it P times along the sequence dimension to form a four-dimensional tensor of (B, P, L, F), which is then flattened into a data matrix of dimension (B, P × L, F).

[0086] f) LSTM network layer: Considering that the pilot symbols adopt a diversity transmission strategy in the time-frequency-space dimensions and have the characteristics of cross-domain parameter coupling, the long short-term memory (LSTM) network is used to achieve temporal resolution enhancement.

[0087] The long short-term memory (LSTM) network layer is a single-layer LSTM network structure, and its parameter interface includes:

[0088] (f1) Input gate weight matrix: F×F (input to hidden layer)

[0089] (f2) Forget gate weight matrix: F×F

[0090] (f3) Output gate weight matrix: F×F

[0091] (f4) Cell state transformation matrix: F×F

[0092] The LSTM output dimension of the long short-term memory (LSTM) network layer is maintained at (B, P × L, F).

[0093] g) Obtain an enhanced channel estimation matrix: The output (B, P × L, F) in the previous layer is restored to a (B, P, L, F) structure through an inverse reshaping operation, and its output is output as the reconstructed and predicted channel matrix result, providing a high-order feature expression for subsequent regression.

[0094] h) MLP perception extraction layer:

[0095] The MLP perception extraction layer is a low-complexity regression interpreter that gradually compresses the feature space through linear transformation until the perception parameter prediction value is obtained. Its input dimensions are (B, P, L, F) and its internal structure is three layers, with the dimensions of each layer being:

[0096] (h1)MLP_1:(B,P,L,F)→(B,P,L,F / 2)

[0097] (h2)MLP_2:(B,P,L,F / 2)→(B,P,L,F / 4)

[0098] (h3)MLP_3:(B,P,L,F / 4)→(B,P,L,F / 8)

[0099] RELU is used between each layer for activation, and finally a fully connected layer is used to extract perception parameters.

[0100] This embodiment includes but is not limited to the following configurations:

[0101] (1) Transmitting link part:

[0102] Code mapping: Use any coding method or no coding; constellation mapping preferably uses 16QAM, 64QAM, 256QAM or higher-order mapping;

[0103] Spatial mapping: uses the mapping method defined by the preamble and data frame;

[0104] IFFT: 1024 or higher order points are preferred;

[0105] Add CP: The CP length is greater than the delay spread constraint of the farthest scattering path;

[0106] Digital-to-analog conversion: select AD9176, AD9177 or AD9175 chips;

[0107] RF front-end processing: For 2D positioning, it must have at least a ULA array structure; for 3D positioning, it must have at least a T-type array structure;

[0108] (2) Dynamic scattering channel part:

[0109] Multipath channel: Prefer the indoor office, urban, rural, or indoor factory models of 3GPP TR38.901V18;

[0110] AWGN channel 2.4: noise injection with a preferred signal-to-noise ratio of 0 to 40 dB;

[0111] (3) Receiving link part:

[0112] RF front-end processing: For 2D positioning, it must have at least a ULA array structure; for 3D positioning, it must have at least a T-type array structure;

[0113] Analog-to-digital conversion: Use AD9689, AD9694 or AD9213 chips;

[0114] Remove CP: The CP length must be the same as the CP adding step;

[0115] FFT: The number of FFT points must be the same as the IFFT step;

[0116] Demapping: Demapping method limited by preamble and data frame;

[0117] Communication data recovery: Demapping and decoding must be done in the same way as the encoding and mapping steps;

[0118] (IV) Transmitting signal part:

[0119] Pilot symbol: Generated using maximum length sequence (MLS), LFSR, or Gold Code.

[0120] Example 4: The overall implementation steps of the target positioning perception method proposed in this example are the same as those in Example 2. Figure 7-8 , further details are given on the two stages of initial channel detection and data frame transmission:

[0121] Phase 1: Initial channel detection

[0122] Step s101) Preamble transmission: The first transmitting link of base station A transmits a preamble to the first and second receiving links via a dynamic scattering channel;

[0123] Step s102) Preamble reception: The base station receiving link receives and decodes the preamble;

[0124] Step s103) Channel estimation:

[0125] The channel estimation mentioned above is that the base station receiving link estimates the initial channel state information by extracting the pilot symbols at fixed positions in the preamble. This process is achieved by minimizing the cost function:

[0126]

[0127] Among them, X is the received signal, Y is the received signal, is the channel estimation result; (·) * is the calculated conjugate, (·) H is to calculate the conjugate transpose, is to calculate the partial derivative.

[0128] Step s104) diagonalization:

[0129] The diagonalization is that the base station receiving link is calculated by The result of singular value decomposition is to obtain the precoding weight W pre and the combined weight W comb , the process is expressed by the formula:

[0130]

[0131] Where U and V are the left and right singular value unitary matrices, S is the singular value diagonal matrix, (·) T is the matrix transpose;

[0132] Step s105) Precoding feedback: The base station receiving link uses the reciprocal base station transmitting link to feed back precoding weights to the base station receiving link via the dynamic scattering channel;

[0133] Step s106) Parameter update: The base station transmit link updates the antenna precoding vector using the fed-back precoding weights;

[0134] Step s107) Phase switching: After the base station transmit link updates the antenna precoding vector, it completes diagonalization and enters the data frame transmission process;

[0135] Phase 2: Data frame transmission

[0136] Step s201) ​​Sending data frames: The base station transmit link sends data frames to the base station receive link via the dynamic scattering channel;

[0137] Step s202) Data symbol reception: The base station receiving link receives data OFDM symbols one by one;

[0138] Step s203) The base station receiving link determines whether a frame of data reception has been completed. If yes, it proceeds to step s204); if not, it proceeds to step s202)

[0139] Step s204) Data frame channel estimation:

[0140] The data frame channel estimation is that the base station receiving link extracts the pilot symbols inserted in the data frame to obtain the original channel state information, and feeds it into the generative AI model to reconstruct and predict the channel state information;

[0141] Step s205) At this time, the base station receiving link completes the precoding weight feedback and data solution steps in parallel:

[0142] a) Precoding weight feedback:

[0143] a.1) The base station receive link calculates the precoding and combining weights using the estimated channel state information;

[0144] a.2) The base station receive link feeds back the precoding weights to the base station receive link via a reciprocal base station transmit link through a dynamic scattering channel;

[0145] a.3) The base station transmit link updates the antenna precoding vector using the fed-back precoding weights;

[0146] a.4) After the base station transmit link updates the antenna precoding vector, it updates the diagonalization parameters, starts transmission of the next data frame, and returns to step s201);

[0147] b) Data solution:

[0148] b.1) The base station receives the link and stores the estimated channel state information, then performs communication and sensing functions in parallel;

[0149] For communication functions: the base station receiving link recovers the communication data through the estimated channel state information;

[0150] For the perception function: the base station receiving link realizes the perception parameter regression through the MLP layer, and then performs polar coordinate positioning based on the perception parameters to obtain the target position; the perception parameters include the departure angle AOD, arrival angle AOA, and path delay TOA. The polar coordinate positioning method is as follows Figure 8 As shown, by pairing the three perception parameters in pairs, three estimated coordinates can be obtained, and the three estimated coordinates can form a prediction area, thereby realizing the target perception and positioning function.

[0151] The effects of the present invention will be further described below in conjunction with simulation experiments.

[0152] 1. Simulation conditions:

[0153] The simulation experiment of the present invention is carried out in the hardware environment of RTX3090Ti and the software environment of Python3.12 and MATLAB2024b.

[0154] 2. Simulation content:

[0155] The MATLAB 5G Toolbox was used to generate a data set, and performance simulation was performed using the generative AI-accurately estimated channel response and the ideal channel response.

[0156] 3. Simulation results:

[0157] Figure 9 This is the performance gain curve result of applying generative AI. It can be seen that the performance curve of the target perception method of the dual-base station communication perception integrated system based on generative AI of the present invention under the full signal-to-noise ratio (rectangular node dotted line) is significantly improved compared with regression without using the generative AI model (X node dotted line); compared with the perception performance under the ideal channel response (circular node solid line), similar performance is achieved.

[0158] In order to highlight the beneficial effects of the present invention, Figure 10 Further description:

[0159] MSE represents the mean square error between the perceived coordinates and the true coordinates. The smaller the value, the closer the perceived target position is to the actual position.

[0160] The mean square error (MSE) indicator of the target perception method of the dual-base station communication perception integrated system based on generative AI of the present invention is the highest in detection accuracy compared with popular methods and ablation models.

[0161] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0162] The above simulation analysis proves the correctness and effectiveness of the method proposed in the present invention.

[0163] Parts of the present invention that are not described in detail belong to common knowledge among those skilled in the art.

[0164] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Obviously, for professionals in this field, after understanding the content and principles of the present invention, they may make various modifications and changes in form and details without departing from the principles and structure of the present invention. However, these modifications and changes based on the ideas of the present invention are still within the scope of protection of the claims of the present invention.

Claims

1. A dual-base station synaesthesia integrated system based on generative AI, characterized by: The system consists of two communication base stations with the same internal structure and equivalent functions, and a dynamic multipath scattering channel connected between the two base stations; wherein the communication base stations include a transmitting link and a receiving link, and a reciprocal communication structure is formed between the two communication base stations; The transmission link of the communication base station includes a code mapping module, a spatial mapping module, an inverse fast Fourier transform IFFT module, a cyclic prefix adding CP module, a digital-to-analog conversion module and a radio frequency front-end processing module in sequence; wherein the code mapping module is used to modulate and encode the original communication bit stream, the spatial mapping module is used to map the modulated signal to the multi-antenna transmission structure, the IFFT module is used to convert the frequency domain signal into a time domain signal, the CP module is used to add a cyclic prefix to alleviate multipath interference, the digital-to-analog conversion module is used to convert the time domain discrete signal into an analog signal, and the radio frequency front-end processing module is used to complete up-conversion and radio frequency signal transmission; The receiving link of the communication base station includes, in sequence, a radio frequency front-end processing module, an analog-to-digital conversion module, a cyclic prefix removal module, a fast Fourier transform (FFT) module, a de-spatial mapping module, and a communication data recovery module and a perception information extraction module connected in parallel; wherein, the radio frequency front-end processing module is used to down-convert and preliminarily process the received radio frequency signal, the analog-to-digital conversion module is used to convert the analog signal into a digital signal, the cyclic prefix removal module is used to remove the CP prefix in the received signal, the FFT module is used to convert the received signal from the time domain to the frequency domain, the de-spatial mapping module is used to recover single-channel communication data from the multi-antenna receiving structure, the communication data recovery module is used to restore the original communication bit information, and the perception information extraction module is used to extract the perception parameters of the environmental target from the received signal, thereby realizing the integration of communication and perception functions; The dynamic multipath scattering channel is connected between two communication base stations to simulate the impact of a complex dynamic multipath environment on the transmitted and received signals. The channel includes a signal input module, a multipath channel module, a multipath merging module, a Gaussian white noise (AWGN) channel module, and a signal output module. The multipath channel module consists of a line-of-sight (LOS) path and N non-line-of-sight (NLOS) paths, where N is a positive integer greater than or equal to 2.

2. The method according to claim 1, wherein: The two communication base stations have a reciprocal communication structure, specifically: let the two communication base stations be base station A and base station B, where base station A includes a first transmission link and a first reception link, and base station B includes a second transmission link and a second reception link, the first transmission link and the second reception link constitute a first communication direction, and the second transmission link and the first reception link constitute a second communication direction, thereby realizing reciprocal communication between the two base stations.

3. The method according to claim 1, wherein: The LOS path includes a free space path loss calculation unit and a path delay calculation unit, which are used to simulate the direct propagation path in free space; each NLOS path includes a scatterer reflection response calculation unit, a Doppler frequency shift calculation unit, a free space path loss calculation unit and a path delay calculation unit, which are used to simulate the complex propagation behavior of the signal after being reflected by multiple scatterers in the environment.

4. A method for target positioning perception according to the system of claim 1, characterized in that: The implementation steps include the following: (1) Let the two communication base stations be base station A and base station B. Base station A sends a preset preamble sequence to base station B's second receiving link through a first transmitting link; Base station B receives the preamble code through the second receiving link and extracts the pilot symbol to estimate the channel state information; (2) Base station B performs singular value decomposition on the estimated channel matrix and calculates the precoding weights and combining weights; and updating the receiving antenna combination weight parameters of the second receiving link according to the combination weight, and simultaneously feeding back the precoding weight to the first receiving link of base station A through the second transmitting link; (3) Base station A updates the antenna precoding vector of the first transmit link based on the precoding weights fed back by base station B; and using the updated antenna precoding vector to send a data frame including a pilot to base station B through the first transmission link; (4) Base station B receives the data frame through the second receiving link and extracts the pilot symbols to complete the initial channel state information estimation; then the estimated channel state information is input into the generative AI model to generate the completed reconstructed channel information; (5) Base station B executes the following two processing steps in parallel to complete the sensing task: Feedback update process: Based on the reconstructed channel information output by the generative AI, the updated precoding and combining weights are solved, and the receiving antenna combination weight parameters of the second receiving link are updated according to the combining weights. At the same time, the precoding weights are fed back to base station A through the second transmitting link. Base station A updates the precoding parameters according to the feedback updated weights and enters the transmission process of the next data frame; Communication perception solution process: Store the reconstructed channel information of the generative AI, perform communication data demodulation and perception parameter regression processing based on this information, input the obtained perception parameters into the target positioning module, and obtain the target position information through polar coordinate solution.

5. The method according to claim 4, characterized in that: The preamble code in step (1) is constructed in a frequency diversity manner. In a multiple-input multiple-output MIMO communication scenario, for a signal transmitted by a single antenna, the preamble code is constructed in a frequency diversity manner. t A pilot symbol is inserted into the subcarriers, where N t Indicates the number of transmitting antennas, used to reduce interference between multiple antennas.

6. The method according to claim 4, wherein: The data frame in step (3) is constructed in a two-dimensional OFDM time-frequency resource grid with a dimension of N. sub ×N sym , where N sym Indicates the number of OFDM symbols in a single frame, N sub Indicates the number of valid subcarriers excluding the guard band.

7. The method according to claim 6, characterized in that: In the MIMO scenario, the pilot symbols of the data frame are inserted at intervals in both the time and frequency dimensions, that is, in the time dimension, every interval S t Insert a pilot symbol every S in the frequency dimension f Insert a pilot symbol; the S t With S f Satisfy the following constraints: where f dmax is the maximum Doppler shift, T sym is the OFDM symbol duration; where τ max is the maximum delay spread, and Δf is the subcarrier spacing.

8. The method according to claim 4, wherein: The generative AI model described in step (4) includes, in sequence, a channel matrix data enhancement module, an embedding coding module, an embedding synthesis module, a BERT network layer, an LSTM network layer, and a channel matrix reconstruction module, which are used to enhance and reconstruct the received channel state information.

9. The method according to claim 8, characterized in that: The channel matrix data enhancement module is used to apply a dynamically adjustable random mask to the input channel matrix data CSI to improve the model's robustness to CSI disturbances; the embedded coding module includes feature embedding, frequency domain embedding, and time domain embedding, which are used to extract spatial dimension features, frequency domain position features, and time series features respectively; the embedding synthesis module is used to add the above embedding results and send them to the BERT network layer; the BERT network layer contains multiple Transformer encoder units, which are used to extract the time-frequency semantic information of CSI in the multi-dimensional feature space; the LSTM network layer is used to characterize the dynamic change characteristics of CSI and output the enhanced hidden state tensor to achieve temporal resolution enhancement; the channel matrix reconstruction module is used to restore the enhanced hidden state to a channel estimation matrix of the target size.

10. The method according to claim 4, characterized in that: The perception parameter regression processing described in step (5) is achieved by extracting perception parameters from the reconstructed channel information through the perception parameter regression model MLP. The model includes a three-layer multi-layer perceptron network structure, each layer gradually compresses the feature dimension and outputs the perception parameter prediction value, wherein each layer adopts the RELU activation function and outputs the perception result through the terminal fully connected layer.