Deep learning-based multi-view fusion environment perception method assisted by intelligent reflecting surface
By deploying intelligent reflecting surfaces (IRS) in the ISAC system and combining it with deep learning technology to dynamically optimize the reflection phase and fuse multi-view point cloud information, the problem of insufficient perception capability of the ISAC system in complex environments is solved, and high-precision three-dimensional environment reconstruction and information fusion are achieved.
Patent Information
- Application Number
- CN202510714526.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-26
Smart Images

Figure CN120707734A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication technology, and in particular to a multi-perspective environment perception method assisted by an intelligent reflective surface based on deep learning. Background Art
[0002] Integrated Sensing and Communication (ISAC) is a key technology in next-generation wireless communication systems. By sharing spectrum and communication infrastructure, it integrates communication and environmental perception functions. Compared to traditional separate systems, ISAC systems offer higher resource utilization efficiency, system integration, and energy efficiency. They are widely used in scenarios requiring high communication and perception performance, such as intelligent transportation and autonomous driving. In ISAC systems, communication and perception tasks share the same spectrum and hardware resources, significantly improving overall system performance and flexibility. Intelligent Reflecting Surfaces (IRS), a key technology in ISAC systems, precisely control the phase, amplitude, and direction of signals by deploying multiple independently adjustable reflective elements, thereby optimizing signal transmission paths. IRS not only significantly enhances signal quality in communication systems but also provides additional reflection paths to overcome environmental obstructions, extend sensing range, and improve accuracy. By leveraging IRS technology, ISAC systems can overcome the limitations of traditional sensing technologies in complex environments, further enhancing environmental perception capabilities and improving overall system performance. As a powerful nonlinear modeling tool, deep learning demonstrates its unique advantages in ISAC systems, enabling them to enhance processing capabilities and generalize better in complex environments. Within ISAC systems, deep learning can be used for signal recovery, environmental perception, target recognition, and IRS control strategy optimization. Deep learning automatically extracts multi-level features from raw signals, optimizing signal processing and enhancing system robustness and accuracy. By training neural networks, ISAC systems can adapt to dynamic channel changes in real time and simultaneously optimize intelligent reflector (IRS) phase control strategies and spectrum / power resource allocation strategies, thereby improving overall system performance and resource efficiency. Summary of the Invention
[0003] The purpose of this invention is to use neural networks in an ISAC system to process received signal sequences from different perspectives, enabling multi-perspective fusion environmental perception. By deploying multiple intelligent reflecting surfaces (IRS) in a wireless environment, the invention acquires and fuses point cloud information from different regions from multiple perspectives to achieve high-precision 3D point cloud reconstruction, enhancing the system's perception capabilities.
[0004] Technical Solution: To achieve the above-mentioned purpose, the present invention adopts the following technical solution, which is implemented in the following steps:
[0005] A multi-view fusion environment perception method based on deep learning and assisted by intelligent reflective surfaces includes the following steps:
[0006] Step 1: To achieve environmental perception, multiple access points (APs) transmit pilot signals s. The signal transmission process is divided into F frames. In each frame, echo signals reflected by scatterers (such as walls and furniture) and smart reflective surfaces in the environment are collected to obtain the received signal y for all AP frames. The transmitted signal s is paired with the corresponding channel matrix H to construct the data set required for neural network training.
[0007] Step 2: Based on the proposed multi-view fusion perception neural network architecture, multiple intelligent reflecting surfaces (IRS) are deployed in the perception scene to obtain point cloud information of scatterers from different perspectives. The central unit of the system receives the pilot signal data and echo signal data transmitted from the AP, and uses an adaptive reflector dynamic optimization network based on a long short-term memory network (LSTM) and a deep neural network (DNN) to calculate and adjust the phase of each IRS frame by frame, and convert the transmitted signal s, received signal y, and reflector phase of each frame into the corresponding data. Spliced into At the end of the F frame, the IRS phase is optimized to a near-ideal state through the adaptive reflection surface optimization network to achieve optimal signal reflection.
[0008] Step 3: Optimize the F-frame data Input is a deep perspective perception network, which consists of a residual network (ResNet), LSTM and DNN. The network processes the echo signal sequences of all time frames to directly obtain regional point cloud information under different perspectives.
[0009] Step 4: The point cloud information from different perspectives is fed into the multi-perspective fusion perception network. This network introduces a 3D convolutional neural network (3DCNN) and a self-attention mechanism to perform feature alignment and weighted fusion on the local point cloud information from multiple perspectives, ultimately restoring a high-precision, complete three-dimensional environment point cloud.
[0010] Furthermore, the step 1 specifically includes:
[0011] 1.1) There are K full-duplex access points (APs) with multiple antennas in the space, each equipped with M transmitting antennas and M receiving antennas. The main functions of each AP include sending pilot signals to other APs or reflective surfaces, and receiving echo signals returned from APs, reflective surfaces, and other scattering objects in the environment. The viewing angle can be divided according to a single AP, or adjacent APs with highly overlapping sensing areas can be divided into one viewing angle. To enhance signal quality and expand the viewing angle of the communication link, the system deploys I intelligent reflective surfaces in the indoor environment, where each IRS consists of J micro-elements, capable of reflecting and adjusting signals in different directions. In addition, a central unit is deployed in the system to coordinate the work of all APs and IRSs. This unit is connected to all APs and IRSs, and its main task is to process the pilot signal data and echo signal data transmitted by the APs and optimize the data.
[0012] 1.2) The perception space where the scatterer is located is called ROI. The ROI will be evenly divided into A small cube, [H x ,W y ,L z ] are the height, width and length of the target space respectively, [h x ,w y ,l z ] are the height, width and length of the small cube respectively; the entire ROI space is represented as x=[x1,x2,...,x n ] T ; where x n ∈[0,1],x n =x1,...,x N , represents the scattering coefficient of the nth scatterer. The final ROI space x n It can be expressed as the point cloud distribution in the specific area where the scatterer is located, and the specific distribution position and state of the scatterer can be determined based on this.
[0013] 1.3) The signal transmission process is divided into F frames. The multipath channel H mainly consists of three parts: H = H LOS +H IRS +H S , H LOS and H IRS is the path that is independent of scatterers, namely the line-of-sight (LOS) AP-AP path and the AP-IRS-AP path, H S It is a scattering path. Since H LOS and H IRS So the signal sending process only considers H S . H SDenotes the scattered multipath channel associated with the scatterer. Three paths are considered here: AP-scatterer-AP, AP-scatterer-IRS-AP, and AP-IRS-scatterer-AP. Other paths with more scattering are ignored due to their high attenuation. The formula for the path AP-scatterer-AP is H AP→S→AP =h AP→S ⊙v AP→S diag(x)h S→AP ⊙v S→AP , h AP→S and h S→AP is the path loss matrix from AP to scatterer and from scatterer to AP, v AP→S and v S→AP is the corresponding occlusion matrix, x is the scattering coefficient vector, and ⊙ is the Hadamard product;
[0014] Similarly, the other two paths can be expressed as follows:
[0015] h AP→R,i represents the channel matrix from all antennas to the i-th IRS, h R,i→S represents the channel matrix from the i-th IRS to the scatterer, v R,i→S is the occlusion matrix from the i-th IRS to the scatterer, Θ i represents the reflection phase matrix of the i-th IRS;
[0016] h S→R,i represents the channel matrix from the scatterer to the i-th IRS, v S→R,i is the corresponding occlusion matrix, h R,i→AP represents the channel matrix from the i-th IRS to the AP;
[0017] The final received signal processed by the central node is y = H S s+w. s is the user's transmitted signal in each frame, and w is Gaussian white noise. Splice to And used as features of the training set, validation set and test set, the scatterer distribution point cloud is used as the label.
[0018] Furthermore, the step 2 specifically includes:
[0019] Construct an adaptive reflector optimization network. Each frame of the adaptive reflector optimization network is composed of a recurrent neural network architecture and a multi-layer perceptron (MLP). The architecture of the recurrent neural network is a bidirectional long short-term memory network (LSTM) and a double-layer stacked LSTM. The network dynamically adjusts the number of nodes of the recurrent neural network according to the number of time frames. That is, in the fth frame, the number of nodes of the recurrent neural network is f. The first frame input of the network is the initial y1, s1, Spliced In the fth frame, all y1:y of the previous f-1 frames f-1 ,s1:s f-1 , Splice to As input. Before the end of each frame, the hidden layer output h of the last layer will pass through the DNN of the R layer and output the reflection surface coefficient That is, at the end of each frame, the IRS is updated to the optimized phase coefficient. This process continuously improves the quality of the received signal.
[0020] Furthermore, the step 3 specifically includes:
[0021] Construct a deep perspective perception network. The network consists of a residual long short-term memory network (RES-LSTM) unit, a deep time module composed of a recurrent neural network, and a multi-layer perceptron (MLP). The RES-LSTM unit is based on the LSTM unit. The output of the hidden layer of the LSTM unit is adjusted to the same dimension as the input through a linear transformation, and then added to the input as the final unit output. The introduction of the LSTM with a residual structure can effectively alleviate the gradient disappearance problem caused by the increase in network depth. At the end of the F frame, the IRS phase adjusts the reflection coefficient through the reflection surface optimization network, and the received echo signal is significantly improved. The optimized F frame data is used The input is a deep perspective perception network, which passes through multi-layer RES-LSTM, bidirectional LSTM, and multi-layer LSTM in sequence. The time series features output by RES-LSTM are fed into the bidirectional LSTM module, which processes data in both forward and backward directions. The output is then concatenated and fed into the multi-layer LSTM. The cell state C of the last layer of LSTM is input into the DNN of the R layer, and the network output is the regional point cloud state of the i-th perspective. in Represents the splicing information of frame 1 to frame F and, The mapping rule is used to directly recover the regional scatterer point cloud state from the received signal, reflection surface phase and other data.
[0022] Furthermore, the step 5 specifically includes:
[0023] Construct a multi-view fusion perception network, which uses a 3D convolutional neural network (3D-CNN) to capture three-dimensional information in space. It has a total of 5 layers of 3D structure, each of which includes a 3D convolution layer, a 3D batch normalization layer, an activation function, and an efficient self-attention module. Through the deep perspective perception network, the system reconstructs the perspective point cloud information from multiple different perspectives. The point cloud information from different perspectives is fed into the network for point cloud fusion. After being spliced, the point cloud states of different perspectives pass through 5 3D structures. The output of the last 3D structure passes through the fully connected layer, and the result is the final fused ROI high-precision point cloud information. The mean square error is used as the loss function in training, and the formula is
[0024] The present invention has the following beneficial effects: By introducing a multi-perspective fusion perception mechanism assisted by an intelligent reflecting surface (IRS), combined with an adaptive reflecting surface optimization network and a deep learning-based multi-perspective point cloud fusion network, the present invention significantly improves the accuracy and integrity of wireless environment perception. Compared to existing perception methods that rely solely on a single perspective or static reflections, the present invention dynamically optimizes the IRS reflection phase, effectively mitigating occlusion effects, and fuses echo information from multiple time frames and perspectives to achieve high-precision, globally consistent 3D environment point cloud reconstruction with enhanced robustness, real-time performance, and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A schematic diagram of a scenario for a multi-view fusion environment perception method assisted by intelligent reflective surfaces;
[0026] Figure 2 Diagram of the multi-view fusion perception neural network architecture for intelligent reflective surface assistance;
[0027] Figure 3 Optimize network graph for adaptive reflective surface;
[0028] Figure 4 Perceive network graphs for a deep perspective;
[0029] Figure 5 It is a multi-view fusion perception network;
[0030] Figure 6 Comparison of the effects of regional spatial point clouds under single-view and multi-view fusion conditions;
[0031] Figure 7 The figure shows the perception error of the scatterer point cloud distribution when the solution of the present invention is compared with the comparative solution under different signal-to-noise ratios. DETAILED DESCRIPTION
[0032] The present invention will be further described below with reference to the accompanying drawings.
[0033] A multi-view fusion environment perception method based on deep learning and assisted by intelligent reflective surfaces, with the following specific steps:
[0034] Scenes such as Figure 1As shown in the figure, the scatterers in the sensing space are divided into voxel units as shown in the figure; the space contains four multi-antenna full-duplex access points (APs), two intelligent reflecting surfaces (IRSs), the space to be sensed ROI and a central unit, which is responsible for coordinating the work of all APs and IRSs;
[0035] like Figure 2 As shown in the figure, to achieve environmental perception, multiple access points (APs) transmit pilot signals s. The signal transmission process is divided into eight time frames. The received signal y is obtained for each frame, and the transmitted signal s and channel matrix H are used to generate a neural network dataset. The IRS reflector coefficient is dynamically adjusted through an adaptive reflector optimization network to optimize the echo signal quality. Then, a single-view depth perception network is used to recover local environmental point cloud information from the received signal. Point cloud information from different perspectives is fused through a multi-view fusion perception network to achieve high-precision 3D point cloud reconstruction. The specific implementation process is as follows:
[0036] Step 1 specifically includes:
[0037] Step 1.1: There is a multi-antenna AP, four single-antenna users (UEs), two intelligent reflecting surfaces (IRSs), and random scatterers in the space. Two adjacent APs near the reflecting surface are merged into one view. The transmitted signal from the UE passes through the scatterers and IRSs to the AP. All receiving antennas of each AP will receive the pilot signals sent by the transmitting antennas of other APs. The received signal of the mth receiving antenna in the kth BS is y k,m =h k,m s+w. Ignore paths with three or more reflections.
[0038] Step 1.2: The designated spatial region where the scatterer is located is called ROI. The ROI will be evenly divided into A small cube, [H x ,W y ,L z ] are the height, width and length of the target space respectively, [h x ,w y ,l z ] are the height, width and length of the small cube respectively; the entire ROI space is represented by x=[x1,x2,...,x n ] T ; where x n ∈[0,1],x n =x1,...,x N , represents the scattering coefficient of the nth scatterer. The final ROI space x n It can be expressed as the point cloud distribution in the specific area where the scatterer is located, and the specific distribution position and state of the scatterer can be determined accordingly.
[0039] Step 1.3: Divide the signal transmission process into F frames. The multipath channel H mainly consists of three parts: H = H LOS +H IRS +H S , H LOS and H IRS is the path that is independent of scatterers, namely the line-of-sight (LOS) AP-AP path and the AP-IRS-AP path, H S It is a scattering path. Since H LOS and H IRS So the signal sending process only considers H S . H S Denotes the scattered multipath channel associated with scatterers. Three paths are considered: AP-scatterer-AP, AP-scatterer-IRS-AP, and AP-IRS-scatterer-AP. Other paths with more scattering are ignored due to their high attenuation. The formula for the AP-scatterer-AP path is H AP→S→AP =h AP→S ⊙v AP→S diag(x)h S→AP ⊙v S→AP , where h AP→S and h S→AP is the path loss matrix from AP to scatterer and from scatterer to AP, v AP→S and v S→AP is the corresponding occlusion matrix, x is the scattering coefficient vector, ⊙ is the Hadamard product, and similarly, the other two paths can be expressed in this way. The final received signal processed by the central node is y = H S s+w. s is the user's transmitted signal in each frame, and w is Gaussian white noise. Splice to And used as features of the training set, validation set and test set, the scatterer distribution point cloud is used as the label.
[0040] Step 2 specifically includes:
[0041] like Figure 3As shown in the figure, an adaptive reflective surface optimization network is constructed, which is responsible for adaptively adjusting the reflection coefficient to improve the effectiveness of environmental perception. The adaptive reflective surface optimization network of each frame is composed of a recurrent neural network architecture and a multi-layer perceptron (MLP). The architecture of the recurrent neural network is a bidirectional long short-term memory network (BiLSTM) and a double-layer stacked LSTM. In the stacked LSTM network structure, the first layer of LSTM receives the input sequence and the initial state (c0, h0), and outputs the hidden state h1 and the memory unit c1 as the input of the second layer of LSTM, which are passed in sequence. The network dynamically adjusts the number of nodes of the recurrent neural network according to the number of time frames, that is, in the fth frame, the number of nodes of the recurrent neural network is f. The first frame input of the network is the initial y1, s1, Spliced In the fth frame, all y1:y of the previous f-1 frames f-1 ,s1:s f-1 , Splice to As input. Before the end of each frame, the hidden layer output h of the last layer will pass through the DNN of the R layer and output the optimized reflection surface coefficient That is, at the end of each frame, the IRS is updated to the optimized phase coefficient. This process continuously improves the quality of the received signal.
[0042] Step 3 specifically includes:
[0043] like Figure 4 As shown in the figure, a deep perspective perception network is constructed, which is responsible for recovering the corresponding local point cloud information from the signal received from a specific perspective; the network consists of a residual long short-term memory network (RES-LSTM) unit, a deep time module composed of a recurrent neural network, and a multi-layer perceptron (MLP). The RES-LSTM unit is based on the LSTM unit. The output of the hidden layer of the LSTM unit is adjusted to the same dimension as the input through a linear transformation (Linear layer), and then added to the input as the final unit output. The introduction of the LSTM with a residual structure can effectively alleviate the gradient disappearance problem caused by the increase in network depth. At the end of the F frame, the IRS phase is optimized to an approximately ideal state through the adaptive reflection surface optimization network, and the received echo signal is significantly improved. The optimized F frame data The input is a deep perspective perception network, which passes through multi-layer RES-LSTM, bidirectional LSTM, and multi-layer LSTM in sequence. The time series feature h output by RES-LSTM f It is fed into the bidirectional LSTM module, which processes the data in both the forward and backward directions and then concatenates the output into Send it to the multi-layer LSTM, and the cell state layer of the last layer of LSTM outputs C outAfter the R layer DNN, the DNN network outputs the regional point cloud state of a single perspective in Represents the splicing information of frame 1 to frame F and, The mapping rule is used to directly recover the regional scatterer point cloud state from the received signal, reflection surface phase and other data.
[0044] Step 4 specifically includes:
[0045] like Figure 5 As shown in the figure, a multi-view fusion perception network is constructed, which is responsible for aligning and weighting the features of local point cloud information from multiple perspectives, and finally restoring the complete three-dimensional environment point cloud; the network uses a 3D convolutional neural network (3D-CNN) to capture the three-dimensional information in space. There are 5 layers of 3D structure, each of which includes a 3D convolution layer, a 3D batch normalization layer, an activation function and a channel attention module (EfficientchannelattentionBlock, ECA-Block). The ECA-Block includes global average pooling, 1D convolution, Sigmoid activation, and finally multiplies the output channel weight with the input feature by channel. After the previous deep perspective perception network, the perspective point cloud information from multiple different perspectives is reconstructed. The point cloud information from different perspectives is fed into the network for point cloud fusion. The point cloud states from different perspectives in the network are stitched together and passed through 5 3D structures. The output of the last 3D structure passes through the fully connected layer, and the result is the final fused ROI high-precision point cloud information. The mean square error is used as the loss function in training, and the formula is
[0046] like Figure 6 As shown in the figure, it can be seen that the point cloud data obtained from a single perspective has problems such as incomplete information and missing details. After multi-perspective information fusion, the spatial structure of the point cloud is more complete and the details are richer.
[0047] like Figure 7 As shown in the figure, Comparison Scheme 1 uses the GAMP-MVSVR algorithm based on a multi-layer factor graph, while Comparison Scheme 2 uses a deep learning approach with randomized reflector coefficients. The figure compares the perceptual errors of each scheme under different signal-to-noise ratio conditions. As the signal-to-noise ratio gradually increases, the perceptual errors of all schemes show a downward trend, but the reduction in the comparison schemes is more limited. Without optimization, the reflector coefficients struggle to fully match the signal propagation characteristics, resulting in relatively large perceptual errors. After optimizing the reflector, performance is significantly improved, demonstrating the rationality of the reflector optimization network design and its important role in perception tasks.
[0048] The above embodiments are only preferred embodiments of the present invention and are not limitations on the technical solutions of the present invention. Any technical solution that can be implemented on the basis of the above embodiments without creative work should be deemed to fall within the scope of protection of the patent of the present invention.
Claims
1. A method for multi-view fusion environment perception based on deep learning and assisted by intelligent reflective surfaces, characterized in that: The steps include: Step 1: To achieve environmental perception, multiple access points (APs) transmit pilot signals s. The signal transmission process is divided into F frames. In each frame, echo signals reflected by scatterers and smart reflective surfaces in the environment are collected to obtain the received signal y for all AP frames. The transmitted signal s is paired with the corresponding channel matrix H to construct the data set required for neural network training. Step 2: Deploy multiple intelligent reflective surfaces (IRSs) in the sensing scene to obtain point cloud information of scatterers from different perspectives; deploy a central unit to process the pilot signal data and echo signal data transmitted from the AP, and use an adaptive reflective surface dynamic optimization network based on a long short-term memory network (LSTM) and a deep neural network (DNN) to calculate and adjust the phase of each IRS frame by frame, and convert the transmitted signal s, received signal y, and reflective surface phase of each frame into the corresponding values. Spliced into At the end of the F frame, the IRS phase is optimized to a near-ideal state through the adaptive reflection surface dynamic optimization network to achieve optimal signal reflection; Step 3: Optimize the F-frame data Input is a deep perspective perception network, which is composed of a residual network ResNet, LSTM and DNN. This network processes the echo signal sequence of all time frames to directly obtain regional point cloud information under different perspectives; Step 4: The regional point cloud information from different perspectives is fed into the multi-perspective fusion perception network. This network introduces a 3D convolutional neural network (3D CNN) and a self-attention mechanism to perform feature alignment and weighted fusion on the local point cloud information from multiple perspectives, ultimately restoring a high-precision, complete 3D environment point cloud.
2. The method of multi-view fusion environment perception based on deep learning assisted by intelligent reflective surfaces according to claim 1 is characterized in that The step 1 specifically includes: 1.1) There are K full-duplex access points (APs) (multi-antenna APs) in a space, each equipped with M transmit antennas and M receive antennas. Each AP sends pilot signals to other APs or reflective surfaces and receives echo signals from APs, reflective surfaces, and other scattering objects in the environment. The field of view is divided according to a single AP, or adjacent APs with overlapping sensing areas are divided into a single field of view. To enhance signal quality and expand the field of view of the communication link, I intelligent reflective surfaces (IRSs) are deployed in the indoor environment. Each IRS consists of J microelements, capable of reflecting and adjusting signals in different directions. In addition, a central unit is deployed in the scene to coordinate the work of all APs and IRSs. The central unit is connected to all APs and IRSs and is responsible for processing the pilot signal data and echo signal data transmitted by the APs and optimizing the data. 1.2) The perception space where the scatterer is located is called ROI: ROI will be evenly divided into [H x ,W y ,L z ] are the height, width and length of the target space respectively, [h x ,w y ,l z ] are the height, width and length of the small cube respectively; the entire ROI space is represented by x=[x1,x2,...,x n ] T ; where x n ∈[0,1],x n =x1,...,x N , represents the scattering coefficient of the nth scatterer; the final ROI space x n It is represented as the point cloud distribution in the specific area where the scatterer is located, and the specific distribution position and state of the scatterer are determined based on this; 1.3) The signal transmission process is divided into F frames. The multipath channel H consists of three parts: H = H LOS +H IRS +H S , H LOS and H IRS It is a path that is independent of scatterers, namely the AP-AP path and AP-IRS-AP path with line of sight LOS, H S It is a scattering multipath channel related to the scatterer; because in the intelligent reflector-assisted environmental communication perception model, H LOS and H IRS So the signal sending process only considers H S Consider the three paths of AP-scatterer-AP, AP-scatterer-IRS-AP and AP-IRS-scatterer-AP. Other paths with more scattering are ignored due to high attenuation, i.e. H S =H AP→S→AP +H AP→IRS→S→AP +H AP→S→IRS→AP ; Among them, H AP→S→AP =h AP→S ⊙v AP→S diag(x)h S→AP ⊙v S→AP , h AP→S and h S→AP is the path loss matrix from AP to scatterer and from scatterer to AP, v AP→S and v S→AP is the occlusion matrix from AP to scatterer and from scatterer to AP, x is the scattering coefficient vector, and ⊙ is the Hadamard product; h AP→R,i represents the channel matrix from all antennas to the i-th IRS, h R,i→S represents the channel matrix from the i-th IRS to the scatterer, v R,i→S is the occlusion matrix from the i-th IRS to the scatterer, Θ i represents the reflection phase matrix of the i-th IRS; h S→R,i represents the channel matrix from the scatterer to the i-th IRS, v S→R,i is the corresponding occlusion matrix, h R,i→AP represents the channel matrix from the i-th IRS to the AP; The final received signal processed by the central node is y = H S s+w; s is the user transmitted signal in each frame, and w is Gaussian white noise.
3. The method for multi-view fusion environment perception based on deep learning and assisted by intelligent reflective surfaces according to claim 1 is characterized in that: The step 2 specifically includes: Construct an adaptive reflector optimization network. Each frame of the adaptive reflector optimization network is composed of a recurrent neural network architecture and a multi-layer perceptron MLP. The architecture of the recurrent neural network is a bidirectional long short-term memory network BiLSTM and a double-layer stacked LSTM. The network dynamically adjusts the number of nodes of the recurrent neural network according to the number of time frames. That is, in the fth frame, the number of nodes of the recurrent neural network is f; the first frame input of the network is the initial y1, s1, Spliced In the fth frame, all the previous f-1 frames Splice to As input; before the end of each frame, the hidden layer output h of the last layer will pass through the DNN of the R layer and output the optimized reflection surface coefficient That is, at the end of each frame, the IRS is updated to the optimized phase coefficient. This process continuously improves the quality of the received signal.
4. The method for multi-view fusion environment perception based on deep learning and assisted by intelligent reflective surfaces according to claim 1 is characterized in that: The step 3 specifically includes: A deep perspective perception network is constructed, which consists of a residual long short-term memory network unit RES-LSTM, a deep time module composed of a recurrent neural network, and a multi-layer perceptron MLP. The RES-LSTM unit is based on the LSTM unit, and the output of the LSTM unit hidden layer is adjusted to the same dimension as the input through a linear transformation, and then added to the input as the final unit output. At the end of the F frame, the IRS phase adjusts the reflection coefficient through the reflection surface optimization network; the optimized F frame data is used The input is a deep perspective perception network, which passes through multi-layer RES-LSTM, bidirectional LSTM, and multi-layer LSTM in sequence. The time series features output by RES-LSTM are sent to the bidirectional LSTM module. The bidirectional LSTM processes data from both forward and backward directions, and then the output is spliced and sent to the multi-layer LSTM. The cell state C of the last layer of LSTM is input into the DNN of the R layer, and the network output is the regional point cloud state under a single perspective. in Represents the splicing information of frame 1 to frame F and, The mapping rule is used to directly recover the regional scatterer point cloud state from the received signal, reflection surface phase and other data.
5. The method for multi-view fusion environment perception based on deep learning and assisted by intelligent reflective surfaces according to claim 1 is characterized in that: The step 4 specifically includes: Construct a multi-view fusion perception network, which uses a 3D convolutional neural network (3D-CNN) to capture three-dimensional information in space. It has a total of M layers of 3D structures, each of which includes a 3D convolution layer, a 3D batch normalization layer, an activation function, and an efficient self-attention module. After the previous deep view perception network, the system reconstructs point cloud information from different perspectives. The point cloud information from different perspectives is fed into the network for point cloud fusion. After splicing, the point cloud states from different perspectives pass through M 3D structures. The output of the last 3D structure passes through the fully connected layer, and the result is the final ROI high-precision point cloud information. The mean square error is used as the loss function during training, and the formula is: