Multi-mode semantic-assisted millimeter wave beam prediction method in haze environment
Through the fusion method of lightweight image defogging network and location data, the accuracy and resource consumption problems of millimeter wave beam prediction in haze environments are solved, and efficient and low-cost beam prediction is achieved to meet the communication needs of industrial scenarios.
Patent Information
- Application Number
- CN202510595894.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-01
AI Technical Summary
In a haze environment, the decline in camera imaging quality leads to a decrease in millimeter wave beam prediction accuracy, and the existing deep convolutional neural network models consume high computing resources and are expensive to use.
The lightweight image defogging network is used to reconstruct clear images, combine location data to extract environmental semantics, and beam prediction is performed through lightweight image feature extraction and beam index inference network, fuse absolute and relative position information, and build a lightweight network using depth-separable convolution and attention mechanism.
It effectively improves beam prediction accuracy, reduces system storage and computing resource consumption, reduces hardware costs, and meets the needs of high-quality millimeter wave communication in complex environments.
Smart Images

Figure CN120238165A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless communication, and relates to a millimeter wave beam prediction method assisted by multi-modal semantics in a haze environment. Background Art
[0002] Millimeter wave communication, with its rich spectrum resources, extremely low communication latency, and ultra-high transmission rate, has become a key enabling technology for achieving ultra-reliable low-latency communication. To achieve efficient and reliable data transmission, millimeter wave communication systems usually need to predict the beamforming vector used to control the antenna array state based on additional sensing data, so as to determine the beam direction, realize beam prediction, and complete the directional transmission or reception of millimeter wave signals.
[0003] In typical industrial application scenarios such as mines and processing plants, the aggregation of suspended particulate matter caused by high humidity and dusty operating environments will trigger a persistent haze effect, and the imaging quality of cameras will seriously decline in such environments. In millimeter wave beam prediction methods based on visual sensing data, this will directly lead to a significant reduction in prediction accuracy and ultimately cause beam misalignment, resulting in serious millimeter wave communication delays. In addition, in order to effectively process visual sensing data, it is usually necessary to construct a network model with a complex structure such as a deep convolutional neural network, and the inference process of such a network usually consumes a large amount of system storage and computing resources, thus resulting in high hardware costs and model overheads.
[0004] Therefore, there is an urgent need for a lightweight and efficient prediction method for millimeter wave beams in a haze environment to support high-quality wireless communication requirements. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a millimeter wave beam prediction method assisted by multi-modal semantics in a haze environment. First, an end-to-end reconstruction of a clear image is realized by combining a lightweight image dehazing network. Further, relative and absolute position environmental semantics are extracted based on the dehazed image and positioning information, and multi-modal semantics are processed through a lightweight image feature extraction network and a beam index inference network to achieve efficient millimeter wave beam prediction in a haze environment.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A millimeter-wave beam prediction method assisted by multi-modal semantics in a haze environment, which is applicable to the millimeter-wave wireless communication process between a base station and a mobile user in industrial scenarios such as mines and processing plants where haze environments are prone to occur. It is considered to achieve accurate prediction of millimeter-wave beams by using visual and position data obtained from additional sensors. First, aiming at the influence of the haze environment on the imaging effect of the camera, an end-to-end reconstruction of clear images is realized through a lightweight image dehazing network. Then, the relative and absolute position environmental semantics in the image and position data are extracted. Finally, the multi-modal semantics are input into the image feature extraction network and the beam index inference network to output the beam index and complete the beam prediction. The present invention considers the millimeter-wave communication process and meets the millimeter-wave communication requirements in complex scenarios. The method specifically includes the following steps:
[0008] S1: Construct a millimeter-wave signal transmission model and a beam selection model between the user and the base station in a haze environment;
[0009] S2: Obtain the position data of the user through a locator installed at the user end and perform normalization processing on it to obtain the absolute position environmental semantics;
[0010] S3: Take a photo of the user through a camera deployed at the base station end. Considering the influence of the haze environment on the imaging effect of the camera, define the task content of end-to-end image reconstruction based on the atmospheric scattering physical model, determine the optimization objective of deep learning, construct a lightweight image dehazing network, output clear images, and perform object detection on them to obtain the relative position environmental semantics;
[0011] S4: Define the task content of millimeter-wave beam prediction based on the obtained multi-modal environmental semantics, determine the optimization objective of deep learning, construct lightweight image feature extraction networks and beam index inference networks for different environmental semantics, and complete the millimeter-wave beam prediction.
[0012] Furthermore, step S1 specifically includes the following steps:
[0013] S11: Construct a millimeter-wave signal transmission model between the user and the base station in a haze environment, which specifically includes: considering the wireless communication scenario between the base station and the mobile user in a haze environment such as a mine or a processing plant, where the base station equipped with an M-element uniform linear antenna array and a camera provides wireless communication services in the millimeter-wave band for the mobile user equipped with an omnidirectional antenna and a locator. During the communication process, the base station uses orthogonal frequency division multiplexing to transmit data x to the mobile user. Let the transmission channel vector of the k-th subcarrier at time t be Then the millimeter-wave signal y k [t] received by the mobile user at time t is represented by the following formula:
[0014]
[0015] Among them, represents the beamforming vector; v k [t] represents the additive white Gaussian noise present in the k-th subcarrier channel at time t.
[0016] S12: Construct a millimeter-wave signal beam selection model between the user and the base station in a haze environment, specifically including: To ensure the stable transmission of the downlink signal, the base station needs to select, from a predefined codebook i containing Q beamforming vectors f as many beamforming vectors as possible that can provide a relatively high signal achievable transmission rate to indicate the beamforming process of the antenna array and achieve beam alignment; among them, the beamforming vector that can maximize the signal achievable transmission rate at time t is the optimal beam vector, which is uniquely indicated by the index i of the vector in the codebook, that is, the optimal beam index i * , which is represented by the following formula:
[0017]
[0018] where Rate fi[t] represents the signal achievable rate based on the beamforming vector f i [t].
[0019] Furthermore, step S2 specifically includes the following steps:
[0020] S21: Obtain the longitude and latitude coordinates g pos = [la, lo] of the user through a locator, where la and lo represent the longitude and latitude of the user's current location respectively;
[0021] S22: For the obtained coordinate position data g pos of the user, map it uniformly to the interval [0, 1] to eliminate the dimensional difference between the data, and finally obtain the position vector of the user at time t as its absolute position environmental semantics.
[0022] Furthermore, step S3 specifically includes the following steps:
[0023] S31: Take a photo of the haze environment containing the user through the camera at the base station side, denoted as the vector where C, H, and W represent the number of channels, height, and width of the photo respectively;
[0024] S32: Use the following formula to represent the atmospheric scattering physical model of light propagation in the atmosphere and its mutual scattering with particles such as dust and water vapor:
[0025] I(z) = J(z) × t(z) + A × (1 - t(z))
[0026] Among them, I represents the haze image captured by the camera at the base station end, J represents the clear image, z represents a certain pixel position in the image, A represents the atmospheric light constant value, and t represents the transmittance map; therefore, the clear image can be reconstructed by obtaining the values of the parameters t(z) and A and through the formula expressed as follows:
[0027] J(z) = X(z) × I(z) - X(z) + b ias
[0028] Among them, the parameters t(z) and A are encapsulated in a parameter X defined by the formula expressed as follows:
[0029]
[0030] Among them, b ias is the default bias term with a value of 1;
[0031] Furthermore, determining the task of end-to-end image dehazing is to estimate the parameter X and remove the haze in the image I based on the mapping function f X to reconstruct the dehazed image This dehazing process is expressed by the following formula:
[0032]
[0033] Among them, the mapping function f X = X(z) × I(z) - X(z) + b ias ;
[0034] Adopting the deep learning method, by constructing a deep neural network and using the co-existing dataset of the haze image I p and the clear image J p for training, so as to fit the mapping function f , where P represents the total number of available samples in the dataset. The optimization objective of deep learning is to minimize the mean square error loss between the clear image of the given sample and the corresponding dehazed image. This optimization process is expressed by the following formula: X Among them,
[0035]
[0036] where represents the optimal mapping function, represents the dehazed image corresponding to the haze image I p .
[0037] S33: For the obtained haze environment photos Based on depthwise separable convolution, through residual and skip connections, and introducing an attention layer composed of a channel attention module and a spatial attention module, a lightweight end-to-end image dehazing network Lite-DHNet is constructed to estimate the value of parameter X and output a dehazed image, represented as a vector
[0038] S34: For the dehazed image Output the bounding box coordinates b′[t] = [x r , y r , x l , y l of the user at time t through the SSDLite object detection network, where x r , y r , x l and y l represent the horizontal and vertical coordinates of the upper left point and the horizontal and vertical coordinates of the lower right point of the bounding box respectively; b′[t] is further calculated to represent the bounding box vector b[t] = [c x , c y , w, h] of the user as its relative position environmental semantics, where c x and c y represent the horizontal and vertical coordinates of the center point of the bounding box respectively, and w and h represent the width and height of the bounding box respectively; the mask image is generated based on the bounding box coordinates b′[t], where the area within the user's bounding box is represented by one color (such as white), while other environmental elements are of another color (such as black), and m′[t] is further shrunk to obtain the user's mask image as its relative position environmental semantics.
[0039] Further, step S4 specifically includes the following steps:
[0040] S41: Define the task of beam prediction as estimating the optimal beam index through the mapping function f Θ based on multi-modal environmental semantics The beam prediction process is expressed as:
[0041]
[0042] where Θ represents the relevant parameters constituting the mapping function;
[0043] S42: Adopt the method of deep learning, and fit the mapping function f by constructing a deep neural network and using a co-existing dataset of sensing data and communication data for network training Θ, where \(U\) represents the number of samples in the dataset. The optimization objective of deep learning is to maximize the probability of accurate prediction of given samples, and this optimization process is expressed as:
[0044]
[0045] Among them, represents the mapping function \(f\) Θ The possibility of accurate prediction;
[0046] S43: Input the user mask image into the mask image feature extraction network composed of two attention convolutional blocks and four fully connected neural network layers, and output the mask image features Among them, the attention convolutional block is composed of a convolutional layer, a max pooling layer and an attention layer. The attention layer applies a channel attention module and a spatial attention module in sequence, so that the network can effectively focus on the important regions in the mask image to achieve efficient and accurate extraction of its features;
[0047] S44: Combine the bounding box vector and the position vector with the output of the mask image feature extraction network, and input it into the beam index inference network composed of a multi-layer perceptron network, and output the predicted beam index Complete the millimeter wave beam prediction process.
[0048] The beneficial effects of the present invention are as follows: The present invention can effectively avoid the influence of the haze environment on the beam prediction accuracy, greatly reduce the system storage and computing resources consumed in the beam prediction process, significantly reduce the hardware cost of the wireless system and the model overhead of beam prediction, and meet the high-quality millimeter wave communication requirements in complex environments. The specific manifestations are as follows:
[0049] (1) Enhanced environmental robustness: Through the lightweight end-to-end image dehazing network (Lite-DHNet), the reconstruction of haze images is realized, and the dehazing process is optimized in combination with the atmospheric scattering physical model, effectively eliminating the interference of haze on camera imaging, providing a high-definition image basis for subsequent semantic extraction, and significantly improving the perception reliability in complex environments.
[0050] (2) Advantages of multi-modal semantic fusion: Innovatively fuse the absolute position semantics (normalized longitude and latitude coordinates) and relative position semantics (bounding box vectors and mask images in dehazed images), strengthen the extraction of key features through the attention mechanism, realize the three-dimensional representation of environmental information, and improve the spatial perception accuracy of beam prediction.
[0051] (3) Lightweight resource optimization: A lightweight network architecture (such as the dehazing network Lite-DHNet and the mask feature extraction network) is constructed using depthwise separable convolution, residual skip connections, and channel / space dual attention modules, reducing the number of parameters and computational resource consumption compared to traditional models, and meeting the deployment requirements of low-power devices in industrial scenarios.
[0052] (4) End-to-end system efficiency improvement: Through the cascaded design of SSDLite lightweight object detection and multi-layer perceptron inference network, automated prediction from raw sensing data to beam index is achieved, shortening the decision-making delay compared to traditional beam scanning methods, while maintaining a high prediction accuracy, and significantly improving the real-time performance of millimeter-wave communication.
[0053] In summary, the present invention deeply integrates computer vision and wireless communication technologies, overcomes the industry pain point of inaccurate channel perception in haze environments, provides a highly reliable and low-cost millimeter-wave communication solution for harsh environments such as mines and factories, and has significant industrial application value. Through lightweight design, it breaks through hardware limitations and provides a new paradigm for 5G / 6G environment adaptive beamforming technology.
[0054] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0056] Figure 1 It is a schematic diagram of a wireless communication scenario, a semantic extraction process, and a beam prediction process;
[0057] Figure 2 It is a schematic diagram of the lightweight end-to-end image dehazing network Lite-DHNet;
[0058] Figure 3 It is a flowchart of the millimeter-wave beam prediction method with multi-modal semantic assistance in the haze environment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] The following specific examples illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only schematically illustrate the basic concept of the present invention. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0060] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams rather than physical diagrams, and should not be construed as limitations on the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0061] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as limitations on the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0062] Please refer to Figures 1 to 3 , the present invention provides a millimeter-wave beam prediction method assisted by multi-modal semantics in a haze environment, aiming to meet the wireless communication requirements of millimeter waves in industrial scenarios such as mines and processing plants where haze environments often occur. Considering the additional sensing information collected by cameras and locators for beam prediction, for the user's haze photos taken by the camera, an end-to-end reconstruction of a clear image is performed through a lightweight deep neural network, and then the user bounding box vector and mask image in the dehazed image are extracted. Further, the mask image features are extracted through a lightweight deep neural network, and the beam index is output through a lightweight deep neural network by combining the bounding box vector and the position vector to complete the millimeter-wave beam prediction.
[0063] Figure 1 Schematic diagrams of the wireless communication scenario, the semantic extraction process, and the beam prediction process. As Figure 1As shown in the figure, consider the wireless communication scenario between a base station and mobile users in a haze environment such as a mine or a processing plant. The base station equipped with a uniform linear antenna array and a camera provides wireless communication services in the millimeter wave band for mobile users equipped with omnidirectional antennas and locators; during the user semantic extraction process, for the position coordinates g of the user at time t pos , through normalizing the data, it is uniformly mapped into the interval [0, 1], and then the position vector g[t] of the user at time t is output. For the haze environment photo r of the user at time t img , first, a clear image is reconstructed through the lightweight end-to-end image dehazing network Lite-DHNet, and a dehazed image is output Then, the SSDLite object detection network is used to extract the bounding box coordinates b′[t] = [x of the user in the dehazed image r , y r , x l , y l , and a corresponding user mask image is generated Furthermore, b′[t] and m′[t] are respectively obtained through calculation operations and reduction processing to get the bounding box vector b[t] = [c x , c y , w, h] and the mask image During the beam prediction process, for the multi-modal environmental semantics First, the mask image m[t] is input into the mask image feature extraction network, and the feature m F is output. This network consists of two attention convolutional blocks and four fully connected neural network layers. Among them, the attention convolutional block consists of a convolutional layer, a max pooling layer, and an attention layer. The attention layer sequentially applies a channel attention module and a spatial attention module. Then, the feature m F is merged with the bounding box vector b[t] and the position vector g[t], and further input into the beam index inference network, and the predicted millimeter wave beam index at time t is output This network consists of a multi-layer perceptron network, which includes a total of 1 input layer, 5 hidden layers, and 1 output layer.
[0064] Figure 2Schematic diagram of the lightweight end-to-end image dehazing network Lite-DHNet. First, Lite-DHNet introduces a splicing layer to retain low-dimensional features and avoid feature loss. That is, splicing layer 1 connects the depthwise separable convolutional layer 1 and the depthwise separable convolutional layer 3 in the channel direction and inputs them into the depthwise separable convolutional layer 4. Splicing layer 2 connects the depthwise separable convolutional layer 4 and the depthwise separable convolutional layer 6 in the channel direction and inputs them into the depthwise separable convolutional layer 7. Splicing layer 3 connects the depthwise separable convolutional layer 2, the depthwise separable convolutional layer 5, and the depthwise separable convolutional layer 7 in the channel direction and inputs them into the depthwise separable convolutional layer 8. Among them, the depthwise separable convolutional layer 2 and the depthwise separable convolutional layer 5 are skipped in the splicing operations of connection layer 1 and connection layer 2. Such a design aims to reduce the dimension of the connection layer and streamline the network architecture while better retaining the differentiated features between different levels. Second, Lite-DHNet enhances the feature transfer and fusion between depthwise separable convolutional layers through residual connections to avoid gradient disappearance or explosion during network training. That is, the output of the depthwise separable convolutional layer 8 is connected to the outputs of splicing layer 1, splicing layer 2, and splicing layer 3 using residual operations. Finally, Lite-DHNet introduces an attention mechanism to retain important picture feature details and structural information and reduce the interference of irrelevant feature regions, further highlighting the key features that play an obvious role in dehazing image reconstruction. That is, an attention layer composed of a channel attention module and a spatial attention module is added between the depthwise separable convolutional layer 8 and the depthwise separable convolutional layer 9.
[0065] Figure 3 Flowchart of the millimeter wave beam prediction method with multi-modal semantic assistance in the haze environment of the present invention, specifically including the following steps:
[0066] V1: The beam prediction process starts.
[0067] V2: Establish a millimeter wave signal transmission model, a millimeter wave beam selection model, and an end-to-end image dehazing model.
[0068] V3-V4: Take a haze image r containing the user through a camera img and obtain the longitude and latitude coordinates g of the user through a locator pos and use them as the original semantic data for describing the user environment.
[0069] V5-V11: Determine whether it is user image data r img If so, process the haze image r through the lightweight end-to-end image dehazing network Lite-DHNet img and output the dehazed image Further use SSDLite for the dehazed image Perform object detection and output the bounding box vector b[t], generate the mask image m[t] as the relative position environmental semantics; if not, it is the coordinate position data g pos , then perform a normalization operation on it and output the position vector g[t] as the absolute position environmental semantics; merge to obtain the multi-modal environmental semantics
[0070] V12 - V16: Determine whether it is the mask image m[t]. If so, extract the image feature m through the mask image feature extraction network F ; if not, it is the user position vector g[t] and the bounding box vector b[t], then merge it with the mask image feature m F and output the predicted beam index through the beam index inference network
[0071] V17: The beam prediction process ends.
[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A multimodal semantically assisted millimeter wave beam prediction method in a haze environment, characterized in that: The method specifically comprises the following steps: S1: Construct the millimeter wave signal transmission model and beam selection model between users and base stations in a haze environment; S2: Obtain the user's location data through the locator installed on the user side, and normalize it to obtain the absolute location environment semantics; S3: Use the camera deployed at the base station to take photos of users, define the task content of end-to-end image reconstruction based on the atmospheric scattering physical model, determine the optimization target of deep learning, build a lightweight image defogging network, output clear images, and detect targets, so as to obtain relative position environment semantics; S4: Define the task content of millimeter wave beam prediction based on the obtained multimodal environmental semantics, determine the optimization goal of deep learning, build a lightweight image feature extraction network and beam index inference network for different environmental semantics, and complete millimeter wave beam prediction.
2. The millimeter wave beam prediction method according to claim 1, characterized in that: In step S1, a millimeter wave signal transmission model between users and base stations in a haze environment is constructed, specifically including: considering the wireless communication scenario between base stations and mobile users in a haze environment, where a base station equipped with an M-element uniform linear antenna array and a camera provides wireless communication services in the millimeter wave frequency band for mobile users equipped with omnidirectional antennas and locators. During the communication process, the base station uses orthogonal frequency division multiplexing to transmit data x to the mobile user. Let the transmission channel vector of the k∈{1,...,K}th subcarrier at time t be Then the millimeter wave signal y received by the mobile user at time t is k [t] is expressed by the following formula: in, represents the beamforming vector; v k [t] represents the Gaussian white noise existing in the k-th subcarrier channel at time t.
3. The millimeter wave beam prediction method according to claim 2, characterized in that: In step S1, a millimeter wave signal beam selection model between the user and the base station in a haze environment is constructed, specifically including: the base station selects a beam containing Q beamforming vectors f i Predefined codebook A beamforming vector that can provide a high signal transmission rate is selected to indicate the beamforming process of the antenna array and realize beam alignment; wherein the beamforming vector that can maximize the signal transmission rate at time t is the optimal beam vector, which is uniquely indicated by the index i of the vector in the codebook, that is, the optimal beam index i * , which is expressed by the following formula: in, Represents the beamforming vector f i [t] is the achievable rate of the signal.
4. The millimeter wave beam prediction method according to claim 3, characterized in that: Step S2 specifically includes the following steps: S21: Obtain the user's latitude and longitude coordinates g through the locator pos = [la,lo], where la and lo represent the longitude and latitude of the user's current location respectively; S22: Obtained user coordinate position data g pos , uniformly map it to the interval [0, 1] to eliminate the dimension difference between the data, and finally obtain the user's position vector at time t as its absolute position environment semantics.
5. The millimeter wave beam prediction method according to claim 4, characterized in that: Step S3 specifically includes the following steps: S31: The camera at the base station takes a photo of the haze environment containing the user, represented as a vector Where C, H, and W represent the number of channels, height, and width of the photo, respectively; S32: The atmospheric scattering physical model that represents the propagation of light in the atmosphere and its interaction with particulate matter is expressed using the following formula: I(z)=J(z)×t(z)+A×(1-t(z)) Where I represents the haze image captured by the camera at the base station, J represents the clear image, z represents a pixel position in the image, A represents the atmospheric light constant value, and t represents the transmittance map; therefore, the clear image is reconstructed by obtaining the values of the parameters t(z) and A and expressing the following formula: J(z)=X(z)×I(z)-X(z)+b ias Here, the parameters t(z) and A are encapsulated in a parameter X defined by the formula expressed as follows: Among them, b ias is the default bias term with a value of 1; Furthermore, the task of end-to-end image dehazing is to estimate the parameter X based on the mapping function f X , remove the haze in image I and reconstruct the dehazed image The dehazing process is expressed by the following formula: Among them, the mapping function f X =X(z)×I(z)-X(z)+b ias ; Using deep learning, we build a deep neural network and use haze images I p and clear image J p Coexisting datasets Train to fit the mapping function f X , where P represents the total number of samples available in the data set. The optimization goal of deep learning is to minimize the mean square error loss between the clear image of a given sample and the corresponding dehazed image. The optimization process is expressed by the following formula: in, represents the optimal mapping function, Represents the haze image I p The corresponding dehazed image; S33: Obtained photos of haze environment Based on the depthwise separable convolution, through residual and skip layer connections, and introducing the attention layer composed of channel attention module and spatial attention module, a lightweight end-to-end image dehazing network Lite-DHNet is constructed to estimate the value of parameter X and output the dehazed image, which is represented as a vector S34: for dehazing images Output the bounding box coordinates of the user at time t through the SSDLite object detection network b′[t]=[x r ,y r ,x l ,y l ], where x r ,y r 、x l and l Respectively represent the horizontal and vertical coordinates of the upper left point and the lower right point of the bounding box; b′[t] is further calculated to be the user's bounding box vector b[t]=[c x ,c y ,w,h], as its relative position environment semantics, where c x and c y Respectively represent the horizontal and vertical coordinates of the center point of the bounding box, w and h represent the width and height of the bounding box respectively; mask image Based on the bounding box coordinates b′[t], the area within the user's bounding box is represented by one color, while other environmental elements are represented by another color. m′[t] is further reduced to obtain the user's mask image As its relative position environment semantics.
6. The millimeter wave beam prediction method according to claim 5, characterized in that: Step S4 specifically includes the following steps: S41: Define the task of beam prediction as a semantically based multimodal environment By mapping the function f Θ Estimating the best beam index The beam prediction process is expressed as: Among them, Θ represents the relevant parameters constituting the mapping function; S42: Using deep learning methods, by building a deep neural network and using sensor data and communication data Coexisting datasets Perform network training to fit the mapping function f Θ , where U represents the number of samples in the data set. The optimization goal of deep learning is to maximize the probability of accurate prediction of a given sample. The optimization process is expressed as: in, Represents the mapping function f Θ the likelihood of accurate predictions; S43: User mask image Input into the mask image feature extraction network and output the mask image features S44: Bounding box vector and the position vector Output of the feature extraction network with the mask image Merge and input into the beam index inference network composed of a multi-layer perceptron network, and output the predicted beam index Complete the mmWave beam prediction process.
7. The millimeter wave beam prediction method according to claim 6, characterized in that: In step S43, the mask image feature extraction network includes two attention convolution blocks and four fully connected neural network layers; the attention convolution block includes a convolution layer, a maximum pooling layer and an attention layer, and the attention layer applies a channel attention module and a spatial attention module in sequence.