Channel prediction method based on image processing and machine learning
By employing a channel prediction method based on image processing and machine learning, and utilizing semantic segmentation and scene recognition models, the problems of poor environmental adaptability and high measurement costs in traditional channel modeling are solved. This method enables channel characteristic prediction and high-precision channel modeling for unknown scenarios, and is suitable for multi-band and multi-scenario coverage in 6G systems.
Patent Information
- Application Number
- CN202311096516.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-28
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2043-08-28
AI Technical Summary
Existing channel modeling methods cannot effectively predict channel characteristics in unknown scenarios, and traditional channel measurement is costly, time-consuming, and labor-intensive. It cannot fully explore the complex relationship between new channel characteristics and frequency bands/scenarios, and it has poor adaptability to real-time environmental changes.
We employ a channel prediction method based on image processing and machine learning. By combining semantic segmentation and scene recognition models with feature extraction and prediction networks, we use scene images to predict channels and construct a mapping relationship between the physical environment and channel characteristics.
It enables real-time channel prediction, improves environmental adaptability and prediction accuracy, reduces measurement costs, better meets the coverage requirements of multiple frequency bands and scenarios in 6G systems, and builds a foundation for environmentally aware channel modeling.
Smart Images

Figure CN117132981B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of channel prediction technology, and in particular to a channel prediction method based on image processing and machine learning. Background Technology
[0002] Future 6G wireless communication technology will achieve full coverage, full spectrum, full application, and strong security. Wireless channel characteristic analysis and modeling are the foundation for communication system design, performance evaluation, optimization, and deployment. However, high-performance channel detectors are expensive, the measurement process is time-consuming and labor-intensive, and it is impossible to exhaustively measure channels in all frequency bands and all scenarios. In addition, channel measurements cannot even be carried out in some extreme scenarios. Therefore, there is an urgent need to learn from existing channel characteristics to better predict channel characteristics in unknown scenarios and reduce measurement costs.
[0003] Existing channel modeling efforts primarily focus on deterministic and statistical modeling, such as channel measurement-based modeling, ray tracing-based deterministic channel modeling, and geometrically stochastic channel modeling. None of these existing models can predict channel characteristics for unknown scenarios, and they lack sufficient prior knowledge of the current environment. The channel models rely only on a subset of physical parameters, ignoring environmental information and failing to adapt to real-time environmental changes. Furthermore, traditional channel measurement and modeling still suffer from long-standing problems, such as the high cost and time-consuming, labor-intensive measurement process of high-performance channel detectors. In some extreme cases, channel measurement may even be impossible. Therefore, only partial channel characteristics of known frequency bands and scenarios can be analyzed, failing to fully explore the complex relationships between new channel characteristics and frequency bands / scenarios. To address these limitations, channel research has gradually evolved from simple statistical channel modeling to more flexible and accurate predictive channel modeling based on large amounts of data. In this regard, artificial intelligence (AI) is expected to serve as an auxiliary tool, given its rapid development and significant success in many fields such as image processing, computer vision, and data mining over the past decade. Utilizing AI for channel prediction is attracting increasing attention from researchers in the communications field. Spatial domain-based predictive channel modeling can delve into the relationship between channel characteristics and physical parameters.
[0004] However, current spatial domain channel prediction modeling largely focuses on site data while ignoring scene images, which inherently contain a wealth of information. As factors influencing channels are continuously explored, parameter acquisition becomes a major challenge. Therefore, more flexible and mobile image processing techniques are needed to extract more information from images and achieve more accurate temporal / spatial channel prediction.
[0005] Current research on image-based predictive channel modeling, while achieving better prediction results than traditional predictive channel modeling, is still limited by physical parameters, using images only as an auxiliary prediction tool. Furthermore, these models cannot explain how images affect channel prediction, thus their interpretability requires further investigation. In contrast, image processing can reduce this dependence to some extent and further extract factors influencing the channel from the image. Summary of the Invention
[0006] This invention provides a channel prediction method based on image processing and machine learning to solve the technical problems of high cost and difficulty in channel measurement and difficulty in obtaining traditional channel modeling parameters.
[0007] This invention provides a channel prediction method based on image processing and machine learning, comprising the following steps:
[0008] Acquire scene images of the channel to be tested;
[0009] The scene image is input into a pre-trained semantic segmentation model to obtain the segmented scene image;
[0010] The segmented scene image is input into a pre-trained scene recognition model to obtain the scene classification result of the segmented scene image;
[0011] Based on the scene classification results, the segmented scene images are input into a pre-trained feature extraction and prediction network for channel prediction, thereby obtaining the channel prediction results of the scene images for the channel to be tested.
[0012] In one embodiment of the present invention, it further includes:
[0013] Acquire channel measurement data and scene images for existing frequency bands and scenarios, and construct a training database for multiple frequency bands and multiple scenarios;
[0014] Based on the types of channel scatterers, the scene images in the training database are labeled to obtain a semantic segmentation dataset. The semantic segmentation dataset is then used to train a pre-constructed semantic segmentation neural network to obtain a trained semantic segmentation model.
[0015] In one embodiment of the present invention, the semantic segmentation neural network is trained using the semantic segmentation dataset to obtain a trained semantic segmentation model, including:
[0016] MobileNetV2 is used as the backbone network of the encoder, and the feature maps are expanded from shallow to deep.
[0017] A DeepLabV3+ semantic segmentation neural network based on the MobileNetV2 backbone network is constructed, in which dilated convolutional layers with different dilation rates are selected to extract features at different scales.
[0018] The semantic segmentation neural network is trained using the semantic segmentation dataset to obtain a trained semantic segmentation model.
[0019] In one embodiment of the present invention, in a trained semantic segmentation model, the segmentation accuracy of the semantic segmentation model is evaluated by pixel accuracy, wherein the pixel accuracy is:
[0020]
[0021] In the formula, TP represents the number of pixels correctly predicted as positive, and FP represents the number of pixels incorrectly predicted as positive.
[0022] In one embodiment of the present invention, it further includes:
[0023] The segmented images in the semantic segmentation database are classified according to the scene classification criteria to obtain the scene classification results of the segmented images;
[0024] Using segmented images and their corresponding scene classification results as training data, a pre-constructed scene recognition model is trained to obtain a trained scene recognition model, wherein the scene recognition model uses the GoogLeNet network as its main structure.
[0025] In one embodiment of the present invention, it further includes:
[0026] The channel measurement data in the training database are normalized.
[0027] Based on the scene classification results, segmented scene images and channel measurement data of the same scene are used as training data to train a pre-constructed feature extraction and prediction network. Multiple rounds of training are performed for different scenes to obtain a trained feature extraction and prediction network. The feature extraction and prediction network consists of five connected convolutional blocks containing different numbers of 3×3 convolutional kernels as the image feature extraction structure. After the last convolutional block, five fully connected layers are connected. The first four fully connected layers reduce the dimensionality of the extracted features, and the last fully connected layer outputs a channel prediction result with a dimension of 1.
[0028] In one embodiment of the present invention, normalization processing is performed on the channel measurement data in the training database, including:
[0029] The channel measurement data were normalized using the min-max normalization method.
[0030]
[0031] Where X represents the normalized data, and x represents the original channel measurement data. min x is the minimum value of the channel measurement data. max This represents the maximum value of the channel measurement data.
[0032] The channel prediction method based on image processing and machine learning in this invention breaks away from the limitations of physical parameters in traditional channel modeling, enabling real-time channel prediction based on scene images. Compared to traditional channel modeling, this method has higher environmental adaptability and can better address the coverage requirements of multiple frequency bands and scenarios in 6G systems. Furthermore, compared to other machine learning-based channel prediction models, it introduces semantic segmentation to further extract information from the image, weakening the influence of non-scattering parts of the image on channel prediction, thus achieving better prediction results. This method constructs a mapping relationship between the physical environment and channel characteristics, laying the foundation for further research on environment-aware channel modeling.
[0033] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0034] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0035] Figure 1 A flowchart illustrating a channel prediction method based on image processing and machine learning according to an embodiment of the present invention;
[0036] Figure 2 The training process for the channel prediction model provided in the embodiments of the present invention;
[0037] Figure 3 This is a structural diagram of a semantic segmentation model provided according to an embodiment of the present invention;
[0038] Figure 4 This is an image semantic segmentation result diagram provided according to an embodiment of the present invention;
[0039] Figure 5 This is a diagram of the feature extraction and prediction network structure provided according to an embodiment of the present invention;
[0040] Figure 6 This is an application process for predictive channel modeling based on a multimodal database, provided according to an embodiment of the present invention;
[0041] Figure 7 This is a schematic diagram of the image processing and machine learning algorithm provided according to an embodiment of the present invention;
[0042] Figure 8 A flowchart illustrating the prediction channel modeling and simulation process for an application provided according to an embodiment of the present invention. Detailed Implementation
[0043] Embodiments of the present invention are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0044] Figure 1 This is a flowchart illustrating a channel prediction method based on image processing and machine learning according to an embodiment of the present invention.
[0045] like Figure 1 As shown, this channel prediction method based on image processing and machine learning includes the following steps:
[0046] In step S101, a scene image of the channel to be tested is acquired.
[0047] In the embodiments of the present invention, there are multiple methods for obtaining images of the channel scene to be tested, and there is no limitation on the specific acquisition method.
[0048] In step S102, the scene image is input into the pre-trained semantic segmentation model to obtain the segmented scene image.
[0049] For the scene image of the channel to be tested, it is input into the semantic segmentation model for annotation, resulting in an annotated scene image.
[0050] In one embodiment of the present invention, the semantic segmentation model is pre-trained, such as... Figure 2 As shown, the specific training process is as follows: acquire channel measurement data and scene images of existing frequency bands and scenes, and construct a multi-frequency band and multi-scene training database; label the scene images in the training database according to the types of channel scatterers to obtain a semantic segmentation dataset; use the semantic segmentation dataset to train the pre-constructed semantic segmentation neural network to obtain a trained semantic segmentation model.
[0051] To meet the needs of various communication coverage scenarios and frequency bands, a training database of channel measurement data and scenario images / videos is constructed. During the construction of the training database, various measurement devices and tools can be used, while the scenario image / video database is used for analysis and optimization.
[0052] In one specific embodiment, the channel measurement environment was a suburban scene near a building. The receiver was placed at a fixed location within the building, while the transmitter continuously changed locations according to the scene. The measurement frequency was 3.7 GHz, and a total of 8392 sets of data were collected across six scenes. Each set of data, with the building as the origin, recorded its location coordinates and the received power of the transceiver, thereby calculating the path loss at each location. Simultaneously, scene images were captured at each location using a drone.
[0053] Based on actual needs, channel scatterers can be categorized. For example, indoor scenes can be categorized into walls, ceilings, human figures, and furniture, while outdoor scenes are more complex and diverse, and can be categorized into buildings, vegetation, and various terrain features. Based on the categorized channel scatterers, the scene images are labeled to generate a semantic segmentation dataset.
[0054] In the specific embodiments described above, the implementation scenarios mainly include trees and buildings. Therefore, the implementation scenarios can be divided into six categories based on the density of trees and the shading effect of buildings, as shown in Table 1.
[0055] Table 1 Image Scene Classification Criteria
[0056] Scene number Classification criteria 1 The trees are not densely packed and there are no buildings to provide shelter. 2 The trees are of moderate density and there are no buildings to provide shelter. 3 The trees are densely packed and there are no buildings to provide shelter. 4 The trees are not densely packed and are shaded by buildings. 5 The trees are of moderate density and are sheltered by buildings. 6 The trees are densely packed and sheltered by buildings.
[0057] Furthermore, a pre-constructed semantic segmentation neural network is trained using a semantic segmentation dataset to obtain a trained semantic segmentation model, including:
[0058] MobileNetV2 is used as the backbone network of the encoder, and the feature maps are expanded from shallow to deep.
[0059] A DeepLabV3+ semantic segmentation neural network based on the MobileNetV2 backbone network is constructed, in which dilated convolutional layers with different dilation rates are selected to extract features at different scales.
[0060] A semantic segmentation neural network is trained using a semantic segmentation dataset to obtain a trained semantic segmentation model.
[0061] like Figure 3 As shown, the structure of the semantic segmentation model is illustrated, such as... Figure 4 As shown, the results of semantic segmentation are presented.
[0062] Selecting dilated convolutional layers with different dilation rates to extract features at different scales is as follows:
[0063] f(x;w)=[f0(x;w0),f1(x;w1),...,f n (x;w n )]
[0064] In the formula, x represents the image features output from MobileNetV2, w is the set of parameters to be trained, f0 represents a 1×1 convolutional layer without dilation, and f1 to f n The dilation rates represent convolutional layers with different dilation rates: f1 represents a 3×3 convolutional layer with a dilation rate of 6, f2 represents a 3×3 convolutional layer with a dilation rate of 12, and f3 represents a 3×3 convolutional layer with a dilation rate of 18. These can be modified as needed. Modifying the dilation rate may affect the model's performance; the optimal combination of dilation rates can be determined experimentally.
[0065] In the trained semantic segmentation model, the segmentation accuracy is evaluated by pixel accuracy, where pixel accuracy is:
[0066]
[0067] In the formula, TP represents the number of pixels correctly predicted as positive, and FP represents the number of pixels incorrectly predicted as positive. The higher the pixel accuracy, the higher the model accuracy.
[0068] In step S103, the segmented scene image is input into a pre-trained scene recognition model to obtain the scene classification result of the segmented scene image.
[0069] In embodiments of the present invention, a scene recognition model is used to classify the segmented scene images. The scene recognition model is pre-built and trained, and the specific training process is as follows: the segmented images in the semantic segmentation database are classified according to scene classification criteria to obtain the scene classification results of the segmented images; the pre-built scene recognition model is trained using the segmented images and their corresponding scene classification results as training data to obtain the trained scene recognition model, wherein the scene recognition model uses the GoogLeNet network as the main structure.
[0070] In step S104, based on the scene classification results, the segmented scene images are input into a pre-trained feature extraction and prediction network for channel prediction, thereby obtaining the channel prediction results of the scene images for the channel to be tested.
[0071] After determining the category of the segmented scene image, it is input into a pre-trained feature extraction and prediction network for channel prediction. The feature extraction and prediction network is pre-built and trained, and its training data consists of known scenes placed into the segmented scene images and corresponding channel measurement data. The specific training process is as follows:
[0072] First, the channel measurement data in the training database is normalized. As a specific implementation, the channel measurement data is normalized using the min-max normalization method. The normalization formula is:
[0073]
[0074] Where X represents the normalized data, and x represents the original channel measurement data. min x is the minimum value of the channel measurement data. max This represents the maximum value of the channel measurement data.
[0075] Min-max normalization is a commonly used data normalization method used to scale data to a specific range [0,1]. This normalization method can help algorithms converge faster and improve their accuracy and stability. Furthermore, when the value ranges of features differ significantly, normalization can prevent certain features from having an excessive impact on the model, thereby improving the model's generalization performance.
[0076] Secondly, based on the scene classification results, the segmented scene images and channel measurement data of the same scene are used as training data to train the pre-constructed feature extraction and prediction network. Multiple rounds of training are performed for different scenes to obtain the trained feature extraction and prediction network.
[0077] Among them, such as Figure 5 As shown, the feature extraction and prediction network consists of five interconnected convolutional blocks containing varying numbers of 3×3 convolutional kernels as the image feature extraction structure. In this embodiment of the invention, to reduce training load, model parameters pre-trained on the ImageNet image dataset are used for image feature extraction. Pre-trained models are typically trained on large-scale datasets, possessing strong feature extraction and generalization capabilities, adapting to different datasets, and improving the model's transferability. Simultaneously, it can reduce training time and the number of model parameters, reducing the risk of overfitting.
[0078] Five fully connected layers are connected after the last convolutional block. The first four fully connected layers reduce the dimensionality of the extracted features, and the last fully connected layer outputs a channel prediction result with a dimension of 1. The weights of the prediction network are updated by the Adam optimizer, which performs well on larger datasets. The activation function is ReLU, and the output nodes of the five fully connected layers are 4096, 2000, 500, 100, and 1, respectively. The parameters of the feature extraction and prediction networks are shown in Table 2.
[0079] Table 2 Feature extraction and prediction network structure settings
[0080]
[0081]
[0082] The feature extraction and prediction network is trained by inputting segmented images from the same scene until the model converges. Images from the same scene share similar environmental features; further feature extraction and prediction channel modeling of images from the same scene allows for the perception of more environmental details with the same training load. This process is repeated for different scenes to obtain the prediction channel model parameters for each scene.
[0083] In embodiments of this invention, the predicted channel model can be divided into three parts: a semantic segmentation model, a scene recognition model, and a feature extraction and prediction network. Image semantic segmentation technology is introduced to identify and segment scatterers in the environment, extracting effective scatterer location information. The image is divided into different scenes based on the segmentation results and input into the model for scene recognition. Through scene recognition, subsequent feature extraction is performed in similar scenes, facilitating the extraction of more subtle environmental features. Semantically segmented images from the same scene are input into the feature extraction network to complete the training of the predicted channel model. The pre-trained predicted channel model can achieve channel prediction for unknown scenes, such as... Figure 6 , Figure 7 and Figure 8 As shown, after obtaining the trained prediction channel model parameters, the performance of the prediction channel model can be analyzed. Specifically, in the embodiments of this invention, the path loss obtained from the prediction channel model and the path loss from channel measurement data are used to analyze the accuracy of channel prediction. The performance of the 3GPP Uma model and the channel prediction model without image processing is compared in this experimental scenario. The simulation results are shown in Table 3, demonstrating the performance comparison between the prediction channel model and the 3GPP Uma model and the image prediction channel model without semantic segmentation. The simulation results show that the performance of the prediction channel modeling method based on image processing and machine learning far exceeds that of the 3GPP Uma model in the experimental scenario. Furthermore, because image semantic segmentation is incorporated into the algorithm, the role of scatterers affecting the channel in channel prediction is strengthened. Therefore, the performance of the prediction channel modeling method based on image processing and machine learning is superior to that without image processing.
[0084] Table 3 Simulation Results
[0085]
[0086] The channel prediction method based on image processing and machine learning proposed in this embodiment of the invention introduces semantic segmentation technology, which allows for more flexible input of environmental information, thereby improving model accuracy and ultimately achieving higher precision than traditional channel models. This helps to better meet the technical requirements of multi-band, multi-scenario full coverage in 6G systems, and solves the problems of parameter limitations and high cost and difficulty of channel measurement in traditional channel modeling. It realizes real-time channel prediction based on scene images and has high environmental adaptability. This model has reference value for the design of channel modeling algorithms based on environmental awareness.
[0087] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0088] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0089] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.
Claims
1. A channel prediction method based on image processing and machine learning, characterized in that, The method comprises the following steps: obtaining a scene picture of a to-be-tested channel; inputting the scene picture into a pre-trained semantic segmentation model to obtain a segmented scene picture; inputting the segmented scene picture into a pre-trained scene recognition model to obtain a scene classification result of the segmented scene picture; according to the scene classification result, inputting the segmented scene picture into a pre-trained feature extraction and prediction network for channel prediction to obtain a channel prediction result of the scene picture of the to-be-tested channel; further comprising: obtaining channel measurement data and scene pictures of existing frequency bands and scenes to construct a training database of multiple frequency bands and multiple scenes; annotating the scene pictures in the training database according to the types of channel scatterers to obtain a semantic segmentation dataset, and training a pre-constructed semantic segmentation neural network using the semantic segmentation dataset to obtain a trained semantic segmentation model.
2. The method of claim 1, wherein, Training the pre-constructed semantic segmentation neural network using the semantic segmentation dataset to obtain the trained semantic segmentation model comprises: using MobileNetV2 as the backbone network of the encoder to expand the feature map from shallow to deep; constructing a DeepLabV3+ semantic segmentation neural network based on the MobileNetV2 backbone network, wherein different dilated convolution layers with different dilation rates are selected to extract features of different scales; training the semantic segmentation neural network using the semantic segmentation dataset to obtain the trained semantic segmentation model.
3. The method of claim 2, wherein, In the trained semantic segmentation model, the segmentation accuracy of the semantic segmentation model is evaluated by pixel accuracy, wherein the pixel accuracy is: ; In the formula, TP represents the number of pixels that are correctly predicted as positive, FP represents the number of pixels that are incorrectly predicted as positive.
4. The method of claim 1, wherein, Further comprising: classifying the segmented pictures in the semantic segmentation database according to the scene classification standard to obtain a scene classification result of the segmented pictures; training a pre-constructed scene recognition model using the segmented pictures and their corresponding scene classification results as training data to obtain a trained scene recognition model, wherein the scene recognition model uses GoogLeNet network as the main structure.
5. The method of claim 1, wherein, Further comprising: normalizing the channel measurement data in the training database; based on the scene classification result, training a pre-constructed feature extraction and prediction network using the segmented scene pictures and channel measurement data of the same scene as training data, and performing multiple rounds of training for different scenes to obtain a trained feature extraction and prediction network, wherein the feature extraction and prediction network comprises five connected convolution blocks containing different numbers of 3*3 convolution kernels as an image feature extraction structure, and five fully connected layers are connected after the last convolution block, wherein the first four fully connected layers reduce the dimension of the extracted features, and the last fully connected layer outputs a channel prediction result with a dimension of 1.
6. The method of claim 5, wherein, The normalization of the channel measurement data in the training database comprises: normalizing the channel measurement data according to the min-max normalization method: ; wherein, X is the normalized data, x is the original channel measurement data, is the channel measurement data minimum value, x max is the channel measurement data maximum value.
Citation Information
Patent Citations
Wireless channel large-scale fading modeling method and device
CN110213003A
Distributed scatterer deformation monitoring method, system and equipment based on deep learning
CN114578356A