Communication index prediction method based on multi-modal large model and related equipment

Through the multimodal large model, the timing, spatial and semantic data are integrated to generate joint feature representations, which solves the problem of insufficient prediction accuracy of the single-modal prediction model in complex network scenarios, realizes high-precision communication index prediction, and supports network optimization decision-making.

CN120090946AActive Publication Date: 2025-06-03SHENZHEN RES INST OF BIG DATA

Patent Information

Application Number
CN202510542918.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-06-03
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing single-modal single-index prediction model is difficult to fully analyze user behavior patterns in complex network scenarios, resulting in limited prediction accuracy of communication network optimization decisions.

Method used

The communication index prediction method based on a multimodal large model is adopted, and by obtaining multi-dimensional timing data, geospatial data and text description data, timing feature extraction, spatial feature modeling and semantic coding are performed. Combined with the cross-modal fusion capability of the large language model, joint feature representations are generated to output communication index prediction information.

Benefits of technology

Effectively model the coupling relationship between multimodal indicators, improve the accuracy of communication indicator prediction, ensure the effectiveness of network optimization decisions, reduce network congestion risks, and improve user experience and operation and maintenance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120090946A_ABST
    Figure CN120090946A_ABST
Patent Text Reader

Abstract

The invention provides a communication index prediction method based on a multi-modal large model and related equipment, relates to the technical field of data processing, and can obtain multi-dimensional time sequence data, geographic space data and corresponding text description data in a user communication process; performing time sequence feature extraction processing on the multi-dimensional time sequence data to obtain a time sequence feature vector, and performing spatial feature extraction processing on the geographic spatial data to obtain a spatial feature vector; encoding the text description data based on a large language model word embedding layer to obtain a semantic feature vector; performing cross-modal attention fusion on the three types of feature vectors based on a text prototype to generate joint feature representation; and the joint feature representation is input into the large language model to generate communication index prediction information, so that the coupling relationship among the multi-modal indexes can be effectively modeled, the prediction precision is improved, and the effectiveness of a network optimization decision is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of data processing, and in particular, to a communication metric prediction method and related devices based on a multimodal large model. Background Art

[0002] In the actual application of communication networks, to ensure service quality and optimize resource allocation, network operators need to predict future network congestion, signal attenuation, and performance issues based on users' historical communication metrics, so as to formulate forward-looking network optimization strategies to improve user experience and reduce operation and maintenance costs.

[0003] Existing user-side performance modeling methods usually analyze based on the traffic time-series characteristics in the internal communication environment. However, this single-modal single-metric prediction model can only capture local information and is difficult to comprehensively analyze the user behavior patterns in complex network scenarios, resulting in limited prediction accuracy and thus affecting the effectiveness of network optimization decisions. Summary of the Invention

[0004] The embodiments of the present application provide a communication metric prediction method and related devices based on a multimodal large model, which can effectively model the coupling relationship between multimodal metrics, improve prediction accuracy, and thus ensure the effectiveness of network optimization decisions.

[0005] In a first aspect, the embodiments of the present application provide a communication metric prediction method based on a multimodal large model, including: obtaining multi-dimensional time-series data, geospatial data, and text description data corresponding to the multi-dimensional time-series data during the user's communication process; performing time-series feature extraction processing on the multi-dimensional time-series data to obtain a time-series feature vector; performing spatial feature extraction processing on the geospatial data to obtain a spatial feature vector; encoding the text description data based on the word embedding layer of a pre-trained large language model to obtain a semantic feature vector; performing cross-modal attention fusion on the time-series feature vector, the spatial feature vector, and the semantic feature vector based on the text prototype of the large language model to generate a joint feature representation; inputting the joint feature representation into the large language model so that the large language model obtains a corresponding inference vector based on the joint feature representation and outputs communication metric prediction information according to the inference vector.

[0006] In some embodiments, the performing time-series feature extraction processing on the multi-dimensional time-series data to obtain a time-series feature vector includes: constructing a convolutional neural network based on the attention mechanism according to the feature dimension of the multi-dimensional time-series data; dividing the multi-dimensional time-series data into multiple local time-series blocks along the time series dimension, and respectively extracting local time-series features from each local time-series block through the convolutional neural network, and obtaining a time-series feature vector according to all the local time-series features.

[0007] In some embodiments, the spatial feature extraction process for the geospatial data to obtain a spatial feature vector includes: dividing the geospatial data into multiple grid cells and marking the grid cells of the serving cells involved in the user communication process; performing convolutional feature extraction on all the grid cells to generate local spatial features; superimposing learnable embedding vectors on the local spatial features corresponding to all the serving cell grid cells to obtain locally enhanced features; and obtaining a spatial feature vector based on the local spatial features of all the unmarked grid cells and the locally enhanced features of all the serving cell grid cells.

[0008] In some embodiments, the cross-modal attention fusion of the temporal feature vector, the spatial feature vector, and the semantic feature vector based on the text prototype of the large language model to generate a joint feature representation includes: generating a first text prototype and a second text prototype based on the text dictionary data of the large language model; using the temporal feature vector as the query vector and the first text prototype as the key-value vector to generate the temporal feature of the text modality through attention weight calculation; using the spatial feature vector as the query vector and the second text prototype as the key-value vector to generate the spatial feature of the text modality; and concatenating the temporal feature of the text modality, the spatial feature of the text modality, and the semantic feature vector to generate a multi-modal joint feature representation.

[0009] In some embodiments, the process of obtaining the geospatial data includes: obtaining the base station location data involved in the user communication process; determining a target map area according to the base station location data and preset engineering parameters, and obtaining the geospatial data corresponding to the target map area from a preset map library.

[0010] In some embodiments, the process of obtaining the text description data corresponding to the multi-dimensional temporal data includes: constructing a dataset introduction text containing the meanings of multiple communication metrics in the multi-dimensional temporal data; constructing a task description text containing the definitions of the historical sequence length and the prediction sequence length in the communication metric prediction; and concatenating the dataset introduction text and the task description text to obtain the text description data.

[0011] In some embodiments, the training process of the large language model includes: dividing the multi-dimensional temporal data in the training data into multiple samples according to a time sliding window, each sample containing an input sequence and an output sequence; calculating the predicted output of the large language model to be trained for the input sequence, and calculating the model loss according to the predicted output and the output sequence; and updating the input-output mapping layer parameters of the large language model based on the model loss to obtain a trained large language model.

[0012] Second aspect, an embodiment of the present application provides a communication metric prediction device based on a multimodal large model, including: a data acquisition module, configured to acquire multi-dimensional time-series data, geospatial data, and text description data corresponding to the multi-dimensional time-series data during the user's communication process; a time-series processing module, configured to perform time-series feature extraction processing on the multi-dimensional time-series data to obtain a time-series feature vector; a spatial processing module, configured to perform spatial feature extraction processing on the geospatial data to obtain a spatial feature vector; a text processing module, configured to encode the text description data based on the word embedding layer of a pre-trained large language model to obtain a semantic feature vector; a feature fusion module, configured to perform cross-modal attention fusion on the time-series feature vector, the spatial feature vector, and the semantic feature vector based on the text prototype of the large language model to generate a joint feature representation; a prediction and inference module, configured to input the joint feature representation into the large language model, so that the large language model obtains a corresponding inference vector based on the joint feature representation and outputs communication metric prediction information according to the inference vector.

[0013] Third aspect, an embodiment of the present application provides an electronic device, including: at least one processor; at least one memory, configured to store at least one program; when at least one of the programs is executed by at least one of the processors, the communication metric prediction method based on the multimodal large model according to any one of the first aspect is implemented.

[0014] Fourth aspect, an embodiment of the present application provides a computer-readable storage medium, storing computer-executable instructions, where the computer-executable instructions are used to execute the communication metric prediction method based on the multimodal large model according to any one of the first aspect.

[0015] A communication metric prediction method and related devices proposed in this application, based on a multimodal large model, can integrate multi-dimensional time-series data, geospatial data, and text description data in the user communication process to construct a multimodal feature fusion framework, so as to solve the problem that existing single-modal single-index prediction models can only capture local information and are difficult to comprehensively analyze the user behavior patterns in complex network scenarios. Among them, this application can use the attention mechanism and convolutional neural network to extract time-series feature vectors from multi-dimensional time-series data respectively, extract spatial feature vectors from geospatial data, and combine the semantic encoding ability of the large language model to analyze the context information of text description data to obtain semantic feature vectors. It can be understood that the time-series feature vectors reflect the communication metrics during communication, the geospatial data reflects the environmental information during communication, and the semantic feature vectors reflect the communication metrics and the background knowledge of the dataset of the prediction task, enabling the large language model to have the semantic reasoning ability for communication metrics and environmental information. Further, this application can, based on the text prototype of the large language model, dynamically interact the time-series features, spatial features, and semantic features through a cross-modal attention mechanism, associate the time-series features and spatial features with the semantic features, and obtain a joint feature representation in the text modality. The joint feature representation can be input into the pre-trained large language model, and its semantic reasoning ability can be used to map multi-modal information including time-series, space, and semantics into future communication metric prediction results, so as to model the spatio-temporal correlation and semantic logic in the network environment, improve the prediction accuracy, enable the operator to optimize the base station resource allocation in advance based on the high-precision prediction results, thereby reducing the network congestion risk and improving the user experience and operation and maintenance efficiency. Description of the Drawings

[0016] Figure 1 It is a method flow chart of the communication metric prediction method based on a multimodal large model provided by an embodiment of this application; Figure 2 It is a method flow chart of obtaining time-series feature vectors in the communication metric prediction method based on a multimodal large model provided by an embodiment of this application; Figure 3 It is a method flow chart of obtaining spatial feature vectors in the communication metric prediction method based on a multimodal large model provided by an embodiment of this application; Figure 4 It is a method flow chart of generating a joint feature representation in the communication metric prediction method based on a multimodal large model provided by an embodiment of this application; Figure 5 It is a method flow chart of obtaining the geospatial data in the communication metric prediction method based on a multimodal large model provided by an embodiment of this application; Figure 6In a communication metric prediction method based on a multimodal large model provided by an embodiment of the present application, a flowchart of a method for obtaining text description data corresponding to the multi-dimensional time series data; Figure 7 In a communication metric prediction method based on a multimodal large model provided by an embodiment of the present application, a flowchart of a method for training the large language model; Figure 8 An overall model flowchart of a communication metric prediction method based on a multimodal large model provided by an embodiment of the present application; Figure 9 In a communication metric prediction method based on a multimodal large model provided by an embodiment of the present application, an overall flowchart for obtaining a spatial feature vector; Figure 10 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0017] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0018] In some embodiments, although functional modules are divided in the system schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the system or the order in the flowchart. Terms such as first and second in the description of the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0019] In addition, unless otherwise clearly specified and limited, the term "connected / linked" should be understood in a broad sense. For example, it can be a fixed connection or a movable connection, or a detachable connection or an inseparable connection, or an integral connection; it can be a mechanical connection, an electrical connection or can communicate with each other; it can be directly connected or indirectly connected through an intermediate medium.

[0020] In the description of the embodiments of the present application, the descriptions with reference to terms such as "one embodiment / embodiment manner", "another embodiment / embodiment manner" or "certain embodiments / embodiment manners", "in the above embodiments / embodiment manners", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least two embodiments or embodiment manners disclosed in the present application. In the disclosure of the present application, the schematic expressions of the above terms do not necessarily refer to the same embodiment or embodiment manner. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the order in the flowchart.

[0021] In the practical application of communication networks, to ensure service quality and optimize resource allocation, network operators need to predict future network congestion, signal attenuation, and performance issues based on users' historical communication metrics, so as to formulate forward-looking network optimization strategies to improve user experience and reduce operation and maintenance costs. Existing user-side performance modeling methods usually analyze based on the traffic time-series characteristics in the internal communication environment and achieve communication metric prediction through isolated modeling. However, in the internal communication environment, in addition to traffic characteristics, there are also other metrics such as signal level, delay, packet loss rate, and retransmission rate. Outside the internal communication environment, there are also metrics of other modalities such as the external communication environment. These metrics are correlated in the spatio-temporal dimension and jointly reflect the dynamic change characteristics of the network environment. However, the existing single-modal single-index prediction model can only capture local information and is difficult to comprehensively analyze the user behavior patterns in complex network scenarios, resulting in limited prediction accuracy and further affecting the effectiveness of network optimization decisions.

[0022] To solve the above problems, this application notes that communication metric prediction needs to integrate the fluctuations in the time dimension, spatial distribution characteristics, and semantic description information. Existing single-modal models cannot effectively represent the correlation of multi-dimensional data, especially lacking quantitative analysis of geospatial characteristics. Through analysis, it is found that the spatial correlation between base station location data and user movement trajectories affects signal strength, and text descriptions can provide semantic interpretations of communication scenarios. Based on this, this application proposes to combine time-series feature extraction, spatial feature modeling, and natural language processing, and use the cross-modal fusion ability of large language models to construct a joint representation. Further exploration reveals that the attention mechanism can effectively align the distributions of different modal features in the semantic space, and through text prototypes to guide cross-modal interactions, realizing the deep coupling of spatio-temporal features and semantic information.

[0023] Therefore, the embodiments of this application provide a communication metric prediction method and related devices based on a multi-modal large model. This application can integrate multi-dimensional time-series data, geospatial data, and text description data in the user communication process to construct a multi-modal feature fusion framework to solve the problem that the existing single-modal single-index prediction model can only capture local information and is difficult to comprehensively analyze the user behavior patterns in complex network scenarios. Among them, this application can use the attention mechanism and convolutional neural network to extract time-series feature vectors from multi-dimensional time-series data, extract spatial feature vectors from geospatial data, and combine the semantic encoding ability of the large language model to analyze the context information of text description data to obtain semantic feature vectors, which can effectively model the coupling relationship between multi-modal metrics, improve prediction accuracy, and further ensure the effectiveness of network optimization decisions.

[0024] The embodiments of this application will be further described below with reference to the accompanying drawings.

[0025] Refer to Figure 1 ,Figure 1 This is a flowchart of a communication metric prediction method based on a multimodal large model provided by an embodiment of the present application; in some embodiments, the present application proposes a communication metric prediction method based on a multimodal large model, which at least includes the following steps: Step S110: Obtain multi-dimensional time-series data, geospatial data, and text description data corresponding to the multi-dimensional time-series data during the user's communication process; Step S120: Perform time-series feature extraction processing on the multi-dimensional time-series data to obtain time-series feature vectors; Step S130: Perform spatial feature extraction processing on the geospatial data to obtain spatial feature vectors; Step S140: Encode the text description data based on the word embedding layer of a pre-trained large language model to obtain semantic feature vectors; Step S150: Perform cross-modal attention fusion on the time-series feature vectors, spatial feature vectors, and semantic feature vectors based on the text prototype of the large language model to generate a joint feature representation; Step S160: Input the joint feature representation into the large language model, so that the large language model obtains corresponding inference vectors based on the joint feature representation and outputs communication metric prediction information according to the inference vectors.

[0026] Among them, the multi-dimensional time-series data can be continuous time-series data containing multiple communication metrics such as traffic, latency, and packet loss rate. Specifically, it can be implemented using minute-level sampling data collected by sensors, which is used to characterize the dynamic changes of the communication network state; the geospatial data can be geospatial information data containing base station locations, terrain elevations, and building distributions. Specifically, it can be implemented using vector layer data provided by a GIS system, which is used to analyze the impact of the spatial environment on signal propagation; the text description data can be natural language text describing the characteristics of the communication scenario. Specifically, it can be implemented using manually annotated or automatically generated scenario description texts, which are used to provide context information at the semantic level; the time-series feature extraction processing can be to capture local patterns in the time series through a convolutional neural network. Specifically, it can be implemented by dividing the time series into segments using a sliding window and extracting periodic features, which is used to identify short-term fluctuation patterns of communication metrics; the spatial feature extraction processing can be to extract local spatial features of the geographic grid through convolutional operations. Specifically, it can be implemented by dividing the geographic grid and overlaying service cell embedding vectors, which is used to enhance the spatial representation ability of key regions; the cross-modal attention fusion can be to calculate the attention weights of different modal features using the text prototype as the key-value vector. Specifically, it can be implemented using a multi-head attention mechanism to align the time-series features with the semantic spatial distribution, which is used to eliminate the semantic gap between modalities.

[0027] It can be understood that the time-series feature vector reflects the communication metrics during communication, the geospatial data reflects the environmental information during communication, and the semantic feature vector reflects the communication metrics and the background knowledge of the dataset for the prediction task, enabling the large language model to have the semantic reasoning ability for communication metrics and environmental information; further, the present application can, based on the text prototype of the large language model, dynamically interact the time-series features, spatial features, and semantic features through a cross-modal attention mechanism, associate the time-series features and spatial features with the semantic features, and obtain a joint feature representation in the text modality. The joint feature representation can be input into the pre-trained large language model, and its semantic reasoning ability can be used to map multi-modal information including time series, space, and semantics into the prediction result of future communication metrics, so as to model the spatio-temporal correlation and semantic logic in the network environment, improve the prediction accuracy, enable the operator to optimize the base station resource allocation in advance based on the high-precision prediction result, thereby reducing the network congestion risk and improving the user experience and operation and maintenance efficiency. In some embodiments, the multi-dimensional time-series data during the communication process is divided into multiple local time-series blocks, and the periodic features and trend features are extracted through an attention convolutional network to form a feature vector with time dependence; the geospatial data is converted into a grid representation, a learnable position embedding is superimposed on the grid where the serving cell is located, and the spatial distribution pattern is extracted through a convolutional layer; the text description data is converted into a semantic vector through the word embedding layer of the large language model, and the context relationship of the natural language description is retained.

[0028] In some embodiments, in the feature fusion stage, the time-series feature vector and the spatial feature vector can respectively perform attention interaction with the text prototype to generate cross-modal features aligned with the semantic space. The fused joint feature representation is input into the large language model for autoregressive decoding, and the prediction value of the communication metrics in the future period is output through the sequence generation method, so as to realize the collaborative modeling of the dynamic changes in the time dimension, the spatial distribution characteristics, and the semantic description information.

[0029] It can be understood that, compared with the prior art, the traditional single-modal prediction method only analyzes the isolated time-series changes of traffic metrics and does not consider the mutual influence between multi-dimensional metrics. For example, the negative correlation between the packet loss rate and the signal strength is not modeled, and the existing spatial analysis methods usually independently process the base station location data and do not perform joint optimization with the time-series features; while in this solution, through the cross-modal attention mechanism, the geospatial features are mapped to the text semantic space, enabling the base station engineering parameter features to be associated with the signal attenuation trend in the time dimension; in addition, the traditional text processing methods are difficult to associate structured communication data, while this solution uses the text prototype of the large language model to guide feature fusion, enabling concepts such as "congestion during peak hours" in the semantic description to be directly associated with the traffic mutation pattern in the time-series features.

[0030] It should be noted that through the above technical solution, the present application can effectively integrate the multi-modal data features in the communication scenario, solve the problem of the disconnection between spatio-temporal features and semantic information, that is, capture the dynamic evolution law of communication metrics through temporal feature extraction, model the spatial features to quantify the impact of the geographical environment on signal quality, provide semantic explanations for the scenario through text encoding, and then eliminate the representational differences between different modalities through cross-modal attention fusion, so that the base station distribution features can interact synergistically with the time series fluctuations, and use the reasoning ability of the large language model to convert the joint features into interpretable prediction results, thereby improving the accuracy of network congestion prediction in complex scenarios and providing a reliable basis for dynamically adjusting the base station load distribution.

[0031] Reference Figure 2 , Figure 2 FIG. is a flowchart of a method for obtaining a temporal feature vector in a communication metric prediction method based on a multi-modal large model provided by an embodiment of the present application; in some embodiments, temporal feature extraction processing is performed on multi-dimensional temporal data to obtain a temporal feature vector, which at least includes the following steps: Step S210, constructing a convolutional neural network based on the attention mechanism according to the feature dimensions of the multi-dimensional temporal data; Step S220, dividing the multi-dimensional temporal data into multiple local temporal blocks along the time series dimension, and respectively extracting local temporal features from each local temporal block through the convolutional neural network, and obtaining a temporal feature vector according to all the local temporal features.

[0032] Among them, the convolutional neural network based on the attention mechanism can refer to a network architecture composed of a convolutional layer and an attention module connected in series. Specifically, a one-dimensional convolutional layer can be used to extract local features and then connect to a self-attention module to implement weight distribution in the temporal dimension. This structure can adaptively focus on the correlation of feature dimensions at key time nodes and solve the problem of insufficient sensitivity of traditional convolutional networks to temporal dynamic changes.

[0033] Among them, the local temporal block can be a continuous data segment divided according to a preset time window length. Specifically, it can be implemented by using a sliding window or a fixed-length block method. Through the block operation, the long sequence can be decomposed into multiple short segments, which is convenient for capturing local time domain features and reducing the modeling difficulty of long-range dependencies.

[0034] In some embodiments, the process of obtaining multi-dimensional time-series data during the user communication process may include collecting user measurement reports of all communication processes between the base station and the user equipment at the base station side, classifying them according to different users, extracting the time-series user measurement report sequence of one user, then selecting multiple key communication metric fields from them to form multi-dimensional time-series data, and constructing a multi-variable time-series prediction task through the multi-dimensional time-series data to capture the user's communication service situation during the corresponding time period; specifically, it may include the following 17 key communication metric fields: DL GrantCount (downlink scheduling grant count), DL RB Num (downlink resource block number), UL Grant Count (uplink scheduling grant count), UL RB Num (uplink resource block number), DL MCS AVG (downlink average modulation and coding strategy), DL RANK (downlink rank), UL MCS AVG (uplink average modulation and coding strategy), UL RANK (uplink rank), CQI (channel quality indicator), DL IBLER (downlink initial block error rate), DL RBLER (downlink retransmission block error rate), UL IBLER (uplink initial block error rate), UL RBLER (uplink retransmission block error rate), UL DMRS RSRP (uplink demodulation reference signal received power), UL DMRS SINR (uplink demodulation reference signal signal-to-noise ratio), UL SRS RSRP (uplink sounding reference signal received power), UL SRS SINR (uplink sounding reference signal signal-to-noise ratio). By constructing a multi-variable time-series prediction task for these fields, the user's communication service situation during the corresponding time period can be captured.

[0035] It can be understood that after the multi-dimensional time-series data is segmented into multiple time segments, local feature extraction is performed along the feature dimension through a convolution kernel. Each convolution kernel corresponds to the local pattern recognition of one feature dimension. The attention module recalibrates the weights of the convolution output, and filters out key features by calculating the correlation degree between time steps. After all local time-series blocks are weighted and fused, a global time-series feature vector is formed, which not only retains local dynamic characteristics but also constructs global time correlations.

[0036] In some embodiments, corresponding to steps S210 to S220, feature processing is performed through the designed convolutional neural network based on the attention mechanism. After the time-series data extracts features through the Attention mechanism of the feature dimension, it is then sliced into multiple patch small blocks along the time-series dimension. For example, the duration of each small block can be 2 time-series records.

[0037] It should be noted that, compared with the prior art, the traditional method uses a single convolutional layer or a fully connected network to process time series data, which cannot distinguish the contribution differences of different feature dimensions to time series changes and lacks the fine-grained modeling ability for local time windows. This solution realizes local feature focusing through block operations and combines the attention mechanism to enhance the correlation modeling between feature dimensions, which can solve the problem of insufficient capture of the coupling relationship of multi-dimensional indicators.

[0038] Reference Figure 3 , Figure 3 FIG. is a flowchart of a method for obtaining a spatial feature vector in a communication metric prediction method based on a multi-modal large model provided by an embodiment of the present application; in some embodiments, performing spatial feature extraction processing on geospatial data to obtain a spatial feature vector includes at least the following steps: Step S310, dividing the geospatial data into a plurality of grid cells and marking the grid cells of the serving cells involved in the user communication process; Step S320, performing convolutional feature extraction on all grid cells to generate local spatial features; Step S330, superimposing learnable embedding vectors on the local spatial features corresponding to all serving cell grids to obtain local enhanced features; Step S340, obtaining a spatial feature vector according to the local spatial features of all unmarked grid cells and the local enhanced features of all serving cell grids.

[0039] Among them, the grid cell can be a discrete area formed by dividing the geospatial data according to a preset resolution, and can be specifically implemented by a square grid with a fixed size or a dynamically adjusted size. The continuous space is converted into structured data through discretization processing. The serving cell grid can be a grid cell corresponding to the coverage area of the base station actually connected during the user communication process, and can be specifically marked through a preset mapping relationship between the base station coordinates and the grid for positioning the core influence area. Convolutional feature extraction can be understood as using a convolutional neural network to perform spatial feature encoding on the local neighborhood of the grid cell, and can be specifically implemented by a stacked structure of multiple convolutional kernels to capture the spatial correlation between adjacent grids through a sliding window operation. The learnable embedding vector can be a feature enhancement parameter automatically optimized through model training, and can be specifically initialized as a random vector and updated during the backpropagation process to assign dynamic weights to the serving cell grids to distinguish their importance.

[0040] It can be understood that the geospatial data is first divided into grid cells. Among them, the service cell grid is marked as a key area through the mapping relationship between the base station location and the grid coordinates. Each grid cell and its neighboring area are input into a convolutional neural network for local spatial feature extraction, so as to capture the influence of spatial attributes such as terrain and building distribution on signal propagation. For the service cell grid, a learnable embedding vector is superimposed on the basis of the extracted local spatial features. This vector is automatically adjusted through gradient descent during the training process, enabling the model to adaptively enhance the spatial feature expression related to the service cell. The unmarked grid cells retain the original convolutional features to avoid excessive interference of the features in non-critical areas. Furthermore, by splicing the features of all grid cells, a complete spatial feature vector containing spatial heterogeneity information can be formed, while retaining the global spatial distribution pattern and highlighting the local influence of the service cell.

[0041] In some embodiments, dividing the geospatial data into multiple grid cells includes: dynamically adjusting the grid size according to the base station coverage density, or using a sliding window method to perform multi-scale division of the map. Among them, small-sized grids are used for high-density coverage areas, and large-sized grids are used for low-density coverage areas. It can be understood that through the dynamic grid division strategy, the communication activity characteristics of different regions are adaptively matched, enhancing the flexibility and environmental adaptability of spatial feature extraction and improving the prediction accuracy.

[0042] It is worth noting that traditional methods usually perform global feature extraction or uniform processing on geospatial data, and cannot distinguish the spatial correlation differences between service cells and the surrounding environment. Through the grid marking and feature enhancement mechanism, this solution dynamically adjusts the weight of the service cell grid while retaining the overall spatial structure, enabling the spatial feature vector to more accurately represent the interaction relationship between the core area and the surrounding environment. Through the combination of grid division and feature enhancement, not only the global spatial distribution characteristics are retained, but also the key information in the area where the service cell is located is strengthened, avoiding the loss of core area features caused by uniform processing in traditional methods, enabling the spatial feature vector to reflect the spatial correlation between the core area covered by the network and the overall environment, and providing a more complete spatial information basis for subsequent cross-modal fusion.

[0043] Reference Figure 4 , Figure 4 FIG. is a flowchart of a method for generating a joint feature representation in a communication metric prediction method based on a multi-modal large model provided in an embodiment of the present application. In some embodiments, based on the text prototype of the large language model, cross-modal attention fusion is performed on the time series feature vector, spatial feature vector, and semantic feature vector to generate a joint feature representation, which at least includes the following steps: Step S410, generating a first text prototype and a second text prototype based on the text dictionary data of the large language model; Step S420: Using the temporal feature vector as the query vector and the first text prototype as the key vector, calculate the temporal features of the text modality through attention weight calculation. Step S430: Using the spatial feature vector as the query vector and the second text prototype as the key vector, generate the spatial features of the text modality. Step S440: Concatenate the temporal features of the text modality, the spatial features of the text modality, and the semantic feature vector to generate a multi-modal joint feature representation.

[0044] Among them, the text prototype can be a set of semantically representative vectors extracted from the pre-trained dictionary of the large language model. Specifically, it can be constructed by selecting the word embedding vectors corresponding to high-frequency words, and is used to establish the semantic mapping relationship of cross-modal features. Cross-modal attention fusion can map different modal features to a unified semantic space through the attention mechanism. Specifically, the query-key-value attention calculation method is adopted to align the non-text modal features with the text prototype. Attention weight calculation can determine the matching degree between the query vector and the key vector through dot product operation combined with scaling operation. Specifically, the softmax function can be used for normalization to obtain the weight distribution, which is used to weighted aggregate the key vector to generate the modal fusion feature.

[0045] In some embodiments, two sets of semantically representative vocabulary vectors are selected from the text dictionary of the large language model to construct the first text prototype and the second text prototype respectively; the temporal feature vector is used as the query vector to perform attention calculation with the first text prototype to obtain the mapping representation of each temporal feature in the text semantic space, forming the temporal features of the text modality; similarly, the spatial feature vector performs attention calculation with the second text prototype to generate the spatial features of the text modality; in this way, the temporal and spatial features are converted into vector representations aligned with the text semantic space, eliminating the semantic gap between different modalities; furthermore, the temporal features, spatial features of the text modality and the original semantic feature vector are concatenated in the feature dimension to form a joint feature representation containing multi-modal information, providing a unified semantic input for subsequent prediction tasks.

[0046] In some embodiments, generating the first text prototype and the second text prototype based on the text dictionary data of the large language model may include: copying the dictionary of the text and its corresponding embedding vectors of the large language model twice as the first text prototype and the second text prototype, and storing them in the proposed model. This step can be achieved only by using the get_input_embeddings() method of the large language model class provided by Huggingface. After obtaining the two text prototypes, during the process of feature extraction from the temporal data and the map data respectively, their modalities can be converted from their respective modalities to the text modality, so that the large language model can understand and reason.

[0047] Furthermore, the process of generating the temporal features of the text modality by using the temporal feature vector as the query vector and the first text prototype as the key-value vector may include: using a one-dimensional convolutional layer and an MLP to extract each patch in the temporal feature vector into a Token vector, and performing Cross Attention processing with the preset first text prototype to obtain the temporal features of the text modality. The formula for Cross Attention can be: , , , , where Q is obtained by linear layer Linear calculation from the previously extracted temporal feature vector K and V are obtained by linear layer Linear calculation from the text prototype, the first text prototype (i.e., P in the formula).

[0048] Furthermore, the process of generating the spatial features of the text modality by using the spatial feature vector as the query vector and the second text prototype as the key-value vector may include: performing Cross Attention processing on the preset second text prototype to obtain the spatial features of the text modality. The Cross Attention here is the same as the Cross Attention process in the temporal data feature extraction stage, only replacing it with the previously obtained spatial feature vector and replacing P with the second text prototype.

[0049] It can be understood that through the above technical solution, the present application can solve the semantic fragmentation problem between multi-modal data, enabling the temporal dynamic change features, geographical spatial distribution features, and text semantic features to be jointly modeled in a unified semantic space. By establishing cross-modal mapping relationships, the comprehensive analysis ability of the model for multi-dimensional influencing factors of communication metrics is enhanced, and the accuracy and reliability of communication metric prediction in complex network environments are improved.

[0050] Refer to Figure 5 , Figure 5 which is the flowchart of the method for obtaining geographical spatial data in the communication metric prediction method based on a multi-modal large model provided by an embodiment of the present application; in some embodiments, the process of obtaining geographical spatial data includes at least the following steps: Step S510: Obtain the base station location data involved in the user's communication process; Step S520: Determine the target map area according to the base station location data and the preset engineering parameters, and obtain the geographical spatial data corresponding to the target map area from the preset map library.

[0051] Among them, the base station location data can be the three-dimensional coordinate information of the base station that establishes a wireless connection with the user communication device. Specifically, it can be obtained by collecting through the base station positioning interface or retrieving from the network topology database, and is used to establish the basic association between the user's communication behavior and the physical location. The preset engineering parameters can be the effective action area of the base station signal propagation. Specifically, it can be calculated and generated by using the honeycomb network hexagonal coverage model or the signal strength attenuation model, and is used to delimit the spatial boundary related to the base station service capacity. The preset map library can be a data set storing standardized geographical information. Specifically, it can be a multi-layer database including elevation data, building contour vector maps, and terrain classification raster maps, and is used to provide the key environmental feature data that affect the wireless signal propagation.

[0052] It can be understood that when the user initiates a communication request, the system obtains the location coordinates of the serving base station by parsing the signaling message. According to the transmission power of the base station and the propagation model, the effective coverage radius of the signal is calculated. For example, based on the free space loss formula, the engineering parameter is determined as the standard coverage area of 1 km. After superimposing this engineering parameter on the geographic coordinate system, a spatial polygon including the center point of the base station and the surrounding affected area is generated, and the rasterized data of all geographical elements within the polygon can be extracted by calling the pre-constructed map database interface, including layer information such as terrain elevation values, building heights, and vegetation distribution densities.

[0053] In some embodiments, the process of obtaining the geospatial data during the user's communication process may include: according to the base station locations involved in the communication process recorded in the single-user measurement report, combining with the engineering parameters (such as longitude and latitude) of the base station, obtaining a map from the open-source database OpenStreetMap website. For example, with the base station location as the center, expand outward to intercept the standard coverage area map (such as 1 km * 1 km). Among them, as an open-source map library, OpenStreetMap only needs to calculate the longitude and latitude range of the target map in advance according to the base station longitude and latitude and the selected range size, input it on the website to frame and obtain the corresponding map, and then use screenshot software to intercept it.

[0054] In some embodiments, the preset engineering parameters include the base station transmission power and the wireless propagation model; determining the specific target map area specifically includes: calculating the effective coverage radius of the signal based on the base station transmission power and the wireless propagation model; taking the base station location as the center, determining the range of the target map area according to the effective coverage radius of the signal; it can be understood that combining the propagation model with the transmission power to accurately calculate the signal coverage range can avoid the error of the artificially preset coverage range, make the division of the geospatial data more conform to the actual network environment, and improve the accuracy of the spatial feature representation.

[0055] It can be understood that through the above technical solution, the present application can solve the problem of the correlation between geospatial data and user communication behavior. By delimiting the target area based on the base station engineering parameters, the interference of geographical features in communication-unrelated areas can be excluded, ensuring that the subsequent spatial feature extraction focuses on the actual signal propagation impact area. The structured data source provided by the preset map library enhances the integrity of geographical elements, enabling key environmental factors such as terrain undulation and building occlusion to be accurately modeled, thereby improving the analytical ability of the spatial attenuation effect in communication metric prediction.

[0056] Reference Figure 6 , Figure 6 FIG. is a flowchart of a method for obtaining text description data corresponding to multi-dimensional time series data in a communication metric prediction method based on a multi-modal large model provided by an embodiment of the present application; in some embodiments, the process of obtaining text description data corresponding to multi-dimensional time series data at least includes the following steps: Step S610, constructing a dataset introduction text containing the meanings of multiple communication metrics in the multi-dimensional time series data; Step S620, constructing a task description text containing the definitions of the historical sequence length and the prediction sequence length in communication metric prediction; Step S630, splicing the dataset introduction text and the task description text to obtain the text description data.

[0057] Among them, the dataset introduction text can be a text paragraph that semantically interprets each communication metric in the multi-dimensional time series data. Specifically, natural language generation technology can be used to convert the metric name, unit, and physical meaning into a readable description. For example, "RSRP" is interpreted as "Reference Signal Received Power, which reflects the signal strength of the base station received by the user, with the unit of dBm". By clarifying the metric meaning, it helps the large language model establish the mapping relationship between data features and semantic concepts. The task description text can be a text paragraph that defines the time window of the time series modeling task. Specifically, a structured template can be used to generate an explanatory statement containing the historical sequence length and the prediction sequence length, and the modeling boundary of the data time series correlation is constrained by limiting the time range.

[0058] Specifically, when generating the text for introducing the dataset, by parsing the technical documents of each communication metric in the multi-dimensional time series data, elements such as the metric name, measurement unit, and physical meaning are extracted to form standardized metric explanation statements. For example, converting "RTT" to "Network round-trip delay, which represents the time for a data packet to travel from the sender to the receiver and back, with the unit of milliseconds". Such statements can map the technical parameters in the original data to semantic concepts that can be understood by the language model. When generating the task description text, based on the preset parameters of the prediction task, the time span of the historical sequence and the time span of the prediction target are clarified. For example, "Predict the metric changes in the next 15 minutes based on the communication metrics in the past 60 minutes". Such descriptions provide context constraints for the model's time series modeling. By splicing these two types of texts, a complete text description that includes both the data structure explanation and the task parameter definition is formed, enabling the large language model to understand the physical meaning of the time series data in combination with prior knowledge of natural language and establish cross-modal feature associations within the preset time window.

[0059] In some embodiments, corresponding to steps S610 to S630, the present application can construct text for introducing the dataset for the meaning of communication data metrics, store it in the model, and introduce the background knowledge of the dataset by introducing the text for introducing the dataset in the prediction, leveraging the large language model's understanding ability of the text modality, thereby improving the prediction accuracy. Then, task description information is spliced after this introduction text. For example, the task description information can include the length of the historical sequence and the length of the future sequence. It can be understood that through the above technical solution, the present application solves the problem of insufficient semantic understanding caused by the lack of text modality in traditional prediction models, enabling the large language model to accurately analyze the technical meanings of each metric in multi-dimensional time series data and clarify the time window constraints for time series modeling based on the task description, thereby improving the reasoning accuracy for complex communication scenarios. For example, when predicting network congestion, the model can focus on extracting the characteristics of short-term bursty traffic through the clear indication of "predicting the next 10 minutes" in the task description, avoiding prediction biases caused by ambiguous time windows.

[0060] Reference Figure 7 , Figure 7 FIG. is a flowchart of a method for training a large language model in a communication metric prediction method based on a multi-modal large model provided by an embodiment of the present application; in some embodiments, the training process of the large language model at least includes the following steps: Step S710, dividing the multi-dimensional time series data in the training data into multiple samples according to a time sliding window, and each sample includes an input sequence and an output sequence; Step S720, calculating the prediction output of the large language model to be trained for the input sequence, and calculating the model loss according to the prediction output and the output sequence; Step S730: Update the input-output mapping layer parameters of the large language model based on the model loss to obtain a trained large language model.

[0061] Among them, dividing the time sliding window into multiple samples can be understood as sliding and splitting continuous time series data according to a fixed window length. Specifically, a time series slicing method can be used to achieve this. Each window contains a historical input sequence and a corresponding future output sequence. This division method can retain the dynamic temporal correlation in the data and provide a mapping relationship between the clear historical state and the future prediction target for the model. Among them, the input-output mapping layer parameters can be the network layer parameters in the large language model responsible for converting input features into prediction outputs. Specifically, a linear transformation layer or a fully connected layer can be used to achieve this. This parameter update method enables the model to focus on learning the non-linear mapping relationship from multi-modal features to prediction results, avoiding over-adjusting the underlying feature extraction layer and damaging the learned cross-modal correlation features.

[0062] In some embodiments, during the training process of the large language model, it also includes: freezing the backbone network parameters of the large language model and only performing gradient updates on the parameters of the input mapping layer and the output mapping layer; the input mapping layer is used to adapt the multi-modal joint feature representation to the embedding space of the large language model, and the output mapping layer is used to map the inference vector of the large language model back to the time series data modality. It can be understood that by freezing the backbone network parameters, while retaining the pre-trained semantic capabilities of the large language model, the training calculation cost is significantly reduced and the model training efficiency is improved.

[0063] Specifically, in the training stage, the original multi-dimensional time series data is cut into multiple samples containing input-output sequences. These samples retain the dynamic temporal patterns in the data through the time sliding window division method. After receiving the input sequence, the model generates a predicted output. By calculating the loss function between the predicted output and the true output sequence, the deviation of the model in capturing temporal features can be quantified. When the model loss is backpropagated, only the parameters of the input-output mapping layer are adjusted, rather than updating the entire network parameters. This selective parameter update strategy enables the model to maintain the stability of the underlying feature extraction layer during the optimization process, while strengthening the mapping ability between the input features and the prediction target, thereby improving the generalization performance of the model in multi-modal data joint modeling.

[0064] In some embodiments, before training the large language model, the present application can obtain the backbone network of the large language model. For example, the large language model can be obtained through the Huggingface website as the prediction backbone. Transformers in Python can be used to interact with the Huggingface website. It provides the model Config class and Model class. In Python, simply setting the Config object and Model object can easily obtain the required large language model, such as LLaMA 3 7B. Further, the Tokenizer of the obtained large language model can be used as the Text Embedder to extract the semantic feature vector. It should be noted that the large language model can provide a directly usable Tokenizer, which is responsible for word segmentation and feature extraction of text data, so as to become the Token that can be directly processed by the backbone part of the large language model. In some embodiments, dividing the multi-dimensional time series data in the training data into multiple samples according to a time sliding window may include: dividing the multi-dimensional time series data in the training data, and dividing multiple samples according to the time sliding window. Each sample is further divided into an input sequence and an output sequence. For example, the input sequence has 96 time points and the output sequence has 24 time points.

[0065] It can be understood that compared with the prior art, the traditional single-modal single-index prediction model usually adopts a global parameter update strategy during training, resulting in the simultaneous adjustment of the underlying feature extraction layer and the high-level mapping layer, which is likely to destroy the learned cross-modal correlation features. Moreover, this method still relies on a large amount of pre-trained data. In this solution, by dividing samples and restricting the parameter update range, the underlying feature extraction layer is kept stable, and the input-output mapping layer specifically learns the mapping rules of multi-modal time series features, effectively avoiding the overfitting problem and improving the model's adaptability to complex time series patterns.

[0066] In some embodiments, during the inference stage of the large language model corresponding to steps S710 to S730, the present application can splice the time series features, spatial features, and semantic feature vectors of the text modality corresponding to the input sequence to generate a multi-modal joint feature representation, and send it to the backbone of the large language model for inference and prediction to obtain the corresponding inference vector. The prediction output obtained from the above process with the true value Calculate the mean square error loss function for training, where N is the number of samples, H is the length of the prediction time window, and D is the number of data feature dimensions. Furthermore, based on the model loss, update the parameters of the input-output mapping layer of the large language model, so that the large language model can map the inference vector back to the time series data modality through the linear layer as the final prediction output of the entire model, in order to obtain the trained large language model.

[0067] Through the above technical solution, this application solves the problem that existing models cannot effectively utilize the dynamic features of multi-dimensional time-series data. By dividing the time sliding window, the time-series correlation is retained. At the same time, by optimizing the parameters of specific network layers, the ability of joint modeling of multi-modal data is enhanced. This training mechanism enables the model to more accurately capture the coupling relationship between multi-modal metrics, thereby improving the accuracy of communication metric prediction and providing a reliable basis for network optimization decisions.

[0068] Reference Figure 8 , Figure 8 FIG. is the overall model flow chart of the communication metric prediction method based on the multi-modal large model provided by an embodiment of this application; in some embodiments, the specific implementation process of the communication metric prediction method based on the multi-modal large model can be as follows: multi-dimensional time-series data, geospatial data, and corresponding text description data during the user communication process can be obtained; time-series feature extraction processing is performed on the multi-dimensional time-series data to obtain a time-series feature vector, and spatial feature extraction processing is performed on the geospatial data to obtain a spatial feature vector; semantic feature vectors are obtained by encoding the text description data based on the large language model word embedding layer; cross-modal attention fusion is performed on the three types of feature vectors based on the text prototype to generate a joint feature representation; the joint feature representation is input into the large language model to generate communication metric prediction information, which can effectively model the coupling relationship between multi-modal metrics, improve the prediction accuracy, and further ensure the effectiveness of network optimization decisions.

[0069] Reference Figure 9 , Figure 9 FIG. is the overall flow chart of obtaining the spatial feature vector in the communication metric prediction method based on the multi-modal large model provided by an embodiment of this application. It can be understood that corresponding to steps S310 to S340, in the map data feature extraction stage: the map in the above text is divided into small blocks, and according to the location information in the user measurement report, a 01 mark of "whether it is a serving cell" is assigned to each small block in the map, which is called a cell mark. If the user measurement report of this user is taken from the location of this small block, it is marked as 1, otherwise it is marked as 0. Then, feature extraction is performed on each small block using the Patch Embedder module respectively, that is, two-dimensional convolution, multi-layer perceptron, and ReLU function. Then, only the small blocks with a cell mark of 1 are further feature-extracted through the Cell Embedder module, that is, adding a learnable CellEmbedding vector to obtain a map Patch vector, and the map Patch vector is the spatial feature vector that appeared in the above embodiment.

[0070] The present application also provides a communication metric prediction device based on a multimodal large model, including: a data acquisition module, configured to acquire multi-dimensional time-series data, geospatial data, and text description data corresponding to the multi-dimensional time-series data during the user's communication process; a time-series processing module, configured to perform time-series feature extraction processing on the multi-dimensional time-series data to obtain a time-series feature vector; a spatial processing module, configured to perform spatial feature extraction processing on the geospatial data to obtain a spatial feature vector; a text processing module, configured to encode the text description data based on the word embedding layer of a pre-trained large language model to obtain a semantic feature vector; a feature fusion module, configured to perform cross-modal attention fusion on the time-series feature vector, spatial feature vector, and semantic feature vector based on the text prototype of the large language model to generate a joint feature representation; and a prediction and inference module, configured to input the joint feature representation into the large language model, so that the large language model obtains a corresponding inference vector based on the joint feature representation and outputs communication metric prediction information according to the inference vector.

[0071] Among them, the multi-dimensional time-series data can be continuous monitoring data including multiple communication metrics such as signal level, delay, packet loss rate, and retransmission rate, and can be specifically stored and called using a time-series database for capturing the dynamic change rules of the communication environment. The geospatial data can be the geographical information of the base station locations and their coverage areas involved in the user's communication process, and can be specifically mapped to spatial coordinates using a geographic information system for characterizing the spatial distribution characteristics of the communication environment. The text description data can be explanatory text for the meaning of communication metrics and the definition of prediction tasks, and can be specifically constructed using natural language generation technology for providing task understanding at the semantic level for the model. The cross-modal attention fusion can use the text prototype of the large language model as the key-value vector, and map the time-series and spatial features to the text semantic space through attention weight calculation, and can be specifically implemented using a multi-head attention mechanism for eliminating the semantic gap between different modalities. The joint feature representation can be a multi-dimensional vector that fuses time-series, spatial, and semantic features, and can be specifically generated using feature splicing and normalization operations for providing a unified input representation for the large language model.

[0072] It can be understood that the data acquisition module synchronously collects multi-source data of the user's communication behavior in terms of time series, space, and text, and realizes structured storage through a multi-dimensional time series database and a geographic information system. The time series processing module uses a convolutional neural network based on the attention mechanism to extract local features from the time series data, divides the data into chunks along the time dimension, and extracts dynamic change patterns. The space processing module divides the geospatial data into grid cells, uses convolutional operations to extract local spatial features, and superimposes learnable embedding vectors on the service cell grids to enhance spatial correlation. The text processing module uses the word embedding layer of the large language model to convert the text description into a semantic vector, and establishes a mapping relationship between communication metrics and natural language descriptions. The feature fusion module uses the text prototype as a semantic bridge, aligns the time series and spatial features to the text semantic space through two cross-modal attention calculations respectively, and then concatenates them with the semantic features to form a joint representation. The prediction and inference module inputs the joint features into the large language model for sequence inference, and outputs prediction results including future network congestion, signal attenuation and other metrics.

[0073] It should be noted that, compared with the prior art, the traditional communication metric prediction device can only process a single type of time series data, uses an isolated model for single metric prediction, and cannot model the spatio-temporal correlation between multi-modal data. This solution realizes the collaborative modeling of time series dynamic features, spatial distribution features and semantic task understanding by constructing a multi-modal data processing module and a cross-modal fusion mechanism. The existing device lacks the ability to structurally process geospatial data and does not introduce text description information to assist the model in understanding the prediction task, resulting in difficulty in capturing the multi-factor coupling relationship in a complex network environment. This solution enables the large language model to comprehensively use spatio-temporal features and task semantics for reasoning and decision-making through grid-based spatial feature extraction and text prototype-guided attention fusion, improving the accuracy of network congestion prediction and signal attenuation warning, and providing a reliable basis for dynamic network resource scheduling.

[0074] The implementation manner of the communication metric prediction device based on the multi-modal large model is the same as the implementation process of the "communication metric prediction method based on the multi-modal large model" described above. For details, please refer to the previous description and will not be elaborated here.

[0075] Reference Figure 10 , Figure 10 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the communication metric prediction method based on the multi-modal large model in any one of the above embodiments. For example, it executes the method steps S110 to S160 described above in Figure 1 and executes the method steps S210 to S220 described above in Figure 2 , Figure 3The method steps S310 to S340 in Figure 4 The method steps S410 to S440 in Figure 5 The method steps S510 to S520 in Figure 6 The method steps S610 to S630 in Figure 7 The method steps S710 to S730 in

[0076] The electronic device 1000 according to an embodiment of the present application includes one or more processors 1010 and a memory 1020. Figure 10 Taking one processor 1010 and one memory 1020 as an example in

[0077] The processor 1010 and the memory 1020 can be connected through a bus or other means. Figure 10 Taking connection through a bus as an example in

[0078] The memory 1020, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory 1020 may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 1020 may optionally include a memory 1020 remotely provided relative to the processor 1010, and these remote memories can be connected to the electronic device 1000 through a network. At the same time, examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0079] In some embodiments, when the processor executes a computer program, it executes the communication metric prediction method based on a multimodal large model according to any one of the above embodiments at a preset interval.

[0080] Those skilled in the art can understand that Figure 10 The device structure shown in

[0081] In Figure 10 In the electronic device 1000 shown, the processor 1010 can be used to call the communication metric prediction method based on a multimodal large model stored in the memory 1020, so as to implement the communication metric prediction based on a multimodal large model.

[0082] Based on the hardware structure of the above-mentioned electronic device 1000, various embodiments of the communication metric prediction device based on the multimodal large model of the present application are proposed. At the same time, the non-transitory software programs and instructions required to implement the communication metric prediction method based on the multimodal large model of the above embodiments are stored in the memory and, when executed by the processor, execute the communication metric prediction method based on the multimodal large model of the above embodiments.

[0083] An embodiment of the present application also provides a computer-readable storage medium storing computer-executable instructions for executing the above-mentioned communication metric prediction method based on the multimodal large model, which enables the above one or more processors to execute the communication metric prediction method based on the multimodal large model of any of the above embodiments. For example, execute the method steps S110 to S160 described above in Figure 1 and execute the method steps S210 to S220 described above in Figure 2 and execute the method steps S310 to S340 described above in Figure 3 and execute the method steps S410 to S440 described above in Figure 4 and execute the method steps S510 to S520 described above in Figure 5 and execute the method steps S610 to S630 described above in Figure 6 and execute the method steps S710 to S730 described above in Figure 7 and execute the method steps S710 to S730 described above in

[0084] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network nodes. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0085] Those of ordinary skill in the art will appreciate that all or some of the steps and systems disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some or all of the physical components can be implemented as a microprocessor, such as a central processing unit, a digital signal processor, or software executed by a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer-readable storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer-readable storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The computer-readable storage medium includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0086] The above is a specific description of the preferred embodiments of the present application, but the present application is not limited to the above embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.

Claims

1. A communication index prediction method based on a multimodal large model, characterized in that: include: Acquire multi-dimensional time series data, geographic space data and text description data corresponding to the multi-dimensional time series data during user communication; Performing time series feature extraction processing on the multi-dimensional time series data to obtain a time series feature vector; Performing spatial feature extraction processing on the geographic spatial data to obtain a spatial feature vector; Encoding the text description data based on a word embedding layer of a pre-trained large language model to obtain a semantic feature vector; Based on the text prototype of the large language model, cross-modal attention fusion is performed on the temporal feature vector, the spatial feature vector and the semantic feature vector to generate a joint feature representation; The joint feature representation is input into the large language model, so that the large language model obtains a corresponding inference vector based on the joint feature representation, and outputs communication indicator prediction information according to the inference vector.

2. The communication index prediction method based on the multimodal large model according to claim 1 is characterized in that: The step of performing time series feature extraction processing on the multi-dimensional time series data to obtain a time series feature vector includes: Constructing a convolutional neural network based on an attention mechanism according to the feature dimensions of the multi-dimensional time series data; The multi-dimensional time series data is divided into a plurality of local time series blocks along the time series dimension, and local time series features are extracted from each of the local time series blocks respectively through the convolutional neural network, and a time series feature vector is obtained according to all the local time series features.

3. The communication index prediction method based on multimodal large model according to claim 1 is characterized in that: The performing of spatial feature extraction processing on the geographic spatial data to obtain a spatial feature vector includes: Dividing the geographic spatial data into a plurality of grid units, and marking the serving cell grids involved in the user communication process; Performing convolution feature extraction on all the grid cells to generate local spatial features; Superimposing a learnable embedding vector on the local spatial features corresponding to all the serving cell grids to obtain local enhanced features; A spatial feature vector is obtained according to the local spatial features of all the unmarked grid units and the local enhanced features of all the serving cell grids.

4. The communication index prediction method based on a multimodal large model according to any one of claims 1 to 3, characterized in that: The text prototype based on the large language model performs cross-modal attention fusion on the temporal feature vector, the spatial feature vector and the semantic feature vector to generate a joint feature representation, including: generating a first text prototype and a second text prototype based on the text dictionary data of the large language model; Taking the time series feature vector as the query vector and the first text prototype as the key value vector, the time series features of the text modality are generated through attention weight calculation; Using the spatial feature vector as the query vector and the second text prototype as the key value vector, the spatial feature of the text modality is generated; The temporal features of the text modality, the spatial features of the text modality and the semantic feature vector are concatenated to generate a multimodal joint feature representation.

5. The communication index prediction method based on multimodal large model according to claim 1 is characterized in that: The process of acquiring the geospatial data includes: Acquiring base station location data involved in the user communication process; A target map area is determined according to the base station location data and preset engineering parameters, and geographic space data corresponding to the target map area is obtained from a preset map library.

6. The communication index prediction method based on multimodal large model according to claim 1 is characterized in that: The process of acquiring the text description data corresponding to the multi-dimensional time series data includes: Constructing a data set introduction text including the meanings of multiple communication indicators in the multi-dimensional time series data; Construct a task description text including the definition of the length of the historical sequence and the length of the predicted sequence in the communication indicator prediction; The data set introduction text and the task description text are concatenated to obtain text description data.

7. The communication index prediction method based on multimodal large model according to claim 1 is characterized in that: The training process of the large language model includes: The multi-dimensional time series data in the training data is divided into multiple samples according to the time sliding window, and each sample contains an input sequence and an output sequence; Calculating the predicted output of the large language model to be trained for the input sequence, and calculating the model loss according to the predicted output and the output sequence; The input-output mapping layer parameters of the large language model are updated based on the model loss to obtain a trained large language model.

8. A communication index prediction device based on a multimodal large model, characterized in that: include: A data acquisition module, used to acquire multi-dimensional time series data, geographic space data and text description data corresponding to the multi-dimensional time series data during user communication; A time series processing module, used for performing time series feature extraction processing on the multi-dimensional time series data to obtain a time series feature vector; A spatial processing module, used for performing spatial feature extraction processing on the geographic spatial data to obtain a spatial feature vector; A text processing module, used for encoding the text description data based on a word embedding layer of a pre-trained large language model to obtain a semantic feature vector; A feature fusion module, configured to perform cross-modal attention fusion on the temporal feature vector, the spatial feature vector and the semantic feature vector based on the text prototype of the large language model to generate a joint feature representation; A prediction and reasoning module is used to input the joint feature representation into the large language model so that the large language model obtains a corresponding reasoning vector based on the joint feature representation, and outputs communication indicator prediction information according to the reasoning vector.

9. An electronic device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When at least one of the programs is executed by at least one of the processors, the communication indicator prediction method based on the multimodal large model as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: Computer executable instructions are stored, and the computer executable instructions are used to execute the communication indicator prediction method based on the multimodal large model as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Large language model prediction method and device based on space-time attention mechanism

    CN117786061A

  • Network traffic prediction method based on space-time diagram multi-attention mechanism

    CN117880871A

  • Multivariable time sequence prediction method, electronic equipment and storage medium

    CN118503711A

  • Communication network performance prediction method and system based on large model

    CN119172259A

  • Telephone traffic demand prediction method and system

    CN119497093A

Cited By

  • Shape element-based large language model interpretability time sequence classification method and system

    CN120561664A

  • Invoice identification method, device and equipment based on multi-modal large model

    CN120808377A

  • Behavior prediction method and device based on multi-dimensional user data, equipment and medium

    CN120849870A

  • Multi-modal feature fusion actuator stroke calibration system

    CN121140872A

  • Power load scheduling method and device based on large language model

    CN121503946A