An environment and intent perception communication link construction method based on a pre-trained large language model
Patent Information
- Application Number
- CN202611077096.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-09-11
AI Technical Summary
[0006]本发明目的在于针对现有技术中难以联合利用物理层环境信息与用户意图、难以实现全局最优决策的问题,提出一种基于预训练大语言模型的环境与意图感知通信链路构建方法
本发明不是简单地将大语言模型应用于通信链路配置任务,而是针对无线物理层链路构建问题,建立了信道状态信息、用户文本意图与物理层动作空间之间的结构化映射关系。通过将连续 CSI 表征、离散用户意图和多个物理层配置动作统一到同一决策框架中,本发明使链路构建过程能够同时感知无线传播环境和用户业务需求,从而避免传统方案仅依据信道质量或固定门限进行单一链路自适应的局限。
Smart Images

Figure CN122734489A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of wireless communication and intelligent information processing technology, and in particular to a method for constructing environment and intent-aware communication links based on a pre-trained large language model. Background Technology
[0002] To advance towards the intelligent evolution of sixth-generation (6G) mobile communication systems, future wireless networks not only need to possess high-speed, low-latency, and high-reliability communication capabilities, but also need to understand service requirements, perceive the communication environment, and make adaptive decisions. Traditional physical layer systems typically employ a modular, layered design paradigm, designing and independently optimizing functions such as channel coding, modulation, precoding, power control, channel estimation, and equalization. While this approach can achieve good results on specific modules, it often fails to achieve globally optimal end-to-end link performance because it ignores the complex nonlinear coupling relationships between modules.
[0003] Most existing link adaptive methods rely primarily on channel state information for strategy selection, such as adjusting modulation and coding schemes, selecting precoding matrices, or configuring transmission power based on CSI. These methods typically only focus on the physical layer environment and struggle to simultaneously characterize users' differentiated needs in terms of throughput, reliability, or energy consumption, thus failing to meet the personalized service goals of "on-demand customization" in future intelligent communication systems.
[0004] In recent years, large language models have demonstrated superior capabilities in contextual understanding, sequence reasoning, and cross-modal semantic fusion, and have begun to be explored for application in wireless communication tasks such as channel prediction, beam prediction, and resource allocation. However, most existing methods are limited to single-task or single-modal application scenarios, and have not fully leveraged the advantages of large language models in multimodal joint perception, unified reasoning, and global decision-making, making it difficult to achieve end-to-end link policy generation that simultaneously considers the channel environment and user intent.
[0005] To address the aforementioned issues, it is necessary to propose a novel approach that can jointly sense channel state information and user textual intent, and generate personalized, executable communication link strategies based on a pre-trained large language model, thereby improving the robustness, flexibility, and global optimization capabilities of communication systems in complex environments. Summary of the Invention
[0006] The purpose of this invention is to address the problem in existing technologies that it is difficult to jointly utilize physical layer environmental information and user intent, and difficult to achieve globally optimal decision-making. The invention proposes a method for constructing an environment and intent-aware communication link based on a pre-trained large language model.
[0007] The objective of this invention is achieved through the following technical solution: a method for constructing an environment- and intent-aware communication link based on a pre-trained large language model, the method comprising: Channel state information and user textual intent are obtained as multimodal inputs, and feature extraction is performed separately to obtain CSI features and textual semantic features. Text semantic features are compressed and filtered using linear projection, self-attention, and cross-modal attention mechanisms. The cross-modal attention mechanism uses CSI features as query vectors and text semantic features as keys and values. The cross-modal aligned text semantic features are concatenated with CSI features and input into a pre-trained large language model for joint inference, outputting a link configuration strategy for multi-physical layer modules. Based on the link configuration strategy, link construction is completed in a communication simulation environment. A reward function is constructed according to the bit error rate, transmission rate and power consumption indicators to optimize the pre-trained large language model, so as to obtain a personalized communication link construction result that matches the user intent and channel environment.
[0008] Furthermore, the link configuration includes one or more of the following: channel coding, rate selection, modulation, precoding, power control, channel estimation, and equalization.
[0009] Furthermore, the feature extraction process for the channel state information includes: mapping the original physical layer channel state to a latent representation space consistent with the semantic space of the large language model, and the processing can be expressed as follows: ; in, , Indicates the length of the CSI sequence. Representing feature dimension, This represents the encoded CSI features. The embedding length is consistent with the hidden dimension of the pre-trained large language model.
[0010] Furthermore, the feature extraction of the user's text intent includes: first, tokenizing the text using a tokenizer, and then mapping it to the semantic space of a large language model through an embedding layer to obtain the original text embedding representation. ; in, , This indicates the length of the text sequence; this text feature is used to characterize the user's high-level goals and preferences in communication services.
[0011] Furthermore, the compression and filtering of text semantic features using linear projection, self-attention, and cross-modal attention mechanisms includes: firstly, compressing the text embedding using linear projection to obtain: ; in, ,and ; Then, a self-attention mechanism is applied to the compressed text representation to establish semantic dependencies within the text and extract global semantic context information, resulting in: ; After completing text semantic compression, a cross-modal attention mechanism is employed to align CSI features with text semantics, using CSI features... As a query matrix Represented in compressed text as a bond matrix Sum matrix The cross-modal attention process can be represented as: ; in, ; in, , and For learnable projection matrices, This is the scaling factor.
[0012] Furthermore, the large language model includes: CSI Feature Representation The sequences are concatenated along the sequence dimension and fed into the pre-trained large language model backbone network to obtain the joint inference result: ; in, This represents the hidden state of the output of the last layer of the large language model; Multiple actuator networks are set up after the pre-trained large language model backbone network. Each actuator network corresponds to a physical layer functional module and is derived from the joint hidden representation. The optimal configuration strategy for the corresponding module is generated in the middle; The output of an actuator network can be represented as: ; The outputs of all actuator networks together constitute a complete link configuration strategy: .
[0013] Furthermore, the pre-training of the large language model includes: A heuristic search algorithm is used to explore in a communication simulation environment and collect high-quality state-action sample pairs. and build an experience pool , where state Composed of textual intent and CSI features, actions A strategy is configured for the link, and then, supervised learning is used to minimize the difference between the model's output strategy and the strategy of the experience pool samples, enabling the model to achieve near-optimal initialization. Its supervised fine-tuning loss function is expressed as: in, The parameter is The strategy network.
[0014] Furthermore, optimizing the pre-trained large language model includes: the model based on the current state. According to strategy Sampling action The generated link configuration strategy was applied to the communication simulation environment to obtain corresponding performance feedback. A reward function composed of bit error rate, transmission rate, and power consumption was constructed, and its form is as follows: in, , and These are non-negative weighting coefficients, and the sum of the three is 1. , and These represent the normalized reward terms corresponding to bit error rate, data rate, and power consumption, respectively. By maximizing the expected reward, the model can adaptively generate better physical layer link configuration strategies under different communication environments and user intentions.
[0015] According to another aspect of the specification, an environment and intent-aware communication link construction device based on a pre-trained large language model is also provided, including a memory and one or more processors. The memory stores executable code, and when the processor executes the executable code, it implements the aforementioned method for constructing an environment and intent-aware communication link based on a pre-trained large language model.
[0016] According to another aspect of the specification, a computer-readable storage medium is also provided, on which a program is stored, which, when executed by a processor, implements the aforementioned method for constructing an environment and intent-aware communication link based on a pre-trained large language model.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention does not simply apply a large language model to communication link configuration tasks. Instead, it addresses the wireless physical layer link construction problem by establishing a structured mapping relationship between channel state information, user textual intent, and the physical layer action space. By unifying continuous CSI representation, discrete user intent, and multiple physical layer configuration actions into a single decision framework, this invention enables the link construction process to simultaneously perceive the wireless propagation environment and user service requirements, thereby avoiding the limitations of traditional solutions that rely solely on channel quality or fixed thresholds for single-link adaptation.
[0018] This invention designs a communication task-oriented connection module in the multimodal data processing stage. Instead of directly inputting text and CSI into a large language model, this module first performs task-related semantic compression and filtering on the user's text intent, retaining semantic information relevant to link optimization objectives such as throughput, reliability, bit error rate, and power consumption. Simultaneously, it uses CSI features to guide the text semantics, enabling the model to dynamically focus on different intent information based on the current channel environment. Thus, this invention achieves cross-modal alignment between environmental state and user intent, reduces redundant semantic interference unrelated to physical layer decisions, and improves the effectiveness of multimodal fusion and the stability of link strategy generation.
[0019] This invention models physical layer link construction as a multi-module joint action generation problem, rather than a traditional module-by-module independent optimization problem. The model outputs actions not as natural language results, but as structured action sequences corresponding one-to-one with physical layer modules such as modulation scheme, coding rate, power compensation, channel coding, channel estimation, equalization, and precoding. By explicitly preserving the dependencies between different physical layer modules in the action space, this invention can coordinate the combined effects of multiple link configuration parameters on bit error rate, transmission rate, power consumption, and computational complexity, thereby achieving superior end-to-end link performance.
[0020] This invention employs a two-stage training mechanism combining heuristic search sample-driven supervised fine-tuning and reinforcement learning. In the first stage, the invention utilizes a heuristic search algorithm oriented towards physical layer link simulation to generate high-quality expert samples, enabling the model to obtain near-optimal policy initialization in the early stages of training, rather than relying on random exploration or manual rule initialization. In the second stage, the invention further constructs a multi-objective reward function related to user intent, weighting and optimizing indicators such as bit error rate, transmission rate, power consumption, and complexity according to different service requirements, and further refining the initial policy through reinforcement learning. This training mechanism allows the model to both inherit effective link configuration experience obtained from heuristic search and adaptively optimize policy quality under different intent constraints, thereby improving training stability and generalization ability.
[0021] Experimental results show that, under different user intent scenarios such as high throughput, high reliability, and energy saving, the physical layer configuration strategy generated by this invention exhibits better overall performance in terms of bit error rate, transmission rate, power consumption, reward value, and inference latency. Compared with random strategies, traditional AMC rules, greedy search, beam search, and ordinary Transformer methods that have not been adapted to the communication domain, this invention can generate more stable and targeted link configuration results under different channel conditions and different service intents, demonstrating its practical application potential for AI-native 6G communication systems. Attached Figure Description
[0022] Figure 1 This is a general block diagram of the environment and intent-aware communication link construction method proposed in this invention; Figure 2 This is a schematic diagram of the two-stage training method used in this invention; Figure 3 This invention provides a performance comparison of high throughput strategy selection methods under different signal-to-noise ratios in its embodiments. Figure 4 This invention provides a performance comparison of high reliability strategy selection methods under different signal-to-noise ratios in its embodiments. Figure 5 This invention provides a performance comparison of energy-saving strategy selection methods under different signal-to-noise ratios in its embodiments. Figure 6 This invention provides a carrier device for the method of constructing environment and intent-aware communication links based on pre-trained large language models. Detailed Implementation
[0023] The present invention will be further described below with reference to the accompanying drawings: The present invention proposes a method for constructing environment and intent-aware communication links based on pre-trained large language models, such as... Figure 1 As shown, the method mainly includes a CSI encoding module, a text embedding module, a connection module, a pre-trained large language model backbone network, an actuator network, and a two-stage training optimization module. This method uses channel state information and user textual intent as joint inputs, and through cross-modal alignment and large language model inference, achieves adaptive construction of communication links for multiple modules at the physical layer.
[0024] In this implementation, the Channel State Information (CSI) and the user-provided textual intent description for the current communication scenario are first acquired. The CSI characterizes the current physical channel environment characteristics, and the textual intent characterizes the user's preferred performance requirements for the communication link, such as high throughput, high reliability, or low power consumption. For the physical layer link construction problem, a set of link configuration strategies is defined as follows: in, Indicates the first Configuration strategies for each physical layer module This indicates the number of modules participating in the joint decision-making process. These modules include, but are not limited to, channel coding, rate selection, modulation, precoding, power control, channel estimation, and equalization. The end-to-end transmission process of the system can be abstractly represented as: in, Indicates sending data. This indicates that the receiving end has resumed data. This indicates channel state information.
[0025] In this invention, the input CSI sequence is first... The input is fed into the CSI encoding module for feature extraction. The CSI encoding module includes a Cross Spatial Self-Attention (CSSA) module, which is used to extract key features related to changes in the wireless environment and link configuration decisions from the original channel state.
[0026] Specifically, for the input complex CSI, its real and imaginary parts are first separated and combined into a real-valued tensor, and then normalized to reduce the impact of amplitude differences between different samples on model training. Subsequently, the normalized CSI is rearranged according to the subcarrier, time slice, and transmit / receive antenna dimensions, and divided into several non-overlapping CSI patches along the time dimension, thus obtaining a serialized CSI representation. Through this processing method, the original high-dimensional CSI is converted into a local time-frequency feature sequence suitable for attention calculation.
[0027] Building upon this foundation, the CSSA module performs cross-spatial self-attention modeling on the serialized CSI representation. Specifically, CSSA calculates the correlations between different CSI patches, subcarriers, and channel features through multi-head self-attention, enabling the model to capture local fading features and inter-channel coupling relationships in the time-frequency dimension. Furthermore, it adaptively recalibrates the importance of different channel features by obtaining channel-level weights through convolution, pooling, fully connected layers, and sigmoid activation. Thus, CSSA enhances the focus on key propagation features and suppresses redundant features that are weakly relevant to the current link decision.
[0028] Therefore, the processing procedure of the CSI encoding module can be represented as follows: in, , Indicates the length of the CSI sequence. Representing feature dimension, This indicates a CSI encoder that includes CSSA. This represents the latent CSI characterization after cross-space self-attention modulation.
[0029] At the same time, the user's text intent The text is input into the text processing module. First, the text is tokenized by a tokenizer, then mapped to the semantic space of a large language model through an embedding layer to obtain the original text embedding representation. in, , This indicates the length of the text sequence. This text feature is used to characterize the user's high-level goals and preferences in communication services.
[0030] Because the original text sequence is quite long, directly inputting it into a pre-trained large language model would result in high computational complexity and inference overhead. Therefore, this invention sets up a connection module before the text enters the backbone of the large language model to perform semantic compression and task-relevant filtering of the text information. The connection module consists of three parts: linear projection, self-attention, and cross-modal attention. First, the text embedding is compressed through linear projection, resulting in: in, ,and This step is used to reduce the length of the text sequence, thereby reducing the cost of subsequent inference.
[0031] Then, a self-attention mechanism is applied to the compressed text representation to establish semantic dependencies within the text and extract global semantic context information, resulting in: This step allows key information in the text to be preserved in a more compact representation, thus providing a more effective semantic carrier for subsequent cross-modal alignment.
[0032] After completing text semantic compression, this invention further employs a cross-modal attention mechanism to align CSI features with text semantics. Specifically, using CSI features... As a query matrix Represented in compressed text as a bond matrix Sum matrix The cross-modal attention process can be represented as: in, in, , and For learnable projection matrices, The scaling factor is used. Through this cross-modal alignment process, the current channel environment represented by CSI can actively filter the semantic information in the text that is most relevant to the current scene, thereby obtaining a task semantic representation under environmental constraints. .
[0033] After obtaining the aligned semantic representation of the text, CSI Feature Representation The sequences are concatenated along the sequence dimension and fed into the pre-trained large language model backbone network to obtain the joint inference result: in, This represents the hidden state of the final layer output of the large language model. Through this step, the large language model can combine environmental information and user intent information to perform unified modeling and global reasoning of the entire communication link.
[0034] To achieve joint decision-making for multiple physical layer functional modules, this invention sets up multiple actuator networks after pre-training a large language model. Each actuator network corresponds to a physical layer functional module and is derived from the joint hidden representation. The optimal configuration strategy for the corresponding module is generated in the middle. The output of an actuator network can be represented as: The outputs of all actuator networks together constitute a complete link configuration strategy: Different actuators correspond to physical layer modules such as channel coding, rate selection, modulation scheme, precoding, power control, channel estimation, and equalization. Finally, the end-to-end physical layer link is constructed based on the generated joint strategy.
[0035] like Figure 2 As shown, in terms of training, this invention employs a two-stage training mechanism to optimize the construction model of the environment and intent-aware communication link. The first stage is the supervised fine-tuning stage. Firstly, a heuristic search algorithm is used to explore the communication simulation environment and collect high-quality state-action sample pairs. and build an experience pool Among them, state Composed of textual intent and CSI features, actions A strategy is configured for the link. Subsequently, supervised learning is used to minimize the difference between the model's output strategy and the strategies in the experience pool, enabling the model to achieve near-optimal initialization. The supervised fine-tuning loss function is expressed as: in, The parameter is The policy network is trained in this stage. Through this stage, the model learns high-quality policy samples generated by heuristics, thus providing stable initial parameters for subsequent reinforcement learning optimization.
[0036] The second stage is the reinforcement learning fine-tuning stage. In this stage, the model adjusts its settings according to the current state. According to strategy Sampling action The generated link configuration strategy is then applied to the communication simulation environment to obtain corresponding performance feedback. To comprehensively consider the performance of the communication link across multiple dimensions, this invention constructs a reward function composed of bit error rate, transmission rate, and power consumption, in the form of: in, , and These are non-negative weighting coefficients, and the sum of the three is 1. , and These represent the normalized reward terms corresponding to bit error rate, data rate, and power consumption, respectively. By maximizing the expected reward, the model can adaptively generate better physical layer link configuration strategies under different communication environments and user intentions.
[0037] In this embodiment, the communication environment can be constructed using a physical layer simulation system conforming to 3GPP standards. This system includes multiple functional modules such as channel coding, rate selection, modulation, precoding, power control, channel estimation, and equalization. By changing channel conditions, signal-to-noise ratio, and user intent type, various link construction tasks can be generated to verify the generalization and adaptive capabilities of this invention. This embodiment selects three types of user intent—high throughput, high reliability, and energy saving—as representative tasks, and statistically analyzes the performance of different methods on three indicators: bit error rate, transmission rate, and power compensation, to verify whether this invention can generate differentiated physical layer configuration strategies according to different service requirements.
[0038] To verify the performance advantages of this invention compared to existing link planning methods, it is compared with randomized policies, greedy search, bundle search, and policy generation methods based on ordinary Transformers. Randomized policies are used to characterize the link configuration effect without intelligent decision-making; greedy search and bundle search are used to characterize the performance of traditional heuristic search methods within a certain search space; and the ordinary Transformer method is used to characterize the baseline of deep models without intent-environment alignment design in the communication domain. Experimental results are as follows: Figures 3 to 5 As shown.
[0039] like Figure 3 As shown, in high-throughput scenarios, user needs primarily manifest as maximizing transmission rate within an acceptable bit error rate range. Figure 3 (a) It can be seen that as the signal-to-noise ratio (SNR) increases, the bit error rate of the present invention generally decreases, and approaches zero under higher SNR conditions, indicating that the present invention can maintain basic reliability while pursuing high transmission rates. Figure 3 (b) It can be seen that the present invention can achieve high transmission rates under various signal-to-noise ratio conditions, especially in the medium-to-high signal-to-noise ratio region, where its rate is significantly higher than that of random strategies, greedy search, and ordinary Transformer methods, and also superior to or close to the beam search method. This indicates that the present invention can effectively identify high-throughput intentions and tends to select joint link configurations with higher-order modulation, higher coding rates, or those more conducive to rate improvement. Figure 3 (c) It can be further seen that the present invention appropriately increases power compensation to support high-speed transmission under low signal-to-noise ratio conditions, while the required power compensation decreases significantly after the channel conditions improve. This indicates that the present invention does not simply increase the rate by continuously increasing power, but can dynamically adjust the combination relationship between multiple modules such as modulation, coding, and power according to the channel conditions. Therefore, in high-throughput scenarios, the present invention demonstrates an adaptive link construction capability that prioritizes rate while taking into account reliability and power consumption.
[0040] like Figure 4 As shown, in high-reliability scenarios, user needs primarily manifest as reducing the bit error rate and improving link transmission stability. Figure 4 (a) It can be seen that in the low signal-to-noise ratio (SNR) region, the random strategy and the ordinary Transformer method have high bit error rates, indicating that they are difficult to generate reliable link configurations under adverse channel conditions. In contrast, the present invention achieves a low bit error rate even under low SNR conditions, and the bit error rate decreases rapidly and approaches zero as the SNR increases. This result shows that the present invention can proactively select a more conservative and robust physical layer configuration based on high reliability intentions, such as lower-order modulation, lower coding rate, or channel processing methods that are more conducive to noise resistance. Figure 4(b) It can be seen that in high-reliability scenarios, the transmission rate of this invention is not always the highest; under certain signal-to-noise ratio conditions, greedy search or beam search can achieve higher rates; however, combined with Figure 4 (a) It is evident that the speed improvement of these methods does not necessarily correspond to better reliability. In contrast, this invention is more in line with the actual needs of high-reliability services, namely, prioritizing bit error rate performance rather than simply pursuing maximum speed. Figure 4 (c) It can be seen that the present invention uses higher power compensation to reduce the risk of bit error rate under low signal-to-noise ratio conditions, and gradually reduces power compensation after the channel conditions improve, thereby achieving a trade-off between reliability and resource consumption. This result shows that the present invention can change the optimization target weight according to the user's intention, forming a link construction strategy of "bit error rate priority, appropriate sacrifice of rate and power consumption" in high reliability scenarios.
[0041] like Figure 5 As shown, in energy-saving scenarios, user needs mainly manifest as reducing power compensation and energy consumption, while maintaining an acceptable bit error rate and transmission rate as much as possible. Figure 5 (c) It can be seen that the present invention maintains a low power compensation level under different signal-to-noise ratio conditions, especially in the medium-to-high signal-to-noise ratio region, where the required power compensation is close to zero, significantly lower than that of random strategies and greedy search methods. This indicates that the present invention can identify energy-saving intentions and avoid unnecessary high-power link configurations. Figure 5 (a) It can be seen that although this invention prioritizes reducing power consumption, its bit error rate still decreases rapidly with increasing signal-to-noise ratio (SNR). It maintains a low bit error rate in the medium-to-high SNR region, indicating that this invention does not simply reduce power output, but rather coordinates the modulation, coding, estimation, equalization, and precoding modules under energy-saving constraints to maintain basic link reliability. Figure 5 (b) It can be seen that the present invention can maintain a relatively stable transmission rate in energy-saving scenarios, and the rate gradually increases with the improvement of the signal-to-noise ratio. In contrast, although the ordinary Transformer method has lower power compensation, it has a higher bit error rate and limited rate improvement under low signal-to-noise ratio conditions, indicating that it lacks effective modeling of the relationship between energy-saving goals and link quality. Therefore, in energy-saving scenarios, the present invention demonstrates a multi-objective balancing capability of "power consumption first, while also considering rate and reliability".
[0042] comprehensive Figures 3 to 5It can be seen that this invention exhibits significantly different strategy preferences under different user intentions: in high-throughput scenarios, this invention tends to increase the transmission rate; in high-reliability scenarios, this invention prioritizes reducing the bit error rate; and in energy-saving scenarios, this invention prioritizes reducing power compensation. This phenomenon indicates that this invention does not learn fixed link configuration rules, but rather can generate personalized combinations of physical layer actions based on CSI, signal-to-noise ratio, and user textual intent. Compared with random strategies, this invention has a clear goal orientation; compared with greedy search and bundle search, this invention can directly generate a complete link configuration through a single model inference, avoiding the computational overhead of repeated searches in the online phase; compared with ordinary Transformer methods, this invention, through multimodal alignment oriented towards the communication domain and a two-stage training mechanism, can more effectively extract task-related semantics and generate link strategies that conform to different communication intentions.
[0043] Corresponding to the aforementioned embodiment of the method for constructing environment and intent-aware communication links based on a pre-trained large language model, the present invention also provides an embodiment of an apparatus for constructing environment and intent-aware communication links based on a pre-trained large language model.
[0044] See Figure 6 The present invention provides an environment and intent-aware communication link construction device based on a pre-trained large language model, comprising a memory and one or more processors. The memory stores executable code, and when the processor executes the executable code, it is used to implement an environment and intent-aware communication link construction method based on a pre-trained large language model as described in the above embodiments.
[0045] The embodiment of the environment and intent-aware communication link construction device based on a pre-trained large language model provided by this invention can be applied to any device with data processing capabilities, such as a computer. The device embodiment can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data-processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution. From a hardware perspective, such as... Figure 6 The diagram shown is a hardware structure diagram of any data processing-capable device, including the environment and intent-aware communication link construction device based on a pre-trained large language model provided by the present invention. (Except for...) Figure 6 In addition to the processor, memory, network interface, and non-volatile memory shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.
[0046] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0047] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0048] This invention also provides a computer-readable storage medium storing a program thereon, which, when executed by a processor, implements a method for constructing an environment and intent-aware communication link based on a pre-trained large language model as described in the above embodiments.
[0049] The computer-readable storage medium can be an internal storage unit of any data processing device described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device of any data processing device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of any data processing device. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.
[0050] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method for constructing an environment and intent-aware communication link based on a pre-trained large language model.
[0051] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.
[0052] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. This application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A pre-trained large language model-based environment and intent perception communication link construction method, characterized in that, The method includes: Channel state information and user textual intent are obtained as multimodal inputs, and feature extraction is performed separately to obtain CSI features and textual semantic features. Text semantic features are compressed and filtered using linear projection, self-attention, and cross-modal attention mechanisms. The cross-modal attention mechanism uses CSI features as query vectors and text semantic features as keys and values. The cross-modal aligned text semantic features are concatenated with CSI features and input into a pre-trained large language model for joint inference, outputting a link configuration strategy for multi-physical layer modules. Based on the link configuration strategy, link construction is completed in a communication simulation environment. A reward function is constructed according to the bit error rate, transmission rate and power consumption indicators to optimize the pre-trained large language model, so as to obtain a personalized communication link construction result that matches the user intent and channel environment.
2. The environment and intent perception communication link construction method based on a pre-trained large language model according to claim 1, characterized in that, The link configuration includes one or more of the following: channel coding, rate selection, modulation, precoding, power control, channel estimation, and equalization.
3. The method for constructing an environment- and intent-aware communication link based on a pre-trained large language model according to claim 1, characterized in that, The feature extraction process for the channel state information includes: mapping the original physical layer channel state to a latent representation space consistent with the semantic space of the large language model, which can be represented as follows: ; in, , Indicates the length of the CSI sequence. Representing feature dimension, This represents the encoded CSI features. The embedding length is consistent with the hidden dimension of the pre-trained large language model.
4. The method for constructing an environment and intent-aware communication link based on a pre-trained large language model according to claim 1, characterized in that, The feature extraction of the user's text intent includes: first, tokenizing the text using a tokenizer, and then mapping it to the semantic space of a large language model through an embedding layer to obtain the original text embedding representation. ; in, , This indicates the length of the text sequence; this text feature is used to characterize the user's high-level goals and preferences in communication services.
5. The method for constructing an environment and intent-aware communication link based on a pre-trained large language model according to claim 1, characterized in that, The compression and filtering of text semantic features using linear projection, self-attention, and cross-modal attention mechanisms includes: firstly, compressing the text embedding through linear projection to obtain: ; in, ,and ; Then, a self-attention mechanism is applied to the compressed text representation to establish semantic dependencies within the text and extract global semantic context information, resulting in: ; After completing text semantic compression, a cross-modal attention mechanism is employed to align CSI features with text semantics, using CSI features... As a query matrix Represented in compressed text as a bond matrix Sum matrix The cross-modal attention process can be represented as: ; in, ; in, , and For learnable projection matrices, This is the scaling factor.
6. The method for constructing an environment and intent-aware communication link based on a pre-trained large language model according to claim 1, characterized in that, The large language model includes: CSI Feature Representation The sequences are concatenated along the sequence dimension and fed into the pre-trained large language model backbone network to obtain the joint inference result: ; in, This represents the hidden state of the output of the last layer of the large language model; Multiple actuator networks are set up after the pre-trained large language model backbone network. Each actuator network corresponds to a physical layer functional module and is derived from the joint hidden representation. The optimal configuration strategy for the corresponding module is generated in the middle; The output of an actuator network can be represented as: ; The outputs of all actuator networks together constitute a complete link configuration strategy: 。 7. The method for constructing an environment and intent-aware communication link based on a pre-trained large language model according to claim 1, characterized in that, The pre-training of the large language model includes: A heuristic search algorithm is used to explore in a communication simulation environment and collect high-quality state-action sample pairs. and build an experience pool , where state Composed of textual intent and CSI features, actions A strategy is configured for the link, and then, supervised learning is used to minimize the difference between the model's output strategy and the strategy of the experience pool samples, enabling the model to achieve near-optimal initialization. Its supervised fine-tuning loss function is expressed as: in, The parameter is The strategy network.
8. The method for constructing an environment and intent-aware communication link based on a pre-trained large language model according to claim 7, characterized in that, Optimizing the pre-trained large language model includes: the model based on the current state According to strategy Sampling action The generated link configuration strategy was applied to the communication simulation environment to obtain corresponding performance feedback. A reward function composed of bit error rate, transmission rate, and power consumption was constructed, and its form is as follows: in, , and These are non-negative weighting coefficients, and the sum of the three is 1. , and These represent the normalized reward terms corresponding to bit error rate, data rate, and power consumption, respectively. By maximizing the expected reward, the model can adaptively generate better physical layer link configuration strategies under different communication environments and user intentions.
9. An apparatus for constructing an environment- and intent-aware communication link based on a pre-trained large language model, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that, When the processor executes the executable code, it implements the method for constructing an environment and intent-aware communication link based on a pre-trained large language model as described in any one of claims 1-8.
10. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the method for constructing an environment and intent-aware communication link based on a pre-trained large language model as described in any one of claims 1-8.