Embodied agent anti-interference communication method based on large model assistance
Patent Information
- Application Number
- CN202610870075.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-08-18
AI Technical Summary
现有方案仅利用大模型推断环境特征,以辅助优化传输多模态感知数据的上行通信策略,未对控制器传输控制指令的下行通信进行优化,导致控制信息延迟反馈或丢失,抓取任务成功率受限
[0009] The embodied agent anti-interference communication method based on large model assistance proposed in this application has the following advantages: based on the environmental features of the embodied agent extracted by the large model, combined with channel gain, received interference intensity and motion state, reinforcement learning is used to simultaneously optimize uplink transmission power, channel and downlink transmission power, which can improve the success rate of the embodied agent in grasping tasks under malicious interference, while reducing communication latency and energy consumption.
Smart Images

Figure CN122602288A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of wireless communication, embodied intelligence, and network security, and specifically relates to a large model-assisted anti-interference communication method for embodied intelligent agents. Background Technology
[0002] In embodied intelligent communication scenarios, multimodal large language models can infer and generate movement trajectories or control commands such as gripper opening and closing based on sensory data such as images and videos, as well as task description text, supporting applications such as parts sorting, human-machine collaborative assembly, and multi-target grasping and placement. However, under malicious interference attacks, the reliability of the transmission of sensory data and control commands decreases, leading to decision-making errors, task failures, or equipment malfunctions, and even threatening personal safety.
[0003] Currently, intelligent wireless anti-jamming communication can optimize transmission strategies based on information such as channel state and interference characteristics, thereby improving the signal-to-interference-plus-noise ratio (SINR), reducing communication energy consumption, and decreasing transmission latency. Existing research has utilized deep reinforcement learning to optimize transmission channels or optimized frequency hopping strategies based on spectrum waterfall plots to combat various types of interference. Terminals with sufficient computing power can also employ methods such as large language models and reinforcement learning to extract environmental features, assisting in optimizing transmit power and channels, and improving communication performance and inference accuracy.
[0004] However, existing technologies still face the following shortcomings: (1) Only uplink communication is optimized, while downlink control decision-making is ignored. Existing solutions only use large models to infer environmental features to assist in optimizing the uplink communication strategy for transmitting multimodal sensing data, without optimizing the downlink communication for the controller to transmit control commands. This results in delayed or lost control information feedback, limiting the success rate of the capture task.
[0005] (2) Lack of a joint uplink and downlink optimization framework. Existing methods do not incorporate uplink sensing data transmission and downlink control command transmission into a unified optimization framework. They cannot coordinately adjust uplink and downlink communication strategies based on information such as channel status, received interference intensity, and motion status, making it difficult to ensure the reliability of end-to-end task execution under malicious interference.
[0006] Therefore, this invention proposes a large-model-assisted embodied intelligent agent anti-interference communication method to address the aforementioned shortcomings. Summary of the Invention
[0007] This invention aims to at least partially solve one of the technical problems in the aforementioned technologies. To this end, one objective of this invention is to propose a large-model-assisted embodied agent anti-interference communication method. In each time slot, the controller obtains channel gain through pilot estimation, measures received interference intensity, receives motion state, task success rate, latency, and energy consumption feedback from the embodied agent, and uses a large model to extract environmental features such as surrounding object types and location coordinates from depth images. This information is constructed into a system state, input into a benefit network and a risk network to evaluate the expected benefits and risks of each strategy, and accordingly jointly optimizes uplink transmit power, channel, and downlink transmit power. Based on the capture task request and depth image, the large model generates control commands and transmits them downlink to the embodied agent. Uplink and downlink coordinated anti-interference communication is achieved by iteratively updating network parameters until convergence.
[0008] To achieve the above objectives, a first aspect of the present invention proposes a large model-assisted embodied agent anti-interference communication method, comprising the following steps: obtaining the state vector of the current time slot, wherein the state vector includes channel gain, received interference intensity, motion state, task feedback information from the previous time slot, and environmental characteristics; inputting the state vector of the current time slot into a benefit network and a risk network to obtain the expected benefits and expected risks of each optional strategy; selecting the optimal communication strategy for the current time slot from the set of optional strategies according to the expected benefits and expected risks of each optional strategy, wherein the optimal communication strategy includes optional uplink transmit power, optional channel, and downlink transmit power; obtaining a grasping task inference request and a depth image, and generating control commands using a large model based on the grasping task inference request and the depth image; The control command, optional uplink transmission power, and optional channel are transmitted to the embodied agent based on the downlink transmission power, so that the embodied agent can perform the grasping action according to the control command and provide feedback on the task feedback information of the current time slot according to the optional uplink transmission power and optional channel; receive the task feedback information of the current time slot and obtain the benefit value and risk value based on the task feedback information of the current time slot; store the complete experience tuple of the current time slot into the experience pool, and randomly sample a batch of experience from the experience pool to update the weight parameters of the benefit network and the risk network respectively. The complete experience tuple includes the benefit value of the current time slot, the risk value of the current time slot, the state vector of the current time slot, the optimal communication strategy of the current time slot, and the state vector of the next time slot; repeat the above steps until convergence to obtain the final optimal communication strategy.
[0009] The embodied agent anti-interference communication method based on large model assistance proposed in this application has the following advantages: based on the environmental features of the embodied agent extracted by the large model, combined with channel gain, received interference intensity and motion state, reinforcement learning is used to simultaneously optimize uplink transmission power, channel and downlink transmission power, which can improve the success rate of the embodied agent in grasping tasks under malicious interference, while reducing communication latency and energy consumption.
[0010] Optionally, before obtaining the state vector of the current time slot, the system may also include: initializing the available power, channels, and policy space for communication between the embodied agent and the controller, as well as the parameters and experience pools for the benefit network and the risk network.
[0011] Optionally, initialize the available power, channels, and policy space for communication between the embodied agent and the controller, as well as the parameters and experience pool of the benefit network and the risk network, including: initializing the optional uplink transmit power of the embodied agent, the number of optional channels of the embodied agent, the downlink power order of the controller, and the optional policies; construct the benefit network and the risk network, each of which includes four fully connected layers, and initialize the learning parameters, including the discount factor, benefit coefficient, benefit network weight, risk network weight, number of experience points for updating risk values, convergence threshold, task success rate, communication latency, communication energy consumption, and experience pool.
[0012] Optionally, the feedback information from the previous time slot task includes the success rate of the previous time slot task, the communication latency of the previous time slot, and the communication energy consumption of the previous time slot. The environmental features are extracted from the depth image using a large model, and the environmental features include the types and location coordinates of objects around the embodied agent.
[0013] Optionally, the benefit value can be calculated using the following formula:
[0014] in, This represents the benefit value of the current time slot. This indicates the success rate of the embodied intelligent agent executing control commands to complete the grasping task. This indicates the communication delay of the current time slot. This indicates the communication energy consumption of the current time slot. , This represents the benefit coefficient.
[0015] Optionally, the risk value can be calculated using the following formula:
[0016] in, This indicates the risk value of the current time slot. Indicates the risk level index. Indicates the number of risk levels. Indicates the risk level threshold. This represents the risk level coefficient. Indicates communication delay.
[0017] Optionally, the above steps are repeated until convergence, including: determining whether the convergence condition is met; if it is met, convergence is determined; if it is not met, the above steps are repeated until convergence is achieved. If both the updated benefit network weight parameters and the updated risk network weight parameters are less than the preset convergence threshold, the convergence condition is met; otherwise, the convergence condition is not met.
[0018] A second objective of this invention is to provide a computer-readable storage medium storing a large model-assisted embodied agent anti-interference communication program, which, when executed by a processor, implements the large model-assisted embodied agent anti-interference communication method as described above.
[0019] The third objective of this invention is to provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements the aforementioned large model-assisted embodied intelligent agent anti-interference communication method. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating the anti-interference communication method for embodied intelligent agents based on large model assistance according to an embodiment of the present invention. Figure 2 The task success rate of the anti-interference method described in the embodiments of the present invention; Figure 3 The transmission delay of the anti-interference method described in the embodiments of the present invention; Figure 4 This refers to the communication energy consumption of the anti-interference method described in the embodiments of the present invention. Detailed Implementation
[0021] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0022] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the invention to those skilled in the art.
[0023] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0024] refer to Figure 1 As shown, the anti-interference communication method for embodied intelligent agents based on large model assistance in this embodiment of the invention includes the following steps: S101, obtain the state vector of the current time slot, where the state vector includes channel gain, received interference intensity, motion state, feedback information of the previous time slot task, and environmental characteristics.
[0025] The feedback information from the previous time slot includes the success rate of the previous time slot task, the communication latency of the previous time slot, and the communication energy consumption of the previous time slot. The environmental features are extracted from the depth image using a large model, and the environmental features include the types and coordinates of objects around the embodied agent.
[0026] Specifically, in the In the time slot, the controller obtains the channel gain with the embodied agent through pilot estimation. Software-defined radio equipment was used to measure the received interference strength. Receive motion status feedback from the embodied intelligent agent Success rate of the previous time slot Communication delay and energy consumption Using a large model based on depth images Extract environmental features such as the types and coordinates of people and vehicles around the embodied intelligent agent. .
[0027] Before obtaining the state vector of the current time slot, the process also includes: initializing the available power, channel, and policy space for communication between the embodied agent and the controller, as well as the parameters and experience pools for the benefit network and the risk network.
[0028] As an example, the available power, channels, and policy space for communication between the embodied agent and the controller, as well as the parameters and experience pool of the benefit network and the risk network, are initialized. This includes: initializing the optional uplink transmit power of the embodied agent, the number of optional channels of the embodied agent, the downlink power order of the controller, and the optional policies; constructing the benefit network and the risk network, each of which includes four fully connected layers; and initializing the learning parameters, including the discount factor, benefit coefficient, benefit network weight, risk network weight, number of experience points for updating the risk value, convergence threshold, task success rate, communication latency, communication energy consumption, and experience pool.
[0029] Specifically, during initialization, the transmit power of the embodied intelligent agent and the quantization order of the available channels are denoted as follows: and The downlink power order of the controller is denoted as Note the selectable transmit power. Number of selectable channels Downlink power is ,in, For the maximum selectable power, the selectable strategy is... .
[0030] Building Neural Networks and Each consists of four fully connected layers, with the input layer containing... One neuron; the second and third layers are hidden layers, each containing... and 10 neurons; the output layer contains 100 neurons; One neuron. Initialize learning parameters, including discount factors. Benefit coefficient and ,network and weight , The number of experience points used to update risk values is Algorithm convergence threshold , Initialization task success rate Initial transmission delay Communication energy consumption Initialize the experience pool .
[0031] S102, input the state vector of the current time slot into the benefit network and the risk network to obtain the expected benefits and expected risks of each optional strategy.
[0032] Specifically, constructing state vectors Input neural network and , obtained in the first Execution under multidimensional vectors in time slots Expected benefits of the strategy and in the Execution under multidimensional vectors in time slots Expected risks of the strategy .
[0033] S103, select the optimal communication strategy for the current time slot from the set of optional strategies based on the expected benefits and expected risks of each optional strategy. The optimal communication strategy includes optional uplink transmit power, optional channel and downlink transmit power.
[0034] Specifically, select according to the following formula
[0035] in, Indicates the first Selecting from the policy matrix of time slots The probability of the strategy This represents the set of all possible strategies.
[0036] S104: Obtain the grasping task inference request and depth image, and generate control commands using a large model based on the grasping task inference request and depth image; transmit the control commands, optional uplink transmission power and optional channel to the embodied agent according to the downlink transmission power, so that the embodied agent can execute the grasping action according to the control commands, and provide feedback on the current time slot task feedback information according to the optional uplink transmission power and optional channel.
[0037] Specifically, based on the grasping task, infer the request and the depth image containing the types of objects such as people and vehicles in the surrounding environment, as well as the layout of the environment. Large model is used to generate control commands. In the control channel with power Link messages Downlink transmission to the embodied intelligent agent.
[0038] S105: Receive the feedback information of the current time slot task and obtain the benefit value and risk value based on the feedback information of the current time slot task.
[0039] The benefit value is calculated according to the following formula:
[0040] in, This represents the benefit value of the current time slot. This indicates the success rate of the embodied intelligent agent executing control commands to complete the grasping task. This indicates the communication delay of the current time slot. This indicates the communication energy consumption of the current time slot. , This represents the benefit coefficient.
[0041] The risk value is calculated using the following formula:
[0042] in, This indicates the risk value of the current time slot. Indicates the risk level index. Indicates the number of risk levels. Indicates the risk level threshold. This represents the risk level coefficient. Indicates communication delay. S106, store the complete experience tuple of the current time slot into the experience pool, and randomly sample a batch of experiences from the experience pool to update the weight parameters of the benefit network and the risk network respectively. The complete experience tuple includes the benefit value of the current time slot, the risk value of the current time slot, the state vector of the current time slot, the optimal communication strategy of the current time slot, and the state vector of the next time slot.
[0043] Specifically, depending on the status ,Strategy ,benefit Risk Value And the observed state of the next time slot Experience in anti-interference transmission Store in experience pool In the middle, random sampling is performed from the experience pool. This experience, updating the network of benefits With risk network :
[0044]
[0045] S107, Repeat the above steps until convergence is achieved, and the final optimal communication strategy is obtained.
[0046] As an example, the above steps are repeated until convergence, including: determining whether the convergence condition is met; if it is met, convergence is determined; if it is not met, the above steps are repeated until convergence is achieved. If both the updated benefit network weight parameters and the updated risk network weight parameters are less than the preset convergence threshold, the convergence condition is met; otherwise, the convergence condition is not met.
[0047] To better understand the above technical solution, a specific embodiment is described in detail below, which includes the following steps: Step 1: Initialize parameters: Let the power order available to the embodied intelligent agent be... Number of available channels The number of transmit power orders available to the controller Maximum selectable power milliwatts, with selectable uplink transmit power of milliwatts, selectable channels are Downlink transmit power is milliwatts, optional strategy Building a benefit network and risk network The number of neurons in the four fully connected layers are as follows: , , ,and Initialize the discount factor. Benefit coefficient , ,network and weight The number of experience points used to update risk values Risk level quantity Risk level threshold Risk level coefficient Algorithm convergence threshold Initialization task success rate Initial transmission delay Communication energy consumption Experience pool .
[0048] Step 2: In the In the time slot, the controller obtains the channel gain with the embodied agent through pilot estimation. Software-defined radio equipment was used to measure the received interference strength. Receive motion status feedback from the embodied intelligent agent Success rate of the previous time slot Communication delay and energy consumption Using a large model based on depth images Extract environmental features such as the types and coordinates of people and vehicles around the embodied intelligent agent. .
[0049] Step 3: Build Status Input neural network and , obtained in the first Execution under multidimensional vectors in time slots Expected benefits of the strategy and in the Execution under multidimensional vectors in time slots Expected risks of the strategy .
[0050] Step 4: Select according to the following formula
[0051] in, Indicates the first Selecting from the policy matrix of time slots The probability of the strategy This represents the set of all possible strategies.
[0052] Step 5: Infer the request and depth image containing the types of objects such as people and vehicles in the surrounding environment, as well as the layout of the environment, based on the grasping task. Large model is used to generate control commands. .
[0053] Step 6: In the control channel with power Link messages Downlink transmission to the embodied intelligent agent.
[0054] Step 7: Calculate the benefits according to the following formula :
[0055] in, This indicates that the embodied intelligent agent executes control commands. Success rate of completing the data capture task. To measure uplink and downlink communication latency based on network time protocol synchronization timestamps, This represents the total energy consumption for communication in transmitting multimodal sensing data and control commands to embodied intelligent agents and controllers.
[0056] Step 8: Based on the threshold Transmission delay of multimodal sensing data and control commands Quantified as Each level, combined with corresponding penalty factors Assess the risk level of policy failures in embodied agent grasping tasks caused by delays in multimodal perception data and control commands:
[0057] Step 9: Based on the status ,Strategy ,benefit Risk Value And the observed state of the next time slot Experience in anti-interference transmission Store in experience pool middle.
[0058] Step 10: Randomly sample 32 experiences from the experience pool and update the benefit network. With risk network :
[0059]
[0060] Step 11: Repeat steps 3-10 until the condition is met. and The algorithm converged.
[0061] In summary, such as Figure 2 As shown, the method of this application rapidly increases the task success rate from approximately 0.45 in the early training phase, and stabilizes in a high success rate range of 0.7~0.75 after approximately 300 time slots, verifying that this application can effectively improve the success rate of grasping tasks by embodied intelligent agents under malicious interference. Figure 3 As shown, the transmission latency rapidly decreased from approximately 67 milliseconds in the early stages of training, and converged to a low-latency stable range of 38-41 milliseconds after approximately 200 time slots, verifying that this application can effectively reduce the communication latency between multimodal sensing data and control commands. Figure 4 As shown, the communication power consumption fluctuated significantly between 2.3 and 3.3 millijoules in the early stage of training, and then gradually decreased, stabilizing in the low power consumption range of 1.55 to 1.8 millijoules after about 800 time slots, which verifies that the present application can effectively reduce communication power consumption.
[0062] In addition, the present invention also proposes a computer-readable storage medium storing a large model-assisted embodied agent anti-interference communication program thereon, which, when executed by a processor, implements the large model-assisted embodied agent anti-interference communication method as described above.
[0063] In addition, this invention also proposes a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements the above-described large model-assisted embodied intelligent agent anti-interference communication method.
[0064] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0065] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1A device that provides the functions specified in one or more boxes.
[0066] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0067] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0068] It should be noted that any reference signs placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0069] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0070] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
[0071] In the description of this invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0072] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0073] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "over," and "on top" of the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.
[0074] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms should not be construed as necessarily referring to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0075] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A large-model-assisted embodied intelligent agent anti-interference communication method, characterized in that, Includes the following steps: Obtain the state vector of the current time slot, wherein the state vector includes channel gain, received interference intensity, motion state, feedback information of the previous time slot task, and environmental characteristics; The state vector of the current time slot is input into the benefit network and the risk network to obtain the expected benefits and expected risks of each optional strategy; The optimal communication strategy for the current time slot is selected from the set of optional strategies based on the expected benefits and expected risks of each optional strategy, wherein the optimal communication strategy includes optional uplink transmit power, optional channel and downlink transmit power; Obtain the grasping task inference request and depth image, and generate control instructions using a large model based on the grasping task inference request and depth image; The control command, optional uplink transmission power, and optional channel are transmitted to the embodied intelligent agent according to the downlink transmission power, so that the embodied intelligent agent can perform a grasping action according to the control command and provide feedback on the current time slot task feedback information according to the optional uplink transmission power and optional channel. Receive the current time slot task feedback information, and obtain the benefit value and risk value based on the current time slot task feedback information; The complete experience tuple of the current time slot is stored in the experience pool, and a batch of experiences is randomly sampled from the experience pool to update the weight parameters of the benefit network and the risk network respectively. The complete experience tuple includes the benefit value of the current time slot, the risk value of the current time slot, the state vector of the current time slot, the optimal communication strategy of the current time slot, and the state vector of the next time slot. Repeat the above steps until convergence is achieved, and the final optimal communication strategy is obtained.
2. The anti-interference communication method for embodied intelligent agents based on large model assistance as described in claim 1, characterized in that, Before obtaining the state vector of the current time slot, the process also includes: initializing the available power, channels, and policy space for communication between the embodied agent and the controller, as well as the parameters and experience pools for the benefit network and the risk network.
3. The anti-interference communication method for embodied intelligent agents based on large model assistance as described in claim 2, characterized in that, Initialize the available power, channel, and policy space for communication between the embodied agent and the controller, as well as the parameters and experience pools for the benefit network and the risk network, including: Initialize the embodied agent's optional uplink transmit power, the number of optional channels for the embodied agent, the controller's downlink power order, and optional policies; Construct a benefit network and a risk network, each consisting of four fully connected layers. Initialize learning parameters, including discount factor, benefit coefficient, benefit network weight, risk network weight, number of experience points for updating risk values, convergence threshold, task success rate, communication latency, communication energy consumption, and experience pool.
4. The anti-interference communication method for embodied intelligent agents based on large model assistance as described in claim 1, characterized in that, The feedback information from the previous time slot includes the success rate of the previous time slot, the communication latency of the previous time slot, and the communication energy consumption of the previous time slot. Environmental features are extracted from the depth image using a large model. Among these features, the environmental features include the types and coordinates of objects around the embodied agent.
5. The anti-interference communication method for embodied intelligent agents based on large model assistance as described in claim 1, characterized in that, Calculate the benefit value using the following formula: in, This represents the benefit value of the current time slot. This indicates the success rate of the embodied intelligent agent executing control commands to complete the grasping task. This indicates the communication delay of the current time slot. This indicates the communication energy consumption of the current time slot. , This represents the benefit coefficient.
6. The anti-interference communication method for embodied intelligent agents based on large model assistance as described in claim 1, characterized in that, Calculate the risk value using the following formula: in, This indicates the risk value of the current time slot. Indicates the risk level index. Indicates the number of risk levels. Indicates the risk level threshold. This represents the risk level coefficient. Indicates communication delay.
7. The anti-interference communication method for embodied intelligent agents based on large model assistance as described in claim 1, characterized in that, Repeat the above steps until convergence, including: Determine whether the convergence condition is met. If it is met, convergence is confirmed. If not, repeat the above steps until convergence is achieved. Specifically, if both the updated benefit network weight parameters and the updated risk network weight parameters are less than the preset convergence threshold, the convergence condition is met; otherwise, the convergence condition is not met.
8. A computer-readable storage medium, characterized in that, It stores a large model-assisted embodied agent anti-interference communication program, which, when executed by the processor, implements the large model-assisted embodied agent anti-interference communication method as described in any one of claims 1-7.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the anti-interference communication method for embodied intelligent agents based on large model assistance as described in any one of claims 1-7.