Communication method and communication device

By improving the information interaction and data processing flow between intelligent agents, the efficiency problem of multimodal data processing in intelligent networks has been solved, thereby enhancing data analysis and network maintenance.

CN120956594APending Publication Date: 2025-11-14HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410601980.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-14
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In intelligent networks based on AGI design, how can we enable agents to efficiently process and interact with multimodal data, thereby improving network intelligence and maintenance efficiency?

Method used

Through information interaction between the first and second intelligent agents, capability information, perception module configuration information, and task description information are transmitted to realize the analysis and processing of multimodal data, including the registration, connection, and data request of perception modules, data processing is performed using processing functions, and processing logs are fed back to adjust the scheduling.

Benefits of technology

It improves the processing efficiency of multimodal data and network maintenance in intelligent networks, and ensures the accuracy and robustness of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956594A_ABST
    Figure CN120956594A_ABST
Patent Text Reader

Abstract

The communication method comprises the steps that a first intelligent agent receives capability information of a second intelligent agent, and related information of first data and configuration information of a sensing module are determined to be sent to the second intelligent agent according to the capability information and task description information. The capability information comprises a processing function and / or modal information of the multi-modal data supported to be processed by the second intelligent agent, and the related information of the first data comprises the content of the first data, the modal information of the first data and the processing priority information of the first data. The configuration information of the sensing module comprises identification information of the sensing module and / or the first secret key. According to the method, through interaction between a first intelligent agent and a second intelligent agent, the first intelligent agent can analyze multi-modal data based on task description information and capability information of the second intelligent agent and indicate related information of the multi-modal data to be processed and configuration information of a sensing module for the second intelligent agent; the intelligent agent can process the multi-modal data, and the network intelligence is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communications, specifically to a communication method and a communication device. Background Technology

[0002] In the current field of computer science (CS), any entity capable of independent thought and interaction with its environment can be abstracted as an intelligent agent. The basic characteristics of an intelligent agent are: the ability to react to changes in the environment and automatically adjust its behavior and state; and the ability for different intelligent agents to interact with each other according to their own intentions.

[0003] Artificial general intelligence (AGI) will be an indispensable part of network architecture in the future. Therefore, how to enable agents to process multimodal data in intelligent networks based on AGI design is currently a research hotspot. Summary of the Invention

[0004] This application provides a communication method to enable intelligent agents to process multimodal data in intelligent networks designed based on AGI.

[0005] Firstly, this application provides a method that can be executed by a first intelligent agent. Unless otherwise specified, the term "first intelligent agent" in this application can refer to the intelligent agent itself, a component within the first intelligent agent (e.g., a processor, chip, or chip system), or a logic module or software capable of implementing all or part of the functions of the first intelligent agent. The following explanation uses a first intelligent agent as an example.

[0006] The method includes: a first agent receiving capability information from a second agent, the capability information including processing functions and / or modal information for multimodal data supported by the second agent; the first agent sending relevant information about first data and configuration information of a perception module, wherein the relevant information about the first data is determined based on the capability information and task description information, the task description information corresponding to the first data, the perception module corresponding to the first data, and the configuration information of the perception module including identification information and / or a first key. The multimodal data includes the first data, and the relevant information about the first data includes: the content of the first data, the modal information of the first data, and the processing priority information of the first data.

[0007] It should be understood that the perception module corresponding to the first data can be understood as the perception module being able to perceive and acquire the first data. For example, if the modality of the first data is "image", the perception module corresponding to the first data can perceive and acquire image-related data; or, for example, if the modality of the first data is "speech", the perception module corresponding to the first data can perceive and acquire speech-related data.

[0008] According to the method provided in this application, the interaction between the first and second intelligent agents enables the first intelligent agent to analyze multimodal data based on task description information and the capability information of the second intelligent agent, and to instruct the second intelligent agent on relevant information of the multimodal data to be processed (e.g., the first data) and the configuration information of the sensing module. This technical solution enables the intelligent agent to process multimodal data through information interaction between the first and second intelligent agents, improving network intelligence. Furthermore, introducing intelligent agents into an AGI-based intelligent network can further improve the efficiency of network maintenance and / or operation.

[0009] In conjunction with the first aspect, in some possible implementations, where the processing function includes a first processing function corresponding to the first data, the relevant information of the first data also includes the first processing function.

[0010] Based on the above technical solution, the capability information includes the processing functions supported by the second intelligent agent. The first intelligent agent determines the first processing function corresponding to the first data in the processing function according to the task description information, and instructs the first processing function to the second intelligent agent through the relevant information of the first data. This allows the second intelligent agent to directly use the first processing function indicated by the first intelligent agent when processing the first data, without having to match the corresponding processing function itself based on the relevant information of the first data.

[0011] In conjunction with the first aspect, in some possible implementations, before sending the relevant information of the first data and the configuration information of the sensing module, the method further includes: receiving registration request information from the sensing module, the registration request information including the data types and data content that the sensing module can perceive, and the configuration information of the sensing module being determined based on the registration request information and the relevant information of the first data.

[0012] It should be understood that the sensing module and the first intelligent agent can be carried on the same communication device or on different communication devices. For example, both the sensing module and the first intelligent agent can be carried on a network device, meaning the information interaction between the sensing module and the first intelligent agent can be an interaction within the network device. Alternatively, the sensing module can be carried on a terminal device, and the first intelligent agent can be carried on a network device, meaning the information interaction between the sensing module and the first intelligent agent can be an over-the-air interaction.

[0013] Secondly, a communication method is provided, which can be executed by a second intelligent agent. Unless otherwise specified, the "second intelligent agent" in this application can refer to the second intelligent agent itself, a component of the second intelligent agent (e.g., a processor, chip, or chip system), or a logic module or software that can implement all or part of the functions of the second intelligent agent. The following description uses a second intelligent agent as an example.

[0014] The method includes: a second agent sending capability information of the second agent, the capability information including processing functions and / or modal information of multimodal data supported by the second agent; the second agent receiving relevant information of first data and configuration information of a perception module, wherein the relevant information of the first data is determined based on the capability information and task description information, the task description information corresponding to the first data, the perception module corresponding to the first data, and the configuration information of the perception module including identification information and / or a first key of the perception module, wherein the multimodal data includes the first data, and the relevant information of the first data includes: the content of the first data, the modal information of the first data, and the processing priority information of the first data.

[0015] It should be understood that the technical effects of the methods shown in the second aspect and its possible designs can be referred to the technical effects in the first aspect and its possible designs, and will not be repeated here.

[0016] In conjunction with the second aspect, in some possible implementations, the method further includes: a second intelligent agent establishing a connection with a first sensing module in the sensing module based on the identification information of the sensing module and / or the first key; the second intelligent agent sending a request message to the first sensing module, the request message being used to request the acquisition of the first data; and the second intelligent agent receiving the first data from the first sensing module.

[0017] In conjunction with the second aspect, in some possible implementations, the request information includes at least one of the following: content description information of the first data, data type of the first data, or data format of the first data.

[0018] In conjunction with the second aspect, in some possible implementations, where the processing function includes a first processing function corresponding to the first data, the relevant information of the first data also includes the first processing function.

[0019] In conjunction with the second aspect, in some possible implementations, where the capability information includes the modal information, the method further includes: a second agent performing function matching on the first data based on relevant information of the first data to determine a first processing function.

[0020] It should be understood that when capability information includes modal information, it can be interpreted as capability information not including processing functions.

[0021] Based on the above technical solution, when the capability information does not include the processing functions supported by the second intelligent agent, when the second intelligent agent receives the relevant information of the first data and the configuration information of the perception module, the second intelligent agent performs function matching on the first data according to the relevant information of the first data received, determines the first processing function corresponding to the first data, and performs multimodal data processing on the first data based on the first processing function.

[0022] In conjunction with the second aspect, in some possible implementations, the method further includes: a second agent performing multimodal data processing on the first data according to the first processing function; the second agent sending a processing log, the processing log including at least one of the following: the processing priority of the first data, the processing completion status of the first data, or the processing result of the first data.

[0023] Based on the above technical solution, after the second intelligent agent completes multimodal data processing of the first data, the second intelligent agent will feed back the processing log to the first intelligent agent, so that the first intelligent agent can adjust the information required for scheduling according to the task plan.

[0024] In conjunction with the second aspect, in some possible implementations, the processing completion status of the first data includes: processing not completed or processing completed.

[0025] In conjunction with the second aspect, in some possible implementations, before the second agent sends its capability information, the method further includes: the second agent acquiring a multimodal model; the second agent decomposing the functions of the multimodal model to identify robust functions; and the second agent determining a processing function for the multimodal data based on the robust functions.

[0026] Based on the above technical solution, the second intelligent agent decomposes the functions of the multimodal model, identifies robust functions, and functions the robust functions to determine the processing functions, thereby ensuring the robustness of multimodal data processing and improving the performance of the second intelligent agent in multimodal data processing.

[0027] Thirdly, a communication method is provided, which can be executed by a communication device. Unless otherwise specified, the term "communication device" in this application can refer to the communication device itself (e.g., core network equipment, access network equipment, terminal equipment, network equipment), a component in the communication device (e.g., processor, chip, or chip system), or a logic module or software that can implement all or part of the communication device.

[0028] The method includes: a second agent sending capability information of the second agent to a first agent, the capability information including processing functions and / or modal information of multimodal data supported by the second agent; the first agent determining relevant information of first data and configuration information of a perception module based on task description information and the capability information of the second agent, the task description information corresponding to the first data and the perception module corresponding to the first data; the first agent sending relevant information of the first data and configuration information of the perception module to the second agent, the configuration information of the perception module including identification information and / or a first key of the perception module, wherein the multimodal data includes the first data, and the relevant information of the first data includes: the content of the first data, the modal information of the first data, and the processing priority information of the first data.

[0029] The technical effects of the methods shown in the third aspect and its possible designs above can be referred to the technical effects in the first aspect, the second aspect and its possible designs.

[0030] In conjunction with the third aspect, in some possible implementations, where the processing function includes a first processing function corresponding to the first data, the relevant information of the first data also includes the first processing function.

[0031] In conjunction with the third aspect, in some possible implementations, before sending the relevant information of the first data and the configuration information of the sensing module, the method further includes: receiving registration request information from the sensing module, the registration request information including the data types and data content that the sensing module can perceive, and the configuration information of the sensing module being determined based on the registration request information and the relevant information of the first data.

[0032] In conjunction with the third aspect, in some possible implementations, the method further includes: establishing a connection with a first sensing module in the sensing modules based on the identification information of the sensing module and / or the first key; sending a request message to the first sensing module, the request message being used to request the acquisition of the first data; and receiving the first data from the first sensing module.

[0033] In conjunction with the third aspect, in some possible implementations, the request information includes at least one of the following: content description information of the first data, data type of the first data, or data format of the first data.

[0034] In conjunction with the third aspect, in some possible implementations, where the capability information includes the modal information, the method further includes: performing function matching on the first data based on the relevant information of the first data to determine a first processing function.

[0035] In conjunction with the third aspect, in some possible implementations, the method further includes: performing multimodal data processing on the first data according to the first processing function; sending a processing log, the processing log including at least one of the following: the processing priority of the first data, the processing completion status of the first data, or the processing result of the first data; and receiving the processing log.

[0036] In conjunction with the third aspect, in some possible implementations, the processing completion status of the first data includes: processing not completed or processing completed.

[0037] In conjunction with the third aspect, in some possible implementations, before sending the capability information of the second agent, the method further includes: acquiring a multimodal model; decomposing the functions of the multimodal model to identify robust functions; and determining a processing function for the multimodal data based on the robust functions.

[0038] Fourthly, a communication apparatus is provided for performing the method provided in the first aspect. Specifically, the communication apparatus may include units and / or modules for performing the method provided in any of the above implementations of the first aspect, such as a processing unit and an acquisition unit.

[0039] In one implementation, the transceiver unit can be a transceiver or an input / output interface; the processing unit can be at least one processor. Optionally, the transceiver can be a transceiver circuit. Optionally, the input / output interface can be an input / output circuit.

[0040] In another implementation, the transceiver unit can be an input / output interface, interface circuit, output circuit, input circuit, pin, or related circuit on the chip, chip system, or circuit; the processing unit can be at least one processor, processing circuit, or logic circuit.

[0041] Fifthly, a communication apparatus is provided for performing the method provided in the second aspect. Specifically, the communication apparatus may include units and / or modules for performing the method provided in the second aspect, such as a processing unit and an acquisition unit.

[0042] In one implementation, the transceiver unit can be a transceiver or an input / output interface; the processing unit can be at least one processor. Optionally, the transceiver can be a transceiver circuit. Optionally, the input / output interface can be an input / output circuit.

[0043] In another implementation, the transceiver unit can be an input / output interface, interface circuit, output circuit, input circuit, pin, or related circuit on the chip, chip system, or circuit; the processing unit can be at least one processor, processing circuit, or logic circuit.

[0044] In a sixth aspect, a communication apparatus is provided for performing the method provided in the third aspect. Specifically, the communication apparatus may include units and / or modules for performing the method provided in the third aspect, such as a processing unit and an acquisition unit.

[0045] In one implementation, the transceiver unit can be a transceiver or an input / output interface; the processing unit can be at least one processor. Optionally, the transceiver can be a transceiver circuit. Optionally, the input / output interface can be an input / output circuit.

[0046] In another implementation, the transceiver unit can be an input / output interface, interface circuit, output circuit, input circuit, pin, or related circuit on the chip, chip system, or circuit; the processing unit can be at least one processor, processing circuit, or logic circuit.

[0047] In a seventh aspect, this application provides a processor for executing the method provided by any of the implementations of the first to third aspects described above.

[0048] Unless otherwise specified, or if it does not contradict its actual function or internal logic in the relevant description, the transmission and acquisition / reception operations involved in the processor can be understood as processor output and reception, input and other operations, or as transmission and reception operations performed by radio frequency circuits and antennas. This application does not limit them in this regard.

[0049] Eighthly, a computer-readable storage medium is provided that stores program code for execution by a device, the program code including a method for performing any of the implementations of the first to third aspects described above.

[0050] Ninth aspect, a computer program product containing instructions is provided, which, when run on a computer, causes the computer to perform the method provided by any one of the implementations of the first to third aspects described above.

[0051] In a tenth aspect, a chip is provided, the chip including one or more processors and a communication interface, wherein the processor reads a computer program or instructions stored in a memory through the communication interface and executes the method provided by any of the implementations of the first to third aspects described above.

[0052] Optionally, as one implementation, the chip also includes a memory storing computer programs or instructions. The processor is used to execute the computer programs or instructions stored in the memory. When the computer programs or instructions are executed, the processor is used to execute the method provided by any of the first to third aspects described above.

[0053] Eleventhly, a communication system is provided, including the aforementioned first intelligent agent and second intelligent agent. Attached Figure Description

[0054] Figure 1 This is a schematic diagram of a communication system applicable to this application.

[0055] Figure 2 This is a schematic diagram of an AI network element built into a communication system.

[0056] Figure 3 This is a schematic diagram of an intelligent agent.

[0057] Figure 4 This is a schematic flowchart of a communication method provided in this application.

[0058] Figure 5 This is a schematic flowchart of another communication method provided in this application.

[0059] Figure 6 This is a schematic flowchart of another communication method provided in this application.

[0060] Figure 7 This is a schematic block diagram of a communication device provided in an embodiment of this application.

[0061] Figure 8 This is a schematic diagram of another communication device provided in an embodiment of this application. Detailed Implementation

[0062] To facilitate understanding of the embodiments of this application, the following points will be explained first.

[0063] First, in this application, "for indicating" can include both direct and indirect indication. When describing an indication information as indicating A, it can include whether the indication information directly indicates A or indirectly indicates A, but does not necessarily mean that the indication information carries A.

[0064] The information indicated by the instruction is called the information to be instructed. In the specific implementation process, there are many ways to indicate the information to be instructed, such as, but not limited to, directly indicating the information to be instructed, such as the information to be instructed itself or its index. It can also be indirectly indicated by indicating other information, where there is a relationship between the other information and the information to be instructed. It can also indicate only a part of the information to be indicated, while the other parts are known or pre-agreed upon. For example, the instruction of specific information can be achieved by using a pre-agreed (e.g., protocol-defined) arrangement of various pieces of information, thereby reducing instruction overhead to some extent. At the same time, common parts of various pieces of information can be identified and indicated uniformly to reduce the instruction overhead caused by individually indicating the same information.

[0065] Second, in this application, "at least one" refers to one or more, and "more than one" refers to two or more (including two). Furthermore, in the embodiments of this application, "first," "second," and various numerical designations (e.g., "#1," "#2," etc.) are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application. The sequence numbers of the processes below do not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. It should be understood that the objects described in this way can be interchanged where appropriate to describe solutions other than those in the embodiments of this application. Moreover, in the embodiments of this application, terms such as "S410" are merely identifiers for descriptive convenience and do not limit the order of execution steps.

[0066] Third, in the embodiments of this application, the words "exemplary" or "for example" are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design that is described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design options. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0067] Fourth, the term "storage" in the embodiments of this application can refer to storage in one or more memories. These memories can be separate installations or integrated into an encoder, decoder, processor, or communication device. Alternatively, some memories can be separately installed, while others are integrated into the processor or communication device. The type of memory can be any form of storage medium, and this application does not limit this.

[0068] Fifth, in the implementation of this application, "protocol" may refer to standard protocols in the field of communications, such as the NR protocol and related protocols applied in future communication systems, and this application does not limit it.

[0069] Sixth, in the embodiments of this application, the terms "of", "corresponding (relevant)", "corresponding", and "associate" can sometimes be used interchangeably. It should be noted that when their differences are not emphasized, their intended meanings are consistent.

[0070] Seventh, in the embodiments of this application, "under the circumstances", "when", and "if" can sometimes be used interchangeably. It should be noted that when the distinction is not emphasized, their intended meanings are consistent.

[0071] Eighth, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0072] Ninth, in this article, "message", "information", or "information element (IE)" can be used interchangeably. There are no restrictions on the name of the message or information, as long as it can achieve the corresponding function.

[0073] In this application, "send" and "receive" indicate the direction of signal transmission. For example, "send information to XX" can be understood as the destination of the information being XX, and "send information" can include direct transmission or indirect transmission through other units or modules. "Receive information from YY" can be understood as the source of the information being YY, and "receive information" can include direct reception from YY or indirect reception from YY through other units or modules. Furthermore, "send" can also be understood as the "output" of a chip interface, and "receive" can be understood as the "input" of a chip interface. In other words, "send" or "receive" can occur between devices, such as network devices and terminal devices transmitting or receiving data via an air interface, or they can occur within a device, such as transmitting or receiving data between components, modules, chips, software modules, or hardware modules within a device via a bus, wiring, or interface.

[0074] The technical solutions in this application will now be described with reference to the accompanying drawings.

[0075] The technical solutions of this application embodiment can be applied to various communication systems, such as: Long Term Evolution (LTE) systems, LTE Frequency Division Duplex (FDD) systems, LTE Time Division Duplex (TDD) systems, Universal Mobile Telecommunication System (UMTS), Worldwide Interoperability for Microwave Access (WiMAX) systems, 5th Generation (5G) systems, New Radio (NR) and future communication networks, vehicle-to-other devices (V2X), where V2X can include vehicle-to-network (V2N), vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I), vehicle-to-pedestrian (V2P), etc., Long Term Evolution-V (LTE-V) technology for vehicle-to-everything (V2X), vehicle-to-everything (V2X), machine-type communication (MTC), and the Internet of Things (IoT). Things (IoT), Long Term Evolution of Machines (LTE-M), Machine to Machine (M2M), etc.

[0076] Furthermore, the embodiments of this application are applicable to both homogeneous and heterogeneous network scenarios, and there are no restrictions on the transmission points. Systems such as multi-point collaborative transmission between macro base stations, micro base stations, and macro base stations are all applicable. The embodiments of this application are applicable to both low-frequency scenarios (sub-6G) and high-frequency scenarios (above 6G), terahertz, optical communication, etc.

[0077] Figure 1 This is a schematic diagram of a communication system applicable to this application. For example... Figure 1 As shown, the communication system 100 includes at least one network device, such as... Figure 1 At least one of network device 111, network device 112, and network device 113 shown; the communication system 100 may also include at least one terminal device, such as Figure 1At least one of the terminal devices 121 and 122 shown; the communication system 100 may also include at least one AI network element, such as Figure 1 The AI ​​network element 131 is shown. In this communication system, network devices and terminal devices can communicate via a wireless link to exchange information. It is understood that network devices and terminal devices can also be referred to as communication devices or communication apparatuses.

[0078] A network device is a network-side device with wireless transceiver capabilities. A network device can be a device in a radio access network (RAN) that provides wireless communication capabilities to terminal devices; this is called a RAN device. An RAN can be a cellular system related to the 3rd Generation Partnership Project (3GPP), such as a 5G mobile communication system, or a future-oriented evolution system (such as Future Communication Networks). An RAN can also be an open radio access network (O-RAN or ORAN), a cloud radio access network (CRAN), or a wireless fidelity (WiFi) system. For example, this network device can be a base station, an evolved NodeB (eNodeB), a next-generation NodeB (gNB) in a 5G mobile communication system, a 3GPP subsequent evolution base station, a transmission reception point (TRP), an access node, a wireless relay node, or a wireless backhaul node in a WiFi system. In communication systems using different radio access technologies (RATs), the name of the device with base station functionality may differ. For example, in an LTE system it may be called an eNB or eNodeB, and in a 5G or NR system it may be called a gNB. This application does not limit the specific name of the base station. Network equipment may include one or more co-located or non-co-located transmitting and receiving points.

[0079] For example, a network device may include at least one of the following: one or more central units (CUs), one or more distributed units (DUs), and one or more radio units (RUs). In different systems, CUs (or CU-control plane (CP), CU-user plane (UP)), DUs, or RUs may have different names, but those skilled in the art will understand their meaning. For example, in an ORAN system, a CU may also be called an open CU (O-CU), a DU may also be called an open DU (O-DU), a CU-CP may also be called an open CU-CP (O-CU-CP), a CU-UP may also be called an open CU-UP (O-CU-UP), and a RU may also be called an open RU (O-RU). Any of the CUs (or CU-CP, CU-UPs), DUs, and RUs in this application may be implemented through software modules, hardware modules, or a combination of software and hardware modules. Exemplarily, the functionality of a CU may be implemented by one entity or different entities. For example, the functions of the CU can be further divided, separating the control plane and user plane and implementing them through different entities: the control plane CU entity (i.e., the CU-CP entity) and the user plane CU entity (i.e., the CU-UP entity). The CU-CP and CU-UP entities can be coupled with the DU to jointly complete the functions of the access network device. For instance, the CU is responsible for handling non-real-time protocols and services, implementing the functions of the radio resource control (RRC) and packet data convergence protocol (PDCP) layers. The DU is responsible for handling physical layer protocols and real-time services, implementing the functions of the radio link control (RLC), media access control (MAC), and physical (PHY) layers. In this way, some functions of the wireless access network device can be implemented through multiple network function entities. These network function entities can be network elements in hardware devices, software functions running on dedicated hardware, or virtualized functions instantiated on a platform (e.g., a cloud platform). Network devices can also include active antenna units (AAUs). AAU implements some physical layer processing functions, radio frequency processing, and related functions of active antennas.Since information from the RRC layer ultimately becomes information from the PHY layer, or is derived from information from the PHY layer, in this architecture, higher-layer signaling, such as RRC layer signaling, can also be considered as being sent by the DU, or by the DU+AAU. It is understood that network devices can be devices including one or more of the following: CU nodes, DU nodes, and AAU nodes. Furthermore, the CU can be classified as a network device in the RAN, or it can be classified as a network device in the core network (CN); this application does not limit this. For example, in vehicle-to-everything (V2X) technology, the access network device can be a roadside unit (RSU). Multiple access network devices in the communication system can be base stations of the same type or different types. Base stations can communicate with terminal devices, or they can communicate with terminal devices through relay stations. In the embodiments of this application, the device used to implement the network device function can be the network device itself, or it can be a device that supports the network device in implementing this function, such as a chip system or a combination of devices or components that can implement the access network device function; this device can be installed in the network device. In this embodiment of the application, the chip system may be composed of chips or may include chips and other discrete devices.

[0080] A terminal device is a user-side device with wireless transceiver capabilities. It can be a fixed device, mobile device, handheld device (e.g., mobile phone), wearable device, in-vehicle device, or a wireless device (e.g., communication module, modem, or chip system) built into the aforementioned devices. Terminal devices are used to connect people, things, and machines, and can be widely used in various scenarios, such as: cellular communication, device-to-device (D2D) communication, V2X communication, machine-to-machine / machine-type communications (M2M / MTC) communication, the Internet of Things (IoT), virtual reality (VR), augmented reality (AR), industrial control, self-driving, remote medical care, smart grids, smart furniture, smart offices, smart wearables, smart transportation, smart cities, drones, and robots. For example, a terminal device can be a handheld terminal in cellular communication, a communication device in D2D, an IoT device in MTC, a surveillance camera in intelligent transportation and smart cities, or a communication device on a drone, etc. Terminal devices are sometimes referred to as user equipment (UE), user terminal, user device, user unit, user station, terminal, access terminal, access station, UE station, remote station, mobile device, or wireless communication device, etc. A terminal device can also be a terminal device in an IoT system. IoT is an important component of future information technology development. Its main technical characteristic is connecting objects to networks through communication technology, thereby realizing an intelligent network of human-machine interconnection and machine-to-machine interconnection. In the embodiments of this application, IoT technology can achieve massive connectivity, deep coverage, and terminal power saving through, for example, narrowband (NB) technology. In the embodiments of this application, the device used to implement the functions of the terminal device can be the terminal device itself, or a device capable of supporting the terminal device to implement the functions, such as a chip system or a combination of devices or components capable of implementing the functions of the terminal device. This device can be installed in the terminal device.

[0081] AI network elements can perform some or all AI-related operations. These AI network elements can also be called AI nodes, AI devices, AI entities, AI modules, AI models, or AI units. An AI model can be considered a specific method for implementing AI functions. An AI model represents the mapping relationship or function between the model's input and output. AI functions can include one or more of the following: data collection, model training (or model learning), model information dissemination, model inference (or model reasoning, inference, or prediction, etc.), model monitoring or model validation, or inference result dissemination, etc. AI functions can also be called AI (related) operations or AI-related functions.

[0082] The AI ​​module is used to implement corresponding AI functions. AI modules deployed in different network elements can be the same or different. Depending on the parameter configuration, the AI ​​module can implement different functions. The AI ​​module model can be configured based on one or more of the following parameters: structural parameters (e.g., at least one of the following: number of neural network layers, neural network width, inter-layer connections, neuron weights, neuron activation function, or bias in the activation function), input parameters (e.g., type and / or dimension of input parameters), or output parameters (e.g., type and / or dimension of output parameters). The bias in the activation function can also be referred to as the neural network bias.

[0083] An AI module can have one or more models. A model can infer an output, which includes one or more parameters. The learning, training, or inference processes of different models can be deployed on different nodes or devices, or they can be deployed on the same node or device.

[0084] For example, the AI ​​network element can be built into the communication system. For instance, the AI ​​network element can be an AI module built into access network equipment, core network equipment, cloud servers, or operation, administration, and maintenance (OAM) management systems to implement AI-related functions. The core network equipment includes, but is not limited to, network elements such as access and mobility management function (AMF), user plane function (UPF), or session management function (SMF). The OAM can be the network management system of the core network equipment and / or the network management system of the access network equipment. Alternatively, the AI ​​network element can also be a network element independently set up in the communication system. Optionally, the terminal or its built-in chip can also include an AI entity to implement AI-related functions.

[0085] For ease of understanding, combined with Figure 2 This section provides a brief introduction to the integration of AI modules and communication systems. For example... Figure 2 As shown, network elements in a communication system are connected via interfaces (e.g., NG, Xn) or air interfaces. These network element nodes, such as core network equipment, access network nodes (RAN nodes), terminals, or one or more devices in the OAM, are equipped with one or more AI modules (for clarity, ...). Figure 2 (Only one AI module is shown in the diagram). The access network node can be a single RAN node or can include multiple RAN nodes, such as CU and DU. The CU and / or DU can also be equipped with one or more AI modules. Optionally, the CU can also be split into CU-CP and CU-UP. One or more AI models are set in CU-CP and / or CU-UP.

[0086] Network devices, terminal devices, and AI network elements can be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; they can also be deployed on water; and they can also be deployed in the air on airplanes, balloons, and satellites. This application does not limit the scenarios in which the network devices, terminal devices, and core network devices are located.

[0087] To facilitate understanding of the embodiments of this application, the basic concepts involved in this application will be explained first.

[0088] 1. AI: AI enables machines to possess human-like intelligence, for example, allowing machines to use computer hardware and software to simulate certain intelligent human behaviors. To achieve artificial intelligence, machine learning methods can be employed. In machine learning, machines learn (or train) models using training data. This model represents the mapping between inputs and outputs. The learned model can be used for reasoning (or prediction), that is, it can be used to predict the output corresponding to a given input. This output can also be called the reasoning result (or prediction result).

[0089] 2. Large Model Techniques: Large models refer to neural network models containing an extremely large number of parameters (usually over one billion), and have the following characteristics:

[0090] 1) Huge scale: Large models contain billions of parameters, and their size can reach hundreds of gigabytes (GB) or even larger. This huge model scale provides powerful expressive and learning capabilities.

[0091] 2) Multi-task learning: Large models typically learn multiple different natural language processing (NLP) tasks, such as machine translation, text summarization, or question answering systems. This allows the model to learn a broader and more generalized language understanding ability.

[0092] 3) Powerful computing resources: Training large models typically requires hundreds or even thousands of graphics processing units (GPUs) and a significant amount of time, usually ranging from weeks to months. This can accelerate the training process while preserving the ability to train large models.

[0093] 4) Abundant data: Large models require a large amount of data for training. Only with a large amount of data can the advantages of the parameter scale of large models be brought into play.

[0094] Large models are widely used in the field of natural language processing (NLP) and are transforming NLP tasks, giving rise to more powerful and intelligent language technologies. Large models are a key direction in AI development. They also excel in various NLP tasks, such as text classification, sentiment analysis, summarization, and translation. Furthermore, large models can be used in multiple application areas, including automated writing, chatbots, virtual assistants, voice assistants, and automated translation.

[0095] It should be understood that the application of large-scale models in networks requires a series of supporting peripheral functions to truly realize their potential. This system engineering can be called an AI Agent. The following is a brief explanation of AI Agents.

[0096] 3. Intelligent Agent: This is a concept in the field of artificial intelligence. Any entity capable of independent thought and interaction with its environment can be abstracted as an intelligent agent. The basic characteristics of an intelligent agent are: it can react to changes in its environment and automatically adjust its behavior and state; different intelligent agents can also interact with other intelligent agents according to their own intentions. Intelligent agents can be considered a type of AI module.

[0097] In this application, an intelligent agent is generally considered to be an agent that can autonomously complete a set goal through action. "Intelligent agent" is inseparable from "intelligence"; an intelligent agent possesses some human-like intelligent abilities and behaviors, such as learning, reasoning, decision-making, and execution capabilities. Optionally, "intelligent agent" can be replaced with other terms, such as artificial general intelligence (AGI), artificial intelligence, or intelligent agent.

[0098] Optionally, in a communication device, the intelligent agent can be integrated into the existing hardware / software of the communication device, or the intelligent agent can be independent of the existing hardware / software of the communication device. For example, the existing hardware / software may include chips, baseband chips, modem chips, system-on-chip (SoC) chips containing modem cores, system-in-package (SIP) chips, communication modules, chip systems, processors, logic modules, or software, etc.

[0099] For example, an intelligent agent can use a large language model (LLM) as its core, comprising a memory module, a tool module, a planning module, and an action module. The memory module implements long-term and / or short-term memory functions; the tool module contains multiple callable external tools; the planning module contains various planning algorithms; for example, the agent can plan externally input tasks based on its memory; and the action module supports the agent in performing actions based on the planning results, such as calling tools.

[0100] like Figure 3 As shown, in an LLM-supported autonomous agent system, the LLM acts as the brain of the agent (or agent), and is supplemented by several key components:

[0101] 1) Planning, including but not limited to:

[0102] Subgoal decomposition: Agents break down large tasks into smaller, manageable subgoals, enabling them to handle complex tasks more efficiently. For example, by instructing the model to "think step by step" through a chain of thoughts (CoT), more testing time is used to compute the breakdown of difficult tasks into smaller, simpler steps. CoT transforms large tasks into multiple manageable tasks and elucidates the explanation of the model's thought process.

[0103] Reflection and Improvement: Intelligent agents can engage in self-criticism and self-reflection on past behaviors, learn from mistakes, and improve future steps, thereby enhancing the quality of the final result.

[0104] 2) Memory, including but not limited to:

[0105] Short-term memory: Learning by utilizing the short-term memory of models.

[0106] Long-term memory: Provides agents with the ability to retain and recall (unlimited) information for a long time, usually by utilizing external vector storage and fast retrieval.

[0107] 3) Tool usage, including but not limited to:

[0108] Agent learning calls external application programming interfaces (APIs) to obtain additional information missing from the model weights (which is usually difficult to change after pre-training), including current information, code execution capabilities, and access to proprietary information sources.

[0109] 4) Task execution (action): The model performs a specific task and records the results.

[0110] It should be understood that the intelligent agent in this application may also be called an AI controller, intelligent unit, or intelligent entity, etc. This application does not limit the name of the intelligent agent, as long as it can achieve the corresponding function.

[0111] For example, the intelligent agent includes an AI model, a tool management unit, and a data management unit, wherein the tool management unit includes tools that can be called by the AI ​​model, and the data management unit is used to manage data related to the operation of the intelligent agent.

[0112] Optionally, the AI ​​model can be understood as the core of the intelligent agent, such as an LLM core. This AI model is used to coordinate other components (e.g., tool management unit, data management unit, etc.) to plan and schedule tasks based on task requirements. This application does not impose any limitations on the name of the AI ​​model, as long as it can achieve the corresponding function.

[0113] Optionally, the tool management unit includes tools that the AI ​​model can call, such as code compilers, interpreters, performance monitoring, digital twins, ray tracing, or network function virtualization tools, to enable the first intelligent agent to perform the corresponding functions.

[0114] Optionally, the data management unit can be used to manage, for example, device operation logs, agent operation logs, domain knowledge, device-supported functions, and the status of device sensing and measurement.

[0115] The above description of the terminology is for ease of understanding only and does not limit the scope of protection of the embodiments of this application.

[0116] The above text combined Figure 1 or Figure 2This paper briefly introduces the scenarios in which the communication method provided in the embodiments of this application can be applied, as well as the basic concepts that may be involved in the embodiments of this application. The concept of intelligent agent is introduced in the basic concepts. As can be seen from the above, intelligent agent can realize a variety of functions and has a high degree of intelligence.

[0117] It should also be understood that the above are exemplary examples of intelligent agent architecture in this application, and this application does not limit the specific architecture of intelligent agents.

[0118] 4. Artificial General Intelligence (AGI)

[0119] AGI refers to an artificial intelligence system capable of thinking, learning, and performing multiple tasks like a human. AGI is considered a higher level of artificial intelligence and represents an important direction and goal in the current development of artificial intelligence technology.

[0120] In AGI, the controlling agent (C-Agent) and the embodied agent (E-Agent) are two important concepts.

[0121] C-Agent refers to the controlling entity of an intelligent system, primarily responsible for decision-making and guiding the behavior of the entire system. A C-Agent can be understood as a high-level decision-maker that can formulate appropriate strategies and plans based on environmental changes and objective requirements, and guide E-Agents to execute corresponding tasks. C-Agents typically possess learning, reasoning, and decision-making abilities, enabling them to make flexible decisions based on different situations.

[0122] An E-Agent is an intelligent agent that performs specific tasks and acts as the executor of the C-Agent. An E-Agent can be a robot, a virtual character, or other form of entity. E-Agents perceive their environment, collect information, and execute corresponding actions based on instructions from the C-Agent. E-Agents typically possess capabilities such as perception, movement, and interaction, enabling real-time interaction and feedback with their environment.

[0123] It is evident that the C-Agent is the controlling entity of the intelligent system, responsible for decision-making and guiding the behavior of the entire system. The E-Agent is the intelligent agent that performs specific tasks, perceiving the environment and executing corresponding actions based on the instructions of the C-Agent.

[0124] 5. Multimodal:

[0125] A modal is a way in which things are experienced and occur, or a way of expressing or perceiving things. Every source or form of information can be called a modality. For example, humans have touch, hearing, vision, and smell; information media include speech, video, and text; and various sensors, such as radar, infrared, and accelerometers, can each be considered a modality. Compared to the classification of multimedia data such as images, speech, and text, "modality" is a more granular concept; different modalities may exist within the same medium. For example, two different languages ​​can be considered two different modalities, and even data collected under two different circumstances can be considered two modalities.

[0126] Multimodality refers to the expression or perception of things in multiple modalities. Multimodality can be categorized into homogeneous modalities and heterogeneous modalities. Homogeneous modalities include, for example, photographs taken by two different cameras. Heterogeneous modalities include, for example, the relationship between images and textual language.

[0127] Large multimodal models (LMMs) are a class of artificial intelligence models capable of processing and understanding various types of data input, such as text, images, audio, and video. LMMs can integrate and understand different data formats. For example, an LMM can analyze news articles (text), related photos (images), and related video clips (video clips) to gain a relatively comprehensive understanding.

[0128] Currently, Language Models (LMMs) are replacing Language Models (LLMs) as the brain of intelligent agents to process multimodal data. Since LLMs are generally evaluated primarily on language understanding and generation tasks, such as through metrics like fluency, coherence, and relevance, LMMs are mainly applied to a wider range of evaluation metrics. LMMs need to be proficient in multiple domains, such as through metrics like image recognition accuracy, audio processing quality, and the model's ability to integrate cross-modal information. With the continuous development of science and technology, LMMs are increasingly replacing LLMs as the brain of intelligent agents to process multimodal data, becoming a current research hotspot.

[0129] The above text combined Figure 1 This paper briefly introduces the scenarios in which the communication method provided in the embodiments of this application can be applied, and introduces the basic concepts that may be involved in the embodiments of this application. Among the basic concepts, the concept of intelligent agent is introduced. As can be seen from the above, intelligent agent can realize a variety of functions and has a high degree of intelligence.

[0130] This application provides a communication method that can be applied to... Figure 1 The communication system shown aims to establish connections between intelligent agents and devices in the communication network, thereby improving the intelligence level of the communication network.

[0131] It should be understood that the embodiments shown below do not particularly limit the specific structure of the execution subject of the method provided in the embodiments of this application, as long as it is possible to communicate according to the method provided in the embodiments of this application by running a program that records the code of the method provided in the embodiments of this application. For example, the execution subject of the method provided in the embodiments of this application can be a device, or a functional module in the device that can call and execute a program.

[0132] Figure 4 This is a schematic flowchart illustrating a communication method provided in this application. It includes the following steps:

[0133] 401. The second agent sends capability information to the first agent. Correspondingly, the first agent receives the capability information from the second agent.

[0134] This capability information includes processing functions and / or modal information for the multimodal data that the second agent supports processing.

[0135] In one possible implementation, the capability information includes processing functions for multimodal data that the second agent supports processing. These processing functions may be functions by which the second agent decomposes the multimodal model functionality and determines the resulting functionality. The capability information may include the processing function or an identification (ID) corresponding to the processing function.

[0136] For example, the second agent supports processing functions #1, #2, and #3 for multimodal processing, where function #1 has a function ID of T1, function #2 has a function ID of T2, and function #3 has a function ID of T3. This capability information includes the identification information of the processing functions supported by the second agent, that is, the reporting format of the processing function is: [T1, T2, T3], and this capability information includes [T1, T2, T3].

[0137] It should be understood that a detailed description of the second agent's determination processing function can be found below. Figure 5 The description in the text.

[0138] In one possible implementation, the capability information includes modal information supported by the second agent.

[0139] For example, if the modal information that the second intelligent agent supports processing includes images and language, then the reporting format of the modal information can be: [image, speech], and the capability information includes [image, speech]; or, if the modal information that the second intelligent agent supports processing includes text, then the reporting format of the modal information can be: [text], and the capability information includes [text].

[0140] 402, The first intelligent agent sends relevant information about the first data and configuration information of the perception module to the second intelligent agent.

[0141] For example, the first intelligent agent determines the relevant information of the first data and the configuration information of the perception module based on the capability information and task description information of the second intelligent agent, and sends the relevant information of the first data and the configuration information of the perception module to the second intelligent agent.

[0142] It should be understood that the task description information may come from a first intelligent agent, a second intelligent agent, or other devices, and this application does not limit the source of such information. For example, the task description information may include improving the user's throughput, increasing the user's transmission rate, or ensuring the user's communication quality, etc.

[0143] It should also be understood that the relevant information of the first data includes any one or more of the following: the content of the first data, the modal information of the first data, and the processing priority information of the first data.

[0144] It should also be understood that the configuration information of the sensing module includes the identification information of the sensing module and / or a first key. The identification information of the sensing module is used to identify a specific sensing module and establish a connection with that sensing module; the first key is used to establish a connection with the corresponding sensing module.

[0145] In one possible implementation, when the capability information includes processing functions for multimodal data supported by the second agent, the first agent determines relevant information about the first data and configuration information for the perception module based on the task description information and the capability information. Specifically, the relevant information about the first data is determined based on the task description information, and the configuration information for the perception module is determined based on the relevant information about the first data.

[0146] For example, the capability information sent by the second agent to the first agent includes processing functions for multimodal data supported by the second agent. Based on the task description information and the processing functions for multimodal data supported by the second agent, the first agent determines the relevant information of the first data and the configuration information of the perception module corresponding to the first data. The relevant information of the first data also includes the first processing function corresponding to the first data.

[0147] In another possible implementation, when the capability information includes modal information of multimodal data that the second agent supports processing, the first agent determines the relevant information of the first data and the configuration information of the perception module based on the task description information and the capability information.

[0148] For example, the capability information sent by the second intelligent agent to the first intelligent agent includes modal information. The first intelligent agent determines the relevant information of the first data and the configuration information of the perception module corresponding to the first data based on the task description information and the modal information of the multimodal data supported by the second intelligent agent.

[0149] It should be understood that before step 402, such as Figure 4 The method shown may further include:

[0150] The first agent receives a registration request from the perception module. Correspondingly, the perception module sends a registration request to the first agent.

[0151] The registration request information is used to request registration with the first intelligent agent. This registration request information includes relevant information about the data that the perception module supports sensing. This data information includes the data type and / or data content.

[0152] It should be understood that the sensing module can be any one or more sensing modules; this step is described using a sensing module as an example.

[0153] The first intelligent agent sends a registration response message to the perception module. Correspondingly, the perception module receives the registration response message from the first intelligent agent.

[0154] For example, the first intelligent agent receives a registration request from the perception module, and assigns a perception module ID and / or key to the perception module based on the registration request. The perception module ID and key correspond one-to-one with the perception module.

[0155] It should be understood that the perception module can be carried on the same communication device as the first intelligent agent, or on different communication devices.

[0156] It should be understood that, in cases where the relevant information for the first data does not include the first processing function, such as Figure 4 The method may further include step 403:

[0157] 403, The second agent performs function matching on the first data.

[0158] For example, when a second agent receives relevant information about first data from a first agent, and this information does not include a processing function corresponding to the first data, the second agent performs function matching on the first data based on the relevant information to determine the corresponding processing function (e.g., a first processing function). This first processing function is a function used for multimodal data processing of the first data.

[0159] As an example, suppose the relevant information of the first data includes the modal information of the first data, which is "image". The second agent obtains the processing function (e.g., function #4) corresponding to the data type "image" from its own stored function library, and uses function #4 as the function to perform multimodal data processing on the first data.

[0160] As another example, suppose the relevant information of the first data includes the content of the first data, which is a number. The second agent obtains the processing function (e.g., function #5) corresponding to the data type "text" from its own stored function library, and uses function #5 as the function to perform multimodal data processing on the first data.

[0161] It should be understood that step 403 is an optional step. If the relevant information of the first data includes the processing function corresponding to the first data, the second agent does not need to perform function matching on the first data, that is, the second agent does not need to execute the operation in step 403.

[0162] 404, The second intelligent agent establishes a connection with the first perception module in the perception module.

[0163] For example, after the second agent receives the relevant information of the first data and the configuration information of the perception module from the first agent, the second agent selects the corresponding perception module (e.g., the first perception module) in the perception module configuration information and / or the first key to establish a connection.

[0164] As an example, suppose the configuration information of the sensing module includes the ID of the first sensing module, meaning the second agent sends a connection establishment request to the first sensing module based on the first sensing module's ID. Correspondingly, the first sensing module receives the connection establishment request from the second agent and, based on the first sensing module's ID carried in the request, sends a response to the second agent. This response indicates that the connection between the second agent and the first sensing module has been successfully established.

[0165] As another example, suppose the configuration information of the sensing module includes a first key, which is the key corresponding to the first sensing module. That is, the second agent sends a request to the first sensing module to establish a connection using the first key. Accordingly, after receiving the request from the second agent using the first key, the first sensing module sends a response to the second agent, indicating that the connection between the second agent and the first sensing module has been successfully established.

[0166] It should be understood that the first sensing module and the second intelligent agent can be carried on the same communication device or on different communication devices. For example, both the first sensing module and the second intelligent agent can be carried on a network device, meaning the information interaction between the first sensing module and the second intelligent agent can be an interaction within the network device. Alternatively, the first sensing module can be carried on a terminal device, and the second intelligent agent can be carried on a network device, meaning the information interaction between the first sensing module and the second intelligent agent can be an over-the-air interaction.

[0167] 405, the second intelligent agent sends a first request message to the first sensing module. Correspondingly, the first sensing module receives the first request message from the second intelligent agent.

[0168] For example, after the second agent successfully establishes a connection with the first sensing module, the second agent sends a first request message to the first sensing module based on the relevant information of the first data obtained in step 402. The first request message is used to request the acquisition of the first data.

[0169] As an example, the second intelligent agent sends a first request message to the first perception module according to the input format of the first processing function corresponding to the first data. For example, the first request message may include any one or more of the following fields: data content description, data type, and data format. The data content description may be, for example, an environmental map, a channel, etc.; the data type may be, for example, an image, video, text, audio, or numbers, etc.; and the data format may be, for example, png, jpg, pdf, txt, etc.

[0170] 406, the second intelligent agent receives the first data from the first sensing module. Correspondingly, the first sensing module sends the first data to the second intelligent agent.

[0171] For example, after the first perception module receives the first request information from the second intelligent agent, the first perception module perceives and obtains the first data based on the first request information, and sends the first data to the second intelligent agent.

[0172] 407. The second intelligent agent performs multimodal data processing on the first data.

[0173] For example, after the second agent receives the first data from the first perception module, the second agent can perform multimodal data processing on the first data according to the processing priority corresponding to the first data and the first processing function corresponding to the first data.

[0174] As an example, suppose the first data includes a user distribution image. The second agent performs multimodal data processing on this user distribution image using a first processing function to obtain user identifiers for clustered areas. The second agent then uses these user identifiers to obtain the user power of the clustered areas. For example, the second agent sends a request to the corresponding user to obtain their user power based on the obtained user identifier.

[0175] It should be understood that after receiving the first data from the first sensing module, the second agent performs multimodal data processing on the first data according to the first processing function corresponding to the first data and the relevant information of the first data. The specific content and steps of the second agent's multimodal data processing of the first data can be found in other reference documents related to multimodal data processing, and will not be elaborated here.

[0176] 408, The second agent sends the processing log to the first agent.

[0177] For example, after the second agent performs multimodal data processing on the first data, the second agent can feed back the processing log to the first agent.

[0178] The processing log includes at least one of the following: first data processing priority, first data processing status, first data processing result, or first data processing error information.

[0179] The first data processing state includes either incomplete processing or completed processing. This first data processing state can be represented by bits. For example, a bit of "0" can represent completed processing, and a bit of "1" can represent completed processing; or, a bit of "1" can represent completed processing, and a bit of "0" can represent completed processing.

[0180] The first data processing result includes the result of multimodal data processing performed on the first data by the second intelligent agent.

[0181] According to the above Figure 4 The method shown involves information exchange between a first intelligent agent and a second intelligent agent. The first agent analyzes multimodal data based on task description information and the capabilities of the second agent, and provides the second agent with relevant information about the multimodal data to be processed (e.g., the first data) and configuration information of the sensing modules. This information exchange between the first and second intelligent agents enables the collection and processing of multimodal data, thereby improving network intelligence.

[0182] In addition, after the second agent completes multimodal data processing of the first data, the second agent will feed back the processing log to the first agent, so that the first agent can adjust the information required for scheduling according to the task plan.

[0183] Figure 5 This is a schematic flowchart illustrating another communication method provided in an embodiment of this application. Figure 5 The method shown may include the following steps:

[0184] 501. The second agent sends a request to the multimodal model library to obtain a multimodal model. Correspondingly, the multimodal model library receives the request from the second agent.

[0185] It should be understood that a multimodal model library can refer to a collection of various multimodal models. The second agent can obtain multimodal models from the multimodal model library through public websites or APIs.

[0186] 502, the second agent receives a multimodal model from the multimodal model database. Correspondingly, the multimodal model database sends a multimodal model to the second agent.

[0187] For example, the multimodal model library receives a request from a second agent to obtain a multimodal model, and sends the multimodal model to the second agent based on the request.

[0188] 503, the function of the second agent to decompose multimodal models.

[0189] For example, the second agent receives a multimodal model from a multimodal model library and decomposes the functionality of the multimodal model.

[0190] 504, Robust function for second agent recognition.

[0191] For example, the second agent decomposes the functionality of the multimodal model and identifies the robust functionality of the multimodal model.

[0192] Robustness is a property of a function, and a robust function indicates stable / robust performance. This robust function ensures the effectiveness of the multimodal model in subsequent use. For example, when the multimodal model is a large visual language model, the robust function of this large visual language model includes finding the coordinates of a specified (or given) object in an image.

[0193] 505, The second agent functions the robust functionality.

[0194] For example, the second agent can functionalize robust functions. The specific functionalization format may include one or more of the following fields: the parameter format of the function call, modal and prompter fields, and the function output and modality.

[0195] 506. The second agent registers functions and updates the function library.

[0196] For example, after the second agent functionalizes the robust function, it registers the resulting function and updates it to the function library so that the function can be used later.

[0197] It should be understood that the fields stored in the function library include one or more of the following: function ID, parameter format of the called function, mode and prompt fields, function output and modality, or function description information. The function description information includes at least one of the following: a functional description of the function, function ID, function registration date, and function registration source.

[0198] It should be understood that the above-mentioned function library can be stored in a storage unit of the second intelligent agent, or in other storage units, and this application does not limit this.

[0199] The above Figure 5 The diagram illustrates how a second agent decomposes the functions of a multimodal model, functionalizes the resulting robust functions, registers and updates these functionalized functions to a function library for subsequent use in processing multimodal data. This method ensures the robustness of the second agent in processing multimodal data based on the processing functions, thereby improving the robustness of multimodal data processing in intelligent networks.

[0200] It should be understood that the above Figure 4 and Figure 5 In the method shown, the first intelligent agent (e.g., C-Agent) and the second intelligent agent (e.g., E-Agent) can be carried on a terminal device, a network device, a core network device, or any other communication device, or the first intelligent agent and the second intelligent agent can be carried on the same communication device at the same time. This application does not limit this.

[0201] Based on the above Figure 4 and Figure 5 The method shown below, combined with Figure 6 The method provided in this application is described by way of example.

[0202] See Figure 6 ,exist Figure 6In the example shown, the method provided in this application embodiment is exemplarily described with the example of a first intelligent agent (e.g., C-Agent) deployed on the core network device side and a second intelligent agent (e.g., E-Agent) deployed on the network device side, and the task request being initiated by the network device.

[0203] 601. The network device sends task description information to the C-Agent. Correspondingly, the C-Agent receives the task description information from the network device.

[0204] Assume the task request is initiated by a network device, which sends a task description to the C-Agent. For example, this task description might include improving the throughput of aggregated users.

[0205] 602. The E-Agent sends capability information to the C-Agent. Correspondingly, the C-Agent receives the capability information from the E-Agent.

[0206] In one possible implementation, the capability information may include processing functions that the E-Agent supports for handling multimodal data, such as a power control function (hereinafter referred to as function #1) and a function for finding the coordinates of a given object in an image (hereinafter referred to as function #0).

[0207] In another possible implementation, the capability information includes modal information of the multimodal data that the E-Agent supports processing, such as [numbers, images].

[0208] 603, the perception module requests to register with C-Agent.

[0209] It should be understood that step 603 is related to the above. Figure 4 The process of the perception module requesting registration with the C-Agent is similar; please refer to the above for details. Figure 4 The details are as follows. In this embodiment, the sensing module is described using Sensor#0 and Sensor#1 as examples.

[0210] It should also be understood that the sensing module can be carried on the same communication device as the C-Agent and / or the E-Agent, or on different communication devices. For example, the sensing module and the C-Agent can be carried on the same network device, meaning the information exchange between the sensing module and the C-Agent can be an internal interaction within that network device. Alternatively, the sensing module and the C-Agent can be carried on different devices, such as the sensing module being carried on a terminal device and the C-Agent being carried on a network device. In this case, the registration request information of the sensing module can be transmitted to the C-Agent via an air interface channel, enabling the sensing module to register with the C-Agent.

[0211] 604. C-Agent sends information related to the first data and the configuration information of the sensing module to E-Agent. Correspondingly, E-Agent receives information related to the first data and the configuration information of the sensing module from C-Agent.

[0212] For example, after receiving the task description information and capability information, the C-Agent analyzes the multimodal data requirements and processing order based on the task description information, which includes improving the throughput of aggregated users, and determines the relevant information of the first data (e.g., data #0 and data #1) and the configuration information of the sensing module.

[0213] The relevant information for the first data may include information related to data #0 and data #1. The relevant information for data #0 includes: user distribution information, image, priority 0, and function #0; the relevant information for data #1 includes: user power information, number, priority 1, and function #1. The configuration information for the sensing module includes the configuration information for the sensing module corresponding to data #0 and the configuration information for the sensing module corresponding to data #1. For example, the configuration information for the sensing module corresponding to data #0 includes: sensing module ID #0 or key #0; the configuration information for the sensing module corresponding to data #1 includes: sensing module ID #1 or key #1.

[0214] As an example, the information related to the first data sent by C-Agent to E-Agent and the configuration information of the sensing module can be represented as follows: user distribution information + image + priority 0 + function #0 + sensing module ID #0 / key #0; user power information + number + priority 1 + function #1 + sensing module ID #1 / key #1.

[0215] As another example, the E-Agent receives relevant information about the first data from the C-Agent. If the relevant information about the first data does not include the processing function, that is, the form in which the C-Agent sends the relevant information about the first data and the configuration information of the sensing module to the E-Agent can be: user distribution information + image + priority 0 + sensing module ID#0 / key#0; user power information + number + priority 1 + sensing module ID#1 / key#1.

[0216] It should be understood that, in cases where the information related to the first data sent by C-Agent to E-Agent does not include the processing function corresponding to the first data, such as... Figure 6 As shown, the method may further include:

[0217] 605. E-Agent performs function matching on the first data to determine the processing function corresponding to the first data.

[0218] For example, E-Agent receives relevant information about the first data from C-Agent, and E-Agent determines the processing function corresponding to the first data based on the relevant information about the first data.

[0219] As an example, E-Agent uses the information in the first data, including the content of data #0 as "user distribution information" and the multimodal information of data #0 as "image", as well as its own stored function library to obtain the processing function (e.g., function #0) corresponding to data #0, and uses function #0 as the function to perform multimodal data processing on user distribution and images.

[0220] As an example, E-Agent uses the information in the first data, including the content of data #1 as "user power information" and the multimodal information of data #1 as "digital", as well as its own stored function library to obtain the processing function (e.g., function #1) corresponding to data #1, and uses function #1 as the function to perform multimodal data processing on user distribution and images.

[0221] 606, E-Agent establishes a connection with the perception module.

[0222] For example, after receiving the perception module configuration information from the C-Agent, the E-Agent determines, based on the content of the perception module configuration information, to establish a connection with the perception module corresponding to the perception module with ID#0 or key#0 (e.g., sensor#0), and to establish a connection with the perception module corresponding to the perception module with ID#1 or key#1 (e.g., sensor#1).

[0223] 607. The E-Agent sends request information #0 to the sensing module sensor #0. Correspondingly, the sensing module sensor #0 receives request information #0 from the E-Agent.

[0224] For example, after the E-Agent establishes a connection with the sensing module, the E-Agent sends request information #0 to the sensing module sensor #0. This request information #0 requests the data information corresponding to data #0 (e.g., user distribution information). This request information #0 may include the following fields: user distribution, image, and jpg. These fields respectively correspond to the content description, data type, and data format of the data requested by request information #0.

[0225] 608, the sensing module sensor#0 sends data#0 to the E-Agent. Correspondingly, the E-Agent receives data#0 from the sensing module sensor#0.

[0226] For example, after receiving request information #0 from the E-Agent, the perception module sensor#0 sends data #0 to the E-Agent based on the request information #0. This data #0 includes the fields corresponding to the requested information #0. This data #0 includes a user distribution image, and its format is jpg.

[0227] 609. The E-Agent sends request information #1 to the sensing module sensor #1. Correspondingly, the sensing module sensor #1 receives request information #1 from the E-Agent.

[0228] For example, after the E-Agent establishes a connection with the sensing module, the E-Agent sends request information #1 to the sensing module sensor #1. This request information #1 requests the acquisition of data information corresponding to data #1 (e.g., user power information). This request information #1 may include the following fields: user power, number, and txt. These fields correspond to the content description, data type, and data format of the data requested by request information #1.

[0229] 610, the sensing module sensor#1 sends data#1 to the E-Agent. Correspondingly, the E-Agent receives data#1 from the sensing module sensor#1.

[0230] For example, after receiving request information #1 from the E-Agent, the sensing module sensor #1 sends data #1 to the E-Agent based on the request information #1. This request information #1 includes data #1 corresponding to the fields listed. This data #1 includes user power, and the format of this data #1 is txt.

[0231] 611, E-Agent performs multimodal data processing on data #0 and data #1.

[0232] For example, E-Agent determines the identifier (ID) of clustered users based on the user distribution image included in data #0. E-Agent determines the power of clustered users based on the user power included in data #1, and adjusts the power of clustered users to improve their throughput.

[0233] 612, E-Agent performs the first operation on the terminal device.

[0234] Assuming the terminal device belongs to a cluster of users, the following explanation uses a terminal device as an example. The E-Agent performs multimodal data processing on data #0 and data #1 to determine the power adjustment needed for the terminal device. For example, the E-Agent performs a first operation on the terminal device. This first operation increases the power of the terminal device, thereby increasing its throughput.

[0235] 613. The E-Agent sends processing logs to the C-Agent. Correspondingly, the C-Agent receives the processing logs from the E-Agent.

[0236] The processing log includes at least one of the following: a first data processing priority, a first data processing status, a first data processing result, and a first data processing error message. For example, the first data includes data #0 and data #1 mentioned above.

[0237] For example, the processing log may include the processing priority of data #0, the processing status of data #0, the processing result of data #0, the processing priority of data #1, the processing status of data #1, and the processing result of data #1.

[0238] Assume that processing priorities are represented using 0 and 1, where a processing priority of 0 has a higher priority than a processing priority of 1, or vice versa. For example, the processing priority for data #0 is 0, and the processing priority for data #1 is 1.

[0239] Assume that the processing status is represented in bits. For example, a bit value of "0" indicates that processing is complete, and a bit value of "1" indicates that processing is not complete, or a bit value of "1" indicates that processing is complete, and a bit value of "0" indicates that processing is not complete. In this case, a bit value of "0" indicates that processing is complete, meaning that the processing status for data #0 and data #1 is both "0".

[0240] Assume that the processing result of the first data represents the result of E-Agent performing multimodal data processing on the first data. Specifically, the processing result corresponding to data #0 includes the ID of the aggregated user, and the processing result corresponding to data #1 includes the adjusted power level of the aggregated user (e.g., the power level after E-Agent performs the first operation on the terminal device).

[0241] It should be understood that C-Agent can also be deployed on the network device side, and E-Agent can also be deployed on the terminal device side. Task requests can be initiated by either the terminal device or the network device. The specific process for deploying C-Agent on the network device side and E-Agent on the terminal device side is the same as described above. Figure 6 The examples in [the document] are similar, and will not be repeated here.

[0242] It should be understood that the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0243] It should also be understood that, in the various embodiments of this application, unless otherwise specified or in case of logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced by each other, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.

[0244] It should also be understood that the above embodiments are mainly illustrated using devices in existing network architectures as examples. It should be understood that the specific form of the device is not limited in the embodiments of this application. For example, any device that can achieve the same function in the future is applicable to the embodiments of this application.

[0245] It is understood that in the above-described method embodiments, the methods and operations implemented by the device (such as the first intelligent agent and the second intelligent agent) can also be implemented by components (such as chips or circuits) that can be used in the device. In the embodiments of this application, the first intelligent agent and the second intelligent agent can be carried on the same device, that is, the device can implement the methods and operations implemented by the first intelligent agent and the second intelligent agent, or it can be implemented by parts (such as chips or circuits) that can be used in the device.

[0246] It is also understood that some optional features in the various embodiments of this application may, in certain scenarios, be independent of other features, or may be combined with other features in other scenarios, without limitation. Furthermore, simple modifications to the embodiments of this application are also within the protection scope of this application.

[0247] The above, combined with Figures 4 to 6 The communication method provided in the embodiments of this application is described in detail. The above communication method is mainly introduced from the perspective of a first intelligent agent and a second intelligent agent. It is understood that, in order to achieve the above functions, the first intelligent agent and the second intelligent agent include hardware structures and / or software modules corresponding to the execution of each function.

[0248] Those skilled in the art will recognize that, based on the units and algorithm steps described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is implemented in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0249] The following, combined with Figure 7 and Figure 8 This application provides a detailed description of the communication device provided in its embodiments. It should be understood that the descriptions of the device embodiments correspond to the descriptions of the method embodiments; therefore, any content not described in detail can be found in the above method embodiments. For brevity, some content is omitted.

[0250] This application embodiment can divide the transmitting or receiving device into functional modules according to the above method examples. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation. The following description uses the division of functional modules according to each function as an example.

[0251] Figure 7 This is a schematic block diagram of a communication device 10 provided in an embodiment of this application. The communication device 10 includes a transceiver module 11 and a processing module 12. The transceiver module 11 can implement corresponding communication functions, and the processing module 12 is used for data processing. In other words, the transceiver module 11 is used to perform receiving and sending related operations, and the processing module 12 is used to perform other operations besides receiving and sending. The transceiver module 11 can also be referred to as a communication interface or a communication unit. The transceiver module 11 may include a receiving module and / or a sending module, whereby the receiving module performs receiving-related operations and the sending module performs sending-related operations.

[0252] Optionally, the communication device 10 may further include a storage module 13, which may be used to store computer programs or instructions and / or data. The processing module 12 may read the computer programs or instructions and / or data in the storage module so that the device can perform the operation of the device in the aforementioned method embodiments. The above modules may also be referred to as units, such as transceiver units, processing units, storage units, etc.

[0253] In one design, the communication device 10 may correspond to the first intelligent agent in the above method embodiments, or to a component of the first intelligent agent (such as a chip).

[0254] The communication device 10 can implement the steps or processes corresponding to those executed by the first intelligent agent in the above method embodiments. The transceiver module 11 can be used to execute the transceiver-related operations of the first intelligent agent in the above method embodiments, and the processing module 12 can be used to execute the processing-related operations of the first intelligent agent in the above method embodiments.

[0255] In one possible implementation, the transceiver module 11 is used to receive capability information from the second intelligent agent, the capability information including processing functions and / or modal information of multimodal data supported by the second intelligent agent; the transceiver module 11 is also used to send relevant information of the first data and configuration information of the perception module, the relevant information of the first data is determined according to the capability information and task description information, the task description information corresponds to the first data, the perception module corresponds to the first data, and the configuration information of the perception module includes the identification information and / or the first key of the perception module, wherein the multimodal data includes the first data, and the relevant information of the first data includes: the content of the first data, the modal information of the first data, and the processing priority information of the first data.

[0256] Optionally, the communication device 10 further includes a processing module 12, which is used to perform other operations in the communication device 10 besides receiving and / or sending.

[0257] When the communication device 10 is used to perform Figure 4 When the method is in use, the transceiver module 11 can be used to execute the steps of sending and receiving information in the method, such as steps 401, 402 and 408; the processing module 12 can be used to execute the processing steps in the method.

[0258] When the communication device 10 is used to perform Figure 6 When the method is in use, the sending and receiving module 11 can be used to execute the steps of sending and receiving information in the method, such as steps 601, 602, 604, and 613; the processing module 12 can be used to execute the processing steps in the method.

[0259] It should be understood that the specific process of each unit performing the above-mentioned corresponding steps has been described in detail in the above method embodiments, and will not be repeated here for the sake of brevity.

[0260] In another design, the communication device 10 may correspond to the second intelligent agent in the above method embodiment, or to a component of the second intelligent agent (such as a chip).

[0261] The communication device 10 can implement the steps or processes corresponding to those executed by the second intelligent agent in the above method embodiments. The transceiver module 11 can be used to perform the transceiver-related operations of the second intelligent agent in the above method embodiments, and the processing module 12 can be used to perform the processing-related operations of the second intelligent agent in the above method embodiments.

[0262] In one possible implementation, the transceiver module 11 is used to send capability information of the second intelligent agent, including processing functions and / or modal information of multimodal data that the second intelligent agent supports processing; the transceiver module 11 is also used to receive relevant information of the first data and configuration information of the perception module, wherein the relevant information of the first data is determined based on the capability information and task description information, the task description information corresponds to the first data, the perception module corresponds to the first data, and the configuration information of the perception module includes the identification information and / or the first key of the perception module, wherein the multimodal data includes the first data, and the relevant information of the first data includes: the content of the first data, the modal information of the first data, and the processing priority information of the first data.

[0263] Optionally, the communication device 10 further includes a processing module 12, which is used to perform other operations in the communication device 10 besides receiving and / or sending.

[0264] When the communication device 10 is used to perform Figure 4 When the method is in use, the transceiver module 11 can be used to execute the steps of sending and receiving information in the method, such as steps 401, 402, 404, 405, 406, and 408; the processing module 12 can be used to execute the processing steps in the method, such as steps 403 and 407.

[0265] When the communication device 10 is used to perform Figure 5 When the method is in use, the sending and receiving module 11 can be used to execute the steps of sending and receiving information in the method, such as steps 501 and 502; the processing module 12 can be used to execute the processing steps in the method, such as steps 503, 504, 505 and 506.

[0266] When the communication device 10 is used to perform Figure 6 When the method is in use, the transceiver module 11 can be used to execute the steps of sending and receiving information in the method, such as steps 601, 602, 604, 606, 607, 608, 609, 610, 612, and 613; the processing module 12 can be used to execute the processing steps in the method, such as steps 605 and 611.

[0267] It should be understood that the specific process by which each module or unit performs the above-mentioned corresponding steps has been described in detail in the above method embodiments, and will not be repeated here for the sake of brevity.

[0268] It should also be understood that the communication device 10 here is embodied in the form of a functional module. The term "module" here can refer to application-specific integrated circuits (ASICs), electronic circuits, processors (e.g., shared processors, proprietary processors, or group processors, etc.) and memories for executing one or more software or firmware programs, combined logic circuits, and / or other suitable components supporting the described functions. In an alternative example, those skilled in the art will understand that device 10 may specifically be a mobility management network element in the above embodiments, and may be used to execute the various processes and / or steps corresponding to the mobility management network element in the above method embodiments; or, device 10 may specifically be a terminal device in the above embodiments, and may be used to execute the various processes and / or steps corresponding to the terminal device in the above method embodiments. To avoid repetition, further details are omitted here.

[0269] The communication device 10 of each of the above schemes has the function of implementing the corresponding steps performed by the devices (such as the first intelligent agent and the second intelligent agent) in the above methods. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions; for example, the transceiver module can be replaced by a transceiver (for example, the sending module in the transceiver module can be replaced by a transmitter, and the receiving module in the transceiver module can be replaced by a receiver), and other units, such as processing modules, can be replaced by processors, which respectively execute the transmission and reception operations and related processing operations in each method embodiment.

[0270] In addition, the transceiver module 11 can also be a transceiver circuit (for example, it may include a receiving circuit and a transmitting circuit), and the processing module can be a processing circuit.

[0271] Figure 8 This is a schematic diagram of another communication device 20 provided in an embodiment of this application. The communication device 20 includes a processor 21, which is used to execute computer programs or instructions stored in a memory 22, or to read data / signaling stored in the memory 22, to perform the methods in the above-described method embodiments. Optionally, there may be one or more processors 21.

[0272] Optionally, such as Figure 8 As shown, the communication device 20 also includes a transceiver 23, which is used for receiving and / or transmitting signals. For example, the processor 21 is used to control the transceiver 23 to receive and / or transmit signals. The transceiver 23 may include a receiver and / or a transmitter, the receiver being used for receiving signals and the transmitter for transmitting signals; if the communication device 20 is a chip, then the transceiver 23 is the chip's input / output interface, where the output corresponds to transmitting and the input corresponds to receiving.

[0273] Optionally, such as Figure 8 As shown, the communication device 20 also includes a memory 22 for storing computer programs or instructions and / or data. The memory 22 may be integrated with the processor 21 or may be separately configured. Optionally, there may be one or more memories 22.

[0274] As one approach, the communication device 20 is used to implement operations performed by, for example, the first intelligent agent and the second intelligent agent in the various method embodiments described above.

[0275] It should be understood that the processor mentioned in the embodiments of this application can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0276] It should also be understood that the memory mentioned in the embodiments of this application can be volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM). For example, RAM can be used as an external cache. By way of example and not limitation, RAM includes the following forms: static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).

[0277] It should be noted that when the processor is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, the memory (storage module) can be integrated into the processor.

[0278] It should also be noted that the memory described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0279] This application provides a chip system. The chip system (or processing system) includes logic circuits and an input / output interface.

[0280] The logic circuit can be a processing circuit in the chip system. The logic circuit can be coupled to a memory cell, calling instructions from the memory cell, enabling the chip system to implement the methods and functions of the embodiments of this application. The input / output interface can be an input / output circuit in the chip system, outputting processed information or inputting data or signaling information to be processed into the chip system for processing.

[0281] As one approach, the chip system is used to implement the operations performed by, for example, the first intelligent agent and the second intelligent agent in the various method embodiments described above.

[0282] For example, the logic circuit is used to implement the processing-related operations performed by the first intelligent agent and the second intelligent agent in the above method embodiments; the input / output interface is used to implement the sending and / or receiving-related operations performed by the first intelligent agent and the second intelligent agent in the above method embodiments.

[0283] This application also provides a computer-readable storage medium storing computer instructions for implementing the methods executed by the first intelligent agent and the second intelligent agent in the above-described method embodiments.

[0284] For example, when the computer program is executed by the computer, it enables the computer to implement the methods executed by the first intelligent agent and the second intelligent agent in the various embodiments of the above methods.

[0285] This application also provides a computer program product comprising a computer program or instructions which, when executed by a computer, implement the methods performed by the first intelligent agent and the second intelligent agent in the above-described method embodiments.

[0286] This application also provides a communication system, including the aforementioned first intelligent agent and second intelligent agent.

[0287] The explanations and beneficial effects of the relevant contents in any of the devices provided above can be found in the corresponding method embodiments provided above, and will not be repeated here.

[0288] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0289] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0290] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0291] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0292] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0293] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0294] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A communication method, characterized in that, include: The second agent sends its capability information to the first agent, the capability information including processing functions and / or modal information for multimodal data that the second agent supports processing; The first intelligent agent determines the relevant information of the first data and the configuration information of the perception module based on the task description information and the capability information, wherein the task description information corresponds to the first data and the perception module corresponds to the first data; The first intelligent agent sends relevant information about the first data and configuration information of the perception module to the second intelligent agent. The configuration information of the perception module includes the identification information of the perception module and / or a first key. The multimodal data includes the first data, and the relevant information of the first data includes: the content of the first data, the modal information of the first data, and the processing priority information of the first data.

2. A communication method, characterized in that, Applied to the first intelligent agent, including: Receive capability information from a second intelligent agent, the capability information including processing functions and / or modal information for multimodal data that the second intelligent agent supports processing; The system transmits relevant information about the first data and configuration information about the sensing module. The relevant information about the first data is determined based on the capability information and task description information. The task description information corresponds to the first data, and the sensing module corresponds to the first data. The configuration information of the sensing module includes its identification information and / or a first key. The multimodal data includes the first data, and the relevant information of the first data includes: the content of the first data, the modal information of the first data, and the processing priority information of the first data.

3. The method according to claim 2, characterized in that, In the case where the processing function includes a first processing function corresponding to the first data, the relevant information of the first data also includes the first processing function.

4. The method according to claim 2 or 3, characterized in that, Before sending the relevant information of the first data and the configuration information of the sensing module, the method further includes: The system receives a registration request from the sensing module. The registration request includes the data types and data content that the sensing module can sense. The configuration information of the sensing module is determined based on the registration request and related information of the first data.

5. A communication method, characterized in that, Applied to second intelligent agents, including: Send capability information of the second agent, the capability information including processing functions and / or modal information of multimodal data that the second agent supports processing; The system receives relevant information about the first data and configuration information about the sensing module. The relevant information about the first data is determined based on the capability information and task description information. The task description information corresponds to the first data, and the sensing module corresponds to the first data. The configuration information of the sensing module includes its identification information and / or a first key. The multimodal data includes the first data, and the relevant information of the first data includes: the content of the first data, the modal information of the first data, and the processing priority information of the first data.

6. The method according to claim 5, characterized in that, The method further includes: A connection is established with the first sensing module in the sensing module based on the identification information of the sensing module and / or the first key; Send a request message to the first sensing module, the request message being used to request the acquisition of the first data; Receive the first data from the first sensing module.

7. The method according to claim 6, characterized in that, The request information includes at least one of the following: content description information of the first data, data type of the first data, or data format of the first data.

8. The method according to any one of claims 5 to 7, characterized in that, In the case where the processing function includes a first processing function corresponding to the first data, the relevant information of the first data also includes the first processing function.

9. The method according to any one of claims 5 to 7, characterized in that, When the capability information includes the modal information, the method further includes: Based on the relevant information of the first data, function matching is performed on the first data to determine the first processing function.

10. The method according to claim 8 or 9, characterized in that, The method further includes: According to the first processing function, the first data is subjected to multimodal data processing; Send a processing log, which includes at least one of the following: the processing priority of the first data, the processing completion status of the first data, or the processing result of the first data.

11. The method according to claim 10, characterized in that, The processing completion status of the first data includes: processing not completed or processing completed.

12. The method according to any one of claims 5 to 11, characterized in that, Before sending the capability information of the second agent, the method further includes: Obtain a multimodal model; The functions of the multimodal model are decomposed to identify robust functions; The processing function for the multimodal data is determined based on the robustness feature.

13. A communication device, characterized in that, include: The method comprises one or more functional modules for performing the method as described in claim 1, or includes one or more functional modules for performing the method as described in any one of claims 2 to 4; or includes one or more functional modules for performing the method as described in any one of claims 5 to 12.

14. A communication device, characterized in that, include: A processor is configured to execute a computer program stored in a memory to cause the apparatus to perform the method as claimed in claim 1; or to cause the apparatus to perform the method as claimed in any one of claims 2 to 4; or to cause the apparatus to perform the method as claimed in any one of claims 5 to 12.

15. A computer program product, characterized in that, The computer program product includes instructions for performing the method as described in any one of claims 1 to 12.

16. A computer-readable storage medium, characterized in that, include: The computer-readable storage medium stores a computer program or instructions; when the computer program or instructions are executed on a computer, the computer causes the computer to perform the method as described in any one of claims 1 to 12.

17. A chip, characterized in that, The chip is installed in a communication device. The chip includes a processor and a communication interface. The processor reads and runs a computer program or instruction through the communication interface, causing the communication device to perform the method as described in any one of claims 1 to 12.

18. A communication system, characterized in that, It includes a first intelligent agent and a second intelligent agent, wherein the first intelligent agent is used to perform the method as described in any one of claims 2 to 4, and the second intelligent agent is used to perform the sending of capability information of the second intelligent agent and the receiving of relevant information of the first data and the configuration information of the sensing module.