Communication method and communication apparatus
By enabling information interaction and data processing between intelligent agents in intelligent networks, the problem of low efficiency in multimodal data processing is solved, and the intelligence and operational efficiency of the network are improved.
Patent Information
- Application Number
- PCT/CN2025/094179
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-14
- Filing Date
- 2025-05-12
- Publication Date
- 2025-11-20
AI Technical Summary
In intelligent networks designed based on Artificial General Intelligence (AGI), how can we effectively process multimodal data to improve network intelligence and operational efficiency?
Through information interaction between the first and second intelligent agents, capability information, configuration information of the sensing modules, and data-related information are transmitted to realize the analysis and processing of multimodal data, including the registration, connection, and data request of the sensing modules, and data processing is performed using processing functions.
It improves the processing efficiency of multimodal data and network maintenance in intelligent networks, and ensures the accuracy and robustness of data processing.
Smart Images

Figure CN2025094179_20112025_PF_FP_ABST
Abstract
Description
Communication method and communication apparatus
[0001] The present application claims priority to the Chinese patent application No. 202410601980.4, filed on May 14, 2024, and entitled "A communication method and communication apparatus", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the field of communication, in particular to a communication method and communication apparatus. BACKGROUND
[0003] The current field of computer sciences (CS) proposes that any entity that can independently think and interact with the environment can be abstracted as an agent. The basic characteristics of an agent are: it can react to changes in the environment and then automatically adjust its behavior and state, and different agents can interact with other agents according to their own intentions.
[0004] Artificial general intelligence (AGI) will be an indispensable part of network architecture in the future. Therefore, how to implement the processing of multi-modal data by agents in the design of intelligent networks based on AGI is a current research hotspot. SUMMARY
[0005] The present application provides a communication method to implement the processing of multi-modal data by agents in the design of intelligent networks based on AGI.
[0006] In the first aspect, the method can be executed by a first agent. In the absence of special description, the "first agent" in the present application can refer to the agent itself, a component (such as a processor, a chip, or a chip system, etc.) in the first agent, or a logic module or software that can realize all or part of the functions of the first agent. The following will be described taking the first agent as an example.
[0007] The method comprises: a first agent receiving capability information from a second agent, the capability information comprising processing functions and / or modality information of multi-modal data supported by the second agent for processing; and the first agent sending related information of first data and configuration information of a perception module, the related information of the first data being determined according to the capability information and task description information corresponding to the first data, the perception module corresponding to the first data, and the configuration information of the perception module comprising identification information of the perception module and / or a first key.
[0008] It should be understood that the perception module corresponding to the first data can be understood as the perception module being capable of perceiving and acquiring the first data. For example, the modality of the first data is “image”, and the perception module corresponding to the first data is capable of perceiving and acquiring image-related data; for another example, the modality of the first data is “voice”, and the perception module corresponding to the first data is capable of perceiving and acquiring voice-related data.
[0009] According to the method provided in the present application, the interaction between the first agent and the second agent enables the first agent to analyze the multi-modal data based on the task description information and the capability information of the second agent, and to indicate the related information of the multi-modal data (for example, the first data) to be processed and the configuration information of the perception module to the second agent. In the technical solution, the information interaction between the first agent and the second agent realizes the processing of the multi-modal data by the agents, and improves the intelligence of the network. In addition, the introduction of the agents in the intelligent network designed based on the AGI can further improve the efficiency of network maintenance and / or operation.
[0010] In combination with the first aspect, in some possible implementation manners, in a case where the processing function comprises a first processing function corresponding to the first data, the related information of the first data further comprises the first processing function.
[0011] Based on the above technical solution, the capability information comprises the processing functions supported by the second agent, the first agent determines the first processing function corresponding to the first data in the processing functions according to the task description information, and indicates the first processing function to the second agent through the related information of the first data, so that the second agent can directly use the first processing function indicated by the first agent when processing the first data, without the need to match the corresponding processing function according to the related information of the first data by itself.
[0012] In some possible implementation manners, the method further includes: receiving registration request information from the perception module, the registration request information including a data type and a data content that the perception module can perceive, and the configuration information of the perception module is determined according to the registration request information and the related information of the first data.
[0013] It should be understood that the perception module can be carried on the same communication device as the first intelligent entity, or carried on different communication devices. For example, the perception module and the first intelligent entity are both carried on a network device, that is, the information interaction between the perception module and the first intelligent entity can be internal interaction of the network device. For another example, the perception module can be carried on a terminal device, and the first intelligent entity is carried on a network device, that is, the information interaction between the perception module and the first intelligent entity can be information interaction over the air.
[0014] In a second aspect, a communication method is provided, which can be executed by a second intelligent entity. In the case of no special description, the "second intelligent entity" in the present application can refer to the second intelligent entity itself, a component (for example, a processor, a chip, or a chip system) in the second intelligent entity, or a logic module or software capable of realizing all or part of the functions of the second intelligent entity. The following is described taking the second intelligent entity as an example.
[0015] The method includes: the second intelligent entity sending capability information of the second intelligent entity, the capability information including processing function and / or modality information of multi-modal data supported by the second intelligent entity for processing; the second intelligent entity receiving related information of first data and configuration information of a perception module, the related information of the first data being determined according to the capability information and task description information, the task description information corresponding to the first data, the perception module corresponding to the first data, and the configuration information of the perception module including identification information and / or a first key of the perception module, wherein the multi-modal data includes the first data, and the related information of the first data includes content, modality information, and processing priority information of the first data.
[0016] It should be understood that the technical effects of the method of the second aspect and possible designs thereof can refer to the technical effects of the first aspect and possible designs thereof, which will not be described herein.
[0017] With reference to the second aspect, in some possible implementation manners, the method further includes: establishing, by the second agent, a connection with a first perception module in the perception modules according to the identification information of the perception modules and / or the first key; sending, by the second agent, request information to the first perception module, the request information being used to request to obtain the first data; and receiving, by the second agent, the first data from the first perception module.
[0018] With reference to the second aspect, in some possible implementation manners, the request information includes at least one of the following: content description information of the first data, a data type of the first data, or a data format of the first data.
[0019] With reference to the second aspect, in some possible implementation manners, in a case where the processing function includes a first processing function corresponding to the first data, the related information of the first data further includes the first processing function.
[0020] With reference to the second aspect, in some possible implementation manners, in a case where the capability information includes the modality information, the method further includes: performing, by the second agent, function matching on the first data according to the related information of the first data, to determine a first processing function.
[0021] It should be understood that, in a case where the capability information includes the modality information, it can be understood that the capability information does not include the processing function.
[0022] Based on the technical solution described above, in a case where the capability information does not include the processing function supported by the second agent, when the second agent receives the related information of the first data and the configuration information of the perception module, the second agent performs function matching on the first data according to the received related information of the first data, to determine a first processing function corresponding to the first data, and performs multi-modal data processing on the first data based on the first processing function.
[0023] With reference to the second aspect, in some possible implementation manners, the method further includes: performing, by the second agent, multi-modal data processing on the first data according to the first processing function; and sending, by the second agent, a processing log, the processing log including at least one of the following: a processing priority of the first data, a processing completion state of the first data, or a processing result of the first data.
[0024] Based on the technical solution described above, after the second agent completes the multi-modal data processing on the first data, the second agent feeds back the processing log to the first agent, so as to facilitate the first agent to adjust information required for scheduling according to the task planning.
[0025] With reference to the second aspect, in some possible implementation manners, the processing completion state of the first data includes: processing not being completed or processing being completed.
[0026] With reference to the second aspect, in some possible implementation manners, before the second agent sends the capability information of the second agent, the method further includes: the second agent obtaining a multi-modal model; the second agent decomposing a function of the multi-modal model and identifying a robust function; and the second agent determining a processing function of the multi-modal data according to the robust function.
[0027] Based on the technical solution described above, the second agent decomposes a function of a multi-modal model, identifies a robust function, and determines a processing function according to the robust function, thereby guaranteeing robustness of multi-modal data processing and improving performance of the second agent in processing multi-modal data.
[0028] In a third aspect, a communication method is provided, which can be performed by a communication apparatus. In the case where no special description is made, the "communication apparatus" in the present application can refer to the communication apparatus itself (for example, a core network device, an access network device, a terminal device, a network device), a component (for example, a processor, a chip, or a chip system) in the communication apparatus, or a logic module or software capable of realizing all or part of the communication apparatus.
[0029] The method includes: a second agent sending capability information of the second agent to a first agent, the capability information including a processing function of multi-modal data supported by the second agent for processing and / or modality information; the first agent determining, according to task description information and the capability information of the second agent, related information of first data and configuration information of a perception module, the task description information corresponding to the first data, and the perception module corresponding to the first data; the first agent sending, to the second agent, the related information of the first data and the configuration information of the perception module, the configuration information of the perception module including identification information of the perception module and / or a first key, wherein the multi-modal data includes the first data, and the related information of the first data includes: content of the first data, modality information of the first data, and processing priority information of the first data.
[0030] The technical effects of the method shown in the third aspect and possible designs thereof can be referred to the technical effects in the first aspect, the second aspect, and possible designs thereof.
[0031] With reference to the third aspect, in some possible implementation manners, in the case where the processing function includes a first processing function corresponding to the first data, the related information of the first data further includes the first processing function.
[0032] In some possible implementation manners, before the sending of the first data related information and the configuration information of the perception module, the method further includes: receiving registration request information from the perception module, the registration request information including a data type and a data content that the perception module can perceive, and the configuration information of the perception module being determined according to the registration request information and the first data related information.
[0033] In some possible implementation manners, the method further includes: establishing a connection with a first perception module in the perception module according to the identification information of the perception module and / or the first key; sending request information to the first perception module, the request information being used to request to obtain the first data; and receiving the first data from the first perception module.
[0034] In some possible implementation manners, the request information includes at least one of the following: content description information of the first data, a data type of the first data, or a data format of the first data.
[0035] In some possible implementation manners, when the capability information includes the modality information, the method further includes: performing function matching on the first data according to the first data related information, to determine a first processing function.
[0036] In some possible implementation manners, the method further includes: performing multi-modal data processing on the first data according to the first processing function; sending a processing log, the processing log including at least one of the following: a processing priority of the first data, a processing completion state of the first data, or a processing result of the first data; and receiving the processing log.
[0037] In some possible implementation manners, the processing completion state of the first data includes: uncompleted processing or completed processing.
[0038] In some possible implementation manners, before the sending of the capability information of the second agent, the method further includes: obtaining a multi-modal model; decomposing a function of the multi-modal model to identify a robust function; and determining a processing function of the multi-modal data according to the robust function.
[0039] In a fourth aspect, a communication apparatus is provided, which is configured to execute the method provided in the first aspect. Specifically, the communication apparatus can include units and / or modules configured to execute the method provided in any of the implementation manners of the first aspect, such as a processing unit and an obtaining unit.
[0040] In an implementation form, the transceiving unit can be a transceiver, or an input / output interface; the processing unit can be at least one processor. Optionally, the transceiver can be a transceiving circuit. Optionally, the input / output interface can be an input / output circuit.
[0041] In another implementation form, the transceiving unit can be an input / output interface, an interface circuit, an output circuit, an input circuit, a pin or related circuitry on the chip, chip system or circuitry; the processing unit can be at least one processor, a processing circuit or a logic circuit.
[0042] In the fifth aspect, a communication apparatus is provided, which is configured to execute the method provided in the second aspect. Specifically, the communication apparatus can include units and / or modules for performing the method provided in the second aspect, such as a processing unit and an obtaining unit.
[0043] In an implementation form, the transceiving unit can be a transceiver, or an input / output interface; the processing unit can be at least one processor. Optionally, the transceiver can be a transceiving circuit. Optionally, the input / output interface can be an input / output circuit.
[0044] In another implementation form, the transceiving unit can be an input / output interface, an interface circuit, an output circuit, an input circuit, a pin or related circuitry on the chip, chip system or circuitry; the processing unit can be at least one processor, a processing circuit or a logic circuit.
[0045] In the sixth aspect, a communication apparatus is provided, which is configured to execute the method provided in the third aspect. Specifically, the communication apparatus can include units and / or modules for performing the method provided in the third aspect, such as a processing unit and an obtaining unit.
[0046] In an implementation form, the transceiving unit can be a transceiver, or an input / output interface; the processing unit can be at least one processor. Optionally, the transceiver can be a transceiving circuit. Optionally, the input / output interface can be an input / output circuit.
[0047] In another implementation form, the transceiving unit can be an input / output interface, an interface circuit, an output circuit, an input circuit, a pin or related circuitry on the chip, chip system or circuitry; the processing unit can be at least one processor, a processing circuit or a logic circuit.
[0048] In the seventh aspect, a processor is provided, which is configured to execute the method provided in any one of the implementation forms of the first to third aspects.
[0049] For the sending and obtaining / receiving operations involved by the processor, if no special description is made, or if it does not contradict with the actual role or internal logic in the related description, it can be understood as the processor output and receiving, input operations, and also can be understood as the sending and receiving operations performed by the radio frequency circuit and the antenna, and the present application does not limit this.
[0050] In an eighth aspect, a computer-readable storage medium storing program code for execution by an apparatus is provided. The program code includes code for performing any of the methods provided by the implementations of the first through third aspects.
[0051] In a ninth aspect, a computer program product containing instructions that, when executed on a computer, cause the computer to perform any of the methods provided by the implementations of the first through third aspects.
[0052] In a tenth aspect, a chip is provided. The chip includes one or more processors and a communication interface. The processor reads a computer program or instructions stored on a memory through the communication interface and executes any of the methods provided by the implementations of the first through third aspects.
[0053] Optionally, as an implementation, the chip further includes a memory. The memory stores the computer program or instructions. The processor is configured to execute the computer program or instructions stored on the memory. When the computer program or instructions are executed, the processor is configured to execute any of the methods provided by the implementations of the first through third aspects.
[0054] In an eleventh aspect, a communication system is provided. The communication system includes the first agent and the second agent. BRIEF DESCRIPTION OF DRAWINGS
[0055] FIG. 1 is a schematic diagram of a communication system suitable for use in the present application.
[0056] FIG. 2 is a schematic diagram of an artificial intelligence (AI) network element built in a communication system.
[0057] FIG. 3 is a schematic diagram of an agent.
[0058] FIG. 4 is a schematic flowchart of a communication method provided by the present application.
[0059] FIG. 5 is a schematic flowchart of another communication method provided by the present application.
[0060] FIG. 6 is a schematic flowchart of another communication method provided by the present application.
[0061] FIG. 7 is a schematic block diagram of a communication apparatus provided by an embodiment of the present application.
[0062] FIG. 8 is a schematic diagram of another communication device according to an embodiment of the present application. DETAILED DESCRIPTION
[0063] In order to facilitate understanding of the embodiments of the present application, the following points are first explained.
[0064] First, in the present application, "for indicating" can include for directly indicating and for indirectly indicating. When describing that certain indication information is for indicating A, it can include that the indication information directly indicates A or indirectly indicates A, and does not mean that A must be carried in the indication information.
[0065] Taking the information indicated by the indication information as to-be-indicated information, in the specific implementation process, there are many ways to indicate the to-be-indicated information, for example but not limited to, the to-be-indicated information can be directly indicated, such as the to-be-indicated information itself or an index of the to-be-indicated information. The to-be-indicated information can also be indirectly indicated by indicating other information, where the other information and the to-be-indicated information have an association relationship. The to-be-indicated information can also be only indicated in part, and the other part of the to-be-indicated information is known or agreed in advance. For example, the indication of a specific information can also be realized by means of the arrangement order of each information agreed in advance (for example, a protocol stipulates), thereby reducing the indication overhead to a certain extent. At the same time, the common part of each information can be identified and uniformly indicated, so as to reduce the indication overhead caused by separately indicating the same information.
[0066] Second, in the present application, "at least one" means one or more, and "multiple" means two or more (including two). In addition, in the embodiments of the present application, "first", "second", and various numerical numbers (for example, "#1", "#2", etc.) are only for distinguishing convenience of description, and do not limit the scope of the embodiments of the present application. The size of the serial number of each process below does not mean the execution order, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. It should be understood that the objects thus described can be interchanged under appropriate circumstances, so as to be able to describe schemes other than the embodiments of the present application. In addition, in the embodiments of the present application, the words such as "S410" are only for distinguishing identification, and do not limit the order of execution steps.
[0067] Third, in the embodiments of the present application, the words "exemplary" or "for example" are used to mean serving as an example, instance, or illustration. Any embodiment or design scheme described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design schemes. Rather, the words "exemplary" or "for example" are used in the sense of presenting a specific example.
[0068] Fourthly, the "storing" in the embodiments of the present application can refer to storing in one or more memories. The one or more memories can be separately arranged or integrated in the encoder or decoder, processor, or communication device. The one or more memories can also be partially separately arranged and partially integrated in the processor or communication device. The type of the memory can be any form of storage medium, which is not limited in the present application.
[0069] Fifthly, in the embodiments of the present application, the "protocol" can refer to a standard protocol in the communication field, which can include the NR protocol and related protocols applied to future communication systems, which is not limited in the present application.
[0070] Sixthly, in the embodiments of the present application, "of", "corresponding", "relevant", "corresponding" and "associate" can be used interchangeably at times. It should be pointed out that the meanings expressed are consistent when the differences are not emphasized.
[0071] Seventhly, in the embodiments of the present application, "in the case of", "when", "if" can be used interchangeably at times. It should be pointed out that the meanings expressed are consistent when the differences are not emphasized.
[0072] Eighthly, the term "and / or" in the present application only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in the present application generally represents an "or" relationship between the associated objects before and after it.
[0073] Ninthly, the "message", "information", or "information element (IE)" in the present application can be used interchangeably, and the name of the message or information is not limited in any way as long as the corresponding function can be realized.
[0074] In the present application, "sending" and "receiving" represent the direction of signal transmission. For example, "sending information to XX" can be understood as that the destination of the information is XX, and "sending information" can include direct sending and indirect sending through other units or modules. "Receiving information from YY" can be understood as that the source of the information is YY, and "receiving information" can include direct receiving from YY and indirect receiving from YY through other units or modules. In addition, "sending" can also be understood as the "output" of the chip interface, and "receiving" can also be understood as the "input" of the chip interface. In other words, "sending" or "receiving" can be carried out between devices, for example, through the air interface between network devices and terminal devices, and "sending" or "receiving" can also be carried out within the device, for example, through the bus, wire or interface between components, modules, chips, software modules or hardware modules within the device.
[0075] The technical solutions in the present application will be described below with reference to the accompanying drawings.
[0076] The technical solutions of the embodiments of the present application can be applied to various communication systems, for example: long term evolution (LTE) system, LTE frequency division duplex (FDD) system, LTE time division duplex (TDD), universal mobile communication system (UMTS), worldwide interoperability for microwave access (WiMAX) communication system, 5th generation (5G) system, new radio (NR), and future communication network, vehicle-to-X (V2X), wherein V2X can include vehicle to network (V2N), vehicle to vehicle (V2V), vehicle to infrastructure (V2I), vehicle to pedestrian (V2P), etc., long term evolution-vehicle (LTE-V), Internet of vehicles, machine type communication (MTC), Internet of things (IoT), long term evolution-machine (LTE-M), machine to machine (M2M), etc.
[0077] In addition, the embodiments of the present application are applicable to homogeneous network and heterogeneous network scenarios, and there is no limitation on transmission points, which can be applicable to systems such as multi-point cooperative transmission between macro base stations and macro base stations, micro base stations and micro base stations, and macro base stations and micro base stations. The embodiments of the present application are applicable to low frequency scenarios (sub 6G) and high frequency scenarios (6G and above), terahertz, optical communication, etc.
[0078] FIG. 1 is a schematic diagram of a communication system applicable to the present application. As shown in FIG. 1, the communication system 100 includes at least one network device, such as at least one of the network devices 111, 112, and 113 shown in FIG. 1; the communication system 100 can also include at least one terminal device, such as at least one of the terminal devices 121 and 122 shown in FIG. 1; the communication system 100 can also include at least one AI network element, such as the AI network element 131 shown in FIG. 1. The network devices and the terminal devices in the communication system can communicate with each other through wireless links, and in turn, exchange information. It can be understood that the network devices and the terminal devices can also be referred to as communication devices or communication apparatuses.
[0079] A network device is a network-side device with wireless transceiving functions. The network device can be a device in a radio access network (RAN) that provides wireless communication functions for terminal devices, referred to as a RAN device. The RAN can be a 3rd Generation Partnership Project (3GPP)-related cellular system, such as a 5G mobile communication system or a future-oriented evolved system (e.g., a future communication network). The RAN can also be an open radio access network (O-RAN or ORAN), a cloud radio access network (CRAN), or a wireless fidelity (WiFi) system. For example, the network device can be a base station, an evolved NodeB (eNodeB), a next-generation NodeB (gNB) in a 5G mobile communication system, a base station in a subsequent evolution of 3GPP, a transmission reception point (TRP), an access node in a WiFi system, a wireless relay node, a wireless backhaul node, etc. In a communication system using different radio access technologies (RATs), the name of the device with the function of a base station can be different. For example, in an LTE system, it can be referred to as an eNB or eNodeB, and in a 5G system or an NR system, it can be referred to as a gNB. The present application does not limit the specific name of the base station. The network device can include one or more co-sited or non-co-sited transmission reception points.
[0080] For another example, the network device can include at least one of one or more central units (CUs), one or more distributed units (DUs), one or more radio units (RUs). The CU (or CU-control plane (CP), CU-user plane (UP)), DU or RU can also have different names in different systems, but the person skilled in the art can understand its meaning. For example, in the ORAN system, the CU can also be referred to as an open CU (O-CU), the DU can also be referred to as an open DU (O-DU), the CU-CP can also be referred to as an open CU-CP (O-CU-CP), the CU-UP can also be referred to as an open CU-UP (O-CU-UP), and the RU can also be referred to as an open RU (O-RU). Any of the CU (or CU-CP, CU-UP), DU and RU in this application can be implemented by a software module, a hardware module, or a combination of a software module and a hardware module. Exemplarily, the functions of the CU can be implemented by one entity or different entities. For example, the functions of the CU are further divided, i.e., the control plane and the user plane are separated and implemented by different entities, which are the control plane CU entity (i.e., the CU-CP entity) and the user plane CU entity (i.e., the CU-UP entity), respectively. The CU-CP entity and the CU-UP entity can be coupled with the DU to jointly complete the functions of the access network device. For example, the CU is responsible for processing non-real-time protocols and services, implementing radio resource control (RRC), and functions of the packet data convergence protocol (PDCP) layer. The DU is responsible for processing physical layer protocols and real-time services, implementing functions of the radio link control (RLC) layer, the media access control (MAC) layer, and the physical (PHY) layer. In this way, the functions of the wireless access network device can be implemented by multiple network function entities. These network function entities can be network elements in hardware devices, or software functions running on dedicated hardware, or virtualized functions instantiated on a platform (e.g., a cloud platform). The network device can also include an active antenna unit (AAU). The AAU implements part of the physical layer processing function, the radio frequency processing function, and the related function of the active antenna.Since the information of the RRC layer will eventually become the information of the PHY layer, or be transformed from the information of the PHY layer, under this architecture, high-layer signaling, such as RRC layer signaling, can also be considered as being sent by the DU, or by the DU+AAU. It can be understood that the network device can be a device including one or more of the CU node, the DU node, and the AAU node. In addition, the CU can be divided into a network device in the RAN, or can be divided into a network device in the core network (CN), which is not limited in the present application. For example, in vehicle to everything (V2X) technology, the access network device can be a road side unit (RSU). The plurality of access network devices in the communication system can be the same type of base station, or different types of base stations. The base station can communicate with the terminal device, or can communicate with the terminal device through the relay station. In the embodiments of the present application, the device for realizing the function of the network device can be the network device itself, or can be a device capable of supporting the network device to realize the function, such as a chip system or a combination device or component capable of realizing the function of the access network device, which can be installed in the network device. In the embodiments of the present application, the chip system can be composed of a chip, or can include a chip and other discrete devices.
[0081] The terminal device is a user-side device with wireless transceiving function, which can be a fixed device, a mobile device, a handheld device (such as a mobile phone), a wearable device, a vehicle-mounted device, or a wireless device (such as a communication module, a modem, or a chip system) built into the above devices. The terminal device is used to connect people, things, machines, etc., and can be widely used in various scenarios, such as cellular communication, device-to-device (D2D) communication, V2X communication, machine-to-machine / machine-type communications (M2M / MTC) communication, Internet of Things, virtual reality (VR), augmented reality (AR), industrial control, self-driving, remote medical treatment, smart grid, smart furniture, smart office, smart wear, smart transportation, smart city, unmanned aerial vehicle, robot, etc. Exemplarily, the terminal device can be a handheld terminal in cellular communication, a communication device in D2D, an Internet of Things device in MTC, a monitoring camera in smart transportation and smart city, or a communication device on an unmanned aerial vehicle, etc. The terminal device can be referred to as user equipment (UE), user terminal, user device, user unit, user station, terminal, access terminal, access station, UE station, remote station, mobile device, or wireless communication device, etc. The terminal device can also be a terminal device in an IoT system. IoT is an important part of future information technology development, and its main technical feature is to connect objects through communication technology and network, so as to realize the intelligent network of man-machine interconnection and object-object interconnection. In the embodiments of the present application, IoT technology can achieve massive connection, deep coverage, and terminal power saving through, for example, narrow band (NB) technology. In the embodiments of the present application, the device for realizing the function of the terminal device can be a terminal device, or a device capable of supporting the terminal device to realize the function, such as a chip system or a combination device or component that can realize the function of the terminal device, which can be installed in the terminal device.
[0082] An AI network element can implement part or all of AI-related operations. Among them, the AI network element can also be referred to as an AI node, an AI device, an AI entity, an AI module, an AI model, or an AI unit, etc. The AI model can be considered as a specific method to implement AI functions. The AI model represents the mapping relationship or function between the input and output of the model. The AI function can include one or more of the following: data collection, model training (or model learning), model information publishing, model inference (or model reasoning, reasoning, or prediction, etc.), model monitoring or model verification, or inference result publishing, etc. The AI function can also be referred to as an AI (related) operation, or an AI-related function.
[0083] The AI module is used to implement the corresponding AI function. The AI modules deployed in different network elements can be the same or different. The model of the AI module can implement different functions according to different parameter configurations. The model of the AI module can be configured based on one or more of the following parameters: structure parameters (such as at least one of the number of neural network layers, the width of the neural network, the connection relationship between layers, the weight of neurons, the activation function of neurons, or the bias in the activation function), input parameters (such as the type of input parameters and / or the dimension of input parameters), or output parameters (such as the type of output parameters and / or the dimension of output parameters). Among them, the bias in the activation function can also be referred to as the bias of the neural network.
[0084] One AI module can have one or more models. One model can infer an output, which includes one parameter or multiple parameters. The learning process, training process, or inference process of different models can be deployed in different nodes or devices, or can be deployed in the same node or device.
[0085] Exemplarily, the AI network element can be built-in in a communication system. For example, the AI network element can be an AI module built-in in: an access network device, a core network device, a cloud server, or an operation, administration and maintenance (OAM), to implement AI-related functions. Among them, the core network device includes but is not limited to network elements such as access and mobility management function (AMF), user plane function (UPF), or session management function (SMF). The OAM can be the network management of the core network device and / or the network management of the access network device. Alternatively, the AI network element can also be a network element independently set in the communication system. Optionally, the terminal or the chip built-in in the terminal can also include an AI entity for implementing AI-related functions.
[0086] For ease of understanding, the integration of AI modules and communication systems is briefly introduced in conjunction with FIG. 2. As shown in FIG. 2, the network elements in the communication system are connected through interfaces (such as NG, Xn) or air interfaces. One or more AI modules (only one AI module is shown in FIG. 2 for clarity) are provided in one or more of the network element nodes, such as core network devices, access network nodes (RAN nodes), terminals, or OAM. The access network node can serve as a separate RAN node or can include multiple RAN nodes, such as a CU and a DU. The CU and / or DU can also be provided with one or more AI modules. Optionally, the CU can also be split into a CU-CP and a CU-UP. One or more AI models are provided in the CU-CP and / or CU-UP.
[0087] Network devices, terminal devices, and AI network elements can be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; can also be deployed on water; and can also be deployed on aircraft, balloons, and satellites in the air. The scenarios in which the network devices, terminal devices, and core network devices are located are not limited in the embodiments of the present application.
[0088] To facilitate understanding of the embodiments of the present application, the basic concepts involved in the present application are first described.
[0089] 1. Artificial intelligence (AI): can enable machines to have human intelligence, for example, can enable machines to apply computer hardware and software to simulate certain intelligent behaviors of humans. To achieve artificial intelligence, a machine learning method can be used. In the machine learning method, the machine learns (or trains) a model using training data. The model represents the mapping between the input and the output. The learned model can be used for inference (or prediction), that is, the model can be used to predict the output corresponding to a given input. The output can also be referred to as an inference result (or prediction result).
[0090] 2. Large model technology: a large model refers to a neural network model containing a super large number of parameters (usually more than one billion), which has the following characteristics:
[0091] 1) Huge size: a large model contains tens of billions of parameters, and the model size can reach hundreds of gigabytes (GB) or even larger. Such a huge model size provides a large model with strong expression and learning capabilities.
[0092] 2) Multi-task learning: Large models often learn multiple different natural language processing (NLP) tasks, such as machine translation, text summarization, or question answering systems. This can enable the model to learn more general and generalized language understanding capabilities.
[0093] 3) Strong computing resources: Training large models often requires hundreds or even thousands of graphics processor units (GPUs) and a large amount of time, usually weeks to months. This can accelerate the training process while preserving the capabilities of large models.
[0094] 4) Rich data: Large models require a large amount of data for training, and only a large amount of data can take advantage of the parameter size advantage of large models.
[0095] Large models are widely used in natural language processing and are changing the state of NLP tasks, giving rise to more powerful and intelligent language technology. Large models are an important direction of AI development. At the same time, large models also have excellent performance in various natural language processing tasks, such as text classification, sentiment analysis, summary generation, or translation. In addition, large models can also be used in automatic writing, chat robots, virtual assistants, voice assistants, or automatic translation in multiple application fields.
[0096] It should be understood that the application of large models in the network requires a series of supporting functions to truly realize their potential. This system engineering can be called AI Agent. The AI Agent is briefly described below.
[0097] 3、Agent: is a concept in the field of artificial intelligence, any entity that can think independently and interact with the environment can be abstracted as an agent. The basic characteristics of the agent are: the agent can respond to changes in the environment and then automatically adjust its behavior and state, and different agents can interact with other agents according to their intentions. The agent can belong to a kind of AI module.
[0098] The agent in this application is generally considered to be an agent that can independently complete the set target through action ability. "Agent" is inseparable from "intelligence"; the agent has some human-like intelligent capabilities and behaviors, such as learning, reasoning, decision-making, and execution capabilities. Alternatively, the agent can be replaced by other terms, such as artificial general intelligence (AGI), artificial intelligence, or intelligent agent.
[0099] Optionally, in the communication device, the agent can be integrated into the existing hardware / software of the communication device, or the agent can be independent of the existing hardware / software of the communication device. For example, the existing hardware / software can include a chip, a baseband chip, a modem chip, a system on chip (SoC) chip containing a modem core, a system in package (SIP) chip, a communication module, a chip system, a processor, a logic module, or software, etc.
[0100] For example, the agent can use a large language model (LLM) as the core, including a memory module, a tool module, a planning module, and an action module, etc. The memory module is used to realize long-term memory and / or short-term memory functions; the tool module contains multiple callable external tools; the planning module contains multiple planning algorithms, for example, the agent can plan the task inputted by the external according to the content in the memory; and the action module supports the agent to make actions according to the planning results, such as calling tools, etc.
[0101] As shown in FIG. 3, in the LLM-supported autonomous agent system, the LLM acts as the brain of the agent (or agent), and is supplemented by several key components:
[0102] 1) Planning, including but not limited to:
[0103] Subgoal decomposition: the agent decomposes large tasks into smaller, manageable subgoals, enabling effective handling of complex tasks. For example, through a chain of thoughts (CoT) indicating the model to “think step by step”, more test time computation is utilized to break down difficult tasks into smaller, simpler steps. CoT transforms large tasks into multiple manageable tasks and elucidates the explanation of the model's thought process.
[0104] Reflection and refinement: the agent can self-criticize and reflect on past behavior, learn from mistakes, and refine future steps to improve the quality of the final result.
[0105] 2) Memory, including but not limited to:
[0106] Short-term memory: using the short-term memory of the model to learn.
[0107] Long-term memory: provides the agent with the ability to retain and recall (infinite) information for a long time, usually by utilizing external vector storage and fast retrieval.
[0108] 3) tools, including but not limited to:
[0109] The agent learns to call external application programming interfaces (APIs) to obtain additional information missing in the model weights (usually difficult to change after pre-training), including current information, code execution capabilities, access to proprietary information sources, etc.
[0110] 4) action, the model performs a specific task and records the results.
[0111] It should be understood that the agent in the present application can also be referred to as an AI controller, an intelligent unit, or an intelligent entity, etc. The name of the agent in the present application is not limited in any way, as long as it can achieve the corresponding function.
[0112] Exemplarily, the agent includes an AI model, a tool management unit, and a data management unit, wherein the tool management unit includes tools called by the AI model, and the data management unit is used to manage data related to the operation of the agent.
[0113] Optionally, the AI model can be understood as the core of the agent, such as an LLM core. The AI model is used to coordinate other components (such as the tool management unit, the data management unit, etc.) based on task requirements to achieve planning and scheduling of tasks. The name of the AI model in the present application is not limited in any way, as long as it can achieve the corresponding function.
[0114] Optionally, the tool management unit contains tools that can be called by the AI model, such as code compiler, interpreter, performance monitoring, digital twin, ray tracing, or network function virtualization, etc. to enable the first agent to achieve the corresponding function.
[0115] Optionally, the data management unit can be used to manage, for example, device operation logs, agent operation logs, domain knowledge, device supported functions, device perception measured states, etc.
[0116] The above description of the terms is only for the convenience of understanding and does not limit the protection scope of the embodiments of the present application.
[0117] The above briefly introduces the scenario to which the communication method provided by the embodiments of the present application can be applied in combination with FIG. 1 or FIG. 2, and introduces the basic concepts that can be involved in the embodiments of the present application, and introduces the concept of an intelligent agent in the basic concepts. As can be known from the above, the intelligent agent can implement multiple functions and has a high degree of intelligence.
[0118] It should also be understood that the above is an exemplary example of the intelligent agent architecture provided by the present application, and the specific architecture of the intelligent agent is not limited by the present application.
[0119] 4. Artificial general intelligence (AGI)
[0120] AGI refers to an artificial intelligence system that can think, learn and perform multiple tasks like a human being. AGI is considered as a higher level of artificial intelligence and is an important direction and goal of the development of current artificial intelligence technology.
[0121] In AGI, the controlling agent (C-Agent) and the decision-making agent (E-Agent) are two important concepts.
[0122] C-Agent refers to the main body of the intelligent system that controls the system, mainly responsible for decision-making and guiding the behavior of the entire system. C-Agent can be understood as a high-level decision maker that can develop appropriate strategies and plans according to changes in the environment and requirements of the target, and guide E-Agent to perform corresponding tasks. C-Agent usually has the ability to learn, reason and make decisions, and can make flexible decisions according to different situations.
[0123] E-Agent refers to an intelligent agent that performs specific tasks and is the executor of C-Agent. E-Agent can be a robot, a virtual character or other forms of entities. E-Agent perceives the environment, collects information, and performs corresponding actions according to the instructions of C-Agent. E-Agent usually has the ability to perceive, move and interact, and can interact and feedback with the environment in real time.
[0124] As can be seen, C-Agent is the main body of the intelligent system that controls the system, responsible for decision-making and guiding the behavior of the entire system. E-Agent is an intelligent agent that performs specific tasks, and perceives the environment and performs corresponding actions according to the instructions of C-Agent.
[0125] 5. Multimodal:
[0126] A modality is a way in which things happen or occur, or a way in which something is expressed or perceived. Each source or form of information can be referred to as a modality. For example, a person has the sense of touch, hearing, vision, smell; the medium of information has voice, video, text, etc.; a variety of sensors, such as radar, infrared, accelerometer, etc., each of which can be referred to as a modality. Compared with the form of multimedia data division such as image, voice, text, etc., "modality" is a more fine-grained concept, and different modalities can exist under the same medium. For example, two different languages can be considered as two different modalities, and even the data collected in two different situations can also be considered as two modalities.
[0127] A multimodal is a way in which something is expressed or perceived by multiple modalities. Multimodal can be classified into homogeneous modalities and heterogeneous modalities. Among them, the homogeneous modalities are, for example, two photos taken by two cameras respectively. The heterogeneous modalities are, for example, the relationship between a picture and a text language.
[0128] A large multimodal model (LMM) is a type of artificial intelligence model that can process and understand multiple different types of data inputs, such as text, images, audio, and video. LMM can integrate and understand different data formats. For example, LMM can analyze a news article (text), related photos (images), and related video (video) clips to obtain a more comprehensive understanding.
[0129] Currently, LMM replaces LLM as the brain of an intelligent agent to process multimodal data. Since LLM is generally used for evaluation indicators mainly focusing on language understanding and generation tasks, such as evaluating LLM through indicators such as fluency, coherence, and relevance. LMM is mainly applied to more extensive indicator evaluation, and LMM needs to be proficient in multiple fields, such as evaluating LMM through indicators such as image recognition accuracy, audio processing quality, and the ability of the model to integrate cross-modal information. With the continuous development of science and technology, LMM replaces LLM as the brain of an intelligent agent to process multimodal data, and becomes a hot research topic.
[0130] The above briefly introduces the scenario in which the communication method provided by the embodiments of the present application can be applied in combination with FIG. 1, and introduces the basic concepts that can be involved in the embodiments of the present application, and introduces the concept of an intelligent agent in the basic concepts. As can be known from the above, the intelligent agent can implement multiple functions and has a high degree of intelligence.
[0131] The present application provides a communication method, which can be applied in the communication system shown in FIG. 1, so as to establish a connection between an intelligent agent and a device in a communication network and improve the intelligence degree of the communication network.
[0132] It should be understood that the embodiments shown below do not particularly limit the specific structure of the execution subject of the method provided by the embodiments of the present application, as long as the execution subject can communicate according to the method provided by the embodiments of the present application by running a program in which the code of the method provided by the embodiments of the present application is recorded. For example, the execution subject of the method provided by the embodiments of the present application can be a device, or a functional module in the device that can call and execute a program.
[0133] FIG. 4 is a schematic flowchart of a communication method provided by the present application. The method includes the following steps:
[0134] 401, the second agent sends capability information to the first agent. Correspondingly, the first agent receives the capability information from the second agent.
[0135] The capability information includes processing functions and / or modality information of the multi-modal data supported by the second agent.
[0136] In a possible implementation, the capability information includes processing functions of the multi-modal data supported by the second agent. The processing functions can be processing functions obtained by the second agent by decomposing the multi-modal model functions and determining the decomposed functions. The capability information can include the processing functions, or include the identification (ID) corresponding to the processing functions.
[0137] For example, the processing functions of the multi-modal supported by the second agent are function #1, function #2 and function #3, wherein the function ID corresponding to function #1 is T1, the function ID corresponding to function #2 is T2, and the function ID corresponding to function #2 is T3. The capability information includes the identification information of the processing functions supported by the second agent, that is, the reporting format of the processing functions is [T1, T2, T3], and the capability information includes [T1, T2, T3].
[0138] It should be understood that the detailed description of the determination of the processing functions by the second agent can be referred to the description in FIG. 5 below.
[0139] In a possible implementation, the capability information includes modality information supported by the second agent.
[0140] For example, the modality information supported by the second agent for processing includes pictures and language, and the reporting format of the modality information can be [picture, voice], and the capability information includes [picture, voice]; or the modality information supported by the second agent for processing includes text, and the reporting format of the modality information can be [text], and the capability information includes [text].
[0141] 402, the first agent sends the related information of the first data and the configuration information of the perception module to the second agent.
[0142] For example, the first agent determines the related information of the first data and the configuration information of the perception module according to the capability information of the second agent and the task description information, and sends the related information of the first data and the configuration information of the perception module to the second agent.
[0143] It should be understood that the task description information can come from the first agent, or from the second agent, or from other devices, which are not limited by the present application. For example, the task description information can include improving the throughput of the user, improving the transmission rate of the user, or guaranteeing the communication quality of the user, etc.
[0144] It should also be understood that the related information of the first data includes any one or more of the content of the first data, the modal information of the first data and the processing priority information of the first data.
[0145] It should also be understood that the configuration information of the perception module includes the identification information of the perception module and / or the first key. The identification information of the perception module is used to determine a specific perception module and establish a connection with the perception module; the first key is used to establish a connection with the corresponding perception module.
[0146] In one possible implementation, when the capability information includes the processing function of the multi-modal data supported by the second agent, the first agent determines the related information of the first data and the configuration information of the perception module according to the task description information and the capability information. The related information of the first data is determined according to the task description information, and the configuration information of the perception module is determined according to the related information of the first data.
[0147] For example, the capability information sent by the second agent to the first agent includes the processing function of the multi-modal data supported by the second agent, and the first agent determines the related information of the first data and the configuration information of the perception module corresponding to the first data according to the task description information and the processing function of the multi-modal data supported by the second agent. The related information of the first data further includes the first processing function corresponding to the first data.
[0148] In another possible implementation, when the capability information includes the modal information of the multi-modal data processed by the second agent, the first agent determines the related information of the first data and the configuration information of the perception module according to the task description information and the capability information.
[0149] For example, the capability information sent by the second agent to the first agent includes modality information, and the first agent determines the related information of the first data and the configuration information of the perception module corresponding to the first data according to the task description information and the modality information of the multi-modal data supported by the second agent.
[0150] It should be understood that, before step 402, the method shown in FIG. 4 can further include:
[0151] The first agent receives the registration request information from the perception module. Accordingly, the perception module sends the registration request information to the first agent.
[0152] The registration request information is used to request registration to the first agent, and the registration request information includes the related information of the data perceived by the perception module. The related information of the data includes the data type and / or the data content.
[0153] It should be understood that the perception module can be any one or more perception modules, and the perception module is described as an example in this step.
[0154] The first agent sends the registration response information to the perception module. Accordingly, the perception module receives the registration response information from the first agent.
[0155] For example, the first agent receives the registration request information from the perception module, and the first agent allocates the perception module ID and / or the key to the perception module according to the registration request information. The perception module ID and the key correspond to the perception module one by one.
[0156] It should be understood that the perception module can be carried on the same communication device as the first agent or carried on different communication devices.
[0157] It should be understood that, in the case that the related information of the first data does not include the first processing function, the method described in FIG. 4 can further include step 403:
[0158] 403. The second agent performs function matching on the first data.
[0159] For example, the second agent receives the related information of the first data from the first agent, and the related information of the first data does not include the processing function corresponding to the first data. The second agent performs function matching on the first data according to the related information of the first data to determine the processing function (for example, the first processing function) corresponding to the first data. The first processing function is a function for processing the first data.
[0160] As an example, it is assumed that the related information of the first data includes the modality information of the first data, and the modality information of the first data is "picture". The second agent obtains the processing function corresponding to the data type "picture" (for example, function #4) from the function library stored by itself, and takes function #4 as the function for processing the first data in a multi-modal data manner.
[0161] As another example, it is assumed that the related information of the first data includes the content of the first data, and the content of the first data is a number. The second agent obtains the processing function corresponding to the data type "text" (for example, function #5) from the function library stored by itself, and takes function #5 as the function for processing the first data in a multi-modal data manner.
[0162] It should be understood that the step 403 is an optional step. In the case where the related information of the first data includes the processing function corresponding to the first data, the second agent does not need to perform function matching on the first data, i.e., the second agent does not need to perform the operation in step 403.
[0163] 404, the second agent establishes a connection with the first perception module in the perception module.
[0164] For example, after the second agent receives the related information of the first data from the first agent and the configuration information of the perception module, the second agent selects the corresponding perception module (for example, the first perception module) in the perception module according to the perception module ID and / or the first key in the configuration information of the perception module, and establishes a connection.
[0165] As an example, it is assumed that the configuration information of the perception module includes the ID of the first perception module, i.e., the second agent sends the request information for establishing a connection to the first perception module according to the ID of the first perception module. Correspondingly, the first perception module receives the request information for establishing a connection from the second agent, and sends response information to the second agent according to the ID of the first perception module carried in the request information, the response information being used to indicate that the connection establishment between the second agent and the first perception module is successful.
[0166] As another example, it is assumed that the configuration information of the perception module includes the first key, which is the key corresponding to the first perception module, i.e., the second agent sends the request information for establishing a connection to the first perception module through the first key. Correspondingly, the first perception module receives the request information for establishing a connection from the second agent through the first key, and sends response information to the second agent, the response information being used to indicate that the connection establishment between the second agent and the first perception module is successful.
[0167] It should be understood that the first perception module can be carried on the same communication device as the second agent, or carried on different communication devices. For example, the first perception module and the second agent are both carried on a network device, that is, the information interaction between the first perception module and the second agent can be an internal interaction of the network device. For another example, the first perception module can be carried on a terminal device, and the second agent is carried on a network device, that is, the information interaction between the first perception module and the second agent can be information interaction over the air.
[0168] 405. The second agent sends first request information to the first perception module. Correspondingly, the first perception module receives the first request information from the second agent.
[0169] For example, after the second agent successfully establishes a connection with the first perception module, the second agent sends first request information to the first perception module according to the related information of the first data obtained in step 402, and the first request information is used to request to obtain the first data.
[0170] As an example, the second agent sends first request information to the first perception module according to the input format of the first processing function corresponding to the first data. For example, the first request information can include any one or more of the following fields: data content description, data type, data format. Among them, the data content description is, for example: environment map, channel, etc.; the data type is, for example: picture, video, text, audio, number, etc.; the data format is, for example: png, jpg, pdf, txt, etc.
[0171] 406. The second agent receives the first data from the first perception module. Correspondingly, the first perception module sends the first data to the second agent.
[0172] For example, after the first perception module receives the first request information from the second agent, the first perception module perceives to obtain the first data according to the first request information, and sends the first data to the second agent.
[0173] 407. The second agent performs multi-modal data processing on the first data.
[0174] For example, after the second agent receives the first data from the first perception module, the second agent can perform multi-modal data processing on the first data according to the first processing function corresponding to the first data according to the processing priority corresponding to the first data.
[0175] As an example, it is assumed that the first data includes a user distribution picture, and the second agent can obtain the user identification of the aggregation area by performing multi-modal data processing on the user distribution picture through the first processing function. It is assumed that the first data includes a user distribution picture, and the second agent can obtain the user identification of the aggregation area by performing multi-modal data processing on the user distribution picture through the first processing function. The second agent obtains the user power of the aggregation area according to the user identification of the aggregation area. For example, the second agent sends request information for obtaining the user power to the corresponding user according to the obtained user identification, so as to obtain the user power.
[0176] It should be understood that after the second agent receives the first data from the first perception module, the second agent performs multi-modal data processing on the first data according to the first processing function corresponding to the first data and the relevant information of the first data. The specific content and related steps of the second agent performing multi-modal data processing on the first data can be referred to other multi-modal data processing related reference files, which will not be described here.
[0177] 408, the second agent sends the processing log to the first agent.
[0178] For example, after the second agent performs multi-modal data processing on the first data, the second agent can feed back the processing log to the first agent.
[0179] The processing log includes at least one of the following: first data processing priority, first data processing state, first data processing result, or first data processing error information.
[0180] The first data processing state includes incomplete processing or completed processing. The first data processing state can be represented by a bit. For example, a bit of "0" can represent that the processing is completed, and a bit of "1" can represent that the processing is completed; or a bit of "1" can represent that the processing is completed, and a bit of "0" can represent that the processing is completed.
[0181] The first data processing result includes the result of the multi-modal data processing performed by the second agent on the first data.
[0182] According to the method shown in FIG. 4, the information interaction between the first agent and the second agent enables the first agent to analyze the multi-modal data based on the task description information and the capability information of the second agent, and to indicate the related information of the multi-modal data (for example, the first data) to be processed and the configuration information of the perception module to the second agent. The information interaction between the first agent and the second agent realizes the collection and processing of multi-modal data between agents, thereby improving the network intelligence.
[0183] In addition, after the second agent completes the multi-modal data processing on the first data, the second agent feeds back the processing log to the first agent, so as to facilitate the first agent to adjust the information required for scheduling according to the task planning.
[0184] FIG. 5 is a schematic flowchart of another communication method provided by an embodiment of the present application. As shown in FIG. 5, the method can include the following steps:
[0185] 501. The second agent sends request information for requesting to obtain a multi-modal model to a multi-modal model library. Accordingly, the multi-modal model library receives the request information for requesting to obtain the multi-modal model from the second agent.
[0186] It should be understood that the multi-modal model library can refer to a collection of various multi-modal models. Among them, the second agent can obtain the multi-modal model in the multi-modal model library through a public website or API, etc.
[0187] 502. The second agent receives the multi-modal model from the multi-modal model library. Accordingly, the multi-modal model database sends the multi-modal model to the second agent.
[0188] For example, the multi-modal model library receives the request information for requesting to obtain the multi-modal model from the second agent, and sends the multi-modal model to the second agent based on the request information.
[0189] 503. The second agent decomposes the function of the multi-modal model.
[0190] For example, the second agent receives the multi-modal model from the multi-modal model library, and decomposes the function of the multi-modal model.
[0191] 504. The second agent identifies a robust function.
[0192] For example, the second agent decomposes the function of the multi-modal model, and identifies a robust function of the multi-modal model.
[0193] Among them, robustness is a property of a function, and the robust function is used to indicate stable / robust performance. The robust function can guarantee the subsequent use effect of the multi-modal model. For example, when the multi-modal model is a visual language large model, the robust function of the visual language large model includes finding the coordinates of a specified (or called given) object in a picture.
[0194] 505. The second agent functions the robust function.
[0195] For example, the second agent functions the robust function, and the specific function format can include one or more fields of the following: parameter format of calling function, modality and prompter field, output and modality of function.
[0196] 506, the second agent registers the function and updates the function library.
[0197] For example, after the second agent functions the robust function, the obtained function is registered and updated in the function library, so as to facilitate subsequent use of the function.
[0198] It should be understood that the fields saved in the function library include one or more of the following: function ID, parameter format for calling function, mode and prompt field, output and mode of function, or description information of function. The description information of function includes at least one of the following: function description, function function ID, registration date of function, and source of function registration.
[0199] It should be understood that the function library described above can be stored in a unit with storage function in the second agent, or stored in other units with storage function, which is not limited by the present application.
[0200] The above Figure 5 shows that the second agent decomposes the function of the multi-modal model, functions the robust function obtained by decomposition, registers and updates the function after the robust function is functioned to the function library, so as to facilitate subsequent use of the multi-modal data processing. This method ensures the robustness of the second agent based on the processing function for processing the multi-modal data, and improves the robustness of the multi-modal data processing in the intelligent network.
[0201] It should be understood that in the method shown in the above Figure 4 and Figure 5, the first agent (such as C-Agent) and the second agent (such as E-Agent) can be carried on a terminal device, a network device, a core network device or any other communication device, or the first agent and the second agent can also be carried on the same communication device at the same time, which is not limited by the present application.
[0202] Based on the method shown in the above Figure 4 and Figure 5, the method provided by the present application will be introduced exemplarily in combination with Figure 6.
[0203] Referring to Figure 6, in the example shown in Figure 6, the first agent (such as C-Agent) is deployed on the core network device side, the second agent (such as E-Agent) is deployed on the network device side, and the task request is initiated by the network device. As an example, the method provided by the embodiment of the present application is introduced exemplarily.
[0204] 601, the network device sends task description information to C-Agent. Correspondingly, C-Agent receives task description information from the network device.
[0205] It is assumed that the task request is initiated by a network device, i.e., the network device sends the task description information to the C-Agent. For example, the task description information includes improving the throughput of the aggregated users.
[0206] 602. The E-Agent sends the capability information to the C-Agent. Correspondingly, the C-Agent receives the capability information from the E-Agent.
[0207] In one possible implementation, the capability information can include the processing functions supported by the E-Agent for processing the multi-modal data, such as a power control function (hereinafter referred to as function #1) and a function for finding the coordinates of a given object in a picture (hereinafter referred to as function #0).
[0208] In another possible implementation, the capability information includes the modality information of the multi-modal data supported by the E-Agent for processing, such as [digital, picture].
[0209] 603. The perception module requests registration to the C-Agent.
[0210] It should be understood that step 603 is similar to the perception module requesting registration to the C-Agent in FIG. 4 described above, and specific details can be referred to the detailed description in FIG. 4 described above. In the embodiments of the present application, the perception module is taken as an example of Sensor #0 and Sensor #1 for description.
[0211] It should also be understood that the perception module can be carried on the same communication device as the C-Agent and / or the E-Agent, or carried on different communication devices. For example, the perception module and the C-Agent are carried on the same network device, i.e., the information interaction between the perception module and the C-Agent can be internal interaction of the network device. For another example, the perception module and the C-Agent are carried on different devices, such as: the perception module can be carried on a terminal device, and the C-Agent is carried on a network device, i.e., the registration request information of the perception module can be transmitted to the C-Agent through an air interface channel to realize the registration of the perception module to the C-Agent.
[0212] 604. The C-Agent sends the related information of the first data and the configuration information of the perception module to the E-Agent. Correspondingly, the E-Agent receives the related information of the first data and the configuration information of the perception module from the C-Agent.
[0213] For example, after the C-Agent receives the task description information and the capability information, the C-Agent analyzes the multi-modal data requirements and processing order based on the improvement of the throughput of the aggregated users included in the task description information, and determines the related information of the first data (such as data #0 and data #1) and the configuration information of the perception module.
[0214] The related information of the first data can include related information of data #0 and related information of data #1. The related information of data #0 includes user distribution information, a picture, priority 0, and function #0; and the related information of data #1 includes user power information, a number, priority 1, and function #1. The configuration information of the perception module includes configuration information of a perception module corresponding to data #0 and configuration information of a perception module corresponding to data #1. For example, the configuration information of the perception module corresponding to data #0 includes perception module ID #0 or key #0; and the configuration information of the perception module corresponding to data #1 includes perception module ID #1 or key #1.
[0215] As an example, the C-Agent can send the related information of the first data and the configuration information of the perception module to the E-Agent in the form of: user distribution information + picture + priority 0 + function #0 + perception module ID #0 / key #0; and user power information + number + priority 1 + function #1 + perception module ID #1 / key #1.
[0216] As another example, the E-Agent receives the related information of the first data from the C-Agent, and in the case that the related information of the first data does not include a processing function, the C-Agent can send the related information of the first data and the configuration information of the perception module to the E-Agent in the form of: user distribution information + picture + priority 0 + perception module ID #0 / key #0; and user power information + number + priority 1 + perception module ID #1 / key #1.
[0217] It should be understood that, in the case that the C-Agent does not include the processing function corresponding to the first data in the related information of the first data sent to the E-Agent, the method can further include:
[0218] 605. The E-Agent performs function matching on the first data to determine the processing function corresponding to the first data.
[0219] For example, the E-Agent receives the related information of the first data from the C-Agent, and determines the processing function corresponding to the first data according to the related information of the first data.
[0220] As an example, the E-Agent determines the processing function (for example, function #0) corresponding to data #0 from a function library stored by the E-Agent according to the content of data #0 included in the related information of the first data, which is “user distribution information”, and the multi-modal information of data #0, which is “picture”, and takes function #0 as a function for multi-modal data processing of user distribution and picture.
[0221] As an example, the E-Agent acquires a processing function (e.g., function #1) corresponding to the data #1 from the function library stored by itself according to the content of the data #1 being "user distribution information" and the multi-modal information of the data #1 being "number" included in the related information of the first data, and takes the function #1 as a function for multi-modal data processing of the user distribution and the picture.
[0222] 606, the E-Agent establishes a connection with the perception module.
[0223] For example, after the E-Agent receives the perception module configuration information from the C-Agent, the E-Agent establishes a connection with the perception module (e.g., sensor #0) corresponding to the ID of the perception module being ID #0 or the key being key #0, and establishes a connection with the perception module (e.g., sensor #1) corresponding to the ID of the perception module being ID #1 or the key being key #1 according to the content in the perception module configuration information.
[0224] 607, the E-Agent sends request information #0 to the perception module sensor #0. Correspondingly, the perception module sensor #0 receives the request information #0 from the E-Agent.
[0225] For example, after the E-Agent establishes a connection with the perception module, the E-Agent sends the request information #0 to the perception module sensor #0, and the request information #0 is used to request to acquire data information (e.g., user distribution information) corresponding to the data #0. The request information #0 can include the following fields: user distribution, picture, and jpg. The above fields correspond to the content description, data type, and data format of the data requested to be acquired by the request information #0 respectively.
[0226] 608, the perception module sensor #0 sends data #0 to the E-Agent. Correspondingly, the E-Agent receives the data #0 from the perception module sensor #0.
[0227] For example, after the perception module sensor #0 receives the request information #0 from the E-Agent, the perception module sensor #0 sends the data #0 to the E-Agent according to the request information #0. The data #0 corresponds to the fields included in the request information #0. The data #0 includes a user distribution picture, and the format of the data #0 is jpg.
[0228] 609, the E-Agent sends request information #1 to the perception module sensor #1. Correspondingly, the perception module sensor #1 receives the request information #1 from the E-Agent.
[0229] For example, after the E-Agent establishes a connection with the perception module, the E-Agent sends a request information #1 to the perception module sensor #1, and the request information #1 is used to request to obtain data information (for example, user power information) corresponding to data #1. The request information #1 can include the following fields: user power, number and txt. The above fields correspond to the content description, data type and data format of the data requested by the request information #1 to obtain the data respectively.
[0230] 610, the perception module sensor #1 sends data #1 to the E-Agent. Correspondingly, the E-Agent receives the data #1 from the perception module sensor #1.
[0231] For example, after the perception module sensor #1 receives the request information #1 from the E-Agent, the perception module sensor #1 sends the data #1 to the E-Agent according to the request information #1. The data #1 corresponds to the fields included in the request information #1. The data #1 includes user power, and the format of the data #0 is txt.
[0232] 611, the E-Agent performs multi-modal data processing on the data #0 and the data #1.
[0233] For example, the E-Agent determines the identification (ID) of the aggregated user according to the user distribution picture included in the data #0. The E-Agent determines the power of the aggregated user according to the user power included in the data #1, and adjusts the power of the aggregated user to improve the throughput of the aggregated user.
[0234] 612, the E-Agent performs a first operation on the terminal device.
[0235] Assuming that the terminal device belongs to the aggregated user, the terminal device is taken as an example for introduction below. The E-Agent performs multi-modal data processing on the data #0 and the data #1 to determine to perform power adjustment on the terminal device. For example, the E-Agent performs a first operation on the terminal device. The first operation is used to improve the power of the terminal device and improve the throughput of the terminal device.
[0236] 613, the E-Agent sends a processing log to the C-Agent. Correspondingly, the C-Agent receives the processing log from the E-Agent.
[0237] The processing log includes at least one of the following: a first data processing priority, a first data processing state, a first data processing result, and a first data processing error information. For example, the first data includes the above-mentioned data #0 and data #1.
[0238] For example, the processing log can include a processing priority of data #0, a processing status of data #0, a processing result of data #0, a processing priority of data #1, a processing status of data #1, and a processing result of data #1.
[0239] For example, the processing priority of data #0 is 0, and the processing priority of data #1 is 1.
[0240] For example, the processing status of data #0 is 0, and the processing status of data #1 is 0.
[0241] For example, the processing result of data #0 includes the ID of the aggregated user, and the processing result of data #1 includes the adjusted power size of the aggregated user (for example, the power size after the E-Agent performs the first operation on the terminal device).
[0242] It should be understood that the C-Agent can also be deployed on the network device side, and the E-Agent can also be deployed on the terminal device side. The task request can be initiated by the terminal device or the network device. The specific process of deploying the C-Agent on the network device side and deploying the E-Agent on the terminal device side is similar to the example in FIG. 6, and will not be described here.
[0243] It should be understood that the size of the serial number of the above processes does not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0244] It should also be understood that in various embodiments of the present application, the terms and / or descriptions of different embodiments are consistent and can be mutually referred to if there is no special description and logical conflict. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.
[0245] It should also be understood that in some of the above embodiments, the existing network architecture is mainly taken as an example for exemplary description, and it should be understood that the specific form of the device is not limited in the embodiments of the present application. For example, devices that can achieve the same function in the future are also applicable to the embodiments of the present application.
[0246] It can be understood that the methods and operations implemented by the devices (e.g., the first agent and the second agent) in the above method embodiments can also be implemented by components (e.g., chips or circuits) available to the devices. The first agent and the second agent in the embodiments of the present application can be carried on the same device, that is, the device can implement the methods and operations implemented by the first agent and the second agent, or can be implemented by parts (e.g., chips or circuits) available to the device.
[0247] It can also be understood that some optional features in the embodiments of the present application can not depend on other features in some scenarios, or can be combined with other features in some scenarios, without limitation. In addition, simple modifications of the embodiments of the present application are also within the scope of protection of the present application.
[0248] The above describes the communication method provided by the embodiments of the present application in detail in combination with FIG. 4 to FIG. 6. The above communication method is mainly introduced from the perspective of the first agent and the second agent. It can be understood that the first agent and the second agent contain corresponding hardware structures and / or software modules for executing various functions in order to implement the above functions.
[0249] Those skilled in the art should appreciate that units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by hardware or a combination of hardware and computer software. Whether a certain function is performed in hardware or computer software driven hardware depends on a specific application and design constraint condition of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but it should not be considered that such implementation is beyond the scope of the present application.
[0250] The following describes the communication apparatus provided by the embodiments of the present application in combination with FIG. 7 and FIG. 8. It should be understood that the description of the apparatus embodiments corresponds to the description of the method embodiments, therefore, the content not described in detail can be referred to the above method embodiments, and some content will not be described again for brevity.
[0251] The embodiments of the present application can divide the function modules of the sending end device or the receiving end device according to the above method examples, for example, each function module can be divided according to each function, or two or more functions can be integrated in one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software function module. It should be noted that the division of the modules in the embodiments of the present application is illustrative, and is only a logical function division. Actual implementation can have another division manner. The following takes dividing each function module according to each function as an example for description.
[0252] FIG. 7 is a schematic block diagram of the communication apparatus 10 according to an embodiment of the present application. The communication apparatus 10 includes a transceiver module 11 and a processing module 12. The transceiver module 11 can implement corresponding communication functions, and the processing module 12 is configured to perform data processing. In other words, the transceiver module 11 is configured to perform receiving and transmitting related operations, and the processing module 12 is configured to perform operations other than receiving and transmitting. The transceiver module 11 can also be referred to as a communication interface or a communication unit. The transceiver module 11 can include a receiving module and / or a transmitting module. The receiving module is configured to perform receiving related operations, and the transmitting module is configured to perform transmitting related operations.
[0253] Optionally, the communication apparatus 10 can further include a storage module 13. The storage module 13 can be configured to store computer programs or instructions and / or data. The processing module 12 can read the computer programs or instructions and / or data in the storage module 13, so that the apparatus implements the actions of the devices in the foregoing various method embodiments. The above modules can also be referred to as units, such as a transceiver unit, a processing unit, a storage unit, and the like.
[0254] In one design, the communication apparatus 10 can correspond to a first intelligent entity in the above method embodiments, or be a component (such as a chip) of the first intelligent entity.
[0255] The communication apparatus 10 can implement the steps or processes performed by the first intelligent entity corresponding to the above method embodiments. The transceiver module 11 can be configured to perform the receiving and transmitting related operations of the first intelligent entity in the above method embodiments, and the processing module 12 can be configured to perform the processing related operations of the first intelligent entity in the above method embodiments.
[0256] In one possible implementation, the transceiver module 11 is configured to receive capability information from a second intelligent entity. The capability information includes processing functions and / or modality information of the second intelligent entity for processing multi-modal data. The transceiver module 11 is further configured to send related information of first data and configuration information of a perception module. The related information of the first data is determined according to the capability information and task description information. The task description information corresponds to the first data. The perception module corresponds to the first data. The configuration information of the perception module includes identification information of the perception module and / or a first key. The multi-modal data includes the first data. The related information of the first data includes content of the first data, modality information of the first data, and processing priority information of the first data.
[0257] Optionally, the communication apparatus 10 further includes the processing module 12. The processing module 12 is configured to perform operations other than receiving and / or transmitting in the communication apparatus 10.
[0258] When the communication apparatus 10 is configured to perform the method in FIG. 4, the transceiver module 11 can be configured to perform the steps of receiving and / or transmitting information in the method, such as steps 401, 402 and 408; the processing module 12 can be configured to perform the processing steps in the method.
[0259] When the communication apparatus 10 is configured to perform the method in FIG. 6, the transceiver module 11 can be configured to perform the steps of receiving and / or transmitting information in the method, such as steps 601, 602, 604, 613; the processing module 12 can be configured to perform the processing steps in the method.
[0260] It should be understood that the specific procedures of the respective units performing the corresponding steps are described in detail in the above method embodiments, and thus are not described herein again for the sake of brevity.
[0261] In another design, the communication apparatus 10 can correspond to the second agent in the above method embodiments, or be a component (e.g., a chip) of the second agent.
[0262] The communication apparatus 10 can implement the steps or procedures performed by the second agent corresponding to the above method embodiments, wherein the transceiver module 11 can be configured to perform the operations related to receiving and / or transmitting of the second agent in the above method embodiments, and the processing module 12 can be configured to perform the operations related to processing of the second agent in the above method embodiments.
[0263] In a possible implementation, the transceiver module 11 is configured to send capability information of the second agent, the capability information comprising processing functions and / or modality information of the second agent supporting processing of multi-modal data; the transceiver module 11 is further configured to receive related information of the first data and configuration information of the perception module, the related information of the first data being determined according to the capability information and task description information, the task description information corresponding to the first data, the perception module corresponding to the first data, and the configuration information of the perception module comprising identification information of the perception module and / or a first key, wherein the multi-modal data comprises the first data, and the related information of the first data comprises content of the first data, modality information of the first data and processing priority information of the first data.
[0264] Optionally, the communication apparatus 10 further comprises a processing module 12, which is configured to perform other operations in the communication apparatus 10 described above except for receiving and / or transmitting.
[0265] When the communication apparatus 10 is configured to perform the method in FIG. 4, the transceiver module 11 can be configured to perform the steps of receiving and / or transmitting information in the method, such as steps 401, 402, 404, 405, 406, 408; the processing module 12 can be configured to perform the processing steps in the method, such as steps 403, 407.
[0266] When the communication apparatus 10 is configured to perform the method in FIG. 5, the transceiver module 11 can be configured to perform the steps of receiving and / or transmitting information in the method, such as steps 501, 502; the processing module 12 can be configured to perform the processing steps in the method, such as steps 503, 504, 505, 506.
[0267] When the communication apparatus 10 is configured to perform the method in FIG. 6, the transceiver module 11 can be configured to perform the steps of receiving and / or transmitting information in the method, such as steps 601, 602, 604, 606, 607, 608, 609, 610, 612, 613; the processing module 12 can be configured to perform the processing steps in the method, such as steps 605, 611.
[0268] It should be understood that the specific processes by which the modules or units perform the corresponding steps described above have been described in detail in the method embodiments described above, and thus will not be described again here for brevity.
[0269] It should also be understood that the communication apparatus 10 herein is embodied in the form of functional modules. The term "module" herein can refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (e.g., shared, dedicated, or group) and memory for executing one or more software or firmware programs, a combinational logic circuit, and / or other suitable components that provide the described functionality. In an alternative example, those skilled in the art can understand that the apparatus 10 can be embodied as the mobility management network element in the embodiments described above, and can be configured to perform the processes and / or steps corresponding to the mobility management network element in the method embodiments described above; or the apparatus 10 can be embodied as the terminal device in the embodiments described above, and can be configured to perform the processes and / or steps corresponding to the terminal device in the method embodiments described above, and thus will not be described again here for brevity.
[0270] The communication apparatus 10 of each of the above-described schemes has the function of performing the corresponding steps performed by the device (e.g., the first agent, the second agent) in the above-described methods. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-described functions; for example, the transceiver module can be replaced by a transceiver (e.g., the transmitting module in the transceiver module can be replaced by a transmitter, and the receiving module in the transceiver module can be replaced by a receiver), and other units, such as the processing module, can be replaced by a processor, which performs the receiving and / or transmitting operations and related processing operations in each of the method embodiments.
[0271] In addition, the transceiver module 11 described above can also be a transceiver circuit (e.g., which can include a receiving circuit and a transmitting circuit), and the processing module can be a processing circuit.
[0272] FIG. 8 is a schematic diagram of another communication apparatus 20 according to an embodiment of the present application. The communication apparatus 20 comprises a processor 21 configured to implement methods according to various embodiments of the present application, for example, by executing instructions stored in a memory 22 or by reading and executing program code stored in the memory 22. Alternatively, the processor 21 can be configured as one or more of various hardware components, such as a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic component, a discrete gate or transistor logic component, discrete hardware components, or any combination thereof.
[0273] Optionally, the communication apparatus 20 further comprises a transceiver 23 configured to receive and / or transmit signals, as shown in FIG. 8. For example, the processor 21 is configured to control the transceiver 23 to receive and / or transmit signals. The transceiver 23 can include a receiver configured to receive signals and / or a transmitter configured to transmit signals. If the communication apparatus 20 is a chip, the transceiver 23 is an input / output interface of the chip, where the output corresponds to transmission and the input corresponds to reception.
[0274] Optionally, the communication apparatus 20 further comprises a memory 22 configured to store program codes and / or instructions and / or data, as shown in FIG. 8. The memory 22 can be integrated with the processor 21 or can be separate from the processor 21. Optionally, the memory 22 is one or more of a buffer, a flash memory, a hard drive, a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a tape, a floppy disk, a CD-ROM, another form of non-transitory computer-readable medium, or any combination thereof.
[0275] As an example, the communication apparatus 20 is configured to implement operations performed by the first agent and / or the second agent in the methods described above.
[0276] It should be understood that the processor mentioned in the embodiments of the present application can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, or the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0277] It should also be understood that the memory referred to in the embodiments of the application can be a volatile memory and / or a non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM). For example, the RAM can be used as an external cache. As an example but not limitation, the RAM includes the following various forms: static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM) and direct rambus RAM (DR RAM).
[0278] It should be noted that when the processor is a general processor, DSP, ASIC, FPGA or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, the memory (storage module) can be integrated in the processor.
[0279] It should also be noted that the memory described herein is intended to include, but not limited to, these and any other suitable type of memory.
[0280] The embodiments of the application provide a chip system. The chip system (or also can be called processing system) includes a logic circuit and an input / output interface.
[0281] Among them, the logic circuit can be a processing circuit in the chip system. The logic circuit can be coupled to the storage unit, call the instruction in the storage unit, so that the chip system can realize the method and function of the embodiments of the application. The input / output interface can be an input / output circuit in the chip system, output the information processed by the chip system, or input the data or signaling information to be processed into the chip system for processing.
[0282] As a solution, the chip system is configured to implement operations performed by the first agent and the second agent in the various method embodiments described above.
[0283] For example, the logic circuit is configured to implement operations related to processing performed by the first agent and the second agent in the various method embodiments described above; and the input / output interface is configured to implement operations related to sending and / or receiving performed by the first agent and the second agent in the various method embodiments described above.
[0284] The embodiments of the present application also provide a computer readable storage medium, having stored thereon computer instructions for implementing the method performed by the first agent and the second agent in the various method embodiments described above.
[0285] For example, the computer program, when executed by a computer, enables the computer to implement the method performed by the first agent and the second agent in the various method embodiments described above.
[0286] The embodiments of the present application also provide a computer program product, comprising a computer program or instructions, which, when executed by a computer, implement the method performed by the first agent and the second agent in the various method embodiments described above.
[0287] The embodiments of the present application also provide a communication system, comprising the first agent and the second agent described above.
[0288] The explanations and beneficial effects of the related contents in any of the apparatuses described above can refer to the corresponding method embodiments described above, and will not be repeated here.
[0289] Those skilled in the art can understand that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0290] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the system, apparatus and unit described above can refer to the corresponding processes in the method embodiments described above, and will not be repeated here.
[0291] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. The division of the units is merely logical function division. There can be other division manners in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0292] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments.
[0293] In addition, each functional unit in the various embodiments of the present application can be integrated into a processing unit, or each unit can be a physically independent unit, or two or more units can be integrated into one unit.
[0294] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0295] The above description is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A communication method characterized by comprising: The method comprises: The second agent sends capability information of the second agent to the first agent, the capability information comprising processing functions and / or modality information of multi-modal data supported by the second agent for processing; The first agent determines, according to task description information and the capability information, related information of first data and configuration information of a perception module corresponding to the first data, the task description information corresponding to the first data; The first agent sends the related information of the first data and the configuration information of the perception module to the second agent, the configuration information of the perception module comprising identification information of the perception module and / or a first key. The multi-modal data comprises the first data, and the related information of the first data comprises content of the first data, modality information of the first data, and processing priority information of the first data.
2. A communication method characterized by comprising: The method is applied to a first agent and comprises: Receiving capability information from a second agent, the capability information comprising processing functions and / or modality information of multi-modal data supported by the second agent for processing; Sending related information of first data and configuration information of a perception module, the related information of the first data being determined according to the capability information and task description information, the task description information corresponding to the first data, the perception module corresponding to the first data, the configuration information of the perception module comprising identification information of the perception module and / or a first key. The multi-modal data comprises the first data, and the related information of the first data comprises content of the first data, modality information of the first data, and processing priority information of the first data.
3. The method of claim 2, wherein, In a case where the processing functions comprise a first processing function corresponding to the first data, the related information of the first data further comprises the first processing function.
4. The method according to claim 2 or 3, characterized in that, Before sending the related information of the first data and the configuration information of the perception module, the method further comprises: Receiving registration request information from the perception module, the registration request information comprising a data type and data content that can be perceived by the perception module, and the configuration information of the perception module being determined according to the registration request information and the related information of the first data.
5. A communication method characterized by comprising: The method is applied to a second agent and comprises: Sending capability information of the second agent, the capability information comprising processing functions and / or modality information of multi-modal data supported by the second agent for processing; Receiving related information of first data and configuration information of a perception module, the related information of the first data being determined according to the capability information and task description information, the task description information corresponding to the first data, the perception module corresponding to the first data, the configuration information of the perception module comprising identification information of the perception module and / or a first key. The multi-modal data comprises the first data, and the related information of the first data comprises content of the first data, modality information of the first data, and processing priority information of the first data.
6. The method of claim 5, wherein, The method further comprises: establishing a connection with a first perception module among the perception modules according to the identification information of the perception module and / or the first key; sending request information to the first perception module, the request information being used for requesting to obtain the first data; receiving the first data from the first perception module.
7. The method of claim 6, wherein, The request information comprises at least one of content description information of the first data, a data type of the first data, or a data format of the first data.
8. The method according to any one of claims 5 to 7, characterized in that, In a case where the processing function comprises a first processing function corresponding to the first data, the related information of the first data further comprises the first processing function.
9. The method according to any one of claims 5 to 7, characterized in that, In a case where the capability information comprises the modality information, the method further comprises: performing function matching on the first data according to the related information of the first data to determine a first processing function.
10. The method according to claim 8 or 9, characterized in that, The method further comprises: performing multi-modal data processing on the first data according to the first processing function; sending a processing log, the processing log comprising at least one of a processing priority of the first data, a processing completion status of the first data, or a processing result of the first data.
11. The method of claim 10, wherein, The processing completion status of the first data comprises: uncompleted processing or completed processing.
12. The method according to any one of claims 5 to 11, characterized in that, Before the sending of the capability information of the second agent, the method further comprises: obtaining a multi-modal model; decomposing a function of the multi-modal model to identify a robust function; determining a processing function of the multi-modal data according to the robust function.
13. A communications device, characterized by comprises: one or more functional modules for performing the method of claim 1, or one or more functional modules for performing the method of any one of claims 2 to 4; or one or more functional modules for performing the method of any one of claims 5 to 12.
14. A communications device, characterized by comprises: a processor for executing a computer program stored in a memory to cause the apparatus to perform the method of claim 1, or to cause the apparatus to perform the method of any one of claims 2 to 4, or to cause the apparatus to perform the method of any one of claims 5 to 12.
15. A computer program product, characterised in that, The computer program product comprises instructions for performing the method of any one of claims 1 to 12.
16. A computer-readable storage medium, characterized in that, comprises: The computer readable storage medium stores a computer program or instructions; the computer program or instructions, when running on a computer, cause the computer to perform the method of any one of claims 1 to 12.
17. A chip, characterized by The chip is installed in a communication apparatus, the chip comprises a processor and a communication interface, and the processor reads and runs the computer program or instructions through the communication interface, so that the communication apparatus performs the method of any one of claims 1 to 12.
18. A communication system, characterized by comprises a first agent and a second agent, the first agent is used to perform the method of any one of claims 2 to 4, and the second agent is used to perform the sending of the capability information of the second agent, the receiving of the related information of the first data, and the configuration information of the perception module.
Citation Information
Patent Citations
Capability reporting method and device and capability determining method and device
CN114430920A
Information processing method, electronic equipment and storage medium
CN117556864A
Channel Fusion for Vision-Language Representation Learning
US20240119713A1
Communication method and apparatus
WO2024061125A1
Cited By
Group agent credible controllable interaction method and device
CN121659328A