An agent network comprising a coordinator and one or more agents and a method for an agent network

CN122767010APending Publication Date: 2026-09-15HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480088120.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-06
Publication Date
2026-09-15

Smart Images

  • Figure CN122767010A_ABST
    Figure CN122767010A_ABST
Patent Text Reader

Abstract

An agent network includes a coordinator and one or more agents, each agent including a sensor and a controller. The controller is configured to receive sensor inputs of a local scene from the sensor, identify features relevant to solving an action, generate a segmentation dataset and a low resolution sensor dataset, and send the segmentation dataset and the low resolution sensor dataset to the coordinator. The controller of the coordinator is configured to receive a first segmentation dataset and a first low resolution sensor dataset, generate a model of a global scene based on a combination of the datasets received from the agents. Further, a query is received from a first agent, a hint is determined for the query based on the global scene model, and the hint is sent to the first agent. The controller of the first agent is configured to receive the hint, reevaluate the sensor data, and perform a desired action.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the field of wireless communications, and more specifically, to an agent network comprising a coordinator and one or more agents. Furthermore, more specifically, this disclosure relates to a method for using an agent network comprising a coordinator and one or more agents. Background Technology

[0002] Advances in wireless communication have played a crucial role in building the capabilities of intelligent agents, particularly in scenarios where real-time decision-making is critical, such as in autonomous vehicles and robotics. Autonomous agents in autonomous vehicles need to perform specific actions (e.g., actions required for safe navigation, industrial assembly, etc.) based on distributed perception data collected in real time from sensors. However, making optimal actions requires a comprehensive understanding of the environment, including the intentions of other active autonomous agents, such as whether another autonomous vehicle intends to change lanes. The fact that autonomous agents only possess partial information about the environment and lack any information about other autonomous agents increases safety risks.

[0003] Existing technologies have made some attempts to enable autonomous vehicles to cooperate effectively in dynamic environments. However, these attempts have failed for many reasons, such as decentralized processing of locally collected data, inability to collect data from mobile sensors in real time, support for heterogeneous data formats and data types, and resource limitations. Therefore, a technical problem exists: how to improve the cooperation and decision-making capabilities of autonomous agents in real-time environments while adapting to dynamic conditions.

[0004] Therefore, in light of the above discussion, it is necessary to overcome the aforementioned drawbacks associated with existing agent networks and methods. Summary of the Invention

[0005] This disclosure provides an agent network including a coordinator and one or more agents. More specifically, this disclosure relates to a method for using the agent network including the coordinator and the one or more agents. This disclosure provides a solution to the existing problem of how to improve the collaborative and decision-making capabilities of autonomous agents in real-time environments while adapting to dynamic conditions. The purpose of this disclosure is to provide a solution that at least partially overcomes the problems encountered in the prior art, and to provide an improved agent network and improved methods for using the agent network, such as by generating a global scene model that can be further used to achieve accurate and reliable decision-making.

[0006] One or more objectives of this disclosure are achieved by means of the solutions provided in the appended independent claims. Advantageous implementations of this disclosure are further defined in the dependent claims.

[0007] In one aspect, this disclosure provides an agent network including a coordinator and one or more agents. Furthermore, each agent includes a sensor and a controller. The controller is configured to: receive sensor input of a local scene from the sensor; identify features related to the solution action; generate a segmentation dataset and a low-resolution sensor dataset; and send the segmentation dataset and the low-resolution sensor dataset to the coordinator. The coordinator includes a controller configured to: receive a first segmentation dataset and a first low-resolution sensor dataset from a first agent. Furthermore, the controller is configured to: receive a second segmentation dataset and a second low-resolution sensor dataset from a second agent; and generate a global scene model based on a combination of the datasets received from the agents. The controller is configured to: receive a query from the first agent; determine a prompt for the query based on the global scene model; and send the prompt to the first agent. Furthermore, the controller of the first agent is also configured to: receive the prompt; re-evaluate the sensor data in conjunction with the received prompt; and execute the desired action.

[0008] Advantageously, the agent network is used to generate a comprehensive global scene model of a dynamic environment based on diverse local views of autonomous agents. Furthermore, the generated global scene model enhances the decision-making capabilities of the one or more agents. Each agent is equipped with sensors and controllers for capturing image data, which can be further analyzed to perform actions. Additionally, the agent network includes a coordinator for combining datasets from one or more agents in the network (i.e., the segmented dataset and the low-resolution sensor dataset) to generate the global scene model. Generating the global scene model enables one or more agents in the agent network to make informed decisions based on a shared understanding of the agent environment, thereby achieving efficient, effective, and coordinated actions. Furthermore, one or more agents in the agent network can freely join or leave the network as needed, making the agent network highly flexible. When needed, the one or more agents in the agent network promote high-bandwidth optimization by sending the data to the coordinator and other agents to minimize the amount of data exchanged within the agent network and improve optimized resource utilization. Furthermore, the privacy of the agent network is enhanced by selectively sending segmented data and low-resolution sensor datasets to the coordinator of the agent network, rather than directly sharing raw data with the coordinator. Additionally, prompts determined based on the global scene model for the query support the agent network's agents in finding additional information, which is then used to efficiently and reliably execute the desired action. Therefore, the agent network provides a scalable, resource-efficient, and reliable solution for accurate and efficient decision-making by the agent network's agents.

[0009] In another aspect, this disclosure provides a method for an agent network including a coordinator and one or more agents, each agent including sensors. Furthermore, the method includes the agents performing the following operations: receiving sensor input of a local scene from the sensors; identifying features related to a solution action; generating a segmentation dataset and a low-resolution sensor dataset; and sending the segmentation dataset and the low-resolution sensor dataset to the coordinator. Additionally, the method includes the coordinator performing the following operations: receiving a first segmentation dataset and a first low-resolution sensor dataset from a first agent; receiving a second segmentation dataset and a second low-resolution sensor dataset from a second agent; generating a model of a global scene based on a combination of the datasets received from the agents; receiving a query from the first agent; and determining a prompt for the query based on the model of the global scene. The prompt is then sent to the first agent, whereby the method further includes the first agent performing the following operations: receiving the prompt; re-evaluating the sensor data in conjunction with the received prompt; and performing a desired action.

[0010] The disclosed method achieves all the advantages and technical effects of agent networks, including a coordinator and one or more agents.

[0011] In another aspect, this disclosure provides an agent in an agent network including a coordinator and one or more agents. Furthermore, each agent includes a sensor and a controller, the controller being configured to: receive sensor input of a local scene from the sensor; identify features related to the solution action; generate a segmentation dataset and a low-resolution sensor dataset; and send the segmentation dataset and the low-resolution sensor dataset to the coordinator. Additionally, the agent is configured to: send a query to the coordinator; receive a prompt for the query based on the global scene; re-evaluate the sensor data in conjunction with the received prompt; and execute the desired action.

[0012] Advantageously, the agent is used to provide dynamic query prompts for interaction with the coordinator, supporting real-time adjustment of the agent's decisions based on received prompts, thereby enhancing its responsiveness and adaptability in complex scenarios within the agent network.

[0013] In another aspect, this disclosure provides a method for agents in an agent network including a coordinator and one or more agents. Furthermore, each agent includes a sensor, and the method includes the agent performing the following operations: receiving sensor input of a local scene from the sensor; identifying features related to a solution action; generating a segmentation dataset and a low-resolution sensor dataset; sending the segmentation dataset and the low-resolution sensor dataset to the coordinator; sending a query to the coordinator; receiving a prompt based on the query and the global scene; re-evaluating the sensor data in conjunction with the received prompt; and performing the desired action.

[0014] The disclosed method realizes all the advantages and technical effects of agent networks, including a coordinator and one or more agents.

[0015] In another aspect, this disclosure provides a coordinator in an agent network comprising one or more agents, each agent including sensors. Furthermore, the coordinator includes a controller configured to: receive a first segmented dataset and a first low-resolution sensor dataset from a first agent; receive a second segmented dataset and a second low-resolution sensor dataset from a second agent; generate a model of a global scene based on a combination of the datasets received from the agents; receive a query from the first agent; determine a prompt for the query based on the global scene model; and send the prompt to the first agent.

[0016] Advantageously, the coordinator is used to generate a comprehensive model of the global scene by combining data from multiple agents, thereby enabling prompt-based interaction and customized guidance for each agent within the network, thus enhancing overall decision-making capabilities.

[0017] In another aspect, this disclosure provides a method for a coordinator in an agent network comprising one or more agents, each agent including sensors, wherein the method includes the coordinator performing the following operations: receiving a first segmented dataset and a first low-resolution sensor dataset from a first agent; receiving a second segmented dataset and a second low-resolution sensor dataset from a second agent; generating a model of a global scene based on a combination of the datasets received from the agents; receiving a query from the first agent; determining a prompt for the query based on the global scene model; and sending the prompt to the first agent.

[0018] The disclosed method achieves all the advantages and technical effects of a coordinator in an agent network that includes a coordinator and one or more agents.

[0019] It should be understood that all of the above implementation methods can be combined.

[0020] It should be noted that all devices, elements, circuits, units, and components described in this application can be implemented in software or hardware elements or any combination thereof. All steps performed by the various entities described in this application, and the functions described for performance by the various entities, are intended to indicate that the respective entities are suitable for or used to perform the corresponding steps and functions. Even in the description of the following specific embodiments, if the specific functions or steps to be performed by an external entity are not reflected in the detailed description of the specific elements of the entity performing the specific steps or functions, it will be apparent to those skilled in the art that these methods and functions can be implemented by the corresponding software or hardware elements or any combination thereof. It should be understood that the features of this disclosure are readily combined in various ways without departing from the scope of this disclosure as defined by the appended claims.

[0021] Additional aspects, advantages, features, and objects of this disclosure will become apparent from the accompanying drawings and the detailed description of illustrative implementations as interpreted in conjunction with the following appended claims. Attached Figure Description

[0022] A better understanding of the above-described invention and the following detailed description of illustrative embodiments can be obtained by reading the accompanying drawings. Exemplary structures of this disclosure are shown in the drawings to illustrate the present disclosure. However, this disclosure is not limited to the specific methods and tools disclosed herein. Furthermore, those skilled in the art will understand that the drawings are not drawn to scale. Where possible, the same elements are represented by the same numbers.

[0023] Embodiments of this disclosure are described below by way of example only with reference to the following accompanying drawings, in which:

[0024] Figure 1 It is a block diagram depicting an agent network including a coordinator and one or more agents according to embodiments of the present disclosure;

[0025] Figure 2A and Figure 2B This is a schematic diagram depicting a flowchart of a method for an agent network including a coordinator and one or more agents according to an embodiment of the present disclosure;

[0026] Figure 3 It is a block diagram depicting agents in an agent network including one or more agents according to embodiments of the present disclosure;

[0027] Figure 4 It is a schematic diagram depicting a flowchart of a method for an agent in an agent network including one or more agents according to an embodiment of the present disclosure;

[0028] Figure 5 This is a block diagram depicting a coordinator in an agent network according to an embodiment of the present disclosure;

[0029] Figure 6 This is a schematic diagram depicting a flowchart of a method for a coordinator in an agent network according to an embodiment of the present disclosure;

[0030] Figure 7A It is a schematic diagram depicting the reconstruction of a global scene model according to an embodiment of the present disclosure;

[0031] Figure 7B It is a schematic diagram depicting a partial view processed by an intelligent agent according to an embodiment of the present disclosure;

[0032] Figure 8 This is a schematic diagram depicting the input processing of an intelligent agent according to an embodiment of the present disclosure;

[0033] Figure 9 This is a schematic diagram illustrating global scene model regeneration using cross-view attention and diffusion according to an embodiment of the present disclosure;

[0034] Figure 10 This is a schematic diagram illustrating the prompting generated by the coordinator according to an embodiment of the present disclosure;

[0035] Figure 11 This is a schematic diagram illustrating scene-aware semantic analysis performed on the agent side according to an embodiment of the present disclosure.

[0036] In the accompanying diagram, underlined numbers indicate the item in which the underlined number is located or the item adjacent to the underlined number, while ununderlined numbers are associated with the item identified by the line that links the ununderlined number to the item. When a number is ununderlined and has an associated arrow, the ununderlined number is used to identify the general item that the arrow points to. Detailed Implementation

[0037] The following detailed description illustrates embodiments of this disclosure and ways in which these embodiments may be implemented. While some modes of implementing this disclosure have been disclosed, those skilled in the art will recognize that other embodiments for implementing or practicing this disclosure may also exist.

[0038] Figure 1 This is a block diagram depicting an agent network including a coordinator and one or more agents according to embodiments of the present disclosure. (See reference...) Figure 1 A. The accompanying figure illustrates an agent network 100 including a coordinator 104, a communication network 106, and one or more agents 102.

[0039] One or more intelligent agents 102 refer to a pool of autonomous intelligent agents capable of perceiving the dynamic environment, processing information, and taking necessary actions. In one implementation, the one or more intelligent agents 102 are autonomous intelligent agents of an autonomous vehicle. Furthermore, the one or more intelligent agents 102 include a first intelligent agent 102A and a second intelligent agent 102B up to an Nth intelligent agent, which are used to collect and process data. Each of the one or more intelligent agents 102 also includes sensors and controllers. For example, the first intelligent agent 102A includes a first controller 108, a first sensor 110, and a first memory 112. Similarly, the second intelligent agent 102B includes a second controller 118, a second sensor 120, and a second memory 122.

[0040] Coordinator 104 is used to generate a global scene model 132 based on a combination of datasets received from the agent. Coordinator 104 includes a third controller 128, a third memory 130 for storing the global scene model 132, and a network interface 134.

[0041] Communication network 106 includes a medium (e.g., a communication channel) through which one or more intelligent agents 102 and coordinator 104 communicate with each other. Examples of communication network 106 may include, but are not limited to, cellular networks (e.g., 2G, 3G, Long-Term Evolution (LTE), 4G, 5G, or 5G New Radio (NR) networks, such as sub-6 GHz, centimeter wave (cmWave), or millimeter wave (mmWave) communication networks), wireless sensor networks (WSN), cloud networks, local area networks (LANs), vehicle-to-network (V2N) networks, metropolitan area networks (MANs), and / or the Internet.

[0042] A first controller 108 of the first agent 102A is used to receive sensor input of the local scene from the first sensor 110. Similarly, a second controller 118 of the second agent 102B is used to receive sensor input of the local scene from the second sensor 120. Furthermore, a third controller 128 of the coordinator 104 is used to generate a global scene model 132 based on a combination of datasets received from one or more agents 102. Examples of the first controller 108, the second controller 118, and the third controller 128 may include, but are not limited to, a central data processing device, a microprocessor, a microcontroller, a complex instruction set computing (CISC) processor, an application-specific integrated circuit (ASIC) processor, a reduced instruction set (RISC) processor, a very long instruction word (VLIW) processor, a central processing unit (CPU), a state machine, a data processing unit, and other processors or circuits.

[0043] The first memory 112 of the first agent 102A is used to store the first segmentation dataset 114 and the first low-resolution sensor dataset 116. Similarly, the second memory 122 of the second agent 102B is used to store the second segmentation dataset 124 and the second low-resolution sensor dataset 126. In addition, the third memory 130 is used to store the global scene model 132. Examples of implementations of the first memory 112, the second memory 122, and the third memory 130 may include, but are not limited to, electrically erasable programmable read-only memory (EEPROM), dynamic random-access memory (DRAM), random access memory (RAM), read-only memory (ROM), hard disk drive (HDD), flash memory, secure digital (SD) card, solid-state drive (SSD), and / or CPU cache memory.

[0044] Network interface 134 may include hardware or software for establishing communication between third controller 128 and third memory 130. Examples of network interface 134 may include, but are not limited to, computer ports, network sockets, network interface controllers (NICs), and any other network interface devices.

[0045] An agent network 100 is provided, comprising a coordinator 104 and one or more agents 102. Furthermore, each of the one or more agents 102 includes a sensor and a controller, the controller being configured to receive sensor input of a local scene, identify features related to the solution action, generate a segmentation dataset and a low-resolution sensor dataset, and send the segmentation dataset and the low-resolution sensor dataset to the coordinator 104. In one example, a first agent 102A includes a first sensor 110 and a first controller 108, the first controller being configured to receive sensor input of the local scene from the first sensor 110. Subsequently, the first controller 108 is configured to identify features related to the solution action and generate a first segmentation dataset 114 and a first low-resolution sensor dataset 116. Furthermore, the first controller 108 is configured to send the first segmentation dataset 114 and the first low-resolution sensor dataset 116 to the coordinator 104. In another example, a second agent 102B includes a second sensor 120 and a second controller 118, the second controller being configured to receive sensor input of the local scene from the second sensor 120. Subsequently, the second controller 118 is used to identify features related to the solving action and generate a second segmentation dataset 124 and a second low-resolution sensor dataset 126. Furthermore, the second controller 118 is used to send the second segmentation dataset 124 and the second low-resolution sensor dataset 126 to the coordinator 104. Additionally, sensor inputs received from each of one or more agents 102 are used to provide a comprehensive and detailed understanding of the dynamic environment in real time, which can be further used to identify objects, patterns, or any other relevant elements in the dynamic environment. Furthermore, the generation of segmentation datasets (e.g., the first segmentation dataset 114 and the second segmentation dataset 124) and low-resolution sensor datasets (i.e., the first low-resolution sensor dataset 116 and the second low-resolution sensor dataset 126) is used to obtain insights into specific aspects of the local scene of the dynamic environment. Furthermore, the generated segmentation datasets and low-resolution sensor datasets are sent to the coordinator 104 to provide a holistic representation of the dynamic environment. According to one embodiment, the sensor is a RADAR sensor, and the sensor data is radar image data. In one implementation, the RADAR sensor is used to use radio waves to detect and determine the distance, velocity, orientation, and other characteristics of objects near the agent. For example, the first sensor 110 of the first agent 102A is a RADAR sensor, and the sensor data is radar image data. Similarly, the second sensor 120 is another RADAR sensor, and the sensor data is another radar image data. Advantageously, the RADAR sensor of each of one or more agents 102 is used to provide improved object detection and object recognition under adverse weather conditions such as rain or fog, which may be challenging for other types of sensors.According to another embodiment, the sensor is an image sensor, and the sensor data is image data. In one implementation, a first sensor 110 of the first agent 102A and a second sensor 120 of the second agent 102B are used to capture snapshots of corresponding local scenes (or local views). Furthermore, a first controller 108 of the first agent 102A and a second controller 118 of the second agent 102B are respectively used to collect the snapshots captured by the first sensor 110 and the second sensor 120. Then, the first controller 108 and the second controller 118 are used to execute a task-aware processing algorithm that highlights information related to the actions that the first agent 102A and the second agent 102B need to take, while suppressing irrelevant details. Furthermore, the task-aware processing algorithm is used to generate segmented images (e.g., Im-Seg) that show detected task-related objects (e.g., cars, pedestrians, street lines, etc.).

[0046] According to one embodiment, the controller of the first agent 102A (i.e., the first controller 108) is further configured to determine that the sensor input is insufficient to solve the action, and in response, to perform full segmentation using a trained neural network. Furthermore, the first controller 108 is configured to filter out segmentation objects irrelevant to the local target of the task and send positional information about the agent's global position and pose, as well as a timestamp indicating the recording time, to the coordinator 104. In one implementation, the first controller 108 of the first agent 102A is configured to perform full segmentation on the input image data using a trained neural network (such as a Mask R-CNN neural network), including filtering out objects irrelevant to the local target and identifying relevant objects as relevant. Target-aware filtering is also used to generate a first segmentation dataset 114 and a second segmentation dataset 124. Additionally, the input image data is compressed using a task-agnostic image compression technique to generate a first low-resolution sensor dataset 116 and a second low-resolution sensor dataset 126. Therefore, a first segmentation dataset 114 and a first low-resolution sensor dataset 116 are generated to provide information about the global position and pose of the first agent 102A, as well as timestamps indicating the time when the corresponding image was captured. Furthermore, a first controller 108 and a second controller 118 are used to execute a task-independent compression algorithm that removes unnecessary information from the perceptual data, generating a low-resolution version (e.g., Im-LowRes) of the input image data. By performing the task-aware processing algorithm and the task-independent compression algorithm on the input image datasets from the sensors (i.e., the first sensor 110 and the second sensor 120), the position embedding and pose of the first agent 102A are enhanced. Moreover, including timestamps indicating the specific time of image capture helps provide a comprehensive and context-rich dataset for the first agent 102A among one or more agents 102 to improve decision-making.

[0047] Furthermore, the coordinator 104 includes a controller (i.e., a third controller 128) for receiving a first segmented dataset 114 and a first low-resolution sensor dataset 116 from a first agent 102A. Subsequently, a second segmented dataset 124 and a second low-resolution sensor dataset 126 are received from a second agent. Furthermore, after receiving the first segmented dataset 114, the first low-resolution sensor dataset 116, the second segmented dataset 124, and the second low-resolution sensor dataset 126, the third controller 128 of the coordinator 104 is used to generate a global scene model 132 based on a combination of datasets received from the agents (i.e., from the first agent 102A and the second agent 102B). For example, the third controller 128 of the coordinator 104 is used to receive processed image data from the agents and combine the processed data with information received from other agents. The received metadata (such as agent-id, timestamp, and location-agent) is used to organize the arrangement of the combined image data. Furthermore, the combined data can be processed continuously by two or more algorithms (such as cross-view attention neural networks, diffuse attention neural networks, etc.) without affecting the scope of this disclosure.

[0048] According to one embodiment, the coordinator's controller (i.e., the third controller 128) is also configured to generate a global scene model 132 based on a combination of datasets received from the agents by running cross-view attention and diffusion. For example, the third controller 128 is configured to generate the global scene model 132 based on a combination of received datasets (such as a first segmented dataset 114, a first low-resolution sensor dataset 116, a second segmented dataset 124, and a second low-resolution sensor dataset 126 received from a first agent 102A and a second agent 102B) by running cross-view attention and diffusion. Cross-view attention enables the association of datasets received from one or more agents 102 to provide a detailed, comprehensive, and nuanced understanding of the dynamic environment. Similarly, diffusion is also used to generate the global scene model 132 by filling in potential information gaps in the data. Therefore, generating the global scene model 132 based on a combination of datasets received from one or more agents 102 by running cross-view attention and diffusion ensures a robust, holistic representation of the dynamic environment, which enhances the decision-making capabilities of the agent network 100 by providing a more refined and accurate global scene model 132.

[0049] Furthermore, the coordinator's controller (i.e., the third controller 128) is also used to receive queries from the first agent 102A, determine prompts for the queries based on the global scene model 132, and send the prompts to the first agent 102A. First, the first agent 102A is used to collect local data and further evaluate whether the local data has sufficient information to perform the desired action. Furthermore, if the local information is insufficient, in this case, the first agent 102A is used to select a query describing the missing information (i.e., a query from a predefined query list). In one example, the received query list for the first agent 102A includes "Can I safely turn right?", "Are there any obstacles ahead?", etc. Subsequently, the first agent 102A is used to send a query to the coordinator 104 including a tuple with agent-id, timestamp, location-agent, and query-id. In one implementation, the query includes information about objects, object attributes and relationships, terrain information, and additional background information. Furthermore, the query is accompanied by locally processed data (e.g., Im-Seg, Im-LowRes, etc.). Subsequently, upon receiving the query and local data from the first agent 102A, the coordinator 104 uses data retrieval techniques to extract all information related to the query and received local data from the global scene model 132 (or global database), such as information about objects, their attributes and relationships, some terrain information about streets and buildings, and additional background information (e.g., adverse weather conditions). Furthermore, the extracted information forms a list of database records called "hints." In one example, the format of such database records includes obj-id, obj-location, obj-type, and obj-attributes, where obj is an object described in the global scene model 132. Finally, the coordinator 104 sends the hints back to the first agent 102A. Furthermore, the hints include id-agent, timestamp, query-id, and hint. Upon receiving a hint, the first agent 102A re-evaluates its own objectives by considering the newly received information. Therefore, the coordinator 104 retrieves relevant information from the global scene model 132 and further organizes the received information into hints shared with the first agent 102A. According to one embodiment, the controller of the first agent (i.e., the first controller 108) is further configured to determine that the sensor input is insufficient to complete the action, and in response, send a query to the coordinator 104 accordingly. First, the first controller 108 of the first agent 102A is configured to determine whether the available sensor input is insufficient to successfully perform the required action. Thereafter, the first controller 108 is configured to generate a query and further send the generated query to the coordinator 104.Therefore, the first controller 108 is used to ensure proactive communication from the first agent 102A, seeking assistance from the coordinator 104, even under data constraints, to enhance the understanding and decision-making capabilities of the first agent 102A. By enabling the first agent 102A to recognize limitations in sensor inputs and to assist through query requests to the coordinator 104, the agent network 100 facilitates efficient problem-solving and decision-making in real-time dynamic environments.

[0050] According to one embodiment, query-related data includes one, some, or all of the following: information about the object, object attributes and relationships, terrain information, and additional contextual information. Including information about the object, object attributes, relationships, terrain details, and additional contextual information in response to a query helps the first agent 102A obtain comprehensive and context-rich data when seeking information from the coordinator 104. Integrating various query-related data ensures that the first agent 102A receives a holistic understanding of the corresponding query, such as by including object details, object attributes, relationships, terrain features, and additional contextual background. Therefore, by including data, the first agent 102A ensures that received sensor data input includes a comprehensive insight that enhances the adaptability of the agent network 100 and further strengthens the decision-making capabilities of the first agent 102A, thereby supporting the first agent 102A in making informed choices in complex and dynamic environments.

[0051] According to one embodiment, the coordinator's controller (i.e., the third controller 128) is also configured to determine hints for a query based on the global scenario model 132 by extracting query-relevant data from the global scenario model 132. Furthermore, the extracted data is then structured into query-specific hints, ensuring that the first agent 102A receives accurate and relevant information to guide its decision-making capabilities. Therefore, by directly extracting query-relevant data from the global scenario model 132, the coordinator 104 provides accurate and reliable hints, enhancing the adaptability and decision-making efficiency of one or more agents 102 in the agent network 100 in dynamic environments.

[0052] According to one embodiment, the coordinator's controller (i.e., the third controller 128) is also used to extract query-relevant data using data retrieval techniques. Query-relevant data is located and extracted from the global scene model 132 using data retrieval techniques (e.g., semantic data association, pattern recognition, or any other complex algorithms). The extracted data is then structured into comprehensive hints, which are sent back to the query agent. Therefore, the coordinator 104 is used to improve the accuracy of the extracted data and enhance the decision-making efficiency of agents throughout the agent network 100.

[0053] Furthermore, the controller of the first agent (i.e., the first controller 108) is also used to receive prompts, re-evaluate sensor data in conjunction with the received prompts, and then execute the desired action. Additionally, after receiving a prompt, the first agent 102A is used to re-evaluate the input data using the newly received information and execute the desired action. By re-evaluating sensor data in response to specific guidance, the first agent 102A improves its responsiveness and decision-making capabilities within the agent network 100, thereby enhancing the overall adaptability and intelligence level of the agent network 100 when navigating in complex and dynamic environments.

[0054] According to one embodiment, the controller of the first agent 102A (i.e., the first controller 108) is further configured to re-evaluate sensor data by combining received cues with the following manner: jointly evaluating the transformed cues in combination with recognized features using a neural network that provides a rich feature set; and then selecting a desired action based on the rich feature set. The received cues are transformed (i.e., vectorized) into a format acceptable to the agent network 100. Furthermore, the transformed cues are further combined with input features and jointly evaluated features, such as selecting a desired action by using a neural network. The neural network uses a neural network that provides a rich feature set to facilitate the combination of cues and recognized features with jointly recognized features. The first controller 108 of the first agent 102A is also configured to select a desired action based on the rich feature set and ensure accurate and reliable selection of the desired action.

[0055] According to one embodiment, the controller of the first agent (i.e., the first controller 108) is also used to identify features through image segmentation. Image segmentation refers to the process of dividing an image into different segments, thereby enabling the identification of individual elements or objects within the image. By utilizing image segmentation, the first controller 108 of the first agent 102A can depict and identify specific features, thereby enhancing the perception of the dynamic environment, which becomes a valuable input for subsequent decision-making processes within the agent network 100.

[0056] Advantageously, the agent network 100 is used to generate a comprehensive global scene model 132 of a dynamic environment based on diverse local views of autonomous agents. Furthermore, the generated global scene model 132 is used to enhance the decision-making capabilities of one or more agents 102. Each agent is equipped with sensors and controllers for capturing image data, which can be further analyzed to perform actions. Additionally, the agent network 100 includes a coordinator 104 for combining datasets (i.e., segmented datasets and low-resolution sensor datasets) from one or more agents 102 of the agent network 100 to generate the global scene model 132. Generating the global scene model 132 enables one or more agents 102 of the agent network 100 to make informed decisions based on a shared understanding of the agent environment, thereby achieving efficient, effective, and coordinated actions. Furthermore, one or more agents 102 of the agent network 100 can freely join or leave the agent network 100 as needed, thus giving the agent network 100 high flexibility. When needed, one or more agents 102 in the agent network 100 facilitate high-bandwidth optimization by sending data to coordinator 104 and other agents 102 to minimize the amount of data exchanged within the agent network 100 and improve optimized resource utilization. Furthermore, one or more agents 102 enhance the privacy of the agent network 100 by selectively sending segmented data and low-resolution sensor datasets to coordinator 104, rather than directly sharing raw data with coordinator 104. Additionally, based on a global scene model 132, prompts for queries are determined to support one or more agents 102 in searching for additional information, which is further used to efficiently and reliably execute the desired action. Therefore, the agent network 100 provides a scalable, resource-efficient, and reliable solution for accurate and efficient decision-making by one or more agents 102 within the agent network 100.

[0057] Figure 2A and Figure 2B This is a schematic diagram of a flowchart depicting a method for an agent network including a coordinator and one or more agents. Referring to Figure 2, this figure shows a flowchart of a method 200 for an agent network including a coordinator 104 and one or more agents 102. Method 200 includes steps 202 to 226. In one implementation, a first controller 108 is used to perform all operations of method 200.

[0058] In step 202, method 200 includes receiving sensor input of a local scene from sensors. Furthermore, a first agent 102A and a second agent 102B are used to receive sensor input from a first sensor 110 and a second sensor 120. The received sensor input of the local scene may include local image data captured by the first sensor 110 and the second sensor 120. In step 204, method 200 includes identifying features related to the solving action. Furthermore, after receiving the sensor input of the local scene from the sensors, the first agent 102A and the second agent 102B are used to identify relevant features from the local scene captured by the first sensor 110 and the second sensor 120. In step 206, method 200 includes generating a segmentation dataset and a low-resolution sensor dataset. Furthermore, a first controller 108 of the first agent 102A and a second controller 120 of the second agent 102B are used to generate a first segmentation dataset 114, a first low-resolution sensor dataset 116, a second segmentation dataset 124, and a second low-resolution sensor dataset 126. In step 208, method 200 includes sending a segmented dataset and a low-resolution sensor dataset to a coordinator. Furthermore, a first agent 102A and a second agent 102B are used to send a first segmented dataset 114, a first low-resolution sensor dataset 116, a second segmented dataset 124, and a second low-resolution sensor dataset 126 to the coordinator 104. In step 210, method 200 includes receiving the first segmented dataset 114 and the first low-resolution sensor dataset 116 from the first agent 102A. Furthermore, the coordinator 104 in the agent network is used to receive the first segmented dataset 114 and the first low-resolution sensor dataset 116 from the first agent 102A.

[0059] In step 212, method 200 includes receiving a second segmented dataset and a second low-resolution sensor dataset from a second agent. Furthermore, a coordinator 104 in the agent network is used to receive a second segmented dataset 124 and a second low-resolution sensor dataset 126 from the second agent 102B. In step 214, method 200 includes generating a model of a global scene based on a combination of the datasets received from the agents. Furthermore, coordinator 104 is used to generate a model 132 of the global scene by combining the received first and second datasets. In step 216, method 200 includes receiving a query from a first agent. Furthermore, coordinator 104 is used to receive a first query from the first agent 102A. In step 218, method 200 includes determining a prompt for the query based on the global scene model. Furthermore, a third controller 128 of coordinator 104 is used to determine a prompt for the query received from the first agent 102A based on the global scene model 132. In step 220, method 200 includes sending a prompt to the first agent. Furthermore, the coordinator 104 is used to send a prompt to the first agent 102A via the communication network 106. In step 222, method 200 includes receiving the prompt. Furthermore, a first controller 108 of the first agent 102A is used to receive the prompt from the coordinator. In step 224, method 200 includes re-evaluating sensor data in conjunction with the received prompt. Furthermore, the first agent 102A is used to re-evaluate the first low-resolution sensor dataset 116 after receiving the prompt. In step 226, method 200 includes performing a desired action. Furthermore, after re-evaluating the sensor data according to the prompt, the first agent 102A performs the desired action.

[0060] Advantageously, method 200 generates a comprehensive global scene model 132 of the dynamic environment based on diverse local views of autonomous agents. Furthermore, the generated global scene model 132 is used to enhance the decision-making capabilities of one or more agents 102. Each agent is equipped with sensors and controllers for capturing image data, which can be further analyzed to perform actions. Additionally, method 200 combines datasets (i.e., segmented datasets and low-resolution sensor datasets) from one or more agents 102 of agent network 100 to generate the global scene model 132. Generating the global scene model 132 enables one or more agents 102 of agent network 100 to make informed decisions based on a shared understanding of the agent environment, thereby achieving efficient, effective, and coordinated actions. Furthermore, one or more agents 102 of agent network 100 can freely join or leave agent network 100 as needed, thus giving agent network 100 high flexibility. When needed, one or more agents 102 in method 200 facilitate high-bandwidth optimization by sending data to coordinator 104 and other agents 102 to minimize the amount of data exchanged within the agent network 100 and improve optimized resource utilization. Furthermore, one or more agents 102 enhance the privacy of the agent network 100 by selectively sending segmented data and low-resolution sensor datasets to coordinator 104 of one or more agents 102, rather than directly sharing raw data with coordinator 104 of one or more agents 102. Additionally, one or more agents 102 in method 200 seek additional information based on global scene model 132 to provide hints for queries, which is further used to efficiently and reliably execute the desired action. Therefore, agent network 100 provides a scalable, resource-efficient, and reliable solution for accurate and efficient decision-making by one or more agents 102. Advantageously, method 200 generates a comprehensive global scene model 132 of the dynamic environment based on the diverse local perspectives of autonomous agents. Furthermore, the generated global scene model 132 is used to enhance the decision-making capabilities of one or more agents 102. Each agent is equipped with sensors and controllers for capturing image data, which can be further analyzed to perform actions. Additionally, method 200 includes a coordinator 104 for combining datasets (i.e., segmented datasets and low-resolution sensor datasets) from one or more agents 102 in the agent network 100 to generate the global scene model 132. Generating the global scene model 132 enables one or more agents 102 in method 200 to make informed decisions based on a shared understanding of the agent environment, thereby achieving efficient, effective, and coordinated actions.Furthermore, one or more agents 102 can freely join or leave the agent network 100 as needed, thus enabling the agent network 100 to possess high flexibility. When needed, one or more agents 102 in method 200 promote high-bandwidth optimization by sending data to coordinator 104 and one or more other agents 102, in order to minimize the amount of data exchanged within the agent network 100 and improve optimized resource utilization. In addition, one or more agents 102 enhance the privacy of the agent network 100 by selectively sending segmented data and low-resolution sensor datasets to the coordinator 104 of one or more agents 102, rather than directly sharing raw data with the coordinator 104 of one or more agents 102. Furthermore, based on the global scene model 132, the prompts for the query determine the additional information that supports one or more agents 102 in the agent network 100 in searching for additional information, which is further used to efficiently and reliably execute the desired action. Therefore, method 200 provides a scalable, resource-efficient, and reliable solution for accurate and efficient decision-making by one or more agents 102 in method 200.

[0061] Steps 202 to 226 are merely illustrative, and other alternatives may be provided, in which one or more steps are added, one or more steps are deleted, or one or more steps are provided in a different order, without departing from the scope of the claims herein.

[0062] A computer program product is also provided, comprising program instructions for performing method 200 when executed by one or more processors in the agent network 100. The computer program product is implemented as an algorithm and embedded in software stored in a non-transitory computer-readable storage medium. The non-transitory computer-readable storage device may include, but is not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. Examples of implementations of the computer-readable storage medium include, but are not limited to, electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), read-only memory (ROM), hard disk drive (HDD), flash memory, secure digital (SD) cards, solid-state drives (SSDs), computer-readable storage media, and / or CPU cache memory.

[0063] Figure 3 This is a block diagram depicting agents in an agent network including one or more agents according to embodiments of the present disclosure. (In conjunction with...) Figure 1 element pairs in Figure 3 Provide a description. (See reference.) Figure 3 The attached figure illustrates an agent (i.e., Figure 1 A schematic diagram 300 of the first intelligent agent 102A is shown, the intelligent agent including a first controller 108, a first sensor 110, and a first memory 112. Furthermore, the first memory includes a first segmented dataset 114 and a first low-resolution sensor dataset 116.

[0064] Provides agents in an agent network 100, including a coordinator 104 and one or more agents 102 (i.e., Figure 1 The first agent 102A in the agent network 100. However, agents in the agent network 100 may correspond to the first agent 102A, the second agent 102B, up to the nth agent, without affecting the scope of this disclosure. Each agent (such as the first agent 102A) includes a sensor (i.e., the first sensor 110) and a controller (i.e., the first controller 108). The first controller 108 is used to receive sensor inputs of the local scene from the sensor (i.e., the first sensor 110) and identify features related to the solving action. The sensor inputs received from each of one or more agents 102 are used to provide a comprehensive and detailed understanding of the dynamic environment in real time, which can be further used to identify objects, patterns, or any other relevant elements in the dynamic environment. The first controller 108 is used to generate a segmented dataset (i.e., the first segmented dataset 114) and a low-resolution sensor dataset (i.e., the first low-resolution sensor dataset 116). A segmented dataset (e.g., a first segmented dataset 114 and a second segmented dataset 124) and a low-resolution sensor dataset (i.e., a first low-resolution sensor dataset 116 and a second low-resolution sensor dataset 126) are generated to gain insights into specific aspects of a local scene within a dynamic environment. Furthermore, a first controller 108 is used to send the first segmented dataset 114 and the first low-resolution sensor dataset 116 to a coordinator 104. Therefore, the generation of the first segmented dataset 114 and the first low-resolution sensor dataset 116 is sent to the coordinator 104 to provide a holistic representation of the dynamic environment.

[0065] Furthermore, the first controller 108 is used to send a query to the coordinator 104, receive a prompt based on the global scene model 132 in response to the query, re-evaluate sensor data in conjunction with the received prompt, and execute the desired action. First, the first agent 102A is used to collect local data and further evaluate whether the local data has sufficient information to execute the desired action. Furthermore, if local information is insufficient, the first agent 102A is used to select a query describing the missing information (i.e., a query from a predefined query list). In one example, the received queries forming the first agent 102A include "Can I safely turn right?", "Are there any obstacles ahead?", etc. Subsequently, the first agent 102A is used to send a query to the coordinator 104 including a tuple with agent-id, timestamp, location-agent, and query-id. In one implementation, the query includes information about objects, object attributes and relationships, terrain information, and additional background information. Furthermore, the query is accompanied by locally processed data (e.g., Im-Seg, Im-LowRes, etc.). Subsequently, upon receiving the query and local data from the first agent 102A, the coordinator 104 uses data retrieval techniques to extract all information related to the query and received local data from the global scene model 132 (or global database), such as information about objects, their attributes and relationships, some terrain information about streets and buildings, and additional background information (e.g., adverse weather conditions). Furthermore, the prompt includes the agent ID, timestamp, query ID, and prompt. Upon receiving the prompt, the first agent 102A re-evaluates its own objectives by considering the newly received information. Therefore, the first controller 108 thereby ensures that the decision-making process of the first agent 102A is efficient and accurate, thereby enhancing the adaptability and responsiveness of the first agent 102A in dynamic environments, and consequently contributing to the overall effectiveness of the agent network 100.

[0066] Figure 4 This is a schematic diagram depicting a flowchart of a method for using agents in an agent network comprising one or more agents according to embodiments of the present disclosure. (Reference) Figure 4 The accompanying figure illustrates a flowchart of a method 400 for a first agent in an agent network. Method 400 includes steps 402 to 416.

[0067] In step 402, method 400 includes: receiving sensor input of a local scene from a sensor. Furthermore, a first agent 102A is used to receive the sensor input of the local scene. In step 404, method 400 includes: identifying features related to the solving action. Furthermore, the first agent 102A is used to identify relevant features from the sensor input of the local scene. In step 406, method 400 includes: generating a segmentation dataset and low-resolution data. Furthermore, a first controller 108 of the first agent 102A is used to generate a first segmentation dataset 114 and a first low-resolution sensor dataset 116. In step 408, method 400 includes: sending the segmentation dataset and low-resolution sensor dataset to a coordinator. Furthermore, the first agent 102A is used to send the first segmentation dataset 114 and the first low-resolution sensor dataset 116 to a coordinator 104 via a communication network 106. In step 410, method 400 includes: sending a query to the coordinator 104. Furthermore, the first agent 102A is used to send a query to the coordinator 104. In step 412, method 400 includes receiving a query-based, global scenario-based prompt. Furthermore, the first agent 102A is configured to receive the query-based, global scenario-based prompt 132. In step 414, method 400 includes re-evaluating sensor data in conjunction with the received prompt. Furthermore, the first agent 102A is configured to re-evaluate the sensor data based on the received prompt. In step 416, method 400 includes performing a desired action. Furthermore, the first agent 102A is configured to perform the desired action. Advantageously, method 400 is used to ensure that the decision-making process of the first agent 102A is efficient and accurate, thereby enhancing the adaptability and responsiveness of the first agent 102A in dynamic environments, and consequently contributing to the overall performance of the agent network 100.

[0068] Steps 402 to 416 are merely illustrative, and other alternatives may be provided, in which one or more steps are added, one or more steps are deleted, or one or more steps are provided in a different order, without departing from the scope of the claims herein.

[0069] A computer program product is also provided, comprising program instructions for performing method 400 when executed by one or more processors in the agent network 100. The computer program product is implemented as an algorithm and embedded in software stored in a non-transitory computer-readable storage medium. The non-transitory computer-readable storage device may include, but is not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. Examples of implementations of the computer-readable storage medium include, but are not limited to, electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), read-only memory (ROM), hard disk drive (HDD), flash memory, secure digital (SD) cards, solid-state drives (SSDs), computer-readable storage media, and / or CPU cache memory.

[0070] Figure 5 This is a block diagram depicting a coordinator in an agent network including one or more agents according to embodiments of the present disclosure. (In conjunction with...) Figure 1 and Figure 3 element pairs in Figure 5 Provide a description. (See reference.) Figure 5 The accompanying drawing depicts ( Figure 1 A schematic diagram 500 of a coordinator 104, the coordinator comprising ( Figure 1 Global scene model 132, ( Figure 1 The third controller 128 and ( Figure 1 Network interface 134.

[0071] A coordinator 104 is provided in an agent network 100 comprising one or more agents 102. Furthermore, each agent (e.g., first agent 102A and second agent 102B) includes sensors (e.g., first sensor 110 and second sensor 120). The coordinator 104 includes a controller (i.e., a third controller 128) configured to receive a first segmented dataset 114 and a first low-resolution sensor dataset 116 from the first agent 102A. Subsequently, the coordinator 104 is configured to receive a second segmented dataset 124 and a second low-resolution sensor dataset 126 from the second agent 102B. The coordinator 104 is also configured to generate a global scene model 132 based on a combination of datasets received from the agents. Furthermore, the coordinator 104 is configured to receive a query from the first agent 102A, determine a prompt for the query based on the global scene model 132, and send the prompt to the first agent (102A). Furthermore, after receiving the first segmentation dataset 114, the first low-resolution sensor dataset 116, the second segmentation dataset 124, and the second low-resolution sensor dataset 126, the third controller 128 is used to generate a global scene model 132 based on a combination of datasets received from agents (i.e., from the first agent 102A and the second agent 102B). For example, the third controller 128 of the coordinator 104 is used to receive processed image data from agents and combine the processed data with information received from other agents. The combined image data is organized using received metadata (e.g., agent-id, timestamp, and location-agent). Furthermore, the combined data is continuously processed by two or more algorithms (such as cross-view attention neural networks, diffuse attention neural networks, etc.), but the scope of this disclosure is not limited. Advantageously, the generated comprehensive global scene model 132 is used as a knowledge base, and the third controller 128 processes queries to determine hints for making informed decisions, effectively improving the scalability and adaptability of one or more agents in the agent network.

[0072] Figure 6 This is a schematic diagram depicting a flowchart of a method for using agents in an agent network according to embodiments of the present disclosure. (See also:) Figure 6 The accompanying figure shows a flowchart of a method 600 of an agent (i.e., the first agent 102A) in an agent network 100.

[0073] In step 602, method 600 includes: coordinator 104 performing the following operations: receiving a first segmented dataset 114 and a first low-resolution sensor dataset 116 from a first agent 102A. In step 604, method 600 includes: receiving a second segmented dataset 124 and a second low-resolution sensor dataset 126 from a second agent 102B. Subsequently, in step 606, method 600 includes: generating a global scene model 132 based on a combination of datasets received from the agents. Then, in step 608, method 600 includes: receiving a query from the first agent 102A; in step 610, method 600 includes: determining a hint for the query based on the global scene model. Finally, in step 612, method 600 includes: sending a hint to the first agent. Advantageously, the generated comprehensive global scene model 132 is used as a knowledge base, and the third controller 128 processes the query to determine hints for making informed decisions, effectively improving the scalability and adaptability of one or more agents in the agent network.

[0074] Steps 602 to 612 are merely illustrative, and other alternatives may be provided, in which one or more steps are added, one or more steps are deleted, or one or more steps are provided in a different order, without departing from the scope of the claims herein.

[0075] A computer program product is also provided, comprising program instructions for performing method 600 when executed by one or more processors in the agent network 100. The computer program product is implemented as an algorithm and embedded in software stored in a non-transitory computer-readable storage medium. The non-transitory computer-readable storage device may include, but is not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. Examples of implementations of the computer-readable storage medium include, but are not limited to, electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), read-only memory (ROM), hard disk drive (HDD), flash memory, secure digital (SD) cards, solid-state drives (SSDs), computer-readable storage media, and / or CPU cache memory.

[0076] Figure 7A This is a schematic diagram depicting the reconstruction of a global scene model according to an embodiment of the present disclosure. Combined with... Figure 1 , Figure 3 and Figure 5 element pairs in Figure 7A Provide a description. (See reference.) Figure 7A The accompanying diagram illustrates a depiction of an intelligent agent (i.e., Figure 1 First intelligent agent 102A) and ( Figure 1 Schematic diagram 700A of the coordinator 104.

[0077] In one implementation scenario, each agent (e.g., first agent 102A) is equipped with a first sensor 110, such as a camera sensor, which provides partial understanding of the dynamic environment. Furthermore, each agent (e.g., first agent 102A) performs action 702 to regenerate a global scene model 132 and query hints for knowledge sharing. First agent 102A is used to capture snapshots of local views (i.e., images 704 from the camera). The captured data is then processed by a task-independent compression algorithm to remove all redundancy from the perceptual data, and an acquisition-perceptual processing algorithm to emphasize all information relevant to the agent's task while suppressing irrelevant information. Furthermore, the output of the task-independent compression algorithm is a low-resolution version of the input view Im-LowRes, and the output of the task-perceptual processing algorithm is a segmented image (i.e., Im-Seg) showing the outlines of detected objects (e.g., cars, pedestrians, street lines, etc.). The first agent 102A processes perceptual data using a visual neural network 706 and sends the processed input data to the coordinator 104 (e.g., at operation 714). The coordinator 104 receives the processed perceptual data, which includes agent-id, timestamp, location-agent, Im-Seg, and Im-LowRes. The first agent 102A generates a query list 708 and sends an action-related query 726 to the coordinator 104, for example, at operation 720. The coordinator 104 generates a prompt 722 based on the action-related query, through a prompt generation operation 718. The coordinator 104 also sends the action-related query 726 along with the action-related prompt to the first agent 102A, for example, at operation 724. Furthermore, the first agent 102A executes an artificial neural network 710 to perform the desired action. At operation 716, coordinator 104 receives processed data from first agent 102A and concatenates the processed data with data received from other agents among one or more agents 102. Furthermore, the order of the concatenated data is maintained by the received metadata (i.e., agent-id, timestamp, location-agent). The concatenated data is further processed continuously using transformers (such as cross-view attention neural networks and diffuse attention neural networks). The cross-view attention model is trained to associate (at the semantic level) data fragments provided by one or more agents 102. Such associations may include identifying objects seen by one or more agents 102 from different angles, discovering relationships between two distant objects observed by different agents, and locating views in a global 3D coordinate system.Furthermore, the diffuse attention neural network is trained to regenerate the global scene model 132 from the output of the cross-view attention neural network and data received from other agents (e.g., at operation 712). The diffuse attention neural network is used to generate new synthetic information that can fill gaps to support the regeneration of the global scene model 132 suitable for machine-type processing and communication.

[0078] Figure 7B A schematic diagram depicts a partial view processed by an intelligent agent according to an embodiment of the present disclosure. (In conjunction with...) Figure 1 , Figure 3 , Figure 5 and Figure 7A element pairs in Figure 7B Provide a description. (See reference.) Figure 7B The accompanying figure illustrates a schematic diagram 700B depicting a partial view processed by an agent (such as a first agent 102A or a second agent 102B in one or more agents 102). In one implementation scenario, an original view 728A, a low-resolution version 728B of the original view (i.e., Im-LowRes), and a segmented version 728C of the original view (i.e., Im-Seg) are shown. Thus, one or more agents 102 are used to make informed decisions by accurately and reliably viewing the dynamic environment in real time without any latency.

[0079] Figure 8 This is a schematic diagram depicting the input processing of an intelligent agent according to an embodiment of the present disclosure. (Reference) Figure 8 The accompanying figure shows a schematic diagram 800 depicting an input image 802, based on which the agent performs a desired action (e.g., at operation 804).

[0080] At operation 806, when the agent fails to perform the desired action, the first controller 108 of the first agent 102A performs task-aware segmentation 808 and task-independent image compression 810. Furthermore, at task-aware segmentation 808, the input image 802 undergoes full image segmentation 814 to generate a segmented image 816, and the segmented image 816 undergoes target-aware filtering 818 to generate a task-filtered segmented image 822. Subsequently, the first agent 102A of one or more agents 102 performs full image segmentation using a trained neural network (e.g., Mask R-CNN). Simultaneously, the input image 802 is compressed using some task-independent image compression technique. Furthermore, at task-independent image compression 810, the input image 802 undergoes lossy image compression 820 to generate a low-resolution image 824. Additionally, the task-filtered segmented image 822 and the low-resolution image 824 undergo position embedding 812, which enables the first agent 102A to take the desired necessary action in the real-time scene.

[0081] Figure 9 This is a schematic diagram illustrating global scene model regeneration with cross-view attention and diffusion according to an embodiment of the present disclosure. Reference Figure 9 The accompanying figure illustrates allowable resolution and segmented image data 902 with location embeddings received from one or more agents 102, which undergoes data fusion 904 based on an artificial neural network and is combined with a global map embedding 908 using a query 910. Low resolution and segmented image data 902 with location embeddings is processed by a cross-view artificial neural network 912. The cross-view artificial neural network 912 receives three types of input: queries, keys, and values. Keys and values ​​are combined tuples (such as Im-Seg, Im-LowRes, and embeddings) from multiple agents. Query is a global map embedding, which can be viewed as an empty model structure of the global scene to be filled with semantic information.

[0082] At operation 906, the low-resolution and segmented image data 902 with location embeddings then undergoes a diffusion-based scene regeneration diffusion artificial neural network 914 to regenerate a global scene model 132. The low-resolution and segmented image data 902 with location embeddings and a cross-view artificial neural network 912 are correlated, and all relevant semantic information is populated into the global map embedding 908. The resulting partially filled model is then fed into the diffusion artificial neural network 914 for training. The diffusion artificial neural network 914 is used to fill information gaps. The output of the diffusion artificial neural network 914 is the complete global scene model 132, which is treated as a database 916. Therefore, the regeneration of the global scene model 132 using the cross-view and diffusion network models is performed.

[0083] Figure 10This is a schematic diagram illustrating prompt generation by the coordinator according to an embodiment of the present disclosure. (In conjunction with...) Figure 1 , Figure 3 , Figure 5 , Figure 7A , Figure 8 and Figure 9 element pairs in Figure 10 Provide a description. (See reference.) Figure 10 The accompanying figure illustrates a schematic diagram 1000 depicting a coordinator 104 and a global scene model 132. Furthermore, the global scene model 132 includes a database for the model, which includes attributes such as id, class, and attr. As shown, the coordinator 104, upon receiving a query from the first agent and its processed local view, extracts facts and features relevant to the first agent 102A from the global scene model 132. These facts and features include information about objects, object attributes and relationships, some terrain information (such as streets and buildings), and additional background information (such as adverse weather conditions). The extraction of relevant information is performed using existing data retrieval techniques. Therefore, based on the relevant information, the first agent 102A takes the desired action.

[0084] Figure 11 This is a schematic diagram illustrating scene-aware semantic analysis performed on the agent side according to an embodiment of the present disclosure. (See reference) Figure 11 The accompanying figure illustrates an agent 1100 (e.g., a first agent 102A or a second agent 102B). Upon receiving a cue 1102A, the agent undergoes cue-aware image segmentation 1104, where the cue 1102A undergoes vectorization 1106 to generate a scene map feature vector 1108. Image 1102B is processed by a region candidate network 1110 to generate a target feature vector 1112. Furthermore, the scene map feature vector 1108 and the object feature vector 1112 are concatenated to generate cross-attention 1114, thereby generating a refined feature vector 1116. This refined feature vector helps the agent 1100 execute actions 1118 accurately and efficiently.

[0085] Modifications to the embodiments of the present disclosure described above may be made without departing from the scope of the present disclosure as defined by the appended claims. Expressions such as “comprising,” “combining,” “having,” “is,” and “are” used to describe and claim this disclosure are intended to be interpreted in a non-exclusive manner, allowing for the presence of items, components, or elements not explicitly described. Singular references should also be interpreted to refer to the plural. The term “exemplary” as used herein means “as an example, instance, or illustration.” Any embodiment described as “exemplary” is not necessarily to be construed as being more preferred or advantageous than other embodiments, or excluding combinations of features from other embodiments. The term “optionally” as used herein means “provided in some embodiments and not in others.” It should be understood that certain features of the present disclosure described in the context of a single embodiment for brevity may also be provided in combination in a single embodiment. Conversely, various features of the invention described in the context of a single embodiment for brevity may also be provided individually or in any suitable combination or appropriately in any other described embodiment of the present disclosure.

Claims

1. A method, an agent network (100) comprising a coordinator (104) and one or more agents (102), characterized in that, Each agent includes sensors and a controller, the controller being used for: Receive sensor input from the local scene from the sensor; Identify and solve for features related to the action; Generate segmented datasets and low-resolution sensor datasets; Sending the segmented dataset and the low-resolution sensor dataset to the coordinator (104), wherein the coordinator (104) includes a controller, the controller being configured to: Receive a first segmented dataset (114) and a first low-resolution sensor dataset (116) from the first intelligent agent (102A). Receive the second segmented dataset (124) and the second low-resolution sensor dataset (126) from the second agent (102B); A global scene model is generated based on a combination of datasets received from the agent (132). Receive queries from the first intelligent agent; The model (132) based on the global scenario is used to determine the suggestions for the query; as well as Send the prompt to the first intelligent agent (102A); Therefore, the controller of the first intelligent agent (102A) is also used for: Receive the prompt; Reassess the sensor data based on the received prompts; and Perform the desired action.

2. The agent network (100) of claim 1, characterized in that The controller of the first intelligent agent (102A) is also configured to re-evaluate the sensor data in conjunction with the received prompts in the following manner: A neural network providing a rich feature set is used to jointly evaluate the transformed cue combined with the identified features; The desired action is selected based on the rich feature set.

3. The agent network (100) according to claim 1 or 2, characterized in that The controller of the first intelligent agent is also used for: If the sensor input is determined to be insufficient to solve the action, the query is sent to the coordinator (104) in response.

4. The agent network (100) according to any one of the preceding claims, characterized in that The controller of the first intelligent agent (102A) is also used for: If the sensor input is determined to be insufficient to solve the action, then a trained neural network is used to perform full segmentation. Filter out segmented objects that are irrelevant to the local objectives of the task; Send location information about the global position and attitude of the agent, as well as a timestamp indicating the recording time, to the coordinator (104).

5. The agent network (100) according to any one of the preceding claims, characterized in that The controller of the coordinator (104) is also configured to generate a model (132) of the global scene based on a combination of datasets received from the agent by running cross-view attention and diffusion.

6. The agent network (100) according to any one of the preceding claims, characterized in that The controller of the coordinator (104) is also used for: By extracting query-related data from the model (132) of the global scenario, a prompt for the query is determined based on the model (132) of the global scenario.

7. The agent network (100) according to any one of the preceding claims, characterized in that The query-related data includes one, some, or all of the following: information about the object, the object's attributes and relationships, terrain information, and additional background information.

8. The agent network (100) of claim 7, characterized in that The controller of the coordinator (104) is also used to: use data retrieval technology to extract query-related data.

9. The agent network (100) according to any one of the preceding claims, characterized in that The sensor is an image sensor, and the sensor data is image data.

10. The agent network (100) according to any one of the preceding claims, characterized in that The sensor is a radar sensor, and the sensor data is radar image data.

11. The agent network (100) according to any one of the preceding claims, characterized in that The controller of the first intelligent agent (102A) is also used to: identify features through image segmentation.

12. A method (200) for an agent network (100) comprising a coordinator (104) and one or more agents (102), characterized in that, Each agent includes a sensor, wherein the method includes the agent performing the following operations: Receive sensor input from the local scene from the sensor; Identify and solve for features related to the action; Generate segmented datasets and low-resolution sensor datasets; and The segmented dataset and the low-resolution sensor dataset are sent to the coordinator (104), and the method includes the coordinator (104) performing the following operations: Receive a first segmented dataset (114) and a first low-resolution sensor dataset (116) from the first intelligent agent (102A). Receive the second segmented dataset (124) and the second low-resolution sensor dataset (126) from the second agent (102B); A model of the global scene is generated based on a combination of datasets received from the agent (132). Receive queries from the first intelligent agent (102A); The model (132) based on the global scenario is used to determine the suggestions for the query; Sending the prompt to the first intelligent agent (102A), thereby the method (200) further includes: The first intelligent agent (102A) performs the following operations: Receive the prompt; Reassess the sensor data based on the received prompts; Perform the desired action.

13. An agent in an agent network (100) comprising a coordinator (104) and one or more agents (102A), characterized in that, Each agent includes sensors and a controller, the controller being used for: Receive sensor input from the local scene from the sensor; Identify and solve for features related to the action; Generate segmented datasets and low-resolution sensor datasets; and The segmented dataset and the low-resolution sensor dataset are sent to the coordinator (104); Send a query to the coordinator (104); Receive prompts based on the query and the global context; Reassess the sensor data based on the received prompts; as well as Perform the desired action.

14. A method (400) for agents in an agent network (100) including a coordinator (104) and one or more agents (102), characterized in that, Each agent includes a sensor, wherein the method (400) includes the agent performing the following operations: Receive sensor input from the local scene from the sensor; Identify and solve for features related to the action; Generate segmented datasets and low-resolution sensor datasets; and The segmented dataset and the low-resolution sensor dataset are sent to the coordinator (104); Send a query to the coordinator (104); Receive prompts based on the query and the global context; Reassess the sensor data based on the received prompts; as well as Perform the desired action.

15. A coordinator (104) in an agent network (100) comprising one or more agents (102), characterized in that, Each agent includes sensors, wherein the coordinator (104) includes a controller, the controller being configured to: Receive a first segmented dataset (114) and a first low-resolution sensor dataset (116) from the first intelligent agent (102A). Receive the second segmented dataset (124) and the second low-resolution sensor dataset (126) from the second agent (102B); A model of the global scene is generated based on a combination of datasets received from the agent (132). Receive queries from the first intelligent agent (102A); Based on the global scene model (132), suggestions for the query are determined; The prompt is sent to the first intelligent agent (102A).

16. A method (600) for a coordinator (104) in an agent network (100) comprising one or more agents (102), characterized in that, Each agent includes sensors, wherein the method (600) includes the coordinator (104) performing the following operations: Receive a first segmented dataset (114) and a first low-resolution sensor dataset (116) from the first intelligent agent (102A). Receive the second segmented dataset (124) and the second low-resolution sensor dataset (126) from the second agent (102B); A model of the global scene is generated based on a combination of datasets received from the agent (132). Receive queries from the first intelligent agent (102A); Based on the global scene model (132), suggestions for the query are determined; The prompt is sent to the first intelligent agent (102A).

17. A computer program product, characterized in that, Includes program instructions that, when executed by one or more processors in the agent network (100), are used to perform the method (200, 400, 600) according to claim 12, 14 or 16.