Autonomous electromagnetic agent system based on programmable metasurfaces

By combining a programmable metasurface with a large-scale language model, the metaAgent system achieves full-coverage perception and control of intelligent scenarios, solves the problems of privacy protection and danger identification, and provides autonomous interaction and real-time response capabilities.

CN119514589BActive Publication Date: 2025-10-24PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411508753.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-28
Publication Date
2025-10-24
Estimated Expiration
2044-10-28

AI Technical Summary

Technical Problem

Existing technologies make it difficult to achieve full coverage perception and control of the entire smart scene, especially in private areas. They are unable to achieve real-time perception and understanding of user intentions and are unable to provide effective assistance in dangerous situations.

Method used

By combining programmable metasurfaces and large-scale language models, a metaAgent system is designed. It uses semantically programmable metasurfaces to manipulate electromagnetic waves, combines multimodal sensing data, and utilizes large-scale language models for task decomposition and execution to achieve autonomous interaction and hazard identification.

Benefits of technology

It enables comprehensive perception and control of intelligent scenarios, protects user privacy, allows for natural language interaction from any location, and provides timely assistance in dangerous situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119514589B_ABST
    Figure CN119514589B_ABST
Patent Text Reader

Abstract

The application discloses an autonomous electromagnetic intelligent agent system based on a programmable metasurface and belongs to the technical field of intelligent robots.The application comprises a cerebellum module based on a programmable metasurface and a brain module based on a large-capacity basic model, the system has autonomous beam control ability of human-level intelligence; the cerebellum module utilizes the programmable metasurface to execute electromagnetic wave control tasks; the brain module is used for generating high-level task strategies, decomposing high-level tasks and issuing executable subtask sequences obtained to the cerebellum module; the brain module and the cerebellum module work on two host computers respectively, and data transmission is carried out between the two through a network communication protocol. The autonomous electromagnetic intelligent agent system of the application exists in an intelligent scene in the form of a wisdom butler and can be widely applied in various intelligent scenes such as wisdom family, wisdom factory, wisdom community and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of intelligent robots, and relates to the cross technology of artificial intelligence, embodied intelligence, robots and electromagnetic perception, large language model technology, human-computer interaction technology, and in particular to an electromagnetic metasurface intelligent agent system based on a programmable metasurface. BACKGROUND

[0002] With the rapid development of modern information technology, autonomous intelligent agent products driven by embodied intelligence technology, large language model technology, Internet of Things technology, 5G communication technology and artificial intelligence technology have gradually been applied to people's daily life. Therefore, with the gradual landing of these technologies, autonomous intelligent agents will also promote the gradual entry of smart homes, smart factories and other intelligent scenarios into our real life. On the other hand, due to the continuous technical improvement and functional enhancement of intelligent devices, how to build an autonomous intelligent agent as an intelligent official of a smart home or factory, so that it can communicate with users through language and be able to understand the environment, and even users only need simple gestures to make the autonomous agent understand the user's needs, has become an urgent need of the current developing intelligent environment.

[0003] The existing technology is currently still at the stage of voice control of various devices, and some devices can only perform simple human body detection on users in a local area. The existing technology is difficult to achieve full coverage perception and control of the entire intelligent scene, especially in some private place areas, and cannot realize the perception and understanding of the user's intention at any time and any place, and even cannot effectively help the user in a dangerous situation.

[0004] Programmable metasurfaces are ultra-thin engineered structures that have been recognized as important media for flexible manipulation of electromagnetic waves and have valuable applications in many fields such as perception, control, communication and computing. Electromagnetic perception technology based on programmable metasurfaces can realize real-time and all-weather intelligent perception of user's posture, behavior and health status in indoor environment, and microwave-based perception technology will not cause the problem of easily invading user privacy as visual perception does. In addition, programmable metasurfaces can also achieve remote and accurate control of wireless devices in the environment, so that intelligent devices can quickly respond to each other and will not cause the risk of user data leakage as traditional communication methods do.

[0005] Recently, large-capacity foundation models (LFMs), especially large language models (LLMs), have achieved remarkable success in realizing human-level intelligence, such as ChatGPT, BERT, Wenxin Yanyan, and so on. LLM technology will bring great impetus to the development of service-oriented autonomous agents. Autonomous agents assisted by large language models will serve as an intelligent butler in future intelligent scenarios, and users can interact with autonomous agents anytime and anywhere. Autonomous agents will dispatch specific execution mechanisms to complete specific tasks issued by users after understanding the user's intent.

[0006] Autonomous agents combining programmable metasurfaces with large language models will be able to play multiple roles in future intelligent environments, such as advanced caregivers, family doctors, security guards, and chefs. No matter where the user is in the environment, the intelligent butler can perceive the user's behavior and state in real time through programmable metasurfaces and make correct decisions with the powerful reasoning and understanding capabilities of large language models, and then decompose and gradually execute the tasks. Therefore, autonomous electromagnetic agent systems based on programmable metasurfaces will become an effective way to perceive complex environments and execute complex tasks in future next-generation smart homes or smart factories. SUMMARY

[0007] To solve the above problems of the prior art, the present application provides an autonomous electromagnetic agent system based on programmable metasurfaces, called metaAgent.

[0008] The metaAgent system of the present application is composed of two parts: one part is a cerebellum module based on programmable metasurfaces, which uses programmable metasurfaces to perform specific electromagnetic manipulation tasks; the other part is a brain module based on large-capacity basic model LFM, which is used to generate high-level task strategies, decompose complex tasks into a series of simple executable subtask sequences, and issue executable subtask sequences to the cerebellum module. In the present application, the programmable metasurfaces used by the metaAgent system are called semantically programmable metasurfaces (SPMs), which are different from traditional programmable metasurfaces. The semantically programmable metasurfaces control beams through semantic coding patterns. In the present application, the semantic coding pattern reflects that the metaAgent system brain uses high-level natural language formulated prompts in the reasoning process, rather than low-level binary digital sequences in traditional programmable metasurfaces. According to multi-modal input data (including text, voice, image, microwave signal), the brain module of metaAgent will start an action plan, which includes commanding each SPM in the cerebellum module to obtain supplementary perception data, commanding robot entities, etc. Therefore, the autonomous electromagnetic intelligent agent system metaAgent of the present application will exist in the form of a smart butler in an intelligent scene, and the user (metaAgent user) and the metaAgent can freely communicate through language. At the same time, when the user falls down or other dangerous behaviors occur, the metaAgent will actively perform behavior posture recognition and health sign detection, quickly perceive the current dangerous situation of the user (such as user falling down, rapid heartbeat, etc.), and make corresponding actions according to the actual situation, so as to help the user get out of danger. The present application can be widely applied in various intelligent scenes (such as smart home, smart factory, smart community, etc.) in the future.

[0009] The technical scheme adopted by the autonomous electromagnetic intelligent agent system of the present application is as follows:

[0010] An autonomous electromagnetic intelligent agent system metaAgent based on programmable metasurfaces, comprising: a brain module for formulating high-level strategies and a cerebellum module for executing specific electromagnetic manipulation tasks; the metaAgent system has autonomous beam manipulation ability with human-level intelligence based on programmable metasurface beam manipulation technology and autonomous intelligent agent technology; specifically as follows:

[0011] The metaAgent brain module comprises a memory submodule, a perception expert submodule, a planning expert submodule, a grounding expert submodule and a coding expert submodule. The memory submodule is composed of an action library and a memory library. The action library is used to store a device library composed of all devices capable of being controlled by the metaAgent and a skill library composed of all skills capable of being executed by the metaAgent. The memory library is used to store all historical records during the operation of the metaAgent system and environment information corresponding to the operation environment, mainly including a visual semantic map of the environment in which the user is located, which is saved in the form of a knowledge graph. The perception expert submodule continuously collects multi-modal data as input, including sound recorded by a microphone, images recorded by a camera, text input by a keyboard and microwave signals collected by a software-defined radio (USRP), and processes the multi-modal data through a deep neural network model (a specific type of model is selected according to a specific task) integrated by the system to output semantic results for the planning expert submodule. The perception expert submodule selects appropriate deep neural networks or other signal processing methods to analyze and understand the input information. In terms of image and sound processing, pre-trained models on large-scale datasets can be used. For example, Xunfei API can identify the identity and voice content of a speaker, and a stereo depth camera can realize multi-camera fusion and human key point detection. Microwave perception of the mmetaAgent system to the environment is based on multiple programmable metasurfaces working at different frequency bands, which can be used to realize various perception tasks (such as skeleton key point detection, user positioning, behavior recognition, breathing and heartbeat monitoring, etc.). Microwave perception uses programmable metasurfaces to detect the environment and uses artificial neural networks trained through supervised learning to interpret the measured signals. The above-mentioned perception methods (including visual perception, voice perception and microwave perception) are also shared with the coding expert submodule. When optical perception is invalid (such as behind a corner or an opaque layer) or unavailable (such as for privacy reasons), microwave perception can be used as a supplement to visual perception. On the other hand, microwave has the ability of penetration perception, so it can be safely applied in some areas where visual sensors are not suitable due to user privacy concerns. The perception expert submodule comprehensively understands the multi-modal input data in natural language and passes it to the planning expert submodule. The planning expert submodule analyzes the natural language input from the perception expert submodule or the coding expert submodule using the knowledge provided by the memory submodule, and decomposes the natural language input into a series of executable subtask sequences. Then, the executable subtask sequences to be executed are passed to the grounding expert submodule in the form of natural language. The planning expert submodule is implemented based on an LLM that provides some context demonstrations. In addition, the planning expert submodule is improved according to human feedback stored in the memory library of the memory submodule. The coding expert submodule inputs a task sequence generated by the planning expert submodule, a function of a required action generated by the grounding expert module and natural language devices into a triple as input.In particular, the encoding expert submodule checks whether various required subtasks can be executed in parallel or need to be executed sequentially. Then, the encoding expert submodule writes python code and automatically runs it on the host computer. The output of the encoding expert submodule is generated in natural language in two kinds of output: one is "external" output, which is used to communicate with human users through chat tools; the other is "internal" output, which is provided to the planning expert submodule and the grounding expert submodule for consideration. The "internal" output is a key part of the metaAgent system to reason in natural language through the above multi-expert (perception expert, planning expert, grounding expert and encoding expert) discussion, and the autonomous self-organizing skills of the metaAgent system are built on this basis. The encoding expert submodule is designed based on the LLM provided with some context demonstrations.

[0012] The cerebellum module of the metaAgent includes programmable metasurfaces working at different frequency bands, FPGA control modules, USRP modules, transmitting antennas and receiving antennas, visual sensing modules, text input modules, and voice modules. The programmable metasurfaces are composed of 32*24 controllable units (the structure of the controllable units mainly consists of a rectangular metal resonant patch, a phase bias line, a dielectric substrate, and a metal reflection plate), and work at 2.4 GHz, 5.5 GHz, and 9.7 GHz frequency bands. Each programmable metasurface is controlled by a FPGA control module, and the USRP modules are responsible for real-time collection of microwave data through the transmitting and receiving antennas. The USRP modules are connected to the host computer, and all control algorithms of the metaAgent system run on the host computer. The visual perception module is composed of multiple ZED2 depth cameras for collecting image information of the environment, the text input module is composed of a keyboard connected to the host computer for collecting text input information of the user, and the voice module is composed of a microphone and a speaker for real-time interaction between the user and the metaAgent through voice.

[0013] The brain module and the cerebellum module of the metaAgent work on two host computers respectively, and data transmission is performed between the two modules through a network communication protocol.

[0014] The above autonomous electromagnetic agent system based on programmable metasurfaces can realize flexible interaction between the metaAgent and the user. When the autonomous electromagnetic agent system works, the following steps can be included:

[0015] 1) The brain module of the autonomous electromagnetic agent system receives the sensing model data transmitted by the cerebellum module;

[0016] 2) The perception expert submodule of the brain module receives multi-modal (text, speech, image, microwave) data, analyzes the multi-modal data through the corresponding deep learning network model, and summarizes the semantic information in a specific format as the output result;

[0017] 3) The planning expert submodule receives the semantic information output by the perception expert submodule, decomposes the specific task according to the knowledge of the memory submodule, and generates a series of executable subtask sequences;

[0018] 4) The grounding expert submodule allocates corresponding executable devices and corresponding skills for each subtask according to the knowledge of the memory submodule;

[0019] 5) The coding expert submodule compiles the python code executable on the host according to the subtask sequence and the allocated executable devices and skills;

[0020] 6) The cerebellum module receives the python code issued by the brain module and executes it on the host;

[0021] 7) The semantic programmable metasurface of the cerebellum module executes the corresponding python code to complete the corresponding task;

[0022] 8) The cerebellum module outputs the task execution result of the semantic programmable metasurface to the user, and also feeds back to the brain module;

[0023] 9) The brain module updates the memory submodule.

[0024] Compared with the prior art, the beneficial effects of the present application are:

[0025] The present application proposes an autonomous electromagnetic intelligent agent system based on programmable metasurface; the present application has the following advantages:

[0026] (I) The present application designs an intelligent agent for autonomous electromagnetic manipulation based on programmable metasurface combined with large language model, which can interact with users anytime and anywhere and flexibly through natural language.

[0027] (II) The present application utilizes the powerful wide field of view and penetrable perception ability of programmable metasurface to realize omnidirectional perception and control of users, robots and environment in intelligent scenes, including human three-dimensional skeleton key point detection, breathing and heartbeat detection, object positioning, signal enhancement and robot motion control, etc.

[0028] (III) The metaAgent system framework proposed in the present application has expandability, which can be customized according to the specific needs in specific intelligent scenes.

[0029] (Fourthly), the metaAgent system provided by the application has data security and can fully protect the personal privacy of users. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 is a structural block diagram of the autonomous electromagnetic agent system based on programmable metasurface of the application.

[0031] Figure 2 is a flowchart of the working of the autonomous electromagnetic agent system based on programmable metasurface in the embodiment of the application.

[0032] Figure 3 is a visual interface designed by the autonomous electromagnetic agent system based on programmable metasurface in the embodiment of the application. DETAILED DESCRIPTION

[0033] The application will be further described by examples in conjunction with the drawings, but the scope of the application is not limited in any way.

[0034] The application provides an autonomous electromagnetic agent system based on programmable metasurface, which comprises a brain module based on a large-capacity basic model and a cerebellum module based on programmable metasurface.

[0035] The autonomous electromagnetic agent system based on programmable metasurface provided by the application is composed of only two core parts.

[0036] The memory module is composed of an action library and a memory library; the action library is used for storing a device library composed of all devices capable of being controlled by the system and a skill library composed of all skills capable of being executed by the system; and the memory library is used for storing all historical records during the operation of the system and environment information corresponding to the operation environment, mainly including a visual semantic map of the environment, which is saved in the form of a knowledge graph.

[0037] The perception expert module is composed of multiple deep learning models for converting perception data of multiple modalities (such as text, sound, vision, microwave, etc.) into semantic results in the language mode.

[0038] The planning expert module is designed based on a large language model, uses the knowledge provided by the memory module, analyzes the natural language input from the perception expert or the coding expert, and decomposes it into a series of executable sub-tasks.

[0039] The grounding expert module is designed based on a large language model, responsible for receiving the sub-task list of the planning expert, and assigning appropriate skills and related equipment for each sub-task.

[0040] The coding expert module is designed based on a large language model, responsible for writing Python code that can run on the host based on the three types of knowledge output by the previous expert: expected goals in natural language input, action skills, and equipment.

[0041] The cerebellum module is composed of multiple semantic programmable metasurfaces, microphone modules, camera modules, and text input modules.

[0042] The semantic programmable metasurface is composed of a programmable metasurface, a micro control module, a universal software radio peripheral (USRP), a transceiver antenna, and a host. The programmable metasurface is composed of 32*24 independently controllable units, working at 2.4GHz, 5.5GHz and 9.7GHz frequency bands; the micro control module is connected with the programmable metasurface and is composed of a field programmable logic gate array (FPGA) module; the universal software radio peripheral is connected with the transceiver antenna and simultaneously connected with the host through a network port, and the semantic programmable metasurface control algorithm runs on the host.

[0043] In the autonomous electromagnetic agent system based on programmable metasurfaces, further, the programmable metasurface includes but is not limited to the programmable metasurface working at 2.4GHz, 5.5GHz and 9.7GHz frequency bands.

[0044] The perception expert, the planning expert, the grounding expert, and the coding expert constitute an expert group chat to formulate high-level strategies in the form of natural language discussion.

[0045] When the autonomous electromagnetic agent system based on programmable metasurfaces is working, the brain perception expert transmits tasks to the planning expert by analyzing environmental data or user instructions, the planning expert is responsible for task decomposition and assigns corresponding equipment and skills by the grounding expert, then the coding expert generates corresponding code, and finally the programmable metasurface of the cerebellum executes, thereby realizing human-computer interaction. Including the following steps:

[0046] 1) The brain receives the data of each sensing model transmitted by the cerebellum;

[0047] 2) Perception experts receive multi-modal (text, speech, image, microwave) data, analyze the data through corresponding deep learning network models, and summarize the output results into semantic information in a specific format;

[0048] 3) Planning experts receive semantic information output by perception experts, decompose specific tasks according to the knowledge of the memory module, and generate a series of executable sub-task sequences;

[0049] 4) Grounding experts assign corresponding executable devices and corresponding skills to each sub-task according to the knowledge of the memory module;

[0050] 5) Coding experts write python code executable on the host according to the sub-task sequence and the assigned devices and skills;

[0051] 6) Cerebellum receives python code issued by brain and executes on the host;

[0052] 7) Semantic programmable metasurface executes corresponding code to complete corresponding tasks;

[0053] 8) Cerebellum outputs task execution results to users and also feeds back to brain;

[0054] 9) Brain updates memory module.

[0055] In steps 4), 5), 6) above, the planning experts, grounding experts and coding experts jointly use the same memory module; the skills stored in the memory module corresponding to the programmable metasurface include but are not limited to: beam focusing, human 3D skeleton recognition, user positioning, user behavior recognition, user health detection, object positioning, robot positioning, communication link enhancement, robot control, etc. Skills; for the cerebellum module (ID meta ,C)) can be represented by a function π:

[0056] (ID meta ,C)=π(l,other inputs;KG) (1)

[0057] Where ID meta is the index of the programmable metasurface; C is the coding mode; l is the natural language prompt based on stored knowledge; other inputs represent other inputs; KG represents the knowledge graph; the function π converts the natural language prompt l based on the stored knowledge (in the form of a knowledge graph (KG)) and other inputs into the index ID meta and the spatial (or spacetime) coding mode C of the selected programmable metasurface.

[0058] Embodiment one

[0059] In this embodiment, the autonomous electromagnetic agent system architecture based on programmable metasurface is as shown in Figure 1 The workflow of the autonomous electromagnetic agent system based on programmable metasurface is as shown in Figure 2 The brain module of the metaAgent is as shown in Figure 1 The cerebellum module of the metaAgent is as shown in Figure 1 The bottom left. Figure 3 The visualization interface designed for the autonomous electromagnetic agent system based on programmable metasurface of the present application.

[0060] In this embodiment, the interaction process of the user and the metaAgent system is as shown in Figure 2 Specifically as follows:

[0061] 1) The user inputs an instruction through voice: "Please detect Alice's breathing frequency!";

[0062] 2) The metaAgent first starts the perception environment (perception expert submodule) to analyze various modal data and obtains the user task language description.

[0063] 3) After the planning expert submodule of the metaAgent receives the task language description of the perception expert submodule, it decomposes it into a subtask sequence: 1) locate Alice; 2) measure Alice's breathing frequency.

[0064] 4) The grounding expert submodule assigns corresponding devices and skills to the subtask sequence: 1) locate Alice (device: 2.4GHz-SPM, skill: user positioning); 2) measure Alice's breathing frequency (device: 5.5GHz-SPM, skill: h breath detection).

[0065] 5) The coding expert submodule writes corresponding python code according to the above results (corresponding devices and skills).

[0066] 6) The cerebellum module receives the python code and executes it.

[0067] 7) If the code execution is successful, the result is output to the user through the microphone, and the result is passed to the planning expert submodule for task summary and analysis. If the planning expert submodule thinks that there is no need for further planning, it outputs "END" and ends this interaction.

[0068] 8) If the code writing is incorrect, it is returned to the coding expert submodule for re-writing.

[0069] Finally, it is to be understood that the embodiments are presented by way of example only, and that the invention is not limited to the embodiments. It will be apparent to persons skilled in the art that various modifications and variations can be made to the present application without departing from the scope or spirit of the application. Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. For this reason it is to be understood that the application is not limited to the specific examples indicated and that various modifications can be made which fall within the scope of the application as defined by the appended claims.

Claims

1. An autonomous electromagnetic agent system based on programmable metasurfaces, characterized by, The system comprises a cerebellum module based on programmable metasurface and a brain module based on large capacity base model, and has autonomous beam steering capability with human-level intelligence; The cerebellum module uses programmable metasurface to perform electromagnetic wave steering tasks; The brain module is used for generating high-level task strategies, and decomposing the high-level tasks to issue executable subtask sequences to the cerebellum module; The brain module and the cerebellum module work on two hosts respectively, and data transmission is performed between the two through a network communication protocol; The brain module comprises a memory submodule, a perception expert submodule, a planning expert submodule, a grounding expert submodule and a coding expert submodule; wherein, The memory submodule is composed of an action library and a memory library; the action library is used for storing a device library composed of all devices capable of being controlled by the autonomous electromagnetic intelligent agent system and a skill library composed of all skills capable of being executed by the system; the memory library is used for storing all historical records during system operation and environment information corresponding to the operating environment; The perception expert submodule is used for taking continuously collected multi-modal data signals as input, and outputting semantic results to the planning expert submodule after processing; the signal processing method adopts a deep neural network, including using a model pre-trained on a large-scale data set to process images and sounds, or using a programmable metasurface to detect the environment and using an artificial neural network trained through supervised learning to perceive microwave signals; The planning expert submodule analyzes natural language input from the perception expert submodule or the coding expert submodule using the knowledge provided by the memory submodule, and decomposes the natural language input into a series of executable subtask sequences; the executable subtask sequences to be executed are transmitted to the grounding expert submodule in the form of natural language; The coding expert submodule is used for taking the task sequence generated by the planning expert submodule, the functions of the required actions generated by the grounding expert module and the natural language devices as input; writing codes and automatically running on the host; the coding expert submodule generates two kinds of outputs in natural language: external output is used for communicating with human users through a chat tool; internal output is reasoning in natural language after discussion by multiple expert submodules and provided to the planning expert submodule and the grounding expert submodule; The cerebellum module comprises programmable metasurfaces working in different frequency bands, field programmable gate array control modules, general-purpose software radio peripheral modules, transmitting antennas and receiving antennas, visual sensing modules, text input modules and voice modules; wherein: The programmable metasurface is composed of a plurality of independently controllable control units and works in multiple frequency bands; each programmable metasurface is controlled by a field programmable gate array control module; The general-purpose software radio peripheral module is used for real-time collection of microwave data through the transmitting and receiving antennas; the general-purpose software radio peripheral module is connected with the host; and the control algorithms are all run on the host; The visual perception module is used for collecting image information of the environment; the text input module is used for collecting text input information of the user; and the voice module is used for real-time interaction between the user and the system through voice.

2. The autonomous electromagnetic agent system based on programmable hypersurfaces of claim 1, wherein, In the brain module, the memory bank in the memory sub-module is used to store system information including a visual semantic map of the environment in which the user is located in the form of a knowledge graph.

3. The autonomous electromagnetic agent system based on programmable hypersurfaces as claimed in claim 1 wherein, The perception expert sub-module continuously collects multi-modal data including sound recorded by a microphone, images recorded by a camera, text input by a keyboard, and microwave signals collected by a software-defined radio.

4. The autonomous electromagnetic agent system based on programmable hypersurfaces as claimed in claim 1, wherein, Further, the planning expert sub-module is implemented based on an LLM that provides context demonstrations, and is improved according to human feedback opinions stored in the memory bank of the memory sub-module.

5. The autonomous electromagnetic agent system based on programmable hypersurfaces as claimed in claim 1 wherein, The structure of the control unit of the programmable metasurface includes a rectangular metal resonant patch, a phase bias line, a dielectric substrate, and a metal reflector plate; and the working frequency bands include 2.4 GHz, 5.5 GHz, and 9.7 GHz.

6. The autonomous electromagnetic agent system based on programmable hypersurfaces as claimed in claim 1 wherein, The visual perception module is composed of multiple ZED2 depth cameras; the text input module is composed of a keyboard connected to the host computer; and the voice module is composed of a microphone and a speaker.

7. The autonomous, electromagnetic agent system based on programmable hypersurfaces as defined in claim 1, wherein, The cerebellum module is represented by a function π as follows: ID meta C) = π(l, other inputs; KG) where ID meta is the index of the programmable hypersurface; C is the encoding pattern; / is the natural language cue based on stored knowledge; other inputs represent other inputs; KG represents the knowledge graph; the function π converts the natural language cue / and other inputs based on the knowledge stored in the form of the knowledge graph into the index ID meta of the selected programmable hypersurface and the spatial or spatio-temporal encoding pattern C.

8. The autonomous, electromagnetic agent system based on programmable hypersurfaces of claim 1, wherein, The working steps of the autonomous electromagnetic intelligent agent system include: 1) The brain module receives various sensing model data transmitted by the cerebellum module; 2) The perception expert sub-module receives multi-modal data, analyzes the data through corresponding deep learning network models, and converts the output results into semantic information in a specific format; 3) The planning expert sub-module receives semantic information output by the perception expert sub-module, decomposes the task according to the knowledge of the memory sub-module, and generates a series of executable sub-task sequences; 4) The grounding expert sub-module allocates corresponding executable devices and corresponding skills for each sub-task according to the knowledge of the memory sub-module; 5) The coding expert sub-module compiles program code executable on the host computer according to the sub-task sequence and the allocated devices and skills; 5) The cerebellum module receives program code issued by the brain module and executes it on the host computer; 7) The semantic programmable metasurface executes the corresponding code to complete the corresponding task; 8) The cerebellum module outputs the task execution result to the user, and also feeds back to the brain module; 9) The brain module updates the memory sub-module.

Citation Information

Patent Citations

  • Intelligent electromagnetic metasurface controlled by voice recognition

    CN114171927A

  • Incremental multi-mode expert large model architecture based on integrated model

    CN118607665A