Multi-agent cooperation method and system based on intelligent glasses

Through the multi-agent collaborative system, smart glasses adaptively select agents for data processing, solving the problem of insufficient computing power, achieving a dynamic balance between low power consumption, high real-time performance and information sufficiency, and supporting extremely fast feedback of multi-task events.

CN120653449AInactive Publication Date: 2025-09-16SHARGE TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511127541.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-09-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Due to insufficient computing power, smart glasses cannot meet the needs of multi-tasking or deep reasoning tasks, and the end-cloud collaboration causes large delays in data training models.

Method used

Through the multi-agent collaborative system, smart glasses obtain real-time data and adaptively select glasses, mobile devices, PCs and cloud agents for reasoning. Real-time tasks are processed locally on the glasses, and deep tasks are processed in the cloud. TTS technology is used to transform feedback results.

Benefits of technology

It achieves a dynamic optimal balance between low power consumption, high real-time performance and sufficient information, and can achieve extremely fast feedback when multiple tasks are running in parallel, without being affected by deep reasoning tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653449A_ABST
    Figure CN120653449A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-agent cooperation method and system based on intelligent glasses. The multi-agent cooperation method comprises the following steps: analyzing real-time data of the intelligent glasses to obtain a task event; obtaining a preset expected delay budget and a task predicted time length of the task event, matching the expected delay budget and the task predicted time length with a preset endpoint capability of each agent, and selecting at least one target agent from the multiple agents; inputting the real-time data and the task event into at least one target agent for reasoning to generate a reasoning result, analyzing the real-time data by using a multi-agent cooperation system to match the task event, and determining the task event according to the expected delay budget of the task event and the task predicted time length. And the intelligent agent is adaptively selected to perform reasoning on the task event and the real-time data, so that dynamic optimal balance among low power consumption, high real-time performance and sufficient information is finally realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and wearable device technology, and in particular to a multi-agent collaboration method and system based on smart glasses. Background Art

[0002] With the increasing popularity of smart wearable devices, smart glasses can acquire external environmental data through built-in sensors, image acquisition devices, and audio acquisition devices. These data is then processed by an AI-powered agent built into the glasses to perform tasks such as voice broadcasting, meeting recording, keyword recognition, and real-time navigation. However, due to their size limitations, the computing power of smart glasses is far less than that of other computing devices. Therefore, when faced with multitasking or deep reasoning tasks, the built-in agent in smart glasses cannot meet the requirements.

[0003] Based on the above defects, existing technologies use "end-cloud collaboration" to solve the problem of insufficient computing power of smart glasses. However, this results in the model trained with the data generated by the current task being unable to solve the current problem in a timely manner. Moreover, due to network limitations, there will be large delays in both data interaction and model calling. Summary of the Invention

[0004] To address the problems of existing smart glasses, the present invention provides a multi-agent collaboration method and system based on smart glasses. Real-time data is acquired through smart glasses, and the multi-agent collaboration system analyzes the real-time data to match task events. Based on the expected delay budget of the task event and the estimated duration of the task, agents are adaptively selected to reason about the task event and the real-time data. This ultimately achieves a dynamic optimal balance between low power consumption, high real-time performance, and sufficient information. Specifically, the present invention is implemented through the following technical solutions: In one aspect, the present invention provides a multi-agent collaboration method based on smart glasses, comprising: On the other hand, the present invention also provides a multi-agent collaboration system based on smart glasses, which includes: S10: Analyze the real-time data of the smart glasses to obtain task events; S20: Obtaining the expected delay budget and the expected task duration preset for the task event, matching the expected delay budget and the expected task duration with the preset endpoint capabilities of each agent, and selecting at least one target agent from the multiple agents; S30: Inputting the real-time data and the task event into at least one target agent for reasoning to generate a reasoning result.

[0005] Furthermore, the multiple agents include: glasses agent, mobile device agent, PC agent and cloud agent.

[0006] Furthermore, when there is more than one target agent, step S30 includes: S301: Splitting the real-time data based on a preset time threshold to obtain multiple data groups; S302.A: Sequentially obtain data from multiple data groups, and call the glasses agent to perform task reasoning according to the task event to obtain real-time reasoning results; S302.B: Transmitting the data sets to a cloud database in sequence; S303: Calling the remaining target agents to perform task reasoning according to the task event and the multiple data groups to obtain deep reasoning results.

[0007] Furthermore, the step S10 includes: S101: Perform intent recognition based on real-time data from smart glasses and intercept key events; S102: Perform task matching on the key events to obtain corresponding task events.

[0008] Furthermore, in step S302.A, after obtaining the real-time reasoning result, the smart glasses convert the real-time reasoning result into one or more feedbacks including vibration, language or image according to TTS technology.

[0009] On the other hand, the present invention also provides a multi-agent collaboration system based on smart glasses, which includes: Smart glasses, a task event triggering unit, an agent matching unit, an agent calling unit, and multiple agents; The smart glasses are used to obtain real-time data and output real-time feedback; The task event triggering unit is used to: analyze the real-time data of the smart glasses to obtain the task event; The agent matching unit is used to: obtain the expected delay budget and the expected task duration preset for the task event, match the expected delay budget and the expected task duration with the preset endpoint capabilities of each agent, and select at least one target agent from the multiple agents; The agent calling unit is used to input the real-time data and the task event into at least one target agent for reasoning, and generate a reasoning result.

[0010] Furthermore, the agent calling unit includes: Data splitting subunit: used for splitting the real-time data according to a preset time threshold to obtain multiple data groups; Glasses agent calling subunit: used to sequentially obtain data from multiple data groups, and call the glasses agent according to task events to perform task reasoning and obtain real-time reasoning results; Data transmission subunit: used for transmitting the data groups to the cloud database in sequence; Deep reasoning agent calling subunit: used to call the remaining target agents to perform task reasoning based on the task event and the multiple data groups to obtain deep reasoning results.

[0011] Furthermore, the task event triggering unit includes: Intent recognition subunit: used to identify intent based on real-time data from smart glasses and intercept key events; Task matching subunit: performs task matching on the key events to obtain corresponding task events.

[0012] Furthermore, the agent calling unit further includes: The TTS conversion subunit is used to convert the real-time reasoning result into one or more feedbacks such as vibration, language or image after obtaining the real-time reasoning result.

[0013] The present invention obtains real-time data through smart glasses, analyzes real-time data with a multi-agent collaborative system to match task events, and adaptively selects agents to reason about task events and real-time data based on the expected delay budget of task events and the estimated duration of tasks. When there are only real-time tasks, the reasoning results of real-time tasks are converted from text to voice form through TTS (Text to Speech) technology. At the same time, when one or more tasks appear, the real-time data is split into multiple groups based on a preset time threshold. On the one hand, the multiple groups of data are processed in real time in sequence, and on the other hand, the multiple groups of data are uploaded to the cloud database synchronously. When all the real-time data are transmitted, the multiple groups of data are transmitted from the cloud to the agent matching the task event for processing, ultimately achieving a dynamic optimal balance between low power consumption, high real-time and sufficient information. When multiple task events are carried out in parallel, the smart glasses can also achieve extremely fast feedback and are not affected by deep reasoning tasks.

[0014] For better understanding and implementation, the present invention is described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a structural diagram of the multi-agent collaboration system based on smart glasses provided by the present invention; Figure 2 To execute Figure 1 Flowchart of the collaboration method of the multi-agent collaboration system shown; Figure 3 It is a structural block diagram of the task event triggering unit of the present invention; Figure 4 This is an execution flow chart of the intelligent agent calling unit of the present invention when triggering a deep reasoning task. DETAILED DESCRIPTION

[0016] The present invention proposes a multi-agent collaboration method and system based on smart glasses. Figure 1 The system includes: smart glasses, a task event triggering unit 10, an agent matching unit 20, an agent calling unit 30 and multiple agents.

[0017] In the present invention, multiple agents include: glasses agent, mobile device agent, PC-side agent and cloud-side agent. When the system is initialized, multiple agents are registered in the hardware devices that realize communication connection with each other. For example, the glasses agent is registered in the system of smart glasses, while the mobile device agent is registered in devices such as Raspberry Pi and mobile phones. The PC-side is a small or medium-sized computer device. Each hardware device is connected to the cloud network or interconnected based on wireless communication technologies such as Bluetooth communication. Based on the above premise, after the smart glasses obtain the audio and video data, the task event triggering unit 10 first analyzes the obtained data in real time. When there is a task event in the real-time data, the agent matching unit 20 adaptively selects the agent according to the expected delay budget and estimated task duration of the pre-stored task event. Then the agent calling unit 30 infers the task event and the real-time data based on the data obtained by the smart glasses and the agent selected by the agent matching unit 20, and finally achieves a dynamic optimal balance between low power consumption, high real-time and sufficient information. Specifically, the working process of each module of the multi-agent collaboration system based on smart glasses is as follows: Figure 2 As shown: The task event triggering unit 10 is configured to execute step S10: parsing the real-time data of the smart glasses to obtain a task event.

[0018] Smart glasses have their own audio and video acquisition devices, and can generate image output through the lenses, as well as audio devices to realize voice broadcasting. Generally, smart glasses are awakened by a keyword trigger mechanism. For example, when the voice input "real-time translation" is input, when the audio device recognizes "translation", it matches the translation task event; or when "enable real-time translation and meeting record analysis", it matches the two task events "translation" and "meeting record analysis". Preferably, in order to prevent the problem of accidental touch and untimely triggering, the present invention can also replace the trigger mechanism with a trigger form based on the semantic model, please refer to Figure 3 At this time, the task event triggering unit 10 includes an intention recognition subunit 101 and a task matching subunit 102; The agent calling unit 30 is used to execute step S30: input the real-time data and the task event into at least one target agent for reasoning, and generate a reasoning result.

[0019] The intention recognition subunit 101 is used to perform step S101: perform intention recognition based on the real-time data of the smart glasses and intercept key events; The task matching subunit 102 is configured to execute step S102: performing task matching on the key event to obtain a corresponding task event.

[0020] The agent matching unit 20 is used to execute step S20: obtain the expected delay budget and expected task duration preset for the task event, match the expected delay budget and expected task duration with the preset endpoint capabilities of each agent, and select at least one target agent from multiple agents.

[0021] The endpoint capabilities of each agent in this application are shown in the following table:

[0022] According to the above content, when only real-time tasks are triggered, the glasses intelligent agent performs inference on the acquired real-time data, such as: wake-up word detection, scene / speaker reminders, and quick shot result broadcasting.

[0023] However, when triggering a deep reasoning task, multiple agents are required. Figure 4 Step S30 includes: S301: Splitting the real-time data based on a preset time threshold to obtain multiple data groups; Due to the computing power limitations of smart glasses, large quantities of data inference cannot be performed simultaneously when acquiring real-time data, data inference, and data transmission. In order to balance the immediacy of the glasses' intelligent body reasoning and the stability of data transmission, the acquired data must first be split and stored.

[0024] S302.A: Sequentially obtain data from multiple data groups, and call the glasses agent to perform task reasoning according to the task event to obtain real-time reasoning results; The glasses intelligent body infers the data of each data group according to the time sequence, obtains the inference result, and outputs the inference result through the output device. Preferably, after obtaining the real-time inference result, the TTS conversion subunit converts the real-time inference result into one or more feedbacks in the form of vibration, language or image.

[0025] S302.B: Transmitting the data sets to a cloud database in sequence; S303: Calling the remaining target agents to perform task reasoning according to the task event and the multiple data groups to obtain deep reasoning results.

[0026] Since deep reasoning tasks do not require immediate processing, they are assigned to corresponding agents through the cloud. Each agent calls the preset API based on the task event to reason on the acquired data, and finally outputs the reasoning results through the hardware device where the agent is located.

[0027] The multi-agent collaborative system illustrates the entire collaborative process of a meeting. Five minutes before the meeting begins, the meeting reminder task is triggered. The glasses agent receives the time signal that triggers the voice announcement and converts the text content into a voice announcement through TTS. During the meeting, the agent accesses the meeting minutes in real time. The mobile agent triggers the "meeting summary" task and transmits cloud data to the mobile device every 1-2 minutes. The mobile agent uses the meeting summary API to perform voice analysis on the text content, summarize it, and generate graphic cards that are displayed on the mobile agent's display. After the meeting, the cloud transmits all the meeting minutes data to the PC. The PC agent obtains the complete meeting minutes data in JSON format and creates Mardown+ charts based on the meeting minutes data to complete the meeting minutes. During its idle time, the cloud agent performs cross-day task processing on the meeting minutes data, implementing knowledge graph induction, to-do link reasoning, and generating a comprehensive meeting report. Ultimately, this allows for real-time meeting broadcast, mid-session summary, full record, and post-event analysis.

[0028] The present invention obtains real-time data through smart glasses, analyzes real-time data with a multi-agent collaborative system to match task events, and adaptively selects agents to reason about task events and real-time data based on the expected delay budget of task events and the estimated duration of tasks. When there are only real-time tasks, the reasoning results of real-time tasks are converted from text to voice form through TTS (Text to Speech) technology. At the same time, when one or more tasks appear, the real-time data is split into multiple groups based on a preset time threshold. On the one hand, the multiple groups of data are processed in real time in sequence, and on the other hand, the multiple groups of data are uploaded to the cloud database synchronously. When all the real-time data are transmitted, the multiple groups of data are transmitted from the cloud to the agent matching the task event for processing, ultimately achieving a dynamic optimal balance between low power consumption, high real-time and sufficient information. When multiple task events are carried out in parallel, the smart glasses can also achieve extremely fast feedback and are not affected by deep reasoning tasks.

[0029] Based on the same inventive concept described above, the present invention further provides an electronic device, which can be a terminal device such as a server, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.). The device includes one or more processors and a memory, wherein the processor is configured to execute a program to implement the multi-agent collaboration method based on smart glasses described above; and the memory is configured to store a computer program executable by the processor.

[0030] Based on the same inventive concept, the present invention also provides a computer-readable storage medium, corresponding to the aforementioned embodiment of a multi-agent collaboration method based on smart glasses, wherein the computer-readable storage medium stores a computer program thereon, and when the program is executed by a processor, the steps of a multi-agent collaboration method based on smart glasses recorded in any of the above embodiments are implemented.

[0031] The present invention may take the form of a computer program product implemented on one or more storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing program code. Computer-usable storage media include both permanent and non-permanent, removable and non-removable media, and may utilize any method or technology for information storage. Information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0032] The above-described embodiments merely represent several implementations of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous modifications and improvements without departing from the spirit of the present invention, and the present invention is intended to encompass such modifications and variations.

Claims

1. A multi-agent collaboration method based on smart glasses, characterized in that: include: S10: Analyze the real-time data of the smart glasses to obtain task events; S20: Obtaining the expected delay budget and the expected task duration preset for the task event, matching the expected delay budget and the expected task duration with the preset endpoint capabilities of each agent, and selecting at least one target agent from the multiple agents; S30: Inputting the real-time data and the task event into at least one target agent for reasoning to generate a reasoning result.

2. The multi-agent collaboration method based on smart glasses according to claim 1, characterized in that: The multi-agent comprises: Glasses intelligent agent, mobile device intelligent agent, PC intelligent agent and cloud intelligent agent.

3. The multi-agent collaboration method based on smart glasses according to claim 2, characterized in that: When there is more than one target agent, step S30 includes: S301: Splitting the real-time data based on a preset time threshold to obtain multiple data groups; S302.A: Sequentially obtain data from multiple data groups, and call the glasses agent to perform task reasoning according to the task event to obtain real-time reasoning results; S302.B: Transmitting the data sets to a cloud database in sequence; S303: Calling the remaining target agents to perform task reasoning according to the task event and the multiple data groups to obtain deep reasoning results.

4. The multi-agent collaboration method based on smart glasses according to claim 3, characterized in that: The step S10 includes: S101: Perform intent recognition based on real-time data from smart glasses and intercept key events; S102: Perform task matching on the key events to obtain corresponding task events.

5. The multi-agent collaboration method based on smart glasses according to claim 4, characterized in that: In step S302.A, after obtaining the real-time reasoning result, the smart glasses convert the real-time reasoning result into one or more feedbacks including vibration, language or image according to TTS technology.

6. A multi-agent collaboration system based on smart glasses, characterized in that: include: Smart glasses, a task event triggering unit, an agent matching unit, an agent calling unit, and multiple agents; The smart glasses are used to obtain real-time data and output real-time feedback; The task event triggering unit is used to: analyze the real-time data of the smart glasses to obtain the task event; The agent matching unit is used to: obtain the expected delay budget and the expected task duration preset for the task event, match the expected delay budget and the expected task duration with the preset endpoint capabilities of each agent, and select at least one target agent from the multiple agents; The agent calling unit is used to input the real-time data and the task event into at least one target agent for reasoning, and generate a reasoning result.

7. The multi-agent collaboration system based on smart glasses according to claim 6, characterized in that: The agent calling unit includes: Data splitting subunit: used for splitting the real-time data according to a preset time threshold to obtain multiple data groups; Glasses agent calling subunit: used to sequentially obtain data from multiple data groups, and call the glasses agent according to task events to perform task reasoning and obtain real-time reasoning results; Data transmission subunit: used for transmitting the data groups to the cloud database in sequence; Deep reasoning agent calling subunit: used to call the remaining target agents to perform task reasoning based on the task event and the multiple data groups to obtain deep reasoning results.

8. The multi-agent collaboration system based on smart glasses according to claim 7, characterized in that: The task event trigger unit includes: Intent recognition subunit: used to identify intent based on real-time data from smart glasses and intercept key events; Task matching subunit: performs task matching on the key events to obtain corresponding task events.

9. The multi-agent collaboration system based on smart glasses according to claim 8, characterized in that: The agent calling unit also includes: The TTS conversion subunit is used to convert the real-time reasoning result into one or more feedbacks such as vibration, language or image after obtaining the real-time reasoning result.

Citation Information

Patent Citations

  • Split type intelligent glasses control method and system and medium

    CN112346862A

  • Video or image analysis method and device based on end-side cloud adaptive collaboration

    CN117278552A