An obesity intervention system and method based on behavioral science multi-agent cooperation

By using a multi-agent collaborative system based on behavioral science, combined with a large language model and IoT sensors, the system can identify and resolve deep-seated behavioral disorders in users, improve the personalization and interactive experience of obesity management, provide real-time environmental intervention, and significantly improve user compliance.

CN122224401APending Publication Date: 2026-06-16ZHEJIANG UNIV CITY COLLEGE

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIV CITY COLLEGE
Filing Date
2026-03-09
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing obesity management systems lack in-depth behavioral intervention, have low personalization, poor interactive experience, cannot effectively address users' deep psychological and environmental barriers, and lack immediate physical feedback.

Method used

It adopts a multi-agent collaborative system based on behavioral science, combined with a large language model and IoT sensors. Data is collected through edge-side intelligent voice interaction devices, and the cloud-based multi-agent system analyzes user behavioral obstacles, generates personalized nutrition suggestions and provides hardware feedback, and integrates streaming interaction technology to reduce latency.

Benefits of technology

It enables the identification and resolution of users' deep-seated behavioral disorders, improves the effectiveness and interactive experience of personalized interventions, provides immediate environmental interventions, and significantly improves adherence to obesity management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122224401A_ABST
    Figure CN122224401A_ABST
Patent Text Reader

Abstract

The application discloses an obesity intervention system and method based on multi-agent cooperation of behavioral science, and relates to the technical field of digital health.The system comprises an end-side intelligent voice interaction device and a cloud-side multi-agent cognitive hub; the end-side intelligent voice interaction device is based on an embedded micro control platform and integrates an offline voice wake-up module, a streaming voice interaction module and an Internet of Things control module; and the cloud-side multi-agent cognitive hub is based on a large language model and comprises a disorder identification agent, a strategy execution agent, a nutrition suggestion agent and a metabolism monitoring agent.The application identifies deep behavioral disorders (ability, opportunity and motivation) of a user through a COM-B behavioral science model, generates a personalized intervention strategy by means of multi-agent cooperation, and executes the strategy in a closed loop in the form of voice dialogue and environmental hardware feedback through end-side interaction.The application solves the problems of lack of deep motivation intervention, delayed interaction feedback and inability to realize soft and hardware linkage intervention in the prior art weight management system, and significantly improves the compliance and effectiveness of obesity intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital health technology, specifically to an obesity nutrition management and behavior intervention system that combines multi-agent technology based on a large language model with an embedded voice interaction terminal. Background Technology

[0002] Overweight and obesity have become a global public health crisis. Effective management of obesity relies not only on medication or surgery, but also on continuous lifestyle interventions and scientific nutritional management. Traditional weight management methods mainly depend on human nutritionists or rule-based mobile applications, typically involving calorie recording, general meal plan recommendations, and weight tracking.

[0003] In recent years, with breakthroughs in deep learning technology, artificial intelligence technologies, represented by large language models, have demonstrated enormous potential in the healthcare field. Large language models possess powerful natural language understanding and generation capabilities, enabling them to process unstructured medical texts and conduct multi-turn dialogues. However, single large language models still have limitations when handling complex medical intervention tasks. To address the limitations of single models, multi-agent systems have emerged. These systems construct multiple agents with independent perception, decision-making, and execution capabilities, utilizing role-playing and collaborative reasoning mechanisms to solve complex problems. In the field of medical diagnosis, research has emerged using multi-agent systems to simulate doctor consultations, demonstrating that debate and collaboration among agents can significantly improve diagnostic accuracy.

[0004] In today's increasingly sophisticated digital healthcare landscape, existing weight management solutions suffer from significant pain points. First, they lack in-depth behavioral intervention. Most systems only offer superficial suggestions like "eat less and move more," failing to analyze the underlying psychological barriers (such as stress eating) or environmental obstacles (such as lack of access to healthy foods) that users cannot overcome based on behavioral science. Second, they lack personalization. Traditional systems struggle to handle complex individual differences and cannot dynamically adjust strategies based on real-time metabolic data and emotional feedback like a human nutritionist, often leading to poor user adherence. Finally, the interactive experience is poor. Pure text-based interactions lack emotional warmth, and traditional cloud-based communication devices often suffer from high latency and lack interaction with the physical world, making it difficult to provide immediate environmental intervention when users experience eating urges. Summary of the Invention

[0005] Based on the above background, the purpose of this invention is to provide an obesity intervention system and method based on behavioral science multi-agent collaboration, which can be used to build an intelligent health manager with professional behavioral cognitive ability and real-time interaction ability.

[0006] Therefore, according to a first aspect of the present invention, the present invention adopts the following technical solution:

[0007] An obesity intervention method based on behavioral science and multi-agent collaboration includes the following steps: S1, collecting user voice commands and dialogue content through an edge-side intelligent voice interaction device, and collecting user physiological metabolic data and environmental data through associated IoT sensors to construct a multimodal input data stream; S2, a cloud-based multi-agent system receives the multimodal input data stream, analyzes the voice dialogue content and physiological data based on a COM-B behavioral science model, and identifies user behavioral impairment characteristics in three dimensions: ability, opportunity, and motivation; S3, based on the identified behavioral impairment characteristics, matching intervention strategies through an impairment-policy mapping database, and generating personalized nutritional advice text and hardware control commands using a multi-agent collaboration mechanism driven by a large language model; S4, converting the nutritional advice text into a voice stream and sending it to the edge-side intelligent voice interaction device for playback, while simultaneously providing physical feedback from the peripheral devices of the edge-side intelligent voice interaction device or associated IoT devices according to the hardware control commands; S5, collecting user feedback information on the intervention, updating the user dynamic profile using a long short-term memory mechanism, and correcting subsequent intervention strategies.

[0008] Furthermore, in step S2, the specific processing procedure for multi-agent obstacle recognition includes:

[0009] Semantic analysis of users' natural language is performed using a pre-trained Large Language Model (LLM); semantic features are mapped to the three dimensions of the COM-B model to identify specific disorder types, including psychogenic eating, lack of planning ability, or environmental triggers; combined with users' historical dialogue records, users' lifestyle preferences are extracted through the LangChain memory module to generate structured, specific disorder diagnoses.

[0010] Furthermore, the multi-agent system includes the following sub-modules:

[0011] Obstacle recognition agent: responsible for analyzing user input and detecting specific obstacles in nutritional activities;

[0012] The strategy execution agent: Based on the obstacle diagnosis report, it retrieves the behavior change strategy library and selects intervention methods including cognitive restructuring, environmental reshaping, or incentive reinforcement.

[0013] A metabolic monitoring agent analyzes data on body fat percentage, blood glucose, and hormone levels to generate physiological constraints.

[0014] Nutritional advice agent: Under the guidance of physiological constraints and intervention methods, it uses thought chains to generate dialogue responses that match the user's emotional state.

[0015] Furthermore, step S3 specifically includes:

[0016] S31. The obstacle recognition agent analyzes user input and detects specific obstacles in nutritional behavior.

[0017] S32. The metabolic monitoring agent analyzes body fat percentage, blood glucose and hormone level data to generate physiological constraints.

[0018] S33. The strategy execution agent retrieves the Behavior Change Wheel (BCW) strategy library based on the obstacle diagnosis report and selects intervention methods including cognitive restructuring, environmental reshaping, or incentive reinforcement.

[0019] S34. Nutritional advice: Under the guidance of physiological constraints and intervention methods, the intelligent agent uses thought chains to generate dialogue responses that match the user's emotional state.

[0020] Furthermore, in step S3, the multi-agent cooperative mechanism adopts a formalized agent definition, wherein each agent... It can be represented as a quintuple:

[0021]

[0022] in The model consists of an LLM architecture, memory modules, and a task-specific adapter. )composition, For the task objectives of the intelligent agent, The environmental context in which the agent exists. The input data received by the intelligent agent. This refers to the output response of the intelligent agent.

[0023] Furthermore, in step S3, the overall collaborative output of the multi-agent system... The following generation process is satisfied:

[0024]

[0025] in The goal is to create a personalized health solution system based on user profiles. For environment context parameters, The total input data received by the system includes user voice and text, as well as IoT sensor data. For a collection of intelligent agents participating in the collaboration, It is a set of collaborative channels between intelligent agents. This is a system generation function under cooperative channel constraints.

[0026] According to a second aspect of the present invention, the present invention adopts the following technical solution:

[0027] An obesity intervention system based on behavioral science multi-agent collaboration includes:

[0028] The device includes an edge-side intelligent voice interaction system for collecting user voice commands and environmental data, and executing voice feedback and physical environment control. Based on an ESP32 microcontroller unit, it integrates offline voice wake-up, streaming voice interaction modules, and an IoT control interface. A cloud-based multi-agent cognitive center, deployed on a cloud server, receives edge-side data, performs reasoning based on behavioral science models, and generates intervention strategies. An obstacle recognition module analyzes user input and detects specific obstacles in nutritional behavior based on a COM-B model. A strategy execution module invokes behavioral science strategies from an obstacle-strategy mapping database based on identified obstacles. A metabolic monitoring module integrates data from medical IoT devices, establishes a dynamic metabolic prediction model, and provides biochemical indicator constraints. A nutrition advice module generates personalized nutrition advice based on user background, biochemical indicator constraints, and strategy execution results.

[0029] Furthermore, the system also includes a user profile database, which is built on MongoDB and uses a long short-term memory mechanism to persistently store users' dynamic profiles and conversation history, ensuring the continuity of long-term management.

[0030] Furthermore, the edge-side intelligent voice interaction device adopts a streaming voice processing architecture, which uploads the user's voice stream to the cloud in real time for intent recognition while receiving the user's voice stream; upon receiving the first token generated in the cloud, it immediately starts voice synthesis and playback to reduce interaction latency.

[0031] Furthermore, the system supports a zero-code access mode, and the multi-agent cognitive center provides a standardized API interface, allowing new behavioral intervention strategies to be defined or new IoT sensors to be connected through configuration files without modifying the underlying code.

[0032] The beneficial effects of this invention are as follows:

[0033] 1. We have created a unique obesity intervention system and method based on behavioral science multi-agent collaboration. For the first time, we have integrated behavioral science theory (COM-B model) into the multi-agent architecture of a large language model, enabling the system to have the analytical logic of a professional health coach and to identify and resolve users' deep-seated behavioral disorders.

[0034] 2. By utilizing the cost-effectiveness and streaming interaction technology based on the open-source ESP32 platform, low-latency natural voice dialogue was achieved, and environmental intervention was creatively combined with hardware feedback, providing a good solution for the management and intervention of obesity.

[0035] 3. It addresses the pain point of traditional methods lacking in-depth motivational analysis and overcomes the limitation of pure software interaction lacking physical feedback, significantly improving compliance and effectiveness of obesity intervention. It can be widely applied in home health management, chronic disease rehabilitation and other fields. Attached Figure Description

[0036] Figure 1 This is a flowchart of an obesity intervention method based on behavioral science multi-agent collaboration, according to an embodiment of the present invention.

[0037] Figure 2 This is a flowchart illustrating the main technical route of an obesity intervention method and system based on behavioral science multi-agent collaboration, according to an embodiment of the present invention. Detailed Implementation

[0038] like Figure 1 and Figure 2 As shown, this invention proposes the construction and implementation of an obesity intervention system and method based on behavioral science multi-agent collaboration. In this embodiment, the end-side device can be a general-purpose smart speaker, a companion robot, or a specially developed health management terminal. For ease of explanation, this embodiment uses a smart voice robot based on the ESP32 platform as an example, but the invention is not limited to this specific hardware model. The main steps include:

[0039] The first step is to construct a cloud-based multi-agent cognitive hub based on behavioral science. The core of this step lies in leveraging the reasoning capabilities of large language models to transform abstract behavioral science theories into executable code logic. The specific implementation is as follows:

[0040] The formal definition and initialization of intelligent agents are crucial steps in transforming abstract behavioral science theories into executable code logic by leveraging the reasoning capabilities of large language models. In practical implementation, the construction of the multi-agent cognitive hub can be rapidly deployed and orchestrated using an intelligent agent development platform (such as Coze or n8n), or it can be customized based on frameworks like LangGraph.

[0041] Formal Definition and Initialization of Intelligent Agents: To address the problem of "illusion" that traditional single models easily produce when handling complex medical logic, this invention adopts a formal definition approach. Each intelligent agent... Defined as a quintuple:

[0042]

[0043] Model( ): The open-source, high-performance generative large language model DeepSeek-R1 was chosen as the base.

[0044] In this embodiment, the model The configuration and deployment are based on the Coze intelligent agent development platform. The platform's visual workflow orchestration function is used to configure the model's system prompts, temperature parameters, and context window size.

[0045] Target( ): Defines the single responsibility of an agent (e.g., identifying COM-B obstacles). In the Coze platform, this is reflected in the "node objectives" or "preset skills" set in the workflow.

[0046] environment( Initialize the context environment of the intelligent agent. Utilize the memory components or database connectors provided by the platform to read the user's static profile (such as gender, age, medical history, etc.) and dynamic dialogue history in real time, as context constraints for model inference.

[0047] enter( ) and output ( Define a standardized data interface. Through a platform-configured Webhook trigger, the input... Defined as a multimodal data stream in JSON format (including speech-to-text and IoT sensor data), the output will be... Defined as structured intervention recommendations and hardware control instructions.

[0048] For obesity intervention scenarios, a knowledge base is built and includes authoritative medical literature such as the "Chinese Dietary Guidelines" and the "Adult Obesity Dietary Guidelines". The model's reasoning accuracy in vertical domains is enhanced through retrieval-enhanced generation technology.

[0049] Construct a core group of intelligent agents. Based on the above definition, instantiate four clearly defined intelligent agents:

[0050] The obstacle recognition agent's system prompts are designed to focus on diagnosing and identifying the patient's primary obstacle. It receives natural language input dialogue from the user and classifies it according to the COM-B model. For example, when a user says, "I only have time tonight to eat a good meal," it identifies it as an "opportunity-physical opportunity (time constraint)" obstacle, rather than simply being greedy.

[0051] The strategy execution agent connects to a pre-built "obstacle-policy mapping database." This database, built on a behavior change wheel framework, stores intervention strategies for different obstacles. For example, for "emotional eating," a "cognitive restructuring" strategy is mapped.

[0052] Metabolic monitoring agent: Responsible for processing numerical data. It analyzes the user's BMI and blood glucose trends in real time. If it detects excessive fluctuations in the user's blood glucose recently, it generates a strong constraint bar.

[0053] Nutritional advice agent: As the final output layer, it receives the output of the above agents as context and uses thought chain technology to generate response text that is both medically compliant and emotionally engaging.

[0054] The second step is to construct and package the edge-side intelligent voice interaction device hardware terminal. The purpose of this step is to create a low-cost, highly reliable physical interaction terminal. Considering the cost requirements for large-scale deployment, this embodiment uses the cost-effective ESP32 chip.

[0055] The third step involves implementing edge-cloud collaboration and streaming interaction through standardized interfaces. This avoids interaction stuttering caused by network latency. Streaming interaction is optimized to achieve near-human conversation speed. The system employs end-to-end streaming processing: while the client is still collecting the second half of the user's speech, the cloud-based automatic speech recognition engine has already begun processing the first half of the text. When the multi-agent system, combined with a large language model, generates the first character, the text-to-speech engine immediately synthesizes it into an audio frame and sends it down. The client plays the first audio packet immediately upon receiving it. Real-world testing shows that this architecture effectively reduces first-character latency.

[0056] Step four: Demonstration of the intervention process in typical application scenarios. The following uses "nighttime binge eating intervention" as an example to explain the system's workflow in detail:

[0057] Scenario Trigger: At 10:30 PM, a user says to the intelligent voice robot, "I just finished working overtime and really want to order some fried chicken." Multimodal Perception: The on-device intelligent voice interaction device uploads the voice stream and converts it into digital data. Simultaneously, it collects relevant historical user profile data. If key data is incomplete, it guides the user through questioning, allowing the user to supplement their answers. Cloud-based Collaborative Reasoning: The obstacle recognition agent analyzes the user's profile and concludes that the user knows fried chicken is unhealthy but still wants to eat it, accompanied by an increased heart rate, classifying it as "motivated-emotional eating (stress compensation)." The metabolic monitoring agent searches the profile and finds that the user has a high risk of hyperlipidemia, and that eating at this time would severely impact metabolism, issuing a "red alert." The strategy execution agent matches a strategy library and decides to adopt a strategy of "environmental boosting (moderating the atmosphere)" + "substitute satisfaction (healthy alternatives)."

[0058] Finally, the AI-generated nutritional advice replied: "You must be exhausted after get off work, and craving something spicy is your body's way of comforting you. But eating fried chicken so late will put extra strain on your body and may even affect your sleep. How about a cup of hot oat milk? It's soothing and helps you sleep. I've already dimmed the lights for you; how about listening to some light music to relax?" The AI ​​voice robot then played the aforementioned caring message.

[0059] Effectiveness Evaluation: The system can ask users if they ordered takeout last night. If the user answers "No, I drank milk and went to sleep," the system will strengthen the weight of this strategy in the user profile.

[0060] Through the above implementation methods, this invention deeply integrates complex behavioral science theories with cutting-edge multi-agent construction technology. Researchers and developers do not need to write complex dialogue logic from scratch; they can quickly adapt to various chronic disease scenarios such as diabetes management and hypertension management simply by configuring cloud-based agent prompts via API. This demonstrates extremely high scalability and practical value.

[0061] The above description of the embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. It should be noted that those skilled in the art can make several improvements and modifications to the present invention without departing from the principles of the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. An obesity intervention method based on behavioral science multi-agent collaboration, characterized in that, Includes the following steps: S1. Collect user's voice commands and dialogue content through the edge intelligent voice interaction device, and collect user's physiological metabolic data and environmental data through the associated IoT sensors to construct a multimodal input data stream; S2. The cloud-based multi-agent system receives the multimodal input data stream, analyzes the voice dialogue content and physiological data based on the COM-B behavioral science model, and identifies the user's behavioral impairment characteristics in the three dimensions of ability, opportunity, and motivation. S3. Based on the identified behavioral disorder features, intervention strategies are matched through the disorder-policy mapping database, and personalized nutrition advice text and hardware control instructions are generated using a multi-agent collaborative mechanism driven by a large language model. S4. Convert the nutrition advice text into a voice stream and send it to the terminal intelligent voice interaction device for playback. At the same time, drive the peripheral device or associated IoT device of the terminal intelligent voice interaction device to make physical feedback according to the hardware control instructions. S5. Collect user feedback on the intervention, update the user dynamic profile using the long short-term memory mechanism, and revise subsequent intervention strategies.

2. The method according to claim 1, characterized in that, In step S2, the specific processing procedure for multi-agent obstacle recognition includes: Using a pre-trained large language model, semantic analysis of the user's natural language is performed, and the semantic features are mapped to the three dimensions of the COM-B model to identify specific obstacle types, including psychogenic eating, lack of planning ability, or environmental triggers. Then, combined with the user's historical dialogue records, the user's lifestyle preferences are extracted through the LangChain memory module to generate a structured and specific classification of obstacle diagnosis.

3. The method according to claim 1, characterized in that, Multi-agent systems include the following sub-modules: Obstacle recognition agent: responsible for analyzing user input and detecting specific obstacles in nutritional activities; The strategy execution agent: Based on the obstacle diagnosis report, it retrieves the behavior change strategy library and selects intervention methods including cognitive restructuring, environmental reshaping, or incentive reinforcement. A metabolic monitoring agent analyzes data on body fat percentage, blood glucose, and hormone levels to generate physiological constraints. Nutritional advice agent: Under the guidance of physiological constraints and intervention methods, it uses thought chains to generate dialogue responses that match the user's emotional state.

4. An obesity intervention system based on behavioral science multi-agent collaboration, used to implement the method described in any one of claims 1-3, characterized in that, The system includes a terminal-side intelligent voice interaction device and a cloud server; the terminal-side intelligent voice interaction device includes: Microcontroller units are used for the core control and signal processing of the system; The audio interaction module, including a microphone array and speakers, is used to enable offline voice wake-up, automatic speech recognition, and text-to-speech. The IoT communication module supports Wi-Fi and Bluetooth connectivity for data uploading and peripheral control. The hardware feedback module includes a display screen and a programmable LED array for visual interaction. The cloud server is deployed with a multi-agent cognitive hub, which communicates with the edge-side intelligent voice interaction device in full-duplex mode via a network.

5. The system according to claim 4, characterized in that: The multi-agent cognitive center can integrate open-source general-purpose large language models (Deepseek-R1, Kimi-32k, etc.) as the reasoning core. The cloud server also includes an obstacle-policy mapping database for storing behavioral intervention strategies based on evidence-based medicine. The system also includes a database for persistently storing users' dynamic profiles and dialogue history. The multi-agent cognitive center adopts a formalized agent definition, where each agent... It can be represented as a quintuple: in The model consists of an LLM architecture, memory modules, and a task-specific adapter. )composition, For the task objectives of the intelligent agent, The environmental context in which the agent exists. The input data received by the intelligent agent. The output response of the intelligent agent; The overall collaborative output of the multi-agent system The following generation process is satisfied: in The goal is to create a personalized health solution system based on user profiles. For environment context parameters, The total input data received by the system includes user voice and text, as well as IoT sensor data. For a collection of intelligent agents participating in the collaboration, It is a set of collaborative channels between intelligent agents. This is a system generation function under cooperative channel constraints.

6. The system according to claim 5, characterized in that, The edge-side intelligent voice interaction device adopts a streaming voice processing architecture. While receiving the user's voice stream, it uploads it to the cloud in real time for intent recognition. Upon receiving the first token generated in the cloud, it immediately starts voice synthesis and playback to reduce interaction latency.

7. The system according to claim 5, characterized in that, The system supports zero-code access mode. The multi-agent cognitive center provides a standardized API interface, which allows new behavioral intervention strategies to be defined or new IoT sensors to be connected through configuration files without modifying the underlying code.