User dining recording method, system and device based on image recognition

By collecting dining data through micro-force sensing arrays, active sonar, and millimeter-wave radar, and combining it with multimodal analysis, the problems of privacy leakage and accurate identification in dining behavior analysis have been solved, achieving privacy-secure, accurate dining behavior analysis and anomaly monitoring.

CN121506399APending Publication Date: 2026-02-10BEIJING CORN YOUPIN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511929800.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies for analyzing dining behavior have problems such as privacy risks, limited visual data collection, inability to accurately identify food texture, and inability to infer physical properties, making it difficult to meet the needs of precise nutritional analysis and dietary management for special populations.

Method used

It employs a micro-force sensor array, an active sonar system, and millimeter-wave radar to collect data on table micro-deformation, oral resonance, and hand movement during dining. Through multimodal data coupling analysis, it identifies dining actions and food attributes, abandoning visual recognition to achieve privacy-preserving and accurate analysis of dining behavior.

Benefits of technology

It achieves privacy-preserving, interference-resistant, and seamless dining behavior analysis, accurately identifies dining content and pace, provides alerts for abnormal behavior, and meets the needs of precise nutrition analysis and dietary management for special populations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506399A_ABST
    Figure CN121506399A_ABST
Patent Text Reader

Abstract

The invention provides a user dining recording method, system and device based on image recognition, and aims to solve the problems that traditional dining records have privacy disclosure risks, cannot capture dynamic behaviors and are large in dining content inference error and the like. Comprising the steps of multi-modal data acquisition, data coupling alignment, feeding rhythm analysis, dining content inference and state feedback. Visual identification is replaced by physical field perception, the recording accuracy is improved while privacy is guaranteed, and the method is suitable for scenes of health management, special crowd monitoring and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of dietary health and relates to a method, system and device for recording user meals based on image recognition. Background Technology

[0002] To ensure users' health, their dining habits are an indispensable part of their dietary health, and good dining habits can play an essential role in users' health.

[0003] Dining behavior analysis requires the direct collection of visual information such as users' faces and food textures, posing a serious risk of privacy breaches and leading to strong user resistance. Furthermore, relevant regulations impose strict restrictions on visual data collection, limiting application scenarios. Traditional methods primarily focus on static states, such as the initial weight of food and the instantaneous posture of the user holding utensils, failing to capture the dynamic and subtle behaviors during the dining process.

[0004] Traditional methods can only infer content from the appearance of food or user records, and cannot distinguish between foods that look similar but have different textures, nor can they infer the physical properties of food; moreover, they rely on active user input, which is highly subjective and prone to errors, making it difficult to meet the needs of precise nutritional analysis and dietary management for special populations. Summary of the Invention

[0005] To address the problems existing in the background technology, this application proposes a user dining record method, system, and device based on image recognition.

[0006] To achieve the above objectives, the technical solution adopted in this application is as follows: On one hand, the present invention provides a user dining record method based on image recognition, comprising the following steps: Data on the microscopic deformation and vibration of the desktop caused by the user dining is collected using a micro-force sensor array; By emitting high-frequency ultrasound waves and receiving echoes through an active sonar system, a cavity resonance model is constructed by analyzing the phase and spectral changes of the echoes, and the characteristics of oral cavity resonance are analyzed. Data on the user's hand movement trajectory is collected using millimeter-wave radar; The above multimodal data are coupled and aligned on a timeline to identify basic dining action units; Based on the identified sequence of action units, the user's eating speed and rhythm are analyzed; By matching motion characteristics with a database of food physical properties, the content of the meal can be inferred. The user status feedback process outputs a list of food consumed, a real-time eating speed curve, an overall eating rhythm assessment, and alerts for abnormal behavior.

[0007] Furthermore, the micro-force sensing array consists of hundreds of miniature capacitive force sensing units, capable of sensing nano-Newton-meter level micro-deformation and vibration of the desktop.

[0008] Furthermore, the basic dining action unit includes at least three of the following: picking up utensils, scooping, clamping, cutting, putting into the mouth, chewing, swallowing, and putting down the utensils.

[0009] Furthermore, the eating speed analysis includes counting the frequency of food delivery actions per unit time and calculating the eating rhythm by combining the average time from food delivery to swallowing.

[0010] Furthermore, the eating rhythm analysis includes calculating the distribution, variance, and autocorrelation of the eating cycle duration to reflect changes in the user's eating habits and state.

[0011] Furthermore, the inference of the meal content includes: Extract sonar texture features and micromechanical vibration features corresponding to chewing events; Map features to the physical attribute space; Graph neural networks are used to match the attribute sequences of multiple consecutive chewing events with a knowledge graph of food physical properties. The food physical property knowledge graph establishes a mapping relationship between food and physical properties, with food as nodes and physical properties such as crispness, hardness, viscosity, and density as edges.

[0012] Furthermore, the multimodal data coupling alignment uses a temporal convolutional network to segment and classify the fused multimodal data stream, identifying basic dining action units.

[0013] On the other hand, the present invention also provides an image recognition-based user dining record system for implementing the above method, comprising: A micro-force sensor array is integrated into the contact surface between the robot base and the dining table to sense microscopic deformation and vibration of the tabletop; An active sonar system, comprising an ultrasonic transmitter and receiver, is integrated inside the robot to transmit and receive high-frequency ultrasonic signals; Millimeter-wave radar module, used to collect the user's hand movement trajectory; The data processing unit is configured to perform multimodal data coupling analysis, action recognition, eating speed and rhythm analysis, and meal content inference. The communication module is used to transmit analysis results.

[0014] Furthermore, the data processing unit includes: The data synchronization module is used to align multimodal data timelines; The action recognition module uses a temporal convolutional network to recognize basic dining actions. The rhythm analysis module calculates eating speed and rhythm parameters; The content inference module uses graph neural networks to match the physical properties of food.

[0015] On the other hand, the present invention also provides a user dining record device based on image recognition, including an electronic device and a computer-readable storage medium thereon storing a computer program, which implements the above-described method when executed by a processor. The electronic device includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the program to implement the method described above.

[0016] Compared with the prior art, this application has the following beneficial effects: This invention achieves privacy-preserving, interference-resistant, accurate, comprehensive, and seamless dining behavior analysis through core innovations such as physical field perception replacing vision, multimodal coupling to capture dynamics, and implicit integration to ensure the experience.

[0017] This invention abandons visual recognition and uses micro-force sensing, active sonar, and millimeter-wave radar to sense changes in the physical field, without collecting any visual data of the user's face or the appearance of food, thus eliminating the risk of privacy leakage from the underlying technology.

[0018] This invention emits high-frequency ultrasonic waves that are inaudible to the human ear and analyzes the echo phase / spectrum, focusing only on the dynamic changes in the oral cavity caused by chewing, and is completely immune to environmental noise.

[0019] Abnormal behaviors can be identified by analyzing indicators such as the variance of the feeding cycle and the swallowing interval. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a schematic diagram of the system structure of the present invention; Figure 3 This is a schematic diagram of the device structure of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0022] Example: On one hand, the present invention provides a user dining record method based on image recognition, comprising the following steps: Data on the microscopic deformation and vibration of the desktop caused by the user dining is collected using a micro-force sensor array; By emitting high-frequency ultrasound waves and receiving echoes through an active sonar system, a cavity resonance model is constructed by analyzing the phase and spectral changes of the echoes, and the characteristics of oral cavity resonance are analyzed. Data on the user's hand movement trajectory is collected using millimeter-wave radar; The above multimodal data are coupled and aligned on a timeline to identify basic dining action units; Based on the identified sequence of action units, the user's eating speed and rhythm are analyzed; By matching motion characteristics with a database of food physical properties, the content of the meal can be inferred. The user status feedback process outputs a list of food consumed, a real-time eating speed curve, an overall eating rhythm assessment, and alerts for abnormal behavior.

[0023] This example uses a 75-year-old elderly person living alone with mild hypertension as a scenario for monitoring their dinner intake and controlling their eating speed and high-salt food consumption. The dinner consists of pan-fried steak with black pepper sauce, stir-fried broccoli, and rice. A robot, shaped like a retro table lamp, is placed in the center of the table and integrates multiple sensors discreetly. The system must accurately identify the food's contents, eating speed, and rhythm without taking any images. It should trigger alerts when abnormalities such as excessive chewing or prolonged periods without eating are observed, while simultaneously protecting the user's privacy.

[0024] The core sensor is integrated into the base (15cm in diameter, 20cm in height), and its appearance is no different from that of an ordinary desk lamp. The specific configuration is as follows: The distributed micro-force sensor array has 300 capacitive force sensor units (1mm spacing, ±10μN range, 1kHz sampling rate) embedded in the circular area where the base contacts the table, sensing nano-Newton-meter level micro-vibrations and deformations of the table.

[0025] The active sonar module contains a miniature ultrasonic transmitter and receiver hidden inside the top lampshade. It operates at a frequency of 60-150kHz, with a chirp signal period of 0.5s and a transmission power of ≤1mW, and can detect changes in oral cavity resonance.

[0026] Millimeter-wave radar with a 24GHz FMCW radar integrated on the edge of the base, ranging from 0.5 to 3 meters, with an angular resolution of 5° and a sampling rate of 10Hz, reconstructs the three-dimensional motion trajectory of the hand.

[0027] The system synchronously collects raw data from the three modes and performs preprocessing: The raw signal of the micro-force array data is a record of the micro-force changes transmitted to the table surface by the vibration of the mandible when Grandpa Wang cuts the steak with a knife, eats with a fork, and chews. It is presented as a voltage signal that fluctuates with time, ranging from ±5mV. It is decomposed into a frequency band of 1-1000Hz by wavelet packet transform, and the high-frequency vibration features of the cutting action (100-300Hz) and the low-frequency features of chewing transmission (10-50Hz) are extracted.

[0028] The raw signal of the active sonar echo data is a 60-150kHz linear sweep frequency signal emitted by the transmitter, and the receiver collects the echo reflected through the oral cavity (amplitude ≤0.1mV), which contains information on acoustic impedance changes caused by oral cavity opening and closing and food breaking. The echo is coherently demodulated, the phase shift at different frequencies is calculated, and the oral cavity resonance spectrum (reflecting the dynamic morphology of the cavity during chewing) is constructed.

[0029] The raw signal from the millimeter-wave radar is point cloud data of the hand's movement (10-20 points per frame, including distance, angle, and radial velocity information). The hand trajectory is tracked using Kalman filtering to extract motion parameters, such as the reciprocating speed of the hand when cutting a steak (0.3 m / s) and the lifting height when "feeding" the food into the mouth (0.2 m).

[0030] A lightweight temporal convolutional network (TCN) is used to classify the fused data, identify different actions, perform multimodal coupling based on the trigger features of the system's built-in rule base, and record the time periods in which different actions occur.

[0031] Using "food intake → chewing → swallowing → next food intake" as a feeding cycle, the cyclical characteristics of a 30-minute meal were statistically analyzed. Total number of eating cycles: 22 (8 times for steak, 7 times for broccoli, and 7 times for rice).

[0032] Eating speed: 22 times / 30 minutes = 0.73 times / minute (normal range 0.5-1 times / minute, no risk of eating too fast).

[0033] Eating rhythm analysis: Cycle duration distribution: shortest 1.2 minutes (5th time, rice), longest 2.1 minutes (3rd time, steak, due to the need for multiple cuts).

[0034] Rhythm stability: The period duration variance is 0.12 (<0.2, judged as "stable rhythm"), with no obvious acceleration or deceleration trend.

[0035] Anomaly monitoring: If there is no "food delivery" action for 3 consecutive minutes between 18:20 and 18:23, the system will trigger a local alert (the desk lamp will flicker slightly), and then the meal will continue (excluding the risk of choking).

[0036] The system constructs a "food physical property knowledge graph" (nodes: steak, broccoli, rice; attributes: hardness, crispness, chewing cycle, etc.), maps action features to the attribute space, and the reasoning process is as follows: Steak identification: Associated action, "cutting" (high-frequency vibration 150Hz, lasting 15 seconds, force > 5μN) → corresponds to the attributes of "high hardness and high toughness". Chewing characteristics, sonar resonance period 1.2s (relatively long), single chewing lasts 9 seconds (requires multiple grinding) → matches the "steak" node (steak chewing period 1-1.5s in the knowledge graph). Confidence level: 92%.

[0037] Broccoli identification: Associated action, "scooping" (light force, no cutting action) → corresponds to the "medium hardness" attribute. Chewing characteristics, high-frequency peak (80Hz) in the sonar resonance spectrum → reflects "high brittleness" (acoustic characteristics of broccoli fiber breakage), single chewing lasts 4 seconds → matches the "broccoli" node. Confidence level, 88%.

[0038] Rice recognition: Associated action, small hand movement before "feeding" (radar trajectory arc length 0.2m) → corresponds to the "flowability / easy to scoop" attribute. Chewing characteristics, smooth sonar resonance spectrum (no high-frequency peaks), chewing cycle 0.8s (short, due to softness and stickiness) → match the "rice" node. Confidence level: 95%.

[0039] After the meal, the system generates an encrypted report and pushes it to the user's terminal. There are no images or sounds recorded, only structured data.

[0040] On the other hand, the present invention also provides an image recognition-based user dining record system for implementing the above method, comprising: A micro-force sensor array is integrated into the contact surface between the robot base and the dining table to sense microscopic deformation and vibration of the tabletop; An active sonar system, comprising an ultrasonic transmitter and receiver, is integrated inside the robot to transmit and receive high-frequency ultrasonic signals; Millimeter-wave radar module, used to collect the user's hand movement trajectory; The data processing unit is configured to perform multimodal data coupling analysis, action recognition, eating speed and rhythm analysis, and meal content inference. The communication module is used to transmit analysis results.

[0041] Furthermore, the data processing unit includes: The data synchronization module is used to align multimodal data timelines; The action recognition module uses a temporal convolutional network to recognize basic dining actions. The rhythm analysis module calculates eating speed and rhythm parameters; The content inference module uses graph neural networks to match the physical properties of food.

[0042] On the other hand, the present invention also provides a user dining record device based on image recognition, including an electronic device and a computer-readable storage medium thereon storing a computer program, which implements the above-described method when executed by a processor; The electronic device includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the program to implement the method described above.

[0043] This embodiment utilizes a robot's multimodal perception system to accurately decode dining behavior without relying on vision, through the coupled analysis of micro-force sensing, active sonar, and millimeter-wave radar. Its core innovation lies in transforming "visual recording" into "physical field perception," addressing the privacy concerns of traditional image recognition while enhancing analytical accuracy through multimodal temporal coupling. This provides a novel technological approach for scenarios such as elderly health monitoring and dietary management.

[0044] Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for recording user dining activities based on image recognition, characterized in that, Includes the following steps: Data on the microscopic deformation and vibration of the desktop caused by the user dining is collected using a micro-force sensor array; By emitting high-frequency ultrasound waves and receiving echoes through an active sonar system, a cavity resonance model is constructed by analyzing the phase and spectral changes of the echoes, and the characteristics of oral cavity resonance are analyzed. Data on the user's hand movement trajectory is collected using millimeter-wave radar; The above multimodal data are coupled and aligned on a timeline to identify basic dining action units; Based on the identified sequence of action units, the user's eating speed and rhythm are analyzed; By matching motion characteristics with a database of food physical properties, the content of the meal can be inferred. User status feedback outputs include a list of food consumed, a real-time eating speed curve, an overall eating rhythm assessment, and alerts for abnormal behavior.

2. The method according to claim 1, characterized in that: The micro-force sensing array consists of hundreds of miniature capacitive force sensing units, capable of sensing nano-Newton-meter level micro-deformation and vibration of the desktop.

3. The method according to claim 1, characterized in that: The basic dining action unit includes at least three of the following: picking up utensils, scooping, pinching, cutting, putting into the mouth, chewing, swallowing, and putting down the utensils.

4. The method according to claim 1, characterized in that: The eating speed analysis includes counting the frequency of food delivery actions per unit time and calculating the eating rhythm by combining the average time from food delivery to swallowing.

5. The method according to claim 1, characterized in that: The eating rhythm analysis includes calculating the distribution, variance, and autocorrelation of the eating cycle duration, reflecting changes in the user's eating habits and state.

6. The method according to claim 1, characterized in that: The inference of the meal content includes: Extract sonar texture features and micromechanical vibration features corresponding to chewing events; Map features to the physical attribute space; Graph neural networks are used to match the attribute sequences of multiple consecutive chewing events with a knowledge graph of food physical properties. The food physical property knowledge graph establishes a mapping relationship between food and physical properties, with food as nodes and physical properties such as crispness, hardness, viscosity, and density as edges.

7. The method according to claim 1, characterized in that: The multimodal data coupling alignment uses a temporal convolutional network to segment and classify the fused multimodal data stream and identify basic dining action units.

8. A user dining record system based on image recognition for implementing the method of any one of claims 1-7, characterized in that, include: A micro-force sensor array is integrated into the contact surface between the robot base and the dining table to sense microscopic deformation and vibration of the tabletop; An active sonar system, comprising an ultrasonic transmitter and receiver, is integrated inside the robot to transmit and receive high-frequency ultrasonic signals; Millimeter-wave radar module, used to collect the user's hand movement trajectory; The data processing unit is configured to perform multimodal data coupling analysis, action recognition, eating speed and rhythm analysis, and meal content inference. The communication module is used to transmit analysis results.

9. The system according to claim 8, characterized in that: The data processing unit includes: The data synchronization module is used to align multimodal data timelines; The action recognition module uses a temporal convolutional network to recognize basic dining actions. The rhythm analysis module calculates eating speed and rhythm parameters; The content inference module uses graph neural networks to match the physical properties of food.

10. A user dining record device based on image recognition, comprising an electronic device and a computer-readable storage medium storing a computer program thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-7; The electronic device includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the program to implement the method as described in any one of claims 1-7.