Visual monitoring method for dietary habits

By combining a mobile terminal camera and microphone, eating videos and audios are collected and analyzed in real time, solving the problems of inconvenient sensor devices and poor image clarity. This enables accurate monitoring of the number of chews, reducing costs and errors.

CN121330752APending Publication Date: 2026-01-13JIAXING QIXU MARKET RESEARCH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511307156.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-13
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

In existing technologies, using additional sensor devices to monitor the number of chews is inconvenient and results in poor image clarity, leading to inaccurate chewing count monitoring, especially with the difficulty in achieving accurate recording due to differences in the performance of cameras on different mobile terminals.

Method used

The system uses a mobile terminal camera to capture real-time video streams of eating, analyzes images of food and facial expressions, and combines this with microphones to capture chewing audio. It calculates the chewing time and number of chews for each bite, dynamically updates the chewing time for each chew, and uses visual and audio recognition technologies for precise monitoring.

Benefits of technology

It achieves relatively accurate chewing count monitoring without the need for additional sensors, reducing costs and inconvenience while improving monitoring accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121330752A_ABST
    Figure CN121330752A_ABST
Patent Text Reader

Abstract

The invention relates to the field of visual monitoring, in particular to a dietary habit visual monitoring method, which specifically comprises the following steps of: acquiring a feeding video stream in real time through a camera, and obtaining a feeding food image and a feeding face opening image; analyzing and processing the eating food image to obtain the standard chewing frequency of the eaten food; analyzing and processing the front and back two frames of eating face mouth opening images to obtain the chewing duration of each mouth of food; according to the chewing duration of each piece of food and the preset duration required by single chewing, the actual chewing frequency of each piece of food is obtained; according to the method, the actual chewing frequency of each food is compared with the standard chewing frequency, and whether the reminding mechanism is triggered is judged, so that when the camera performance of the user mobile terminal is poor, the chewing frequency of each food can be relatively accurately judged, an additional sensor device is not needed, and convenience is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of visual monitoring, and in particular to a method for visual monitoring of dietary habits. Background Technology

[0002] There are many ways to lose weight nowadays. The common methods are to increase exercise and reduce food intake, but few people pay attention to the weight loss effect of chewing slowly. The main reason is that most people do not know how many times to chew different foods, and they generally do not consciously count the number of chews when chewing food.

[0003] To this end, some systems or methods for monitoring the number of chewings have been developed and designed. For example, an intelligent eating reminder device with publication number CN113974611A provides multiple monitoring methods. For the visual monitoring part, it takes pictures of food to obtain the corresponding food type, and takes pictures of the face. It judges whether food is put into the mouth for chewing by the difference of image pixels when the mouth is open and closed.

[0004] The specific number of chews needs to be determined through electromyography, sound, or vibration, which requires the use of corresponding sensors. However, for the general population, wearing a specific sensor device before eating is very inconvenient. If real-time photography is taken directly through the camera built into a mobile device, the image clarity may be low due to differences in the camera performance of different mobile devices. Furthermore, the mouth is generally closed during chewing, making it difficult to accurately analyze low-resolution images of small facial changes, thus making it difficult to record the number of chews normally. Summary of the Invention

[0005] In order to achieve relatively accurate statistics on the number of chews without the need for additional sensor devices, this application provides a visual monitoring method for dietary habits.

[0006] The method for visual monitoring of dietary habits provided in this application adopts the following technical solution.

[0007] A visual monitoring method for dietary habits includes the following steps.

[0008] S1. Real-time video stream of eating is captured via camera to obtain images of the food being eaten and images of the face with open mouths while eating;

[0009] S2. Analyze and process images of the ingested food to obtain the standard number of chews for the ingested food.

[0010] S3. Analyze and process the two consecutive frames of the face with open mouths while eating to obtain the chewing time for each bite of food.

[0011] S4. Based on the chewing time of each bite of food and the preset time required for a single chew, obtain the actual number of chews for each bite of food;

[0012] S5. Compare the actual number of chews for each bite of food with the standard number of chews and determine whether to trigger the reminder mechanism.

[0013] By adopting the above technical solution, it is only necessary to identify images with large mouth openings and significant differences during eating, and calculate the actual number of chews based on the time difference between two consecutive images of the face openings during eating, as well as the duration of a single chew. Compared to the poor shooting performance of mobile terminal cameras, it can still monitor the number of chews relatively accurately, and there is no need to wear additional sensor devices during eating, reducing costs and inconvenience.

[0014] Optionally, the formula for calculating the chewing time of each bite of food in S3 is:

[0015] ;

[0016] in, The frame number is the later frame of the image showing the face with its mouth open while eating. The frame number is the earliest chronologically significant image of a person eating with their mouth open. The video capture frame rate of the mobile terminal camera. The duration of pausing chewing during the chewing of a single bite of food.

[0017] By employing the above technical solution, the time that may occur during the chewing process of a mouthful of food is eliminated, thereby ensuring the final... It is closer to the actual chewing time of a bite of food.

[0018] Optionally, the acquisition in S3 Specifically, the following steps are included:

[0019] S31. Obtain the audio of eating via microphone;

[0020] S32. Process all data points in each signal frame of the eating audio to obtain the RMS value.

[0021] S33. Perform bandpass filtering on all RMS values ​​within 2 seconds from 1 to 4 Hz and extract the envelope.

[0022] S34. Perform a 0.5s moving average on the envelope.

[0023] S35. Compare the envelope after the moving average processing with the baseline. If both are lower than the baseline within 1.5s, it is determined that no food was chewed within the corresponding 2s time period.

[0024] S36, will The sum of all 2-second intervals within a given time period that were judged as not being chewed food is obtained. .

[0025] By adopting the above technical solution, chewing sounds are collected using a mobile terminal microphone or a microphone connected to a mobile terminal for further processing. The acquisition of this information is intended to assist vision in achieving more accurate monitoring.

[0026] Optionally, after performing a 1-4Hz bandpass filter in S33, noise reduction processing is performed to suppress noise in the 0.3-3kHz range.

[0027] By adopting the above technical solutions, the surrounding voices and environmental noise can be suppressed to a certain extent, reducing misjudgments.

[0028] Optionally, in S34, based on the envelope... The and the first The time interval between peaks The preset time required for a single chew. Perform dynamic updates to obtain the first [number] on the envelope. Time required for a single chewing session at peak times .

[0029] By adopting the above technical solution, the duration of a single chewing session can be made more accurate, and the final calculated number of chewing sessions can also be more precise.

[0030] Optionally, the The calculation formula is:

[0031] ;

[0032] in, This is the weighting coefficient, with a value range of 0.15-0.3.

[0033] By adopting the above technical solution, the single chewing time can be dynamically updated, so that the single chewing time can be updated more accurately and synchronously with the chewing speed when eating in different states.

[0034] Optionally, in step S35, if 3-5 consecutive envelope lines are below the baseline, step S36 will... Set it to 0.

[0035] By adopting the above technical solution, when the microphone malfunctions and cannot collect chewing sounds, the number of chewing sounds will be calculated normally through visual detection.

[0036] Optionally, the baseline acquisition in step S35 specifically includes the following steps:

[0037] S351. Collect audio under natural closure within a predetermined time period in a quiet environment;

[0038] S352. Perform 1-4Hz bandpass filtering on the audio signal and extract the envelope;

[0039] S353. Sample the envelope at intervals to obtain several amplitude values;

[0040] S354. Use the 90th percentile of all amplitude values ​​as the baseline.

[0041] By adopting the above technical solution, the baseline setting can be better suited to the actual situation of different microphones, and the 90th percentile setting can better filter out occasional coughs or external noise during the process, ensuring that the baseline setting is more accurate.

[0042] Optionally, the specific steps in S1 for acquiring the image of a face with its mouth open while eating include:

[0043] S11. Obtain the pixel coordinates of four fixed points in the face image: the midpoint of the upper lip, the midpoint of the lower lip, the left corner of the mouth, and the right corner of the mouth.

[0044] S12. Obtain the mouth height based on the pixel distance between the midpoint of the upper lip and the midpoint of the lower lip, and obtain the mouth width based on the pixel distance between the left corner of the mouth and the right corner of the mouth;

[0045] S13. Based on the mouth height and mouth width, obtain the vertical proportion and horizontal stretch of the open mouth and compare it with the standard value to determine whether the corresponding face image is an open mouth image of a face eating.

[0046] By adopting the above technical solution, the vertical ratio and horizontal stretching values ​​are used to make a more accurate judgment on whether the mouth opening is for eating or for general speaking, so as to more accurately calculate the actual total chewing time of a mouthful of food.

[0047] Optionally, the formula for calculating the transverse tension is:

[0048] ;

[0049] in, This is the width of the mouth when it is naturally closed. This corresponds to the mouth width value in the image.

[0050] By adopting the above technical solution, the lateral width ratio of the mouth when it is open can be calculated more accurately.

[0051] In summary, this application includes at least the following beneficial effects.

[0052] 1. It only needs to identify images with large mouth openings and significant differences during eating, and calculate the actual number of chews based on the time difference between two consecutive images of the face with open mouths during eating, as well as the duration of a single chew. Compared to the poor shooting performance of mobile terminal cameras, it can still monitor the number of chews relatively accurately, and there is no need to wear additional sensor devices during eating, reducing costs and inconvenience.

[0053] 2. Use the microphone of the mobile terminal or a microphone connected to the mobile terminal to collect chewing sounds for analysis. The acquisition of this information is intended to assist vision in achieving more accurate monitoring. Attached Figure Description

[0054] Figure 1 This is a flowchart of the main steps of this application; Detailed Implementation

[0055] The present application will be further described in detail below with reference to the accompanying drawings.

[0056] This application discloses a method for visual monitoring of dietary habits, referring to... Figure 1 Specifically, it includes the following steps.

[0057] S1. Real-time video stream of eating is captured via camera to obtain images of the food being eaten and images of the face with its mouth open while eating.

[0058] The camera can be a front-facing or rear-facing camera built into the mobile device, such as a smartphone, tablet, or smartwatch. The captured video stream resolution can be 640×480, and the frame rate can be 30fps, ensuring that most mobile device cameras on the market today meet the requirements.

[0059] The specific steps for obtaining the image of a face with its mouth open while eating in S1 include the following.

[0060] S11. Obtain the pixel coordinates of four fixed points in the face image: the midpoint of the upper lip, the midpoint of the lower lip, the left corner of the mouth, and the right corner of the mouth. Specifically, a face key point model, such as a 468-point FaceMesh, can be used to obtain the pixel coordinates.

[0061] S12. Obtain the mouth height based on the pixel distance between the midpoint of the upper lip and the midpoint of the lower lip, and obtain the mouth width based on the pixel distance between the left corner of the mouth and the right corner of the mouth.

[0062] S13. Based on the mouth height and mouth width, obtain the vertical proportion and horizontal stretch of the open mouth and compare it with the standard value to determine whether the corresponding face image is an open mouth image of a face eating.

[0063] The vertical proportion is calculated by dividing the mouth height by the mouth width, while the formula for calculating the horizontal stretch is as follows.

[0064] ;

[0065] in, This is the mouth width value when the mouth is naturally closed, which can be the mouth width corresponding to the minimum mouth height in all images. This corresponds to the mouth width value in the image.

[0066] During the comparison process, the vertical aspect ratio must be greater than 0.4 and the horizontal stretching must be less than 0.15. Only when both conditions are met can the corresponding image be identified as an image of a face with its mouth open while eating.

[0067] S2. Analyze and process the images of the food being eaten to obtain the standard number of chews for the food being eaten.

[0068] Specifically, the process can begin by identifying the hands in an image of a face with its mouth open while eating, using a MediaPipe Hand algorithm. The area between the hands and mouth is then cropped to obtain a cropped image. This cropped image is then set to a standard size with a width and height of 224 pixels to obtain an adjusted image. The adjusted image is then subjected to background removal using algorithms such as GrabCut or a lightweight U-Net model to obtain a denoised image. This denoised image is then input into a food recognition and classification model to output the specific food type. The food recognition and classification model can be configured by using MobileNetV3-Small as the backbone network, combined with a Squeeze-and-Excitation attention module to enhance feature representation capabilities, and finally outputting fine-grained food category predictions through a custom classification head.

[0069] After obtaining the specific types of food to be eaten, the standard number of chews for the corresponding food is extracted according to the preset food category-chew count mapping table.

[0070] S3. Analyze and process the two consecutive frames of the face with open mouths while eating to obtain the chewing time for each bite of food.

[0071] The formula for calculating the chewing time for each bite of food is as follows.

[0072] ;

[0073] in, The frame number is the later frame of the image showing the face with its mouth open while eating. The frame number is the earliest chronologically significant image of a person eating with their mouth open. The video capture frame rate of the mobile terminal camera. The duration of pausing chewing during the chewing of a single bite of food.

[0074] Get Specifically, it includes the following steps.

[0075] S31. Obtain eating audio via microphone.

[0076] The microphone can be a built-in microphone of the mobile terminal, or a microphone from a wired or wireless headset connected to the mobile terminal, and the audio format is a 16kHz sampling rate, mono recording, and 16bit PCM storage format.

[0077] S32. Process all data points in each signal frame of the eating audio to obtain the RMS value.

[0078] The specific processing procedure can be as follows: take 400 sampling points every 25ms in the eating audio as a signal frame, normalize the amplitude of each sampling point to the range of [-1,1], and then calculate the RMS amplitude of each signal frame according to the root mean square formula.

[0079] S33. Perform bandpass filtering on the latest 80 RMS values ​​within 2 seconds from 1 to 4 Hz and extract the envelope.

[0080] The envelope of the filtered signal can be extracted using methods such as Hilbert transform, peak detection, or moving average. Furthermore, a 2-second window provides a good balance between signal timeliness and smoothness.

[0081] In addition, after bandpass filtering at 1-4 Hz, the energy of speech and ambient noise in the range of 300 Hz–3 kHz is attenuated by spectral subtraction or the RNNoise lightweight model, while retaining the chewing envelope at 1–4 Hz to achieve noise reduction.

[0082] S34. Perform a 0.5s moving average on the envelope.

[0083] Define a time window of 0.25 seconds before and after any data point on the envelope as the center point. Move the time window sequentially along the time axis on the envelope and calculate the arithmetic mean of the amplitudes of all data points within the time window as the new amplitude at the center point.

[0084] S35. Compare the envelope after the moving average processing with the baseline. If both are lower than the baseline within 1.5s, it is determined that no food was chewed within the corresponding 2s time period.

[0085] The baseline acquisition process includes the following steps.

[0086] S351. Collect audio in a quiet environment within a predetermined time period, such as 5 seconds, under natural closed-mouth conditions.

[0087] S352. Perform 1-4Hz bandpass filtering on the audio signal and extract the envelope.

[0088] S353. Sample the envelope at intervals to obtain several amplitude values. The sampling interval can be 25ms.

[0089] S354. Use the 90th percentile of all amplitude values ​​as the baseline.

[0090] The 90th percentile of all amplitude values ​​means that 90% of the amplitude values ​​in all data are lower than or equal to this value, in order to filter out extremely high amplitude values ​​such as coughing or instantaneous environmental noise.

[0091] S36, will The sum of all 2-second intervals within a given time period that were judged as not being chewed food is obtained. .

[0092] Additionally, if 3-5 consecutive envelope lines in S35 are below the baseline, it indicates a microphone problem or that chewing has not been performed for an extended period. Therefore, in S36, the chewing process of the corresponding bite of food will be... Set to 0, and the mobile terminal will issue a reminder that the microphone is damaged or that the next bite of food needs to be chewed properly, until the next frame of the eating face with open mouth image is detected. Resume testing.

[0093] S4. Based on the chewing time of each bite of food and the preset time required for a single chew, obtain the actual number of chews for each bite of food.

[0094] The preset chewing time can be adjusted on the mobile device, with an adjustment range of 0.3-1.0s and a default value of 0.85s.

[0095] And it can be based on the envelope in S34. The and the first The time interval between peaks The preset time required for a single chew. Perform dynamic updates to obtain the first [number] on the envelope. Time required for a single chewing session at peak times . The calculation formula is as follows.

[0096] ;

[0097] in, This is the weighting coefficient; the larger the value, the higher the weighting coefficient. The faster the update, the value range is 0.15-0.3.

[0098] And only when Values ​​within 0.4-1.2 seconds are considered valid; values ​​outside this range are discarded to reduce the impact of abnormal chewing on the updated measurement of the actual chewing time.

[0099] S5. Compare the actual number of chews for each bite of food with the standard number of chews and determine whether to trigger the reminder mechanism.

[0100] When the actual number of chews per bite is greater than or equal to the standard number of chews, the mobile terminal does not trigger the reminder mechanism; otherwise, the reminder mechanism is triggered, and the mobile terminal reminds the user that the previous bite was not chewed enough and the next bite needs to be chewed more thoroughly through one or more of the following methods: sound, image pop-up, and vibration.

[0101] Furthermore, the mobile device can also refer to the standard number of chews for eating food and the real-time update of the time required for a single chew. The system calculates the required chewing time for each bite of food, and when the target chewing time is reached, it prompts the user to take the next bite.

[0102] When the last bite of food is chewed, the monitoring stops on the mobile device to obtain the chewing status of all types of food consumed in this meal.

[0103] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A method for visually monitoring dietary habits, characterized in that: Specifically, the following steps are included: S1. Real-time video stream of eating is captured via camera to obtain images of the food being eaten and images of the face with open mouths while eating; S2. Analyze and process images of the ingested food to obtain the standard number of chews for the ingested food. S3. Analyze and process the two consecutive frames of the face with open mouths while eating to obtain the chewing time for each bite of food. S4. Based on the chewing time of each bite of food and the preset time required for a single chew, obtain the actual number of chews for each bite of food; S5. Compare the actual number of chews for each bite of food with the standard number of chews and determine whether to trigger the reminder mechanism.

2. The method for visual monitoring of dietary habits according to claim 1, characterized in that: The formula for calculating the chewing time for each bite of food in S3 is as follows: ; in, The frame number is the later frame of the image showing the face with its mouth open while eating. The frame number is the earliest chronologically significant image of a person eating with their mouth open. The video capture frame rate of the mobile terminal camera. The duration of pausing chewing during the chewing of a single bite of food.

3. The method for visual monitoring of dietary habits according to claim 2, characterized in that: The acquisition in S3 Specifically, the following steps are included: S31. Obtain the audio of eating via microphone; S32. Process all data points in each signal frame of the eating audio to obtain the RMS value. S33. Perform bandpass filtering on all RMS values ​​within 2 seconds from 1 to 4 Hz and extract the envelope. S34. Perform a 0.5s moving average on the envelope. S35. Compare the envelope after the moving average processing with the baseline. If both are lower than the baseline within 1.5s, it is determined that no food was chewed within the corresponding 2s time period. S36, will The sum of all 2-second intervals within a given time period that were judged as not being chewed food is obtained. .

4. The method for visual monitoring of dietary habits according to claim 3, characterized in that: After performing a 1-4Hz bandpass filter, the S33 performs noise reduction processing to suppress noise in the 0.3-3kHz range.

5. The method for visual monitoring of dietary habits according to claim 3, characterized in that: In S34, the first step is based on the envelope. The and the first The time interval between peaks The preset time required for a single chew. Perform dynamic updates to obtain the first [number] on the envelope. Time required for a single chewing session at peak times .

6. The method for visual monitoring of dietary habits according to claim 5, characterized in that: The The calculation formula is: ; in, This is the weighting coefficient, with a value range of 0.15-0.

3.

7. The method for visual monitoring of dietary habits according to claim 3, characterized in that: When 3-5 consecutive envelope lines in S35 are below the baseline, S36 will... Set it to 0.

8. The method for visual monitoring of dietary habits according to claim 3, characterized in that: The baseline acquisition in S35 specifically includes the following steps: S351. Collect audio under natural closure within a predetermined time period in a quiet environment; S352. Perform 1-4Hz bandpass filtering on the audio signal and extract the envelope; S353. Sample the envelope at intervals to obtain several amplitude values; S354. Use the 90th percentile of all amplitude values ​​as the baseline.

9. The method for visual monitoring of dietary habits according to claim 1, characterized in that: The specific steps in S1 for obtaining the image of a face with its mouth open while eating include: S11. Obtain the pixel coordinates of four fixed points in the face image: the midpoint of the upper lip, the midpoint of the lower lip, the left corner of the mouth, and the right corner of the mouth. S12. Obtain the mouth height based on the pixel distance between the midpoint of the upper lip and the midpoint of the lower lip, and obtain the mouth width based on the pixel distance between the left corner of the mouth and the right corner of the mouth; S13. Based on the mouth height and mouth width, obtain the vertical proportion and horizontal stretch of the open mouth and compare it with the standard value to determine whether the corresponding face image is an open mouth image of a face eating.

10. The method for visual monitoring of dietary habits according to claim 9, characterized in that: The formula for calculating transverse tension is: ; in, This is the width of the mouth when it is naturally closed. This corresponds to the mouth width value in the image.

Citation Information

Patent Citations

  • Intelligent feeding reminding device

    CN113974611A

  • Mastication frequency calculation device

    JP2014083279A

  • Eating monitoring device, method, and program

    JP2024142210A

  • Meal supporting system

    JP2025087332A