Artificial intelligence-based user nutrition detection and intervention system and method

By integrating multimodal data acquisition hardware and a multimodal fusion neural network model, the invasiveness and singularity of existing nutritional testing and intervention methods have been solved, enabling rapid and accurate nutritional risk assessment and personalized intervention recommendations, thereby improving user compliance and testing accuracy.

CN122290995APending Publication Date: 2026-06-26HUNAN LONGYI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUNAN LONGYI TECH CO LTD
Filing Date
2026-03-16
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing nutritional testing and intervention methods suffer from problems such as cumbersome testing procedures, high intrusiveness, limited data, poor model generalization ability, limited interaction methods, and lack of personalized intervention recommendations, resulting in low user compliance, inaccurate test results, and difficulty in implementing recommendations.

Method used

Employing a hardware platform integrating a foot electrode array, millimeter-wave radar module, 3D optical scanning module, and voice interaction module, combined with a multimodal fusion neural network model, it achieves non-invasive multimodal data acquisition and personalized nutritional intervention recommendations, supports multiple interaction methods, and generates personalized nutritional intervention reports.

Benefits of technology

It enables non-invasive, rapid multimodal detection, improves user compliance, increases the accuracy of nutritional risk assessment and the implementation rate of intervention recommendations, meets clinical-grade testing accuracy, and ensures data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290995A_ABST
    Figure CN122290995A_ABST
Patent Text Reader

Abstract

This invention provides an AI-based user nutrition detection and intervention system and method. The system enables non-invasive multimodal detection without the need for blood sample collection, addressing the invasiveness and inefficiency issues of traditional clinical testing. Through feature-level fusion and deep neural network processing, it achieves collaborative analysis of BIA body composition, respiratory features, body shape data, and voice interaction information, improving the accuracy of risk level classification compared to single-modal models. The voice interaction process is optimized for special groups such as the elderly and children, employing beamforming noise reduction technology and the ERNIE-Speech voice recognition model to overcome the operational barriers of traditional devices. Intervention plans are generated based on user dietary preferences and physiological characteristics. AES encryption and blockchain notarization technologies ensure the immutability of user privacy data, complying with the Personal Information Protection Law and addressing data security vulnerabilities of existing devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent medical device technology, and in particular to an artificial intelligence-based user nutrition detection and intervention system and method, electronic device and computer-readable storage medium. Background Technology

[0002] Currently, nutritional health issues among residents are becoming increasingly prominent. Unhealthy dietary structures and lack of exercise are leading to a year-on-year increase in the incidence of nutrition-related diseases such as obesity, protein deficiency, and vitamin deficiency. However, existing nutritional testing and intervention methods have many limitations: 1. Traditional clinical testing methods rely on biochemical analysis, which requires the collection of blood, urine and other samples. The testing process is cumbersome and time-consuming (it usually takes 1-3 days to get results), and it involves invasive procedures, resulting in low user compliance and making it difficult to achieve routine nutritional monitoring.

[0003] 2. Most existing home nutrition testing devices (such as body fat scales and smart bracelets) can only collect data from a single modality. For example, body fat scales only obtain body composition data through bioelectrical impedance analysis (BIA), and cannot integrate multi-dimensional information such as the user's respiratory characteristics, body shape distribution, diet and exercise habits. The accuracy and comprehensiveness of the test results are insufficient, making it difficult to accurately assess nutritional risks.

[0004] 3. Existing AI nutrition assessment models are mostly trained on single-modal data and lack an effective mechanism for fusing multimodal data. This results in poor generalization ability of the models, low accuracy in identifying nutritional risks, and an inability to provide users with accurate and personalized intervention suggestions.

[0005] 4. The interaction methods are limited. Existing devices mostly rely on touch screens or mobile apps for operation, which poses a barrier to entry for people with weaker operational skills, such as the elderly and children, and cannot achieve convenient real-time interaction and data collection.

[0006] 5. Nutritional intervention recommendations lack personalization and are mostly based on general nutrition guidelines without taking into account the user's specific nutritional risks, physical characteristics, and lifestyle habits. This results in insufficient feasibility and effectiveness of the recommendations, making it difficult for users to adhere to them. Summary of the Invention

[0007] To address the technical problems existing in the prior art, the present invention provides the following technical solution: On the one hand, an artificial intelligence-based user nutrition detection and intervention system is provided, the system comprising: a detection platform, a backend server, and a user terminal, wherein: The detection platform is a hardware entity that integrates a foot electrode array, a millimeter-wave radar module, a three-dimensional optical scanning module, and a voice interaction module. It is used to simultaneously collect bioelectrical impedance data, respiratory characteristic data, three-dimensional volumetric mesh data, and interactive voice data of users while they are standing barefoot. The backend server is connected to the detection platform via a 5G or WiFi network and is equipped with a data storage unit, a multimodal fusion model inference unit, and an intervention suggestion generation unit. It is used to perform fusion inference on the received encrypted data, generate a risk assessment result containing nutritional risk level and specific risk points, and generate a personalized nutritional intervention suggestion report accordingly. The user terminal connects to the backend server via the HTTPS protocol for user authentication, receiving and visualizing the intervention suggestion report, and interacting with the detection platform.

[0008] Preferably, the foot electrode array of the detection platform specifically includes: The capacitive electrode array has eight electrodes evenly distributed in the left and right foot placement areas, with an electrode diameter of 10 mm and a spacing of 20 mm. The electrode array is used to automatically identify the user's foot contact. When the impedance value of the detected human tissue is lower than 1000Ω, it triggers a bioelectrical impedance analysis scan. It is also used to apply sinusoidal alternating currents at three frequencies of 5 kHz, 50 kHz, and 100 kHz to the human body and measure its impedance to extract body composition characteristic data, including body fat percentage, muscle mass, water content, and basal metabolic rate.

[0009] Preferably, the millimeter-wave radar module of the detection platform is a 77GHz frequency-modulated continuous wave radar with a sampling rate of 200Hz and a detection range of 0.5-2m; This module is used to capture chest wall fluctuation data for at least 10 complete respiratory cycles within 60 seconds. It calculates chest wall displacement by analyzing the frequency changes of reflected waves, and after removing motion noise using a Kalman filter algorithm, it extracts the user's respiratory rate, respiratory rhythm variation coefficient, and chest wall fluctuation amplitude data to assess the user's metabolic status.

[0010] Preferably, the three-dimensional optical scanning module of the detection platform includes an infrared structured light projector and two high-definition cameras, wherein the cameras have a resolution of 2MP and a frame rate of 30fps; The module is used to project coded infrared patterns onto the human body and calculate the three-dimensional coordinates of the point cloud using the triangulation principle to construct a three-dimensional volumetric mesh model for the user. The point cloud density of the model reaches 1000 points / cm². After being generated by the Poisson reconstruction algorithm and statistically filtered to remove noise, it can calculate body shape characteristic indicators including waist circumference, hip circumference, waist-to-hip ratio, and body fat distribution uniformity.

[0011] Preferably, the detection platform further includes a local data processing unit and a voice interaction module, wherein: The local data processing unit, built on an ARM processor, is used to perform time alignment, noise reduction, and encryption processing on the collected bioelectrical impedance data, millimeter-wave radar data, and three-dimensional optical data, and uploads the encrypted data via the network after packaging it using the AES-256 encryption algorithm. The voice interaction module integrates a microphone array, an automatic speech recognition module, and a text-to-speech module. It is used to ask users questions related to diet, exercise, etc. in voice form during the later stages of data collection, and convert the user's voice response into text data through a speech recognition model. At the same time, it broadcasts the reports and suggestions generated by the backend server to the user.

[0012] Preferably, the multimodal fusion model inference unit of the backend server is, at its core, a multimodal fusion neural network model deployed on a GPU cluster. This model outputs a nutritional risk assessment through the following processing flow: (1) Modal feature encoding: The input BIA body component feature vector X1, breathing feature vector X2, three-dimensional optical body shape feature vector X3, and voice interaction feature vector X4 are multiplied by the corresponding weight matrices W1~W4 and bias terms b1~b4 are added. The higher-order feature vectors are obtained by mapping through the Sigmoid activation function σ. (2) Multimodal fusion: The encoded high-order features are subjected to Hadamard product (element-level multiplication) operation and then concatenated with some features to form an aggregated feature vector. The final fusion feature vector Z is generated again through linear transformation and nonlinear activation. (3) Classification output: The fused feature vector Z is calculated by the classification layer, and the nutritional risk level vector Y_risk is output by the Softmax function, and the probability vector Y_point of the specific nutritional risk point is output by the Sigmoid function, thereby realizing multi-dimensional nutritional risk assessment at the same time.

[0013] Preferably, the intervention suggestion generation unit is coupled with the multimodal fusion model inference unit and the data storage unit, and is used to generate personalized comprehensive intervention suggestions based on the risk points inferred from the model, combined with a pre-set nutritional knowledge graph and the "Dietary Guidelines for Chinese Residents" standard; The intervention recommendations include at least dietary advice, exercise programs, and lifestyle adjustments, and can automatically match and output dedicated report templates based on the user's age and health status.

[0014] Preferably, the data communication between the various parts of the system specifically includes: The local data processing unit of the detection platform encrypts the collected data using AES-256 and uploads it to the backend server via 5G NR or WiFi 6 network. After the backend server decrypts and processes the data, it securely pushes the report and suggestions to the user terminal via HTTPS protocol. The detection platform and the user terminal establish a direct connection via Bluetooth 5.0 protocol for issuing control commands and displaying real-time data.

[0015] Preferably, the data storage unit of the backend server is a distributed storage architecture, including an HDFS file system for storing massive amounts of time-series data, and a MongoDB database for managing user reports and model metadata, while also performing anonymization processing on user data. On the other hand, an artificial intelligence-based user nutrition detection and intervention method for the aforementioned system is provided, the method comprising the following steps: S1: Synchronous acquisition of multimodal data When a user stands barefoot in a designated area of ​​the detection platform, multimodal data acquisition is automatically triggered. S2: Local data preprocessing and encrypted transmission The local data processing unit of the detection platform receives the raw data collected in step S1, performs noise reduction, filtering and time alignment operations in sequence, and then uses the AES-256 algorithm to encrypt the integrated multimodal data packets and transmits the encrypted data packets to the backend server through the 5G / WiFi network. S3: Risk Assessment Based on Multimodal Fusion Model The backend server receives and decrypts the data packet, and inputs the cleaned multimodal feature vector into a pre-trained multimodal fusion neural network for inference: The model calculates layer by layer through modal feature encoding layer, multimodal fusion layer and classification output layer to obtain the probability distribution of nutritional risk level (low, medium and high) and the binary judgment probability of specific risk points. Based on the risk point assessment results, combined with the built-in nutritional standards and knowledge graph, the system initially identifies the main types of nutritional risks. S4: Personalized Intervention Report Generation and Feedback The report generation unit on the backend server, based on the risk assessment results of step S3, retrieves a report template that matches the user's age and risk characteristics, and generates a structured nutrition risk report with illustrations and text; the intervention suggestion generation unit, based on specific risk points, integrates dietary preferences and health goals to generate actionable dietary, exercise, and lifestyle suggestions, and packages them together. Finally, the server will push the generated full report and recommendations to the user's mobile app or mini-program via an HTTPS secure channel. At the same time, it can also broadcast the report to the user via the text-to-speech module of the detection platform, completing the complete closed-loop process from detection to intervention.

[0016] On the other hand, an electronic device is provided, comprising: a processor; and a memory storing computer-readable instructions, which, when executed by the processor, implement the method described above.

[0017] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored therein, the at least one instruction being loaded and executed by a processor to implement the method described above.

[0018] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: 1. Non-invasive multimodal detection: Utilizing non-contact technologies such as BIA, millimeter-wave radar, and 3D optical scanning, no blood sample collection is required. The detection process takes only 60 seconds, improving user compliance by more than 80% and solving the problems of invasiveness and inefficiency in traditional clinical testing.

[0019] 2. Multimodal data fusion mechanism: The MFN model achieves collaborative analysis of BIA body composition, respiratory features, body shape data, and voice interaction information through feature-level fusion (element-wise multiplication + concatenation) and deep neural network processing, with a risk level classification accuracy of 94%, which is 23% higher than that of the single-modal model.

[0020] 3. Adaptive Interaction for All Ages: The voice interaction process is optimized for special groups such as the elderly and children. Beamforming noise reduction technology and ERNIE-Speech voice recognition model are used, with a recognition accuracy of 98%, solving the problem of the operation threshold of traditional devices.

[0021] 4. Personalized Intervention Engine: Based on the rule engine (risk point-intervention measure matching) and knowledge graph (food-nutrient association), intervention plans are generated by combining users' dietary preferences and physiological characteristics, such as the high-protein diet for the elderly in Example 1 and the vitamin D supplementation plan for children in Example 2, which is expected to improve the execution rate by 65%.

[0022] 5. Secure and reliable data management: Employing 128-bit AES encryption and blockchain notarization technology, it ensures that user privacy data is tamper-proof, complies with the requirements of the Personal Information Protection Law, and resolves existing data security risks.

[0023] 6. Clinical-grade testing accuracy: Through multi-frequency BIA (5kHz / 50kHz / 100kHz) and 3D structured light scanning (error <5mm), the body fat percentage measurement error is ≤1.2%, and the muscle mass measurement error is ≤0.5kg, meeting clinical-grade testing standards and can replace some biochemical test indicators. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a schematic diagram of the system composition provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the interactive control logic provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the method flow provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0026] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0027] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0028] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0029] In this embodiment of the invention, sometimes a subscript such as W1 may be mistakenly written as a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0030] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0031] I. System Overall Architecture like Figure 1 As shown, this system adopts a three-tier architecture of "detection platform - backend server - user terminal" to realize the acquisition, transmission, analysis and feedback of multimodal data. The specific module composition is as follows: (a) Testing Platform The detection platform is the core of the system's front-end data acquisition, integrating a foot electrode array, millimeter-wave radar module, 3D optical scanning module, voice interaction module, and local data processing unit (the hardware models and deployment methods of each electronic device can be selected by the user from the market). Specific parameters and functions are as follows: 1. Foot Electrode Array: Utilizing 16 capacitive electrodes, evenly distributed across the foot placement area of ​​the detection platform (8 for each foot), with an electrode diameter of 10mm and a spacing of 20mm. The system automatically identifies the user's foot contact status. When the user stands barefoot, the electrodes detect impedance changes in human tissue (impedance value below 1000Ω), triggering the BIA scan process. Simultaneously, alternating currents of different frequencies (5kHz, 50kHz, 100kHz) are applied to measure the impedance of different human tissues, extracting body composition characteristics (body fat percentage, muscle mass, water content, basal metabolic rate, etc.). The technology is based on the differences in electrical conductivity among human tissues. Adipose tissue has a high resistivity (approximately 1000 Ω·cm), while muscle tissue and water have lower resistivity (approximately 100 Ω·cm). By measuring the impedance values ​​at different frequencies, body composition indicators are calculated using formulas, such as the body fat percentage formula: BF% = a*R + b*W + c*H + d*A + e, where R is the impedance value, W is weight, H is height, A is age, and a, b, c, d, and e are calibration coefficients that are adjusted according to gender and racial differences.

[0032] 2. Millimeter-wave radar module: Employs a 77GHz frequency-modulated continuous wave (FMCW) millimeter-wave radar with a sampling rate of 200Hz, a detection range of 0.5-2m, and an angular resolution of 1°. It captures the user's respiratory characteristics, completing data acquisition for 10 full respiratory cycles within 60 seconds, extracting features such as respiratory rate, chest wall amplitude, and respiratory rhythm variability coefficient to assess the user's metabolic state. The millimeter-wave radar emits high-frequency electromagnetic waves, which are received after reflection from the surface of the human chest. By analyzing the frequency changes of the reflected waves (Doppler effect), the displacement changes of the chest wall are calculated, thereby extracting respiratory characteristics. A Kalman filter algorithm is used to remove environmental noise and interference from slight user movements, ensuring an accuracy rate of over 99% in respiratory cycle detection.

[0033] 3. 3D Optical Scanning Module: Employing structured light 3D scanning technology, this module integrates one infrared structured light projector and two high-definition cameras (2MP resolution, 30fps). It generates a 3D volumetric mesh model of the user, extracting body features (waist circumference, hip circumference, upper arm circumference, body fat distribution areas, etc.) for assessing body fat distribution. The structured light projector projects an coded infrared pattern onto the human body, and the cameras capture the pattern's deformation. Triangulation is used to calculate the 3D coordinates of each point, generating point cloud data. Statistical filtering removes outliers, and a Poisson reconstruction algorithm generates a refined volumetric mesh model with a point cloud density of 1000 points / cm² and a body measurement error not exceeding 5mm.

[0034] 4. Voice Interaction Module: Integrates a microphone array (4 microphones), speaker, ASR (Automatic Speech Recognition) module, and TTS (Text-to-Speech) module, supporting real-time interaction in Mandarin Chinese. It enables voice dialogue with users, allowing them to ask nutrition-related questions and converting their voice responses into text information to supplement nutrition screening data. Simultaneously, it converts the system-generated nutrition risk assessment results and intervention suggestions into voice feedback to the user. The microphone array uses beamforming technology to focus the user's voice and remove environmental noise; the ASR module uses a pre-trained model based on the Transformer architecture (such as ERNIE-Speech), achieving a speech recognition accuracy of over 98%; the TTS module uses end-to-end speech synthesis technology to generate natural and fluent speech with a near-human tone.

[0035] 5. Local Data Processing Unit: Employs an ARM Cortex-A76 processor, equipped with 8GB of RAM and 64GB of storage. It preprocesses the acquired multimodal data (denoising, filtering, time alignment) to achieve initial data processing and caching; simultaneously, it is responsible for communication between the detection platform and the backend server, enabling encrypted data transmission.

[0036] (ii) Backend server The backend server is the core processing unit of the system, deployed in the cloud, and integrates data storage, MFN model inference, report generation, user management, and intervention suggestion generation units. Its specific functions are as follows: 1. Data Storage Unit: Employs a distributed storage system (HDFS + MongoDB) to support the storage and retrieval of massive amounts of multimodal data. It stores basic user information, detection data, model inference results, nutritional risk reports, and intervention recommendations. Simultaneously, the data undergoes anonymization processing to remove user privacy information, retaining only anonymized user IDs and detection data.

[0037] 2. MFN Model Inference Unit: The MFN multimodal fusion model is deployed using a GPU cluster (NVIDIA A100), supporting real-time inference (single-sample inference time not exceeding 100ms). It loads preprocessed multimodal data, performs nutritional risk assessment, and outputs nutritional risk levels (low, medium, high) and specific risk points.

[0038] 3. Report Generation Unit: Based on the inference results of the MFN model and combined with the user's basic information, a structured nutrition risk report is generated, including body composition analysis, respiratory characteristic analysis, body shape characteristic analysis, nutrition risk assessment, and personalized intervention recommendations. The report supports both PDF and HTML formats, and users can view or download it through the terminal.

[0039] 4. User Management Unit: Responsible for user authentication, permission management, and data permission allocation. Users can log in by scanning a QR code or entering their identity information via voice on the terminal. Each user has an independent personal data space, and only the user can view and manage their own test data and reports.

[0040] 5. Intervention Recommendation Generation Unit: Combining nutrition risk reports, nutrition screening results, and users' health goals, this unit generates personalized intervention recommendations, including dietary recommendations (daily calorie intake, nutrient ratios, food recommendations), exercise recommendations (type of exercise, duration, frequency), and lifestyle recommendations (sleep duration, water intake recommendations). The recommendations are generated based on the "Chinese Dietary Guidelines (2022)" and clinical nutrition intervention standards to ensure their scientific validity and feasibility.

[0041] (iii) User terminal User terminals include a mobile app and a WeChat mini-program, enabling user interaction with the system. Specific functions include: 1. Identity Verification: Users can bind their identity via QR code scanning, mobile phone number login, or voice verification. 2. Report Viewing: Users can view nutrition risk reports, nutrition screening results, and personalized intervention recommendations, and share and download reports. 3. Voice Interaction: Users can engage in real-time voice dialogue with the system, answer nutrition-related questions, and inquire about specific details of intervention recommendations. 4. Data Tracking: The system records users' testing history, the implementation status of intervention recommendations, and changes in health indicators, generating health trend charts.

[0042] II. System Data Communication Methods The system adopts a secure communication architecture of "local encryption - 5G / WiFi transmission - cloud decryption" to ensure data transmission security and real-time performance: 1. Detection platform and backend server: Data transmission is conducted via 5G NR network or WiFi 6 network. The local data processing unit encrypts the pre-processed multimodal data using the AES-256 symmetric encryption algorithm. The encrypted data is uploaded to the backend server in the form of data packets, with a packet size of 10MB and an upload time of no more than 2 seconds. After receiving the data packet, the server decrypts it using the corresponding key and stores it in the data storage unit. 2. Backend server and user terminal: Data transmission is conducted using the HTTPS protocol. All data is encrypted using SSL to ensure that the data is not stolen or tampered with during transmission. After the user terminal sends a request to the server, the server returns the corresponding result within 1 second. 3. Detection platform and user terminal: Short-range communication is conducted using the Bluetooth 5.0 protocol. Users can directly control the start and stop of the detection platform and view the real-time collected data through the terminal.

[0043] III. Interactive Control Process like Figure 2As shown, the system's interactive control flow is divided into four stages: user authentication, data acquisition, processing and analysis, and feedback intervention. The specific steps are as follows: 1. User Authentication Stage: The user scans the QR code on the detection platform through their user terminal to complete identity binding; or directly inputs "My ID is XXX" on the detection platform via voice. The voice interaction module recognizes the user ID and completes authentication. The foot electrode array on the detection platform detects that the user is standing barefoot, and the impedance value is below the threshold, triggering the detection process. All hardware modules complete initialization (taking no more than 2 seconds). 2. Data Acquisition Stage: The foot electrode array initiates BIA scanning, applying AC currents of 5kHz, 50kHz, and 100kHz, measuring the impedance value of each electrode at a sampling rate of 100Hz for 30 seconds to acquire body composition data. The millimeter-wave radar module initiates respiratory data acquisition at a sampling rate of 200Hz for 60 seconds, capturing chest rise and fall data for 10 complete respiratory cycles and extracting respiratory features. The 3D optical scanning module initiates a 3D volumetric mesh scan, the structured light projector projects a coded pattern, the camera captures the image at a sampling rate of 30fps, and the scan time is 20 seconds, generating 3D point cloud data. The voice interaction module initiates a voice query in the last 10 seconds of data acquisition, such as "How much vegetables have you consumed each day for the past week?" The user answers via voice, and the ASR module converts the speech into text. The local data processing unit performs time alignment on the multimodal data, using the start time of the BIA scan as a reference, adjusting the timestamps of the radar and optical data to the same timeline with an error not exceeding 1ms. 3. Processing and Analysis Stage: The local data processing unit performs wavelet transform denoising on the BIA data to remove 50Hz power frequency interference; performs Kalman filtering on the millimeter-wave radar data to remove motion noise; and performs statistical filtering on the 3D optical data to remove outliers. The preprocessed multimodal data is encrypted with AES-256 and then uploaded to the backend server via a 5G network. The backend server's data storage unit stores data in a distributed storage system. The MFN model inference unit loads the data for inference, generating a nutritional risk assessment result. The report generation unit combines the inference result with the user's basic information to generate a nutritional risk report; the intervention suggestion generation unit combines the report with voice interaction information to generate personalized intervention suggestions. 4. Feedback Intervention Phase: The backend server pushes the nutritional risk report and intervention suggestions to the user's terminal, and simultaneously converts them into voice feedback to the user through the testing platform's TTS module, such as "Your nutritional risk assessment result is moderate, indicating a risk of protein deficiency. It is recommended to increase your intake of lean meat by 50g per day." Users can view detailed reports on their terminals or ask specific details about the intervention suggestions via voice, such as "Which foods are rich in protein?" The system provides real-time answers through the voice interaction module.

[0044] IV. Introduction to Nutritional Detection and Intervention Methods Based on the MFN Model (Core Technical Points) like Figure 3 As shown, this method includes the following steps: Step 1: Synchronous Acquisition of Multimodal Data (I) BIA Body Composition Data Acquisition: After the foot electrode array of the detection platform detects the user's foot contact, it triggers the BIA scanning process. The system sequentially applies sinusoidal alternating currents of 5kHz, 50kHz, and 100kHz (current intensity of 0.5mA), measures the voltage value of each electrode, and calculates the impedance value R = U / I according to Ohm's law. At the same time, the user's weight W is measured by a pressure sensor (accuracy of 0.1kg). Combined with the user's pre-input height H, age A, and gender G, the body composition index is calculated. Different frequencies of current penetrate human tissue to varying depths. A 5kHz current primarily penetrates the extracellular fluid, a 50kHz current penetrates both intracellular and extracellular fluids, and a 100kHz current penetrates muscle tissue. By measuring the impedance values ​​at different frequencies, indicators such as body fat percentage (BF%), muscle mass (MM), and basal metabolic rate (BMR) can be calculated. Examples of the calculation formulas are as follows: Male body fat percentage: BF% = 0.546*R + 0.13*W -0.012*H + 0.05*A + 1.43; Female body fat percentage: BF% = 0.56*R + 0.12*W - 0.01*H + 0.04*A + 5.7; Male basal metabolic rate: BMR = 66.5 + 13.75*W + 5.003*H - 6.775*A; Female basal metabolic rate: BMR = 655.1 + 9.563*W + 1.85*H - 4.676*A, through multi-frequency BIA scanning, the error in body fat percentage measurement does not exceed 2%, and the error in basal metabolic rate measurement does not exceed 5%, meeting the accuracy requirements for clinical testing.

[0045] (II) Millimeter-wave radar respiratory data acquisition: After the millimeter-wave radar module is activated, it transmits a 77GHz frequency-modulated continuous wave and receives electromagnetic waves reflected from the human chest cavity. The sampling rate is 200Hz, and data is continuously acquired for 60 seconds. The system uses a peak detection algorithm to identify the start and end times of each respiratory cycle, ensuring that 10 complete respiratory cycles are captured. The formula for calculating the respiratory cycle is T = t_end - t_start, where t_start is the start time of the respiratory cycle and t_end is the end time; the respiratory rate f = 60 / T_avg, where T_avg is the average value of 10 respiratory cycles; the respiratory rhythm variation coefficient... %,in The standard deviation is 10 respiratory cycles. These characteristics reflect the user's metabolic status; for example, a respiratory rate exceeding 20 breaths / minute may indicate hypermetabolism, and a respiratory rhythm variability coefficient exceeding 15% may indicate autonomic nervous system dysfunction. The respiratory cycle detection accuracy reaches 99%, and the respiratory rate measurement error is no more than 0.5 breaths / minute, effectively assessing the user's metabolic status.

[0046] (III) 3D Optical Body Shape Data Acquisition: After the 3D optical scanning module is activated, the structured light projector projects a coded infrared pattern with a resolution of 1024×768. Two cameras capture deformed images of the pattern from different angles. The system calculates the three-dimensional coordinates of each point using the triangulation principle, generates point cloud data, and then generates a 3D volumetric mesh model using the Poisson reconstruction algorithm. Indicators such as waist circumference (WC), hip circumference (HC), and upper arm circumference (AC) are extracted. The calculation formula is WC = 2*π*r, where r is the average radius of the waist, calculated from the point cloud coordinates of the volumetric mesh model. The coded pattern projected by the structured light is unique; the coded value of each pixel is different. By comparing the coded values ​​of the patterns captured by the two cameras, the depth information of each point can be calculated. After statistical filtering to remove outliers, the point cloud density reaches 1000 points / cm², and the measurement error of the body shape indicators does not exceed 5mm. It can accurately extract users' body shape characteristics and assess body fat distribution. For example, the waist-to-hip ratio (WHR) = WC / HC. A WHR greater than 0.9 (men) or 0.85 (women) indicates central obesity and a risk of cardiovascular disease.

[0047] (iv) Voice Interaction Data Acquisition: In the last 10 seconds of data acquisition, the voice interaction module initiates voice queries, such as "What has been your staple food for the past week—rice, noodles, or something else?" and "How much exercise do you do each day?" The user answers via voice, the microphone array captures the user's voice, and the ASR module converts the voice into text information, extracting key features (such as staple food type and exercise intensity). The ASR module uses the Transformer architecture's ERNIE-Speech pre-trained model. Through pre-training on massive amounts of Chinese voice data, it achieves high-accuracy voice recognition. Even in the presence of environmental noise (noise intensity not exceeding 60dB), the recognition accuracy still reaches over 95%. It can quickly obtain information on the user's dietary and exercise habits, supplementing the dimensions of nutritional screening and improving the accuracy of nutritional risk assessment.

[0048] Step 2: Local Data Preprocessing (I) BIA Data Denoising: Wavelet transform is used to remove power frequency interference. The db4 wavelet is selected as the basis function. The BIA data is decomposed into 5 layers. The power frequency noise (50Hz) of the 5th layer is removed, and then the data is reconstructed, improving the signal-to-noise ratio by more than 15dB. (II) Millimeter-Wave Radar Data Filtering: Kalman filtering is used to remove motion noise. The state equation X(k) = A*X(k-1) + B*U(k-1) + W(k-1) and the observation equation Z(k) = H*X(k) + V(k) are established, where X(k) is the thoracic displacement state at time k, A is the state transition matrix, B is the control matrix, U(k-1) is the control input, W(k-1) is the process noise, Z(k) is the observation value, H is the observation matrix, and V(k) is the observation noise. The optimal thoracic displacement estimate is obtained through iterative calculation to remove the influence of motion noise. (III) 3D Optical Data Denoising: Outliers are removed using statistical filtering. The average distance to the neighboring points of each point cloud is calculated, and points whose distance exceeds three times the standard deviation of the average distance are considered outliers and removed. The number of noisy points in the point cloud is reduced by more than 90%. (IV) Time Alignment: Based on the start time of the BIA scan, the timestamps of the millimeter-wave radar data and the 3D optical data are adjusted to the same time axis with an error of no more than 1ms, ensuring the time synchronization of multimodal data and laying the foundation for subsequent multimodal fusion.

[0049] Step 3: MFN Model Inference (a) Model Input: The input to the MFN model is the multimodal feature vector X = ,in: ∈ R 8 BIA body composition feature vector, including body fat percentage, muscle mass, water content, basal metabolic rate, protein content, fat content, bone mass, and visceral fat level; ∈ R 3 : Respiratory feature vector of millimeter-wave radar, including respiratory rate, respiratory rhythm variation coefficient, and chest wall rise and fall amplitude; ∈ R 6 3D optical body shape feature vector, including waist circumference, hip circumference, upper arm circumference, waist-to-hip ratio, body surface area, and body fat distribution uniformity; ∈ R 5 Voice interaction feature vectors include encoding of staple food type, vegetable intake, meat intake, exercise intensity, and medical history.

[0050] (II) Model Structure and Formulas: 1. Modal Feature Encoding Layer: Encodes the feature vector of each modality and extracts higher-order features, as shown in the following formula:

[0051] Where W1∈R64×8 , W2 ∈ R 64×3 , W3 ∈ R 64×6 , W4 ∈ R 64×5 Let b1, b2, b3, b4 ∈ R be the weight matrix for each modality. 64 Let σ be the bias vector, and σ be the Sigmoid activation function. 1. Mapping features to the (0,1) interval enhances the non-linear expressive power of features. 2. Multimodal fusion layer: Employing an element-wise multiplication + feature concatenation fusion method, the encoded multimodal features are fused using the following formula: Z_concat = [ ⊙ ⊙ ⊙ ], , where ⊙ is the Hadamard product (element-wise multiplication), corresponding to element-wise multiplication, capturing the interaction between BIA features and radar features; [] is the feature concatenation operation, and the dimension of the concatenated feature vector Z_concat is 2*64+6+5=139; The fused feature vector is the final feature representation of multimodal data after encoding, fusion, and nonlinear transformation, and serves as the input to the subsequent classification output layer; W5 ∈ R 128×139 Here, b5 represents the weight matrix of the fusion layer, which consists of parameters learned during model training (where 139 is the dimension of the input feature Z_concat, and 128 is the dimension of the output feature Z); 128 The bias vector is a parameter learned during model training, used to adjust the offset of the linear transformation and improve the model's fitting ability.

[0052] This formula achieves deep fusion of multimodal features: firstly through... The concatenated features undergo a linear transformation, followed by the introduction of a non-linearity through the Sigmoid activation function σ, ultimately outputting the fused feature Z, which provides input for subsequent classification of nutritional risk levels and risk points. 3. Classification Output Layer: The fused features are classified, outputting the nutritional risk level and specific risk point, using the following formulas: Y_risk = softmax(W6*Z + b6), Y_point = σ(W7*Z + b7), where softmax is the Softmax activation function.

[0053] The Softmax activation function is used to transform the raw output (logits) in a multi-class classification problem into a probability distribution, mapping the input vector to the (0,1) interval, and the sum of all output values ​​is 1, which can be directly interpreted as class probability; The i-th element of the input vector corresponds to the model's original output for the i-th category (e.g., "low risk" or "medium risk" in nutritional risk levels). Example: In nutritional risk assessment... This can represent the model's raw score for the "high-risk" category; Natural exponential function pairs The calculation result is used to amplify the differences between input values, enhance class discrimination, and the monotonicity of the exponential function ensures a large... A larger molecular value corresponds to a higher probability of that category; For all categories j Summation, used as the normalized denominator, scales the exponential result of the numerator to the (0,1) interval by summing, ensuring that the output meets the basic requirements of the probability distribution (non-negativity, summation to 1). Mapping the output to the (0,1) interval, with the sum of all outputs being 1, Y_risk ∈ R 3 The probabilities corresponding to low risk, medium risk, and high risk; Y_point ∈ R 5 The probability of corresponding protein deficiency, fat excess, vitamin deficiency, mineral deficiency, and energy deficiency is considered to exist if the probability value exceeds 0.5; W6 ∈ R 3×128 W7∈R 5×128 Let b6 be the weight matrix, and b6 ∈ R. 3 b7 ∈ R 5 This is the bias vector.

[0054] (III) Model Training and Optimization: A multimodal dataset was used for training. The dataset included multimodal data from 10,000 users and corresponding clinical nutrition risk assessment results (NRS2002 scores), divided into a training set (80%), a validation set (10%), and a test set (10%). The loss function was a combination of the cross-entropy loss function and the weighted loss function: L = L_risk + λ*L_point, where L_ The cross-entropy loss function for risk level classification, L_p The cross-entropy loss function is used for risk point classification, with λ=0.8 as the weight coefficient. The optimizer is the Adam optimizer with a learning rate of 0.001, a batch size of 32, and 100 training epochs. After each training epoch, the model's accuracy is evaluated on the validation set. Training stops when the validation set accuracy fails to improve for five consecutive epochs to prevent overfitting. On the test set, the model achieves 94% accuracy in nutritional risk level classification and 92% accuracy in risk point classification, significantly higher than the accuracy of the unimodal model.

[0055] Step 4: Generate Nutrition Risk Report The report structure includes seven parts: cover page, user basic information, body composition analysis, respiratory characteristic analysis, body shape characteristic analysis, nutritional risk assessment, and personalized intervention recommendations. The generation process involves the report generation unit calling a predefined template, filling the template with model inference results and user basic information, and generating reports in PDF and HTML formats. Data in the report is presented using charts; for example, body composition analysis uses a pie chart to show the proportions of body fat percentage, muscle mass, and water content; respiratory characteristic analysis uses a line graph to show the tidal volume change trend over 10 respiratory cycles; and body shape characteristic analysis uses a color heatmap of a 3D mesh model to present differences in body fat distribution. The specific implementation process is as follows: 1. Template Calling and Data Binding: The system has 5 built-in professional report templates (adult standard version, senior citizen version, child growth version, athlete enhanced version, and chronic disease management version), automatically matching the optimal template based on the user's basic information. The templates adopt an XML structured design, binding the 128-dimensional feature vector output by the MFN model to the template fields via XPath paths. For example, mapping the volume component feature vector [0.23, 0.35, 0.42] to the pie chart data labels. <data id="body_composition">.

[0056] 2. Chart Rendering Engine: Integrates the ECharts visualization library to achieve dynamic chart generation, supporting 3 categories and 12 chart types. The body composition analysis module calls the pieChart() function, with input parameters {data:[{name:'body fat percentage',value:23},{name:'muscle mass',value:35},{name:'water',value:42}],radius:['40%','70%']}, automatically generating a pie chart with percentages and marking the normal reference range (body fat percentage 18-25%). The respiratory feature analysis uses the lineChart() function to draw a respiratory waveform graph, with the X-axis representing the time series (0-60 seconds) and the Y-axis representing tidal volume (ml), simultaneously displaying derived indicators such as respiratory rate (breaths / minute) and respiratory depth coefficient of variation.

[0057] 3. Risk Level Visualization: The nutrition risk assessment module uses a dashboard chart (gaugeChart) to display the overall risk score (0-10 points). The red warning zone (>7 points), yellow alert zone (5-7 points), and green safe zone (<5 points) correspond to high, medium, and low risk levels, respectively. Risk point analysis uses a radar chart to present the Z-score standardized values ​​of five core indicators (protein, fat, vitamins, minerals, and energy), intuitively displaying the multidimensional nutritional imbalance.

[0058] 4. Personalized suggestion generation: Based on a two-layer reasoning mechanism of rule engine and knowledge graph, risk points and intervention measures are intelligently matched. For example, when a protein deficiency risk is detected (Y_point[0]>0.5), the system automatically calls the food database to select high-quality sources of protein with a biological value ≥85 (eggs, milk, fish), and generates weekly menu suggestions based on the user's dietary preferences (obtained through voice interaction), and presents the types, weights and nutritional content of food for three meals a day in tabular form.

[0059] 5. Multi-format output and encrypted transmission: The report generation unit uses the Apache FOP tool to convert XML templates into PDF format and employs 128-bit AES encryption to protect user privacy data. It also generates responsive HTML reports, supporting adaptive display on mobile and PC devices, and includes interactive charts (supporting zooming and hovering over data points for details). The generated reports are pushed to the user's terminal via HTTPS and simultaneously stored in a blockchain-based notarization system to ensure data immutability.

[0060] Example 1: Detection and intervention of protein deficiency risk in elderly users 1. User Scenario: A 68-year-old male user, 172cm tall, weighing 65kg, with a history of hypertension, whose daily diet consists mainly of staple foods and has insufficient protein intake. The user stands barefoot on the testing platform, and the system automatically initiates the multimodal data acquisition process.

[0061] 2. Data acquisition process: (1) BIA module: Apply three frequency currents of 5kHz, 50kHz and 100kHz, measure the impedance value R=520Ω, combine the weight W=65kg, height H=172cm and age A=68 years, substitute into the male body fat percentage formula BF%=0.546×R+0.13×W-0.012×H+0.05×A+1.43, calculate BF%=28.3% (higher than the normal range of 18-25%); at the same time, extract muscle mass of 28.5kg (normal reference value 30-35kg), which is judged to be low muscle mass.

[0062] (2) Millimeter-wave radar: 12 respiratory cycles were captured within 60 seconds, and the respiratory rate was extracted as 18 breaths / minute (normal range 12-16 breaths / minute). The coefficient of variation of respiratory depth was 15% (normal range <10%), indicating abnormal metabolic rate.

[0063] (3) 3D optical scanning: Generate a volume mesh model, calculate waist circumference of 92cm (normal <85cm for males), hip circumference of 101cm, waist-to-hip ratio of 0.91 (normal <0.9), and body fat distribution heat map shows abdominal fat accumulation.

[0064] (4) Voice interaction: The system initiates a question through the TTS module: "How much meat, eggs or soy products do you consume on average every day?" The user answers in voice: "I drink one carton of milk every day and rarely eat eggs or meat." The ASR module converts the voice into text and extracts the dietary feature "low protein intake frequency".

[0065] 3. MFN Model Inference: (1) Feature encoding: BIA feature X1 (dimension 64), radar feature X2 (dimension 64), optical feature X3 (dimension 6), and speech feature X4 (dimension 5) are encoded and then generated as Z_concat (dimension 139) by element-wise multiplication (X1⊙X2⊙X3⊙X4) and feature concatenation.

[0066] (2) Feature fusion: The fused feature Z (128-dimensional) is generated by formula Z=σ(W5×Z_concat+b5), where W5 is a 128×139 weight matrix, b5 is a 128-dimensional bias vector, and σ is the Sigmoid function.

[0067] (3) Risk classification: Y_risk=softmax(W6×Z+b6) outputs the risk level probability [0.12,0.75,0.13], which is judged as medium risk; Y_point=σ(W7×Z+b7) outputs the risk point probability [0.89,0.32,0.65,0.41,0.28], with significant risks of protein deficiency (0.89>0.5) and vitamin deficiency (0.65>0.5).

[0068] 4. Report Generation and Intervention Recommendations: (1) Report content: Actively matched the "elderly special version" template. The body composition analysis pie chart showed a body fat percentage of 28.3% (red mark indicates excessive) and muscle mass of 28.5kg (yellow mark indicates low); the respiratory characteristic line chart indicated that the respiratory rate was too fast; the 3D heat map marked the high body fat area in the abdomen.

[0069] (2) Personalized suggestions: The rule engine calls the food database to select foods with a protein biological value ≥85 and suitable for the elderly to digest (tofu, fish, low-fat milk), and generates a weekly menu based on the user's dietary preferences: add 1 egg to breakfast, add 100g of steamed fish to lunch, and add 200g of tofu to dinner; at the same time, it recommends 30 minutes of Tai Chi exercise every day (based on the breathing characteristics detected by millimeter-wave radar, avoid strenuous exercise).

[0070] 5. Results: One month after the intervention as recommended, the user underwent a follow-up examination. BIA showed that muscle mass increased to 30.2 kg, the risk of protein deficiency decreased to 0.32, respiratory rate recovered to 15 breaths / minute, and body fat percentage decreased to 25.1%.

[0071] Example 2: Detection and Intervention of Vitamin D Deficiency Risk in Children 1. User Scenario: A 5-year-old girl, 115cm tall and weighing 22kg, whose parents reported frequent colds and night sweats recently. The monitoring platform automatically switches to the "Children's Growth Version" mode, enabling a simplified voice interaction process.

[0072] 2. Data acquisition process: (1) BIA module: Using children-specific electrode pads, the measured impedance value R=480Ω, and substituting into the children's body fat percentage formula BF%=0.42×R+0.08×W-0.005×H+0.02×A+3.2, the calculated BF%=18.7% (normal range 15-20%); muscle mass and water content are normal, but the mineral characteristic vector [0.21,0.18,0.35] indicates abnormal calcium and phosphorus metabolism.

[0073] (2) Millimeter-wave radar: captures 15 respiratory cycles / minute, respiratory depth variation coefficient of 8% (normal), but the average tidal volume is 180ml (lower than the reference value of 220ml for children of the same age).

[0074] (3) 3D optical scan: a body model was generated, and the upper arm circumference was calculated to be 16cm (normal), but the lower edge of the rib protrusion was 1.2cm (higher than the normal threshold of 0.8cm), suggesting that there may be abnormal bone development caused by vitamin D deficiency.

[0075] (4) Voice interaction: The system asks questions in a cartoon voice: "How long do you play outdoors every day? Do you eat milk and dried fish?" The user answers: "I play indoors every day and I don't like to eat dried fish." The features extracted are "insufficient sunlight + insufficient fish intake".

[0076] 3. MFN Model Inference: (1) Feature fusion: Z_concat fuses features specific to children (such as bone age prediction features). After processing by Z=σ(W5×Z_concat+b5), the weight of vitamin D deficiency-related features in the output Z vector reaches 0.72.

[0077] (2) Risk classification: Y_risk=[0.65,0.33,0.02] is judged as low risk, but the probability of vitamin deficiency in Y_point is 0.78>0.5 and the probability of mineral deficiency is 0.61>0.5. Combined with the bone characteristics, it is judged as calcium absorption disorder caused by vitamin D deficiency.

[0078] 4. Report Generation and Intervention Recommendations: (1) Report content: The 3D model is marked with the rib protrusion area and a "vitamin D-calcium metabolism correlation analysis diagram" is generated, showing a negative correlation between sunshine duration and serum 25-OH-VD level (r=-0.68).

[0079] (2) Personalized recommendations: Knowledge graph matching children’s exclusive intervention plan: daily supplementation of 400 IU vitamin D drops (taken twice a day, morning and evening), eating salmon 3 times a week (50g each time, cooked in a cartoon shape), and ensuring 1 hour of outdoor sunlight every day (recommended during the suitable ultraviolet radiation period of 9:00-10:00).

[0080] 5. Results: Two months after the intervention, 3D scans showed that the rib protrusion decreased to 0.6cm, the BIA mineral feature vector returned to normal, and parents reported a significant reduction in the frequency of their children's colds.

[0081] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 4 As shown, optionally, electronic device 410 may include a first processor 2001.

[0082] Optionally, the electronic device 410 may also include a memory 2002 and a transceiver 2003.

[0083] The first processor 2001, memory 2002, and transceiver 2003 can be connected via a communication bus.

[0084] The following is combined Figure 4 A detailed description of each component of electronic device 410 is provided below: The first processor 2001 is the control center of the electronic device 410. It can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).

[0085] Optionally, the first processor 2001 can perform various functions of the electronic device 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.

[0086] In a specific implementation, as one example, the first processor 2001 may include one or more CPUs, for example... Figure 4 CPU0 and CPU1 are shown in the diagram.

[0087] In a specific implementation, as one example, the electronic device 410 may also include multiple processors, for example... Figure 4 The first processor 2001 and the second processor 2004 are shown in the diagram. Each of these processors can be a single-core processor or a multi-core processor. Here, a processor can refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).

[0088] The memory 2002 is used to store the software program that executes the present invention, and is controlled by the first processor 2001 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.

[0089] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and be connected via the interface circuit of the electronic device 410. Figure 4 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.

[0090] The transceiver 2003 is used to communicate with network devices or with terminal devices.

[0091] Alternatively, transceiver 2003 may include a receiver and a transmitter. Figure 4 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.

[0092] Optionally, the transceiver 2003 can be integrated with the first processor 2001, or it can exist independently and be connected via the interface circuit of the electronic device 410. Figure 4 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.

[0093] It should be noted that, Figure 4 The structure of the electronic device 410 shown does not constitute a limitation on the router. Actual knowledge structure identification devices may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0094] Furthermore, the technical effects of the electronic device 410 can be referred to the technical effects of the XXX method described in the above method embodiments, and will not be repeated here.

[0095] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0096] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0097] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0098] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0099] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.

[0100] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0101] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0102] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0103] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0104] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0105] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0106] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0107] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.< / data>

Claims

1. A user nutrition detection and intervention system based on artificial intelligence, characterized in that, The system includes: a detection platform, a backend server, and user terminals, wherein: The detection platform is a hardware entity that integrates a foot electrode array, a millimeter-wave radar module, a three-dimensional optical scanning module, and a voice interaction module. It is used to simultaneously collect bioelectrical impedance data, respiratory characteristic data, three-dimensional volumetric mesh data, and interactive voice data of users while they are standing barefoot. The backend server is connected to the detection platform via a 5G or WiFi network and is equipped with a data storage unit, a multimodal fusion model inference unit, and an intervention suggestion generation unit. It is used to perform fusion inference on the received encrypted data, generate a risk assessment result containing nutritional risk level and specific risk points, and generate a personalized nutritional intervention suggestion report accordingly. The user terminal connects to the backend server via the HTTPS protocol for user authentication, receiving and visualizing the intervention suggestion report, and interacting with the detection platform.

2. The user nutrition detection and intervention system based on artificial intelligence according to claim 1, characterized in that, The foot electrode array of the detection platform specifically includes: The capacitive electrode array has eight electrodes evenly distributed in the left and right foot placement areas, with an electrode diameter of 10 mm and a spacing of 20 mm. The electrode array is used to automatically identify the user's foot contact. When the impedance value of the detected human tissue is lower than 1000Ω, it triggers a bioelectrical impedance analysis scan. It is also used to apply sinusoidal alternating currents at three frequencies of 5 kHz, 50 kHz, and 100 kHz to the human body and measure its impedance to extract body composition characteristic data, including body fat percentage, muscle mass, water content, and basal metabolic rate.

3. The user nutrition detection and intervention system based on artificial intelligence according to claim 1, characterized in that, The millimeter-wave radar module of the detection platform is a 77GHz frequency-modulated continuous wave radar with a sampling rate of 200Hz and a detection range of 0.5-2m. This module is used to capture chest wall fluctuation data for at least 10 complete respiratory cycles within 60 seconds. It calculates chest wall displacement by analyzing the frequency changes of reflected waves, and after removing motion noise using a Kalman filter algorithm, it extracts the user's respiratory rate, respiratory rhythm variation coefficient, and chest wall fluctuation amplitude data to assess the user's metabolic status.

4. The user nutrition detection and intervention system based on artificial intelligence according to claim 1, characterized in that, The detection platform's three-dimensional optical scanning module includes an infrared structured light projector and two high-definition cameras, each with a resolution of 2MP and a frame rate of 30fps. The module is used to project coded infrared patterns onto the human body and calculate the three-dimensional coordinates of the point cloud using the triangulation principle to construct a three-dimensional volumetric mesh model for the user. The point cloud density of the model reaches 1000 points / cm². After being generated by the Poisson reconstruction algorithm and statistically filtered to remove noise, it can calculate body shape characteristic indicators including waist circumference, hip circumference, waist-to-hip ratio, and body fat distribution uniformity.

5. The user nutrition detection and intervention system based on artificial intelligence according to claim 1, characterized in that, The detection platform also includes a local data processing unit and a voice interaction module, wherein: The local data processing unit, built on an ARM processor, is used to perform time alignment, noise reduction, and encryption processing on the collected bioelectrical impedance data, millimeter-wave radar data, and three-dimensional optical data, and uploads the encrypted data via the network after packaging it using the AES-256 encryption algorithm. The voice interaction module integrates a microphone array, an automatic speech recognition module, and a text-to-speech module. It is used to ask users questions related to diet, exercise, etc. in voice form during the later stages of data collection, and convert the user's voice response into text data through a speech recognition model. At the same time, it broadcasts the reports and suggestions generated by the backend server to the user.

6. The user nutrition detection and intervention system based on artificial intelligence according to claim 1, characterized in that, The multimodal fusion model inference unit of the backend server is, at its core, a multimodal fusion neural network model deployed on a GPU cluster. This model outputs a nutritional risk assessment through the following processing flow: (1) Modal feature encoding: The input BIA body component feature vector X1, breathing feature vector X2, three-dimensional optical body shape feature vector X3, and voice interaction feature vector X4 are multiplied by the corresponding weight matrices W1~W4 and bias terms b1~b4 are added. The higher-order feature vectors are obtained by mapping through the Sigmoid activation function σ. (2) Multimodal fusion: The encoded high-order features are subjected to Hadamard product (element-level multiplication) operation and then concatenated with some features to form an aggregated feature vector. The final fusion feature vector Z is generated again through linear transformation and nonlinear activation. (3) Classification output: The fused feature vector Z is calculated by the classification layer, and the nutritional risk level vector Y_risk is output by the Softmax function, and the probability vector Y_point of the specific nutritional risk point is output by the Sigmoid function, thereby realizing multi-dimensional nutritional risk assessment at the same time.

7. The user nutrition detection and intervention system based on artificial intelligence according to claim 1, characterized in that, The intervention suggestion generation unit is coupled with the multimodal fusion model reasoning unit and the data storage unit. It is used to generate personalized comprehensive intervention suggestions based on the risk points inferred by the model, combined with the pre-set nutritional knowledge graph and the "Chinese Dietary Guidelines" standard. ‌ The intervention recommendations include at least dietary advice, exercise programs, and lifestyle adjustments, and can automatically match and output dedicated report templates based on the user's age and health status.

8. The user nutrition detection and intervention system based on artificial intelligence according to claim 1, characterized in that, The data communication between the various parts of the system specifically includes: The local data processing unit of the detection platform encrypts the collected data using AES-256 and uploads it to the backend server via 5G NR or WiFi 6 network. After the backend server decrypts and processes the data, it securely pushes the report and suggestions to the user terminal via HTTPS protocol. The detection platform and the user terminal establish a direct connection via Bluetooth 5.0 protocol for issuing control commands and displaying real-time data.

9. The user nutrition detection and intervention system based on artificial intelligence according to any one of claims 1 to 8, characterized in that, The data storage unit of the backend server is a distributed storage architecture, which includes an HDFS file system for storing massive amounts of time-series data, and a MongoDB database for managing user reports and model metadata, while also de-identifying user data.

10. A user nutrition detection and intervention method based on artificial intelligence applied to the system described in any one of claims 1-9, characterized in that, The method includes the following steps: S1: Synchronous acquisition of multimodal data When a user stands barefoot in a designated area of ​​the detection platform, multimodal data acquisition is automatically triggered. S2: Local data preprocessing and encrypted transmission The local data processing unit of the detection platform receives the raw data collected in step S1, performs noise reduction, filtering and time alignment operations in sequence, and then uses the AES-256 algorithm to encrypt the integrated multimodal data packets and transmits the encrypted data packets to the backend server through the 5G / WiFi network. S3: Risk Assessment Based on Multimodal Fusion Model The backend server receives and decrypts the data packet, and inputs the cleaned multimodal feature vector into a pre-trained multimodal fusion neural network for inference: The model calculates layer by layer through modal feature encoding layer, multimodal fusion layer and classification output layer to obtain the probability distribution of nutritional risk level (low, medium and high) and the binary judgment probability of specific risk points. Based on the risk point assessment results, combined with the built-in nutritional standards and knowledge graph, the system initially identifies the main types of nutritional risks. S4: Personalized Intervention Report Generation and Feedback The report generation unit on the backend server, based on the risk assessment results of step S3, retrieves a report template that matches the user's age and risk characteristics, and generates a structured nutrition risk report with illustrations and text; the intervention suggestion generation unit, based on specific risk points, integrates dietary preferences and health goals to generate actionable dietary, exercise, and lifestyle suggestions, and packages them together. Finally, the server will push the generated full report and recommendations to the user's mobile app or mini-program via an HTTPS secure channel. At the same time, it can also broadcast the report to the user via the text-to-speech module of the detection platform, completing the complete closed-loop process from detection to intervention.