A multifunctional mirror integrating health monitoring and intelligent interaction

By integrating multimodal sensing components and deep learning technology, the smart mirror achieves high-precision physiological parameter monitoring and personalized interaction in complex lighting and noise environments, solving the problems of insufficient monitoring accuracy and weak interaction capabilities of existing smart mirrors under lighting and noise conditions, and improving the user experience.

CN122074784APending Publication Date: 2026-05-26BEIJING LINKE INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING LINKE INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-02-06
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing smart mirrors suffer from insufficient accuracy in non-contact physiological monitoring under complex lighting and motion scenarios, lack contextualized interaction strategies based on user identity and behavioral characteristics, and have weak anti-interference capabilities for voice recognition in noisy bathroom environments.

Method used

It employs integrated multimodal sensing components, including a camera, ambient light sensor, microphone array, and central processing unit, and combines deep learning feature extraction, signal separation technology, and image pyramid processing to achieve high-precision physiological parameter monitoring and personalized interaction.

Benefits of technology

Achieving high-precision heart rate monitoring and voice recognition in complex environments, providing personalized information push, and improving user experience and the effectiveness of information delivery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122074784A_ABST
    Figure CN122074784A_ABST
Patent Text Reader

Abstract

This application relates to the field of smart home technology and discloses a multifunctional mirror integrating health monitoring and intelligent interaction. The device includes a semi-transparent, semi-reflective mirror body with a display module, a multispectral camera, and a control panel. The system operates data processing logic, using visual algorithms to locate the facial ROI area, and employing orthogonal projection in color space to separate pulse signals to monitor heart rate and body temperature. It identifies the user's identity through deep learning feature matching and automatically switches between adult and child modes. Based on skeletal key point tracking, it determines brushing behavior and times it. The system intelligently pushes information or educational content based on user attributes, time period, and actions, and combines visual coordinates to guide microphone array beamforming, using a logarithmic mapping model to adaptively adjust the backlight. This invention effectively improves the robustness of non-contact physiological monitoring under complex lighting conditions and achieves personalized health guidance and interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart home technology, specifically to a multifunctional mirror that integrates health monitoring and intelligent interaction. Background Technology

[0002] With the popularization of IoT and AI technologies, smart home devices are gradually evolving from simple control terminals into intelligent nodes with sensing and decision-making capabilities. Smart mirrors, as core devices in bathrooms and bedrooms, utilize their semi-transparent and semi-reflective optical properties to hide the display screen behind the mirror surface. While providing traditional mirroring functions, they also present users with time, weather, and multimedia information, becoming an important carrier for modern family health management and information interaction.

[0003] While existing smart mirror products are relatively mature in displaying basic information, significant technical bottlenecks remain in practical applications such as non-contact health monitoring and contextualized interaction. In health monitoring, current solutions often employ facial video analysis technology based on ordinary cameras to extract physiological parameters such as heart rate. However, bathroom environments typically involve complex and dynamically changing lighting conditions, such as shadows from direct overhead lights or facial highlights. Furthermore, the unavoidable body movements during washing introduce strong non-physiological noise into the optical signals. Traditional signal processing algorithms often struggle to effectively separate weak pulse wave signals from environmental noise, resulting in poor stability and low accuracy in physiological parameter measurements, failing to meet users' expectations for the reference value of health data.

[0004] Furthermore, existing devices lack targeted contextual awareness capabilities in their human-computer interaction logic. Most smart mirrors cannot accurately distinguish the current user's identity attributes and specific behavioral state, typically employing only a single broadcast-style information push mode. This means the system cannot provide timely action guidance and educational content while children brush their teeth, nor can it automatically filter high-priority schedules and news summaries during busy morning hours for adults, resulting in limited efficiency of information services and a poor user experience. Simultaneously, the noise from running water and exhaust fans commonly found in bathroom settings severely interferes with the accuracy of voice interaction. Ordinary voice acquisition solutions often experience wake-up difficulties or command recognition errors in such high-noise environments, limiting the practicality of contactless interaction. Therefore, developing a smart mirror that can adapt to complex environmental lighting and noise interference, and possesses high-precision physiological monitoring and hierarchical contextual interaction capabilities, has become an urgent problem to be solved in the current technological field. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a multifunctional mirror that integrates health monitoring and intelligent interaction. It solves the problems faced by existing smart mirror technologies, such as insufficient accuracy of non-contact physiological monitoring in complex lighting and motion scenarios, lack of contextualized interaction strategies based on user identity and behavioral characteristics, and weak anti-interference capability of voice recognition in noisy bathroom environments.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a multifunctional mirror integrating health monitoring and intelligent interaction, comprising a support structure and a hardware interaction terminal. The mirror body is fixed to the front of the support frame, and its structure includes a semi-transparent, semi-reflective glass layer and a display module attached to the back of the glass layer, enabling optical state switching between mirror reflection and information display. The system integrates multimodal sensing components, including a camera mounted on the mirror body. This camera integrates an infrared imaging sensor or a multispectral sensor to acquire video stream data including visible and invisible light bands. The system also includes an ambient light sensor to sense the external lighting environment. The control panel integrates a central processing unit, a microphone array, a speaker assembly, and a network communication unit, forming the central hub for data processing and interaction.

[0007] At the control logic level, the central processing unit operates with a control logic architecture, achieving intelligent operation through multiple functional modules: Regarding data acquisition and image preprocessing: To ensure the accuracy of subsequent analysis, the data acquisition and preprocessing module performs image standardization operations. The system first sets the camera's sampling resolution and frame rate parameters, and activates automatic white balance and automatic exposure control logic to adapt to changes in ambient light. In the face detection stage, a feature-based classification detection algorithm scans image frames in the original video stream. A sliding window is used to traverse the image, and a non-maximum suppression algorithm is used to eliminate redundant windows, thereby accurately locating the facial region's bounding box and cropping out a facial sub-image. Subsequently, histogram equalization is performed on this facial sub-image, and the cumulative distribution function is used to calculate the mapping relationship from the original gray level to the target gray level, constructing an image pyramid sequence to enhance the local contrast of the image.

[0008] Regarding biometric identification and physiological parameter monitoring: The biometric recognition and analysis module undertakes the dual tasks of identity verification and physiological signal extraction. In terms of identity recognition, the system calls a deep learning feature extraction model to map the facial sub-image into a high-dimensional feature vector in the feature space. The system determines the user's identity by calculating the cosine similarity between this vector and the pre-stored user template feature vector. When the similarity meets the identity confirmation threshold, the system automatically locks the user's identity and loads associated attributes, such as adult mode or child mode labels.

[0009] In non-contact heart rate monitoring, this invention employs a signal separation technique based on a chromaticity model. The system divides the cheek region from a facial sub-image as the region of interest, calculates the average grayscale value of this region in each of the RGB channels, and converts the RGB three-channel signal into two orthogonal components (e.g., specular reflection and diffuse reflection components) in the chromaticity space through linear projection. The standard deviation of the two orthogonal components is calculated using a sliding time window, and the projection coefficient is calculated based on the ratio of the standard deviations. Subsequently, motion artifacts and illumination noise are eliminated using weighted subtraction logic to construct a clean pulse wave signal. Finally, a fast Fourier transform is performed on the pulse wave signal to search for the frequency point with the highest power spectral density within the effective frequency range of human heart rate (e.g., 0.7Hz to 4Hz), thereby calculating the real-time heart rate value.

[0010] Regarding multidimensional health status assessment: To provide comprehensive health indicators, the system has constructed a multidimensional health status comprehensive index, which integrates three dimensions: sleep quality, body temperature deviation, and heart rate variability. Specifically, the gray-level co-occurrence matrix algorithm is used to extract texture features (such as contrast feature values) of the eye area to quantify sleep quality; facial temperature is inverted using infrared or multispectral data, and its absolute deviation from the baseline reference temperature is calculated; the standard deviation of the time interval sequence of adjacent heartbeat cycles is statistically analyzed to obtain the heart rate variability index; finally, a comprehensive index is generated through a weighted nonlinear combination to intuitively reflect the user's current physical function status.

[0011] Regarding behavioral posture monitoring and brushing guidance: The behavior and posture monitoring module uses computer vision technology to perform motion analysis. The system extracts a set of upper limb skeletal key points (including the tip of the nose, neck, wrist, and elbow) from the video stream. To eliminate the influence of the user's standing position, the system calculates the Euclidean distance from the neck key point to the nose key point as a normalized reference distance. When monitoring brushing behavior, the system calculates the normalized relative distance between the wrist key point and the nose key point. When this distance is less than a preset spatial judgment threshold, a position signal is triggered. Subsequently, the system establishes a sliding time window to analyze the wrist movement pattern, comprehensively considering the movement intensity index and the number of directional flips of the displacement vector. Only when the movement frequency is within a preset range and the number of reciprocations meets the threshold requirement is it determined to be a brushing state. The system is also configured with fault-tolerant timing logic, allowing brief interruptions during the brushing process (such as rinsing the mouth or switching hands). The completion command is triggered only when the cumulative time of the effective action reaches the standard duration.

[0012] Regarding contextual decision-making and content delivery: The contextual decision-making and content delivery module enables personalized information services. The system maps the current time, user attributes, and action status to contextual state variables. For adult users, during the morning hours, the system calculates ranking weights based on the timeliness of information release and the frequency of the user's historical interactions, prioritizing the delivery of social media updates and schedule information through the network communication unit. For children, during the morning hours, the system randomly selects educational resources based on a weighted knowledge point recommendation index. This recommendation index is negatively correlated with the user's historical mastery level, meaning that knowledge points with weak mastery are prioritized. After the morning hours have passed, the system automatically switches strategies, using a keyword blacklist to filter entertainment content and only selecting highly popular news summaries, thus adapting to the user's psychological expectations and attention characteristics at different times.

[0013] Regarding multimodal human-computer interaction: The human-computer interaction interface module enhances the anti-interference capability of voice interaction through visual assistance. After receiving multi-channel audio signals from the microphone array and performing adaptive echo cancellation, the system calls the user's facial center coordinates output by the vision module to calculate the direction of arrival of the sound source relative to the center of the microphone array. Using this direction as the guide vector, time delay compensation is applied to the multi-channel signals and weighted summation is performed to achieve beamforming, thereby extracting pure target speech in noisy environments. In addition, in terms of visual display, the system establishes a nonlinear mapping model between ambient light intensity and backlight driving signal, so that the target duty cycle of the display backlight increases logarithmically with the increase of ambient light, which conforms to the human eye's perception of brightness and ensures that the displayed content is clear and not dazzling under different lighting conditions.

[0014] This invention provides a multifunctional mirror that integrates health monitoring and intelligent interaction. It has the following beneficial effects: 1. This invention achieves highly interference-resistant non-contact physiological parameter monitoring. By employing a signal separation algorithm based on a chromaticity model, the facial RGB signal is projected into orthogonal chromaticity components. The differences in optical properties between different components are used to effectively eliminate non-physiological noise caused by changes in ambient light and specular reflection. Combined with automatic face region tracking and image pyramid technology, the system can still acquire high signal-to-noise ratio pulse wave signals and calculate accurate heart rate and heart rate variability in dynamic scenarios such as when the user is washing up or making slight facial movements, thus overcoming the shortcomings of traditional optical monitoring that are highly dependent on static states.

[0015] 2. This invention constructs an adaptive contextual interaction system based on identity attributes and behavioral states. The system automatically identifies the user's identity and loads adult or child modes through deep learning feature extraction technology. At the same time, it uses skeletal key point detection technology to determine the user's brushing teeth or washing face actions in real time. Based on this logic, the system can push efficient news and schedule summaries to adult users in the morning, and push fun educational content and brushing action guidance to children. This hierarchical and classified information push strategy solves the problem of homogeneous information display in traditional smart mirrors, improves the effectiveness of information transmission and the guiding effect on children's healthy habits.

[0016] 3. This invention enhances the audiovisual interaction experience in complex bathroom environments. In terms of voice, the system uses the user's angle obtained by visual positioning to assist the microphone array in beamforming, using the direction of the sound source as the guide vector to suppress water flow noise and exhaust fan noise, which greatly improves the recognition rate of voice commands in noisy environments. In terms of display, the system establishes a logarithmic backlight adjustment model that conforms to the physiological characteristics of the human eye, so that the screen brightness adapts to the ambient light intensity, ensuring that the displayed content is clear and readable without glare under dim or strong reflective conditions in the bathroom. Attached Figure Description

[0017] Figure 1 This is a perspective view of the present invention; Figure 2 This is a rear view of the present invention; Figure 3 This is a block diagram of the hardware connection module of the system of the present invention; Figure 4 This is the main control logic flowchart of the system of the present invention; Figure 5 This is a comparison chart of the anti-interference performance of the heart rate monitoring system of the present invention; Figure 6 This is a schematic diagram illustrating the speech recognition rate under different noise environments according to the present invention. Figure 7 This is a schematic diagram of the adaptive brightness response curve of the display screen of the present invention.

[0018] The system includes: 1. Mirror body; 2. Camera; 3. Support frame; 10. Data acquisition and preprocessing module; 20. Biometric recognition and analysis module; 30. Behavior and posture monitoring module; 40. Contextual decision-making and content push module; and 50. Human-computer interaction interface module. Detailed Implementation

[0019] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Please see the appendix Figure 1 - Appendix Figure 3 This invention provides a multifunctional mirror that integrates health monitoring and intelligent interaction. Its hardware structure mainly includes a mirror body 1, a camera 2, a control panel, and a support frame 3.

[0021] The support frame 3 forms the mechanical load-bearing foundation of the equipment. The mirror body 1 is fixed to the front side of the support frame 3 by mechanical connectors. The mirror body 1 has optical reflection function and photoelectric display function. It includes a semi-transparent and semi-reflective glass layer and a display module attached to the back of the semi-transparent and semi-reflective glass layer. When the display module is in the off state, the mirror body 1 presents a mirror reflection effect. When the display module is in the on state, the display light penetrates the semi-transparent and semi-reflective glass layer and forms a visual information interaction interface on the mirror surface.

[0022] To adapt to the high humidity characteristics of the bathroom environment, the back of the semi-transparent and semi-reflective glass layer of the mirror body 1 is also attached with an electrically heated defogging film or coated with a nano hydrophobic coating. The central processing unit is connected to a humidity sensor. When the relative humidity of the environment exceeds the threshold (such as 85%), the power supply of the defogging film is automatically turned on, and the Joule heating is used to raise the temperature of the mirror surface to prevent water vapor from condensing and blocking the camera's field of vision or affecting the display effect.

[0023] The camera 2 is located at the top center of the mirror body 1 or embedded in the upper edge area of ​​the mirror body 1. The light-sensing component of the camera 2 faces the user standing area directly in front of the mirror body 1 to collect video stream data containing the user's facial and body features. The camera 2 uses a high-resolution wide-angle lens and integrates an infrared imaging sensor or a multispectral sensor to support image capture and body temperature sensing under different lighting conditions.

[0024] The control panel is fixedly installed on the back of the mirror body 1 or on one side of the support frame 3. The control panel integrates a central processing unit, a power management circuit and a signal transmission interface. The central processing unit is electrically connected to the camera 2 and the display module through an internal bus. The control panel also integrates a microphone array and a speaker assembly. The microphone array is used to collect ambient voice signals and the speaker assembly is used to output voice broadcast audio.

[0025] The control panel also integrates a network communication unit, which connects to the Internet via a wireless LAN or cellular data network for data interaction with external servers, social media platforms, and news data sources. In addition, the network communication unit also integrates a short-range wireless communication module (including but not limited to Bluetooth 5.0, ZigBee, or NFC modules) to build a personal area network (PAN). This communication module actively searches for and pairs with smart bracelets, smartwatches, and smart body fat scales placed nearby, establishing a low-power data transmission channel to read step count data, blood oxygen saturation, body fat percentage, and long-term sleep monitoring data collected by external devices in real time.

[0026] To achieve health monitoring, behavior recognition, and intelligent push functions, the central processing unit operates multiple functional modules. The system control logic architecture includes: a data acquisition and preprocessing module 10, a biometric recognition and analysis module 20, a behavior and posture monitoring module 30, a contextual decision-making and content push module 40, and a human-computer interaction interface module 50.

[0027] The data acquisition and preprocessing module 10 is connected to the camera 2 and the microphone array signal to receive raw video and audio stream data in real time, and to perform noise reduction, frame synchronization and formatting processing on the raw data. At the same time, the data acquisition and preprocessing module 10 is also equipped with an IoT device interface layer to parse heterogeneous data packets from external smart health devices, and to align the accelerometer sensor data transmitted by the smart bracelet, the bioelectrical impedance data transmitted by the smart scale and the visual data collected by the camera with timestamps to form a multi-source heterogeneous health dataset.

[0028] The biometric identification and analysis module 20 is connected to the data acquisition and preprocessing module 10. It is used to locate the face region from the processed video data, extract the facial feature vector, and analyze the user's physiological parameters based on the image signal, including heart rate, body temperature trend and skin condition, and then generate health status assessment data.

[0029] The behavior and posture monitoring module 30 is used to track key points of the human skeleton within the field of view, calculate the spatial positional relationship and motion trajectory between key points, and determine whether the user is currently performing a specific action, including brushing teeth, washing face, and static observation.

[0030] The contextual decision-making and content push module 40 is the logical control center of the system. It is connected to the system clock and network communication unit. Based on the current time information, the user identity information output by the biometric recognition and analysis module 20, and the action status output by the behavior and posture monitoring module 30, the module retrieves relevant information content from local storage or the Internet. This information content includes weather data, memo data, social media updates, news, and educational materials.

[0031] The human-computer interaction interface module 50 is connected to the display module and the speaker assembly. The human-computer interaction interface module 50 receives instructions from the context decision-making and content push module 40, renders visual information onto the display area of ​​the mirror body 1, and converts audio information into voice signals for playback through the speaker assembly.

[0032] See attached document Figure 4 , Figure 4 This is a system overall workflow diagram according to an embodiment of the present invention. The present invention provides an intelligent interaction and health monitoring method based on a multifunctional mirror, including the following steps: S10, System Initialization and Image Acquisition: After the system is powered on, it enters the monitoring state. Camera 2 continuously acquires real-time video stream data of the area in front of the mirror body 1 and transmits the video stream data to the central processing unit in the control panel. The control panel simultaneously synchronizes the current network time, weather data and memo data through the network communication unit and establishes a basic information display layer on the display screen of the mirror body 1.

[0033] S20, User identification and attribute determination: The central processing unit analyzes the video stream data frame by frame, locates the face region and extracts facial feature data; compares the extracted facial feature data with the pre-stored user templates in the local database; if the match is successful, the current user's identity ID is determined, and the user attributes are determined according to the preset user profile, including adult mode and child mode; if the match fails, it is marked as visitor mode.

[0034] S30, Multi-dimensional Health Data Detection and Display: After locking onto the user's facial area, the central processing unit calculates the user's real-time heart rate, estimated body surface temperature, and skin condition parameters based on the spectral and texture features of the facial image. Simultaneously, the system polls the connected smart wearable device and body fat scale via Bluetooth to obtain the user's deep / light sleep duration distribution, current blood oxygen saturation, and weight and body fat data from the previous night. The central processing unit executes a data fusion algorithm, using historical heart rate data from the smart bracelet to perform baseline calibration on the visual heart rate detection results, and combines the visually analyzed skin condition with the sleep structure recorded by the wearable device to comprehensively generate a health status score and sleep quality assessment results. The fused health data is rendered in chart or numerical form on a designated area of ​​the display screen of the mirror body 1.

[0035] S40, Behavior and Posture Monitoring and Brushing Recognition: The central processing unit uses a human posture estimation algorithm to extract the coordinates of the key skeletal points of the user's upper limbs from the video stream, focusing on tracking the spatial positional relationship between the wrist node, elbow node, and mouth node; by calculating the Euclidean distance between the wrist node and the mouth node and the movement frequency of the wrist node, it determines whether the user is brushing their teeth; if the determination result is yes, it triggers the brushing countdown logic, displays a 3-minute countdown progress bar on the display screen, and emits a prompt sound through the speaker when the countdown ends.

[0036] S50, Contextualized Information Decision-Making and Content Push: The central processing unit schedules the displayed content and voice broadcast content according to the current time, user attributes and user behavior status, based on a preset priority strategy. When the current time is earlier than 08:00 AM and the user attribute is in adult mode, the control panel requests the bound social software API through the network interface to obtain the updates of the followed objects, controls the display screen to scroll to display the update summary and broadcast it by voice. When the current time is earlier than 08:00 AM and the user is in child mode, the control panel randomly retrieves English words, poems, or famous quotes from the local or cloud-based educational database, controls the display screen to show the text and images, and guides the user to read aloud. When the current time is later than or equal to 08:00 AM, and the behavior and posture monitoring determines that the user is in the process of washing up, the control panel blocks the social software interface and instead requests data from the news aggregation platform to obtain information on major domestic and international events, and controls the speaker to broadcast news summaries.

[0037] S60, Intelligent Voice Interaction: During the above steps, the microphone array continuously monitors the ambient audio; when a preset wake word or a voice command issued by the user is detected, the system suspends the current non-urgent voice broadcast task, calls the natural language processing model to parse the user's intent, and provides feedback on the dialogue content or performs memo query and input operations through speech synthesis.

[0038] Please see the appendix Figure 3 This invention provides an intelligent magic mirror system based on multimodal perception. The system mainly includes a data acquisition and preprocessing module 10, a user identity and attribute recognition and physiological analysis module 20 based on visual perception, a behavior and posture monitoring module 30, a contextual decision-making and content push module 40, and a human-computer interaction interface module 50. These modules work collaboratively, and the specific processing flow and operating mechanism are as follows: The data acquisition and preprocessing module 10 serves as the system's front-end sensing entry point, and its operation mechanism specifically includes the following steps: S101, Video Stream Parameter Configuration and Raw Data Acquisition: The central processing unit in the control panel sends an initialization command to camera 2, configuring its sampling resolution and frame rate parameters. To balance the accuracy requirements for extracting subtle changes in physiological signals with the load balance of real-time computing, the sampling resolution is set to 1920×1080 pixels to 3840×2160 pixels, and the frame rate is set to 25 frames / second to 60 frames / second. Considering the complex lighting conditions in the bathroom or bedroom where mirror 1 is located, the central processing unit activates the automatic white balance and automatic exposure control logic inside camera 2 to ensure the color accuracy of the acquired image. Camera 2 outputs raw video stream data. The data acquisition and preprocessing module 10 calls the image decoder to convert the raw data into an RGB color space image suitable for computer vision analysis, defining the first... The original image frames captured at each moment are ,in and These represent the horizontal and vertical coordinates of a pixel in the image plane, respectively. Indicates the timestamp of the sampling.

[0039] S102, Region of Interest (ROI) localization for faces: To reduce the computational load of subsequent algorithms and improve detection accuracy, the system uses full-frame original image frames. The system extracts regions of interest (ROIs) containing facial features. The data acquisition and preprocessing module 10 uses a feature-based face detection algorithm (such as MTCNN) to scan the image. The processing includes converting the image to grayscale, using a sliding window to traverse the image to detect potential facial feature regions, and removing redundant windows with overlap exceeding a preset threshold (e.g., IoU > 0.7) using a non-maximum suppression algorithm. After localization, the system obtains the bounding box coordinates of the facial regions and uses this information to extract features from the original image frame. The facial sub-image is cropped from the image and denoted as... For facial sub-images The acquisition of this information can effectively eliminate dynamic interference from the background environment and focus on the physiological characteristic areas of the human body.

[0040] S103, Adaptive Light Enhancement Processing: To address potential issues such as backlighting, sidelighting, or low illumination during user washing, the data acquisition and preprocessing module 10 compares cropped facial sub-images. Histogram equalization is performed, which involves statistically analyzing the frequency of each gray level in the facial sub-image and constructing a gray-level histogram. The cumulative distribution function is then used to calculate the mapping relationship between the original gray levels and the target gray levels, stretching the dense gray-level distribution in the original image to the entire gray-level range. This non-linear mapping enhances the local contrast of the image, ensuring that the color change features and texture details of the facial skin remain clear under different lighting conditions, and providing standardized input data for subsequent pulse wave extraction.

[0041] S104, Multi-scale image pyramid construction: Because the user's position in front of the mirror changes, the proportion of the face in the image is not fixed. To achieve scale-invariant feature extraction, the data acquisition and preprocessing module 10 constructs an image pyramid based on the enhanced facial image. The construction process of the image pyramid includes applying Gaussian kernel smoothing filtering to the bottom layer image to remove high-frequency noise, followed by downsampling (e.g., removing even-numbered rows and columns) to obtain the image of the next level. The above steps are repeated to generate an image sequence containing different resolution levels. The system selects the corresponding level image from the image pyramid as input according to the specific resolution requirements of each subsequent functional module: for the heart rate monitoring module that needs to capture subtle skin color changes, the system selects the bottom high-resolution image that retains complete color information; for the posture monitoring module that mainly identifies limb contours, the system selects the top layer image that has undergone downsampling to reduce computational resource consumption.

[0042] After acquiring the preprocessed facial sub-image, the vision-based biometric recognition and analysis module 20 first performs the following steps to determine the operator's specific identity and associated attribute permissions: S201, High-dimensional facial feature vector extraction: The biometric recognition and analysis module 20 calls a pre-built deep learning feature extraction model, which is built on a convolutional neural network architecture. The system then uses the facial sub-images output by the data acquisition and preprocessing module 10. The input is fed into the feature extraction model. The model extracts deep biological feature information from the image through multi-layer convolution operations, maps this biological feature information into a numerical representation in a high-dimensional feature space, and outputs a fixed-dimensional feature vector, denoted as . This feature vector It is a digital projection of a facial image in a feature space. Its numerical distribution is invariant to changes in illumination and slight pose deviations, and is used to uniquely represent the facial identity features of the current user.

[0043] S202, Feature Space Similarity Matching Calculation: The control panel maintains a registered user database in its local storage, which pre-stores... Let the template feature vectors of known users be denoted as set. The system executes the traversal matching logic, calculating the currently extracted feature vectors in sequence. With each template feature vector in the database The similarity value between them is denoted as In this embodiment, the similarity value is calculated using the cosine similarity algorithm. The specific calculation process is as follows: calculate the current feature vector. With the template feature vectors The dot product of the two vectors is calculated, and their respective Euclidean norms (i.e., magnitudes) are also calculated. The normalized similarity score is then obtained by dividing the dot product by the product of the magnitudes of the two vectors. The closer this value is to 1, the more consistent the directions of the two vectors are in the feature space. The system compares all calculated similarity values ​​and selects the maximum value. This serves as the final identity matching score.

[0044] S203, User Attribute Determination and Permission Association: The system uses an identity verification threshold. The matching results are binarized for verification; this identity confirmation threshold is then determined. The value is set to a range of 0.75 to 0.90, and this range is used to filter out false matches of unregistered users.

[0045] Before the system determines that the similarity meets the standard and locks the user's identity ID, it triggers a liveness detection verification logic to prevent spoofing attacks using high-resolution photos or videos. This logic includes: using the infrared sensor integrated into the camera to obtain facial depth information, calculating the depth gradient map of the three-dimensional facial structure, and determining whether planar features exist; at the same time, analyzing the micro-expression changes and the natural frequency of blinking in continuous video frames. Only when the depth gradient matches the three-dimensional features and non-periodic natural biological movements are detected is the user's identity finally confirmed as legitimate.

[0046] When the calculated maximum value At this time, the system determines that the current user is a registered user and locks the corresponding user identity identifier. The system uses the locked user identity identifier Index the associated user attribute fields in the registered user database. This user attribute field These are predefined category tags when a user registers for the first time, including an adult mode tag to identify adults and a child mode tag to identify minors.

[0047] When the calculated maximum value When the system determines that the current user is an unregistered visitor, it will enter the user attribute fields. Marked as general mode, it will not trigger subsequent personalized information push logic.

[0048] After confirming the user's identity, the biometric identification and analysis module 20 continues to quantitatively assess the user's physiological characteristics based on non-contact visual sensing technology, specifically including the following steps: S301, Dynamic segmentation of regions of interest for physiological signals: To separate signal sources with different physiological parameters, the system uses the facial images acquired in step S102. Using facial landmark localization technology, multiple independent functional regions are divided. The system selects the facial cheek region as the photoplethysmography (PPG) monitoring region, denoted as... This area has a dense distribution of capillaries and is less affected by facial muscle movements; the system selected the central area of ​​the forehead as the body temperature monitoring area, denoted as... The system selects the area below both eyelids as the region for extracting eye fatigue features, denoted as... The system performs spatial averaging on the pixel values ​​within each ROI region to calculate the average grayscale value of that region across all RGB channels, thereby suppressing random noise in individual pixels.

[0049] S302, Pulse wave extraction and heart rate calculation based on chromaticity model: The system utilizes the optical property that hemoglobin in blood has a high absorption rate for green light (approximately 500-600 nm) to monitor... To analyze the temporal variation of the average pixel value in the region and to separate the weak pulse signal from the raw signal containing ambient lighting variations and head motion noise, the system employs a signal separation algorithm based on chroma space. First, the system converts the RGB three-channel signal into two orthogonal components in the chroma space through linear projection (e.g., using a preset projection matrix or projecting the RGB channels onto the chroma plane), denoted as... and Next, pulse wave signals are constructed. The calculation formula is as follows: ; In the formula, Indicates the sampling time point. The projection coefficient is calculated based on the ratio of the standard deviations of the two chromaticity components within a time window. The calculation logic utilizes the difference in the distribution of specular reflection and diffuse reflection components in the color space to eliminate non-physiological optical interference through weighted subtraction.

[0050] Before performing the Fast Fourier Transform, the system analyzes the pulse wave signal. The system performs signal quality assessment by calculating the skewness and kurtosis of the signal within the current time window and detecting whether there is oversaturation or zero-value interruption in the signal. If the calculated signal quality score is lower than the preset reliability threshold (indicating that there may be strenuous exercise or strong light interference), the system will automatically discard the data in the current time window, keep the heart rate value displayed at the previous moment, and prompt the user to remain still through the UI interface, thereby avoiding outputting incorrect physiological parameters to mislead the user.

[0051] Obtain the denoised pulse wave signal The system then performs a Fast Fourier Transform (FFT) on the signal, converting the time-domain signal into a frequency-domain power spectrum. The system searches for the frequency point with the highest power spectral density within a preset effective frequency range for human heart rate (e.g., 0.7Hz to 4.0Hz), denoted as... Then calculate the real-time heart rate. : ; In addition, the system extracts the peak position of the pulse wave signal, calculates the time interval sequence of adjacent heartbeat cycles, and calculates the standard deviation of the sequence to obtain the heart rate variability index. .

[0052] S303, Multimodal Body Temperature Estimation and Calibration: For body temperature monitoring, the system prioritizes using the infrared thermal imaging sensor integrated into the camera module to read... The system collects regional thermal radiation data and converts it into estimated core body temperature values. As an alternative or supplementary solution, with only a visible light camera, the system employs multispectral skin color analysis. By analyzing changes in facial skin rosiness at specific wavelengths and combining this with an environmental temperature compensation model, the system indirectly inversely calculates the trend of body surface temperature changes. The system sets a baseline reference temperature. (e.g., 36.5 degrees Celsius), calculate the current measured temperature. Absolute deviation from the reference temperature : ; in, This serves as basic physiological data used to display the user's real-time body temperature readings on the interactive interface; This status determination data is used for internal system logic analysis, and the system bases on... The numerical value rather than The absolute value is used to determine the level of abnormality of body temperature (such as normal, low fever risk, high fever alarm).

[0053] S304, Construction of Sleep Quality and Comprehensive Health Index: The system analyzes the eye area The system uses image texture features to quantify sleep quality. Specifically, it extracts contrast feature values ​​of the eye region using a gray-level co-occurrence matrix algorithm. This feature value can keenly capture local texture gradient changes caused by puffy eyes or dark circles; the higher the feature value, the more obvious the eye fatigue characteristics, and the lower the corresponding sleep quality assessment.

[0054] To provide users with intuitive quantitative feedback, the system constructs a multi-dimensional comprehensive health status index. This index is a weighted nonlinear combination of sleep characteristics, body temperature status, and cardiac function indicators. Its mathematical model is as follows: ; In the formula, This is a preset extreme value for normalized eye texture contrast, used to map texture features to the [0,1] interval; This is a constant representing the permissible normal range of body temperature fluctuations (e.g., 1.5 degrees Celsius). This serves as the standard reference threshold for heart rate variability. This is the slope parameter of the Sigmoid function, which uses the Sigmoid function to map heart rate variability indicators to a health score. The higher the value, the closer the score is to 1; , , The weighting coefficients for sleep, body temperature, and heart rate are respectively, and satisfy the following conditions: The system is based on The system matches the numerical range of the data with corresponding health advice and displays it on the screen. When an external smart device is detected, the system automatically expands the calculation dimension of Hindex by introducing external variables w4·BMIindex (the normalized body mass index value provided by the smart scale) and w5·SpO2index (the blood oxygen health index provided by the smartwatch). At this time, the system dynamically adjusts the weight allocation strategy, reducing the weight of a single visual sensor and using contact measurement data from external devices to weight the non-contact visual estimation data with confidence. For example, when there is a significant discrepancy between the visually estimated sleep quality and the sleep data recorded by the smart band, the sleep stage data recorded by the smart band is used first to correct the calculation result of Hindex, thereby providing a more objective comprehensive health assessment.

[0055] The behavior posture monitoring module 30 analyzes the spatial position and movement timing characteristics of key points in the user's upper limb bones to determine and automatically time the brushing action. The specific execution steps are as follows: S401, Key Point Construction and Coordinate Tracking of Human Upper Limb Skeleton: The system uses a pre-built deep learning pose estimation model based on real-time video streams to extract skeletal key points from human targets within the field of view. The system identifies a set of upper limb key points strongly correlated with the brushing motion and defines this key point set. Including key points of the tip of the nose Neck key points Key points of the right wrist Key points of the right elbow Key points of the left wrist and key points of the left elbow The system obtains the two-dimensional pixel coordinates of each of the above key points in the image plane; to solve the problem of image scale differences caused by the varying distances of the user's standing position from the mirror, the system calculates the neck key points. Key point to the tip of the nose The Euclidean distance between them is defined as the normalized reference distance. The normalized reference distance It represents the projection length of the physical scale of the current user's face in the image, serving as a dynamic benchmark for subsequent motion amplitude determination.

[0056] S402, Initial motion screening based on spatial geometric constraints: The system monitors the user's wrist position in real time to ensure it is within the effective operating area. The system calculates key points on the left wrist separately. and key points of the right wrist Relative to the key point of the tip of the nose To eliminate the influence of distance factors, the system divides the above Euclidean distance by the normalized reference distance. The normalized relative distance is obtained and denoted as . The system sets a spatial judgment threshold. The threshold value is set between 1.2 and 1.8, meaning the distance between the wrist and the mouth should not exceed a multiple of the distance from the center of the face to the neck. If the relative distance between any two wrists... Less than the spatial determination threshold The system determines that the user's hand has entered the brushing area near the face, generates a position trigger signal, and records it as... Otherwise, place .

[0057] S403, Frequency Domain Analysis and Confirmation Based on Time Domain Motion Characteristics: Trigger signal at the specified position Under the premise that the system establishes a length of Frames (e.g.) The system uses a sliding time window to analyze the motion patterns of wrist key points entering the operation area. It calculates the average displacement vector magnitude of the wrist key point between adjacent frames within the sliding time window, defining it as the motion intensity index. Simultaneously, the system counts the number of times the displacement vector direction flips within the sliding time window. If the motion intensity index... If the frequency is within a preset range and the number of reversals in the displacement vector direction exceeds a preset reciprocating threshold, it indicates that the user is performing high-frequency, small-amplitude reciprocating motion, and the system determines the current behavioral state. This indicates the state of brushing teeth. This decision logic can effectively distinguish the action of brushing teeth from the action characteristics of washing the face (usually a large-span unidirectional movement) or grooming (usually a low-frequency movement).

[0058] While monitoring user posture, the system runs a gesture recognition algorithm in parallel to control information display without touching the mirror. The system defines specific hand spatial trajectories as control commands: when the system detects a user's palm making a horizontal waving motion in the air, it generates a page-turning command to switch the displayed content (such as switching news items); when the system detects that the user's palm is facing the camera and remains still for more than a preset time (such as 1.5 seconds), it generates a pause / play command to control the playback status of multimedia content. This contactless interaction effectively solves the problem of not being able to operate the device with wet hands.

[0059] S404, countdown logic control and status feedback: The system is configured with a standard brushing time threshold. (For example, 180 seconds). When the behavior state When the system first confirms that the user is brushing their teeth, it starts a timer to accumulate the effective brushing time. .

[0060] The system incorporates hysteresis comparison logic to handle brief interruptions in the action (such as rinsing the mouth or switching hands). If a behavioral state is detected... When the system switches to a non-brushing state, it pauses the accumulation of effective brushing time. And start the pause wait timer. .

[0061] If pause the wait timer The value is less than the preset fault tolerance threshold. (e.g., 10 seconds), and behavioral state The system will resume accumulating effective brushing time once the brushing mode is restored. And reset the pause wait timer. ; If pause the wait timer The value exceeds the fault tolerance threshold The system determines that the current brushing process is complete and records the effective brushing time. Reset to zero.

[0062] During the timing process, the control panel drives the display to render a progress bar animation in real time; when the effective brushing time is reached... Reaching the standard brushing time threshold When the system completes the task, it will output a prompt tone through the speaker and display a completion indicator on the screen.

[0063] The contextual decision-making and content delivery module 40, as the system's strategy scheduling unit, is responsible for dynamically adjusting the human-computer interaction strategy based on multi-dimensional input variables. Its specific execution steps are as follows: S501, Aggregation of Multidimensional Contextual Feature Data and Time Period Determination: The system acquires the current time data, user attribute fields determined by the biometric recognition and analysis module 20, and behavioral status output by the behavior posture monitoring module 30 in real time. The system maps the data of the above three dimensions of time, user attributes, and behavioral status into a set of contextual state variables. The system presets a morning time period judgment threshold, which can be configured to 08:00 every day. The system compares the current time data with the morning time period judgment threshold; if the current time is earlier than the threshold, the system sets the time identifier to morning mode; if the current time is later than or equal to the threshold, the system sets the time identifier to normal mode.

[0064] S502, social and efficiency information aggregation based on adult mode: When the user attribute field is identified as adult mode and the time is identified as morning mode, the system executes the adult morning information feed push strategy. The control panel calls the pre-built application programming interface through the network communication unit to initiate data requests to the bound social media platform and calendar management server, and parses the fields of the returned standard format data.

[0065] To present key content within the limited time available for washing up, the system establishes an information priority scoring model. This model calculates ranking weights based on the timeliness and engagement of the information: for timeliness, the system calculates the difference between the information's publication timestamp and the current time, setting the weight to be inversely proportional to this time difference—the more recent the publication time, the higher the weight; for engagement, the system tracks the frequency of user interactions with the information source historically, setting the weight to be directly proportional to the interaction frequency. The system sorts the calculated weighted scores in descending order and selects a preset number of items (e.g., the top 5) for scrolling text display in the sidebar area of ​​the screen.

[0066] S503, Adaptation and guidance of educational content based on children's mode: When the user attribute field is set to Child Mode and the time field is set to Morning Mode, the system activates the educational guidance logic. The system connects to a local storage or cloud-based educational resource database, which stores English vocabulary, classical Chinese poems, and encyclopedic knowledge data categorized by difficulty level.

[0067] The system uses a weighted random sampling algorithm based on the forgetting curve to select content for the day's push notifications. The system reads the historical learning records for each knowledge point and calculates a recommendation index. This recommendation index is calculated based on the following conditions: the recommendation index is negatively correlated with the user's historical mastery rating of the knowledge point (the lower the mastery, the higher the recommendation index); and positively correlated with the time interval since the knowledge point was last displayed (the longer the time since the last review, the higher the recommendation index). The system performs weighted random sampling based on the recommendation index of each knowledge point, rendering the selected content as a graphic card in the center of the display screen, and controlling the speaker to play a standard pronunciation audio guide. If the behavior and posture monitoring module 30 determines that the child is brushing their teeth, the system switches the static graphic content to a brushing instructional video with encouraging animation.

[0068] S504, news aggregation and entertainment blocking outside of morning hours: When the time stamp is in normal mode, the system executes a wide-area information aggregation strategy, obtaining real-time information through RSS feeds or news aggregation interfaces. To filter invalid information, the system matches a preset keyword blacklist, removing content items tagged with entertainment, gossip, etc., and retaining only items belonging to the politics, finance, or technology categories.

[0069] The system calculates a heat weighting value based on the social attention of news items. This heat weighting value is the weighted sum of click data and share data. The system only selects news summaries with heat weighting values ​​exceeding a preset heat threshold for push. At this time, if the system detects that the user is washing up, the system automatically switches the text display mode on the screen to a voice synthesis broadcast mode so that the user can still obtain information when their eyes leave the mirror area.

[0070] S505, Arbitration between display layer rendering and human-computer interaction: The context decision-making and content push module 40 generates rendering instructions based on the content determined in the above steps. The system uses layered rendering technology to manage the display output. The bottom layer renders the background and basic time and weather controls, the middle layer renders the context-pushed content (social information, educational cards, or news summaries) determined in the above steps, and the top layer renders the system status indicator controls (including a brushing countdown progress bar and a device battery icon).

[0071] When receiving high-priority emergency information (such as weather disaster warnings) or detecting user voice interaction commands through the microphone array, the system executes interrupt arbitration logic, immediately suspends the current middle-layer information push task, and prioritizes rendering emergency prompts or voice interaction feedback content in the top-layer area to ensure that key information can be perceived by users in real time.

[0072] The human-computer interaction interface module 50 serves as the system's human-computer interaction interface, responsible for implementing contactless voice command response and dynamic rendering of mirror display content. Its specific operating mechanism includes the following steps: S601, Vision-Assisted Microphone Array Beamforming: In order to extract clear user voice commands in bathroom environments with water flow noise or exhaust fan noise, the system is equipped with a linear microphone array consisting of multiple omnidirectional microphones. The system receives the multi-channel raw audio signals collected by the microphone array and performs adaptive echo cancellation logic to filter out audio echoes generated by the device's own speakers.

[0073] The system calls the user's facial center coordinate information obtained in the preceding step S102, calculates the sound source arrival direction of the user relative to the center of the microphone array, and uses the sound source arrival direction as the guide vector to perform weighted processing on the multi-channel signals using delay summation beamforming technology. Specifically, the system applies time delay compensation to the microphone signals that are far from the sound source, so that the target speech signals of each channel are aligned in the time domain and added together to enhance them. At the same time, the noise signals in non-target directions are mutually attenuated due to phase inconsistency, thereby outputting a single-channel target speech signal with a high signal-to-noise ratio.

[0074] S602, Local wake word detection and cloud-based semantic recognition: The system performs speech endpoint detection on the enhanced single-channel target speech signal, removes silent segments by judging short-time energy and zero-crossing rate, extracts effective speech frames, extracts acoustic features (such as Mel frequency cepstral coefficients) of the effective speech frames, and inputs them into the local low-power wake-up engine, which compares the similarity between the acoustic features and the preset wake-up words in real time.

[0075] Once the wake-up engine confirms the detection of the wake-up word, the system activates the speech recognition engine. To balance response speed and recognition accuracy, the system adopts a hierarchical recognition strategy: for preset short control commands (such as start brushing teeth), the system performs offline matching and recognition locally; for open-domain natural language queries, the system encrypts and uploads the speech data to the cloud server for recognition via the network communication unit and receives the returned text data.

[0076] S603, Semantic Intent Understanding and Command Distribution: The system uses a natural language processing module to parse the identified text data. First, the system categorizes the text by domain, classifying user intent into device control, lifestyle services, or health consultation. Then, the system performs key information slot extraction, extracting actions, objects, and time parameters from the text.

[0077] After parsing, the system generates standardized control commands. If the command involves hardware layer operations (such as turning on the lights), the system directly sends control signals to the underlying driver. If the command involves information acquisition, the system sends a data request to the context decision-making and content push module 40. The system maintains a context state machine to cache the intent state of the previous round of dialogue in order to support multi-round continuous interaction.

[0078] S604, an adaptive display driver based on ambient light perception: The display driver interface module is electrically connected to the display controller behind the mirror via a MIPI or LVDS interface. The system writes the user interface data into the video memory for refresh display based on the rendering instructions generated by the context decision and content push module 40.

[0079] To ensure that the displayed content remains clearly visible and glare-free under different ambient lighting conditions, taking advantage of the semi-transparent and semi-reflective optical characteristics of the mirror body 1, the system periodically reads the ambient light intensity values ​​collected by the ambient light sensor. The system establishes a nonlinear mapping model between ambient light intensity and backlight driving signal. This model, based on the Weber-Fechner law, is used to calculate the target duty cycle of the display backlight. The calculation formula is as follows: ; In the formula, This is the minimum maintenance duty cycle for the display screen in a completely dark environment, used to ensure basic visibility; This is the luminance response coefficient, used to adjust the sensitivity to changes in luminance. This is the normalized reference light intensity constant; This formula represents the natural logarithm, which causes the backlight brightness to increase logarithmically with increasing ambient light, consistent with the physiological characteristics of human eye perception of brightness. The system then uses this calculated value... A pulse width modulation signal is generated to dynamically adjust the drive current of the display screen backlight circuit.

[0080] In addition, the system also introduces a color temperature compensation mechanism based on the human body's circadian rhythm. The system dynamically adjusts the gain weight of the RGB sub-pixels of the display screen according to the current time information: in the morning (06:00-10:00), the system enhances the blue spectral component, making the display color temperature tend towards cool white light (about 6000K), suppressing melatonin secretion and playing a waking-up role; in the evening (after 20:00), the system attenuates the blue spectral component and enhances the red spectral component, making the display color temperature tend towards warm yellow light (about 3000K), reducing the interference of blue light on the user's sleep preparation period.

[0081] Specific application examples: To more intuitively illustrate the application process of the technical solution of this invention in a real-world scenario, the following description uses a typical scenario of morning washing at home: Scene background: User A (father, 35 years old, registered user, set to adult mode) and User B (son, 8 years old, registered user, set to child mode) used the smart mirror with this system installed one after the other.

[0082] Process 1: Efficient Morning Interaction in Adult Mode Seamless wake-up and recognition: At 7:30 AM, User A walks in front of the mirror. The system detects the approaching human body using an infrared sensor and illuminates the screen background. The data acquisition and preprocessing module 10 acquires a facial image, which is then analyzed by the biometric recognition and analysis module 20 to extract feature vectors. A successful match was found with the database; similarity score. User A has been identified and adult mode has been loaded.

[0083] Health check-up: During user A's shaving process, the system automatically locked onto the ROI area on their face.

[0084] The biometric recognition and analysis module 20 calculated a real-time heart rate of 72 beats per minute by analyzing subtle changes in facial color, and measured a body temperature of 36.6℃ using infrared thermal imaging. Simultaneously, eye texture analysis revealed… Low overall health index The score was 85 (good), and a green health icon was displayed in the upper right corner of the screen. At the same time, User A stood barefoot on the smart body fat scale that was linked to the mirror. The lower right corner of the mirror immediately displayed the current weight (75.2kg) and body fat percentage (21.5%), and plotted a weight change curve for the past seven days based on historical records. In addition, the system read the data from the smart bracelet worn by User A and prompted that the deep sleep time last night was insufficient, suggesting that he go to bed earlier tonight, achieving a comprehensive health snapshot from facial complexion to body composition.

[0085] Intelligent information push: The context-based decision-making and content delivery module 40 determines that it is currently morning. Today's schedule automatically appears on the left side of the screen: 09:00 department meeting, 14:00 client visit. At the same time, based on user preferences, two breaking technology news items are scrolled and displayed.

[0086] Noise-resistant voice interaction: At this moment, the sink faucet is running water, and the ambient noise level is approximately 65dB. User A issues a voice command: "What's the weather like today?" The human-computer interaction interface module 50, combined with the face positioning results from the S102, uses beamforming technology to point towards the user's mouth, suppressing water flow noise and accurately recognizing the command. A weather card for today immediately pops up in the center of the screen: Sunny turning cloudy, 22℃.

[0087] Process Two: Behavioral Guidance in Child Mode Mode switching: User A leaves, and at 7:45 AM, User B walks to the mirror. The system re-detects and identifies the user, successfully confirming the user as User B, and automatically switches to child mode.

[0088] Brushing guidance: User B picks up a toothbrush, and the behavior and posture monitoring module 30 detects key points on the wrist. Enter the mouth area ( The system detects reciprocating motion and determines that brushing has begun. The static cartoon wallpaper originally displayed in the center of the screen is instantly switched to a brushing tutorial animation, accompanied by a progress bar countdown (target 180 seconds).

[0089] Error correction and incentives: When brushing teeth for 45 seconds, user B stopped brushing and seemed to be spacing out. The behavior and posture monitoring module 30 detected this. Switch to non-brushing mode and pause the timer. The system then plays a voice prompt: "Keep it up, make sure you brush your teeth clean!" The timer resumes after User B resumes brushing.

[0090] Feedback completed: When the countdown ends, the system emits a cheerful completion sound effect and displays a brushing star badge on the screen. At the same time, the brushing time data is synchronized to user A's mobile app for parents to view.

[0091] Experimental verification and effect comparison: To verify the effectiveness of the multimodal perception and adaptive driving algorithm proposed in this invention, the applicant built a test platform in a standard semi-anechoic chamber and a simulated bathroom environment, and conducted comparative experiments with existing technical solutions.

[0092] Validation of the accuracy of heart rate monitoring under complex lighting and motion interference: Experimental conditions: Light source environment: Simulates the mixed lighting commonly found in bathrooms (cool LED light from the top + warm light from the side), with light intensity dynamically varying between 200 Lux and 800 Lux.

[0093] Test subjects: 10 volunteers, who performed normal head rotation (angle of deflection) during the test. (and changes in facial expressions)

[0094] Comparison benchmark: Medical-grade fingertip pulse oximeter was used as the true value.

[0095] Comparison of options: Solution A (Invention): Executed by the biometric recognition and analysis module 20, employing S102's stable face ROI tracking + S302's colorimetric model ( (Signal separation algorithm).

[0096] Option B (existing technology): adopts the traditional green light averaging method, without introducing color space separation and dynamic ROI correction.

[0097] Analysis of experimental results: See attached document Figure 5 (Comparison chart of anti-interference performance of heart rate monitoring) In this chart, the horizontal axis represents the continuous test time of the experiment, in seconds; the vertical axis represents the number of heartbeats detected per minute, in beats / minute.

[0098] The solid black line in the figure represents the medical-grade true value, that is, the actual heart rate measured using a finger-clip pulse oximeter, which serves as a standard baseline for comparison.

[0099] The gray shaded areas in the image mark the time period during which artificial head movements and changes in lighting were introduced, from the 20th to the 40th second.

[0100] The dotted line marked with a triangle in the figure represents the existing technical solution (Solution B). It can be seen that under the interference of the gray shaded area, the line fluctuates violently and deviates from the true value, with the maximum absolute error exceeding 15 BPM, indicating that it has poor stability under motion interference.

[0101] The dotted line marked with dots in the figure represents the solution of the present invention (Solution A). Thanks to the suppression of non-physiological optical interference by the chromaticity model, even in the gray shadow interference area in the middle, the curve still closely follows the solid line (true value), and the average absolute error is controlled within 3 BPM. The experiment proves the robustness of step S302 of the present invention in non-stationary environments.

[0102] Speech recognition rate test in noisy environments: Experimental conditions: Noise sources: Play recorded water flow sounds, hair dryer sounds, and exhaust fan sounds to create a gradient test environment with a signal-to-noise ratio from 10dB to -10dB.

[0103] Test content: Includes 500 standard commands such as equipment control and weather query.

[0104] Comparison of options: Solution C (the present invention): Step S601 is executed by the human-computer interaction interface module 50, which is based on the microphone array beamforming assisted by visual positioning.

[0105] Option D (Prior Art): Traditional fixed beamforming or single microphone system.

[0106] Analysis of experimental results: See attached document Figure 6 (Speech recognition rate graph under different noise environments).

[0107] The horizontal axis in the graph represents the environmental signal-to-noise ratio, measured in decibels. The axis from left to right indicates that the environment becomes increasingly quiet (positive value) as it progresses from very noisy (negative value, such as -10 decibels, which means the noise is much louder than human voices).

[0108] The vertical axis in the graph represents the success rate of the system in understanding commands (speech recognition accuracy), expressed as a percentage.

[0109] The dashed line marked with a square in the figure represents the existing technical solution (Solution D). Its curve drops sharply as the signal-to-noise ratio decreases. In noisy environments on the left (such as when a hair dryer is turned on at -5dB), the recognition rate is less than 60%, indicating that the system has difficulty recognizing commands once there is noise such as water flow.

[0110] The solid line marked with dots in the figure represents the present invention (Solution C). This curve is always above the prior art, especially in noisy environments on the left (such as -10 dB), the recognition rate is still above 80%. Experiments have shown that by using the precise sound source angle provided by visual positioning, the present invention achieves narrower beam directivity and significantly improves noise resistance.

[0111] Ambient light adaptive display comfort evaluation: Experimental conditions: Controlling ambient light intensity Gradually increase from 0 Lux to 1000 Lux.

[0112] Measuring the duty cycle of the display screen backlight The output response.

[0113] Comparison of options: Scheme E (the present invention): Step S604 is executed by the human-computer interaction interface module 50, based on the logarithmic mapping model of Weber-Fechner's law.

[0114] Option F (existing technology): Traditional linear light sensor adjustment.

[0115] Analysis of experimental results: See attached document Figure 7 (Adaptive brightness response curve of the display screen).

[0116] The horizontal axis in the diagram represents the ambient light intensity around the mirror, measured in lux, from left (dark) to right (bright).

[0117] The vertical axis in the graph represents the intensity of screen illumination (backlight drive brightness), and the unit is percentage.

[0118] The black dashed line in the diagram represents the existing technical solution (Solution F), which shows a linear growth trend. This may result in excessive brightness and glare in the dark areas on the left, while the contrast may not be sufficiently improved in the bright areas on the right.

[0119] The solid black line in the figure represents the present invention (Solution E), and shows the logarithmic adjustment curve characteristics in step S604: In the low-light zone on the left (0-200 lux), the curve rises gently, which means that in a dark bathroom at night, the screen will not suddenly become very bright, effectively protecting the eyes (anti-glare comfort zone).

[0120] In the right-hand high-light area (>200 lux), the curve rises rapidly, meaning that when the ambient light is very strong, the system can quickly increase the brightness to counteract reflections and ensure that the user can see the content clearly (high light counteracts visible area); according to subjective rating, the visual comfort score of the present invention is 40% higher than that of the traditional solution.

Claims

1. A multi-functional mirror integrating health monitoring and intelligent interaction, characterized in that, Comprise: Support frame (3), mirror body (1) containing semi -transparent semi -reflection glass layer and display screen module, camera (2) integrated with infrared or multispectral sensor and control panel; The control panel is integrated with intelligent control system, ambient light sensor, microphone array, loudspeaker assembly and network communication unit; The intelligent control system runs: Data acquisition and pretreatment module (10) for receiving audio and video data and obtaining face sub-image; Biological feature recognition and analysis module (20) for locating face, extracting feature vector to identify identity and analyzing physiological parameters; Behavior and posture monitoring module (30) for tracking key points of skeleton and calculating their spatial relationship and trajectory to determine action state; Context decision and content pushing module (40) for scheduling content through network according to time, identity and action; Man-machine interactive interface module (50) for driving the display screen module and loudspeaker assembly to interact.

2. The multi-functional mirror with integrated health monitoring and smart interaction of claim 1, wherein, The data acquisition and pretreatment module (10) is used to perform the following operations: Set the sampling resolution and frame rate parameters of the camera (2), and activate the automatic white balance and automatic exposure control logic; The original image frame of the video stream data in the audio and video data is scanned by using the face detection algorithm based on feature classification, the face region bounding box is located by sliding window traversal and non-maximum suppression algorithm, and the face sub-image is obtained by cutting; The histogram equalization processing is performed on the face sub-image, the mapping relationship from the original gray level to the target gray level is calculated by using the cumulative distribution function, and the image pyramid sequence is constructed.

3. The multi-functional mirror with integrated health monitoring and smart interaction of claim 1, wherein, The biological feature recognition and analysis module (20) is used to perform the following operations: Obtain the face sub-image output by the data acquisition and pretreatment module (10), call the deep learning feature extraction model, and map the face sub-image to a high-dimensional feature vector in the feature space; Calculate the cosine similarity value between the high-dimensional feature vector and the pre-stored user template feature vector; When the maximum similarity value is greater than or equal to the identity confirmation threshold, lock the user identity and index the associated user attribute field, and the user attribute field includes adult mode label and child mode label.

4. The multi-functional mirror with integrated health monitoring and smart interaction of claim 1, wherein, The biological feature recognition and analysis module (20) is also used to calculate the heart rate value based on the chrominance model: Divide the face cheek area from the face sub-image output by the data acquisition and pretreatment module (10) as a photoplethysmography monitoring area, and calculate the average gray value of the photoplethysmography monitoring area in each channel of RGB; Convert the RGB three-channel signal to two orthogonal components in the chrominance space by linear projection; Calculate the projection coefficient according to the ratio of the standard deviations of the two orthogonal components in the preset time window, and construct the pulse wave signal by using the weighted subtraction logic; Perform fast Fourier transform on the pulse wave signal, search for the frequency point with the maximum power spectral density in the human heart rate effective frequency interval to calculate the real-time heart rate.

5. The multi-functional mirror with integrated health monitoring and smart interaction of claim 1, wherein, The behavior and posture monitoring module (30) is used to perform the following operations: extracting, from the audio and video data, a set of upper limb key points including a nose tip key point, a neck key point, a wrist key point, and an elbow key point; calculating a Euclidean distance from the neck key point to the nose tip key point as a normalized reference distance; calculating a normalized relative distance of the wrist key point relative to the nose tip key point, and generating a position trigger signal when the normalized relative distance is less than a spatial determination threshold; after the position trigger signal is generated, establishing a sliding time window to analyze a motion pattern of the wrist key point, and determining that a current behavior state is a tooth brushing state when a motion intensity index is in a preset frequency interval and a direction of a displacement vector is flipped more than a reciprocation threshold.

6. The multi-functional mirror with integrated health monitoring and smart interaction of claim 5, wherein, The behavior posture monitoring module (30) is further used to: when it is determined for the first time that the current behavior state is the tooth brushing state, starting a timer to accumulate an effective tooth brushing time; if it is monitored that the current behavior state becomes a non-tooth brushing state, starting a pause waiting timer; if a value of the pause waiting timer is less than a fault tolerance threshold and the current behavior state returns to the tooth brushing state, continuing to accumulate the effective tooth brushing time; when the effective tooth brushing time reaches a standard tooth brushing time threshold, triggering a completion instruction.

7. The multi-functional mirror with integrated health monitoring and smart interaction of claim 3, wherein, The context decision and content pushing module (40) is used to perform the following operations: mapping current time information, user attributes recognized by the biological feature recognition and analysis module (20), and action states determined by the behavior posture monitoring module (30) into context state variables; when the user attributes are adult mode and the current time is earlier than a morning period determination threshold, selecting social media dynamics and schedule information for pushing based on sorting weights of information publishing time and user interaction frequency through the network communication unit; when the user attributes are child mode and the current time is earlier than the morning period determination threshold, randomly selecting educational resource content for pushing according to a knowledge point recommendation index, the knowledge point recommendation index being negatively correlated with a historical mastery degree score; when the current time is later than or equal to the morning period determination threshold, filtering entertainment content by matching a keyword blacklist, and only selecting news abstracts with a heat weight value exceeding a heat threshold for pushing.

8. The multifunctional mirror with integrated health monitoring and smart interaction of claim 1, wherein, The human-computer interaction interface module (50) is used to perform microphone array beamforming based on visual assistance: receiving multi-channel original audio signals collected by the microphone array, and performing adaptive echo cancellation; calling user face center coordinate information output by the biological feature recognition and analysis module (20), and calculating a sound source arrival direction of the user relative to a center of the microphone array; taking the sound source arrival direction as a guide vector, applying time delay compensation to the multi-channel signals and performing weighted summation, and outputting a single-channel target speech signal.

9. The multifunctional mirror with integrated health monitoring and smart interaction of claim 1, wherein, The human-computer interaction interface module (50) is further used to perform display adaptive driving: periodically reading an ambient light intensity value collected by the ambient light sensor integrated in the control panel; establishing a nonlinear mapping model of the ambient light intensity and a backlight driving signal, and calculating a target duty cycle of the display screen backlight; the nonlinear mapping model is used to make the target duty cycle logarithmically increase with the increase of the ambient light intensity value.

10. The multi-functional mirror with integrated health monitoring and smart interaction of claim 4, wherein, The biometric identification and analysis module (20) is also used to construct a multidimensional health status comprehensive index: An eye region is segmented from the facial sub-image, and the contrast feature value of the eye region is extracted using the gray-level co-occurrence matrix algorithm to quantify sleep quality; Calculate the absolute deviation between the temperature value read by the infrared imaging sensor integrated in the camera or the temperature value retrieved by multispectral skin color analysis and the reference temperature. The standard deviation of the time interval sequence between adjacent heartbeat cycles is used to obtain the heart rate variability index. The multidimensional health status comprehensive index is generated by weighted nonlinear combination of sleep quality, absolute temperature deviation, and heart rate variability indicators.