Method and device for determining target state parameter, storage medium and electronic device
By combining a dynamic correlation model of image acquisition and supplementary lighting status signals, the accuracy problem of non-contact physiological parameter monitoring systems in low-light environments was solved, enabling accurate determination and stable monitoring of the status parameters of the target object.
Patent Information
- Application Number
- CN202610413491.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-30
- Publication Date
- 2026-07-24
- Estimated Expiration
- 2046-03-30
AI Technical Summary
Existing non-contact physiological parameter monitoring systems cannot accurately determine the state parameters of the target object in low-light environments, resulting in unreliable monitoring results.
By acquiring the target facial video stream output by the image acquisition module and the supplementary lighting status signal of the supplementary lighting module, the supplementary lighting conditions of each frame of the image are determined, and pixel-by-pixel correction is performed based on the reference light intensity offset to eliminate illumination interference. A dynamic correlation model is established to distinguish between physiological signals and non-physiological interference, thereby achieving accurate determination of state parameters.
In low-light environments, it significantly improves the accuracy and stability of physiological parameter calculations, reduces errors, and enables precise monitoring of target subjects' heart rate, blood oxygen, blood pressure trends, and mood.
Smart Images

Figure CN121937707B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart home technology, and more specifically, to a method and apparatus for determining target state parameters, a storage medium, and an electronic device. Background Technology
[0002] With the popularization of home-based elderly care, non-contact physiological parameter monitoring technology has gradually become an important means of health monitoring for the elderly. Currently, non-contact physiological parameter monitoring systems generally use image acquisition modules to obtain facial video streams of the target subject and extract physiological parameters such as heart rate and blood oxygenation from the weak light intensity fluctuations in the skin area using remote photoplethysmography (rPPG). These systems rely on the analysis of changes in reflective signals caused by facial texture and vascular pulsation under visible light conditions, and their algorithm design is not optimized for low-light scenarios. In low-light environments, due to the significantly reduced image signal-to-noise ratio and weakened light intensity signals in facial feature areas, the noise proportion in the rPPG signal increases, leading to inaccurate feature extraction, increased errors in physiological parameter calculation, and unreliable monitoring results. Existing technologies do not provide effective solutions suitable for low-light environments, only addressing the issue by increasing exposure time or improving gain, but this easily introduces image blurring or saturation distortion, failing to guarantee monitoring accuracy.
[0003] There is currently no effective solution to the problem of being unable to accurately determine the state parameters of the target object in related technologies.
[0004] Therefore, it is necessary to improve the relevant technology to overcome the aforementioned defects. Summary of the Invention
[0005] This invention provides a method, apparatus, storage medium, and electronic device for determining target state parameters, to at least solve the problem of being unable to accurately determine the state parameters of a target object.
[0006] According to one aspect of the present invention, a method for determining target state parameters is provided, comprising: acquiring a target facial video stream of a target object output by an image acquisition module, and acquiring a supplementary lighting state signal output by a supplementary lighting module after supplementing lighting on the space where the target object is located, wherein the supplementary lighting state signal includes: the output power of the supplementary lighting lamp group of the supplementary lighting module, and a timing marker for video stream frame rate synchronization; determining the supplementary lighting conditions corresponding to each frame image in the target facial video stream according to the supplementary lighting state signal, and determining the reference light intensity offset between the forehead region and the root of the nose region in each frame image according to the supplementary lighting conditions corresponding to each frame image; correcting the pixel brightness value of each frame image pixel by pixel according to the reference light intensity offset between the forehead region and the root of the nose region in each frame image to obtain a corrected video stream, and determining the target state parameters of the target object according to the corrected video stream.
[0007] In an exemplary embodiment, before acquiring the supplementary lighting status signal output by the supplementary lighting module, the method further includes: detecting the illuminance value of the space where the target object is located; controlling the supplementary lighting module not to perform supplementary lighting when the illuminance value is greater than or equal to a preset threshold; and controlling the supplementary lighting module to perform supplementary lighting on the space where the target object is located based on the acquired facial video stream of the target object when the illuminance value is less than the preset threshold.
[0008] In an exemplary embodiment, controlling the fill light module to provide fill light to the space where the target object is located based on the acquired facial video stream of the target object includes: extracting the pixel brightness distribution of the facial region of the target object from the facial video stream, and calculating the average gray value of the facial region based on the pixel brightness distribution; determining the output power of the fill light group of the fill light module based on the average gray value to obtain a target power value; and, upon receiving the next frame image acquisition synchronization signal from the image acquisition module, controlling the fill light module to adjust the output power of the fill light group to the target power value.
[0009] In an exemplary embodiment, determining the target state parameters of the target object based on the corrected video includes: extracting a first light intensity change signal of the forehead region of the target object from the corrected video stream based on remote photoplethysmography; separating a DC component and an AC component from the first light intensity change signal, and determining the blood oxygen saturation of the target object based on the DC component and the AC component; determining the heart rate of the target object based on the signal period of the first light intensity change signal; and / or extracting a second light intensity change signal of the nasal root region of the target object from the corrected video stream based on remote photoplethysmography; analyzing the pulsation amplitude characteristics of the nasal root region based on the second light intensity change signal; determining the blood pressure trend of the target object based on the pulsation amplitude characteristics of the nasal root region and the heart rate of the target object using a preset blood pressure trend prediction model; and / or determining the features of multiple preset facial feature points from the corrected video stream, and determining the emotion of the target object based on the features of the multiple preset facial feature points.
[0010] In an exemplary embodiment, after determining the target state parameters of the target object based on the corrected video, the method further includes: obtaining the historical average state parameters of the target object over a preset historical time period; if the ratio of the historical average state parameters to the target state parameters is greater than a preset ratio, determining whether a power step event exists based on the output power of the supplementary lighting group recorded in the supplementary lighting state signal; if a power step event exists, replacing the target state parameters with the weighted average of the historical average state parameters and the target state parameters.
[0011] In an exemplary embodiment, before determining the supplementary lighting conditions corresponding to each frame of the target face video stream based on the supplementary lighting state signal, the method further includes: performing Gaussian filtering on each frame of the target face video stream to obtain a filtered image frame sequence; identifying the facial region of the target object using a facial feature point detection algorithm based on the filtered image frame sequence, and determining the position coordinates of multiple preset facial feature points; determining the spatial transformation parameters of each frame image relative to a reference frame based on the changes in the position coordinates of the multiple preset facial feature points between consecutive frames, and remapping the pixel coordinates of each frame image based on the spatial transformation parameters.
[0012] According to another aspect of the present invention, a device for determining target state parameters is also provided, comprising: an acquisition module, configured to acquire a target facial video stream of a target object output by an image acquisition module, and to acquire a supplementary lighting state signal output by a supplementary lighting module after supplementing lighting on the space where the target object is located, wherein the supplementary lighting state signal includes: the output power of the supplementary lighting lamp group of the supplementary lighting module, and a timing marker for the synchronization of the video stream frame rate; a first determination module, configured to determine the supplementary lighting conditions corresponding to each frame image in the target facial video stream according to the supplementary lighting state signal, and to determine the reference light intensity offset between the forehead region and the root of the nose region in each frame image according to the supplementary lighting conditions corresponding to each frame image; and a second determination module, configured to perform pixel-by-pixel correction on the pixel brightness value of each frame image according to the reference light intensity offset between the forehead region and the root of the nose region in each frame image to obtain a corrected video stream, and to determine the target state parameters of the target object according to the corrected video stream.
[0013] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, wherein the computer program is configured to execute the above-described method for determining the target state parameters at runtime.
[0014] According to another aspect of the present invention, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the method for determining the target state parameters through the computer program.
[0015] According to yet another embodiment of the present invention, a computer program product is also provided, comprising a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0016] This invention solves the problem of inaccurate target object state parameter determination under low light conditions by coordinating the input of the target facial video stream and the supplementary lighting status signal. Existing technologies use only the facial video stream as a single input, which leads to rPPG signal distortion and large parameter calculation errors due to the low image signal-to-noise ratio and strong non-physiological fluctuations in light intensity under low light conditions. This application abandons the single signal processing mode and introduces the supplementary lighting status signal as a key input parameter, enabling the state parameter determination process to sense and respond to the real-time state of external lighting intervention, thereby accurately determining the target object's state parameters. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the hardware environment for a method of determining target state parameters according to an embodiment of this application;
[0020] Figure 2 This is a flowchart of a method for determining target state parameters according to an embodiment of the present invention;
[0021] Figure 3 This is a structural block diagram of a target state parameter determination device according to an embodiment of the present invention;
[0022] Figure 4 This is a schematic diagram of the monitoring process of a target state parameter determination device according to an embodiment of the present invention;
[0023] Figure 5 This is a structural block diagram of a target state parameter determination device according to an embodiment of the present invention. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] According to one aspect of the embodiments of this application, a method for determining a target state parameter is provided. This method for determining the target state parameter is widely applicable to whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligencehouse ecosystems. Optionally, in this embodiment, the above-mentioned method for determining the target state parameter can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.
[0027] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.
[0028] To address the aforementioned problems, this embodiment provides a method for determining target state parameters. Figure 2 This is a flowchart of a method for determining target state parameters according to an embodiment of the present invention, the process including the following steps S202-S206:
[0029] Step S202: Obtain the target face video stream of the target object output by the image acquisition module, and obtain the supplementary lighting status signal output by the supplementary lighting module after supplementing the space where the target object is located. The supplementary lighting status signal includes: the output power of the supplementary lighting group of the supplementary lighting module and the timing mark of the video stream frame rate synchronization.
[0030] Optionally, a video sequence of the target object's facial region is continuously acquired through an image acquisition module to ensure that it contains skin light intensity change information that can be used for remote photoplethysmography (rPPG) analysis; at the same time, the working status signal of the supplementary light module is acquired synchronously, which reflects physical parameters such as the current supplementary light intensity, wavelength, or start / stop status, and is used to characterize the real-time conditions of external light intervention.
[0031] Step S204: Determine the supplementary lighting conditions corresponding to each frame of the target face video stream based on the supplementary lighting state signal, and determine the reference light intensity offset between the forehead area and the root of the nose area in each frame of the image based on the supplementary lighting conditions corresponding to each frame of the image.
[0032] In other words, it is possible to analyze the supplementary lighting status signal, obtain the output power of the supplementary lighting group corresponding to each frame of video (such as 0mW, 100mW, 500mW) and the timing mark synchronized with the image acquisition, establish a "frame-lighting" mapping table, clarify the lighting intervention status at the time of each frame acquisition, and provide a basis for subsequent difference compensation.
[0033] Optionally, based on this mapping relationship, the local average light intensity of the forehead region (the main source of rPPG signal) and the nasal root region (the blood pressure trend sensitive area) in each frame of the image can be extracted separately, and the reference light intensity offset relative to the lightless environment reference under different supplementary light intensities can be calculated (i.e. the non-physiological light intensity increment caused by infrared supplementary light). This offset is a quantitative characterization of light interference.
[0034] Step S206: Based on the reference light intensity offset between the forehead region and the nasal root region in each frame of the image, the pixel brightness value of each frame of the image is corrected pixel by pixel to obtain the corrected video stream, and the target state parameters of the target object are determined based on the corrected video.
[0035] Optionally, based on the reference light intensity offset of each frame, the brightness value of each pixel in the original video stream is subtracted frame by frame and region by region to eliminate the DC offset and nonlinear response distortion introduced by the supplementary light, restore the weak light intensity change signal caused by the pulsation of real skin microvessels, and preserve the physiological dynamic components.
[0036] It should be noted that the above steps achieve accurate modeling and pixel-by-pixel compensation of illumination interference, effectively eliminating non-physiological light intensity shifts caused by supplemental lighting, and avoiding parameter drift caused by illumination fluctuations in traditional methods.
[0037] It should be noted that this application does not rely solely on video streams for parameter inference. Instead, it uses the target facial video stream and the supplementary lighting status signal as joint inputs to establish a dynamic correlation model between the two. Based on the supplementary lighting status, it performs environmental attribution analysis on the light intensity fluctuations in the video signal, thereby effectively distinguishing between non-physiological interference caused by supplementary lighting and real physiological signal changes. Ultimately, it achieves accurate and stable determination of the target object's heart rate, blood oxygen, mood, and other state parameters.
[0038] The above steps address the problem of inaccurate target object state parameter determination under low light conditions by coordinating the input of the target facial video stream and the supplementary lighting status signal. Existing technologies use only the facial video stream as a single input, which leads to rPPG signal distortion and large parameter calculation errors due to low image signal-to-noise ratio and strong non-physiological light intensity fluctuations in low-light environments. This application, however, abandons the single signal processing mode and introduces the supplementary lighting status signal as a key input parameter. This allows the state parameter determination process to sense and respond to the real-time state of external lighting intervention, thereby accurately determining the target object's state parameters.
[0039] It should be noted that the execution subject of the method for determining the target state parameters in this application is the device for determining the target state parameters. This device adopts an embedded integrated structure (embedded and installed inside the home appliance): it uses a snap-on base and a detachable decorative panel to achieve physical integration of the camera device with devices such as televisions and massage chairs, without the need for additional installation space.
[0040] In an exemplary embodiment, before acquiring the supplementary lighting status signal output by the supplementary lighting module, the method further includes: detecting the illuminance value of the space where the target object is located; controlling the supplementary lighting module not to perform supplementary lighting when the illuminance value is greater than or equal to a preset threshold; and controlling the supplementary lighting module to perform supplementary lighting on the space where the target object is located based on the acquired facial video stream of the target object when the illuminance value is less than the preset threshold.
[0041] Optionally, an ambient light sensor deployed within the target state parameter determination device collects the ambient illuminance value of the space where the target object is located in real time. This illuminance value is continuously sampled in lux at a sampling frequency of no less than 1 Hz to reflect the dynamic changes in actual lighting. When the detected ambient illuminance value is greater than or equal to a preset threshold (e.g., 300 lux), it is determined that the lighting is sufficient, and an "supplementary lighting off" status signal is automatically output to control the supplementary lighting module to stop working, avoiding visual interference or energy waste caused by visible or infrared light to the user. When the ambient illuminance value is lower than the preset threshold, it is determined to be a low-light environment, and the supplementary lighting module is triggered to start. This control logic realizes a closed-loop response between supplementary lighting behavior and ambient light and video quality, improving monitoring robustness and user experience.
[0042] In an exemplary embodiment, before acquiring the supplementary lighting status signal output by the supplementary lighting module, the method further includes: detecting the illuminance value of the space where the target object is located; controlling the supplementary lighting module not to perform supplementary lighting when the illuminance value is greater than or equal to a first preset threshold; controlling the supplementary lighting module to perform supplementary lighting on the space where the target object is located according to a first power value when the illuminance value is greater than or equal to a second preset threshold and less than a first predetermined threshold; and controlling the supplementary lighting module to perform supplementary lighting on the space where the target object is located according to a second power value when the illuminance value is less than the second preset threshold, wherein the second power value is greater than the first power value, and both the first power value and the second power value are the output power values of the supplementary lighting lamp group of the supplementary lighting module.
[0043] For example, when the illuminance value is greater than or equal to the first preset threshold (e.g., 300 lux), it is determined that natural light or ambient lighting is sufficient, and a "supplementary light off" status signal is output, completely disabling the supplementary light module to avoid unnecessary energy consumption and light pollution. When the illuminance value is between the second preset threshold (e.g., 50 lux) and the first preset threshold, it is determined to be a moderately low light environment, and the supplementary light module is controlled to start the infrared LED light group at the first power value (e.g., 100mW) to provide low-intensity, low-interference, uniform supplementary light, ensuring that facial details are discernible and there is no overexposure. When the illuminance value is lower than the second preset threshold, the system determines to be an extremely low light environment (e.g., getting up at night), and then starts the second power value (e.g., 500mW) to significantly enhance the infrared light output, maintaining a sufficient signal-to-noise ratio in the facial area to support stable rPPG signal extraction. The first and second power values are precisely adjusted by the drive circuit of the supplementary light group, and their switching is fed back in real time by the supplementary light status signal, achieving adaptive matching between light intensity and monitoring needs, balancing energy efficiency and accuracy.
[0044] In an exemplary embodiment, controlling the fill light module to fill light on the space where the target object is located based on the acquired facial video stream of the target object can be achieved through the following steps S11-S13:
[0045] Step S11: Extract the pixel brightness distribution of the facial region of the target object from the facial video stream, and calculate the average gray value of the facial region based on the pixel brightness distribution;
[0046] In other words, in each frame of the facial video stream, the pixel range of the target object's facial region is accurately extracted using a face detection algorithm (such as facial key point localization based on Haar features or deep learning), and grayscale values are sampled for all pixels in the region. Highlights, shadows, or non-facial interference areas are removed, and only neutral skin region data is retained. The weighted average grayscale value is calculated as the core indicator reflecting the current facial lighting level.
[0047] Step S12: Determine the output power of the fill light group of the fill light module based on the average gray value to obtain the target power value;
[0048] Optionally, the average grayscale value can be input into a preset power mapping function. This function is constructed based on the optimal rPPG signal-to-noise ratio-grayscale value relationship curve calibrated in the laboratory, and outputs a target power value that matches it (e.g., 500mW when the average grayscale is below 40, and 100mW when it is between 40 and 70), to ensure that the supplementary light intensity strictly matches the actual lighting needs of the face and avoid underexposure or overexposure.
[0049] Step S13: Upon receiving the next frame image acquisition synchronization signal from the image acquisition module, control the fill light module to adjust the output power of the fill light group to the target power value.
[0050] Optionally, at the instant the image acquisition module completes the acquisition of a frame and outputs a synchronous trigger signal, the control circuit updates the drive current of the fill light group in real time, so that the output power of the infrared LED completes a smooth transition within the inter-frame period, achieving strict timing synchronization between fill light and video acquisition, ensuring that each frame of image is acquired under optimal lighting conditions, thereby improving the quality of the rPPG signal from the source.
[0051] In an exemplary embodiment, determining the target state parameters of the target object based on the corrected video includes the following steps S31, and / or S32, and / or S33:
[0052] Step S31: Based on remote photoplethysmography, extract the first light intensity change signal of the forehead region of the target object from the corrected video stream; separate the DC component and the AC component from the first light intensity change signal, determine the blood oxygen saturation of the target object based on the DC component and the AC component; and determine the heart rate of the target object based on the signal period of the light intensity change signal.
[0053] Optionally, blood oxygen saturation can be calculated using the AC / DC ratio and the Lambert-Beer law, and heart rate can be extracted using Fourier transform or zero-crossing detection.
[0054] It should be noted that the DC component represents the average optical absorption characteristics of skin tissue, while the AC component is modulated by the periodic changes in microarterial blood volume caused by heartbeats, and its frequency is consistent with the heartbeat cycle of the target object.
[0055] Step S32: Based on remote photoplethysmography, extract the second light intensity change signal of the nasal root region of the target object from the corrected video stream; analyze the pulsation amplitude characteristics of the nasal root region based on the second light intensity change signal; determine the blood pressure trend of the target object based on the pulsation amplitude characteristics of the nasal root region and the heart rate of the target object using a preset blood pressure trend prediction model;
[0056] Optionally, the second light intensity change signal of the nasal root area (above the philtrum) can be extracted, its pulsation amplitude (reflecting systolic blood pressure-related vascular tension) can be analyzed, and combined with the synchronously acquired heart rate data, it can be input into a lightweight LSTM blood pressure trend prediction model trained with multiple samples, and output three-level trend labels of "stable", "slightly elevated" and "slightly decreased" to achieve non-invasive dynamic blood pressure monitoring.
[0057] Step S33: Determine the features of multiple preset facial feature points from the corrected video stream, and determine the emotion of the target object based on the features of the multiple preset facial feature points.
[0058] Optionally, 68 facial feature points can be tracked based on Dlib or MediaPipe algorithms to quantify micro-motion parameters such as drooping corners of the eyes, curvature of the corners of the mouth, and opening and closing of the eyelids. The SVM classifier is used to match three emotional states: "calm," "anxious," and "fatigued," to achieve a non-perceptible assessment of psychological state and construct a "physiological-psychological" dual-dimensional health profile.
[0059] The above steps S31–S33, based on the corrected video stream, achieve four-dimensional simultaneous extraction of heart rate, blood oxygen, blood pressure trends, and emotional state: blood oxygen and heart rate are calculated by accurately separating AC / DC signals through rPPG, and a blood pressure trend prediction model is constructed by combining the amplitude of nasal root pulsation with heart rate, breaking through the limitation of traditional methods that only measure 1–2 parameters; facial micro-expression analysis is integrated to achieve imperceptible recognition of emotional state; and the monitoring dimensions are significantly improved.
[0060] In an exemplary embodiment, after determining the target state parameters of the target object based on the corrected video, the method further includes: obtaining the historical average state parameters of the target object over a preset historical time period; if the ratio of the historical average state parameters to the target state parameters is greater than a preset ratio, determining whether a power step event exists based on the output power of the supplementary lighting group recorded in the supplementary lighting state signal; if a power step event exists, replacing the target state parameters with the weighted average of the historical average state parameters and the target state parameters.
[0061] For example, after extracting target state parameters (such as heart rate and blood oxygen) from the target facial video stream and supplementary lighting status signal, a dynamic calibration mechanism can be further activated. First, the historical average state parameters of the target object continuously collected in the past 30 minutes are extracted (e.g., the average heart rate of the last 6 times is 78 bpm). If the ratio of the currently extracted target state parameter to this value exceeds a preset threshold (e.g., 1.2, meaning the current value is 20% higher than the historical average), anomaly diagnosis will be triggered—querying the power change records in the supplementary lighting status signal to determine whether there is a "power step event" (e.g., the supplementary light suddenly increases from 100mW to 500mW, or the power jump is ≥300mW within 3 frames). If such a sudden change in illumination is confirmed, it indicates that the current parameter anomaly may be caused by supplementary lighting interference rather than a real physiological change. The sudden change value will not be used directly, but a weighted fusion strategy will be adopted: the target state parameter and the historical average value will be fused with a linearly decaying weight.
[0062] Optionally, the weighting coefficient is set to decrease linearly based on the number of frames elapsed after the power step event. Specifically, when a sudden change in infrared illumination power is detected (e.g., a sudden increase from 100mW to 500mW), the abnormal physiological parameters are not immediately adopted. Instead, a gradual correction mechanism is initiated. The weighting coefficient is set to 0.1 after one frame (i.e., the historical average accounts for 90%, and the target state parameter accounts for 10%). After two frames (approximately 33ms, calculated at 30fps), the weighting coefficient increases linearly by 0.05 until it reaches 0.6 after 10 frames (approximately 333ms) (i.e., the historical average accounts for 40%, and the target state parameter accounts for 60%), after which it remains stable and does not change.
[0063] It should be noted that the above mechanism effectively suppresses the instantaneous distortion of rPPG signal caused by sudden changes in supplementary light power, reduces the misjudgment rate caused by environmental interference by more than 65%, significantly improves the stability and reliability of monitoring, and at the same time retains the response capability of real physiological changes, achieving an intelligent balance between "anti-interference" and "sensitivity".
[0064] In an exemplary embodiment, before determining the supplementary lighting conditions corresponding to each frame of the target face video stream based on the supplementary lighting state signal, the method further includes: performing Gaussian filtering on each frame of the target face video stream to obtain a filtered image frame sequence; identifying the facial region of the target object using a facial feature point detection algorithm based on the filtered image frame sequence, and determining the position coordinates of multiple preset facial feature points; determining the spatial transformation parameters of each frame image relative to a reference frame based on the changes in the position coordinates of the multiple preset facial feature points between consecutive frames, and remapping the pixel coordinates of each frame image based on the spatial transformation parameters.
[0065] It should be noted that before extracting physiological parameters from the target facial video stream, high-precision video preprocessing is first performed to eliminate signal distortion caused by head micro-movements, posture shifts, or device jitter: For each frame of the original video stream, a 3×3 Gaussian kernel is used for low-pass filtering to effectively suppress environmental noise and high-frequency artifacts, while retaining the main frequency component of microvascular light intensity changes; Subsequently, based on the Dlib or MediaPipe facial key point detection algorithm, 68 facial feature points (such as the center of the eyebrows, tip of the nose, corners of the eyes, and corners of the mouth) are accurately located to construct the geometric topology of the face; To eliminate non-rigid displacement interference, the first frame is used as the reference frame, and the affine transformation parameters (including translation, rotation, and scaling) of each feature point in each subsequent frame relative to the reference frame are calculated, and bilinear interpolation is used to spatially remap (warping) the image pixels, so that the target facial region achieves "deformation correction" in time, ensuring that the rPPG signal is always taken from the same anatomical region (such as the forehead and root of the nose), avoiding signal attenuation or aliasing caused by facial shift.
[0066] It should be noted that this preprocessing procedure reduces rPPG signal noise caused by slight head movement by about 60%, stabilizes blood oxygen detection error from ±8% to ±3%, and improves heart rate recognition accuracy to over 98%.
[0067] In an exemplary embodiment, after determining the target state parameters, the target state parameter determining device will delete the target facial video stream of the target object from the target state parameter determining device, and encrypt the target state parameters (e.g., AES-256 encryption) before sending them to the IoT cloud or the target object's mobile terminal, thereby avoiding the leakage of the target object's privacy.
[0068] Obviously, the embodiments described above are merely some embodiments of the present invention, and not all embodiments. To better understand the above method, the following description, in conjunction with embodiments, illustrates the process, but is not intended to limit the technical solutions of the embodiments of the present invention. Specifically:
[0069] The device for determining the target state parameters in this application is embedded inside a home appliance (such as a TV bezel or massage chair armrest). The overall structure consists of five main modules, and the functions and connections of each module are as follows:
[0070] 1) Embedded camera module:
[0071] Core components: 1 / 2.7-inch high-resolution RGB camera (2 megapixels), wide-angle lens (120° field of view).
[0072] Function Description: Captures facial video streams of elderly people, supports 30fps frame rate, and meets the requirements for rPPG signal extraction.
[0073] Connections with other modules: Output video stream to the low-light adaptive lighting module and the multimodal data processing module.
[0074] 2) Low-light adaptive supplemental lighting module:
[0075] Core components: infrared LED light group (wavelength 850nm, avoiding visible light stimulation), ambient light sensor (detection range 0-1000lux), and power regulation chip.
[0076] Function Description: Based on the illuminance value detected by the ambient light sensor, the infrared LED power is automatically adjusted (power is increased to 500mW when illuminance < 50 lux, and reduced to 100mW when illuminance 50-300 lux).
[0077] Connection with other modules: Receive ambient light data → output supplementary lighting control signal to LED light group; receive video stream synchronization signal from camera module to ensure supplementary lighting matches the acquisition frame rate.
[0078] 3) Multimodal data processing module:
[0079] Core components: ARM Cortex-A75 processor (supports floating-point operations), algorithm storage unit (stores optimized rPPG algorithm, blood pressure trend extraction algorithm, emotion recognition algorithm), and error calibration unit;
[0080] Function Description: a) Preprocessing: Denoise the video stream (Gaussian filtering) and align frames (eliminating the effects of hand shakiness / sitting posture); b) Parameter Extraction: Extract heart rate and blood oxygen by optimizing the rPPG algorithm; calculate blood pressure trends by facial vascular distribution features and pulsation amplitude; identify emotions (calm / anxiety / fatigue) by micro-movements of 68 facial feature points (such as the corners of the eyes and mouth); c) Error Calibration: Correct real-time parameter errors by fusing nearly 30 minutes of historical data with multiple frames (taking the average of 5 frames).
[0081] Connections with other modules: a) Input: Video stream from the camera module, illumination status signal from the low-light illumination module; b) Output: Structured parameter data (heart rate, blood oxygen, blood pressure trend curve, emotion tags) to the data encryption transmission module.
[0082] 4) Data encryption transmission module:
[0083] Core components: AES-256 encryption chip, Wi-Fi 6 communication unit.
[0084] Functional description: a) It does not store the original video stream, but only receives the structured parameters output by the multimodal data processing module; b) It encrypts the structured data with AES-256; c) It transmits the data to the home appliance control center or mobile APP via Wi-Fi 6 (latency <100ms).
[0085] Connections with other modules: a) Input: Structured parameters of the multimodal data processing module; b) Output: Encrypted parameter data to an external terminal.
[0086] 5) Appliance compatibility and installation structure:
[0087] Core components: snap-on base (compatible with different home appliances with thicknesses of 5-20mm), silicone sealing gasket, and detachable decorative panel (color matches the outer shell of the home appliance).
[0088] Functional description: a) The snap-on base requires no drilling and directly snaps into the pre-reserved installation position of the home appliance; b) The silicone sealing gasket prevents dust from entering the interior of the home appliance.
[0089] Connection with other modules: Physically connect the camera module and the fill light module, fix them inside the home appliance, and ensure that the modules are flush with the outer shell of the home appliance.
[0090] It should be noted that, Figure 3 The diagram illustrates the structural block diagram of the device for determining the target state parameters. The embedded camera module and the low-light adaptive supplementary lighting module are bidirectionally connected (synchronous acquisition and supplementary lighting), and both output signals unidirectionally to the multimodal data processing module. The multimodal data processing module outputs structured data unidirectionally to the data encryption transmission module. The adapter installation structure is a physical support module, which is fixedly connected to the other four functional modules.
[0091] To better understand, such as Figure 4 As shown, the monitoring process of the target state parameter determination device is as follows:
[0092] Step S1: Device initialization and power supply integration (T0-T1, time <30s):
[0093] The device is installed on the pre-reserved position on the upper bezel of the TV using a snap-on base. The decorative panel is the same color as the TV casing. The device's power supply is integrated with the TV's power supply, and the device automatically starts when the TV is turned on, without the need for an additional switch.
[0094] The data encryption transmission module automatically connects to the home Wi-Fi and pairs with the mobile app (manual confirmation is required for the first use; subsequent connections are automatic).
[0095] Step S2: Ambient light detection and supplemental lighting adjustment (T1-T2, executed in real time):
[0096] The ambient light sensor of the low-light adaptive supplementary lighting module collects the indoor illuminance value (denoted as L) once per second. If L < 50 lux, the power adjustment chip controls the infrared LED light group to increase the power to 500mW and turn on the supplementary lighting. If 50 lux ≤ L ≤ 300 lux, the power is reduced to 100mW to maintain low-intensity supplementary lighting. If L > 300 lux, the supplementary lighting is turned off, and the video stream is collected only by relying on ambient light.
[0097] Step S3: Facial video stream acquisition (T2-T3, continuous acquisition):
[0098] The embedded camera module is aimed at the seating area in front of the TV (120° field of view covering a range of 2-3 meters) and captures facial video streams at a frame rate of 30fps; the video stream is transmitted to the multimodal data processing module in real time, without storing the original frames locally to avoid privacy leaks.
[0099] Step S4: Video stream preprocessing (T3-T4, executed synchronously with acquisition):
[0100] The multimodal data processing module performs Gaussian filtering (3×3 kernel size) on each frame of image to remove environmental noise; it locates the face region through facial feature point detection (based on the Dlib library); if the face is offset (such as head rotation), it automatically adjusts the frame alignment parameters to ensure stable acquisition of microvascular pulsation signals.
[0101] Step S5: Multimodal physiological parameter extraction (T4-T5, results output every 5 seconds):
[0102] Heart rate / blood oxygen extraction: An optimized rPPG algorithm is used to extract light intensity change signals from the forehead area of the face, separate the DC component (DC) and AC component (AC), calculate blood oxygen saturation through the AC / DC ratio, and calculate heart rate through the signal period;
[0103] Blood pressure trend extraction: Analyze the amplitude of vascular pulsation at the root of the nose (positively correlated with systolic blood pressure), combine with heart rate data, and output blood pressure trends (such as "stable", "slightly elevated", "slightly decreased") through a pre-trained model (trained based on 1000 elderly samples).
[0104] Emotion recognition: By tracking the drooping of the corners of the eyes and the upturn of the corners of the mouth through 68 facial feature points, it matches three emotion labels: "calm" (horizontal corners of the mouth, no drooping of the corners of the eyes), "anxious" (downturned corners of the mouth, tense corners of the eyes), and "fatigued" (drooping corners of the eyes, relaxed facial muscles).
[0105] Step S6: Error calibration (T5-T6, synchronized with parameter extraction):
[0106] The error calibration unit calls up historical parameters from the past 30 minutes (such as the heart rate values of the past 6 times) and calculates the average value as a benchmark. It performs multi-frame fusion on the currently extracted parameters (taking the average value of the parameters of 5 consecutive frames). If the deviation from the historical benchmark exceeds 5%, it automatically corrects it (for example, if the current heart rate value is 100 beats / minute and the historical benchmark is 80 beats / minute, the deviation is 25%, and it is determined whether there is noise interference based on the supplementary lighting status. If it is noise, it is corrected to 85 beats / minute).
[0107] Step S7: Encrypted Data Transmission and Anomaly Warning (Continuously Executed):
[0108] The data encryption transmission module encrypts the calibrated structured parameters (such as "heart rate: 75 beats / minute, blood oxygen: 96%, blood pressure trend: stable, mood: calm") using AES-256; the encrypted data is transmitted to the mobile APP via Wi-Fi 6, and the APP displays the parameter curves in real time; if the parameters exceed the preset threshold (such as heart rate > 100 beats / minute, blood oxygen < 93%), the APP triggers an audible and visual alarm and pushes an SMS to the emergency contact.
[0109] It should be noted that steps S2 (supplementary light adjustment), S3 (video acquisition), S4 (preprocessing), S5 (parameter extraction), S6 (calibration), and S7 (transmission warning) are executed in a loop until the TV is turned off (the device is turned off synchronously); step S1 (initialization) is executed only once when the TV is turned on for the first time.
[0110] It should be noted that, addressing the shortcomings of existing technologies such as "low compliance with standalone devices," this application achieves "monitoring upon powering on the TV / massage chair," eliminating the need for elderly users to actively operate or wear the device. This non-intrusive design significantly improves compliance and ensures continuous and complete monitoring data. Regarding the issue of "insufficient accuracy in low light," infrared adaptive supplemental lighting and an optimized rPPG algorithm ensure that blood oxygenation accuracy in low light is ≥90% consistent with finger-clip devices, and heart rate error is ≤3 beats / minute, meeting monitoring requirements. Addressing the issue of "single parameter," four parameters are extracted simultaneously, enabling a comprehensive assessment of the elderly's chronic diseases and psychological state, providing comprehensive data support for home-based elderly care. Regarding the issue of "privacy risks," original images are not stored, and transmission is encrypted, protecting privacy throughout the entire data acquisition and transmission chain. Regarding the issue of "poor adaptability," the snap-on installation and decorative panel design require no environmental modifications, improving user acceptance.
[0111] This application outperforms existing technologies in terms of monitoring continuity, accuracy, parameter dimensions, privacy protection, and adaptability, and can be widely applied in health monitoring scenarios for the elderly.
[0112] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.
[0113] This embodiment also provides a device for determining target state parameters, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0114] Figure 5 This is a structural block diagram of a target state parameter determination device according to an embodiment of the present invention. The device includes:
[0115] The acquisition module 502 is used to acquire the target face video stream of the target object output by the image acquisition module, and to acquire the supplementary lighting status signal output by the supplementary lighting module after supplementing the space where the target object is located.
[0116] The first determining module 504 is used to determine the supplementary lighting conditions corresponding to each frame of the target face video stream according to the supplementary lighting state signal, and to determine the reference light intensity offset between the forehead area and the root of the nose area in each frame of the image according to the supplementary lighting conditions corresponding to each frame of the image.
[0117] The second determining module 506 is used to perform pixel-by-pixel correction on the pixel brightness value of each frame image based on the reference light intensity offset between the forehead region and the nasal root region in each frame image to obtain a corrected video stream, and to determine the target state parameters of the target object based on the corrected video.
[0118] The aforementioned device addresses the problem of inaccurate target object state parameter determination under low light conditions by coordinating the input of the target facial video stream and the supplementary lighting status signal. Existing technologies use only the facial video stream as a single input, which leads to rPPG signal distortion and large parameter calculation errors due to the low image signal-to-noise ratio and strong non-physiological fluctuations in light intensity under low light environments. This application, however, abandons the single signal processing mode and introduces the supplementary lighting status signal as a key input parameter. This allows the state parameter determination process to sense and respond to the real-time state of external lighting intervention, thereby accurately determining the target object's state parameters.
[0119] In an exemplary embodiment, the device further includes: a control module, configured to detect the illuminance value of the space where the target object is located before acquiring the supplementary lighting status signal output by the supplementary lighting module; control the supplementary lighting module not to perform supplementary lighting if the illuminance value is greater than or equal to a preset threshold; and control the supplementary lighting module to perform supplementary lighting on the space where the target object is located based on the acquired facial video stream of the target object if the illuminance value is less than the preset threshold.
[0120] In an exemplary embodiment, the control module is further configured to extract the pixel brightness distribution of the facial region of the target object from the facial video stream, and calculate the average gray value of the facial region based on the pixel brightness distribution; determine the output power of the fill light group of the fill light module based on the average gray value to obtain a target power value; and, upon receiving the next frame image acquisition synchronization signal from the image acquisition module, control the fill light module to adjust the output power of the fill light group to the target power value.
[0121] In an exemplary embodiment, the second determining module 506 is further configured to: extract a first light intensity change signal of the forehead region of the target object based on the corrected video stream using remote photoplethysmography; separate a DC component and an AC component from the first light intensity change signal; determine the blood oxygen saturation of the target object based on the DC component and the AC component; determine the heart rate of the target object based on the signal period of the first light intensity change signal; and / or extract a second light intensity change signal of the nasal root region of the target object based on the corrected video stream using remote photoplethysmography; analyze the pulsation amplitude characteristics of the nasal root region based on the second light intensity change signal; determine the blood pressure trend of the target object based on the pulsation amplitude characteristics of the nasal root region and the heart rate of the target object using a preset blood pressure trend prediction model; and / or determine the features of multiple preset facial feature points from the corrected video stream, and determine the emotion of the target object based on the features of the multiple preset facial feature points.
[0122] In an exemplary embodiment, the device further includes: a correction module, configured to: after determining the target state parameters of the target object based on the corrected video, obtain the historical average state parameters of the target object within a preset historical time period; if the ratio of the historical average state parameters to the target state parameters is greater than a preset ratio, determine whether a power step event exists based on the output power of the supplementary lighting group recorded in the supplementary lighting state signal; if a power step event exists, replace the target state parameters with the weighted average of the historical average state parameters and the target state parameters.
[0123] In an exemplary embodiment, the apparatus further includes: a preprocessing module, configured to perform Gaussian filtering on each frame of the target face video stream before determining the supplementary lighting conditions corresponding to each frame of the target face video stream based on the supplementary lighting state signal, to obtain a filtered image frame sequence; to identify the facial region of the target object using a facial feature point detection algorithm based on the filtered image frame sequence, and to determine the position coordinates of multiple preset facial feature points; to determine the spatial transformation parameters of each frame of the image relative to a reference frame based on the changes in the position coordinates of the multiple preset facial feature points between consecutive frames, and to perform pixel coordinate remapping on each frame of the image based on the spatial transformation parameters.
[0124] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to perform the steps in any of the above method embodiments when executed.
[0125] Optionally, in this embodiment, the storage medium may be configured to store a computer program for performing the following steps:
[0126] S1, acquire the target face video stream of the target object output by the image acquisition module, and acquire the supplementary lighting status signal output by the supplementary lighting module after supplementing the space where the target object is located, wherein the supplementary lighting status signal includes: the output power of the supplementary lighting group of the supplementary lighting module and the timing mark of the video stream frame rate synchronization.
[0127] S2, determine the supplementary lighting conditions corresponding to each frame of the target face video stream based on the supplementary lighting state signal, and determine the reference light intensity offset between the forehead area and the root of the nose area in each frame of the image based on the supplementary lighting conditions corresponding to each frame of the image.
[0128] S3, based on the reference light intensity offset between the forehead region and the nasal root region in each frame of the image, the pixel brightness value of each frame of the image is corrected pixel by pixel to obtain the corrected video stream, and the target state parameters of the target object are determined based on the corrected video.
[0129] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0130] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0131] Embodiments of the present invention also provide an electronic device including a memory and a processor, the memory storing a computer program and the processor being configured to run the computer program to perform the steps in any of the above method embodiments.
[0132] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0133] S1, acquire the target face video stream of the target object output by the image acquisition module, and acquire the supplementary lighting status signal output by the supplementary lighting module after supplementing the space where the target object is located, wherein the supplementary lighting status signal includes: the output power of the supplementary lighting group of the supplementary lighting module and the timing mark of the video stream frame rate synchronization.
[0134] S2, determine the supplementary lighting conditions corresponding to each frame of the target face video stream based on the supplementary lighting state signal, and determine the reference light intensity offset between the forehead area and the root of the nose area in each frame of the image based on the supplementary lighting conditions corresponding to each frame of the image.
[0135] S3, based on the reference light intensity offset between the forehead region and the nasal root region in each frame of the image, the pixel brightness value of each frame of the image is corrected pixel by pixel to obtain the corrected video stream, and the target state parameters of the target object are determined based on the corrected video.
[0136] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0137] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0138] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0139] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0140] The embodiments described herein also provide a computer program that includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in any of the above method embodiments.
[0141] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those described herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0142] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for determining target state parameters, characterized in that, include: The system acquires a target facial video stream of the target object output by the image acquisition module, and acquires a supplementary lighting status signal output by the supplementary lighting module after supplementing the space where the target object is located. The supplementary lighting status signal includes: the output power of the supplementary lighting group of the supplementary lighting module and the timing marker for the synchronization of the video stream frame rate. The supplementary lighting conditions corresponding to each frame of the target face video stream are determined based on the supplementary lighting state signal, and the reference light intensity offset between the forehead region and the root of the nose region in each frame is determined based on the supplementary lighting conditions corresponding to each frame. Based on the reference light intensity offset between the forehead region and the nasal root region in each frame of the image, the pixel brightness value of each frame of the image is corrected pixel by pixel to obtain the corrected video stream, and the target state parameters of the target object are determined based on the corrected video. The process of determining the supplementary lighting conditions corresponding to each frame of the target face video stream based on the supplementary lighting state signal includes: parsing the supplementary lighting state signal and obtaining the output power of the supplementary lighting group corresponding to each frame of the image. The supplementary lighting conditions corresponding to each frame of the image include the output power of the supplementary lighting group, which is used to clarify the illumination intervention state when each frame of the image is acquired. The determination of the reference light intensity offset of the forehead region and the root of the nose region in each frame image according to the supplementary lighting conditions corresponding to each frame image includes: extracting the local average light intensity of the forehead region and the root of the nose region in each frame image, and calculating the reference light intensity offset of the forehead region and the root of the nose region in each frame image relative to the reference light intensity of the no-light environment under the supplementary lighting intensity corresponding to the output power of the supplementary lighting group based on the local average light intensity. The target state parameters include: blood oxygen saturation, heart rate, blood pressure trend, and mood.
2. The method for determining target state parameters according to claim 1, characterized in that, Before acquiring the supplementary lighting status signal output by the supplementary lighting module, the method further includes: Detect the illuminance value of the space where the target object is located; If the illuminance value is greater than or equal to a preset threshold, the supplementary lighting module is controlled not to provide supplementary lighting. When the illuminance value is less than the preset threshold, the supplementary lighting module is controlled to supplement the space where the target object is located based on the acquired facial video stream of the target object.
3. The method for determining target state parameters according to claim 2, characterized in that, Controlling the fill light module to fill light on the space where the target object is located based on the acquired facial video stream of the target object includes: Extract the pixel brightness distribution of the facial region of the target object from the facial video stream, and calculate the average gray value of the facial region based on the pixel brightness distribution; The output power of the fill light group of the fill light module is determined based on the average gray value to obtain the target power value; Upon receiving the next frame image acquisition synchronization signal from the image acquisition module, the fill light module is controlled to adjust the output power of the fill light group to the target power value.
4. The method for determining target state parameters according to claim 1, characterized in that, The target state parameters of the target object are determined based on the corrected video, including: Based on remote photoplethysmography, a first light intensity change signal of the forehead region of the target object is extracted from the corrected video stream; a DC component and an AC component are separated from the first light intensity change signal, and the blood oxygen saturation of the target object is determined based on the DC component and the AC component; and the heart rate of the target object is determined based on the signal period of the first light intensity change signal; and / or Based on remote photoplethysmography, a second light intensity change signal of the nasal root region of the target object is extracted from the corrected video stream; the pulsation amplitude characteristics of the nasal root region are analyzed based on the second light intensity change signal; the blood pressure trend of the target object is determined based on the pulsation amplitude characteristics of the nasal root region and the heart rate of the target object using a preset blood pressure trend prediction model; and / or The features of multiple preset facial feature points are determined from the corrected video stream, and the emotion of the target object is determined based on the features of the multiple preset facial feature points.
5. The method for determining target state parameters according to claim 1, characterized in that, After determining the target state parameters of the target object based on the corrected video, the method further includes: Obtain the historical average state parameters of the target object within a preset historical time period; If the ratio of the historical average state parameter to the target state parameter is greater than a preset ratio, the existence of a power step event is determined based on the output power of the supplementary lighting group recorded in the supplementary lighting state signal. In the presence of a power step event, the target state parameter is replaced with a weighted average of the historical average state parameter and the target state parameter.
6. The method for determining target state parameters according to claim 1, characterized in that, Before determining the supplementary lighting conditions corresponding to each frame of the target face video stream based on the supplementary lighting state signal, the method further includes: Each frame of the target face video stream is subjected to Gaussian filtering to obtain a filtered image frame sequence. Based on the filtered image frame sequence, the facial region of the target object is identified by a facial feature point detection algorithm, and the position coordinates of multiple preset facial feature points are determined. Based on the changes in the position coordinates of the multiple preset facial feature points between consecutive frames, the spatial transformation parameters of each frame image relative to the reference frame are determined, and the pixel coordinates of each frame image are remapped according to the spatial transformation parameters.
7. A device for determining target state parameters, characterized in that, include: The acquisition module is used to acquire the target facial video stream of the target object output by the image acquisition module, and to acquire the supplementary lighting status signal output by the supplementary lighting module after supplementing the space where the target object is located. The supplementary lighting status signal includes: the output power of the supplementary lighting lamp group of the supplementary lighting module and the timing mark of the video stream frame rate synchronization. The first determining module is used to determine the supplementary lighting conditions corresponding to each frame of the target face video stream based on the supplementary lighting state signal, and to determine the reference light intensity offset between the forehead region and the root of the nose region in each frame of the image based on the supplementary lighting conditions corresponding to each frame of the image. The second determining module is used to perform pixel-by-pixel correction on the pixel brightness value of each frame image based on the reference light intensity offset between the forehead region and the nasal root region in each frame image to obtain a corrected video stream, and to determine the target state parameters of the target object based on the corrected video. The device is further configured to analyze the supplementary lighting status signal and obtain the output power of the supplementary lighting group corresponding to each frame of image. The supplementary lighting conditions corresponding to each frame of image include the output power of the supplementary lighting group, which is used to determine the lighting intervention state when each frame of image is acquired. The device is further configured to extract the local average light intensity of the forehead region and the root of the nose region in each frame of the image, and calculate the reference light intensity offset of the forehead region and the root of the nose region in each frame of the image relative to the reference light intensity of the no-light environment under the supplementary light intensity corresponding to the output power of the supplementary light group based on the local average light intensity. The target state parameters include: blood oxygen saturation, heart rate, blood pressure trend, and mood.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 6.
9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 6 through the computer program.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Light supplementing method and device and storage medium
CN111031255A
Facial physiological detection method and system based on signal quality driving ROI selection
CN121545733A